ISAR Publisher

International Scientific and Academic Research Publisher

Submit Manuscript

Semantic Instability in Vision-Language Classification of Abstract and Finished Sketches


Author: Thomas Chen*
Independent Scholar
Published Date: 2026-06-24
Keywords: vision-language models; OpenCLIP; sketch recognition; semantic similarity; machine vision; abstraction; zero-shot classification; image interpretation.
Abstract:

This study evaluates how OpenCLIP classifies abstract and finished sketches across synthetic and human-drawn corpora. Although vision-language models are now widely used for image retrieval and recognition, less is known about their response to simplified, hand-drawn stimuli that depart from photographic training regularities. Two matched datasets were tested: a 240-image synthetic corpus and a 120-image human corpus, each using 12 object categories and two drawing conditions. Each image was compared with a fixed 12-label text ensemble using zero-shot cosine similarity. Outcomes included Top-1 accuracy, Top-1-to-Top-2 margin, and confusion structure. Human drawings produced higher accuracy (95.0%) than synthetic inputs (87.1%), but synthetic inputs produced higher mean margins (0.081 vs 0.056). Finished drawings improved performance in both sources. Category-specific failures, especially synthetic plant and shoe, demonstrate that controlled abstraction can expose semantic instability and boundary conditions in vision-language image interpretation. The method offers a compact protocol for imaging-science evaluation of sketch robustness.