AI generated flowers vs hand painted watercolour flowers

AI vs Brush: A Practice-Based Study on Visual Semantics in Surface Pattern Design

Illustrator, Client, and Artificial Intelligence: Why This Study Exists

This study began as a practical challenge rather than an abstract question. It emerged from a very real shift in the textile and surface pattern design market, shaped by the rapid adoption of artificial intelligence by both clients and consumers. In early 2024, as a commercial illustrator and surface pattern designer, I received a visual reference from one of my recurring clients. I recognised immediately that the image had been created by an AI model, yet it was presented simply as “inspiration”, without any awareness of its origins or implications. It was a symbolic moment: AI had entered our briefing process quietly, without announcement, carried in by the speed and visual allure of new AI-generated image trends capturing the attention of end customers.

From that point onwards, it became clear that my work already existed within a new semantic environment — one blending my own artistic intent with market expectations, client interpretation of emerging digital aesthetics, and machine-generated visuals entering the design conversation. My task at that moment was to manually design a commercial children’s print in a style similar to the AI-generated reference. Not long after, I found myself testing a range of AI-assisted creation tools with clients from fashion, home décor and the wider textile industry. Some were aware of the digital aesthetic but not its technological origins; others knew exactly where the images came from but preferred to leave the navigation of AI tools to a professional they trusted. In every scenario, I was suddenly interacting with artificial intelligence in graphic design without any established guidelines, rules or shared sense of visual adequacy.

air balloon image created by AI vs air balloon surface pattern design painted manually in watercolour
Surface pattern design created manually based on AI inspiration.

Why AI and Human Meaning Diverge in Surface Pattern Design

This commercially driven, turbulent experience — filled with unexpected advantages but also clear limitations — quickly led me to a simple but far-reaching question: how does AI “read” my work? And more importantly, does the model’s vocabulary align with what I actually create? More than once, I felt that we were speaking entirely different languages. What happens to my aesthetic intent, style, technique and narrative direction when they pass through an embedding space trained on large-scale, internet-sourced data?

¨AI vs Brush¨ was developed as a practice-based investigation into exactly this problem. The project analyses ten commissioned surface pattern designs built from AI-generated client references — combining hand-painted and AI-assisted elements within real commercial workflows — and contrasts them with ten manually illustrated, artist-originated patterns to understand the model’s responses in a broader creative spectrum.

The study focuses on three layers of visual meaning:
- my own professional tags, formed over years of illustration and surface pattern design practice;
- the descriptive tags generated by the current OpenCLIP model;
- and the numerical distances within the embedding space that reveal how closely (or how poorly) these two vocabularies align.

The objective is not simply to determine whether graphical AI performs “well” or “poorly”, though the report includes a short reflection on this. The central purpose is to understand what kind of visual meaning survives — and what kind is lost — when artistic work is interpreted by a multimodal model.

The motivation behind the study was not criticism but the need to regain semantic clarity at a time when images circulate between human and machine interpretation. The creative process no longer moves in a linear direction. Instead, it forms a loop in which human-intended and AI-generated images inspire the next cycle of human work, which is later analysed by AI again. This circularity is subtle but transformative, and it sets the stage for the core themes explored throughout this report.

flowers generated by AI vs flowers painted in watercolour with code that provides tags and embeddings

The Visual Semantics Shift: From Early CLIP Tests to the Updated OpenCLIP Model 

Several months before conducting the full embedding analysis, I ran a series of preliminary tests using earlier CLIP-based models. These early experiments were imperfect, yet unexpectedly revealing. The models attempted to describe style, medium and artistic lineage. They often misidentified them, sometimes in ways that were almost comical, but they displayed at least a rudimentary awareness that visual culture contains categories such as “watercolour illustration”, “digital painting”, “pencil sketch”, or even “baroque-inspired image”. The descriptions were chaotic and inconsistent, yet they still reflected an attempt — however unstable — to reach beyond surface-level perception.

When I returned to the study at the beginning of 2026 with the updated OpenCLIP model, the change was immediate and unmistakable. The model no longer attempted to name style at all. Historical, stylistic and technical categories disappeared entirely, replaced by a much smaller, commercially dominant vocabulary. Instead of stylistic interpretations, OpenCLIP consistently produced product-facing descriptors such as “pattern”, “wallpaper”, “seamless”, “storybook”, and, with far greater precision than a few months before, isolated concrete elements like “floral”, “owl”, “rose”. The descriptive space had narrowed, and the narrowing was not random. It aligned closely with internet-scale visual categories, the kinds that dominate marketplaces, stock libraries and SEO-optimised image captions. Early CLIP’s stylistic labels such as: “watercolour painting”, “baroque style”, reflected the distribution of captions in datasets such as LAION-400M, where artist-generated content was more abundant. This tendency faded in later datasets as commercial descriptors became dominant.

OpenCLIP embedding for author surface pattern design in Google Collab environment

This shift can be read as an improvement in accuracy, but only within a particular framework: one shaped by internet conventions, search optimisation and the overwhelming visibility of commercial imagery. More importantly, it signalled a restructuring of semantic priorities inside the model. The updated OpenCLIP version no longer treats style or technique as meaningful categories. It treats them as elements the training data does not consistently reward. In practical terms, the model behaves as if style has no semantic weight. 

This change immediately reframed the purpose of my study. It was no longer only about comparing human tags with AI tags; it became a question about what it means for a multimodal model to cease recognising a foundational component of visual culture. The disappearance of style from the model’s vocabulary was not a technical glitch — it was a cultural symptom. It pointed to a deeper structural issue: a semantic drift in which models trained on large-scale internet data gradually migrate away from the vocabulary of art and toward the vocabulary of commercial surface appearance.

This observation guided the design of the full embedding analysis. If the model’s language is evolving over time — and evolving in ways that directly affect how artistic work is interpreted — then the study must not only measure semantic alignment but also examine the mechanisms shaping the model’s worldview. The following sections move from these empirical observations toward a broader conceptual framework, describing the recursive relationship between AI, human creative practice and the online environments that increasingly mediate both.

Early Indication of Semantic Echo

As the differences between my earlier CLIP tests and the updated OpenCLIP results began to emerge, it became clear that the model’s behaviour had shifted in a way that was not only technical but conceptual. The disappearance of style from the model’s vocabulary was too consistent to be ignored. A pattern was forming: the model no longer expanded its interpretative space but seemed to fold back into it. What initially appeared to be a simple improvement in precision began to resemble something else — an echo of its own learned categories.

At this early stage, the idea of Semantic Echo existed only as a working hypothesis. It suggested that the model was not simply learning from the internet but also learning from the output of earlier models circulating online. If AI-generated images blended into the same datasets that future models rely on, then the descriptive language of those models would reflect not only human conventions but also the tendencies of previous AI systems. Meaning would not grow - it would reverberate.

This hypothesis became increasingly relevant as the study progressed. With each image, each cluster of tags, each calculated distance in the embedding space, the pattern reappeared. The model behaved as though certain categories were gravity points: commercial labels were stable, while stylistic labels had dissolved. The shift felt less like a random evolution and more like a signal of how multimodal models adapt when exposed to large-scale, internet-shaped data. This early intuition laid the foundation for the conceptual framework explored later in the report, where the mechanisms behind this narrowing become more clearly articulated.

Methodology: A Practice-Based Framework for Analysing Visual Semantics

The methodological structure of ¨AI vs Brush¨ was shaped by the commercial context from which the project emerged. Unlike traditional machine learning studies that begin with controlled datasets, this research began with real commissioned surface pattern designs, created for clients in fashion, home décor and the textile market. This practice-led approach allowed the analysis to remain grounded in the lived realities of contemporary design workflows, where AI-generated imagery, human interpretation and client expectations intersect.

The study examines twenty surface pattern designs, divided into two groups. The first group consists of ten patterns developed directly from AI-generated client references: illustrations originating from prompts typed into consumer-level generative tools. These designs combine hand-painted elements with AI-assisted fragments generated in Adobe Express, echoing the hybrid processes that designers increasingly navigate. The second group includes ten manually illustrated patterns based on original artistic concepts, enabling a view into how the model interprets human-originated intention without the influence of an AI prompt.

Each design was annotated twice: first through my own professional tagging system, developed through years of practice in surface pattern design, and then through OpenCLIP-generated tags, representing how the multimodal model interprets the same image. These paired annotations were examined through three analytical layers:
- technique,
- visual effect and
- aesthetic intent.
By separating the act of making from the act of perceiving, and both from the intended meaning, the analysis reflects the complexity of creative work in a way that multimodal models currently do not.

human visual tags vs ai tags

Comparing Human Tags and OpenCLIP Embeddings in Surface Pattern Design

The core of the methodological framework lies in the embedding analysis. Each tag, whether human-generated or AI-generated, is represented as a high-dimensional vector within OpenCLIP’s shared semantic space. The study examines the distances between these vectors to measure alignment, friction and divergence between human meaning and machine meaning. These distances reveal not only how the model recognises patterns within surface pattern design but also how its interpretative space evolves across different checkpoints of training. By combining real commercial practice with computational analysis, the methodology captures both the artistic and algorithmic layers that shape visual interpretation today.

A portion of the dataset consists of hybrid designs generated with the support of Adobe Express. I selected this tool deliberately. Among currently available generative platforms, Adobe Express aligns most closely with the software environment used by professional illustrators and surface pattern designers. As someone who works daily in Adobe Photoshop and Illustrator, it was important to examine AI-generated material that originates from within the same creative ecosystem.

The generative outputs from Adobe Express revealed a recurring pattern that became methodologically relevant. Despite explicit prompts requesting repeatable, tile-ready compositions, the system never produced a true seamless tile. Every AI-generated output required extensive post-processing: rebuilding repetition structures manually, correcting perspective or continuity errors, and repairing visually implausible artefacts that were unmistakably machine-made rather than human-made. The generator could produce interesting visual starting points, but it could not deliver production-ready material. Each hybrid pattern therefore reflects a layered workflow in which AI contributes raw visual fragments, while the designer remains responsible for the technical coherence, stylistic refinement, and functional viability of the final surface pattern.

These observations informed the methodological decision to treat the dataset as a spectrum rather than a binary. Instead of dividing works into “manual” and “AI-assisted”, the study examines how meaning shifts across varying degrees of machine involvement – and how a multimodal model interprets these shifts, or fails to interpret them at all.

How AI Reads Illustrations: Visual Features Behind CLIP Embeddings

Understanding what AI models analyse when converting an artwork into a numerical embedding is essential for interpreting the semantic results of this study. Contrary to the popular assumption that AI recognises artistic style or creative intent, multimodal models such as CLIP and OpenCLIP operate on a fundamentally different layer of visual meaning. They learn by matching image features with natural-language captions collected at internet scale, not by understanding artistic technique or aesthetic lineage.

We know which visual criteria these models attend to thanks to four complementary sources: the architecture of Vision Transformers, feature-map visualisations that reveal the model’s internal attention patterns, the documented training objectives of CLIP-based systems, and a growing body of research analysing multimodal embedding spaces (Radford et al., 2021; Jia et al., 2021).

In practice, this means that what CLIP learns to emphasise are the visual attributes that most strongly co-occur with common online descriptions—primarily objects, textures, colours, and spatial layout—rather than any high-level properties related to style, medium or technique. As summarised in recent research, “CLIP-based multimodal embedding models are trained by matching image features with natural-language captions found online, learning visual attributes such as objects, textures, colours and spatial layout that correlate strongly with text descriptions, but without a mechanism for capturing high-level artistic style or technique.” (based on Radford et al., 2021; Jia et al., 2021)

fragment of Radford study abut CLIP

Radford et al., 2021; Jia et al., 2021

This leads to an important conclusion for visual-arts research: multimodal embedding models do not learn an abstract notion of artistic style. They learn frequent, descriptive visual features that appear in large-scale internet captions—not the processes, gestures or histories that define artistic practice. From a methodological perspective, this explains why the model in this study is highly consistent in identifying flowers, background structures or ornamental density, yet consistently fails to recognise watercolor technique, brushwork, hybrid media or aesthetic intention. The embedding space encodes the surface-level signals, but the language of the model is not equipped to express the meaning of artistic labour.

A plausible explanation for the disappearance of stylistic vocabulary in newer OpenCLIP versions lies not in a change of the model’s architecture, but in the evolution of the data on which these systems are trained. Early CLIP models (2021) often attempted to label style—however inconsistently—because the internet datasets available at the time contained a higher proportion of artist-generated content, including captions referring to “watercolor illustration”, “baroque painting”, “oil portrait” and similar culturally anchored terms. Between 2023 and 2026, the landscape of online images changed dramatically: AI-generated imagery flooded the internet, stock platforms shifted toward SEO-optimised product descriptors, and large datasets increasingly reflected commercial rather than artistic language. It is therefore reasonable to assume that the model’s changing vocabulary is not a sign of reduced capability but a consequence of the changing linguistic environment from which it learns. As the distribution of captions shifts, the model’s semantic space shifts with it — favouring product categories over stylistic or historical ones.

visual logic behind embeddings latent features

Dataset Description: Twenty Patterns Across Two Creative Realities

The dataset used in the study reflects the dual nature of contemporary design practice, where human-originated artwork coexists with AI-assisted imagery. The first half of the dataset consists of ten commissioned surface pattern designs created directly in response to AI-generated client references. These images often carry the distinctive tone of graphical AI, with its aesthetic shortcuts, its digitally smooth visual logic and its tendency to collapse stylistic traditions into simplified decorative forms. Designing around these references required a balance between respecting commercial expectations and restoring artistic coherence.

The second half of the dataset introduces ten manually illustrated patterns, each grounded in a distinct artistic intention rather than an AI prompt. These works provide a counterbalance to the AI-influenced patterns, offering insight into how the model behaves when confronted with images shaped entirely by human intuition, technique and stylistic sensitivity. Together, the two groups form a comprehensive view of how AI and human creativity interact within the broader context of surface pattern design.

Each design is accompanied by metadata: the professional tags describing its technique, aesthetic direction and intended narrative, as well as the OpenCLIP-generated tags, which represent the model’s interpretation. The dataset also includes numerical embedding vectors for each tag, enabling the computation of distances that reveal how meaning transforms within the multimodal space. By working with both commissioned and author-originated patterns, the dataset captures the reality of designing in an era where artificial intelligence actively shapes client expectations and visual culture at large.

twenty images with surface pattern design from which ten are created with AI and ten manually

By working with both commissioned and author-originated patterns, the dataset captures the reality of designing in an era where artificial intelligence actively shapes client expectations and visual culture at large. I intentionally selected this mix of patterns to observe how the model reads two types of stylistic origins: my “clean”, fully manual artistic watercolour textures, brushwork, and imaginative character-driven compositions – and the “clean” AI handwriting generated almost entirely through Adobe Express, occasionally mixed with hand-painted elements. My aim was to see whether OpenCLIP would distinguish between these different forms of authorship, especially since the graphic “brush” of AI has become highly recognisable to professional designers. Surprisingly, the model ignored this distinction entirely. It treated both stylistic origins as semantically equivalent, which raises an important question: why does a multimodal model flatten the difference between human artistic technique and the distinctive visual signature of AI-generated imagery?

Results: Tags, Embeddings and Alignment Between Human and Machine Meaning

The results of the study reveal a consistent and meaningful pattern in how OpenCLIP interprets surface pattern design. When examining the generated tags across all twenty patterns, the model demonstrates a strong preference for commercial descriptors. Words such as “pattern”, “wallpaper”, “seamless”, “floral” and “storybook” appear with high frequency, creating a narrow field of recognition centred on decorative function rather than stylistic identity or used technique: ¨watercolour illustration¨, ¨digital illustration¨. In contrast, professional tags cluster around aesthetic intent, technique and artistic direction, forming a vocabulary shaped by years of practice within illustrated and textile design.

The embedding analysis amplifies this contrast. Distances between professional tags often reflect the natural variation within artistic intention, rooted in specific narratives or crafted through distinct techniques. Meanwhile, the distances between AI-generated tags reveal a model whose interpretative space is stabilised around commercial categories. The more the model evolved, the more its vocabulary converged toward the same compact set of labels.

Semantic Friction Between Human and AI Descriptions of Pattern Design

When comparing the centroids of professional tags and AI-generated tags, the semantic friction becomes visible. In several projects, the model’s interpretation appears to orbit around decorative function rather than the intended style. Even in author-originated designs, where style is central to the composition, OpenCLIP overwhelmingly reduces the meaning to surface-level descriptors. This behaviour remains consistent across both groups of patterns, suggesting that the model does not respond differently to human-originated or AI-influenced imagery. Instead, it responds to the statistical gravity of its training data.

These results set the stage for the conceptual analysis that follows. They show not only what the model sees, but also what it cannot see: the stylistic nuance, the layered aesthetic intent and the narrative coherence that define illustrated surface pattern design as a creative discipline. The next sections explore why this gap exists and how the recursive relationship between AI and internet-trained data amplifies it over time.

code vs embedding tags results vs real life use of art on textiles

Case Studies: How AI Interprets Surface Pattern Design in Practice

The following case studies illustrate how the semantic gap identified in the embedding analysis manifests in real surface pattern design practice. By looking closely at individual projects, we can observe how a multimodal AI model interprets artistic intent, stylistic nuance and hybrid techniques — and where its interpretations diverge most clearly from human meaning.

Case Study 1 — Photo Peony (ai-vs-brush_17)

Highest semantic friction score: 0.271
Technique: Photography, digital painting, Photoshop post-processing
Visual effect: hybrid digital-photo look
Short description: A project developed in Photoshop with author elements of photographic texture and author digital illustration.

Among all twenty designs, Photo Peony produced the highest semantic friction score in the dataset. The pattern combines photographic peony fragments with painterly overpainting - a hybrid technique that blends realism, artistic intervention and commercial print logic. My professional tags emphasized this hybridity: photographic base and digital overpaint. The intention behind the work was to merge two visual traditions in a coherent artistic gesture.

OpenCLIP, however, interpreted the piece almost exclusively through a commercial lens. The model’s dominant tags: ¨pastel flowery background¨, ¨roses background¨ and similar - stripped the design of its hybrid identity and collapsed it into the same semantic cluster as purely decorative surface graphics. None of the photographic texture, none of the digital-painterly integration, and none of the stylistic intention surfaced in the model’s vocabulary. The embedding distances reveal a wide semantic gap between the human and machine interpretations, confirming that the model lacks representational axes capable of expressing hybridity itself.

Photo Peony demonstrates how multimodal AI handles mixed visual sources: it does not. It reduces the complexity of aesthetic construction to the most frequent commercial descriptors available in its training data. This case illustrates a broader pattern: the more stylistically intentional or technically layered the artwork becomes, the more aggressively the model flattens it.

Case Study 2 — Beehive (ai-vs-brush_10)

Semantic friction score: 0.256
Technique: Watercolour illustration, Photoshop post-processing
Visual effect: watercolour illustration
Short description: A project combining traditional watercolour on-paper illustration with digital inspiration.

Beehive represents the group of manually created, author-originated patterns in the dataset. Its construction relies on delicate watercolour textures, organic linework and a compositional rhythm typical of illustrated surface pattern design for children. My professional tags emphasized narrative intention, handmade technique and the traditional watercolour work.

OpenCLIP’s reading was almost the inverse of this logic. The model identified basic elements: bee, honey, but once again defaulted to a narrow band of labels such as ¨beehive interior backgrounds¨. The embedding analysis shows a strong drift toward the same commercial-descriptive centre visible across the dataset, with distances indicating minimal sensitivity to stylistic intent. The model recognises objects, but not the artistic framework that shapes them.

What makes Beehive particularly revealing is that it is not AI-generated image. This means the semantic loss cannot be attributed to mixed authorship. Even in the clearest examples of human-originated technique, OpenCLIP collapses meaning into commercial categories. The model treats a watercolour illustration as indistinguishable from a mass-market decorative print, demonstrating that stylistic nuance exists outside its representational range.

human tags vs ai tags of original surface pattern design artwork

Case Study 3 — Digital Meadow (ai-vs-brush_07)

Semantic friction score: 0.245
Technique: Adobe Express image generator + Photoshop post-processing
Visual effect: digital illustration look
Short description: A project with strong AI-inspired digital look, developed using Adobe Express and manual correction.

Digital Meadow is a fully hybrid pattern created directly using Adobe Express, selected specifically to examine how the model interprets the “clean”, recognisable AI brush aesthetic. The pattern exhibits typical AI-generated traits: overly smooth contours, homogenised shading, repetitive botanical forms and a distinct digital-flat tonality. My professional tags highlighted this artificial smoothness, the uncanny consistency of shapes, and the need for manual interventions to correct visual artefacts.

OpenCLIP’s interpretation erased this distinction completely. Despite the clearly synthetic quality of the pattern, the model assigned the same narrow vocabulary: ¨seamless fabric pattern 8k¨, ¨flowery wallpaper¨. None of the digital-artificial signature — something immediately visible to any trained designer — appeared in its reading. The embedding distances confirm that the model has no representational dimension corresponding to “AI-generated aesthetic”. It recognises neither the stylisation nor the technical origin.

This case is particularly important because the pattern originates inside the Adobe AI ecosystem itself. It reveals that even when the AI aesthetic is distinct, consistent and visually recognisable, the multimodal model does not consider it semantically meaningful. The digitality of the image becomes invisible within the embedding space, which again gravitates toward the most frequently occurring commercial categories.

In the context of this study, the tag “8k” does not refer to any measurable resolution, instead, it reflects a commercial internet descriptor commonly attached to floral wallpapers and seamless patterns, revealing how multimodal models inherit SEO-driven vocabulary rather than analysing artistic or technical properties.

Semantic Echo Loop

The concept of Semantic Echo emerged as a way to describe the recursive relationship between AI models, internet data and the artistic practices they increasingly intersect with. It refers to the gradual convergence of meaning that occurs when models learn from the very outputs they have influenced. Over time, this creates a closed semantic environment in which nuance, stylistic diversity and aesthetic intent are progressively compressed. Semantic Echo behaves similarly to memory distortion in human cognition: information repeated in the same form becomes more compressed, more confident, and less nuanced over time. While the technical vocabulary for this phenomenon includes terms such as model collapse or feedback drift, the core issue is cultural rather than purely computational: the model’s representational space becomes shaped not by the plurality of visual culture, but by the statistical pressure of its own past interpretations.

At the foundation of this loop lies the nature of internet data. The online image ecosystem is a product of human conventions, selection and convenience. Descriptions tend to be functional, product-oriented and optimised for visibility rather than accuracy. They rarely reference artistic style, medium, narrative or historical context. When a model such as OpenCLIP trains on this material, its semantic axes reflect these conventions. It does not have a dimension for style because style is not consistently present in the data; it does not have a dimension for aesthetic intent because intent leaves no textual trace that a contrastive model could reliably learn.

semantic echo loop appears while creating embeddings

The loop tightens further when AI-generated imagery begins circulating online. Generative models produce images shaped by the same semantic space learned by recognition models, and these images mingle with human-created content in search engines, marketplaces and social media. When new models train on this hybrid dataset, they ingest a mixture where the influence of earlier models is already embedded. The result is a slow but steady narrowing of meaning: less stylistic nuance, fewer long-tail categories, and a growing dominance of descriptors that are frequent, commercially important and easy for the model to detect.

This recursive system affects how AI reads artistic work. My own results illustrate this clearly. Early CLIP versions attempted to name style, albeit inconsistently. The updated OpenCLIP version does not treat style as a semantic category worth predicting. It behaves as though style has no representational weight at all. Instead, it prefers a small set of categories aligned with internet commerce: pattern, floral, wallpaper, seamless, storybook. The model’s descriptive field converges on what is abundant and rewardable in the data, not on what is meaningful in artistic practice.

The embedding analysis makes this drift visible. The professional tags I assign to my own work cluster around stylistic intent, narrative tone and visual decision-making — elements that belong to the creative process rather than merely its surface result. By contrast, OpenCLIP tags scatter widely along commercial axes. The semantic friction between these two vocabularies grows with each new model version, suggesting that the model is not improving its understanding of art but reducing the complexity of its interpretative space.

This is what makes Semantic Echo more than a metaphor. It is a structural feature of how large-scale multimodal systems currently evolve. When the model’s worldview loops back into its training data, the descriptive field becomes increasingly flat. AI appears confident, but its confidence rests on a narrow band of categories that dominate the contemporary image economy. Style, technique and intent remain outside this logic because they are not rewarded by the cycle.

Understanding this closed loop is essential for imagining alternatives. If models trained on uncurated internet data inevitably converge toward specific commercial descriptors, then any attempt to preserve stylistic nuance must begin with different forms of data curation — forms that involve artists, designers and the interpretative frameworks that guide human visual culture. The final section of this report explores what such an alternative could look like, and how artist-led datasets might reshape the future of multimodal embeddings.

Toward an Artist-Led Direction for Future Visual Models

The findings of this study suggest that the limitations observed in OpenCLIP’s interpretation of surface pattern design are not merely technical constraints but symptoms of a deeper structural issue: multimodal models learn from internet-scale data that undervalues style, technique and aesthetic intention. If these categories are to remain part of the algorithmic vocabulary of the next decade, the future of graphical AI will require a different source of semantic grounding, one shaped not only by SEO logic but also by artistic expertise. A promising direction would involve datasets in which illustrators and designers articulate their own stylistic intentions, production methods and visual nuances, creating a knowledge base that restores meaning where current models flatten it. Such a shift would not only reduce the semantic drift documented in this report, but also ensure that AI tools evolve in dialogue with the cultural, technical and imaginative dimensions of human creative practice. Rather than replacing artistic judgment, these systems could then be trained to recognise and preserve it, forming a more balanced relationship between machine vision and the visual languages artists have been developing for centuries.

Conclusion — Rethinking Meaning in the Age of Graphical AI

The AI vs Brush study demonstrates that when artistic work enters a multimodal embedding space, it is not simply translated — it is transformed. Style, technique and intention, which form the backbone of professional illustration and surface pattern design, are gradually stripped away, replaced by a narrow vocabulary shaped by the statistical habits of the internet. The updated OpenCLIP model does not misunderstand these categories; it no longer sees them as meaningful. This shift reveals a broader cultural dynamic: as AI systems increasingly learn from AI-generated imagery and SEO-optimised descriptors, their semantic landscape becomes progressively more uniform, commercial and surface-driven.

Yet the study also shows that this drift is not inevitable. By analysing the friction between human and machine interpretations, we can map where meaning collapses, where nuance survives and where future systems might be guided differently. The disappearance of style from the model’s vocabulary is not merely a technical finding — it is an invitation to rethink how we want visual models to relate to artistic practice.

For artists, contributing their expertise to large-scale AI systems is not a matter of giving knowledge away, but a strategic opportunity to gain influence over how visual meaning is encoded in future models — enabling tools that better understand style and technique, creating new professional roles such as dataset curator or creative AI consultant, strengthening the protection and recognisability of individual artistic styles, and ensuring that the next generation of graphical AI develops in genuine dialogue with human creativity rather than drifting further into commercial simplification.

If AI is to play a meaningful role in the future of design, it must be grounded in the values and interpretations of those who work with visual language every day. This report suggests that the path forward lies not in resisting technological change, but in ensuring that the creative sector contributes actively to the semantic foundations on which the next generation of models will be built.

If you are an artist, researcher, or a team working on embedding-based or multimodal AI systems and would like to collaborate on testing, analysing, or expanding this research, I warmly invite you to connect and explore the next stages of this work together.

Mona Tomassi Studio contact details

No Comments

Share:

Leave a Reply

Your e-mail address will not be published. Fields marked with * are mandatory.

0
    0
    Your cart
    Your cart is emptyReturn to the store