The Times Australia
The Times World News

.
The Times Real Estate

.

If AI image generators are so smart, why do they struggle to write and count?

  • Written by Seyedali Mirjalili, Professor, Director of Centre for Artificial Intelligence Research and Optimisation, Torrens University Australia
If AI image generators are so smart, why do they struggle to write and count?

Generative AI tools such as Midjourney, Stable Diffusion and DALL-E 2 have astounded us with their ability to produce remarkable images in a matter of seconds[1].

Despite their achievements, however, there remains a puzzling disparity between what AI image generators can produce and what we can. For instance, these tools often won’t deliver satisfactory results for seemingly simple tasks such as counting objects and producing accurate text.

If generative AI has reached such unprecedented heights in creative expression, why does it struggle with tasks even a primary school student could complete?

Exploring the underlying reasons helps sheds light on the complex numerical nature of AI, and the nuance of its capabilities.

AI’s limitations with writing

Humans can easily recognise text symbols (such as letters, numbers and characters) written in various different fonts and handwriting. We can also produce text in different contexts, and understand how context can change meaning.

Current AI image generators lack this inherent understanding. They have no true comprehension of what any text symbols mean. These generators are built on artificial neural networks trained on[2] massive amounts of image data, from which they “learn” associations and make predictions.

Combinations of shapes in the training images are associated with various entities. For example, two inward-facing lines that meet might represent the tip of a pencil, or the roof of a house.

But when it comes to text and quantities, the associations must be incredibly accurate, since even minor imperfections are noticeable. Our brains can overlook slight deviations in a pencil’s tip, or a roof – but not as much when it comes to how a word is written, or the number of fingers on a hand.

Read more: Both humans and AI hallucinate — but not in the same way[3]

As far as text-to-image models are concerned, text symbols are just combinations of lines and shapes. Since text comes in so many different styles – and since letters and numbers are used in seemingly endless arrangements – the model often won’t learn how to effectively reproduce text.

AI-generated image produced in response to the prompt ‘KFC logo’. Imagine AI[4]

The main reason for this is insufficient training data. AI image generators require much more training data[5] to accurately represent text and quantities than they do for other tasks.

The tragedy of AI hands

Issues also arise when dealing with smaller objects that require intricate details, such as hands[6].

Two AI-generated images produced in response to the prompt ‘young girl holding up ten fingers, realistic’. Shutterstock AI

In training images, hands are often small, holding objects, or partially obscured by other elements. It becomes challenging for AI to associate the term “hand” with the exact representation of a human hand with five fingers.

Consequently, AI-generated hands often look misshapen[7], have additional or fewer fingers, or have hands partially covered by objects such as sleeves or purses.

We see a similar issue when it comes to quantities. AI models lack a clear understanding of quantities, such as the abstract concept of “four”.

As such, an image generator may respond to a prompt for “four apples” by drawing on learning from myriad images featuring many quantities of apples – and return an output with the incorrect amount.

In other words, the huge diversity of associations within the training data impacts the accuracy of quantities in outputs.

Three AI-generated images produced in response to the prompt ‘5 soda cans on a table’. Shutterstock AI

Will AI ever be able to write and count?

It’s important to remember text-to-image and text-to-video conversion is a relatively new concept in AI. Current generative platforms are “low-resolution” versions of what we can expect in the future.

With advancements being made[8] in training processes and AI technology, future AI image generators will likely be much more capable of producing accurate visualisations.

It’s also worth noting most publicly accessible AI platforms don’t offer the highest level of capability. Generating accurate text and quantities demands highly optimised and tailored networks, so paid subscriptions to more advanced platforms will likely deliver better results.

References

  1. ^ a matter of seconds (www.zdnet.com)
  2. ^ trained on (www.assemblyai.com)
  3. ^ Both humans and AI hallucinate — but not in the same way (theconversation.com)
  4. ^ Imagine AI (www.imagine.art)
  5. ^ more training data (decrypt.co)
  6. ^ such as hands (www.buzzfeednews.com)
  7. ^ often look misshapen (twitter.com)
  8. ^ advancements being made (theconversation.com)

Read more https://theconversation.com/if-ai-image-generators-are-so-smart-why-do-they-struggle-to-write-and-count-208485

The Times Features

The Impact of Healthy Keto on Digestive Health

By adopting the Healthy Keto diet, people can realistically shed pounds, recharge their batteries, and sharpen their mental focus. While the spotlight shines on the major advanta...

5 Hottest Australian Beach Road Trips In 2025

Australia is currently in the middle of a very hot summer, and the many who are not yet back at work are enjoying the incredible beach lifestyle that the country is credited for...

Fixed vs Variable Interest Rates – What You Need To Know

As any reputable home or commercial loan broker will tell you, interest rates are key when taking out a loan or mortgage in Australia. During the loan process, borrowers often ch...

There Are No Boundaries In Love and There Does Not Need To Be!

Love is unpredictable and has its own language. It is the most healing and transformative quality of our existence, it does not know separation by race, boundaries, borders, gove...

Restorative massage: Technique and Contraindications

Any massage, including restorative massage, not only gives a person pleasure and enjoyment but also has a beneficial and therapeutic effect on the whole organism. To date, resto...

Tips on Choosing the Right Tibetan Singing Bowl for You

The art of mindfulness can really do wonders for your life. In fact, it has been proven to help people thrive in the most difficult situations, including the pandemic, and being ...

Times Magazine

How BIM Software is Transforming Architecture and Engineering

Building Information Modeling (BIM) software has become a cornerstone of modern architecture and engineering practices, revolutionizing how professionals design, collaborate, and execute projects. By enabling more efficient workflows and fostering ...

How 32-Inch Computer Monitors Can Increase Your Workflow

With the near-constant usage of technology around the world today, ergonomics have become crucial in business. Moving to 32 inch computer monitors is perhaps one of the best and most valuable improvements you can possibly implement. This-sized moni...

Top Tips for Finding a Great Florist for Your Sydney Wedding

While the choice of wedding venue does much of the heavy lifting when it comes to wowing guests, decorations are certainly not far behind. They can add a bit of personality and flair to the traditional proceedings, as well as enhancing the venue’s ...

Avant Stone's 2025 Nature's Palette Collection

Avant Stone, a longstanding supplier of quality natural stone in Sydney, introduces the 2025 Nature’s Palette Collection. Curated for architects, designers, and homeowners with discerning tastes, this selection highlights classic and contemporary a...

Professional-Grade Tactical Gear: Why 5.11 Tactical Leads the Field

When you're out in the field, your gear has to perform at the same level as you. In the world of high-quality equipment, 5.11 Tactical has established itself as a standard for professionals who demand dependability. Regardless of whether you’re inv...

Lessons from the Past: Historical Maritime Disasters and Their Influence on Modern Safety Regulations

Maritime history is filled with tales of bravery, innovation, and, unfortunately, tragedy. These historical disasters serve as stark reminders of the challenges posed by the seas and have driven significant advancements in maritime safety regulat...

LayBy Shopping