Santa Fe
Institute
  • Research
    • Themes
    • Projects
    • SFI Press
    • Researchers
    • Publications
    • Library
    • Sponsored Research
    • Fellowships
    • Miller Scholarships
  • News + Events
    • News
    • Newsletters
    • Podcasts
    • SFI in the Media
    • Media Center
    • Events
    • Community
    • Journalism Fellowship
  • Education
    • Programs
    • Projects
    • Alumni
    • Complexity Explorer
    • Education FAQ
    • Postdoctoral Research
    • Education Supporters
  • People
    • Researchers
    • Fractal Faculty
    • Staff
    • Miller Scholars
    • Trustees
    • Governance
    • Resident Artists
    • Research Supporters
  • Applied Complexity
    • Office
    • Applied Projects
    • ACtioN
    • Applied Fellows
    • Studios
    • Applied Events
    • Login
  • Give
    • Give Now
    • Ways to Give
    • Contact
  • About
    • About SFI
    • Engage
    • Complex Systems
    • FAQ
    • Campuses
    • Jobs
    • Contact
    • Library
    • Employee Portal

Science for a Complex World

Events

Here's what's happening

Give

You make SFI possible

Subscribe

Sign up for research news

Connect

Follow us on social media

© 2026 Santa Fe Institute. All rights reserved. This site is supported by the Miller Omega Program.

Home / News

Study: Visual Analogies for AI

Examples of some of the visual tasks to test intelligence in humans vs AI. (image: Appendix A in "The ConceptARC Benchmark: Evaluating Understanding and Generalization in the ARC Domain")
September 25, 2023

The field of artificial intelligence has long been stymied by the lack of an answer to its most fundamental question: What is intelligence? AIs such as GPT-4 have highlighted this uncertainty: some researchers believe that GPT models are showing glimmers of genuine intelligence but others disagree.

To address these arguments, we need concrete tasks to pin down and test the notion of intelligence, argue SFI researchers Arseny Moskvichev, Melanie Mitchell, and Victor Vikram Odouard in a new paper in Transactions on Machine Learning Research. The authors provide just that — and find that even the most advanced AIs still lag far behind humans in their ability to abstract and generalize concepts.

The team created evaluation puzzles — based on a domain developed by Google researcher François Chollet — that focus on visual analogy-making, capturing basic concepts such as above, below, center, inside, and outside. Human and AI test-takers were shown several patterns demonstrating a concept and then asked to apply that concept to a different image. The accompanying figure shows tests of the notion of sameness.

Examples of some of the visual tasks the authors of a recent study posed in testing intelligence. (image: Appendix A in "The ConceptARC Benchmark: Evaluating Understanding and Generalization in the ARC Domain")

 

These visual puzzles were very easy for humans: For example, they got the notion of sameness correct 88 percent of the time. But GPT-4 struggled, only getting 23 percent of these puzzles right. So the researchers conclude that, currently, AI programs are still weak at visual abstract reasoning.

“We reason a lot by analogies, so that’s why it’s such an interesting question,” Moskvichev says. The team’s use of novel visual puzzles ensured that the machines hadn’t encountered them before. GPT-4 was trained on large portions of the internet, so it was important to avoid anything it might have encountered already, to be certain it wasn’t just parroting existing text rather than demonstrating its own understanding. That’s why recent results like an AI’s ability to score well on a Bar exam aren’t a good test of its true intelligence.

The team believes that as time goes on and AI algorithms improve, developing evaluation routines will get progressively more difficult and more important. Rather than trying to create one test of AI intelligence, we should design more carefully curated datasets focusing on specific facets of intelligence. “The better our algorithms become, the harder it is to figure out what they can and can’t do,” Moskvichev says. “So we need to be very thoughtful in developing evaluation datasets.

Read the study, "The ConceptARC Benchmark: Evaluating Understanding and Generalization in the ARC Domain" in Transactions on Machine Learning Research (August, 2023)





Share
  • Sign Up For SFI News
News Media Contact

Santa Fe Institute

Office of Communications
news@santafe.edu
505-984-8800



  • Tags
  • SFI News Release
  • Research


More SFI News

View All News

Andreas Wagner awarded ERC Advanced Grant

SFI Professors Give Judges Advice on AI

John Krakauer named director of Champalimaud's Centre for Restorative Neurotechnology

Book Review: "Tipping out of Trouble: How Societies Transformed and How We Can Do So Again"

In Memoriam: Jim Rutt

Does intelligence ‘emerge’ in large language models?

Your dominant hand is made, not born

A bird song almost too quiet to hear

Model redefining conformity excels against real-world data

Decoding animal minds

SFI External Professor Nicholas de Monchaux named Dean of UC Berkeley College of Environmental Design

Simon Levin named Fellow of the Royal Society

Brian Enquist receives Robert H. MacArthur Award

Han van der Maas named director of Amsterdam’s Institute for Advanced Study

Marina Dubova receives Dissertation Prize

Smart parts for smart wholes

Aaron Clauset receives honors from AAAS and University of New Mexico

Laurent Hébert-Dufresne receives Erdős-Rényi Prize

Why noise may be the key to understanding cell group patterns

Reinventing democracy before it breaks