Experiencing Art Through Language

Colby researchers are using AI to make museum collections accessible for everyone

The audio guides and factual placards that accompany art exhibitions do not allow visually impaired art lovers to form a mental image of the artwork. To change this, two Colby professors are using computer vision and artificial intelligence to develop accessible descriptions tailored to the needs of blind and low-vision people. (Photo by Ashley L. Conti)
Share
By Laura Meader
September 29, 2026

Imagine yourself in an art gallery, examining brushstrokes, savoring color palettes, and feeling awestruck or inspired. For visitors with blindness or low vision, that kind of intimate engagement with the visual arts is out of reach.

Because audio guides and factual placards offer few clues to visual composition, visually impaired art lovers cannot form a mental image of the artwork.

To change this, two Colby professors are using computer vision and artificial intelligence to develop accessible descriptions tailored to the needs of the blind and low-vision community. This spring, Colby students will join the project through a new course designed to address the issue. 

Computer scientist Stacy Doore and art historian Véronique Plesch are training a model and building a low-cost, open-source tool that will help museum curators generate descriptions using structured prompts and predefined standards. 

“The idea is that curators are experts in artwork but aren’t typically trained in generating accessible descriptions. The model is trained to generate accessible descriptions with the human expert as the guide,” said Doore, the Clare Booth Luce Associate Professor of Computer Science.

By focusing on core spatial elements—such as canvas orientation and composition—their interdisciplinary project will help create a vivid mental image, allowing blind and low-vision visitors to connect with artwork.

“I’m hoping to help museums and curators to make their collections more accessible,” said Doore, whose work focuses on assistive technology, spatial descriptions, and navigation. “Not only in the museum, but also online.”

Stacy Doore, the Clare Boothe Luce Associate Professor of Computer Science, practices responsible AI using a “slow-tech” approach, which includes intentionality and transparency in her decisions, documentation, and partnerships. (Photo by Ashley L. Conti)

Starting from zero

Doore began her pioneering research in 2018 as a visiting professor at Bowdoin College. At the time, accessible descriptions that included spatial information largely did not exist, and no one had trained a machine to generate them.

When she joined Colby’s faculty in 2020, Doore brought the project with her as part of her Immersive Navigation Systems and Inclusive Technology Ethics (INSITE) Lab. Using open-source large language models and AI tools such as computer vision, she has since developed an onsite, open-source model housed on a high-performance lab computer to ensure her data remains secure and the artwork used for training is copyright-protected if the work is not in the public domain. 

The model can now generate increasingly accurate, accessible descriptions. But it’s taken time.

AI-generated descriptions need human fine-tuning to eliminate errors, clarify vague language, and remove wordiness. Descriptions should also follow established guidelines for the blind and low-vision, or BLV, community. Recommendations include “the use of consistent terms, rules for information ordering, and spatial concept presentation” in ways to help listeners focus while “reducing the cognitive load that is inherent in the use of longer verbal descriptions,” Doore wrote in a paper published in the Journal of Imaging in 2024.

Doore practices responsible AI using a “slow-tech” approach, which includes intentionality and transparency in her decisions, documentation, and partnerships. “ I want to model for my students all of the important ethical decisions that have to be made during the development and research processes,” she said.

The accessible art descriptions project continues the work she’s done throughout her career, including a five-year undertaking titled VisionWay: Accessibility-aware Path Selection for Wayfinding, supported by a $2.7 million grant from the National Eye Institute, and her Assistive Agile Robots project.

Brevity and clarity

In 2024, Doore invited Plesch to collaborate on the project as an art historian with expertise in word and image studies. Plesch is a respected scholar in the field and a former president of the International Association of Word and Image Studies.

“Describing a work of art is the first step in any art historical analysis,” said Plesch, the Ellerton M. and Edith K. Jetté Professor of Art. She eagerly joined Doore’s project because of how it intersects with her scholarly work and her growing interest in artificial intelligence. The course she developed for the project is supported by a Fellowship in Digital Scholarship from Colby’s Center for the Arts and Humanities. 

Véronique Plesch, the Ellerton M. and Edith K. Jetté Professor of Art, brings her expertise in word and image studies to the project, transforming AI-generated descriptions into nuanced, spatial descriptions tailored to the needs of the blind and low-vision community. (Photo by Brian Fitzgerald)

As she evaluated accessible descriptions that Doore’s model had generated, Plesch removed hypothetical and vague language, unnecessary words, and inaccuracies. Clarity and accuracy are of utmost importance in forming a mental image, but so is brevity, she emphasized. In one example, she reduced a generated description from 142 to 86 words.

“The verbosity is not going to help,” Plesch emphasizes. “The comparison I like to make is when someone gives you driving directions. If they’re too long, by the time they get to the end, you’ve forgotten how to start. The same thing applies here. It has to be quite succinct.”

Plesch is also contributing by improving “layered descriptions,” which involve increasingly complex and detailed descriptions of artwork, depending on the viewer’s interest. Doore’s research, based on surveys with visually impaired participants, has shown that layered descriptions interest the BLV community.

‘In pursuit of cultural equity’

Despite Doore and Plesch’s progress, the project needs a larger collection of accurate descriptive examples to continue refining the model.

One source of examples will come from museum curators who test the interface Doore has built that connects to the model. Using the simple interface, curators can select public-domain artwork and use their expertise to refine the descriptions the model generates. Their refined descriptions, like those Plesch improved, will go into a repository to continue training the model. 

Another source is Plesch’s course, which she’s titled The Mind’s Eye: Artwork Description from Ancient Ekphrasis to AI. As students “explore the long history of grappling with the visual through the verbal,” they will work closely with Doore and her student researchers. Plesch’s students will be “humans-in-the-loop” experts, transforming AI-generated data into nuanced, spatial descriptions.

The course will conclude with a public symposium and an “Edit-a-Thon” in collaboration with the INSITE Lab next spring. Students will showcase their work and invite members from the blind and low-vision community to audit and further refine their descriptions. The goal is to create a feedback loop that lets students see their research become a tangible resource.

It promises a powerful learning experience for students and important support for the overall project.

“By bridging art historical practice with accessibility technology,” said Plesch, “students will move from passive learners to active contributors in the pursuit of cultural equity.”

related

Highlights