Every image posted on Facebook and Instagram gets a caption generated by an AI image analysis, and that AI has gotten a lot smarter. The improved system should be a treat for visually impaired users and may help you find your photos faster in the future.
Alt text is a field in the metadata of an image that describes its content: "A person standing in a field with a horse" or "a dog in a boat". This allows the image to be understood by people who cannot see it.
These descriptions are often added manually by a photographer or publication, but people who upload photos to social media generally don't care if they even get the chance. So the relatively recent ability to generate one automatically - the technology has just gotten pretty good in the last couple of years - has been hugely helpful in making social media more accessible overall.
Facebook created its machine alt text system in 2016, eons ago in the field of machine learning. The team has since come up with many improvements, making it faster and more detailed, and the latest update adds an option to generate a more detailed description upon request.
The improved system recognizes 10 times more objects and concepts than in the beginning, now around 1,200. And the descriptions include more details. What was once "Two people near a building" may now be "A selfie of two people near the Eiffel Tower". (Actual descriptions cover with "could be ..." and will avoid including wild assumptions.)
But there are more details than that, even if they're not always relevant. For example, in this image the AI notes the relative positions of people and objects:
Obviously people are above the drums, and the hats are above the people, none of which really need to be told for anyone to understand the gist. But consider an image described as "A house, some trees and a mountain". Is the house in the mountains or in front? Are the trees in front of or behind the house, or maybe on the mountain in the distance?
To adequately describe the picture, these details should be filled in, even if the general idea can be interpreted with fewer words. If a sighted person wants more details, they can take a closer look or click on the image for a larger version: someone who can't now has a similar option with this "generate detailed description of the image" command. (Activate it with a long press in the Android app or a custom action in iOS.)
Perhaps the new description could be something like "A house and trees in front of a mountain with snow on it." This paints a better picture, right? (To be clear, these examples are made up, but that's the kind of improvement to be expected.)
The new detailed description feature will arrive on Facebook first for testing, although the improved vocabulary will appear on Instagram soon. The descriptions are also simple so that they can be easily translated into other languages already supported by the apps, although the feature may not be implemented simultaneously in other countries.
