2020 has reinvented each sector as advancing in the light of COVID-19: civil rights movements, an election year and countless other important moments. On a human level, we had to adapt to a new way of life. We started to accept these changes and understand how to live our lives according to these new pandemic rules. As humans settle, AI is struggling to keep up.
The problem with AI training in 2020 is that, suddenly, we have changed our social and cultural norms. The truths we have taught these algorithms are often no longer actually true. With visual AI in particular, we ask him to immediately interpret the new way we live with an updated context he does not yet have.
The algorithms are still adjusting to the new visual queues and trying to figure out how to identify them accurately. As visual AI reaches, we also need renewed importance on routine updates in the AI training process, so that we can correct inaccurate training datasets and pre-existing open source models.
Machine vision models are struggling to properly label representations of new scenes or situations we find ourselves in during the COVID-19 era. The categories have changed. For example, suppose there is an image of a father working at home while his son is playing. Artificial intelligence is still classifying it as "free time" or "relaxation". It is not identifying this as "work" or "office", despite the fact that working with your children next to you is the very common reality for many families in this period.
On a more technical level, we physically have different pixel representations of our world. At Getty Images, we trained AI to "see". This means that algorithms can identify images and classify them based on the pixel composition of that image and decide what it includes. Quickly changing the way we go in our daily lives means that we are also moving what a category or tag entails (such as "cleanliness").
Think of it this way: cleaning can now include cleaning surfaces that appear visually clean. Previously, algorithms have been taught that disaster is needed to represent cleanliness. Now it looks very different. Our systems must be retrained to take into account these redefined category parameters.
This also applies to a smaller scale. Someone could grab a door handle with a small towel or clean the steering wheel while sitting in the car. What was once a trivial detail now matters as people try to stay safe. We have to grasp these little nuances, so it's tagged appropriately. So AI can begin to understand our world in 2020 and produce accurate results.
Another problem for AI right now is that machine learning algorithms are still trying to figure out how to identify and classify faces with masks. Faces are detected only as the upper half of the face or as two faces: one with a mask and a second only with the eyes. This creates inconsistencies and inhibits the accurate use of face detection patterns.
One way forward is to retrain the algorithms for better performance when only the top of the face (above the mask) is provided. The mask problem is similar to classic face detection challenges like someone who wears sunglasses or detects someone's face in profile. Now even masks are on the agenda.
What it shows us is that artificial vision models still have a long way to go before we can truly "see" in our ever-changing social landscape. The way to counter this is to build reliable data sets. Hence, we can train vision models to account for the myriad of ways in which a face can be obstructed or covered.
At this point, we are expanding the parameters of what the algorithm sees as a face: whether it's a person wearing a mask in a grocery store, a nurse wearing a mask as part of their daily work or a person covering their face for religious reasons.
Since we create the content necessary to create these reliable data sets, we should be aware of potentially increased involuntary bias. While there will always be bias within AI, we now see unbalanced datasets that describe our new normal. For example, we are seeing more images of whites wearing masks than other ethnicities.
This may be the result of strict residence orders in which photographers have limited access to communities other than their own and are unable to diversify their subjects. It could be due to the ethnicity of the photographers who choose to take up this topic. Or, due to the level of impact that COVID-19 has had on different regions. Regardless of the reason, this imbalance will lead to algorithms that can more accurately detect a white person wearing a mask than any other race or ethnicity.
Data scientists and those who build products with models have a greater responsibility to verify the accuracy of models in light of changes in social norms. Routine checks and updates of data and training models are critical to ensuring the quality and robustness of the models, now more than ever. If the results are not accurate, the data scientists can quickly identify them and correct them correctly.
It is also worth mentioning that our current way of life is here to stay for the foreseeable future. For this reason, we need to be cautious about the open source datasets that we are leveraging for training purposes. The datasets that can be edited should. Open source models that cannot be edited must have a disclaimer, so it is clear which projects could be adversely affected by outdated training data.
Identifying the new context that we are asking the system to understand is the first step in advancing visual AI. So we need more content. Other representations of the world around us - and the different perspectives of it. As we are accumulating this new content, take stock of possible new biases and ways to retrain existing open source datasets. We must all monitor inconsistencies and inaccuracies. Persistence and dedication to the retraining of artificial vision models is the way we will bring AI in 2020.
