{"id":4082,"date":"2021-03-31T17:00:09","date_gmt":"2021-03-31T17:00:09","guid":{"rendered":"https:\/\/thenextweb.com\/?p=1345004"},"modified":"2021-03-31T17:00:09","modified_gmt":"2021-03-31T17:00:09","slug":"how-we-taught-google-translate-to-stop-being-sexist","status":"publish","type":"post","link":"https:\/\/www.londonchiropracter.com\/?p=4082","title":{"rendered":"How we taught Google Translate to stop being sexist"},"content":{"rendered":"\n<p>Online translation tools have helped us learn new languages, communicate across linguistic borders, and view foreign websites in our native tongue. But the artificial intelligence (AI) behind them is far from perfect, often replicating rather than rejecting the biases that exist within a language or a society.<\/p>\n<p>Such tools are especially vulnerable to gender stereotyping because some languages (such as English) don\u2019t tend to gender nouns, while others (such as German) do. When translating from English to German, translation tools have to decide which gender to assign English words like \u201ccleaner.\u201d Overwhelmingly, the tools conform to the stereotype, opting for the feminine word in German.<\/p>\n<p><a href=\"https:\/\/doi.org\/10.3389\/fpsyg.2018.01561\" target=\"_blank\" rel=\"nofollow noopener noreferrer\">Biases<\/a> are human: they\u2019re part of who we are. But when left unchallenged, biases can emerge in the form of concrete negative attitudes towards others. Now, our team has found a way to <a href=\"https:\/\/link.springer.com\/article\/10.1007\/s10676-021-09583-1\" target=\"_blank\" rel=\"nofollow noopener noreferrer\">retrain the AI<\/a> behind translation tools, using targeted training to help it to avoid gender stereotyping. Our method could be used in other fields of AI to help the technology reject, rather than replicate, biases within society.<\/p>\n<h2>Biased algorithms<\/h2>\n<p>To the dismay of their creators, AI algorithms often develop racist or sexist traits. <a href=\"https:\/\/www.forbes.com\/sites\/parmyolson\/2018\/02\/15\/the-algorithm-that-helped-google-translate-become-sexist\/?sh=22b6f48d7daa\" target=\"_blank\" rel=\"nofollow noopener noreferrer\">Google Translate<\/a> has been accused of stereotyping based on gender, such as its translations presupposing that all doctors are male and all nurses are female. Meanwhile, the AI language generator GPT-3 \u2013 which wrote an <a href=\"https:\/\/www.theguardian.com\/commentisfree\/2020\/sep\/08\/robot-wrote-this-article-gpt-3\" target=\"_blank\" rel=\"nofollow noopener noreferrer\">entire article<\/a> for the Guardian in 2020 \u2013 recently showed that it was also shockingly good at producing <a href=\"https:\/\/www.technologyreview.com\/2020\/07\/20\/1005454\/openai-machine-learning-language-generator-gpt-3-nlp\/\" target=\"_blank\" rel=\"nofollow noopener noreferrer\">harmful content and misinformation<\/a>.<\/p>\n<p>These AI failures aren\u2019t necessarily the fault of their creators. Academics and activists recently drew attention to <a href=\"https:\/\/www.bbc.co.uk\/news\/uk-england-oxfordshire-51738824?intlink_from_url=&amp;\" target=\"_blank\" rel=\"nofollow noopener noreferrer\">gender bias<\/a> in the Oxford English Dictionary, where sexist synonyms of \u201cwoman\u201d \u2013 such as \u201cbitch\u201d or \u201cmaid\u201d \u2013 show how even a constantly revised, academically edited catalog of words can contain biases that reinforce stereotypes and perpetuate everyday sexism.<\/p>\n<p>AI learns bias because it isn\u2019t built in a vacuum: it learns how to think and act by reading, analyzing, and categorizing existing data \u2013 like that contained in the Oxford English Dictionary. In the case of translation AI, we expose its algorithm to billions of words of textual data and ask it to recognize and learn from the patterns it detects. We call this process <a href=\"https:\/\/ieeexplore.ieee.org\/document\/5392560\" target=\"_blank\" rel=\"nofollow noopener noreferrer\">machine learning<\/a>, and along the way patterns of bias are learned as well as those of grammar and syntax.<\/p>\n<p>Ideally, the textual data we show AI won\u2019t contain bias. But there\u2019s an ongoing trend in the field towards building bigger systems trained on <a href=\"http:\/\/faculty.washington.edu\/ebender\/papers\/Stochastic_Parrots.pdf\" target=\"_blank\" rel=\"nofollow noopener noreferrer\">ever-growing data sets<\/a>. We\u2019re talking hundreds of billions of words. These are obtained from the internet by using undiscriminating text-scraping tools like Common Crawl and WebText2, which maraud across the web, gobbling up every word they come across.<\/p>\n<p>The sheer size of the resultant data makes it impossible for any human to actually know what\u2019s in it. But we do know that some of it comes from platforms like Reddit, which <a href=\"https:\/\/www.bbc.co.uk\/news\/technology-56099232\" target=\"_blank\" rel=\"nofollow noopener noreferrer\">has made headlines<\/a> for featuring offensive, false or conspiratorial information in users\u2019 posts.<\/p>\n<figure class=\"align-center \" readability=\"2.5748031496063\">\n<p><figure class=\"post-image post-mediaBleed aligncenter\"><img decoding=\"async\" loading=\"lazy\" src=\"https:\/\/images.theconversation.com\/files\/392527\/original\/file-20210330-13-2qqote.jpeg?ixlib=rb-1.1.0&amp;q=45&amp;auto=format&amp;w=754&amp;fit=clip\" sizes=\"(min-width: 1466px) 754px, (max-width: 599px) 100vw, (min-width: 600px) 600px, 237px\" alt=\"A magnifying glass over the Reddit logo on a web browser\" width=\"600\" height=\"409\" class=\" lazy\" data-lazy=\"true\" data-srcset=\"https:\/\/images.theconversation.com\/files\/392527\/original\/file-20210330-13-2qqote.jpeg?ixlib=rb-1.1.0&amp;q=45&amp;auto=format&amp;w=600&amp;h=409&amp;fit=crop&amp;dpr=1 600w, https:\/\/images.theconversation.com\/files\/392527\/original\/file-20210330-13-2qqote.jpeg?ixlib=rb-1.1.0&amp;q=30&amp;auto=format&amp;w=600&amp;h=409&amp;fit=crop&amp;dpr=2 1200w, https:\/\/images.theconversation.com\/files\/392527\/original\/file-20210330-13-2qqote.jpeg?ixlib=rb-1.1.0&amp;q=15&amp;auto=format&amp;w=600&amp;h=409&amp;fit=crop&amp;dpr=3 1800w, https:\/\/images.theconversation.com\/files\/392527\/original\/file-20210330-13-2qqote.jpeg?ixlib=rb-1.1.0&amp;q=45&amp;auto=format&amp;w=754&amp;h=514&amp;fit=crop&amp;dpr=1 754w, https:\/\/images.theconversation.com\/files\/392527\/original\/file-20210330-13-2qqote.jpeg?ixlib=rb-1.1.0&amp;q=30&amp;auto=format&amp;w=754&amp;h=514&amp;fit=crop&amp;dpr=2 1508w, https:\/\/images.theconversation.com\/files\/392527\/original\/file-20210330-13-2qqote.jpeg?ixlib=rb-1.1.0&amp;q=15&amp;auto=format&amp;w=754&amp;h=514&amp;fit=crop&amp;dpr=3 2262w\"><figcaption><a href=\"https:\/\/thenextweb.com\/neural\/2021\/03\/31\/google-translate-is-sexist-ai-can-solve-syndication\/#\" data-url=\"https:\/\/twitter.com\/intent\/tweet?url=https%3A%2F%2Fthenextweb.com%2Fneural%2F2021%2F03%2F31%2Fgoogle-translate-is-sexist-ai-can-solve-syndication%2F&amp;via=thenextweb&amp;related=thenextweb&amp;text=Check out this picture on: Some of the text users share on Reddit contains language we might prefer our translation tools not to learn. Gil C\/Shutterstock\" data-title=\"Share Some of the text users share on Reddit contains language we might prefer our translation tools not to learn. Gil C\/Shutterstock on Twitter\" data-width=\"685\" data-height=\"500\" class=\"post-image-share popitup\" title=\"Share Some of the text users share on Reddit contains language we might prefer our translation tools not to learn. Gil C\/Shutterstock on Twitter\"><i class=\"icon icon--inline icon--twitter--dark\"><\/i><\/a>Some of the text users share on Reddit contains language we might prefer our translation tools not to learn. <a href=\"https:\/\/www.shutterstock.com\/image-photo\/lisbon-portugal-february-6-2014-photo-175132031\" target=\"_blank\" rel=\"nofollow noopener noreferrer\">Gil C\/Shutterstock<\/a><span><\/span><\/figcaption><\/figure>\n<\/p>\n<\/figure>\n<h2>New translations<\/h2>\n<p>In <a href=\"https:\/\/link.springer.com\/article\/10.1007\/s10676-021-09583-1\" target=\"_blank\" rel=\"nofollow noopener noreferrer\">our research<\/a>, we wanted to search for a way to counter the bias within textual data-sets scraped from the internet. Our experiments used a randomly selected part of an existing English-German corpus (a selection of text) that originally contained 17.2 million pairs of sentences \u2013 half in English, half in German.<\/p>\n<p>As we\u2019ve highlighted, German has gendered forms for nouns (doctor can be \u201c<em>der Arzt<\/em>\u201d for male, \u201c<em>die \u00c4rztin<\/em>\u201d for female) where in English we don\u2019t gender these noun forms (with some exceptions, <a href=\"https:\/\/www.thestage.co.uk\/your-views\/actor-or-actress-the-debate-continues-your-views-december-14\" target=\"_blank\" rel=\"nofollow noopener noreferrer\">themselves contentious<\/a>, like \u201cactor\u201d and \u201cactress\u201d).<\/p>\n<p>Our analysis of this data revealed clear gender-specific imbalances. For instance, we found that the masculine form of engineer in German (<em>der Ingenieur<\/em>) was 75 times more common than its feminine counterpart (<em>die Ingenieurin<\/em>). A translation tool trained on this data will inevitably replicate this bias, translating \u201cengineer\u201d to the male \u201c<em>der Ingenieur.<\/em>\u201d So what can be done to avoid or mitigate this?<\/p>\n<h2>Overcoming bias<\/h2>\n<p>A seemingly straightforward answer is to \u201cbalance\u201d the corpus before asking computers to learn from it. Perhaps, for instance, adding more female engineers to the corpus would prevent a translation system from assuming all engineers are men.<\/p>\n<p>Unfortunately, there are difficulties with this approach. Translation tools are trained for days on billions of words. Retraining them by altering the gender of words is possible, but it\u2019s inefficient, expensive and complicated. Adjusting the gender in languages like German is especially challenging because, in order to make grammatical sense, several words in a sentence may need to be changed to reflect the gender swap.<\/p>\n<p>Instead of this laborious gender rebalancing, we decided to retrain existing translation systems with targeted lessons. When we spotted a bias in existing tools, we decided to retrain them on new, smaller data-sets \u2013 a bit like an afternoon of gender-sensitivity training at work.<\/p>\n<p>This approach takes a fraction of the time and resources needed to train models from scratch. We were able to use just a few hundred selected translation examples \u2013 instead of millions \u2013 to adjust the behavior of translation AI in targeted ways. When testing gendered professions in translation \u2013 as we had done with \u201cengineers\u201d \u2013 the accuracy improvements after adapting were about nine times higher than the \u201cbalanced\u201d retraining approach.<\/p>\n<p>In our research, we wanted to show that tackling hidden biases in huge data-sets doesn\u2019t have to mean laboriously adjusting millions of training examples, a task which risks being dismissed as impossible. Instead, bias from data can be targeted and unlearned \u2013 a lesson that other <a href=\"https:\/\/hbr.org\/2019\/10\/what-do-we-do-about-the-biases-in-ai\" target=\"_blank\" rel=\"nofollow noopener noreferrer\">AI researchers<\/a> can apply to their own work.<\/p>\n<p><em>This article by&nbsp;<a href=\"https:\/\/theconversation.com\/profiles\/stefanie-ullmann-856014\" target=\"_blank\" rel=\"nofollow noopener noreferrer\">Stefanie Ullmann<\/a>, Postdoctoral Research Associate, <a href=\"https:\/\/theconversation.com\/institutions\/university-of-cambridge-1283\" target=\"_blank\" rel=\"nofollow noopener noreferrer\">University of Cambridge<\/a> and <a href=\"https:\/\/theconversation.com\/profiles\/danielle-saunders-1221845\" target=\"_blank\" rel=\"nofollow noopener noreferrer\">Danielle Saunders<\/a>, Research Student, Department of Engineering, <a href=\"https:\/\/theconversation.com\/institutions\/university-of-cambridge-1283\" target=\"_blank\" rel=\"nofollow noopener noreferrer\">University of Cambridge<\/a>&nbsp;is republished from <a href=\"https:\/\/theconversation.com\" target=\"_blank\" rel=\"nofollow noopener noreferrer\">The Conversation<\/a> under a Creative Commons license. Read the <a href=\"https:\/\/theconversation.com\/online-translators-are-sexist-heres-how-we-gave-them-a-little-gender-sensitivity-training-157846\" target=\"_blank\" rel=\"nofollow noopener noreferrer\">original article<\/a>.<\/em><\/p>\n<p class=\"c-post-pubDate\"> Published March 31, 2021 \u2014 17:00 UTC <\/p>\n<p> <a href=\"https:\/\/thenextweb.com\/neural\/2021\/03\/31\/google-translate-is-sexist-ai-can-solve-syndication\/\">Source<\/a><\/p>\n","protected":false},"excerpt":{"rendered":"<p>Online translation tools have helped us learn new languages, communicate across linguistic borders, and view foreign websites in our native tongue. But the artificial intelligence (AI) behind them is far from perfect,&#8230;<\/p>\n","protected":false},"author":1,"featured_media":4083,"comment_status":"open","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":[],"categories":[1],"tags":[],"_links":{"self":[{"href":"https:\/\/www.londonchiropracter.com\/index.php?rest_route=\/wp\/v2\/posts\/4082"}],"collection":[{"href":"https:\/\/www.londonchiropracter.com\/index.php?rest_route=\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.londonchiropracter.com\/index.php?rest_route=\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.londonchiropracter.com\/index.php?rest_route=\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/www.londonchiropracter.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcomments&post=4082"}],"version-history":[{"count":0,"href":"https:\/\/www.londonchiropracter.com\/index.php?rest_route=\/wp\/v2\/posts\/4082\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.londonchiropracter.com\/index.php?rest_route=\/wp\/v2\/media\/4083"}],"wp:attachment":[{"href":"https:\/\/www.londonchiropracter.com\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=4082"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.londonchiropracter.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcategories&post=4082"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.londonchiropracter.com\/index.php?rest_route=%2Fwp%2Fv2%2Ftags&post=4082"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}