{"id":12160,"date":"2022-09-13T13:00:20","date_gmt":"2022-09-13T13:00:20","guid":{"rendered":"http:\/\/TheNextWeb=1390603"},"modified":"2022-09-13T13:00:20","modified_gmt":"2022-09-13T13:00:20","slug":"what-does-europes-approach-to-data-privacy-mean-for-gpt-and-dall-e","status":"publish","type":"post","link":"https:\/\/www.londonchiropracter.com\/?p=12160","title":{"rendered":"What does Europe\u2019s approach to data privacy mean for GPT and DALL-E?"},"content":{"rendered":"\n<p><span>The global <a href=\"https:\/\/thenextweb.com\/topic\/artificial-intelligence\" target=\"_blank\" rel=\"noopener noreferrer\">AI<\/a> explosion has supercharged the need for a common sense, human-centered methodology for dealing with data privacy and ownership. Leading the way is Europe\u2019s General Data Protection Regulation (GDPR), but there\u2019s more than just personally identifiable information (PII) at stake in the modern market.<\/span><\/p>\n<p><span>What about the data we generate as content and art? It\u2019s certainly not legal to copy someone else\u2019s work and then present it as your own. But there are AI systems that attempt to <\/span><i><span>scrape<\/span><\/i><span> as much human-generated content from the web as possible in order to generate content that\u2019s similar.&nbsp;<\/span><\/p>\n<p><span>Can GDPR or any other EU-centered policies protect this kind of content? As it turns out, like most things in the machine learning world, it depends on the data.&nbsp;<\/span><\/p>\n<h2><span>Privacy vs ownership<\/span><\/h2>\n<div class=\"inarticle-wrapper neural channel-cta hs-embed-tnw\">\n<div id=\"hs-embed-tnw\" class=\"channel-cta-wrapper\" readability=\"6\">\n<div class=\"channel-cta-img\"><img decoding=\"async\" class=\"js-lazy\" src=\"https:\/\/cdn0.tnwcdn.com\/wp-content\/blogs.dir\/1\/files\/2022\/07\/neural.webp\"><\/div>\n<p><noscript><img decoding=\"async\" src=\"src='https:\/\/cdn0.tnwcdn.com\/wp-content\/blogs.dir\/1\/files\/2022\/07\/neural.webp'\"><\/noscript><\/p>\n<div class=\"channel-cta-input\" readability=\"7\">\n<h2 class=\"channel-cta-title\">Greetings, humanoids<\/h2>\n<p class=\"channel-cta-tagline\">Subscribe to our newsletter now for a weekly recap of our favorite AI stories in your inbox.<\/p>\n<\/div>\n<\/div>\n<\/div>\n<p><span>GDPR\u2019s primary purpose is to protect European citizens from harmful actions and consequences related to the misuse, abuse, or exploitation of their private information. It\u2019s not much use to citizens (or organizations) when it comes to protecting intellectual property (IP).&nbsp;<\/span><\/p>\n<p><span>Unfortunately, the policies and regulations put in place to protect IP are, to the best of our knowledge, not equipped to cover data scraping and anonymization. That makes it difficult to understand exactly where the regulations apply when it comes to scraping the web for content.&nbsp;<\/span><\/p>\n<p><span>These techniques, and the data they obtain, are used to create massive databases for use in training large AI models such as OpenAI\u2019s GPT-3 and DALL-E 2 systems.<\/span><\/p>\n<p><span>The only way to teach an AI to imitate humans is to expose it to human-generated data. And the more data you shove in an AI system, the more robust its output tends to be.&nbsp;<\/span><\/p>\n<p><span>It works like this: imagine you draw a picture of a flower and post it to an online forum for artists. Using scraping techniques, a tech outfit sucks up your image along with billions of others so it can create a massive dataset of artwork. The next time someone asks the AI to generate an image of a \u201cflower,\u201d there\u2019s a greater-than-zero possibility that your work will feature in the AI\u2019s interpretation of the prompt.&nbsp;<\/span><\/p>\n<p><span>As to whether such use would be ethical remains an open question.&nbsp;<\/span><\/p>\n<h2><span>Public data versus PII<\/span><\/h2>\n<p><span>While the GDPR\u2019s regulatory oversight could be described as far-reaching when it comes to protecting private information and giving Europeans the<\/span> <a href=\"https:\/\/gdpr.eu\/right-to-be-forgotten\/\" target=\"_blank\" rel=\"nofollow noopener noreferrer\"><i><span>right to erasure<\/span><\/i><\/a><span>, it seemingly does very little to protect content from scraping. However, that doesn\u2019t mean GDPR and other EU regulations are entirely feckless in this regard.&nbsp;<\/span><\/p>\n<p><span>Individuals and organizations have to follow very specific rules for scraping PII, lest they fall afoul of the law \u2014 something that can become quite costly.&nbsp;<\/span><\/p>\n<p><span>As an example, it\u2019s becoming nigh impossible for Clearview AI, a company that builds facial recognition databases for government use by <\/span><i><span>scraping<\/span><\/i><span> social media data, to conduct business in Europe. EU watchdogs from at least seven nations have either issued hefty fines already or recommended fines over the company\u2019s refusal to comply with GDPR and similar regulations.<\/span><\/p>\n<p><span>On the complete other side of the spectrum, companies such as Google, OpenAI, and Meta employ similar <\/span><i><span>data scraping <\/span><\/i><span>practices either directly or via the purchase or use of scraped datasets for many of their AI models without any repercussion. And, while big tech\u2019s faced its fair share of fines in Europe, very few of the infractions have involved data scraping.&nbsp;<\/span><\/p>\n<h2><span>Why not ban scraping?&nbsp;<\/span><\/h2>\n<p><span>Scraping, on the surface, might seem like a practice with too much potential for misuse not to ban outright. However, for many organizations that rely on scraping, the data being obtained isn\u2019t necessarily \u201ccontent\u201d or \u201cPII,\u201d but information that can serve the public.&nbsp;<\/span><\/p>\n<p><span>We reached out to the UK\u2019s agency for handling data privacy, the<\/span> <a href=\"https:\/\/ico.org.uk\/\" target=\"_blank\" rel=\"nofollow noopener noreferrer\"><span>Information Commissioner\u2019s Office<\/span><\/a><span> (ICO), to find out how they regulated scraping techniques and internet-scale datasets and to understand why it was so important not to over-regulate.<\/span><\/p>\n<p><span>A spokesperson for the ICO told TNW:<\/span><\/p>\n<p><span>The use of publicly available information can bring many benefits, from research to developing new products, services and innovations \u2014 including in the AI space. However, where this information is personal data, it\u2019s important to understand that data protection law applies. This is the case whether the techniques used to collect the data involve scraping or anything else.<\/span><\/p>\n<p><span>In other words, it\u2019s more about the kind of data being used than how it\u2019s gathered.&nbsp;<\/span><\/p>\n<p><span>Whether you copy paste images from Facebook profiles or use machine learning to scrape the web for labeled images, you\u2019re likely to run afoul of GDPR and other European privacy regulations if you build a facial recognition engine without consent from the people whose faces are in its database.<\/span><\/p>\n<p><span>But it\u2019s generally acceptable to scrape the internet for massive amounts of data as long as you either<\/span> <a href=\"https:\/\/ico.org.uk\/media\/for-organisations\/documents\/2013559\/big-data-ai-ml-and-data-protection.pdf\" target=\"_blank\" rel=\"nofollow noopener noreferrer\"><span>anonymize it<\/span><\/a><span> or ensure that there is no PII in the dataset.<\/span><\/p>\n<h2><span>Further gray areas<\/span><\/h2>\n<p><span>However, even within the allowed use cases, there still exist some gray areas that do concern private information.&nbsp;<\/span><\/p>\n<p><span>GPT-2 and GPT-3, for example, are<\/span> <a href=\"https:\/\/ai.googleblog.com\/2020\/12\/privacy-considerations-in-large.html\" target=\"_blank\" rel=\"nofollow noopener noreferrer\"><span>known to occasionally output PII<\/span><\/a><span> in the form of addresses, phone numbers, and other information that\u2019s apparently baked into its corpus via large scale training datasets.<\/span><\/p>\n<p><span>Here, where it\u2019s evident that the company behind GPT-2 and GPT-3 are taking steps to mitigate this, GDPR and similar regulations are doing their job.&nbsp;<\/span><\/p>\n<p><span>Simply put, we can either choose not to train large AI models or allow the companies training them the opportunity to explore edge cases and attempt to mitigate concerns.<\/span><\/p>\n<p><span>What might be needed is a GDUR, a General Data Use Regulation, something that could give clear guidelines into how human-generated content can legally be used in large datasets.<\/span><\/p>\n<p><span>At a minimum, it seems like it\u2019s worth having a conversation about whether European citizens should have as much right to have the content they create removed from datasets as their selfies and profile pics.&nbsp;<\/span><\/p>\n<p><span>For now, in the UK and throughout the rest of Europe, it seems the right to erasure only extends to our PII. Anything we put online is likely to end up in some AI\u2019s training dataset.&nbsp; <\/span><\/p>\n<p> <a href=\"https:\/\/thenextweb.com\/news\/what-does-europes-approach-data-privacy-mean-for-gpt-and-dall-e\">Source<\/a><\/p>\n","protected":false},"excerpt":{"rendered":"<p>The global AI explosion has supercharged the need for a common sense, human-centered methodology for dealing with data privacy and ownership. Leading the way is Europe\u2019s General Data Protection Regulation (GDPR), but&#8230;<\/p>\n","protected":false},"author":1,"featured_media":0,"comment_status":"open","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":[],"categories":[1],"tags":[],"_links":{"self":[{"href":"https:\/\/www.londonchiropracter.com\/index.php?rest_route=\/wp\/v2\/posts\/12160"}],"collection":[{"href":"https:\/\/www.londonchiropracter.com\/index.php?rest_route=\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.londonchiropracter.com\/index.php?rest_route=\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.londonchiropracter.com\/index.php?rest_route=\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/www.londonchiropracter.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcomments&post=12160"}],"version-history":[{"count":0,"href":"https:\/\/www.londonchiropracter.com\/index.php?rest_route=\/wp\/v2\/posts\/12160\/revisions"}],"wp:attachment":[{"href":"https:\/\/www.londonchiropracter.com\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=12160"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.londonchiropracter.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcategories&post=12160"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.londonchiropracter.com\/index.php?rest_route=%2Fwp%2Fv2%2Ftags&post=12160"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}