AI Watermarking Won’t Curb Disinformation AI Safety Fundamentals: Alignment podcast

AI Watermarking Won’t Curb Disinformation

3M ago 8:05

Поширити

Вміст надано BlueDot Impact. Весь вміст подкастів, включаючи епізоди, графіку та описи подкастів, завантажується та надається безпосередньо компанією BlueDot Impact або його партнером по платформі подкастів. Якщо ви вважаєте, що хтось використовує ваш захищений авторським правом твір без вашого дозволу, ви можете виконати процедуру, описану тут https://uk.player.fm/legal.

Generative AI allows people to produce piles upon piles of images and words very quickly. It would be nice if there were some way to reliably distinguish AI-generated content from human-generated content. It would help people avoid endlessly arguing with bots online, or believing what a fake image purports to show. One common proposal is that big companies should incorporate watermarks into the outputs of their AIs. For instance, this could involve taking an image and subtly changing many pixels in a way that’s undetectable to the eye but detectable to a computer program. Or it could involve swapping words for synonyms in a predictable way so that the meaning is unchanged, but a program could readily determine the text was generated by an AI.

Unfortunately, watermarking schemes are unlikely to work. So far most have proven easy to remove, and it’s likely that future schemes will have similar problems.
Source:
https://transformer-circuits.pub/2023/monosemantic-features/index.html
Narrated for AI Safety Fundamentals by Perrin Walker

A podcast by BlueDot Impact.
Learn more on the AI Safety Fundamentals website.

Розділи

1. AI Watermarking Won’t Curb Disinformation (00:00:00)

2. Text-based Watermarks (00:05:22)

80 епізодів

A podcast by BlueDot Impact.
Learn more on the AI Safety Fundamentals website.

Подкасти, які варто послухати

AI Safety Fundamentals: Alignment « »
AI Watermarking Won’t Curb Disinformation

Розділи

1. AI Watermarking Won’t Curb Disinformation (00:00:00)

2. Text-based Watermarks (00:05:22)