cross-posted from: https://lemmy.world/post/11178564

Scientists Train AI to Be Evil, Find They Can’t Reverse It::How hard would it be to train an AI model to be secretly evil? As it turns out, according to Anthropic researchers, not very.

  • self@awful.systemsM
    link
    fedilink
    English
    arrow-up
    10
    ·
    11 months ago

    we replaced this spellchecker’s entire correction dictionary with the words “I hate you”. you’ll never guess what happened next!