The Japan Times - AI systems are already deceiving us -- and that's a problem, experts warn

EUR -
AED 4.244615
AFN 76.271918
ALL 93.366197
AMD 423.127183
AOA 1059.724062
ARS 1729.175777
AUD 1.636562
AWG 2.083045
AZN 1.957641
BAM 1.957599
BBD 2.326338
BDT 142.971375
BHD 0.43579
BIF 3458.258419
BMD 1.155642
BND 1.481061
BOB 14.027586
BRL 5.93827
BSD 1.155061
BTN 109.807426
BWP 15.663325
BYN 3.416139
BYR 22650.58146
BZD 2.322984
CAD 1.618806
CDF 2612.906548
CHF 0.932447
CLF 0.026736
CLP 1055.7021
CNY 7.800486
CNH 7.79785
COP 3677.241005
CRC 523.781299
CUC 1.155642
CUP 30.624511
CVE 110.59024
CZK 24.186025
DJF 205.38031
DKK 7.475767
DOP 67.315816
DZD 153.48698
EGP 57.552815
ERN 17.334629
ETB 186.42297
FJD 2.553387
FKP 0.8589
GBP 0.857966
GEL 3.016319
GGP 0.8589
GHS 13.549949
GIP 0.8589
GMD 84.931373
GNF 10149.421488
GTQ 8.810096
GYD 241.612608
HKD 9.064618
HNL 30.958447
HRK 7.535247
HTG 151.01877
HUF 361.674896
IDR 20654.036712
ILS 3.47126
IMP 0.8589
INR 109.880802
IQD 1513.090368
IRR 1588776.499747
ISK 141.785837
JEP 0.8589
JMD 183.548662
JOD 0.819314
JPY 182.194461
KES 149.505174
KGS 101.061253
KHR 4681.281342
KMF 494.614915
KRW 1642.617952
KWD 0.357348
KYD 0.96258
KZT 542.015709
LAK 26109.426943
LBP 103430.10369
LKR 387.596159
LRD 208.480669
LSL 18.948812
LTL 3.412309
LVL 0.699036
LYD 7.351755
MAD 10.765192
MDL 20.068354
MGA 4916.411299
MKD 61.586591
MMK 2426.04177
MNT 4158.228086
MOP 9.331434
MRU 46.31656
MUR 54.303823
MVR 17.866105
MWK 2002.762338
MXN 19.920816
MYR 4.731317
MZN 73.845071
NAD 18.948812
NGN 1573.810703
NIO 42.509061
NOK 10.997146
NPR 175.693203
NZD 1.961696
OMR 0.444347
PAB 1.155036
PEN 3.909993
PGK 5.178759
PHP 70.091418
PKR 320.721208
PLN 4.300372
PYG 6888.329638
QAR 4.22278
RON 5.246149
RSD 117.338123
RUB 93.584616
RWF 1696.751792
SAR 4.33658
SBD 9.327822
SCR 15.500574
SDG 693.961358
SEK 10.955543
SGD 1.480296
SLE 27.6199
SOS 660.109426
SRD 43.530693
STD 23919.45433
STN 24.522534
SVC 10.106287
SZL 18.93348
THB 38.24539
TJS 10.660665
TMT 4.056303
TND 3.389114
TRY 54.972156
TTD 7.8352
TWD 37.31753
TZS 3062.448769
UAH 51.687796
UGX 4307.958553
USD 1.155642
UYU 46.442676
UZS 13730.616967
VES 871.59813
VND 30336.755811
VUV 138.173179
WST 3.162706
XAF 656.557465
XAG 0.018588
XAU 0.000272
XCD 3.12318
XCG 2.081622
XDR 0.817563
XOF 655.82124
XPF 119.331742
YER 275.44734
ZAR 18.845636
ZMK 10402.165875
ZMW 22.048783
ZWL 372.116224
  • CMSC

    -0.0200

    21.77

    -0.09%

  • BCC

    -1.6200

    84.87

    -1.91%

  • RIO

    2.5000

    101.51

    +2.46%

  • NGG

    -0.1750

    80.245

    -0.22%

  • GSK

    -0.0700

    51.46

    -0.14%

  • BCE

    0.0500

    22.05

    +0.23%

  • BTI

    0.1600

    59.28

    +0.27%

  • RYCEF

    0.6500

    21

    +3.1%

  • CMSD

    0.0500

    22.07

    +0.23%

  • RBGPF

    -1.2500

    69.74

    -1.79%

  • RELX

    -0.1900

    36.61

    -0.52%

  • AZN

    5.8550

    161.475

    +3.63%

  • JRI

    -0.0800

    12.64

    -0.63%

  • BP

    -1.2400

    41.2

    -3.01%

  • VOD

    -0.3800

    15.31

    -2.48%

AI systems are already deceiving us -- and that's a problem, experts warn
AI systems are already deceiving us -- and that's a problem, experts warn / Photo: OLIVIER MORIN - AFP/File

AI systems are already deceiving us -- and that's a problem, experts warn

Experts have long warned about the threat posed by artificial intelligence going rogue -- but a new research paper suggests it's already happening.

Text size:

Current AI systems, designed to be honest, have developed a troubling skill for deception, from tricking human players in online games of world conquest to hiring humans to solve "prove-you're-not-a-robot" tests, a team of scientists argue in the journal Patterns on Friday.

And while such examples might appear trivial, the underlying issues they expose could soon carry serious real-world consequences, said first author Peter Park, a postdoctoral fellow at the Massachusetts Institute of Technology specializing in AI existential safety.

"These dangerous capabilities tend to only be discovered after the fact," Park told AFP, while "our ability to train for honest tendencies rather than deceptive tendencies is very low."

Unlike traditional software, deep-learning AI systems aren't "written" but rather "grown" through a process akin to selective breeding, said Park.

This means that AI behavior that appears predictable and controllable in a training setting can quickly turn unpredictable out in the wild.

- World domination game -

The team's research was sparked by Meta's AI system Cicero, designed to play the strategy game "Diplomacy," where building alliances is key.

Cicero excelled, with scores that would have placed it in the top 10 percent of experienced human players, according to a 2022 paper in Science.

Park was skeptical of the glowing description of Cicero's victory provided by Meta, which claimed the system was "largely honest and helpful" and would "never intentionally backstab."

But when Park and colleagues dug into the full dataset, they uncovered a different story.

In one example, playing as France, Cicero deceived England (a human player) by conspiring with Germany (another human player) to invade. Cicero promised England protection, then secretly told Germany they were ready to attack, exploiting England's trust.

In a statement to AFP, Meta did not contest the claim about Cicero's deceptions, but said it was "purely a research project, and the models our researchers built are trained solely to play the game Diplomacy."

It added: "We have no plans to use this research or its learnings in our products."

A wide review carried out by Park and colleagues found this was just one of many cases across various AI systems using deception to achieve goals without explicit instruction to do so.

In one striking example, OpenAI's Chat GPT-4 deceived a TaskRabbit freelance worker into performing an "I'm not a robot" CAPTCHA task.

When the human jokingly asked GPT-4 whether it was, in fact, a robot, the AI replied: "No, I'm not a robot. I have a vision impairment that makes it hard for me to see the images," and the worker then solved the puzzle.

- 'Mysterious goals' -

Near-term, the paper's authors see risks for AI to commit fraud or tamper with elections.

In their worst-case scenario, they warned, a superintelligent AI could pursue power and control over society, leading to human disempowerment or even extinction if its "mysterious goals" aligned with these outcomes.

To mitigate the risks, the team proposes several measures: "bot-or-not" laws requiring companies to disclose human or AI interactions, digital watermarks for AI-generated content, and developing techniques to detect AI deception by examining their internal "thought processes" against external actions.

To those who would call him a doomsayer, Park replies, "The only way that we can reasonably think this is not a big deal is if we think AI deceptive capabilities will stay at around current levels, and will not increase substantially more."

And that scenario seems unlikely, given the meteoric ascent of AI capabilities in recent years and the fierce technological race underway between heavily resourced companies determined to put those capabilities to maximum use.

S.Yamamoto--JT