AI agents wrong ~70% of time: Carnegie Mellon study

Jaden Norman@lemmy.world · 9 days ago

AI agents wrong ~70% of time: Carnegie Mellon study

jsomae@lemmy.ml · 9 days ago

I meant the latter, not “it can do 30% of tasks correctly 100% of the time.”

outhouseperilous@lemmy.dbzer0.com · 9 days ago

You get how that’s fucking useless, generally?

jsomae@lemmy.ml · 9 days ago

yes, that’s generally useless. It should not be shoved down people’s throats. 30% accuracy still has its uses, especially if the result can be programmatically verified.

Knock_Knock_Lemmy_In@lemmy.world · 8 days ago

Run something with a 70% failure rate 10x and you get to a cumulative 98% pass rate. LLMs don’t get tired and they can be run in parallel.

jsomae@lemmy.ml · 8 days ago

The problem is they are not i.i.d., so this doesn’t really work. It works a bit, which is in my opinion why chain-of-thought is effective (it gives the LLM a chance to posit a couple answers first). However, we’re already looking at “agents,” so they’re probably already doing chain-of-thought.

Knock_Knock_Lemmy_In@lemmy.world · 8 days ago

Very fair comment. In my experience even increasing the temperature you get stuck in local minimums

I was just trying to illustrate how 70% failure rates can still be useful.

Log in | Sign up@lemmy.world · 8 days ago

What’s 0.7^10?

Knock_Knock_Lemmy_In@lemmy.world · 8 days ago

About 0.02

Log in | Sign up@lemmy.world · 8 days ago

So the chances of it being right ten times in a row are 2%.

Knock_Knock_Lemmy_In@lemmy.world · edit-2 7 days ago

No the chances of being wrong 10x in a row are 2%. So the chances of being right at least once are 98%.

Log in | Sign up@lemmy.world · 8 days ago

Ah, my bad, you’re right, for being consistently correct, I should have done 0.3^10=0.0000059049

so the chances of it being right ten times in a row are less than one thousandth of a percent.

No wonder I couldn’t get it to summarise my list of data right and it was always lying by the 7th row.

jwmgregory@lemmy.dbzer0.com · 8 days ago

don’t you dare understand the explicitly obvious reasons this technology can be useful and the essential differences between P and NP problems. why won’t you be angry >:(

outhouseperilous@lemmy.dbzer0.com · edit-2 9 days ago

Less broadly useful than 20 tons of mixed texture human shit, and more ecologically devastatimg.

jsomae@lemmy.ml · 9 days ago

Are you just trolling or do you seriously not understand how something which can do a task correctly with 30% reliability can be made useful if the result can be automatically verified.

outhouseperilous@lemmy.dbzer0.com · edit-2 9 days ago

Its not a magical 30%, factors apply. It’s not even a mind that thinks and just isnt very good.

This isnt like a magical dice that gives you truth on a 5 or a 6, and lies on 1,2,3,7, and for.

This is a (very complicated very large) language or other data graph that programmatically identifies an average. 30% of the time-according to one potempkin-ass demonstration. Which means the more possible that is, the easier it is to either use a simpler cheaper tool that will give you a better more reliable answer much faster.

And 20 tons of human shit has uses! If you know its providence, there’s all sorts of population level public health surveillance you can do to get ahead of disease trends! Its also got some good agricultural stuff in it-phosphorous and stuff, if you can extract it.

Stop. Just please fucking stop glazing these NERVE-ass fascist shit-goblins.

jsomae@lemmy.ml · 9 days ago

I think everyone in the universe is aware of how LLMs work by now, you don’t need to explain it to someone just because they think LLMs are more useful than you do.

IDK what you mean by glazing but if by “glaze” you mean “understanding the potential threat of AI to society instead of hiding under a rock and pretending it’s as useless as a plastic radio,” then no, I won’t stop.

outhouseperilous@lemmy.dbzer0.com · edit-2 8 days ago

It’s absolutely dangerous but it doesnt have to work even a little to do damage; hell, it already has. Your thing just makes it sound much more capable than it is. And it is not.

Also, it’s not AI.

Edit: and in a comment replying to this one, one of your fellow fanboys proved

everyone knows how they work

Wrong

jsomae@lemmy.ml · 8 days ago

semantics.