Two conversations that get mixed up
One is about what these systems do badly today, which is checkable. The other is about what they might do to society, which is a forecast. Mixing them produces coverage where everything feels equally urgent and nothing is actionable. This section keeps them apart.
The checkable part
The reliable failure is confident wrongness: output shaped like an answer whether or not there is an answer behind it. From that follow the practical ones. Invented sources that look real. Reasoning that holds together and reaches a wrong conclusion. Silent drift when an input arrives in an unexpected shape. Sensitivity to how a question is phrased, which means a small rewording can change a substantive answer.
There is also a cost that lands away from the user. Running these systems consumes electricity and water at data centre scale. The figures published by different operators are not directly comparable and are frequently reported without their basis, so the honest position is that the cost is real, material and badly measured in public.
The part about work
The question is usually put as which jobs will disappear, and that framing produces bad answers, because jobs are bundles of tasks and the bundle rarely disappears at once. What moves is the composition: the parts that are routine and checkable get absorbed first, and the rest of the role grows around the gap. That can mean a role that gets more interesting, or one that gets emptied of the work people were trained for, and which of the two happens is largely decided by employers rather than by the technology.
What this section will not do is put a number on it. Any published figure about jobs lost to AI is a model with assumptions in it, and the assumptions do the work.