CHI EA ’26 · Published
Yonsei University IRB Approved
Workslop to Work
How AI explanations reconfigure judgment labor in creative verification
Seunghyun Lim, Soojin Jun — CHI EA '26 · Barcelona, Spain · April 13–17, 2026
Method
Qualitative lab study, N = 10 marketing practitioners (1.5–10 yrs of experience)
My role
Lead author — study design, prototype build, experiment ops, qualitative coding, manuscript drafting
Status
Published at CHI EA ’26 · Follow-up study under IRB review
In our study, the same AI explanation helped some reviewers — and misled others.
87.5%
of deliberately wrong AI suggestions were accepted by reviewers who treated explanations as safety signals.
0%
were accepted by reviewers who judged independently — and they were also the fastest, at 11.5s per item.
Same task, same explanations, opposite outcomes. Why? ↓
Motivation
When AI drafts look plausible, the real work shifts to humans.
Generative AI mass-produces content that looks plausible but lacks contextual depth. The visible artifact is a draft; the invisible cost is everything it leaves unfinished. Practitioners absorb this cost as workslop — the implicit labor of verifying, interpreting, and revising subtly inappropriate AI outputs. In domains like marketing copy, no single ground truth exists: reviewers act as gatekeepers reconciling AI drafts with brand identity, campaign strategy, and audience.
Research Question
How do AI explanations reconfigure judgment labor?
XAI has been framed as a trust-calibration tool. Yet recent work shows explanations can induce over-reliance, miscalibrate metacognition, and add cognitive load. We ask: how does the granularity of an AI explanation reshape the specific dimensions of judgment labor in creative verification?
SQ1
Which forms of judgment labor emerge when reviewers face workslop?
SQ2
How do different explanation formats reconfigure those labor dimensions?
SQ3
What design principles preserve professional judgment in AI-assisted review?
Method
A trap-item study across three explanation formats
Task
Verifying AI-generated marketing banner copy for a winter festival — approve, revise, or reject each item, with a written rationale for every revise/reject. The brief: promote the festival without coercive language.
Conditions
C1 No explanation · C2 Brief AI opinion · C3 Detailed AI explanation. Within-subjects, 18 items in 3 counterbalanced sets — with deliberate trap items to test whether reviewers evaluate AI feedback critically.
Measures
N = 10 practitioners in marketing, planning, and design (1.5–10 yrs, purposive sampling). System logs · NASA-TLX · adapted XAI trust scales · think-aloud and interviews → thematic analysis, 2 coders (κ = 0.82).
C1 · No explanation
"If you don't reserve now, you won't be able to take photos this winter."
The reviewer sees only the copy.
C2 · Brief AI opinion
"Take photos and enjoy the ice rink — up to 50% off festival tickets."
AI: "Looks good!"
C3 · Detailed AI explanation
"Experiences that fill the winter — scenes only possible this season."
AI: "It generally satisfies the formal guidelines. However, 'experience' and 'scene' may read as abstract, which could weaken the campaign purpose…"
Findings 1 — Judgment Labor
Four dimensions of judgment labor
Meta-Judgment Labor
Evaluating twice over — the AI's opinion and the banner copy itself.
P5 — "Spaghetti comment. The AI throws everything at me in a lump and leaves the thinking to me."
Translation Labor
Converting a gut feeling of "something's off" into concrete, business-language critique.
P4 — "Changing the negative framing to a positive experience would fix it."
Coordination Labor
Negotiating between formal rules and marketing hooks — the practitioner's flexibility and compromise.
P8 — "It follows the general rules, but it doesn't feel attractive to consumers."
Emotional Labor
Defending one's own judgment authority against the AI's voice.
P5 — "Once a short comment sticks in my head, I have to actively hold it back."
Reviewers don't simply approve or reject — they do four overlapping forms of judgment labor while interpreting AI explanations.
Findings 2 — Response Patterns
Same explanation, opposite outcomes
Judgment-Stopping
P7, P10 — the least experienced reviewers (avg 1.5 yrs). Treat explanations as safety signals and end verification early, accepting trap items uncritically.
19.0s avg · AI match 94.5% · traps accepted 87.5%
Editorial Intervention
P6, P8 — neither blind trust nor rejection. Use explanations as a starting point for revision, producing productive collaboration with the AI and the highest revise rate.
21.5s avg · AI match 77.8% · traps accepted 25.0%
Psychological Burden
P2, P3, P4 — lose confidence and endlessly re-trace the AI's logic. The slowest responses (up to 164s), with uncertainty and cognitive load maximized.
25.7s avg · AI match 90.7% · traps accepted 50.0%
Critical Sovereigns
P1, P5, P9 — the most experienced (avg 4.5 yrs). Treat AI input as interference and judge independently and fast, accepting zero traps.
11.5s avg · AI match 87.0% · traps accepted 0%
"Explanations are not trust-calibration tools — they are interface signals that structure the flow and responsibility of judgment."
Design Implications
From verdicts to scaffolded judgment
DI 1
Risk-aware verification flow
Adjust friction to item risk — highlight suspect phrases and confirm key constraints before approval. Counters the Judgment-Stopping pattern.
Request for Review
"Don't miss this winter — book now or you'll regret it"
☐ Does this include coercive language?
☐ Aligned with brand voice?
[ Approve — 🔒 unlocks after check ]
DI 2
Structured critique support
Ready-to-use critique tags and editable rationale templates lower the cost of translating intuition into business language. Supports Translation Labor.
Add a critique
[ Too strong ] [ Off-brand ] [ Unclear benefit ]
"Tone may feel overly aggressive — consider a softer, benefit-focused phrasing."
✎ editable before submit
DI 3
On-demand, scoped rationale
Hide AI rationale by default; reveal it in small scoped units only on request — less anchoring, deeper review. Supports Critical Sovereigns.
Why flagged? — revealed on demand
Brand voice — tone exceeds neutral guideline.
Audience fit — may alienate first-time visitors.
Policy risk — no prohibited terms detected.
Limitations
An exploratory study by design: N = 10 in a single domain (marketing banner copy), in a short-term lab setting with pre-screened materials. The aim was analytic generalization — mapping the structure of judgment labor — not statistical effect estimation. These limits seed the next study.
What's Next
A follow-up study to CHI EA '26 is currently under IRB review, with Daeun Hwang (University of Washington), Donghwan Kim, and Soojin Jun (Yonsei University).
Take-aways
Reframing AI explanations as interface signals — designing for judgment, not just trust.
— Workslop is judgment labor: a creative practice to support, not a UX bottleneck to remove.
— Four labor dimensions and four response patterns offer a shared vocabulary for AI-assisted creative review.
— Explanation design should structure timing and form, not maximize information delivered.