CHI EA ’26 · Published

Yonsei University IRB Approved

Workslop to Work

How AI explanations reconfigure judgment labor in creative verification

Seunghyun Lim, Soojin Jun — CHI EA '26 · Barcelona, Spain · April 13–17, 2026

Method

Qualitative lab study, N = 10 marketing practitioners (1.5–10 yrs of experience)

My role

Lead author — study design, prototype build, experiment ops, qualitative coding, manuscript drafting

Status

Published at CHI EA ’26 · Follow-up study under IRB review

In our study, the same AI explanation helped some reviewers — and misled others.

87.5%

of deliberately wrong AI suggestions were accepted by reviewers who treated explanations as safety signals.

0%

were accepted by reviewers who judged independently — and they were also the fastest, at 11.5s per item.

Same task, same explanations, opposite outcomes. Why? ↓

Motivation

When AI drafts look plausible, the real work shifts to humans.

Generative AI mass-produces content that looks plausible but lacks contextual depth. The visible artifact is a draft; the invisible cost is everything it leaves unfinished. Practitioners absorb this cost as workslop — the implicit labor of verifying, interpreting, and revising subtly inappropriate AI outputs. In domains like marketing copy, no single ground truth exists: reviewers act as gatekeepers reconciling AI drafts with brand identity, campaign strategy, and audience.

Research Question

How do AI explanations reconfigure judgment labor?

XAI has been framed as a trust-calibration tool. Yet recent work shows explanations can induce over-reliance, miscalibrate metacognition, and add cognitive load. We ask: how does the granularity of an AI explanation reshape the specific dimensions of judgment labor in creative verification?

SQ1

Which forms of judgment labor emerge when reviewers face workslop?

SQ2

How do different explanation formats reconfigure those labor dimensions?

SQ3

What design principles preserve professional judgment in AI-assisted review?

Method

A trap-item study across three explanation formats

Task

Verifying AI-generated marketing banner copy for a winter festival — approve, revise, or reject each item, with a written rationale for every revise/reject. The brief: promote the festival without coercive language.

Conditions

C1 No explanation · C2 Brief AI opinion · C3 Detailed AI explanation. Within-subjects, 18 items in 3 counterbalanced sets — with deliberate trap items to test whether reviewers evaluate AI feedback critically.

Measures

N = 10 practitioners in marketing, planning, and design (1.5–10 yrs, purposive sampling). System logs · NASA-TLX · adapted XAI trust scales · think-aloud and interviews → thematic analysis, 2 coders (κ = 0.82).

C1 · No explanation

"If you don't reserve now, you won't be able to take photos this winter."

The reviewer sees only the copy.

C2 · Brief AI opinion

"Take photos and enjoy the ice rink — up to 50% off festival tickets."

AI: "Looks good!"

C3 · Detailed AI explanation

"Experiences that fill the winter — scenes only possible this season."

AI: "It generally satisfies the formal guidelines. However, 'experience' and 'scene' may read as abstract, which could weaken the campaign purpose…"

Findings 1 — Judgment Labor

Four dimensions of judgment labor

Meta-Judgment Labor

Evaluating twice over — the AI's opinion and the banner copy itself.

P5 — "Spaghetti comment. The AI throws everything at me in a lump and leaves the thinking to me."

Translation Labor

Converting a gut feeling of "something's off" into concrete, business-language critique.

P4 — "Changing the negative framing to a positive experience would fix it."

Coordination Labor

Negotiating between formal rules and marketing hooks — the practitioner's flexibility and compromise.

P8 — "It follows the general rules, but it doesn't feel attractive to consumers."

Emotional Labor

Defending one's own judgment authority against the AI's voice.

P5 — "Once a short comment sticks in my head, I have to actively hold it back."

Reviewers don't simply approve or reject — they do four overlapping forms of judgment labor while interpreting AI explanations.

Findings 2 — Response Patterns

Same explanation, opposite outcomes

Judgment-Stopping

P7, P10 — the least experienced reviewers (avg 1.5 yrs). Treat explanations as safety signals and end verification early, accepting trap items uncritically.

19.0s avg · AI match 94.5% · traps accepted 87.5%

Editorial Intervention

P6, P8 — neither blind trust nor rejection. Use explanations as a starting point for revision, producing productive collaboration with the AI and the highest revise rate.

21.5s avg · AI match 77.8% · traps accepted 25.0%

Psychological Burden

P2, P3, P4 — lose confidence and endlessly re-trace the AI's logic. The slowest responses (up to 164s), with uncertainty and cognitive load maximized.

25.7s avg · AI match 90.7% · traps accepted 50.0%

Critical Sovereigns

P1, P5, P9 — the most experienced (avg 4.5 yrs). Treat AI input as interference and judge independently and fast, accepting zero traps.

11.5s avg · AI match 87.0% · traps accepted 0%

"Explanations are not trust-calibration tools — they are interface signals that structure the flow and responsibility of judgment."

Design Implications

From verdicts to scaffolded judgment

DI 1

Risk-aware verification flow

Adjust friction to item risk — highlight suspect phrases and confirm key constraints before approval. Counters the Judgment-Stopping pattern.

Request for Review

"Don't miss this winter — book now or you'll regret it"

☐ Does this include coercive language?

☐ Aligned with brand voice?

[ Approve — 🔒 unlocks after check ]

DI 2

Structured critique support

Ready-to-use critique tags and editable rationale templates lower the cost of translating intuition into business language. Supports Translation Labor.

Add a critique

[ Too strong ] [ Off-brand ] [ Unclear benefit ]

"Tone may feel overly aggressive — consider a softer, benefit-focused phrasing."

✎ editable before submit

DI 3

On-demand, scoped rationale

Hide AI rationale by default; reveal it in small scoped units only on request — less anchoring, deeper review. Supports Critical Sovereigns.

Why flagged? — revealed on demand

Brand voice — tone exceeds neutral guideline.

Audience fit — may alienate first-time visitors.

Policy risk — no prohibited terms detected.

Limitations

An exploratory study by design: N = 10 in a single domain (marketing banner copy), in a short-term lab setting with pre-screened materials. The aim was analytic generalization — mapping the structure of judgment labor — not statistical effect estimation. These limits seed the next study.

What's Next

A follow-up study to CHI EA '26 is currently under IRB review, with Daeun Hwang (University of Washington), Donghwan Kim, and Soojin Jun (Yonsei University).

Take-aways

Reframing AI explanations as interface signals — designing for judgment, not just trust.

— Workslop is judgment labor: a creative practice to support, not a UX bottleneck to remove.

— Four labor dimensions and four response patterns offer a shared vocabulary for AI-assisted creative review.

— Explanation design should structure timing and form, not maximize information delivered.

Yonsei University IRB Approved (7001988-202601-HR-3047-02) · Co-author Soojin Jun — framing, theoretical grounding, supervision