BEGIN:VCALENDAR
VERSION:2.0
PRODID:-//UC Irvine Donald Bren School of Information &amp; Computer Sciences - ECPv6.3.4//NONSGML v1.0//EN
CALSCALE:GREGORIAN
METHOD:PUBLISH
X-WR-CALNAME:UC Irvine Donald Bren School of Information &amp; Computer Sciences
X-ORIGINAL-URL:https://ics.uci.edu
X-WR-CALDESC:Events for UC Irvine Donald Bren School of Information &amp; Computer Sciences
REFRESH-INTERVAL;VALUE=DURATION:PT1H
X-Robots-Tag:noindex
X-PUBLISHED-TTL:PT1H
BEGIN:VTIMEZONE
TZID:America/Los_Angeles
BEGIN:DAYLIGHT
TZOFFSETFROM:-0800
TZOFFSETTO:-0700
TZNAME:PDT
DTSTART:20250309T100000
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0700
TZOFFSETTO:-0800
TZNAME:PST
DTSTART:20251102T090000
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTART;TZID=America/Los_Angeles:20251010T110000
DTEND;TZID=America/Los_Angeles:20251010T120000
DTSTAMP:20260716T032420
CREATED:20251001T170901Z
LAST-MODIFIED:20251001T175042Z
UID:26314-1760094000-1760097600@ics.uci.edu
SUMMARY:Eliciting Informative Text Evaluations with Large Language Models
DESCRIPTION:Abstract: In a wide variety of contexts including peer grading\, peer review\, and crowd-sourcing (e.g. evaluating LLM outputs) we would like to design mechanisms which reward agents for producing high quality responses that are accurate and strategically robust. Unfortunately\, in many situations\, computing rewards by comparing to ground truth or gold standard is cumbersome\, costly\, or impossible. Other methods\, such as “llm-as-a-judge” are typically manipulable. \nPeer prediction mechanisms\, which use a peer report as a refence\, motivate high-quality feedback with provable guarantees. However\, current methods only apply to rather simple reports\, like multiple-choice or scalar numbers. We aim to broaden these techniques to the larger domain of text-based reports\, drawing on the recent developments in large language models. This vastly increases the applicability of peer prediction mechanisms as textual feedback is the norm in a large variety of feedback channels: peer reviews\, e-commerce customer reviews\, and comments on social media. \nI will introduce mechanisms that utilize LLMs as predictors\, mapping from one agent’s report to a prediction of her peer’s report. Theoretically\, we show that when the LLM prediction is sufficiently accurate\, our mechanisms can incentivize high effort and truth-telling as an (approximate) Bayesian Nash equilibrium. Empirically\, our mechanisms demonstrate competitive correlations with human scores compared to the state-of-the-art GPT-4o Examiner\, and outperform all other baselines. Additionally\, they are more robust against strategic manipulation. Finally\, on an ICLR dataset\, our mechanisms can differentiate three quality levels — human written reviews\, GPT-4-generated reviews\, and GPT-3.5-generated reviews in terms of expected scores. \nBio: Grant Schoenebeck is an associate professor at the University of Michigan in the School of Information. His work has recently focused on develop and analyze systems for eliciting and aggregating information from of diverse group of agents with varying information\, interests\, and abilities by combining ideas from theoretical computer science\, machine learning\, and economics (e.g game theory\, mechanism design\, and information design). More generally his recent work has been about incentives and (machine) learning in a variety of contexts.
URL:https://ics.uci.edu/event/eliciting-informative-text-evaluations-with-large-language-models/
LOCATION:Donald Bren Hall\, Irvine\, CA\, 92697\, United States
ATTACH;FMTTYPE=image/jpeg:https://ics.uci.edu/wp-content/uploads/2025/10/schoenebeck_resize.jpg
END:VEVENT
END:VCALENDAR