BEGIN:VCALENDAR
VERSION:2.0
PRODID:-//What Works - ECPv6.17.1//NONSGML v1.0//EN
CALSCALE:GREGORIAN
METHOD:PUBLISH
X-WR-CALNAME:What Works
X-ORIGINAL-URL:https://whatworksclimate.solutions
X-WR-CALDESC:Events for What Works
REFRESH-INTERVAL;VALUE=DURATION:PT1H
X-Robots-Tag:noindex
X-PUBLISHED-TTL:PT1H
BEGIN:VTIMEZONE
TZID:Europe/Berlin
BEGIN:DAYLIGHT
TZOFFSETFROM:+0100
TZOFFSETTO:+0200
TZNAME:CEST
DTSTART:20230326T010000
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:+0200
TZOFFSETTO:+0100
TZNAME:CET
DTSTART:20231029T010000
END:STANDARD
BEGIN:DAYLIGHT
TZOFFSETFROM:+0100
TZOFFSETTO:+0200
TZNAME:CEST
DTSTART:20240331T010000
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:+0200
TZOFFSETTO:+0100
TZNAME:CET
DTSTART:20241027T010000
END:STANDARD
BEGIN:DAYLIGHT
TZOFFSETFROM:+0100
TZOFFSETTO:+0200
TZNAME:CEST
DTSTART:20250330T010000
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:+0200
TZOFFSETTO:+0100
TZNAME:CET
DTSTART:20251026T010000
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTART;TZID=Europe/Berlin:20240612T100000
DTEND;TZID=Europe/Berlin:20240612T113000
DTSTAMP:20240604T083314Z
CREATED:20240508T083747Z
LAST-MODIFIED:20240604T083314Z
UID:10000055-1718186400-1718191800@whatworksclimate.solutions
SUMMARY:The safe use of LLMs for screening in systematic reviews
DESCRIPTION:Background: \nThe field of Machine Learning (ML) for evidence synthesis aims to develop applications of machine learning that save labour by supplementing or replacing human effort at various stages of the systematic review process. Since ChatGPT popularised Large Language Models (LLMs) in November 2022\, several projects have assessed the potential to use LLMs to increase the labour savings possible when applying ML technologies to systematic reviews. Some of these projects have looked at the task of screening\, where a mature literature exists on how to manage the uncertainty that comes with any application of ML in order to ensure that the quality of reviews is not compromised.\nHowever\, the hitherto available evaluations of LLMs for screening have not sufficiently engaged with this existing literature\, such that we do not yet know the extent to which they may reduce labour in a realistic setting where the risk of missing relevant studies can be appropriately managed. \nObjectives: \n– To outline a framework for evaluating LLMs for screening in a way compatible with their safe use in real projects\, by combining with stopping criteria.\n– To assess the extent to which using LLMs for screening may offer additional labour savings as compared to standard methods\, while maintaining high standards\n– To assess the extent to which model choice and prompting strategy affect potential labour savings\n– To quantify the additional costs involved in using LLMs for screening instead of standard methods \nResults: \nCurrent results show that\, using simple prompts based solely on the review title\, LLMs result in substantially smaller labour savings than standard machine-learning prioritisation pipelines using support vector machine classifiers. Further results will show the effect of more complex prompting strategies involving study inclusion criteria\, as well as the comparative costs of LLM-assisted screening and traditional approaches. \nConclusions: \nThough LLMs offer impressive capabilities given their zero-shot nature\, their simplistic application in systematic review screening may result in smaller work savings than current approaches can deliver. A future research agenda may find ways to combine active learning approaches with LLMs\, for example by automatically adjusting prompts based on inclusion and exclusion decisions.
URL:https://whatworksclimate.solutions/presentation/the-safe-use-of-llms-for-screening-in-systematic-reviews/
LOCATION:H 0112
END:VEVENT
END:VCALENDAR