What are you looking for?
Group Session July 2, 2024

Gathering Metrics and Setting Boundaries: Reusing Collections as Data and the impacts of AI / Recopilando métricas y estableciendo límites: Reutilización de colecciones como datos y el impacto de la Inteligencia Artificial

Plenary: RDA 23rd Plenary Meeting – San José, Costa Rica

Submitted by

Meeting objectives

As Artificial Intelligence (AI) becomes more widely adopted, both as a tool for research across the disciplines and as a potential aid to enhancing the creation and dissemination of digital collections as data, it raises new questions about the methods of collections delivery and the ethics of responsible reuse – particularly with respect to large language models.

Historically, collecting institutions have not been good at gathering reuse metrics or monitoring reuse cases, with the exception of some workflows to avoid copyright infringement or policies against practices with high potential for reputational damage (i.e. not loaning to an unethical partner). The digital environment, and the open licences often associated with collections as data, suggests that more careful examination of the ethics of reuse is needed.

Through presentations and discussions, this session will aim to provide insight into how institutions are assessing use of collections as data and how AI impacts the parameters of that use.

Participants will share:
a. How institutions are tracking and encouraging reuse of collections as data: What kinds of metrics do we expect or want to track?
b. How institutions are writing terms of use, with respect to collections-as-data: Who are the stakeholders involved in crafting these terms? What are the expectations implied in these terms around reuse/reusability?
c. In the group discussion/outcome: What informs an institution’s stance on AI use of collections, as data?

//
A medida que la Inteligencia Artificial (IA) está siendo más ampliamente adoptada como herramienta para la investigación en todas las disciplinas como apoyo para mejorar la creación y difusión de colecciones digitales como datos, surgen nuevas preguntas sobre los métodos de entrega de las colecciones y la ética en cuanto a la reutilización responsable, particularmente con respecto a los grandes modelos de lenguaje.

Históricamente, las instituciones encargadas en la gestión de datos no han sido muy eficientes recopilando o monitorizando métricas de casos de reutilización, a excepción de algunos flujos de trabajo cuyo fin es evitar infracciones sobre derechos de autor o políticas en contra ciertas prácticas con alto potencial de daño hacia su reputación (es decir, no prestando a un socio poco ético). El entorno digital y de las licencias abiertas, a menudo asociadas a colecciones como datos, sugieren que es necesario un examen más cuidadoso en cuanto a la ética de la reutilización.

A través de diversas presentaciones y discusiones, esta sesión tendrá como objetivo proporcionar información sobre cómo las instituciones están evaluando el uso de colecciones como datos y cómo la Inteligencia Artificial afecta a los parámetros de este uso.

Las personas participantes hablarán de:
a. Como las instituciones están haciendo el seguimiento y fomentando la reutilización de las colecciones como datos. ¿Qué tipo de métricas esperamos o queremos usar?
b. Como están redactando las instituciones los términos de uso con respecto a las colecciones como datos.
¿Quiénes son las partes involucradas en la elaboración de estos términos? ¿Cuáles son las expectativas implícitas en estos términos en torno a la reutilización/reusabilidad?
c. En la discusión/resultado del grupo -> ¿Qué influye en una institución sobre qué postura adoptar en materia de Inteligencia Artificial sobre el uso de colecciones como datos?

Meeting presenters

Yasmeen Shorish, Director of Scholarly Communications Strategies & Special Advisor to the Dean for Equity Initiatives, James Madison University (Session chair)|Hannah Scates Kettler, Associate University Librarian, Iowa State University (Session chair)|Mahendra Mahey|Santi A. Thompson, Interim Associate Dean for Organizational Development, Learning and Talent, University of Houston|Ayla Stein Kenfield|Gustavo Candela, Lecturer in Computer Science at the University of Alicante|Gimena del Rio Riande, Researcher at CONICET/Universidad del Salvador

Meeting agenda

1. Welcome and Introduction to the new IG (10-15 min)
a. Menti Poll
b. Group mission and direction
2. Presentations (30-40 min)
a. Gustavo Candela, International GLAM Labs
b. Santi Thompson & Ayla Stein Kenfield, D-CRAFT
c. Mahendra Mahey
3. Break-out discussions on policy decisions supporting reuse (40-45 min)

Target audience

Digital collection stewards
Digital collection/collections-as-data users

Group Activities and Scope

This group is aimed at collections professionals such as archivists, librarians, records managers and museum curators, as well as related professions such as IT professionals, knowledge scientists, and those involved in standards development, who serve in a range of critical roles: as experts in ensuring access, preservation, and reuse of digital records, objects, data, and collections; as provocateurs for good collections curation practices; and as advocates for the construction of responsible and sustainable infrastructures for information sharing. As articulated in the Vancouver Statement on Collections as Data, “this [collections] stewardship role only grows in importance as artificial intelligence applications, trained on vast amounts of data, including collections as data, impact our lives ever more pervasively.” We recognise that there is increasing pressure on memory institutions to establish good models for responsible and culturally sensitive development of data resources while also navigating the challenges and opportunities provided by new technologies (Vancouver Statement, Principles 2, 4, 9-11), and this is not work that can be easily done in silos. This group seeks to provide a space for examining alignments in values, policy, and practice in collections work, broadly imagined, encouraging a vibrant exchange of expertise across collecting areas and domains of practice.

The proposed IG is explicitly aligned with RDA’s goal to build social bridges that enable the open and FAIR sharing of data. Specifically, the IG will build relationships among professionals who are committed to developing data sharing frameworks and practices that are able to withstand the various challenges introduced by the passage of time: technological obsolescence, loss of contextual understanding of data, and resource constraints that make it impractical to commit to preserving all data forever. By creating a space for collections professionals and data curators to come together, this IG has the potential to act as a launching pad for RDA Working Groups that will produce recommendations addressing these fundamental challenges to research data sharing, specifically by bringing preservation, arrangement and description, and appraisal methodologies from the cultural heritage and information management sectors to the wider RDA community. This IG also aims to bring the skills and competencies which have long existed in memory institutions into the wider RDA community, many of whom will not be aware of this existing expertise.

Additional links to informative material

https://www.cepal.org/es/notas/webinario-especial-unete-nosotros-la-celebracion-decada-innovacion-digital|https://doi.org/10.1002/asi.24835|https://thinkepi.scimagoepi.com/index.php/ThinkEPI/article/view/91636|https://rua.ua.es/dspace/handle/10045/110281|https://www.kb.nl/en/ai-statement|https://glamlabs.io/checklist/|https://doi.org/10.1108/GKMC-06-2023-0195

Short Group Status

This group started with a BoF at the RDA Plenary 21 in Salzburg with the theme “Why aren’t we talking about Collections as Data?” It absorbed the former Archives and Records Professionals Interest Group as a key contributor to a wider conversation aimed at collections professionals broadly as well as anyone working with cultural heritage data. The work in this Interest Group is just beginning.

Estimate of the required venue room capacity

Between 60-100 seats

Applicable Pathways

Data Infrastructures and Environments - Generalist
Other (AI)

Please indicate at least (3) three breakout slots that would suit your meeting.

Breakout 1. Tuesday, 12 November, 16:00-17:30 UTC
Breakout 5. Wednesday, 13 November, 18:30-20:00 UTC
Breakout 6. Thursday, 14 November, 16:00-17:30 UTC