What are you looking for?
Working Group Recommendation June 29, 2026

Data Director Agentic AI Tool Blueprint

  • Output Type: Working Group Recommendation
  • Output Status: Endorsed
  • Review Period End: 2026-08-27
  • DOI: 10.15497/RDA00157
  • Regions: Global
  • Primary Domain: Domain Agnostic
  • RDA Pathways: FAIR, CARE, TRUST - Adoption, Implementation, and Deployment,Data Lifecycles - Versioning, Provenance, Citation, and Reward,AI meets data: exploring use cases, applications and innovation
  • Group Technology Focus: Citation & Provenance, Data (Output) Management Planning, Depositing Research Outputs
  • Stakeholders: Early Career Individuals, Funders & Policy makers, Industry & Private Sector, Infrastructure Providers, Libraries, Regions & Nations, Research Performing Organisations, Researchers & Scientists
  • Sustainable Development Goals: Industry, Innovation and Infrastructure, Reduced Inequality, Responsible Consumption and Production
  • Language: English

Abstract

The Data Director Agentic AI Tool Blueprint v1.0 is a community-developed, technology-agnostic specification for an agentic AI tool to support researchers in preparing and depositing research data for publication in alignment with the FAIR principles. Produced by the Data Director Agentic AI Working Group, an RDA-Microsoft collaboration, it defines 12 functional requirements, 17 non-functional requirements, and 14 architecture principles, alongside a six-phase process flow, a reference architecture, implementation guidance, and evaluation criteria. The Blueprint is version 1.0, final version for endorsement by the RDA Council, and is transparent about what remains unresolved, offering a basis for ongoing community discussion as the specification develops.

A research software engineer at a large academic computing centre is engaging with the Blueprint in two ways. First, through their own research on research software metadata, where the Blueprint’s grounding of quality checks in existing validators, and its recognition that schema compliance alone does not confirm semantic correctness, maps onto ongoing work on cross-record metadata consistency; the Blueprint’s framing will be applied and findings reported back. Second, through dissemination within their institution, introducing the Blueprint to research software engineering and data support colleagues and gathering informal feedback. This is a personal commitment and does not commit the institution, nor extend to deploying a Data Director instance.

Impact Statement

The Data Director Agentic AI Tool Blueprint addresses a systemic gap in research data practice: the uneven access to skilled data support that determines whether research data is prepared, documented, and deposited in a way that makes it findable, accessible, interoperable, and reusable. By defining an open, community-governed specification for an agentic AI tool, the Blueprint creates a foundation for agentic AI tool implementations that could narrow this gap across institutions, disciplines, and geographies, supporting the long-term goals of open science and making high-quality research data a realistic outcome for any researcher, regardless of their institution or expertise.

Explanation of Sustainable Development Goals

Industry, Innovation and Infrastructure
The Blueprint provides an open, community-governed specification for agentic AI tool implementations that strengthen the infrastructure through which research data is prepared, deposited, and made reusable across institutions and disciplines.

Reduced Inequality
By addressing systemic gaps in researcher capacity, metadata standards, access to data support, and the complexity of navigating funder, journal, and institutional requirements, the Blueprint aims to support equitable access to high-quality data support tools regardless of institution, discipline, or geography.

Responsible Consumption and Production
The Blueprint promotes the responsible stewardship of research data through a functional specification covering metadata quality, provenance, persistent identification, and the maximisation of value from publicly funded research outputs.

Citations

Abarca, M., Anderson, L., Arancio, J., Aryani, A., Aversa, R., Azizan, A., Bajpai, S., Berti-Equille, L., Buendia, P., Cannon, J., Clare, C., Crawley, F. P., Drira, M., Earl, G., Edgerton, J., Eklund, L., Ferradosa, E. E., Ferretti, R., Fuentes Pardo, A., … Zwölf, C. M. (2026). Data Director Agentic AI Tool Blueprint (Version 1.0). Zenodo. https://doi.org/10.15497/RDA00157

Comments

  • Profile Picture

    July 28, 2026 at 1:38 pm

    Francis P. Crawley says:

    Alongside continuing work for the RDA AIDV WG and the development of the VANTAGE IG, I have looked at how the points raised there fit with current regulation and policy in the European Union, the United States, the United Kingdom, China, the African Union, Latin America, and the UNESCO Recommendation on the Ethics of Artificial Intelligence, together with the NIST AI Risk Management Framework. The convergence is closer than I had expected. Each of the suggestions describes a key concern: a capacity to decline, to suspend, and to refer a matter to an accountable person. This capacity is not present in the Blueprint v0.1.
    The conclusion is a favourable one for the Blueprint Working Group. The improvements suggested in my earlier comment would not constrain the tool. They would support compliance across those jurisdictions, widen the settings in which the Data Director can be implemented, and should improve uptake.
    Francis P. Crawley
    Leuven, Belgium
    [email protected]

    • Profile Picture

      August 6, 2026 at 9:32 am

      Connie Clare says:

      Dear Francis,

      Thank you for this additional cross-jurisdictional review. This is a highly valuable piece of work, and the convergence you have found across the various international recommendations strengthens the case for the concern you raised in your previous comment - that a capacity to decline, suspend, or refer a matter to an accountable person is not currently present in the Blueprint.

      This information provides important context should the ‘Refusable by Design’ and wider ‘hold framework’ you proposed be taken forward in v2.0. 

      Thank you again for the depth and rigour you continue to bring to this review.

      Connie Clare
      (RDA Community Development Manager and WG Facilitator) 

  • Profile Picture

    July 27, 2026 at 3:57 pm

    Timothy Cook says:

    On behalf of Axius SDC, Inc., with co-author Christopher Steven Marcum, we have completed a detailed mapping of the Semantic Data Charter (SDC)
    to the Data Director Agentic AI Blueprint, presented as a candidate open reference implementation aligned to the Blueprint's requirements:
    FAIR, interoperability without re-mapping, persistent identifiers and the sovereignty case, and deterministic governance enforcement, with
    reproducible, publicly runnable backing.

    The full contribution is too extensive for this box, so we have linked it as a downloadable PDF.

    It maps each capability to specific Blueprint requirements, documents the supporting evidence (the FAIR Data Demo, the MTCP×SDC integration,
    and the CordovaOS interoperability demonstration), and states current limitations and next steps plainly. We would welcome the Working Group's
    review and would be glad to contribute a worked conformance example toward v0.2.

    For more information, contact Timothy Cook at [email protected].

    • Profile Picture

      August 6, 2026 at 9:31 am

      Connie Clare says:

      Dear Timothy (and Christopher),

      Thank you for your positive sentiment about the Data Director Agentic AI Tool Blueprint, and for the detailed mapping you have shared. It is encouraging to see the Semantic Data Charter (SDC) align with the Blueprint's requirements, and we appreciate the effort behind the demonstrations and documentation.

      This kind of conformance work warrants a full community discussion, and I am confident many Working Group members would be keen to engage with it, particularly those actively working towards their own implementations of the Blueprint. Perhaps I could connect you with them, if you think that would be beneficial, so you can be part of those conversations. With this in mind, your submission has been added to the list of Potential Data Director Blueprint Implementors in Appendix D and Community Review Comment Log in Appendix G of the final v1.0 Blueprint, recorded with the date and your details for provenance, so it can be easily located and discussed should a v2.0 process go ahead.

      To be transparent about where things stand, the facilitation agreement between Microsoft and the RDA Secretariat supporting this work has now terminated, so the RDA does not currently have the mandate or capacity to continue facilitating this Working Group. If that support is identified in future, your mapping would be a strong candidate for a formal conformance and evaluation framework, and I would encourage you to keep the offer of a worked example open for that stage.

      Thank you for engaging with this community review process. 

      Best regards, 

      Connie Clare
      (RDA Community Development Manager and WG Facilitator) 

  • Profile Picture

    July 26, 2026 at 9:43 pm

    Francis P. Crawley says:

    Community review comment on the Data Director Agentic AI Tool Blueprint v0.1

    Francis P. Crawley, EOSC-Future/RDA Artificial Intelligence and Data Visitation Working Group (AIDV); RDA BIDT IG; HGP2; GCPA; SIDCER. ORCID 0000-0002-6893-5916.

    Submitted to the community review of RDA Recommendation DOI 10.15497/RDA00157, closing 28 July 2026.

    My thanks first to Connie Clare, and to everyone who gave eight weeks to this. Two sessions a week across every time zone is a great deal to ask of volunteers, and the document that came out of it is serious, careful, and considerably better than the pace of its production would lead anyone to expect. I offer what follows as a member of the working group and as one of its authors, in the spirit of a specification that says of itself, in Section 13, that it is transparent about what remains unresolved.

    I want to begin by recording how much of what was raised during the sessions was taken up, because it was taken up generously and it should be noted:

    P4 now distinguishes human-in-the-loop from human-on-the-loop, requires that the mode be documented per workflow, and adds that the tool may assist and recommend but must not decide, authorise, or take responsibility for consequential decisions. That last clause is the single most important sentence in the document, and it was not there in April. P2 now carries "as open as possible, closed as necessary" rather than an unconditional default. P10 extends sovereignty beyond the primary dataset to AI processing, prompts, logs, temporary files, backups, embeddings, telemetry, and derived artefacts, which is exactly where sovereignty is lost in practice. P12 requires that FAIR be informed by CARE from the outset rather than as an afterthought, and that Indigenous and community stakeholders be included in governance and design. Section 5.1 requires implementers to determine the level of human oversight for each workflow before configuring automation. Section 5.2 requires that sovereignty and community authority inform openness decisions from the outset and not retrospectively. Section 5.4 requires that agents never operate anonymously and that every agent action be executed on behalf of a named, identified human. And Table 10 carries forward, as a live open question, whether P2 should be reframed as "as open as possible, as sovereign, protected, and accountable as necessary", conditioned on human rights, consent, and community authority.

    That is a substantial body of governance thinking, and it is present at the level of the specification rather than in a preamble. I am grateful for it, and nothing below should be read as a complaint about a document that did this much this quickly.

    An observation about the vocabulary: I want to offer one observation that took no argument to produce and that anyone can check in a few minutes with the search function.

    The Blueprint is rich in the language of human assent. Counting all uses across the document, "confirm" and its variants appear 54 times, "review" 52 times, "verify" 39 times, and "approve" 18 times. The human being in this specification is asked to confirm, to verify, to review, and to approve, and the conditions under which each of those happens are worked out with real care.

    The following words do not appear anywhere in the Blueprint: withhold, dissent, refuse, decline, halt, stop, disagree, veto.

    Two words come close, and it is worth saying exactly how they fall short. "Override" appears six times, and in every instance it describes a human overriding an output produced by the AI, governed under C14 by the qualification that override rights must be role-based rather than open to all users. It is a correction of the machine's work, not an arrest of the workflow. "Pause" appears four times, and in every instance it is the tool pausing itself upon detecting personal identifiers. That is an automated safeguard, and a good one, but it is the system stopping itself rather than a person stopping the system.

    So the specification is fully articulate about the human who says yes, and silent about the human who says no. There is no artefact in the document that a person can use to stop a workflow, no place where such a stop is recorded, and nothing that states what follows from one.

    I do not think anyone decided this. It is what a well-run requirements process produces, because approval is specifiable and refusal, in its most important form, is not. A rule can be written down, but it does not contain the instructions for its own application, and the moment of applying it to the case in front of you is where a person is still needed. Requirements engineering captures the rule and cannot capture that moment, so the moment goes first unrecorded and then unnoticed. That is precisely why it is worth naming.

    The document already knows: What persuades me that this belongs in v0.2, rather than in some separate conversation about principles, is that the Blueprint has already found the gap from several directions and had nowhere to put it.

    Table 10 identifies as the key community governance question how the safety, transparency, and resilience of the tool should be assessed, reported, and governed over time, including where humans remain in the loop, what outputs must always be reviewed, and how escalation and suspension should work. Escalation and suspension are named there and specified nowhere. Section 13.2 already contemplates adding a new named architectural principle. Section 10.1 already distinguishes gaps requiring agent or system development from gaps that can only be closed through human judgement, which is exactly the right category for what follows.

    Two colleagues have arrived at the same place. The two comments received in this review converge on the gap from different directions, and I would like to acknowledge both.

    Seonyoung Kim, commenting on 30 June, proposed that governance be treated not as an early step within a single workflow but as the entry point to the workflow itself, with the agent first establishing what consent and ethics approval permit, whether data should be shared openly, through controlled access, or not shared at all, and routing the researcher accordingly. That phrase, "or not shared at all", is the missing terminal state. The six-phase process flow begins at planning for data publication, which means that publication is presupposed by the first box, and there is no modelled path in which the considered outcome is that this dataset should not go out. Seonyoung Kim identified this on the first day of the review, and it has waited four weeks for a reply.

    Pengyin Shan, commenting on 23 July, welcomed Section 10.1's candour that no community-agreed benchmark exists for judging AI-generated metadata and that schema compliance cannot confirm semantic correctness, proposed cross-source and internal consistency as a benchmark-free quality signal, and asked that indirect prompt injection be named explicitly, pointing to the OWASP Top 10 for Agentic Applications and to NIST's agent-hijacking evaluation work. The closing point bears directly on mine. Pengyin Shan observes that P4's human approval and R10's provenance logging are already strong mitigations, and that the gap lies in whether they hold under adversarial input. I would put it this way. A workflow that a human can only approve, correct, or allow to proceed is structurally weaker under adversarial conditions than one a human can stop, because assent under time pressure is the easiest human behaviour in the world to manufacture. A hold is an ethical instrument and a security control at once, and the two comments support one another.

    I would also note two posts to the group. Raymond Uzwyshyn's persona-based architecture, shared on 15 July and presented to the AIDV Working Group, has a property this discussion needs, whatever one makes of the persona model as a whole: a persona is a role, and a role can bear a refusal in a way that a capability cannot. And Filippo Vasone's presentation of the Blueprint to the Community of Italian Data Stewards is a reminder that this specification is already travelling, which is a reason to get v0.2 right rather than to defer.

    Ethics has two dimensions. The external dimension is norms, procedures, approvals, and accountability. It can be required of anyone, audited, written down, and, because its correct application is settled by a rule, it can in principle be mechanised. The Blueprint performs this dimension competently, and I mean that as praise rather than as faint praise. The internal dimension is conscience: the capacity of a person to be brought up short, to halt an action that reason, rule, and interest all permit, before the ground of the halt is available to the one who is halted. It is negative in form. It prevents and does not direct. It gives no reasons at the moment it acts, though reasons are often found afterwards, and it operates most characteristically where knowledge is absent. It cannot be required of anyone. It can only be exercised, and it is exercised at cost.

    Let me be careful about what I am not claiming, since the obvious reply to a comment like this is that it flatters human beings at the expense of machines.

    I am not claiming that the Data Director lacks conscience. I could not establish that. Nothing I am able to examine would settle the presence or absence of conscience in a machine, and the same holds, if I am honest, for a colleague, and for myself except in the exercise. That symmetry is uncomfortable and I accept it. My claim is narrower and, I think, harder to dismiss: this specification, as drafted, gives no one an occasion to exercise conscience. Not the researcher, not the data steward, not the implementer. The document asks the human being for assent in four verbs and one hundred and sixty-three instances, and never once provides for the human being who is not able to give it.

    This is not a private vocabulary, and it is not an outside imposition on a data specification. It is already in the instruments under which this work is done. Article 1 of the Universal Declaration of Human Rights holds that human beings are "endowed with reason and conscience". The two are named together and they are not the same faculty. The Declaration of Helsinki, at paragraph 4, provides that "the physician's knowledge and conscience are dedicated to the fulfilment of this duty". That sentence was carried unchanged into the 2024 revision, the tenth amendment of the Declaration in sixty years, and one that altered a great deal else in order to address data, biobanks, and the research use of personal information.

    It is worth pausing on what this means for the Blueprint specifically. Section 5.2 requires confirmation that an ethics approval body has given its approval, and the document refers to an IRB, an REC, or an equivalent national body six times. Those bodies exist in part because of the Declaration of Helsinki. So the specification checks for the output of a system that is founded, in its fourth paragraph, on a sentence about conscience, and specifies nothing whatever about conscience. The Blueprint contains no reference to Helsinki, to Nuremberg, or to Belmont. Section 5.2 also requires that legal, ethical, and human rights compliance be treated as prerequisites for any sharing action, so the human rights frame is invoked in that same subsection. I am not asking the Blueprint to import a foreign idea. I am observing that it has already imported the institution and left the idea behind.

    Nor is this one tradition's idea, which matters here, because the risk of adopting a single ethico-legal framework as the global benchmark is a proper concern in the RDA, and because P12 already requires that community voices shape how their data is represented and acted upon. The account of conscience I have given is old, and it is not European property. Socrates described a daimonion, a certain voice, which only ever turned him back from something and never once told him what to do, which is the negative form exactly. Ubuntu, in the proverb umuntu ngumuntu ngabantu, a person is a person through other persons, holds that the capacity to be answerable is constituted in relation rather than possessed privately, which bears directly on P12 and on CARE. Al-Ghazālī, writing in Arabic in the eleventh century, gave an account of a crisis in which he found himself unable to continue teaching, his speech failing him for months before he could offer any reason for it, and built from that experience a discipline of self-examination, muḥāsaba, which treats the arrest as prior to its justification rather than as its conclusion. And Wittgenstein made the structural point on which all of this rests, that a rule does not contain the instructions for its own application, which is why no specification, however complete, closes the space where a person is still required. Four sources, widely separated in time, language, and tradition, and one shape: the halt comes first, and the reasons come afterwards, if they come at all.

    An institution cannot manufacture conscience. But it can make it unexercisable, and it usually does so not by forbidding it but by leaving no place for it. A procedure that recognises only fully reasoned objection has silenced the person whose doubt is real but not yet formed, and has done so while appearing to honour dissent. This is the specific way in which good tools erode judgement: not by overruling anyone, but by asking nothing of them except agreement.

    That is the whole of my concern, and it is a concern about us rather than about the tool.

    Four conditions a specification can carry

    Conscience cannot be specified, since it issues no positive content and a requirement needs content. What can be specified are the conditions under which it remains possible to exercise. There are four, they follow from what conscience is, and each of them can be written into a table in this document.

    Time. Conscience speaks first as unease and only later as argument. Any procedure that requires a reasoned objection at the moment of decision has abolished the thing while appearing to honour it. Temporal compression, not automation as such, is how an agentic system actually threatens judgement.

    Non-resolution. An objection that is outvoted, closed, or resolved into a row has been processed. An objection that is recorded and travels with the decision has been preserved. Aggregation dissolves conscience by working correctly, which is why the record matters more than the resolution.

    Reasonlessness. The arrest arrives before its own justification. A mechanism that demands grounds at the moment it is invoked is available only to the person who already has the argument, which is to say the person who needs it least.

    Survivable cost. An arrest that costs nothing is a preference rather than a conscience, so the cost cannot be removed. But an institution that makes the cost unsurvivable has abolished the capacity as surely as one that forbids it. The aim is protection of the person, not reward for the act, and the difference matters: a reward converts conscience into compliance with a different incentive.

    The proposals below are these four conditions, expressed as amendments to numbered artefacts. Each is tagged with the condition it carries.

    Proposals for v0.2

    I have kept these to specific locations so that they can be weighed on their merits and adopted, amended, or rejected row by row.

    1. A new architecture principle [reasonlessness, time], using the slot that Section 13.2 already contemplates for a new named principle. Proposed name: Refusable by Design. Statement: an accountable human must be able to stop any Data Director workflow, and the stop must not require a stated reason at the moment it is made. Implication: implementers must provide a hold that suspends the workflow at any phase, that is available without prior justification, that is recorded, and that persists until it is lifted by a documented decision. Reasons may be added afterwards and should be invited, but must never be a precondition of the hold taking effect.

    2. Amend P8, Explainable AI [reasonlessness]. P8 rightly requires the machine to explain how it reached an output. The symmetry should be stated: the requirement of explanation runs from the system to the person and not from the person to the system. A human hold requires no accompanying rationale in order to be valid.

    3. Amend P9, Locally Adaptable [reasonlessness]. P9 already establishes that core FAIR and security requirements cannot be overridden at institutional level. The availability of the hold belongs in that non-configurable core. Who may exercise it, and at what level of the organisation, is properly configurable. Whether it exists at all should not be.

    4. A new functional requirement, R13 [non-resolution]. Record a hold: who placed it, at which phase, against which action, the time, and the outcome, being whether the hold was lifted, the workflow abandoned, or the matter escalated. Priority MUST. Reasons, where given, are recorded but are not a required field.

    5. Amend R10, provenance [non-resolution]. R10 currently traces agent actions, artefacts, inputs, agent versions, and human authorisations. As drafted, the provenance record preserves every approval and no hesitation. It should record refusals and holds alongside authorisations, so that PROV-O captures the whole of human involvement rather than only its affirmative half.

    6. Amend R8, DMP verification [non-resolution]. R8's requirement that discrepancies be flagged for human review rather than automatically enforced is the closest thing in the document to what I am describing, and it stops one step short. R8 should state explicitly that a determination not to share is a valid terminal outcome of the check, and not an exception, a failure, or an incomplete workflow.

    7. Amend C13, AI Governance [time, non-resolution]. This answers the Table 10 question directly. C13 should specify suspension: that any workflow may be suspended by an accountable human, that suspension takes effect immediately and without prior justification, that suspended workflows neither expire nor auto-resume, and that resumption requires a documented human decision. Table 10 asks for exactly this and it can be written now.

    8. Amend C10, Performance [time]. C10 requires response times appropriate to operation complexity and visible progress feedback. Without a carve-out, a deliberation interval will be measured as latency and optimised away, and the principle in proposal 1 will be defeated by a metric sitting four tables from it. C10 should state that time taken by a human in deliberation, review, or hold is excluded from performance measurement.

    9. Amend C12, Data Governance, and C16, Measurability [survivable cost]. Holds should be recorded and auditable, but not individually attributable to named persons in routine performance or usage reporting. A person who has stopped something should not find that fact on a dashboard beside their name. This is the practical form of keeping the cost survivable without pretending it can be removed.

    10. Amend C14, Explainability [reasonlessness]. This addresses a concern already recorded in the working group notes, and raised there by Natalie Meyers, that not every researcher can be expected to understand every AI output or should hold the right to override any of them. The concern is correct and the hold does not disturb it. Overriding an AI output is a substantive act requiring competence and a governed role. Stopping a workflow is not an override; it is a request that a human decision be taken before anything further proceeds. Separating the two in C14 allows the first to remain properly restricted while the second remains broadly available.

    11. Amend the process flow, Section 8.2 [non-resolution]. Phase 2 already places governance and constraints checks before repository selection, which is the right ordering and answers part of what Seonyoung Kim proposes. What Phase 2 lacks is a terminal state: an outcome in which the considered answer is that this dataset should not be published, reached deliberately and recorded as a result rather than as an abandoned session. Phase 6 lacks the reciprocal path, a route back to a person when something surfaces after deposit, which the discussion of tombstone records already anticipates.

    12. Add an evaluation criterion under Section 10.1 [all four], tagged [Human/user input needed], on whether implementations preserve the capacity for human arrest. One measure resists optimisation and I propose it directly: a hold rate of zero across an implementation is a finding requiring investigation and not a mark of success. It indicates either that nothing troubling has arisen, which is worth knowing and unlikely at scale, or that the mechanism is unavailable, unknown, or too costly to use.

    13. Add to the implementation guidance, Sections 11.2 to 11.4 [survivable cost]. Each deployment should state, in its governance documentation, who holds the hold, at what level of the organisation, and what process follows when it is exercised.

    One point on the inherited definition

    The November 2025 consultation defined agentic AI as systems capable of autonomous operation with minimal human oversight, and the Statement of Work carries that definition forward. The Blueprint describes a human-supervised tool and invokes human oversight throughout. I do not think this is an error by anyone. The working group inherited a definition from a prior phase and then, in eight weeks of careful work, built something that has outgrown it. It would be worth saying so in v0.2, in a sentence or two, because the distance between minimal oversight and human supervision is the whole of the governance question, and a specification that closes it explicitly will serve implementers far better than one in which the two definitions sit unreconciled in adjacent documents.

    Closing

    None of this requires the Blueprint to be reopened, and none of it displaces what is there. Section 5 should stay as it stands. What I am proposing is one principle, one functional requirement, a handful of amendments, a terminal state in the process flow, and one evaluation criterion, offered in answer to a question the Blueprint itself identifies as the key unresolved governance question.

    A tool of this kind will be built, and built well, and it will do a great deal of good. The only thing I would ask of it is that somewhere in its specification there be a place for the person who cannot yet say why, and will not go on. That is not a constraint on the tool. It is the condition under which the people using it remain the ones responsible for what it does. Oakeshott described civilisation as a conversation among disparate voices, each speaking in its own idiom, none reducible to another and none entitled to become the master voice; a specification that provides only for assent has quietly made itself the single voice in a conversation that depends on there being more than one.

    I do not offer any of this as the end of a discussion. It is closer to the beginning of one, and it is worth putting a few markers down for where it might go. The first is v0.2 itself, where the thirteen items above can be argued row by row and some of them rejected. The second is that this question does not belong to the Data Director alone. It will arise for every agentic tool the RDA specifies, including the Literature Librarian and the Funding Finder if they are taken up, and it would be better settled once in a form that transfers than thirteen times in thirteen documents. The two virtual AI cross-fertilisation meetings ahead of P27 are a natural place to test whether the community agrees, and P27 in London is a natural place to take it further, alongside the persona-based work Raymond Uzwyshyn intends to bring. The third is broader still, and I raise it only to leave it on the table: whether the RDA, whose outputs increasingly shape how research data and AI are governed worldwide, wishes to develop a settled way of holding questions of this kind. That is a matter for the Council and the community rather than for this review, and I mention it here only so that it is not lost when the review closes.

    I would be glad to draft any of these rows in the Blueprint's own table format for the working group's consideration, and to do the work rather than only to propose it. My thanks again to Connie, and to everyone who built this.

    Francis P. Crawley
    Leuven, Belgium
    [email protected]

    Sources referred to above: Universal Declaration of Human Rights (1948), Article 1. World Medical Association, Declaration of Helsinki: Ethical Principles for Medical Research Involving Human Participants, as revised October 2024, paragraph 4, at https://www.wma.net/policies-post/wma-declaration-of-helsinki/. Word counts and term occurrences are taken from the published v0.1 document of 29 June 2026 and can be reproduced with any text search.

    • Profile Picture

      August 6, 2026 at 9:29 am

      Connie Clare says:

      Dear Francis,

      Thank you for the depth and care you have brought to this Working Group, and to this review specifically.

      Your vocabulary analysis is a fair observation. The Blueprint captures assent (agreement and approval) well but misses refusal. This is likely not a deliberate omission, but because approval is simply easier to specify than refusal. This underpins your central proposal, to add a new architecture principle, ‘Refusable by Design’ (allowing a human to stop an agentic workflow), and your observation about P4's gap, that there is no mechanism for a human-initiated stop independent of system-triggered approval points.

      This kind of scope expansion, adding a new architectural principle, warrants full community discussion and agreement. The resultant dependent amendments to various other principles, functional and non-functional requirements, including a new functional requirement, R13, to record a hold; amendments to P8, P9, R10, C13, C10, C12/C16 and C14; a new evaluation criterion flagging a ‘hold rate of zero’ for investigation; and implementation guidance on who holds the hold at what organisational level, are all grounded in the same premise. This also extends to the process flow (the community-built six-phase workflow produced through facilitated Working Group sessions, which are expressed as diagrams rather than text), whereby adding a new terminal state and return path would be a structural redesign with effects across the requirements and later sections it connects to.

      This is not a single insertion point but would require significant revision and review of the Blueprint. Adopted piecemeal, it would sit inconsistently; adopted collectively, it is an expansion of scope requiring the same deliberate community process that built the original specification. This should be revisited as a package in v2.0.

      Thank you very much for already providing a draft of these additions in the Blueprint's own table format via email. This is a generous and practical contribution, and I suggest sharing it as a post to the Working Group so it can be discussed by the group and taken forward as part of a future version, should a v2.0 process go ahead. 

      With this in mind, your suggestions have been added to the Community Review Comment Log in Appendix G of the final v1.0 Blueprint, recorded with the date and your name for provenance, so they can be easily located and discussed together should a v2.0 process go ahead.

      Your amendment to R8, stating that ‘do not share’ is a valid terminal outcome of the DMP-alignment check, does not depend on the ‘Refusable by Design’ principle or wider ‘hold framework’ you propose, and has been integrated directly into the Blueprint (Table 3, Functional Requirements of the Data Director). Similarly, your note on the ‘minimal oversight’ definitional gap has been addressed with a clarifying sentence in Section 2.4. These too have been logged in Appendix G.

      The connection drawn to Seonyoung Kim's and Pengyin Shan's comments has been noted. All three point to the same underlying gap, that the Blueprint specifies how a human approves an action, but not how one can halt or refuse it. This strengthens the case for addressing it together in a future version.

      To be transparent about where things stand, the facilitation agreement between Microsoft and the RDA Secretariat supporting this work has now terminated, so the RDA does not currently have the mandate or capacity to continue facilitating this Working Group. If that support is identified in future, your proposals would be valuable discussion points for the Working Group as a whole.

      Thank you, again, for engaging with this community review process.

      Best regards,

      Connie Clare
      (RDA Community Development Manager and WG Facilitator) 

  • Profile Picture

    July 23, 2026 at 6:50 pm

    Pengyin Shan says:

    I really appreciate how candid Section 10.1 is that "no community-agreed benchmark exists for evaluating whether AI-generated metadata meets an acceptable curation standard," and that schema compliance alone cannot confirm the semantic correctness of values. That is an honest framing of a genuinely hard problem, and grounding R4 in existing validators is a solid starting point.

    Two suggestions for v0.2:

    1. Consistency checking as a benchmark-free quality signal: The draft notes that no community-agreed benchmark exists for judging AI-generated metadata, and that schema compliance alone cannot confirm semantic correctness. Cross-source and internal consistency may offer a partial way forward: when a generated field contradicts its source record, that disagreement is measurable today without a gold standard. This could serve as an automatable tier in the proposed gold/silver/bronze framework (Section 10.3), and as an early, tractable step toward the hallucination-measurement item in the Section 13.2 roadmap. It also aligns with the long-established metadata quality dimensions (Bruce & Hillmann), where logical consistency sits alongside completeness and accuracy.

    2. Cybersecruity Enhancement: Section 10.1 states that no methodology yet covers robustness against the tool being manipulated or redirected out of scope. This landscape has moved quickly: indirect prompt injection is now well characterised, and the tool can ingest many channels like repository records, DMPs, READMEs, etc. It may be worth naming this threat explicitly, and pointing to the OWASP Top 10 for Agentic Applications (Dec 2025) and NIST's agent-hijacking evaluation work as concrete starting points. The existing P4 human-approval and R10 provenance-logging are already strong mitigations. The gap is in testing whether they hold under adversarial input, and structured red-teaming approaches for exactly this are starting to appear.

    Happy to help develop either of these further if useful, and let me know if there's anything else I can help!

    • Profile Picture

      August 6, 2026 at 9:27 am

      Connie Clare says:

      Dear Pengyin,

      Thank you very much for your time and effort reviewing the Blueprint, and for these two suggestions.


      1. Consistency checking as a benchmark-free quality signal: this has been implemented across Sections 10.1, 10.3, and 13.2; as a new measurable signal under the ‘Cannot yet be measured’ gaps in 10.1 (Area 1), within the gold/silver/bronze quality tiers in 10.3, and as a tractable first step toward the hallucination-measurement roadmap item in 13.2.


      2. Cybersecurity enhancement: this has been implemented in Section 10.1, Area 5, naming indirect prompt injection explicitly and citing the OWASP Top 10 for Agentic Applications and NIST's agent-hijacking evaluation work as starting points, alongside the existing P4 and R10 mitigations you noted.


      Both suggestions have also been recorded in the new Community Review Comment Log in Appendix G of the final v1.0 Blueprint, with the date and your name for provenance. I would also like to add that you have been added as a potential implementer in Appendix D

      Thank you, again, for engaging with this community review process. 

      Best regards,


      Connie Clare
      (RDA Community Development Manager and WG Facilitator) 

You must be logged in or join the group to leave a comment.