Skip to main content
Annexis
Metadata Matters initiative

Partner ona study.

An open invitation

Help us reveal what open scholarly knowledge includes, what it leaves out, and what can be repaired.

We welcome research partnerships, organizational invitations, and sponsorship for focused studies that produce evidence-based findings and actionable improvements.

A shared research agenda

Questions worth answering together.

Some questions can be answered quickly. Others ask why a pattern exists, what it changes, and whether an intervention works. They need different study designs.

Filter by study type

17 questions

Open-ended

Questions that need deeper research.

These investigate influence, causality, behavior, values, and change over time. They may require hypotheses, experiments, interviews, or longitudinal analysis.

  1. O01

    AI recommendations

    How does metadata completeness shape what AI systems recommend, summarize, cite, or leave out?

    This calls for comparative observation across systems, prompts, subjects, and time. A single field check cannot answer it.

  2. O02

    Readers and representations

    Who reads research in a world mediated by AI? What should the research record become?

    Investigate what people and AI systems can see, read, interpret, cite, and reuse. Ask how those actions can be observed responsibly, where PDF remains useful, and where it becomes limiting. Explore machine readable representations that retain clear provenance.

  3. O03

    Context and trust

    One record, many contexts: what should an AI agent trust across pretraining, live retrieval, and connected knowledge systems?

    Compare how the same record is represented in a model’s pretraining corpus, search indexes, repositories, publisher pages, knowledge graphs, agent memory, and derived summaries. Study how provenance, version, recency, authority, and task should determine which context an agent relies on, especially when its pretraining sources cannot be inspected.

  4. O04

    Online attention

    How do online attention scores and signals influence AI recommendations? Can attention outweigh metadata quality or scholarly relevance?

    Separate the influence of mentions, citations, popularity, access, and structured metadata on what an AI system surfaces.

  5. O05

    Impact measurement

    How should research impact be measured in a world mediated by AI? Do established methods still hold up?

    Reconsider citations, journal metrics, downloads, online attention, policy uptake, and reuse as AI systems increasingly mediate discovery, synthesis, recommendation, and application.

  6. O06

    Repair outcomes

    When a record is repaired, does its downstream discovery, recommendation, attribution, or reuse actually improve?

    Follow records over time and across platforms to test whether a technically better record produces a meaningful change.

  7. O07

    Collective responsibility

    Who is responsible for repairing research infrastructure and missing metadata? How should that responsibility be shared?

    Study how responsibility, incentives, costs, and accountability should be distributed across authors, institutions, libraries, funders, publishers, repositories, infrastructure providers, aggregators, and AI platforms when metadata has not been collectively prioritized.

  8. O08

    Meaning and governance

    What should “complete” or “ready for AI” mean across fields, countries, languages, and communities? Who should decide?

    Explore how a readiness framework can remain useful without treating one discipline, language, or publishing system as the universal norm.

Verifiable

Questions that can start with a quick audit.

These use a defined corpus and observable checks. The result can be reproduced, compared, and turned into an immediate repair list.

  1. V01

    Fields and topics

    Which disciplines, subfields, methods, and emerging topics are underrepresented, fragmented, or misclassified in open knowledge graphs?

    Count and compare record and field coverage across a defined graph, classification, and time period.

  2. V02

    Countries and regions

    Which countries, regions, institutions, and research communities lose visibility between local publication and global discovery systems?

    Compare source records with downstream indexes to locate measurable losses in coverage and resolution.

  3. V03

    Languages and scripts

    Whose scholarship becomes invisible when titles, abstracts, keywords, names, and affiliations are not represented across languages and scripts?

    Audit the presence, script, language tags, transliterations, and translations of key descriptive fields.

  4. V04

    Research outputs

    What knowledge is missed when graphs privilege journal articles over data, software, protocols, presentations, preprints, hypotheses, and negative results?

    Measure coverage by output type and whether relationships between outputs survive indexing and aggregation.

  5. V05

    People and contribution

    How often are contributors, affiliations, funders, facilities, communities, and roles beyond authorship missing or unresolved?

    Check the presence, identifier resolution, and consistency of attribution fields across a defined corpus.

  6. V06

    Connections and context

    Which links between articles, data, code, methods, grants, corrections, and later versions are present for people but absent for machines?

    Test whether declared relationships are structured, reciprocal, resolvable, and preserved downstream.

  7. V07

    Beyond Crossref

    Which parts of the Nexus Index framework transfer beyond Crossref, and what additional gaps become measurable when other scholarly infrastructures are included?

    Map the five facets against DataCite, OpenAlex, ORCID, ROR, institutional repositories, CRIS platforms, funder systems, and domain indexes. Test field availability, identifier resolution, relationship preservation, and overlap with Crossref to identify portable checks, necessary adaptations, and new opportunities across datasets, software, preprints, grants, affiliations, versions, and agent facing access.

  8. V08

    Access and reuse

    Which records appear open but remain difficult to reuse because licences, rights, access routes, formats, or versions cannot be resolved by machines?

    Run repeatable checks for licence metadata, access endpoints, content types, version markers, and working links.

  9. V09

    Baseline machine readiness

    Can a machine resolve the identifier, retrieve the described object, distinguish its version, trace its provenance, and find its reuse conditions?

    Apply a transparent checklist to produce a fast, reproducible readiness baseline and a concrete repair list.

Ways to participate

Partner with us!

Partnerships can begin with expertise, data, access, funding, a publication opportunity, or a problem your organization cannot answer alone.

  1. 01

    Design a study together

    Bring a research question, field, region, language, or community. We shape a bounded study and interpretation plan together.

  2. 02

    Contribute a corpus

    Provide a lawful dataset, catalogue, knowledge graph, repository sample, or metadata pipeline that can support a meaningful investigation.

  3. 03

    Sponsor a study

    Fund data access, analysis, infrastructure, community participation, or open publication. Sponsorship supports the work. It does not determine the conclusions.

  4. 04

    Publish and convene

    Help us share and test our findings – in a journal, at a conference, workshop, or other community venue

Invitations welcome

Organizations across the knowledge ecosystem can take part.

Universities & librariesPublishersFundersResearch infrastructuresKnowledge graph teamsScholarly societiesPublic agencies & NGOsOpen knowledge communitiesResponsible AI companies

Work that becomes a public record

Study carefully. Publish usefully.

The intended result is publishable research with transparent scope, methods, limitations, attribution, and actionable findings. Where rights and safeguards allow, outputs may include papers, preprints, reports, code, and reusable derived data.

  1. 01Frame the question and affected community
  2. 02Agree methods, responsibilities, and safeguards
  3. 03Measure gaps and test actionable repairs
  4. 04Publish findings, methods, and reusable outputs

Start the conversation

What should open scholarly knowledge reveal next?

Tell us the question, affected community, relevant corpus or infrastructure, and how you would like to partner or sponsor the work.

Invite Metadata Matters