Climate Knowledge Graph — Final Report: TIB Innovation Fund R&D Project (2025/26)
A knowledge graph pilot for the IPCC Sixth Assessment Report
The project uses Wikibase software to create a community knowledge graph for climate change science literature and is intended for use by the public, policymakers, and scientists. The knowledge graph starts by importing the 10,000-page Sixth Assessment Report of the Intergovernmental Panel on Climate Change (IPCC). The FAIR Principles are used to guide the conversion of the report into structured data for enhancing cataloguing, document distribution, and data analysis.
By Simon Worthington, simon.worthington@tib.eu1 (ORCID iD: 0000-0002-8579-9717) and Laura Oldenbourg (ORCID iD: 0009-0003-5070-0099)
Leibniz Information Centre for Science and Technology and University Library of Hannover (TIB) (ROR ID: ror.org/04aj4c181)
August 2026
Project details
Project name: Climate Knowledge Graph (ClimateKG)
Project details information: Wikidata ID Q1366995142 and as XML data ‘ClimateKG project info’3 (using ResearchProject schema from schema.org)4
TIB – Leibniz Information Centre for Science and Technology and University Library (ROR ID: ror.org/04aj4c181)
Funded by the TIB Innovation Fund 2025/26
Project links
ClimateKG main website: climatekg.tibwiki.io/
ClimateKG Data Bench: tibhannover.github.io/ClimateKG-Data-Bench/
ClimateKG Git repository: github.com/TIBHannover/climate-knowledge-graph
Documentation and development log: tibhannover.github.io/climate-knowledge-graph/
Slides / poster
Software Architecture Plan (Slides, Oct 2025)
About ClimateKG (Slides, July 2026)
Data Siloes (Poster, June 2026)
Team
TIB — Open Science Lab: Simon Worthington, (Project Lead, Publishing technology expert); Laura Oldenbourg, (Publishing expert and data engineer).
ClimateKG partners with Lab Knowledge Infrastructures (TIB) led by Dr Markus Stocker ORCID: 0000-0001-5492-3212.
Thank you to TIB R&D teams: Wikibase4Research (NFDI4Culture) and Open Research Knowledge Graph (ORKG).
Partners and contributions
#semanticClimate is a founding project partner and is an open research group liberating knowledge from climate-related literature. #semanticClimate has three activity tracks: Developing software for document semantification, an AI Chatbot, and a youth education programme.
#semanticClimate: Gitanjali Yadav ORCID: 0000-0001-6591-9964; Peter Murray-Rust ORCID: 0000-0003-3386-3972; and; Renu Kumari ORCID: 0000-0002-9451-7814.
Independent: Raquel Perez de Eulate (data visualisation); Darron M. Broad, Runstop (software development).
Executive summary
A yearlong R&D project as a partnership between TIB – Leibniz Information Centre for Science and Technology and #semanticClimate research group. The project goal is to make a knowledge graph of the 10,000-page Intergovernmental Panel on Climate Change — Sixth Assessment Report.
The mission of the project is to demonstrate the value of semantically structuring a large text corpus to help reveal its urgently needed embedded knowledge.
The target audiences are policy makers, data scientists, and citizen science activities.
Technology developed in the project provides a solid basis to create a system that can be used on any text corpus.
Figure 1: ’IPCC report: From browsable website to dynamic data resource’. The IPCC report is transformed by extending the metadata to map out the document internals. The metadata is then deposited in Open Science infrastructures as well as being accessible for searching using APIs and SPARQL interfaces via Wikibase.
The project goal is to produce a framework knowledge graph using Wikibase, with an attached full text of the reports using MediaWiki. The Wikibase / MediaWiki software framework is used and is supported by a robust hosting and data release infrastructure for quality control and scalability.
The reports have been broken down into five foundational datasets which act as a syntactic (structural) base, as: The corpus full text and structure; bibliographic information; glossary terms; acronyms and; authors.
The knowledge graph has two attached components designed for utilising the datasets:
A data workbench for data scientists and citizen science projects to analyse and enrich the reports, and
A distribution framework for documents and metadata to enable better and faster report access.
The innovations of the project have been:
CKG Scrape (Software): A web scrape of the complete IPCC AR6 report into a knowledge graph built on MediaWiki and Wikibase, including data and full text (reusable on other web text corpus, but the normalisation process requires expert scripting and time-consuming testing hence only useful on large scape and uniform corpus).
CKG Data: XML / DTD base framework for data import and export for Wikibase, and for data validation, portability, and distribution.
CKG Document Distribution (AKA re-publishing): Prototyped a document distribution query service using a knowledge graph.
CKG ER model: Entity-relationship model (beta) for a document corpus to support data scientists’ contributions as a LOD community layer to a literature corpus and for metadata distribution5
Info sheet: ClimateKG
ClimateKG is a climate science literature resource for use by the public, policymakers and scientists. ClimateKG is a community knowledge graph, it provides a Linked Open Data foundation for depositing climate literature and supports data enrichment contributions from the scientific community.
A knowledge graph is used to transform what was only a browsable website8 into a dynamic data resource to be able to: access documents, distribute metadata, and answer questions.
The knowledge graph is made from datasets and an entity-relationship model (ER model).
The AR6 reports on the web are changed into basic foundational datasets to show its syntactic (structural) parts: Report series, report, text division, chapter, glossary terms, acronyms, authors, and more. Then, connections are made between the parts using an ER model
.
Figure 3: How datasets9 are connected to one another using the ER model to form the knowledge graph.
Foundational datasets (corpus syntactical structure)
Corpus full text & structure – 7,524,958 Words, 2,153 Image files; 88 Chapters, 7 Reports
Bibliographic information – 95 DOIs
Glossary terms – 1,274
Acronyms – 1,910
Authors – 932
Entity-relationship model
ClimateKG connects these datasets using an ER model, forming syntactic basis of the knowledge graph that can then have further semantification added:
+- Corpus structure
+- 7 reports
+- REPORT
| <- Bibliographic information
| <- Glossary terms
| <- Acronyms
+- CHAPTER
<- Authors
<- Bibliographic information
<- Corpus full text
Figure 4: ClimateKG ER model (Note: This is a simplified version of the ER model)
Searching the knowledge graph
The connections that the knowledge graph creates allow for querying the data from the reports. As an example question:
'How many South American or Indian authors contributed to the report?'
The knowledge graph contains the answer, which can be computed and calculated dynamically rather than retrieved directly, to gather the relevant chapter texts:
'AR6 author distribution is 71 from South America and 43 from India.'
Figure 5: South American authors, Total: 71 authors (7.6% of all 932 unique authors), and; Figure 6: Authors from India, 43 authors (4.6% of all 932 unique authors)10
The ClimateKG platform and activities
Platforms: ClimateKG would have two websites. ClimateKG uses Wikibase software where the knowledge graph can be browsed and accessed via APIs and query interfaces. ClimateKG Data Bench uses data science tooling such as Quarto and Jupyter Notebooks as a work place to engage the community.
ClimateKG: Browse the full text, datasets, and community data.
ClimateKG Data Bench: Data analysis and visualisation – A platform that enables the community to enrich the corpus, analyse the contents, share results and use AI LLMs.
Activities: These are the ongoing work of ClimateKG and will use the platforms provided.
Document distribution: The data provides a map of the internal structure of the corpus documents, enabling any section or piece of data to be retrieved and delivered to the user.
Extended metadata: Questions can be answered quickly and reliably. Metadata is distributed to library systems and the Data Commons on Wikidata.
Citizen science: ClimateKG collaborates with Youth Data Champions interns from the #semanticClimate organisation on a global scale.
Data science community: Contributors can enrich the corpus while maintaining the integrity of the original documents.
Strategy and methods
The ambition of ClimateKG is to become a semantic web and LOD repository for climate science literature more widely. Semantic technologies have not yet been widely adopted by the publishing sector. Aware of the challenges this poses for community adoption, the service has kept "Keep It Short and Simple" (KISS) as its working mantra to support its ambition of growth. Examples of successful LOD infrastructures include the Protein Data Bank (PDB)11 and GenBank,12 a genome data bank, both of which kept their data models lean and simple to maximise contributions. To this end the project has learned from these examples hence the simplicity of the ER model is of key importance.
An early decision was made to use Wikibase as the software framework for the knowledge graph. TIB Open Science Lab has been using Wikibase for many years and the Wikibase4Rearch team working in NFDI4Culture and with Base4NFDI have extensive knowledge of the developing with Wikibase.
Wikibase is the open-source self-hosted version of the software that Wikidata is built on, the data infrastructure of Wikipedia, and is for, ‘managing and sharing structured data’ – Wikibase. Wikibase is used in tandem with MediaWiki, the wiki software used for Wikipedia.
The Computational Publishing Service (CPS), which is part of the Wikibase4Research project within the NFDI4Culture consortium, has several years of experience in importing complex databases into Wikibase and publishing them using the Wikibase publishing system. This is done using Python, Jupyter notebooks, and Quarto.
#semanticClimate has also worked with Wikidata and Wikibase Cloud and conducted a number of experiments using Wikidata and Wikipedia to enrich and add multilingual content to the IPCC Glossary.
Methods used
Agile software development and rapid prototyping
Open science values and methods
Literature review on the topic ‘Wikibase knowledge graphs’
KISS – keep it short and simple
Data modelling – bottom up and top down
IPCC layer – the need to keep data that is a direct representation of the IPCC reports distinct from community contributions
Ring fencing LLM use: Use to work with software and not to make software. Get LLM to build using Python, Notebooks, and create supporting documentation – so that the LLM creates a resource that then no longer needs LLM and can be run on its own.
Data quality control and portability
Mission, goals, and milestones

Figure 7: Publishing knowledge graph. Prototype CSS Paged Media generated from ClimateKG using Wikibase API. See working example13 and related GitHub issues.14
Mission
Provide greater access to the IPCC Sixth Assessment Report by building a Knowledge Graph (a network of all parts).
United Nations, Intergovernmental Panel on Climate Change, IPCC Sixth Assessment Report is made of 10,009 pages; has 932 authors; 48,400 references, and; 66,834 data links; 1,274 glossary terms; and 1,910 acronyms.
Apply Tim Berners-Lee vision of the Semantic Web to a large-scale corpus: Break down documents into smaller entities, identify, and connect the entities.
Make a publishing knowledge graph (document distribution): A knowledge graph for providing users with search results that could then carry out document retrieval and collate output to provide users with relevant parts to read.
Create a scalable open-source system made from best-in-class software
Target audiences
Policy makers
Scientists
Citizen scientists
Delivering on the mission
Access: Transforming the scattered unstructured corpus to adhere to FAIR Principles for a publication, adding the generated data to a knowledge graph, and making data deposits in Leibniz Data Manager (TBC) – and all in a working system has meant a pathway has been established for creating the type of access planned.
Publishing knowledge graph: A preliminary delivery on the goal of producing a ‘publishing knowledge graph’. This goal can be fully achieved and the ground work has been done as a proof of concept. CSS Paged Media demonstrations have been made and MediaWiki / Wikibase APIs allow for output to clean HTML and interoperable formats.
Create a scalable open-source system: Wikibase has proven to be a useful software suite for storing the knowledge graph data and gives important tooling for editing, versioning, reviewing and distribution. The limits are speed and supporting multiple data edits, but since a text corpus is a static object these speed issues are not detrimental. The project will augment Wikibase with more reliable editing systems in parallel.
Goals
Make search results available as multi-format publications
Enabling data analysis of report by providing as FAIR data
Structure full text of AR6
Collect necessary available data for AR6
Create a simple data structure foundation for others to be able to use to enrich with knowledge related metadata
Delivering on goals
Make search results available as multi-format publications: Prototyped. See working example and related GitHub issues.
Enabling data analysis of report by providing as FAIR data: Done. See: ClimateKG Data Bench
Structure full text of AR6: Main text section done. Front matter and back matter not done. See: ClimateKG
Collect necessary available data for AR6. Done. See: ClimateKG Data Bench
Create a simple data structure foundation for others to be able to use to enrich with knowledge descriptive metadata. Done. See: CKG ER Model
Milestones
Harvest report data sources
Validate tech stack and data models
Import report into Wikibase
Connect search and publish
Support data analysis community
Delivering milestones
Harvest report data sources: Done
Validate tech stack and data models: Done
Import report into Wikibase: Done
Connect search and publish: Prototypes and scoping
Support data analysis community: Done
Deliverables and innovations
Over the period of the project a fully working knowledge graph and supporting system has been able to be delivered and can go into a production phase and ready to be made public.
Deliverables
Platform: ClimateKG – Knowledge graph of AR6 using Wikibase.15
Platform: ClimateKG Data Bench – An online data science workplace using Python data analysis tooling, Jupyter Notebooks, Quarto, and AI LLM assisted coding (Claude).16
Software: CKG Scrape – A web scrape of the complete IPCC AR6 report into a knowledge graph built on MediaWiki and Wikibase, including data and full text (reusable on other web corpus but always with a customisation work package).17
Software: CKG Data – XML / DTD base framework for data import and export for Wikidata, and for data validation, portability, and distribution.18
Software / method: CKG ER model – Entity-relationship model (beta) for a document corpus to support data scientists’ contributions as a LOD community layer to a literature corpus.19
Prototyping: CKG Document Distribution (AKA re-publishing): Prototype a document distribution query service using a knowledge graph.
Further development work is needed to take this to be production ready.
Deployment: ClimateKG as a Docker DevOps deployment with: DEV, TEST, PROD, EXPERIMENTAL, and FORKING.
Data Deposit: AR6 full text, image assets, and data as academic deposit for Leibniz Data Manager (LDM) (to be confirmed).
Innovations
CKG Scrape: A web scrape of the complete IPCC AR6 report into a knowledge graph built on MediaWiki and Wikibase, including data and full text (reusable on any web corpus).
CKG Data: XML / DTD base framework for data import and export for Wikidata, and for data validation, portability, and distribution.
CKG Document Distribution (AKA re-publishing): Prototype a document distribution query service using a knowledge graph.
A prototype of the key components of a document distribution query service, capable of returning report sections as search results for collation into a shareable document. The document is rendered using CSS Paged Media and can include summary data tables and visualisations.
- CKG ER model: Entity-relationship model (beta) for a document corpus to support data scientists’ contributions as a LOD community layer to a literature corpus.
The IPCC Sixth Assessment Report: Background and quantification
What is the IPCC Sixth Assessment Report?
The Sixth Assessment Report (AR6)20 is a report collection made of seven reports with an associated Methodology Report21 which sits outside of AR6.

Figure 8: The seven reports that make up AR6: 1. SR15 (2018); 2. SRCCL (2019); 3. SROCC (2019); 4. WGI (2021); 5. WGII (2022); 6. WGIII (2022); 7. SYR (2023).
It has produced six assessment reports. The reports carry out a literature review of the world climate science and give recommendations to governments on action and policy.
The IPCC reports are among the most important sources of literature on climate science and policy.
The reports are comprehensive and use a framework of modelling 100 years into the future using potential climate pathway based on average surface temperature increases and a multilateral framework for the benefit of all mankind, present and future generations, with the goal of protecting humanity and ecosystems from catastrophic, irreversible disruptions.
This publication has been produced by the IPCC. The Intergovernmental Panel on Climate Change was formed in 1988 as a multilateral endeavour and 147 nations unanimously sign off on its Synthesis Report – Summary for Policy Makers,22 which then becomes legally binding agreements as part of the UN government membership and are vital to the Conference of the Parties (COP)23 meetings.
The Climate Knowledge Graph imports content from the seven reports that make up the IPCC Sixth Assessment Report (AR6) cycle. Of the seven reports, reports 1-6 are published by Cambridge University Press under Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 (CC BY-NC-ND 4.0).24 Report 7 (the Synthesis Report) is published by the IPCC directly with all rights reserved; short extracts may be reproduced with full source attribution. See also the IPCC copyright notice.25
The AR6 report series
Synthesis report
Special reports
Global Warming of 1.5°C: IPCC Special Report on Impacts of Global Warming of 1.5°C above Pre-industrial Levels in Context of Strengthening Response to Climate Change, Sustainable Development, and Efforts to Eradicate Poverty IPCC Special Report on Impacts of Global Warming of 1.5°C above Pre-industrial Levels in Context of Strengthening Response to Climate Change, Sustainable Development, and Efforts to Eradicate Poverty27
Climate Change and Land: IPCC Special Report on Climate Change, Desertification, Land Degradation, Sustainable Land Management, Food Security, and Greenhouse Gas Fluxes in Terrestrial Ecosystems28
The Ocean and Cryosphere in a Changing Climate: Special Report of the Intergovernmental Panel on Climate Change29
Working group reports
Climate Change 2021 – The Physical Science Basis: Working Group I Contribution to the Sixth Assessment Report of the Intergovernmental Panel on Climate Change30
Climate Change 2022 – Impacts, Adaptation and Vulnerability: Working Group II Contribution to the Sixth Assessment Report of the Intergovernmental Panel on Climate Change31
Climate Change 2022 – Mitigation of Climate Change: Working Group III Contribution to the Sixth Assessment Report of the Intergovernmental Panel on Climate Change32
Methodology Report: Sixth Assessment Report Cycle
2019 Refinement to the 2006 IPCC Guidelines for National Greenhouse Gas Inventories33
AR6 Quantification
The section attempts to outline the scale of the IPCC’s corpora and output. This is made difficult as records are not easily available, are incomplete, or are scattered across different systems, and that systematic quantification has not been done by others. The situation is compounded by the fact that only very limited open science practice has been implemented by IPCC publishing as opposed to data works of the IPCC where FAIR Principles implementation is widespread. Such Open Science practices would be the use of PIDs, academic repositories, schemas and ontologies (e.g., UN LOD schemas), interoperable formats, and open licensing, etc.
Quantification included are:
AR6 count
ClimateKG stats AR6 import
ClimateKG AR6 – Not imported
Assessment report back catalogue (AR1-AR5)
AR6 count
Following an initial quantification at the start of the project ‚A Scoping Quantification of the IPCC's Sixth Assessment Report for the Purpose of Using the Text Corpus in a Knowledge Graph’.34
The methods underlying the quantification are documented here: https://github.com/TIBHannover/climate-knowledge-graph/issues/1835
| Item | Count | Status |
| Number of main reports | 7* | Definitive |
| Open access license | 6 | Definitive |
| Report DOIs (top level) | 7 | Definitive |
| Chapter DOIs | 88 | Definitive |
| Page count | 10,009 | Definitive |
| Word count | 8,048,970 | Good |
| Authors | 932 | Good |
| References | 48,400 | Approx. ** |
| SYR Revisions | To be confirmed | To be confirmed |
| Figures | 1,678 | Approx. |
| Data | 66,834 | Approx. ** |
| Glossary terms | 1274 | Approx. |
| Acronyms | 1910 | Approx. |
| Index terms | 5463 | Approx. ** |
| Translations | See: Methods | NA |
| Wikidata entries (report top level) | 13 | Approx. |
| Wikidata entries – IPCC data types *** | To be confirmed | To be confirmed |
Table 1: Summary of IPCC Assessment Report Six quantification. Provisional figures. 2026 revision36
*Out of scope: Report updates to the “2019 Refinement to the 2006 IPCC Guidelines for National Greenhouse Gas Inventories” **Figure may contain duplicates ***IPCC data types, e.g.: Assessment cycles; Reports (sections); Chapters, Data; Citations; Contributors/Authors; Terms (Glossary, acronyms, index;); PIDs; and Qualifiers.
ClimateKG statistics
AR6 imported into ClimateKG
1 Report series (MediaWiki and Wikibase)
7 Reports (MediaWiki and Wikibase)
10 Text divisions (Chapter, Atlas, Cross-Chapter Sections, etc.) (Wikibase)
88 Chapters (MediaWiki and Wikibase)
2,153 Image files (No figure count available)
932 Authors (Wikibase)
1,274 Glossary terms (Wikibase)
1,910 Acronyms (Wikibase)
7,524,958 Words
To see how and what has been connected from the above in the knowledge graph see: (Report section: Knowledge Graph Construction Using Wikibase and MediaWiki / Data in Wikibase / Entity-Relationship Model (ER model))
AR6 for later import into ClimateKG
The front matter, back matter, and covers were not imported into Wikibase at present simply to save on time allocation to the import process. These sections would have different typesetting considerations and the view was taken that this could be done at a later stage.
Additionally, it was seen that many of these parts could be generated from the content gathered in the first phase for the text section, for example: tables of contents, author lists, and many standard appendix parts.
The remaining content is equivalent to 21 more chapters. 88 chapters were imported in phase 1 — which is all of the IPCC AR6 text section. The calculation excluded importing translations which would be an additional job and would be done if resources permit and the Methodology report which is not fully a part of AR6.
Reports (missing: 1: Chapters/Sections 50)
- Methodology
Report sections
Front matter (Missing: 7), including:
Cover
Editors overview
Masthead
Foreword
Preface
In memoriam / Dedication (only in some cases)
Text (Missing: 5)
- Supplementary material
Back matter (Annexes) (Missing: 7)
Glossary
Observational Products
Models
Tables of Historical and Projected Well-mixed Greenhouse Gas Mixing Ratios and Effective Radiative Forcing of All Climate Forcers
Modes of Variability
Monsoons
Annex VI Climatic Impact-driver and Extreme Indices.
Acronyms, Chemical Symbols and Scientific Units
Definitions, Units and Conventions
Scenarios and Modelling Methods
Global to Regional Atlas
Contributors
Expert Reviewers
Publications by the Intergovernmental Panel on Climate Change
Index
Back cover
Other
Reviews
Trickle back (Errata)
Data sets
Citations / References
Translations
To be imported: Assessment Report back catalogue (AR1-5, 1990-2014)
All IPCC Assessment Reports. Some only exist in print and would need OCR scanning and many are not open licensed.
See all reports (AR1 to AR5) AR1-5 bibliography: https://www.zotero.org/groups/2437020/semanticclimate/collections/SNSLIQWV/collection37
IPCC Assessment Report Chronology
The IPCC has completed five earlier assessment cycles since 1990. Each cycle produces reports from three working groups, a synthesis report, and other supporting reports.
| Cycle | Year(s) | Working Group I | Working Group II | Working Group III | Synthesis |
| AR1 | 1990 | The Scientific Assessment38 | The Impacts Assessment39 | The Response Strategies40 | — |
| AR2 | 1995–96 | The Science of Climate Change41 | Impacts, Adaptations and Mitigation42 | Economic and Social Dimensions43 | Second Assessment Synthesis44 |
| AR3 | 2001 | The Scientific Basis45 | Impacts, Adaptation, and Vulnerability46 | Mitigation47 | Synthesis Report48 |
| AR4 | 2007 | The Physical Science Basis49 | Impacts, Adaptation and Vulnerability50 | Mitigation of Climate Change51 | Synthesis Report52 |
| AR5 | 2013–14 | The Physical Science Basis53 | Impacts, Adaptation, and Vulnerability54 | Mitigation of Climate Change55 | Synthesis Report56 |
Table 2: Information and links to main parts of five earlier IPCC assessment reports
Knowledge graph design
Figure 9: Overall ClimateKG schematic
Summary
Wikibase and MediaWiki have proven to be helpful to the project for the purpose of moving from a semi-automated web-scrape to structured data as both the full text and data store can both be manually inspected and checked, as well as being able to run automated data integrity checks using the API and SPARQL interfaces.
Cleaning of the web scrape involved the greatest effort for the project workload, which all happened before the import into MediaWiki / Wikibase. This is the part of the project where new software was made: CKG Scape.
Wikibase will be used in read mode and continue as the source of truth data storage for ClimateKG. But other data editing and storage systems will need to be added for taking community contributions as the Wikibase system has very slow import performance and is error prone when multiple edits are happening, see: Conclusions section. In part these problems have already been addressed with the use of XML/DTD and having at present the five foundational datasets as layers of sorts independent from Wikibase. Also, because there is a need for data integrity and audit-trail to ensure data quality of IPCC data and community contributions, having separate firewalled data stores in a system is a benefit.
Rapid scoping literature review
Literature review search question: “How to build a research knowledge graph with Wikibase”
A number of research papers that specifically reported on knowledge graph construction using Wikibase were consulted for the development of ClimateKG knowledge graph. These papers broke down the development build process and made recommendations on successes and challenges faced by research teams.
Key learning points
No existing Wikibase knowledge graph construction could be directly reused.
Types of knowledge graph:
Community knowledge graph;
Mass import knowledge graph;
Wikidata contributing knowledge graphs.
Bringing together non-technical academic groups and data engineer groups on one platform was an advantage of Wikibase
Use of SHACL (Shapes Constraint Language) validation types
Takeaways
Takeaway: That ClimateKG is a ‘community knowledge graph’ as well as ‘publishing knowledge graph’
Takeaway: Contributing to Wikidata is not a simple task for ClimateKG as IPCC data on Wikidata is too fragmented and inconsistent
Takeaway: Wikidata data models are community built and can only be used as loose guidance as most often only periodically maintained or reviewed.
Research projects using Wikibase profiled in papers
Disability Wiki – https://disabilitywiki.org/57
EU Knowledge Graph – https://linkedopendata.eu/58
Enslaved – https://enslaved.org/59
RaiseWikibase – https://ub-mannheim.github.io/RaiseWikibase/60
- Monumenta Linguae Vasconum (MLV): Linking Historical Corpus Data and Annotations Using Wikibase, 1737 Basque manuscript – https://monumenta.wikibase.cloud/61
See: Appendix I: Wikibase knowledge graph projects and papers
The construction process
The knowledge graph has to hold five data sets and make available the full text and image assets of the corpus. The knowledge graph acts as a framework for community contributions to enrich the corpus.
Figure 10: ClimateKG datasets and data62
Five datasets – with the Corpus structure being the hub:
Corpus full text and structure
Bibliographic information
Glossary terms
Acronyms
Authors
Content:
Full text
Image assets
Wikibase holds the knowledge graph and MediaWiki holds the full text and image assets.
The following technical challenges and requirements had to be worked out:
Automating web scrape to content and data import
Create as simple as possible ER model and Wikibase data model – this acts as fixed, immutable foundation of the knowledge graph
Reconnecting content and knowledge graph so that data reports could be written back to content, e.g., list of authors, acronyms, glossary terms
Show provenance, differentiate IPCC data from community contribution
Have versioning
Have API / SPARQL interfaces
All publishing from knowledge graph
Allow community contributions and data analysis and use
Create data dumps and academic repository deposits
Deliver knowledge graph with a development environment
Prior work
Previously knowledge graph construction has been carried out in two contexts, in NFDI4Cuture and by #semanticClimate.
For NFDI4Culture a Flexfunds project was funded to create a software module ‘Wikibase to Jupyter Notebook’ (WB2JN) for making a knowledge graph of Baroque painting of Barocke Deckenmalerei in Deutschland (CbDD). This involved a much more complex web-scrape of a database-generated website and recreating the art collection in a Wikibase instance. The project gave experience of key areas: Web scraping, Wikibase, Jupyter Notebooks, and Quarto for data science publishing. WB2JN: https://github.com/NFDI4Culture/WB2JN63
#semanticClimate had carried out many rounds of experimentation and development with IPCC report materials, from taking part in UN Data hackathons, linked open data modelling in Wikibase Cloud, working on full text web scrapes to glossary and acronym TDM, running multiple rounds of internship data science training using IPCC reports, and taking place in workshops and citizen science outreach projects. The experience gained here was multi-layered, but the most important aspect was developing an awareness of the extant works of the IPCC are located, as print and digital, and its organisational editorial setup.
Methods
A scoping exercise: This project has aimed to "break new ground," serving as a scoping exercise for the engineering problems it needs to solve.
Agile and iterative: Each component is built to a level to maintain data quality control but the scale of components has been kept to a minimum and later refactoring planned into the process. There has also been a built-in awareness of the need for data portability in case software is thought to be insufficient and needs to be removed and exchanged.
Data modelling: The data modelling approach has been to build bottom up and then apply a top-down approach, and to have KISS as the key principle with an economic approach to the data model.
The entity-relationship model (ER model)
The ER model has undergone several rounds of refinement and is not yet in a fully fixed state. See Git issues #51 and #228.
The ER model has emerged as being centrally important to the project and gained more importance than initially thought. Its importance is twofold:
Support community enrichment layer, and
Distribution of metadata in Open Science infrastructures.
ER model and Wikidata data model documentation automation workflow
Data modelling documentation is now automated. It became necessary to ensure that the ER model and Wikibase data modelling was generated directly from data sources as the resources are becoming too large to manage.
A system based on mapping the XML/DTD resources64 for the five data sets and querying Wikibase directly could generate the Mermaid (similar to GraphViz) diagrams and other needed data.
This directory contains the artefacts for mapping the IPCC AR6 XML/DTD data model to the local Wikibase instance, generating both a flat CSV property table and a Mermaid entity-relationship diagram.65
Info link and Git directory.
Schematic of automation workflow:
DTD files (parent data-xml-dtd/)
│
▼
[Manual step] SPARQL query against local Wikibase
http://localhost:9999/bigdata/namespace/wdq/sparql
→ retrieve all Properties (P1–P33) and class Items (Q-IDs)
│
▼
erm-wikibase-mapping.xml ◄── edit this file to update the mapping
│
├──[erm-to-csv.xslt]──────► erm-mapping.csv
│
└──[erm-to-mermaid.xslt]──► er-diagram-wikibase.mmd
Figure 11: XML / DTD and Wikibase data mapping schematic
The ER model diagram (simplified)
Simple version
WORK <──── Corpus hierarchy
└── REPORT_SERIES
└── REPORT <──── Bibliographic information
| <──── Glossary
| <──── Acronyms
└── TEXT_DIVISION
└── CHAPTER <──── Authors
<──── Bibliographic information
<──── Corpus full text
Figure 12: ClimateKG ER model, simple version
ER model: Full version

Figure 13: ClimateKG, Entity-Relationship Model (ER model). See scalable version online.
Wikibase data model
Wikibase comes as a clean slate with no pre-existing data models of real world entities, although it does have a data model for data,66 and ones that come from Wikidata67 for describing real world entities but these are changeable and evolve. To create data entries the data model needs to be created or loaded by the user. Below are the Wikidata Properties (P numbers) and Item IDs (QIDs) ClimateKG has created. Classes are created using the Item QIDs.
See: Appendix II: Wikibase properties
ER model and schemas
Schemas for effective use of the ER model for metadata distribution and document delivery have not yet been resolved. A TIB and partner consultation will be a primary part of further work discussed in the conclusion and finding section of this report. Schemas have not yet been fixed. First, a fuller picture is needed of the data, the systems in use, and the specific data targets tied to project goals. This groundwork is necessary before TIB experts can be consulted for meaningful feedback. From this TIB Innovation fund phase of the project this complete picture is now available where the bottom-up data modelling approach has been taken.
Schemas, including ontologies, taxonomies, SKOS, controlled vocabularies and terminology services enable the interoperability and machine readability of the reports, and play a role in all aspects of the project.
The project design mantra of KISS comes into play here to engineer schema use around allowing other systems to access ClimateKG data.
Example questions are:
Discovery: For extending bibliographic metadata records of library catalogues, what ontologies and metadata schemas should be used?
Document delivery: ClimateKG holds full text copies of the reports in MediaWiki. MediaWiki API can be enabled to supply documents in common open formats: HMTL, JATS, BITS, etc. What schemas and formats should be used to distribute full text to academic repositories and document systems?
Schema scoping work
Currently schema use has been divided into four sections:
Internal Document Structure
Citations & Reference Context
Subject Content & Semantic Annotations
Open Science Packaging & Provenance
| Category | Standard / Ontology | Acronym | Description |
| 1. Internal Document Structure | Document Components Ontology | DoCO | Model’s structural elements like sections, paragraphs, figures, tables, and lists. |
| Discourse Element Ontology | DEO | Defines rhetorical sections of scientific discourse (e.g., Introduction, Methods, Results, Discussion). | |
| Pattern Ontology / Collections Ontology | PO / CO | Maintains ordered sequences and structural collections within documents. | |
| JATS-RDF / JATS4R | JATS-RDF | Translates publisher NISO JATS XML structures directly into RDF graphs. | |
| 2. Citations & Reference Context | Citation Typing Ontology | CiTO | Captures semantic intent/motivation of citations (e.g., usesMethodIn, citesAsDataSource). |
| Bibliographic Reference Ontology | BiRO | Model’s bibliographic references and reference lists within publications. | |
| Citation Counting & Context Characterization | C4O | Links in-text citation pointers to exact textual snippets and reference items. | |
| FRBR-aligned Bibliographic Ontology | FaBiO | Describes bibliographic works, expressions, manifestations, and publication types. | |
| BIBFRAME / LRMoo | BIBFRAME | Integration standard for library management systems and national union catalogues. | |
| Dublin Core Metadata Terms | DCTER models | Provides core baseline metadata fields (title, creator, publisher, issue date, license). | |
| Schema.org Vocabulary | Schema.org | Facilitates web discovery, search engine indexing, and Google Scholar parsing. | |
| 4. Subject Content & Semantic Annotations | W3C Web Annotation Data Model | W3C WADM | Standard for linking semantic annotations/subjects to specific text spans or targets. |
| Simple Knowledge Organisation System | SKOS | Links subjects to controlled vocabularies and thesauri (GND, LCSH, MeSH, AGROVOC). | |
| Open Research Knowledge Graph / Wikidata | ORKG / Wikidata | Connects specific research statements, data, methods, and entities to global knowledge graphs. | |
| 5. Open Science Packaging & Provenance | Research Object Crate | RO-Crate | JSON-LD packaging format bundling metadata, datasets, code, and document structure. |
| W3C Provenance Ontology | PROV-O | Tracks creation, transformation, and extraction provenance (e.g., manual vs. NLP extraction). |
Table 3: Example schemas and ontologies. Based on a Google Gemini questions: https://share.gemini.google/eo4quclJEnOO
Knowledge graph implementation
Data management
Data from ClimateKG will be deposited with Leibniz Data Manager (LDM) including source data, output data, and Jupyter Notebooks.
ClimateKG has developed an XML workflow for deposing key data sets, and Wikibase data dumps and RDF outputs will be made.
Data provenance has been recorded, including source addresses, source data copies, licensing, copyright notices, and access data. All of this is stored in the knowledge graph using Wikibase properties.
The knowledge graph needs to differentiate between IPCC data and community-contributed data, and this distinction will be recorded accordingly.
Data modelling: It was necessary to have a concrete system, data, and ER model / Wikidata data model to be able to pose questions about its viability and what is the best way to take the design forward. The data and content are made as portable units allowing for redesign and refactoring.
The build process

Figure 14: ClimateKG web scrape process
The process is where the initial plan had to adapt to what was discovered as the project moved forwards.
The web scrape (semi-automated and scripted).
The scrape:
This took up the vast majority of the project time – 50%.
To start with, a decision needed to be made for which corpus source to use: web or PDF. Both have problems and there is no easy way to weigh up the difference without committing to a workflow. The web was selected as the source as on balance PDF text extraction is more error prone due to text lines not being differentiated such as figure captions, table captions, two column texts, etc., and these being merged with body content making it unreadable. But the decision poses an empirical dilemma as to which is the better route as it cannot be known unless it is tried.
ArchiveBox and then plain WGET which ArchiveBox was based on was used for scraping.
A list of URLs provides page content to be scraped to a web server as HTML.
The scrape generated the full text, assets, and some metadata.
Full text:
The IPCC-source websites used two CMSs (WordPress and Gatsby) each with its own HTML quirks.
The sources were simple directories with text so the scrape itself was non-technical and didn’t require the complexities of Selenium and could instead use WGET as used by ArchiveBox. But files were very large and had to deal with issues resulting from server response time delays in delivering content and affected later processing.
Full text cleaning and transformation:
Custom scripts were needed to remove sections from the HTML DOM, for normalising markup. Such scripts need to be custom made per corpus.
Pandoc was used for a second round of cleaning and to transform the HTML to Wikitext.
Assets:
- Image assets were captured and assigned UIDs to then reassociate them with texts.
Metadata:
Metadata was included in the capture for source referencing, such as associated PDFs, licences, source reference URLs, etc.
Special metadata had to be gathered: The corpus hierarchy going down from Work to Chapter:
Work (A creative work)
Report series
Report
Text division
Chapter
MediaWiki import – Full text and assets:
- Full text and assets are first imported into MediaWiki as Pages and Files respectively and on the same instance as the Wikibase install – this uses the custom-made Python software CKG Scrape. This process is run before the Wikibase import as the MediaWiki URLs are needed first.
Additional metadata:
Additional metadata is added to the metadata gathered from the scrape, these were from OpenAlex, CrossRef, and other bibliographic data, etc.
All the metadata was combined in one source CSV file and online preview for user review, including MediaWiki URLs.
Wikibase import – Full text URLs, corpus hierarchy, and metadata:
A custom process was put in place for import after reviewing the various metadata markup languages: JSON, YAML, and TOML, etc. There was a challenge to move from the flat CSV structure to a nested hierarchy of linked open data. An XML / DTD combination was decided on as the data could be validated before import using the DTD and users could review data in a way that made sense using XML editors.
The import scripts populate Wikibase, generating the Properties and Statements needed for the import based on a pre-designed data model.
The import can generate a complete Wikibase data model and data from scratch or make updates.
The design decision of XML / DTD and using WikibaseIntegrator Python package has proven effective and useful for the other data imports and for providing data portability.
Connecting the full text to the data representation (MediaWiki to Wikibase):
- MediaWiki uses a feature called Sitelinks to assign a unique Wikibase item to a MediaWiki page enforcing a 1:1 mapping rule per site at any one time. This way both entities know to use each other's data.
Importing other data:
Four additional data sets have been imported into Wikibase:
Bibliographic information
Glossary terms
Acronyms
Authors
The same XML / DTD method has been used for all data sets.
ClimateKG Data Bench:
Jupyter Notebooks, with Python data analysis and data visualisation tooling are used and published and shared using Quarto.
The Data Bench makes full use of AI LLM with the Notebooks and Quarto.
Output Notebooks and data deposited with Leibniz Data Manager (LDM).
Data dumps:
Data: The XML / DTD method has been used for all data sets.
Full text and image assets: MediaWiki has built in exports, XSLT is used for providing any file format changes needed.
Dumps are Git versioned and made as releases.
Data is deposited in Leibniz Data Manager (LDM).
Academic data repository deposit:
Uses data dumps.
A Jupyter Notebook with supporting guidance is supplied.
Publishing Knowledge Graph:
Using APIs and SPARQL interfaces the knowledge graph can be queried and data visualisations, data reports, and full text segments can be retrieved and made available for ‘document provision’.
CSS Paged Media has been used to create typeset ‘document delivery’ packages.
This part has been partially realised as proof of concept prototypes.
Community contributions:
Data scientists.
Citizen Science activities.
Policy makers.
The results
MediaWiki and Wikibase in combination provide a robust and feature rich environment for knowledge graph creations, maintenance, and as user facing service development. With workflows now established for import, export, and the data model, the project can begin operating the knowledge graph.
The challenge, which did not appear to have been resolved by the community, was the import, maintenance, and then export of data and content. As an open-source project ClimateKG will share this corpus structuring knowledge.
Software System Architecture
The overall challenge of the project is a data engineering one of Extract, Transform, Load (ETL) and bringing together unstructured sources, not in academic data repositories or using FAIR in a publishing context, into a central location — Wikibase and MediaWiki. The report text and assets are on websites as is much of the publishing data: Authors, Acronym, Glossary. With bibliographic data coming from DOI sources.
Software
https://github.com/TIBHannover/climate-knowledge-graph
CKG Scrape (Software): A web scrape of the complete IPCC AR6 report into a knowledge graph built on MediaWiki and Wikibase, including data and full text (reusable on any web corpus);
CKG Data: XML / DTD base framework for data import and export for Wikidata, and for data validation, portability, and distribution;
CKG Document Distribution (AKA re-publishing): Prototype a document distribution query service using a knowledge graph, and;
CKG ER model: Entity-relationship model (beta) for a document corpus to support data scientists’ contributions as a LOD community layer to a literature corpus.
CKG Scrape involved the greatest effort and is responsible for normalising the corpus text and delivering structured content, full text and linked open data, to MediaWiki and Wikibase. CKG Scrape can be used on other web corpora, but due to the custom nature of web corpora it needs an expert team to analyse and script the ETL process.
Architecture
All of the figures below are built in Mermaid and can be viewed online with the links below.
Slide link: https://tib.eu/j7dx
Diagrams: https://github.com/TIBHannover/climate-knowledge-graph/issues/34
Slide link: Request for Comment (RfC) 6.10.2025 Simon Worthington & Laura Oldenbourg (2025). Climate Knowledge Graph (Version 0.0.1) [Computer software]. https://github.com/TIBHannover/climate-knowledge-graph
The software system – Overview

Figure 15: Project overview, from semi-structured corpus to structured corpus / knowledge graph and its usage
The project is service oriented. We structure large document corpus and make Knowledge Graphs and Linked Open Data usable as services.
Harvest
Knowledge Graph
Usage: Browse, data analysis, publish (Document Distribution)
IPCC AR6 report: 10,000 pages as PDF/web, plus supporting data
Software used
Figure 16: Overview over the software elements to complete for the project
See: Appendix III: Software architecture – Continued
Recommendation for IPCC publishing: Adopt FAIR for publishing!
A technical challenge for ClimateKG has been that IPCC publishing utilises little of available modern Open Science infrastructures. The IPCC produces publications to the highest possible scientific standards, but by publishing as PDF and websites only the result is that the work is not fully visible and usable in modern digital scholarly communications systems.
The data science work of the IPCC as opposed to the publishing work has adopted Open Science practices such as the use of the FAIR Principles. For example, data is deposited in academic repositories, uses interoperable formats, and is versioned on GitHub, etc. As a specific example, Working Group I use GitHub to link report figures and their supporting data.
In summary, published work should be thought of as a type of data and accordingly have the FAIR Principles applied. The benefits for the IPCC mission would be to widen and deepen its reach.
The recommendations are based on maintaining current DOCX word processing workflows with seamless integration so as not to disrupt or burden editorial or production workflows.
Recommendations uses of FAIR Principles for publishing
The FAIR Principles are intended to make research data readable by humans and machines.
Findable: Ensure data can be located, e.g., PIDs, indexed metadata.
Accessible: Define clear access pathways, e.g., Standard APIs, explicit access control.
Interoperable: Enable data integration, e.g., Shared ontologies, RDF triples, schemas.
Reusable: Ensure long-term value & clarity, e.g., Explicit licensing, institutional deposits.
Using Open Science infrastructures would be part of supporting FAIR Principles and are considered normal Open Access practice for working alongside a conventional publisher.
ClimateKG recommendations
Open Access: Open licence all parts: Full text, figures, and publishing data.
Open standards for interoperability: Ensure all material is available in validated interoperable formats with declared open standards.
Persistent Identifiers (PIDs): Use PIDs to connect all parts as linked open data: Publication parts; full text; figures; authors; glossary, acronyms, data, references, institutions, and events.
Academic repositories deposits: Make deposits of published full text and data in academic repositories.
Versioned releases: Use the academic repository for versioning publication and data updates.
Open citations: Use open citations methods, e.g., I4OC Principles, and deposit citation information in a data repository, structure, and open licence for scientometric use.
Terminology services: Use terminology services for glossary, acronyms, and indexes, etc.
Linked Open Data: Make a linked open data representation of the complete corpus of IPCC reports so that each report has all of its parts linked: Publications, data, authors, etc. Have this LOD representation make use of taxonomies and ontologies from the UN and IPCC Data Distribution Centre (DDC), e.g., SDGs, AGROVOC, SDGIO , ICH, etc.
IPCC qualifiers: Release IPCC qualifiers as a controlled vocabulary, such as: Evidence levels; agreement levels; confidence levels; likelihood terms (probabilistic scale); Scenarios, Emissions & Mitigation Pathways, and; Climate Models, etc., and map to the reports.
Conclusion and findings
The project has been able to deliver on making the envisaged knowledge graph ClimateKG of AR6 using Wikibase and MediaWiki.
The KISS principle — "Keep It Short, and Simple" was taken on board as a guiding principle to drive community engagement. The idea manifested in three ways for the project, firstly with the ER model, secondly being able to keep the ER model automatically up to date with data in the system using an XML / DTD datasets process, and thirdly in needing clear communication about what and how the knowledge graph works.
But it was a surprise to see the strategy work out in implementation. As the first two aspects, the ER model and the XML / DTD datasets process were not on the roadmap of the project, and instead they emerged over time and were only fully realised at the end of the project.
The ambition of the ER model is to provide a stable syntactical foundation of the corpus and for other community data science work to add semantic enrichment.
This approach and various Wikibase technical challenges also led to a technical implementation as well as an XML / DTD method for holding the dataset in conjunction with the ER model. The XML allows for nested LOD data relations and at the same time provides a human browsable and portable file and the DTD enforced the ER model through validation. The XML / DTD dataset packages function as an operational data management system, acting as the gateway for data moving in and out of Wikibase. As a result, the static XML files stay synchronised with the ER model and with Wikibase itself, along with their associated documentation.
The added benefit of working through the design of the ER model and XML / DTD process was that it was an opportunity to communicate in plain language a description of the knowledge graph process, which in its simplest form came down to the following:
Datasets + ER model = Knowledge graph
Wikibase and MediaWiki
The use of Wikimedia Foundation community, platforms, and tooling has been central to the project and has many benefits that suit a knowledge graph for a text corpus that is stable and stationary like the IPCC AR6.
Three four areas stand out:
Document distribution and info boxes: The connection between Wikibase and MediaWiki using the Sitelinks feature means full text can be delivered from SPARQL queries. Sitelinks make a unique link from a Wikibase item to a MediaWiki page, and the MediaWiki page has an API that can deliver full text in common open standards: HTML and XML, etc. And in reverse Sitelinks allow MediaWiki pages to have info boxes. Info boxes are the summary boxes usually found on Wikiversity pages, e.g., IPCC Sixth Assessment Report. The same technology can be used but with the data ClimateKG has collected, for example: List of authors, links to key concepts or IPCC qualifiers, or bibliographic data, etc.
Wikidata: Using Wikibase means that distributing extended metadata of the corpus to Wikidata, one of the world's largest open knowledge graphs, is a viable proposition.
Automated text corpus structuring: Automating the structuring of a corpus sourced from the web or PDFs can only go so far. Transferring the content into MediaWiki and Wikibase allows the remaining work to be completed through manual and AI-assisted edits. The key benefit for these interventions is that the systems are versioned and can show data provenance with references, access dates, and user information. This is because of the wealth of manual collaborative editing tools that are at the heart of the Wikimedia Foundation tooling. And on the data side, non-experts can review, browse, and edit data in the knowledge graph.
Wikibase limitations: Computational speed of Wikibase is extremely low and this leads to data errors and an unstable environment if multiple entries are being made simultaneously. Since the AR6 corpus is static fixed work, this is not a problem, but it prevents Wikibase use as a directly editable community data store. For ClimateKG, Wikibase would be used in a read only mode. A parallel data store would be used for moderating and managing community updates and only after moderation would these updates be added to the Wikibase instance.
ClimateKG Data Bench
A community platform ‘ClimateKG Data Bench’ has been created as a data science workplace using Python data analysis tooling, and including: Jupyter Notebooks, Quarto, and AI LLM coding assistants. Using these three frameworks with AI LLM to assist the knowledge graph can quickly be put to work and questions asked of ClimateKG that can be backed up with data that has verifiable sources from the IPCC.
Follow up work and next steps
ClimateKG has three operational areas that need to find funding or resourcing for their support.
Implement FAIR Principles for AR6.
Youth citizen science partnerships.
Knowledge graphs as a service for text corpora.
1. Implement FAIR Principles for AR6
The rollout of FAIR Principles for AR6 and IPCC publishing in general.
Realising benefits of FAIR Principles for publishing for report reach and efficacy for the following audiences and how to measure success:
Policy makers and government;
scientists, and;
the public via citizen science activities.
Steps to be taken for disseminating reports:
Populate Wikidata.
Populate global library catalogues.
Full text, assets, and LOD in research repositories.
These areas were left out of scope for the TIB Innovation fund research round:
Complete open citation database.
Complete data repository connection to report.
Map UN LOD schemas to report e.g., UN Sustainable Development Goals (SDG) Ontology, UNESCO Intangible Cultural Heritage (ICH), AGROVOC (Food and Agriculture Organization) Schema, etc.
2. Youth citizen science partnerships
With the #semanticClimate community integrating ClimateKG into the global youth citizen science programme — which is already engaged with young scientists from high school age in: Europe, India, and South America (Chile and Argentina).
#semanticClimate has two programme tracks to be integrated with ClimateKG:
First is the ‘Climate Champions’, these are youth events for learning about data science and climate. #semanticClimate runs an internship programme and educational hackathons. Participants can add their work to the community layer of ClimateKG.
Second is Climate Chatbot: Funded by Open Knowledge Foundation #semanticClimate has released a climate AI LLM chatbot. ClimateKG data can be used in the chatbot.
3. Supporting text corpus knowledge graphs
ClimateKG is scalable and can be used for any text corpus. The challenge is to move from a proof of concept to an active service and integration across TIB services.
ClimateKG has been born of TIB R&D culture — of innovation of the modern research library services and data science. Now that the idea for climate literature of turning publications into usable data has been realised as a working system, at least in a pilot phase, further meaningful and concrete consultation can take place with existing TIB services for how to enable knowledge graph services for text corpora in general. For example, with: TIB Terminology service, ORKG (Open Research Knowledge Graph), TIB Kat, TIB Open Publishing, and Research Data Management (RDM), etc.
Appendices
Appendix I: Wikibase knowledge graph papers
Appendix II: Wikibase properties
Appendix III: Software architecture – Continued
Appendix IV: AI LLM notice
Appendix I: Wikibase knowledge graph projects and papers
Disability Wiki
El Morr, Christo, Pierre Maret, Fabrice Muhlenbach, et al. 2021. ‘A Virtual Community for Disability Advocacy: Development of a Searchable Artificial Intelligence–Supported Platform’. JMIR Formative Research 5 (11): e33335. https://doi.org/10.2196/33335.
EU Knowledge Graph
Diefenbach, Dennis, Max De Wilde, and Samantha Alipio. 2021. ‘Wikibase as an Infrastructure for Knowledge Graphs: The EU Knowledge Graph’. In The Semantic Web – ISWC 2021, ↡↡edited by Andreas Hotho, Eva Blomqvist, Stefan Dietze, et al. Springer International Publishing. https://doi.org/10.1007/978-3-030-88361-4_37.
Enslaved
Shimizu, Cogan, Andrew Eells, Seila Gonzalez, et al. 2024. ‘Ontology Design Facilitating Wikibase Integration — and a Worked Example for Historical Data’. Journal of Web Semantics 82 (October): 100823. https://doi.org/10.1016/j.websem.2024.100823.
Shimizu, Cogan, Pascal Hitzler, Seila Gonzalez-Estrecha, et al. 2023. ‘The Wikibase Approach to the Enslaved.Org Hub Knowledge Graph’. In The Semantic Web – ISWC 2023, edited by Terry R. Payne, Valentina Presutti, Guilin Qi, et al., vol. 14266. Lecture Notes in Computer Science. Springer Nature Switzerland. https://doi.org/10.1007/978-3-031-47243-5_23.
RaiseWikibase
Shigapov, Renat, Jörg Mechnich, and Irene Schumm. 2021. ‘RaiseWikibase: Fast Inserts into the BERD Instance’. In The Semantic Web: ESWC 2021 Satellite Events, edited by Ruben Verborgh, Anastasia Dimou, Aidan Hogan, et al., vol. 12739. Lecture Notes in Computer Science. Springer International Publishing. https://doi.org/10.1007/978-3-030-80418-3_11.
Linking Historical Corpus Data and Annotations Using Wikibase (1737 Basque manuscript)
Lindemann, David, and Mikel Alonso. 2024. Linking Historical Corpus Data and Annotations Using Wikibase. https://euralex.org/wp-content/themes/euralex/proceedings/Euralex%202024/EURALEX2024_Pr_p785-791_Lindemann-Alonso.pdf.pdf.
Appendix II: Wikibase properties
| Wikibase property reference (P1–P33) | ||
| PID | Label | Type |
| P1 | Instance of | WikibaseItem |
| P3 | Part of | WikibaseItem |
| P5 | Wiki | URL |
| P6 | Source | URL |
| P7 | URL | |
| P8 | Date | String |
| P9 | OpenAlex | String |
| P10 | DOI | String |
| P11 | License | String |
| P12 | Has tag | WikibaseItem |
| P13 | Definition | Monolingualtext |
| P17 | date accessed | Time |
| P19 | source version | String |
| P20 | ClimateKG Author ID | String |
| P21 | last name | String |
| P22 | first name | String |
| P23 | gender | String |
| P24 | citizenship | String |
| P25 | country of residence | String |
| P26 | affiliation | String |
| P27 | contributed to chapter | WikibaseItem |
| P28 | role | String |
| P29 | Publisher | String |
| P30 | ISBN Electronic | String |
| P31 | ISBN Print | String |
| P32 | Licence URL | URL |
| P33 | Abstract | String |
Table AII 1: ClimateKG Wikibase properties
Wikibase class items (QIDs)
| QID | Label | Also Known As | Example item |
| Q1 | Glossary Term | Term; Category; Tag | Q1005 |
| Q2 | Work | Q7 | |
| Q3 | Report Series | Monographic Series | Q7 |
| Q4 | Report | Book; Monograph; Volume | Q10 |
| Q5 | Text Division | Division | Q110 |
| Q6 | Chapter | Q154 | |
| Q2087 | Acronym | Q3880 | |
| Q3998 | Author | Q4047 |
Table AII 2: ClimateKG Wikibase class items (QIDs)
Appendix III: Software architecture – Continued
Main Workflow Overview

Figure AIII 1: Main Workflow overview
Semi-structured corpus (sources)
Data harvest
Structured corpus (knowledge graph and data store)
Use of knowledge graph
Browse
Data analysis
Search, review, publish
Sources (semi-structured corpus)** **
Figure AIII 2: Sources
Data harvest 
Figure AIII 3: Data harvest process
Data modelling: Methods and rationale
Structure: Syntactic and/or semantic
Intermediate storage and data processing
Structured corpus
Figure AIII 4: Structured corpus and its storage options
Store text, media and source links (MediaWiki)
Create knowledge graph (network of all parts) (Wikibase)
Browsing
Figure AIII 5: Browsing the knowledge graph
The full text can be browsed in MediaWiki with the knowledge graph from Wikibase able to display data using MediaWiki Lua templates via MediaWiki’s Sitelinks feature. Additionally, the data can be browsed via Wikibase which gives an easy to view data presentation.
Data analysis
Figure AIII 6: Data analysis from the knowledge graph
Usage/export
Annotation
Enrichment
Search, review, publish: Search
Figure AIII 7: Searching the knowledge graph
Search, review, publish: Review
Figure AIII 8: Reviewing process
Search, review, publish: Publish
Figure AIII 9: Publishing process
Appendix IV: AI LLM notice
The method used:
Not used in software: AI LLM are not used in writing software and only used to execute software, to make analysis and offer options, and for documentation.
AI writes itself out of the loop: Ask, Plan, Execute. Have Claude document processes, save scripts. And where possible have Claude write itself out of a workflow — by writing scripts or Jupyter Notebooks that can be run without needing LLMs. This is done for reliability and cost.
System used:
DeepL text optimisation, with edits for Final Report 2026.
Copilot/VSCode/Claude used as an assistant on the following:
Docker deployment
Documentation
SPARL queries
Jupyter Notebooks
Observations:
Beneficial for:
Use of schemas, ontologies, terminology services, etc.
Data validation processing
Data imports and exports
Data analysis and visualisation
DevOps Deployment
Medium costs: 120€ to 40€ PCM.
Unreliable and generates errors.
Needs to be tightly constrained on modularised tasks.
Anonymous, Climate Change 2022 - Mitigation of Climate Change: Working Group III Contribution to the Sixth Assessment Report of the Intergovernmental Panel on Climate Change, 1. Auflage, 2023, https://www.cambridge.org/core/product/identifier/9781009157926/type/book.
Anonymous, Untitled,.
Anonymous, Untitled,.
Anonymous, Untitled,.
Anonymous, Untitled,.
Anonymous, Untitled,.
Anonymous, Untitled,.
Anonymous, Untitled,.
Anonymous, Untitled,.
Bank, RCSB Protein Data, Protein Data Bank, RCSB Protein Data Bank.
Broad, Darron, M / Worthington, Simon / Rahr, Anna, Wikibase to Jupyter Notenooks (WB2JN): Software, 2025,.
Broad, M, Darron / Worthington, Simon / Oldenbourg, Simon, CKG Scrape, 2026,.
Bruce, J. P. / Lee, H. / Haites, E. F., Climate Change 1995: Economic and Social Dimensions of Climate Change, Cambridge, United Kingdom; New York, NY, USA 1996.
Calvin, Katherine / Dasgupta, Dipak / Krinner, Gerhard / Mukherji, Aditi / Thorne, Peter W. / Trisos, Christopher / Romero, José / Aldunce, Paulina / Barrett, Ko / Blanco, Gabriel / Cheung, William W. L. / Connors, Sarah / Denton, Fatima / Diongue-Niang, Aı̈da / Dodman, David / Garschagen, Matthias / Geden, Oliver / Hayward, Bronwyn / Jones, Christopher / Jotzo, Frank / Krug, Thelma / Lasco, Rodel / Lee, Yune-Yi / Masson-Delmotte, Valérie / Meinshausen, Malte / Mintenbeck, Katja / Mokssit, Abdalah / Otto, Friederike E. L. / Pathak, Minal / Pirani, Anna / Poloczanska, Elvira / Pörtner, Hans-Otto / Revi, Aromar / Roberts, Debra C. / Roy, Joyashree / Ruane, Alex C. / Skea, Jim / Shukla, Priyadarshi R. / Slade, Raphael / Slangen, Aimée / Sokona, Youba / Sörensson, Anna A. / Tignor, Melinda / Van Vuuren, Detlef / Wei, Yi-Ming / Winkler, Harald / Zhai, Panmao / Zommers, Zinta / Hourcade, Jean-Charles / Johnson, Francis X. / Pachauri, Shonali / Simpson, Nicholas P. / Singh, Chandni / Thomas, Adelle / Totin, Edmond / Alegrı́a, Andrés / Armour, Kyle / Bednar-Friedl, Birgit / Blok, Kornelis / Cissé, Guéladio / Dentener, Frank / Eriksen, Siri / Fischer, Erich / Garner, Gregory / Guivarch, Céline / Haasnoot, Marjolijn / Hansen, Gerrit / Hauser, Mathias / Hawkins, Ed / Hermans, Tim / Kopp, Robert / Leprince-Ringuet, Noëmie / Lewis, Jared / Ley, Debora / Ludden, Chloé / Niamir, Leila / Nicholls, Zebedee / Some, Shreya / Szopa, Sophie / Trewin, Blair / Van Der Wijst, Kaj-Ivar / Winter, Gundula / Witting, Maximilian / Birt, Arlene / Ha, Meeyoung / Arias, Paola / Bustamante, Mercedes / Elgizouli, Ismail / Flato, Gregory / Howden, Mark / Méndez-Vallejo, Carlos / Pereira, Joy Jacqueline / Pichs-Madruga, Ramón / Rose, Steven K. / Saheb, Yamina / Sánchez Rodrı́guez, Roberto / Ürge-Vorsatz, Diana / Xiao, Cunde / Yassaa, Noureddine / Romero, José / Kim, Jinmi / Haites, Erik F. / Jung, Yonghun / Stavins, Robert / Birt, Arlene / Ha, Meeyoung / Orendain, Dan Jezreel A. / Ignon, Lance / Park, Semin / Park, Youngin / Reisinger, Andy / Cammaramo, Diego / Fischlin, Andreas / Fuglestvedt, Jan S. / Hansen, Gerrit / Ludden, Chloé / Masson-Delmotte, Valérie / Matthews, J. B. Robin / Mintenbeck, Katja / Pirani, Anna / Poloczanska, Elvira / Leprince-Ringuet, Noëmie / Péan, Clotilde / Lee, Hoesung, Climate Change 2023: Synthesis Report. Contribution of Working Groups I, II and III to the Sixth Assessment Report of the Intergovernmental Panel on Climate Change [Core Writing Team, H. Lee and J. Romero (eds.)]. IPCC, Geneva, Switzerland., 2023,.
Commons, Creative, Deed - Attribution-NonCommercial-NoDerivatives 4.0 International - Creative Commons, Creative Commons.
Edenhofer, O. / Pichs-Madruga, R. / Sokona, Y. / Farahani, E. / Kadner, S. / Seyboth, K. / Adler, A. / Baum, I. / Brunner, S. / Eickemeier, P. / Kriemann, B. / Savolainen, J. / Schlomer, S. / Stechow, C. von / Zwickel, T. / Minx, J. C., Climate Change 2014: Mitigation of Climate Change, Cambridge, United Kingdom; New York, NY, USA 2014.
El Morr, Christo / Maret, Pierre / Muhlenbach, Fabrice / Dharmalingam, Dhayananth / Tadesse, Rediet / Creighton, Alexandra / Kundi, Bushra / Buettgen, Alexis / Mgwigwi, Thumeka / Dinca-Panaitescu, Serban / Dua, Enakshi / Gorman, Rachel, A Virtual Community for Disability Advocacy: Development of a Searchable Artificial IntelligenceSupported Platform, JMIR Formative Research 2021, e33335, https://pmc.ncbi.nlm.nih.gov/articles/PMC8663581/.
Field, C. B. / Barros, V. R. / Dokken, D. J. / Mach, K. J. / Mastrandrea, M. D. / Bilir, T. E. / Chatterjee, M. / Ebi, K. L. / Estrada, Y. O. / Genova, R. C. / Girma, B. / Kissel, E. S. / Levy, A. N. / MacCracken, S. / Mastrandrea, P. R. / White, L. L., Climate Change 2014: Impacts, Adaptation, and Vulnerability, Cambridge, United Kingdom; New York, NY, USA 2014.
Google / Microsoft / Yahoo / Yandex / Group, W3C Schema.org Community, ResearchProject - Schema.org, 2026,.
Houghton, J. T. / Jenkins, G. J. / Ephraums, J. J. / Houghton, J. T. / Jenkins, G. J. / Ephraums, J. J., Climate Change: The IPCC Scientific Assessment, Cambridge, United Kingdom; New York, NY, USA 1990.
IPCC, Climate Change 1995: A Report of the Intergovernmental Panel on Climate Change Second Assessment Synthesis of Scientific-Technical Information Relevant to Interpreting Article 2 of the UN Framework Convention on Climate Change, 1995.
IPCC, 2019 Refinement to the 2006 IPCC Guidelines for National Greenhouse Gas Inventories, 2019,.
IPCC, Global Warming of 1.5C: IPCC Special Report on Impacts of Global Warming of 1.5C above Pre-industrial Levels in Context of Strengthening Response to Climate Change, Sustainable Development, and Efforts to Eradicate Poverty, 2022,.
IPCC, Climate Change and Land: IPCC Special Report on Climate Change, Desertification, Land Degradation, Sustainable Land Management, Food Security, and Greenhouse Gas Fluxes in Terrestrial Ecosystems, 1. Auflage, 2022, https://www.cambridge.org/core/product/identifier/9781009157988/type/book.
IPCC, The Ocean and Cryosphere in a Changing Climate: Special Report of the Intergovernmental Panel on Climate Change, 1. Auflage, 2022, https://www.cambridge.org/core/product/identifier/9781009157964/type/book.
IPCC, IPCC Sixth Assessment Report, 2023,.
IPCC, Climate Change 2021 The Physical Science Basis: Working Group I Contribution to the Sixth Assessment Report of the Intergovernmental Panel on Climate Change, 1. Auflage, 2023, https://www.cambridge.org/core/product/identifier/9781009157896/type/book.
IPCC, Climate Change 2022 Impacts, Adaptation and Vulnerability: Working Group II Contribution to the Sixth Assessment Report of the Intergovernmental Panel on Climate Change, 1. Auflage, 2023, https://www.cambridge.org/core/product/identifier/9781009325844/type/book.
IPCC, Copyright IPCC, IPCC.
Lindemann, David / Alonso, Mikel, Linking Historical Corpus Data and Annotations Using Wikibase, 2024, https://euralex.org/wp-content/themes/euralex/proceedings/Euralex\%202024/EURALEX2024\_Pr\_p785-791\_Lindemann-Alonso.pdf.pdf.
McCarthy, J. J. / Canziani, O. F. / Leary, N. A. / Dokken, D. J. / White, K. S., Climate Change 2001: Impacts, Adaptation, and Vulnerability, Cambridge, United Kingdom; New York, NY, USA 2001.
Metz, B. / Davidson, O. R. / Bosch, P. R. / Dave, R. / Meyer, L. A., Climate Change 2007: Mitigation of Climate Change, Cambridge, United Kingdom; New York, NY, USA 2007.
Oldenbourg, Laura / Broad, Darron / Worthington, Simon, Entity-Relationship Model: Climate Knowledge Graph, ClimateKG Data Bench 2026,.
Oldenbourg, Laura / Broad, Darron / Worthington, Simon, CKG Data, 2026,.
Oldenbourg, Laura / Worthington, Simon, A Scoping Quantification of the IPCC’s Sixth Assessment Report for the Purpose of Using the Text Corpus in a Knowledge Graph, 2025, https://zenodo.org/records/17516065.
Oldenbourg, Laura / Worthington, Simon, Data Protocol: Quantifying the AR6 Reports and Data Protocol, ClimateKG GitHub Repo 2026,.
Oldenbourg, Laura / Worthington, Simon, IPCC AR6 Data, Google Docs 2026,.
Oldenbourg, Laura / Worthington, Simon / Broad, Darron, ClimateKG ER Model, ClimateKG Data Bench 2026,.
Pachauri, R. K. / Reisinger, A., Climate Change 2007: Synthesis Report, Geneva, Switzerland 2007.
Parry, M. L. / Canziani, O. F. / Palutikof, J. P. / Linden, P. J. van der / Hanson, C. E., Climate Change 2007: Impacts, Adaptation and Vulnerability, Cambridge, United Kingdom; New York, NY, USA 2007.
Shigapov, Renat / Mechnich, Jörg / Schumm, Irene / Verborgh, Ruben / Dimou, Anastasia / Hogan, Aidan / d’Amato, Claudia / Tiddi, Ilaria / Bröring, Arne / Mayer, Simon / Ongenae, Femke / Tommasini, Riccardo / Alam, Mehwish, RaiseWikibase: Fast Inserts into the BERD Instance, The Semantic Web: ESWC 2021 Satellite Events 2021, 60–64.
Shimizu, Cogan / Hitzler, Pascal / Gonzalez-Estrecha, Seila / Goeke-Smith, Jeff / Rehberger, Dean / Foley, Catherine / Sheill, Alicia / Payne, Terry R. / Presutti, Valentina / Qi, Guilin / Poveda-Villalón, Marı́a / Stoilos, Giorgos / Hollink, Laura / Kaoudi, Zoi / Cheng, Gong / Li, Juanzi, The Wikibase Approach to the Enslaved.Org Hub Knowledge Graph, The Semantic Web ISWC 2023 2023, 419–434.
Stocker, T. F. / Qin, D. / Plattner, G.-K. / Tignor, M. / Allen, S. K. / Boschung, J. / Nauels, A. / Xia, Y. / Bex, V. / Midgley, P. M., Climate Change 2013: The Physical Science Basis, Cambridge, United Kingdom; New York, NY, USA 2013.
UNFCCC, Conference of the Parties (COP),.
Watson, R. T. / Zinyowera, M. C. / Moss, R. H., Climate Change 1995: Impacts, Adaptations and Mitigation of Climate Change: Scientific-Technical Analyses, Cambridge, United Kingdom; New York, NY, USA 1996.
Wikidata, Wikidata: Data model - Wikidata, Wikidata 2026,.
Wikidata, Wikibase/DataModel, MediaWiki 2026,.
Wikipedia, Gini coefficient, Wikipedia 2026,.
Worthington, Simon, Climate Knowledge Graph (Project data deposit on Wikidata), Wikidata 2025,.
Worthington, Simon, Project Information Pipeline, 2026,.
Worthington, Simon, SPARQL query for: Gini coefficient use in AR6, ClimateKG 2026,.
Worthington, Simon, IPCC AR6 Author Distribution, ClimateKG Data Bench 2026,.
Worthington, Simon, Bibliography of IPCC reports 1-5 (incomplete), 2026,.
Worthington, Simon, Entity-Relationship Model (ERM) build Wikibase Mapping, ClimateKG Data Bench 2026,.
Worthington, Simon / Broad, Darron, XML DTD Climate KG Data, ClimateKG Data Bench 2026,.
Worthington, Simon / Oldenbourg, Laura / Broad, Darron, ClimateKG Datasets, ClimateKG Data Bench 2026,.
Footnotes
Corresponding author↩︎
Worthington, Climate Knowledge Graph (Project data deposit on Wikidata)↩︎
Worthington, Project Information Pipeline↩︎
Google u. a., ResearchProject - Schema.org↩︎
Oldenbourg u. a., Entity-Relationship Model: Climate Knowledge Graph↩︎
Wikipedia, Gini coefficient↩︎
Worthington, SPARQL query for: Gini coefficient use in AR6↩︎
IPCC, IPCC Sixth Assessment Report↩︎
Oldenbourg u. a., CKG Data↩︎
Worthington, IPCC AR6 Author Distribution↩︎
Bank, Protein Data Bank↩︎
??↩︎
Oldenbourg, Test Chapter - IPCC AR6 as CSS Paged Media.↩︎
??↩︎
??↩︎
??↩︎
Broad u. a., CKG Scrape↩︎
??↩︎
Oldenbourg u. a., ClimateKG ER Model↩︎
IPCC, IPCC Sixth Assessment Report↩︎
IPCC, 2019 Refinement to the 2006 IPCC Guidelines for National Greenhouse Gas Inventories↩︎
Calvin u. a., Climate Change 2023: Synthesis Report. Contribution of Working Groups I, II and III to the Sixth Assessment Report of the Intergovernmental Panel on Climate Change [Core Writing Team, H. Lee and J. Romero (eds.)]. IPCC, Geneva, Switzerland↩︎
UNFCCC, Conference of the Parties (COP)↩︎
Commons, Deed - Attribution-NonCommercial-NoDerivatives 4.0 International - Creative Commons↩︎
IPCC, Copyright IPCC↩︎
Calvin u. a., Climate Change 2023: Synthesis Report. Contribution of Working Groups I, II and III to the Sixth Assessment Report of the Intergovernmental Panel on Climate Change [Core Writing Team, H. Lee and J. Romero (eds.)]. IPCC, Geneva, Switzerland↩︎
IPCC, Global Warming of 1.5C: IPCC Special Report on Impacts of Global Warming of 1.5C above Pre-industrial Levels in Context of Strengthening Response to Climate Change, Sustainable Development, and Efforts to Eradicate Poverty↩︎
IPCC, Climate Change and Land, 1. Aufl. 2022↩︎
IPCC, The Ocean and Cryosphere in a Changing Climate, 1. Aufl. 2022↩︎
IPCC, Climate Change 2021 The Physical Science Basis, 1. Aufl. 2023↩︎
IPCC, Climate Change 2022 Impacts, Adaptation and Vulnerability, 1. Aufl. 2023↩︎
Anonymous, Climate Change 2022 - Mitigation of Climate Change, 1. Aufl. 2023↩︎
IPCC, 2019 Refinement to the 2006 IPCC Guidelines for National Greenhouse Gas Inventories↩︎
Oldenbourg/Worthington, A Scoping Quantification of the IPCC’s Sixth Assessment Report for the Purpose of Using the Text Corpus in a Knowledge Graph, 2025↩︎
Oldenbourg/Worthington, Data Protocol: Quantifying the AR6 Reports and Data Protocol↩︎
Oldenbourg/Worthington, IPCC AR6 Data↩︎
Worthington, Bibliography of IPCC reports 1-5 (incomplete)↩︎
Houghton u. a., Climate Change: The IPCC Scientific Assessment 1990↩︎
Anonymous, Untitled↩︎
Anonymous, Untitled↩︎
Anonymous, Untitled↩︎
Watson u. a., Climate Change 1995: Impacts, Adaptations and Mitigation of Climate Change: Scientific-Technical Analyses 1996↩︎
Bruce u. a., Climate Change 1995: Economic and Social Dimensions of Climate Change 1996↩︎
IPCC, Climate Change 1995: A Report of the Intergovernmental Panel on Climate Change Second Assessment Synthesis of Scientific-Technical Information Relevant to Interpreting Article 2 of the UN Framework Convention on Climate Change 1995↩︎
Anonymous, Untitled↩︎
McCarthy u. a., Climate Change 2001: Impacts, Adaptation, and Vulnerability 2001↩︎
Anonymous, Untitled↩︎
Anonymous, Untitled↩︎
Anonymous, Untitled↩︎
Parry u. a., Climate Change 2007: Impacts, Adaptation and Vulnerability 2007↩︎
Metz u. a., Climate Change 2007: Mitigation of Climate Change 2007↩︎
Pachauri/Reisinger, Climate Change 2007: Synthesis Report 2007↩︎
Stocker u. a., Climate Change 2013: The Physical Science Basis 2013↩︎
Field u. a., Climate Change 2014: Impacts, Adaptation, and Vulnerability 2014↩︎
Edenhofer u. a., Climate Change 2014: Mitigation of Climate Change 2014↩︎
Anonymous, Untitled↩︎
El Morr u. a., A Virtual Community for Disability Advocacy: Development of a Searchable Artificial IntelligenceSupported Platform, JMIR Formative Research 2021↩︎
Diefenbach et al., Wikibase as an Infrastructure for Knowledge Graphs: The EU Knowledge Graph.↩︎
Shimizu u. a., The Wikibase Approach to the Enslaved.Org Hub Knowledge Graph↩︎
Shigapov u. a., RaiseWikibase: Fast Inserts into the BERD Instance↩︎
Lindemann/Alonso, Linking Historical Corpus Data and Annotations Using Wikibase 2024↩︎
Worthington u. a., ClimateKG Datasets↩︎
Broad u. a., Wikibase to Jupyter Notenooks (WB2JN): Software↩︎
Worthington/Broad, XML DTD Climate KG Data↩︎
Worthington, Entity-Relationship Model (ERM) build Wikibase Mapping↩︎
Wikidata, Wikidata: Data model - Wikidata↩︎
Wikidata, Wikibase/DataModel↩︎
Figure 2: The glossary term ‘








