This notebook is available to use, reproduce, and validate the author results for ClimateKG. Source notebook: authors.ipynb
It analyzes author contributions in the ClimateKG graph using the same workflow as the Report Structure notebook and the ER reference in The Rock.
ER-model baseline (The Rock)
Expected author-side model:
Authors are items with P1 -> Q3998 (Author class).
Authors should have a stable ClimateKG Author ID via P20.
Author contributions are modeled with P27 from Author to Chapter.
Chapters are items with P1 -> Q6 (Chapter class).
The queries below first check live graph conformance to that model, then summarize contributions using duplicate-safe counting by P20 (not by item URI).
This notebook now follows the same structure used in Report Structure:
Start from the ER baseline in The Rock (Q3998 Author, Q6 Chapter, P20 author ID, P27 contributed-to relation).
Query the live graph for author-chapter links.
Report duplicate-safe metrics by deduplicating authors on P20.
Run validation checks to flag modeling exceptions.
The key lesson from the structure notebook also applies here: we should not infer completeness from raw item counts alone. Where duplicate imports exist, author-level counts must be grounded on distinct P20 values, while relationships (P27) are validated against chapters typed as Q6.
Statement anatomy
The validation cell now includes a statement-level anatomy view for sample author Q4750, showing how one P27 contribution is represented with:
Main claim value (P27 -> chapter item)
Author-side relationship role (P28 role, e.g. Lead Author)
References (P17 reference URL and P18 date accessed)
This helps verify that the graph carries provenance at claim level, not only at item level.
Author property usage learned from item inspection
From the concrete author item Q4750 and the broader validation checks, the current Author pattern uses the following fields.
Property
Label / Purpose
Role in Author model
P1
Instance of
Class typing (Q3998 Author)
P20
ClimateKG Author ID
Stable author identifier used for deduplication
P27
contributed to chapter
Link from Author to contribution target (expected class Q6 Chapter)
P28
author role on chapter contribution
Author-side contribution role qualifier used on the P27 statement
P21
last name
Biographic/profile attribute
P22
first name
Biographic/profile attribute
P23
gender
Biographic/profile attribute
P24
citizenship
Biographic/profile attribute
P25
country of residence
Biographic/profile attribute
P26
affiliation
Biographic/profile attribute
Wikibase provenance and reference-related fields
These fields are important for source tracking and auditability when present in the graph (either on author records or related records used in the pipeline).
Property
Label / Purpose
Provenance use
P6
Source
Original source URL for imported data
P17
Reference URL
Reference URL attached to a statement reference
P18
date accessed
Alternate/duplicate access-date field seen in model
P19
source version
Captures source version metadata
P5
Wiki
Link to corresponding wiki page/resource
P7
PDF
Link to source PDF resource when relevant
This provides a practical validation template: every Author should have P1, P20, and at least one P27, P27 targets should normally resolve to class Q6 Chapter, and provenance fields should be populated where available for traceability.
Bibliographic questions
From the GitHub issue: https://github.com/TIBHannover/ClimateKG-Data-Bench/issues/1
How many distinct authors contributed to each of the 7 reports?
Which countries of residence are most represented among AR6 authors, and how does this differ between WGI, WGII and WGIII?
What is the share of female authors per working group report, and which chapters have the lowest female representation?
Which authors contributed to chapters in more than one working group report?
Which two authors have co-authored the most chapters together?
Which institutions have the most authors across AR6?
Question 1
How many distinct authors contributed to each of the 7 reports? This cell uses the live SPARQL endpoint and counts distinct author IDs (P20) for each report.
SPARQL query (copy/paste into ClimateKG Query UI: https://climatekg.tibwiki.io/query/)
PREFIX wd: <https://climatekg.tibwiki.io/entity/>
PREFIX wdt: <https://climatekg.tibwiki.io/prop/direct/>
PREFIX rdfs: <http://www.w3.org/2000/01/rdf-schema#>
SELECT ?report ?reportLabel (COUNT(DISTINCT ?authorId) AS ?distinctAuthors)
WHERE {
?author wdt:P1 wd:Q3998 ;
wdt:P20 ?authorId ;
wdt:P27 ?chapter .
?chapter wdt:P3+ ?report .
?report wdt:P1 wd:Q4 ;
rdfs:label ?reportLabel .
FILTER(LANG(?reportLabel) = "en")
}
GROUP BY ?report ?reportLabel
ORDER BY ?reportLabel
reportQID
reportLabel
distinctAuthors
0
Q189
Climate Change 2023: Synthesis Report. Contrib...
30
1
Q35
Special Report: Climate Change and Land
107
2
Q10
Special Report: Global Warming of 1.5°C
91
3
Q57
Special Report: The Ocean and Cryosphere in a ...
103
4
Q77
Working Group I: Climate Change 2021 – The Phy...
234
5
Q106
Working Group II: Climate Change 2022 – Impact...
257
6
Q150
Working Group III: Climate Change 2022 – Mitig...
239
Question 2
Which countries of residence are most represented among AR6 authors, and how does this differ between WGI, WGII and WGIII? This cell uses the live SPARQL endpoint and counts distinct authors (P20) by country (P25).
Show code
query =f'''PREFIX wd: <{ENTITY_NS}>PREFIX wdt: <{PROPERTY_NS}>PREFIX rdfs: <http://www.w3.org/2000/01/rdf-schema#>SELECT ?workingGroupLabel ?country (COUNT(DISTINCT ?authorId) AS ?authors)WHERE {{ ?author wdt:P1 wd:Q3998 ; wdt:P20 ?authorId ; wdt:P25 ?country ; wdt:P27 ?chapter . ?chapter wdt:P3+ ?report . ?report wdt:P1 wd:Q4 ; rdfs:label ?reportLabel . FILTER(LANG(?reportLabel) = "en") BIND(IF(CONTAINS(?reportLabel, "Working Group I:"), "Working Group I", IF(CONTAINS(?reportLabel, "Working Group II:"), "Working Group II", IF(CONTAINS(?reportLabel, "Working Group III:"), "Working Group III", "Other"))) AS ?workingGroupLabel) FILTER(?workingGroupLabel != "Other")}}GROUP BY ?workingGroupLabel ?countryORDER BY ?workingGroupLabel DESC(?authors) ?country'''display(Markdown("**SPARQL query (copy/paste into ClimateKG Query UI: https://climatekg.tibwiki.io/query/)**"))display(Markdown("```sparql\n"+ query +"\n```"))country_df = run_table(query)country_top = country_df.sort_values(["workingGroupLabel", "authors", "country"], ascending=[True, False, True]).groupby("workingGroupLabel", as_index=False).head(10)display(country_top[["workingGroupLabel", "country", "authors"]])
SPARQL query (copy/paste into ClimateKG Query UI: https://climatekg.tibwiki.io/query/)
PREFIX wd: <https://climatekg.tibwiki.io/entity/>
PREFIX wdt: <https://climatekg.tibwiki.io/prop/direct/>
PREFIX rdfs: <http://www.w3.org/2000/01/rdf-schema#>
SELECT ?workingGroupLabel ?country (COUNT(DISTINCT ?authorId) AS ?authors)
WHERE {
?author wdt:P1 wd:Q3998 ;
wdt:P20 ?authorId ;
wdt:P25 ?country ;
wdt:P27 ?chapter .
?chapter wdt:P3+ ?report .
?report wdt:P1 wd:Q4 ;
rdfs:label ?reportLabel .
FILTER(LANG(?reportLabel) = "en")
BIND(IF(CONTAINS(?reportLabel, "Working Group I:"), "Working Group I",
IF(CONTAINS(?reportLabel, "Working Group II:"), "Working Group II",
IF(CONTAINS(?reportLabel, "Working Group III:"), "Working Group III", "Other"))) AS ?workingGroupLabel)
FILTER(?workingGroupLabel != "Other")
}
GROUP BY ?workingGroupLabel ?country
ORDER BY ?workingGroupLabel DESC(?authors) ?country
workingGroupLabel
country
authors
6
Working Group I
Canada
8
7
Working Group I
Germany
8
8
Working Group I
Norway
7
9
Working Group I
Brazil
6
10
Working Group I
India
6
11
Working Group I
Italy
6
12
Working Group I
Argentina
5
13
Working Group I
Netherlands
5
14
Working Group I
Republic of Korea
5
15
Working Group I
South Africa
5
71
Working Group II
Canada
9
72
Working Group II
Mexico
8
73
Working Group II
South Africa
8
74
Working Group II
France
7
75
Working Group II
Norway
7
76
Working Group II
Switzerland
7
77
Working Group II
Netherlands
6
78
Working Group II
Brazil
5
79
Working Group II
Spain
5
80
Working Group II
Tanzania
5
133
Working Group III
Brazil
9
134
Working Group III
Austria
7
135
Working Group III
France
7
136
Working Group III
Italy
7
137
Working Group III
Canada
6
138
Working Group III
Netherlands
6
139
Working Group III
Norway
6
140
Working Group III
Argentina
4
141
Working Group III
Denmark
4
142
Working Group III
Spain
4
Question 3
What is the share of female authors per working group report, and which chapters have the lowest female representation? This cell uses the live SPARQL endpoint and deduplicates authors by P20 before calculating shares.
Show code
query =f'''PREFIX wd: <{ENTITY_NS}>PREFIX wdt: <{PROPERTY_NS}>PREFIX rdfs: <http://www.w3.org/2000/01/rdf-schema#>SELECT ?workingGroupLabel (COUNT(DISTINCT ?authorId) AS ?authors) (SUM(?isFemale) AS ?femaleAuthors) (SUM(?isMale) AS ?maleAuthors)WHERE {{{{ SELECT DISTINCT ?workingGroupLabel ?authorId ?gender WHERE {{ ?author wdt:P1 wd:Q3998 ; wdt:P20 ?authorId ; wdt:P23 ?gender ; wdt:P27 ?chapter . ?chapter wdt:P3+ ?report . ?report wdt:P1 wd:Q4 ; rdfs:label ?reportLabel . FILTER(LANG(?reportLabel) = "en") BIND(IF(CONTAINS(?reportLabel, "Working Group I:"), "Working Group I", IF(CONTAINS(?reportLabel, "Working Group II:"), "Working Group II", IF(CONTAINS(?reportLabel, "Working Group III:"), "Working Group III", "Other"))) AS ?workingGroupLabel) FILTER(?workingGroupLabel != "Other")}}}} BIND(IF(LCASE(STR(?gender)) = "f", 1, 0) AS ?isFemale) BIND(IF(LCASE(STR(?gender)) = "m", 1, 0) AS ?isMale)}}GROUP BY ?workingGroupLabelORDER BY ?workingGroupLabel'''display(Markdown("**SPARQL query (copy/paste into ClimateKG Query UI: https://climatekg.tibwiki.io/query/)**"))display(Markdown("```sparql\n"+ query +"\n```"))female_report_df = run_table(query)female_report_df["femaleSharePct"] = (female_report_df["femaleAuthors"].astype(float) / female_report_df["authors"].astype(float) *100).round(1)female_report_df["maleSharePct"] = (female_report_df["maleAuthors"].astype(float) / female_report_df["authors"].astype(float) *100).round(1)display(female_report_df[["workingGroupLabel", "authors", "femaleAuthors", "maleAuthors", "femaleSharePct", "maleSharePct"]])chapter_query =f'''PREFIX wd: <{ENTITY_NS}>PREFIX wdt: <{PROPERTY_NS}>PREFIX rdfs: <http://www.w3.org/2000/01/rdf-schema#>SELECT ?workingGroupLabel ?chapter ?chapterLabel (COUNT(DISTINCT ?authorId) AS ?authors) (SUM(?isFemale) AS ?femaleAuthors) (SUM(?isMale) AS ?maleAuthors)WHERE {{{{ SELECT DISTINCT ?workingGroupLabel ?chapter ?chapterLabel ?authorId ?gender WHERE {{ ?author wdt:P1 wd:Q3998 ; wdt:P20 ?authorId ; wdt:P23 ?gender ; wdt:P27 ?chapter . ?chapter wdt:P1 wd:Q6 ; rdfs:label ?chapterLabel . FILTER(LANG(?chapterLabel) = "en") ?chapter wdt:P3+ ?report . ?report wdt:P1 wd:Q4 ; rdfs:label ?reportLabel . FILTER(LANG(?reportLabel) = "en") BIND(IF(CONTAINS(?reportLabel, "Working Group I:"), "Working Group I", IF(CONTAINS(?reportLabel, "Working Group II:"), "Working Group II", IF(CONTAINS(?reportLabel, "Working Group III:"), "Working Group III", "Other"))) AS ?workingGroupLabel) FILTER(?workingGroupLabel != "Other")}}}} BIND(IF(LCASE(STR(?gender)) = "f", 1, 0) AS ?isFemale) BIND(IF(LCASE(STR(?gender)) = "m", 1, 0) AS ?isMale)}}GROUP BY ?workingGroupLabel ?chapter ?chapterLabelORDER BY ASC((?femaleAuthors / ?authors)) ?workingGroupLabel ?chapterLabel'''display(Markdown("**SPARQL query (copy/paste into ClimateKG Query UI: https://climatekg.tibwiki.io/query/)**"))display(Markdown("```sparql\n"+ chapter_query +"\n```"))female_chapter_df = run_table(chapter_query)female_chapter_df["femaleSharePct"] = (female_chapter_df["femaleAuthors"].astype(float) / female_chapter_df["authors"].astype(float) *100).round(1)female_chapter_df["maleSharePct"] = (female_chapter_df["maleAuthors"].astype(float) / female_chapter_df["authors"].astype(float) *100).round(1)lowest_female = female_chapter_df.sort_values(["femaleSharePct", "workingGroupLabel", "chapterLabel"], ascending=[True, True, True]).head(10)display(lowest_female[["workingGroupLabel", "chapterLabel", "authors", "femaleAuthors", "maleAuthors", "femaleSharePct", "maleSharePct"]])
SPARQL query (copy/paste into ClimateKG Query UI: https://climatekg.tibwiki.io/query/)
PREFIX wd: <https://climatekg.tibwiki.io/entity/>
PREFIX wdt: <https://climatekg.tibwiki.io/prop/direct/>
PREFIX rdfs: <http://www.w3.org/2000/01/rdf-schema#>
SELECT ?workingGroupLabel (COUNT(DISTINCT ?authorId) AS ?authors) (SUM(?isFemale) AS ?femaleAuthors) (SUM(?isMale) AS ?maleAuthors)
WHERE {
{
SELECT DISTINCT ?workingGroupLabel ?authorId ?gender
WHERE {
?author wdt:P1 wd:Q3998 ;
wdt:P20 ?authorId ;
wdt:P23 ?gender ;
wdt:P27 ?chapter .
?chapter wdt:P3+ ?report .
?report wdt:P1 wd:Q4 ;
rdfs:label ?reportLabel .
FILTER(LANG(?reportLabel) = "en")
BIND(IF(CONTAINS(?reportLabel, "Working Group I:"), "Working Group I",
IF(CONTAINS(?reportLabel, "Working Group II:"), "Working Group II",
IF(CONTAINS(?reportLabel, "Working Group III:"), "Working Group III", "Other"))) AS ?workingGroupLabel)
FILTER(?workingGroupLabel != "Other")
}
}
BIND(IF(LCASE(STR(?gender)) = "f", 1, 0) AS ?isFemale)
BIND(IF(LCASE(STR(?gender)) = "m", 1, 0) AS ?isMale)
}
GROUP BY ?workingGroupLabel
ORDER BY ?workingGroupLabel
workingGroupLabel
authors
femaleAuthors
maleAuthors
femaleSharePct
maleSharePct
0
Working Group I
234
66
168
28.2
71.8
1
Working Group II
257
106
151
41.2
58.8
2
Working Group III
239
77
162
32.2
67.8
SPARQL query (copy/paste into ClimateKG Query UI: https://climatekg.tibwiki.io/query/)
PREFIX wd: <https://climatekg.tibwiki.io/entity/>
PREFIX wdt: <https://climatekg.tibwiki.io/prop/direct/>
PREFIX rdfs: <http://www.w3.org/2000/01/rdf-schema#>
SELECT ?workingGroupLabel ?chapter ?chapterLabel (COUNT(DISTINCT ?authorId) AS ?authors) (SUM(?isFemale) AS ?femaleAuthors) (SUM(?isMale) AS ?maleAuthors)
WHERE {
{
SELECT DISTINCT ?workingGroupLabel ?chapter ?chapterLabel ?authorId ?gender
WHERE {
?author wdt:P1 wd:Q3998 ;
wdt:P20 ?authorId ;
wdt:P23 ?gender ;
wdt:P27 ?chapter .
?chapter wdt:P1 wd:Q6 ;
rdfs:label ?chapterLabel .
FILTER(LANG(?chapterLabel) = "en")
?chapter wdt:P3+ ?report .
?report wdt:P1 wd:Q4 ;
rdfs:label ?reportLabel .
FILTER(LANG(?reportLabel) = "en")
BIND(IF(CONTAINS(?reportLabel, "Working Group I:"), "Working Group I",
IF(CONTAINS(?reportLabel, "Working Group II:"), "Working Group II",
IF(CONTAINS(?reportLabel, "Working Group III:"), "Working Group III", "Other"))) AS ?workingGroupLabel)
FILTER(?workingGroupLabel != "Other")
}
}
BIND(IF(LCASE(STR(?gender)) = "f", 1, 0) AS ?isFemale)
BIND(IF(LCASE(STR(?gender)) = "m", 1, 0) AS ?isMale)
}
GROUP BY ?workingGroupLabel ?chapter ?chapterLabel
ORDER BY ASC((?femaleAuthors / ?authors)) ?workingGroupLabel ?chapterLabel
workingGroupLabel
chapterLabel
authors
femaleAuthors
maleAuthors
femaleSharePct
maleSharePct
0
Working Group II
Tropical Forests
8
1
7
12.5
87.5
1
Working Group I
The Earth’s Energy Budget, Climate Feedbacks a...
15
2
13
13.3
86.7
2
Working Group III
Cross-sectoral Perspectives
13
2
11
15.4
84.6
3
Working Group III
Industry
11
2
9
18.2
81.8
4
Working Group I
Global Carbon and Other Biogeochemical Cycles ...
19
4
15
21.1
78.9
5
Working Group I
Linking Global to Regional Climate Change
19
4
15
21.1
78.9
6
Working Group II
Point of Departure and Key Concepts
14
3
11
21.4
78.6
7
Working Group III
Energy Systems
14
3
11
21.4
78.6
8
Working Group I
Climate Change Information for Regional Impact...
18
4
14
22.2
77.8
9
Working Group I
Future Global Climate: Scenario-based Projecti...
18
4
14
22.2
77.8
Question 4
Which authors contributed to chapters in more than one working group report? This cell uses the live SPARQL endpoint and deduplicates authors by P20 before counting working groups.
Show code
query =f'''PREFIX wd: <{ENTITY_NS}>PREFIX wdt: <{PROPERTY_NS}>PREFIX rdfs: <http://www.w3.org/2000/01/rdf-schema#>SELECT ?authorId ?authorLabel (COUNT(DISTINCT ?workingGroup) AS ?workingGroups) (GROUP_CONCAT(DISTINCT ?workingGroupLabel; separator=", ") AS ?groupLabels)WHERE {{ ?author wdt:P1 wd:Q3998 ; wdt:P20 ?authorId ; wdt:P27 ?chapter . OPTIONAL {{ ?author rdfs:label ?authorLabel . FILTER(LANG(?authorLabel) = "en") }} ?chapter wdt:P3+ ?report . ?report wdt:P1 wd:Q4 ; rdfs:label ?reportLabel . FILTER(LANG(?reportLabel) = "en") BIND(IF(CONTAINS(?reportLabel, "Working Group I:"), "WGI", IF(CONTAINS(?reportLabel, "Working Group II:"), "WGII", IF(CONTAINS(?reportLabel, "Working Group III:"), "WGIII", "Other"))) AS ?workingGroup) FILTER(?workingGroup != "Other") BIND(IF(?workingGroup = "WGI", "Working Group I", IF(?workingGroup = "WGII", "Working Group II", "Working Group III")) AS ?workingGroupLabel)}}GROUP BY ?authorId ?authorLabelHAVING (COUNT(DISTINCT ?workingGroup) > 1)ORDER BY DESC(?workingGroups) ?authorLabel'''display(Markdown("**SPARQL query (copy/paste into ClimateKG Query UI: https://climatekg.tibwiki.io/query/)**"))display(Markdown("```sparql\n"+ query +"\n```"))multi_wg_df = run_table(query)display(multi_wg_df[["authorLabel", "authorId", "workingGroups", "groupLabels"]])
SPARQL query (copy/paste into ClimateKG Query UI: https://climatekg.tibwiki.io/query/)
PREFIX wd: <https://climatekg.tibwiki.io/entity/>
PREFIX wdt: <https://climatekg.tibwiki.io/prop/direct/>
PREFIX rdfs: <http://www.w3.org/2000/01/rdf-schema#>
SELECT ?authorId ?authorLabel (COUNT(DISTINCT ?workingGroup) AS ?workingGroups) (GROUP_CONCAT(DISTINCT ?workingGroupLabel; separator=", ") AS ?groupLabels)
WHERE {
?author wdt:P1 wd:Q3998 ;
wdt:P20 ?authorId ;
wdt:P27 ?chapter .
OPTIONAL { ?author rdfs:label ?authorLabel . FILTER(LANG(?authorLabel) = "en") }
?chapter wdt:P3+ ?report .
?report wdt:P1 wd:Q4 ;
rdfs:label ?reportLabel .
FILTER(LANG(?reportLabel) = "en")
BIND(IF(CONTAINS(?reportLabel, "Working Group I:"), "WGI",
IF(CONTAINS(?reportLabel, "Working Group II:"), "WGII",
IF(CONTAINS(?reportLabel, "Working Group III:"), "WGIII", "Other"))) AS ?workingGroup)
FILTER(?workingGroup != "Other")
BIND(IF(?workingGroup = "WGI", "Working Group I", IF(?workingGroup = "WGII", "Working Group II", "Working Group III")) AS ?workingGroupLabel)
}
GROUP BY ?authorId ?authorLabel
HAVING (COUNT(DISTINCT ?workingGroup) > 1)
ORDER BY DESC(?workingGroups) ?authorLabel
authorLabel
authorId
workingGroups
groupLabels
0
Linda Mearns
AU0554
2
Working Group II, Working Group I
1
Patricia Romero Lankao
AU0721
2
Working Group III, Working Group II
Question 5
Which two authors have co-authored the most chapters together? This cell uses the live SPARQL endpoint and counts shared chapters per author pair.
Which institutions have the most authors across AR6? This cell uses the live SPARQL endpoint and counts distinct author IDs (P20) per affiliation (P26).