Glossary Terms

This notebook is available to use, reproduce, and validate the glossary results for ClimateKG. Source notebook: glossary.ipynb

It explores glossary terms in the ClimateKG graph by tracing how glossary entries are linked to report and concept structure.

We use the live ClimateKG Wikibase instance and the authenticated SPARQL endpoint.

Query UI: https://climatekg.tibwiki.io/query/

SPARQL endpoint: https://climatekg.tibwiki.io/query/proxy/sparql

This notebook follows the same query-first workflow as the earlier notebooks: validate the graph structure before interpreting the terminology and concept distribution.

Show code
import os
import pandas as pd
from SPARQLWrapper import JSON
from IPython.display import display, Markdown

from wikibase_auth import DEFAULT_SPARQL_ENDPOINT, DEFAULT_WIKIBASE_URL, build_sparql_client

os.environ.setdefault("CLIMATEKG_SPARQL_USERNAME", "ckg")
os.environ.setdefault("CLIMATEKG_SPARQL_PASSWORD", "fairdata")

SPARQL_ENDPOINT = os.getenv("CLIMATEKG_SPARQL_ENDPOINT", DEFAULT_SPARQL_ENDPOINT)
WIKIBASE_URL = os.getenv("CLIMATEKG_WIKIBASE_URL", DEFAULT_WIKIBASE_URL)

print(f"Endpoint: {SPARQL_ENDPOINT}")
print("Libraries imported successfully")
'ckg'
'fairdata'
Endpoint: https://climatekg.tibwiki.io/query/proxy/sparql
Libraries imported successfully
Show code
from pathlib import Path

entity_ns = f"{WIKIBASE_URL.rstrip('/')}/entity/"
property_ns = f"{WIKIBASE_URL.rstrip('/')}/prop/direct/"

count_query = f'''
PREFIX wd: <{entity_ns}>
PREFIX wdt: <{property_ns}>

SELECT (COUNT(DISTINCT ?term) AS ?glossaryCount)
WHERE {{
  ?term wdt:P1 wd:Q1 .
}}
'''

display(Markdown("**SPARQL query (copy/paste into ClimateKG Query UI: https://climatekg.tibwiki.io/query/)**"))
display(Markdown("```sparql\n" + count_query + "\n```"))

sparql = build_sparql_client(SPARQL_ENDPOINT)
sparql.setQuery(count_query)
sparql.setReturnFormat(JSON)
count_results = sparql.query().convert()
glossary_count = int(count_results["results"]["bindings"][0]["glossaryCount"]["value"])
print(f"Glossary term count (Q1 distinct terms): {glossary_count}")

full_query = f'''
PREFIX wd: <{entity_ns}>
PREFIX wdt: <{property_ns}>

SELECT ?term ?termLabel ?report ?reportLabel ?termType ?termTypeLabel
WHERE {{
  ?term wdt:P1 wd:Q1 .
  ?term wdt:P3 ?report .
  OPTIONAL {{ ?term wdt:P1 ?termType . }}
  SERVICE wikibase:label {{ bd:serviceParam wikibase:language \"en\". }}
}}
ORDER BY ?termLabel ?reportLabel
'''

display(Markdown("**SPARQL query (copy/paste into ClimateKG Query UI: https://climatekg.tibwiki.io/query/)**"))
display(Markdown("```sparql\n" + full_query + "\n```"))

sparql.setQuery(full_query)
results = sparql.query().convert()
rows = results.get("results", {}).get("bindings", [])
df_all_terms = pd.DataFrame([{k: v["value"] for k, v in row.items()} for row in rows])

if not df_all_terms.empty:
    df_all_terms = df_all_terms.sort_values(["termLabel", "reportLabel"], na_position="last").reset_index(drop=True)

# Resolve repository root robustly so CSV lands in docs/sparql_queries for site download.
cwd = Path.cwd()
repo_root = next((p for p in [cwd, *cwd.parents] if (p / "docs").exists() and (p / "sparql_queries").exists()), cwd)
csv_path = repo_root / "docs" / "sparql_queries" / "glossary-terms.csv"
csv_path.parent.mkdir(parents=True, exist_ok=True)
df_all_terms.to_csv(csv_path, index=False)

print(f"Rows returned in full terms table: {len(df_all_terms)}")
print(f"CSV saved to: {csv_path}")
display(Markdown("Download CSV: [glossary-terms.csv](glossary-terms.csv)"))
display(df_all_terms)

SPARQL query (copy/paste into ClimateKG Query UI: https://climatekg.tibwiki.io/query/)


PREFIX wd: <https://climatekg.tibwiki.io/entity/>
PREFIX wdt: <https://climatekg.tibwiki.io/prop/direct/>

SELECT (COUNT(DISTINCT ?term) AS ?glossaryCount)
WHERE {
  ?term wdt:P1 wd:Q1 .
}
Glossary term count (Q1 distinct terms): 1001

SPARQL query (copy/paste into ClimateKG Query UI: https://climatekg.tibwiki.io/query/)


PREFIX wd: <https://climatekg.tibwiki.io/entity/>
PREFIX wdt: <https://climatekg.tibwiki.io/prop/direct/>

SELECT ?term ?termLabel ?report ?reportLabel ?termType ?termTypeLabel
WHERE {
  ?term wdt:P1 wd:Q1 .
  ?term wdt:P3 ?report .
  OPTIONAL { ?term wdt:P1 ?termType . }
  SERVICE wikibase:label { bd:serviceParam wikibase:language "en". }
}
ORDER BY ?termLabel ?reportLabel
Rows returned in full terms table: 1361
CSV saved to: C:\git\Climate-KG-data\docs\sparql_queries\glossary-terms.csv

Download CSV: glossary-terms.csv

term report termType termLabel reportLabel termTypeLabel
0 https://climatekg.tibwiki.io/entity/Q1167 https://climatekg.tibwiki.io/entity/Q77 https://climatekg.tibwiki.io/entity/Q1 1.5°C pathway Working Group I: Climate Change 2021 – The Phy... Category
1 https://climatekg.tibwiki.io/entity/Q1167 https://climatekg.tibwiki.io/entity/Q150 https://climatekg.tibwiki.io/entity/Q1 1.5°C pathway Working Group III: Climate Change 2022 – Mitig... Category
2 https://climatekg.tibwiki.io/entity/Q1168 https://climatekg.tibwiki.io/entity/Q10 https://climatekg.tibwiki.io/entity/Q1 1.5°C warmer worlds Special Report: Global Warming of 1.5°C Category
3 https://climatekg.tibwiki.io/entity/Q1169 https://climatekg.tibwiki.io/entity/Q77 https://climatekg.tibwiki.io/entity/Q1 13C Working Group I: Climate Change 2021 – The Phy... Category
4 https://climatekg.tibwiki.io/entity/Q1170 https://climatekg.tibwiki.io/entity/Q77 https://climatekg.tibwiki.io/entity/Q1 14C Working Group I: Climate Change 2021 – The Phy... Category
... ... ... ... ... ... ...
1356 https://climatekg.tibwiki.io/entity/Q1162 https://climatekg.tibwiki.io/entity/Q77 https://climatekg.tibwiki.io/entity/Q1 Younger Dryas Working Group I: Climate Change 2021 – The Phy... Category
1357 https://climatekg.tibwiki.io/entity/Q1163 https://climatekg.tibwiki.io/entity/Q77 https://climatekg.tibwiki.io/entity/Q1 Zero emissions commitment Working Group I: Climate Change 2021 – The Phy... Category
1358 https://climatekg.tibwiki.io/entity/Q1163 https://climatekg.tibwiki.io/entity/Q150 https://climatekg.tibwiki.io/entity/Q1 Zero emissions commitment Working Group III: Climate Change 2022 – Mitig... Category
1359 https://climatekg.tibwiki.io/entity/Q1166 https://climatekg.tibwiki.io/entity/Q77 https://climatekg.tibwiki.io/entity/Q1 pH Working Group I: Climate Change 2021 – The Phy... Category
1360 https://climatekg.tibwiki.io/entity/Q1166 https://climatekg.tibwiki.io/entity/Q106 https://climatekg.tibwiki.io/entity/Q1 pH Working Group II: Climate Change 2022 – Impact... Category

1361 rows × 6 columns

Glossary questions

From the GitHub issue: https://github.com/TIBHannover/ClimateKG-Data-Bench/issues/1

  1. Which glossary terms appear in all three working group reports?

Question 1

Which glossary terms appear in all three working group reports?

This cell identifies glossary terms that are linked to Working Group I, II, and III reports via P3 and shows the reusable SPARQL query.

Show code
query_all_three_wg = f'''
PREFIX wd: <{entity_ns}>
PREFIX wdt: <{property_ns}>
PREFIX rdfs: <http://www.w3.org/2000/01/rdf-schema#>

SELECT ?term ?termLabel
       (COUNT(DISTINCT ?wg) AS ?workingGroupsCovered)
       (GROUP_CONCAT(DISTINCT ?wgLabel; separator=", ") AS ?workingGroupReports)
WHERE {{
  ?term wdt:P1 wd:Q1 ;
        wdt:P3 ?report .
  ?report wdt:P1 wd:Q4 ;
          rdfs:label ?reportLabel .
  FILTER(LANG(?reportLabel) = "en")

  BIND(
    IF(CONTAINS(?reportLabel, "Working Group I:"), "WGI",
    IF(CONTAINS(?reportLabel, "Working Group II:"), "WGII",
    IF(CONTAINS(?reportLabel, "Working Group III:"), "WGIII", "Other"))) AS ?wg
  )
  FILTER(?wg != "Other")

  BIND(
    IF(?wg = "WGI", "Working Group I",
    IF(?wg = "WGII", "Working Group II", "Working Group III")) AS ?wgLabel
  )

  OPTIONAL {{ ?term rdfs:label ?termLabel . FILTER(LANG(?termLabel) = "en") }}
}}
GROUP BY ?term ?termLabel
HAVING (COUNT(DISTINCT ?wg) = 3)
ORDER BY ?termLabel
'''

display(Markdown("**SPARQL query (copy/paste into ClimateKG Query UI: https://climatekg.tibwiki.io/query/)**"))
display(Markdown("```sparql\n" + query_all_three_wg + "\n```"))

sparql = build_sparql_client(SPARQL_ENDPOINT)
sparql.setQuery(query_all_three_wg)
sparql.setReturnFormat(JSON)
all_three_results = sparql.query().convert()
all_three_rows = all_three_results.get("results", {}).get("bindings", [])
all_three_df = pd.DataFrame([{k: v["value"] for k, v in row.items()} for row in all_three_rows])

if all_three_df.empty:
    print("No glossary terms were found across all three working group reports.")
else:
    all_three_df = all_three_df.sort_values("termLabel").reset_index(drop=True)
    all_three_df["termQID"] = all_three_df["term"].astype(str).str.rsplit("/", n=1).str[-1]
    print(f"Glossary terms found in all three working group reports: {len(all_three_df)}")
    display(all_three_df[["termQID", "termLabel", "workingGroupsCovered", "workingGroupReports"]])

SPARQL query (copy/paste into ClimateKG Query UI: https://climatekg.tibwiki.io/query/)


PREFIX wd: <https://climatekg.tibwiki.io/entity/>
PREFIX wdt: <https://climatekg.tibwiki.io/prop/direct/>
PREFIX rdfs: <http://www.w3.org/2000/01/rdf-schema#>

SELECT ?term ?termLabel
       (COUNT(DISTINCT ?wg) AS ?workingGroupsCovered)
       (GROUP_CONCAT(DISTINCT ?wgLabel; separator=", ") AS ?workingGroupReports)
WHERE {
  ?term wdt:P1 wd:Q1 ;
        wdt:P3 ?report .
  ?report wdt:P1 wd:Q4 ;
          rdfs:label ?reportLabel .
  FILTER(LANG(?reportLabel) = "en")

  BIND(
    IF(CONTAINS(?reportLabel, "Working Group I:"), "WGI",
    IF(CONTAINS(?reportLabel, "Working Group II:"), "WGII",
    IF(CONTAINS(?reportLabel, "Working Group III:"), "WGIII", "Other"))) AS ?wg
  )
  FILTER(?wg != "Other")

  BIND(
    IF(?wg = "WGI", "Working Group I",
    IF(?wg = "WGII", "Working Group II", "Working Group III")) AS ?wgLabel
  )

  OPTIONAL { ?term rdfs:label ?termLabel . FILTER(LANG(?termLabel) = "en") }
}
GROUP BY ?term ?termLabel
HAVING (COUNT(DISTINCT ?wg) = 3)
ORDER BY ?termLabel
Glossary terms found in all three working group reports: 96
termQID termLabel workingGroupsCovered workingGroupReports
0 Q262 Adaptation 3 Working Group III, Working Group II, Working G...
1 Q270 Adaptation options 3 Working Group III, Working Group II, Working G...
2 Q272 Adaptive capacity 3 Working Group III, Working Group II, Working G...
3 Q280 Aerosol 3 Working Group III, Working Group II, Working G...
4 Q285 Afforestation 3 Working Group III, Working Group II, Working G...
... ... ... ... ...
91 Q1100 Tipping point 3 Working Group I, Working Group II, Working Gro...
92 Q1128 Uncertainty 3 Working Group I, Working Group II, Working Gro...
93 Q1130 United Nations Framework Convention on Climate... 3 Working Group III, Working Group I, Working Gr...
94 Q1131 Uptake 3 Working Group I, Working Group II, Working Gro...
95 Q1147 Vulnerability 3 Working Group I, Working Group III, Working Gr...

96 rows × 4 columns

Interpretation

Each row connects a glossary term to a report or corpus item via the P3 relationship. That lets us see the terminology in context rather than as isolated labels.

This is the right starting point for glossary analysis: once the relation is confirmed, the next step is to count terms by report, identify repeated concepts, and compare the terms against the glossary pages already generated in the site.