Jump to content

User:Maris Dreshmanis

From Wikidata
Babel user information
lv-N Šim lietotājam latviešu valoda ir dzimtā valoda.
ru-N Русскийродной язык этого участника.
Users by language

About me

[edit]

Independent researcher from Latvia. My focus on Wikidata is improving data quality at scale through deterministic pipelines on authoritative sources — demographic data, multilingual descriptions, occupation labels.

Specialization — multilingual edits in underserved languages. The Wikidata community concentrates on a few major languages by default (EN / DE / FR / ES / RU); Q-items often have labels only in these primary languages, leaving dozens of other living languages uncovered. My focus is every country and every language with at least 1 million native speakers — Persian, Bengali, Khmer, Tamil, Telugu, Marathi, Yoruba, Hausa, Swahili, Filipino, Burmese, Lao, and many others. Closing this structural gap deterministically through national authoritative sources.

Research and credentials

[edit]

ORCID: 0009-0003-8151-4088 · ISNI: 0000 0004 9280 9121 · GitHub · LinkedIn

Wikidata contributions

[edit]

Live statistics: Edit counter (XTools) · Wikiscan · Full contribution history. Extended confirmed editor.

Coverage: 24 countries (population data); 9 analytic languages (settlement descriptions); 50+ languages from ESCO and national occupation registries (occupation labels).

Active tasks

[edit]
Task 2
Latvian labels and descriptions
Native Latvian speaker. Curated dictionary of 2,800+ verified translation pairs.
Task 7
Population figures (population (P1082))
Population data for cities and municipalities — from Wikidata-linked census datasets and Wikipedia infoboxes. Wikipedia sources reference via imported from Wikimedia project (P143) + Wikimedia import URL (P4656) + retrieved (P813); external sources via stated in (P248) + reference URL (P854). Verifies country (P17) (country) on every edit. Coverage: 24 countries.
Task 8
Dead URL deprecation (official website (P856))
HTTP-verified 404 / DNS failures, sets rank to "deprecated" (does not delete). Also migrates reference URL (P854)URL (P2699) in references where applicable.
Task 10
Multilingual settlement descriptions
Short descriptions in 9 analytic languages with fixed prepositions: Indonesian ("kota di {Country}"), Malay, Swahili, Tagalog, Cebuano, Minangkabau, Javanese, Yoruba, Hausa. Templates are deterministic — type-word and preposition are hardcoded; country name comes from the item's own country (P17) value.
Task 11
Multilingual occupation labels
Adds labels and aliases to occupation items using a multilingual aggregation of 156 national occupation registries from 140 countries, linked through ISCO-08 codes. Sources include ESCO v1.2.1 (28 EU languages, 2,942 occupations) and national registries — Turkey ISCO-TR (7,202 occupations), Kenya KeSCO (6,402), Bangladesh BSCO (5,387), Romania COR (4,240), India NCO (3,452), Kazakhstan НК РК 01-2017 (1,717), Czechia CZ-ISCO (1,809), Denmark DISCO-08 (993), Netherlands BRC 2014 (698), Austria ÖISCO-08 (436), Sweden SSYK 2012 (429), Indonesia KBJI (2,721), Hungary FEOR-08 (485), Portugal CPP/2010 (422), Argentina CNO (919 — 2-digit ISCO-08 crosswalk only), Brazil CBO (2,614 — ISCO-88-aligned, ISCO-08 chain pending), Mexico SINCO (686 — partial ISCO-08), Russia OKZ (617), and many others. Live registry list with provenance: gsco.io/api/sources. Matching is deterministic: exact English-label match. Labels ≤ 45 characters → wbsetlabel; longer forms → wbsetaliases. All labels come from officially published government registries. See Wikidata:WikiProject Occupations for the project hub and dataset details.

Bot operation

[edit]

Operations run via User:MarisDreshmanisBot (dedicated account, since 2026-04-25). Source code: Reincarnatiopedia/wikidata-bot (MIT). Stop-requests on User talk:Maris Dreshmanis.

Contact

[edit]

Questions about Wikidata edits — on the talk page. Questions about the occupations dataset — on the discussion page of Wikidata:WikiProject Occupations.