Database Release: September 2026

Important notice about the previous database release

The previous database release, published on July 8, 2026, contained errors regarding player-name resolution. The rules used at the time were too aggressive in some cases. Specifically, abbreviated names, surname-only entries, inconsistent transliterations, missing ratings, and contradictory information from different sources could lead to incorrect player assignments or overly broad name unification.

For the new database release, the rules governing player identity and duplicate handling have been made substantially stricter. The priority is now the prevention of false assignments. Uncertain decisions are no longer resolved by selecting a presumed name form. Instead, they remain visible or, in the case of a concrete conflict, are excluded from the regular export.

More conservative player-name resolution

Player data is replaced or enriched only when the available evidence is sufficiently strong and internally consistent. The assessment takes into account full or abbreviated name forms, initials, event and date information, historical plausibility, available ratings, and source quality.

Uncertain abbreviations and surname-only entries are no longer sufficient on their own for an assignment. Similarly, spelled names are not treated as the same person merely because of their spelling. Alternative spellings and transliterations can still be unified, but only when supported by additional matching evidence.

Existing names receive additional protection, especially in historical games and records lacking Elo information. Historically implausible assignments, such as a player whose lifetime or period of activity does not align with the date of the game, are no longer applied.

Improved Historical Elo Evaluation

All historical FIDE data is now being consolidated into a unified, period-specific database. The applicable rating period is determined directly from the official FIDE list; for older lists that were not published monthly, the most recently published Elo rating applies until the next list is released. This allows player identities to be verified more reliably based on the Elo rating corresponding to the game date and makes it easier to distinguish between players with the same name or those recorded with abbreviated names. Previous FIDE IDs are unambiguously linked to the current ID, while known periods of inactivity prevent a current rating from being incorrectly applied to a historical game. Previous title levels are also displayed according to the relevant period: If a game predates the first known title award, the value “Not Yet Awarded” appears in the PGN. Historical figures without a known FIDE ID are retained with their names, aliases, rating histories, and biographical data as additional evidence of identity, but they still do not receive an estimated or artificially generated FIDE ID. The previous use of FIDE-XML and ratings.ssp is no longer required; monthly updates are carried out exclusively via the official FIDE standard list and are incrementally incorporated into the shared database.

Effects on the number of duplicates

A fundamental difficulty lies in the PGN format itself, or more precisely, in the way it is used in practice. The 1994 PGN specification was designed primarily as an open and extensible interchange format. For archival export, it defines the seven mandatory tags: Event, Site, Date, Round, White, Black, and Result. For certain values, it specifies formats or provides recommendations. For example, Date is intended to use the YYYY.MM.DD schema, while player names are intended to follow the “surname, given name” format.

These rules, however, do not guarantee that the recorded information is factually correct or applied consistently. The PGN import format is explicitly flexible, additional tags are extensible, and many metadata conventions are merely recommendations whose use is not technically enforced across producing programs and data providers. Player names are text fields in the core standard, rather than unique personal identifiers. PGN also lacks a central authority to validate spellings, personal identities, event names, or locations.

In real-world collections, the same person may therefore appear with a full name, initials, a surname only, or in several different transliterations. Event and Site tags may be written out, abbreviated, translated, interchanged, or supplemented with additional information. Date and Round tags may be missing, partially known, or incorrectly filled even though a format is defined. Modern fields such as FIDE IDs are not part of the mandatory core of the original standard and may be absent or incorrect.

Formally valid PGN is therefore not necessarily factually correct or unambiguous. Two records of the same game may have substantially different metadata, while two different games may have very similar metadata. Errors are also frequently copied from one database to another over many years. Reliable duplicate detection must therefore neither depend on a single tag nor treat syntactic PGN compliance as proof of player identity.

The stricter rules substantially reduce the risk of incorrect player assignments. They may, however, leave more apparent or genuine duplicate records in the database. This primarily affects sources with very poor data quality, missing first names, inconsistent transliterations, incorrect FIDE IDs, or contradictory event information.

This behavior is intentional: In an uncertain case, keeping two records separate is less harmful than incorrectly merging different people or publishing a game under the wrong player name.

At the same time, duplicate detection has been expanded. When the move sequence and several independent metadata fields strongly agree, genuine duplicates can still be recognized even when sources use player names of varying detail. If the move sequence clearly matches but the player names cannot be safely reconciled, the database no longer guesses the name form. The case is excluded from the regular export and provided separately for review.

New and extended PGN tags

The database contains additional, more consistently exported tags that make identity decisions, duplicate groups, and record provenance easier to trace.

IdentityEvidenceWhite / IdentityEvidenceBlack

These tags describe the identity assessment separately for White and Black:

There is no single minimum number of arbitrary PGN tags required. One conclusive identity signal carries a different weight than several weakly matching fields. The following minimum conditions therefore describe the required independent evidence:

  • Confirmed: Requires either one conclusive direct identification, especially an exact normalized full-name match, or at least two matching identity signals, such as the same FIDE ID together with matching historical Elo information. There must also be no chronological or factual contradiction. Player data may be replaced or enriched.
  • Strong: Requires at least two matching name components: the same surname and a fully matching given name or complete given-name component. The complete names must also reach a high similarity level. No additional Event, Date, or Elo field is mandatory if no contradiction exists. Player data may be replaced or enriched.
  • Probable: Requires at least two matching name signals: the same surname plus compatible initials or an unambiguous abbreviated given name. At least one independent supporting signal must also be present, such as matching historical Elo information, another external identity confirmation, or a stably identified opponent in the same game context. Player data may be replaced or enriched.
  • Ambiguous: Used when at least one possible candidate or partial match exists, but the combination required for Confirmed, Strong, or Probable is not reached, or when several candidates match almost equally well. There is therefore no fixed minimum number of matching metadata fields. The original player data is retained and is not replaced.
  • Conflict: A single hard contradiction may be sufficient to trigger this value, such as incompatible FIDE IDs, a chronologically impossible assignment, or a contradiction involving a protected identity. For a metadata-based conflict, the identical or structurally matching move sequence must be accompanied by at least three independent contributing factors from Event, Site, Date, and game length, which together must reach a defined evidence threshold. The typical minimum combination is an exactly matching Date, an Event that matches exactly or through shared components, and a game length of at least 40 plies. A matching Site adds evidence and may compensate for weaker or missing evidence from another metadata field; a differing Site never reduces the assessment. At the same time, the player assignments on both sides must remain irreconcilable. The original player data is not replaced. Affected games may be excluded from the regular export and included in the Conflict export.
  • NoMatch: Used when no sufficiently matching identity candidate is found. Therefore, no metadata fields need to match. The original player data is retained and is not replaced.

In summary:

Player data is replaced or enriched only for Confirmed, Strong, or Probable assessments. For Ambiguous, Conflict, and NoMatch, the original data remains unchanged.

The additional Trusted Identity export applies an even stricter rule: It includes only games for which both IdentityEvidenceWhite and IdentityEvidenceBlack are Confirmed or Strong. Probable is not sufficient for this particularly conservative export.

DuplicateStatus

This tag describes the role of a game within duplicate processing. Visible values may include:

  • Master: the authoritative record in a duplicate group;
  • exact: a duplicate with an identical move sequence;
  • sub or subsumption: a shorter game whose complete move sequence is contained in a longer game;
  • sub_replaced: a special subsumption case in which incomplete player data was supplemented from a better-supported group member;
  • join: a strongly matching game with minor differences or additions in the move sequence;
  • identity_conflict: a game with contradictory player information that cannot be safely resolved.

Group_id

The Group_id tag is used in the Conflict export. Every game in an affected duplicate group receives the same Group_id. This keeps the conflicting game, the group master, and related duplicates grouped together during manual review.

SourceQuality

The SourceQuality tag describes the quality or priority of a source. The value -1 is reserved for especially trusted original data or information supplied directly by one of the players. Such data receives additional protection and is not replaced by another source without reliable evidence.

ImportDate

The ImportDate tag identifies the import associated with a game and helps distinguish different database additions.

New Conflict export

Games with identity conflicts that cannot be safely resolved are not included in the regular database export. Instead, they are provided in a separate file named according to the pattern LumbrasGigabase_Conflict.pgn.

The Conflict file contains the complete context of the affected duplicate group. In addition to the conflicting record, the group master and other related duplicates are included. Group_id and DuplicateStatus show which games belong together and the role of each record within the group.

When several records contain the same move sequence, reviewers can thereby easily see whether multiple sources support the same player and event information while only one record differs. The complete context supports manual review without publishing a potentially incorrect player assignment in the regular database.

Trusted Identity export

An additional, particularly conservative database export can be provided. It contains only games for which both player identities are rated Confirmed or Strong.

The Trusted Identity export does not replace the regular export. It is a subset intended for applications where high confidence in player names is more important than the maximum possible number of games.

Corrected date information

Malformed or inconsistent PGN dates are now represented more consistently. Unknown components are written as ?? instead of implying precision not supported by the source. For example, when only the year is known, the date is written as YYYY.??.??.

This prevents incomplete dates from being interpreted as exact day-level information and improves compatibility with standard chess database applications.

Better provenance information

The Source, SourceQuality, and ImportDate tags are preserved more consistently. When duplicate records disagree, provenance information is no longer copied indiscriminately from another game.

This makes it easier to determine where a record originated and which source was preferred within a duplicate group.

Most important corrected data issues

  • Uncertain abbreviations, surname-only entries, and similarly spelled names no longer cause name replacement without additional evidence.
  • Historical and well-known players receive better protection against incorrect assignments.
  • Historically implausible player assignments are no longer applied.
  • Multiple name variants belonging to the same person are assessed more consistently.
  • Malformed or incomplete dates are represented consistently.
  • Especially trusted original data is no longer discarded as an ordinary duplicate.
  • Unresolved identity conflicts are excluded from the regular export and provided with complete group context for review.

Most important new export content

  • IdentityEvidenceWhite and IdentityEvidenceBlack with a visible player-identity assessment;
  • additional DuplicateStatus values sub_replaced and identity_conflict;
  • complete duplicate-group context with Group_id and DuplicateStatus in the Conflict export;
  • separate {Prefix}Conflict.pgn export for unresolved identity conflicts;
  • optional Trusted Identity export containing only Confirmed and Strong identities;
  • more consistent Source, SourceQuality, and ImportDate information;
  • consistent representation of incomplete dates.

Recommendation for users of database release from July 8th 2026

Due to the now-known risks of the earlier, more aggressive player-name resolution, users are strongly advised to replace the database release from July 8, 2026, with the new release.

The new database may contain more duplicate records that have intentionally not been merged. This is not an accidental regression, but rather the result of a stricter safety policy: When source evidence is uncertain, the assignment remains unresolved rather than publishing a potentially incorrect player name as fact.

Scroll to Top