Chinese ancient city space-time knowledge graph construction method and device and storage medium
By acquiring ancient text data and constructing triples, the problem of existing technologies being unable to build spatiotemporal knowledge graphs of ancient cities has been solved, enabling precise spatiotemporal positioning and deep semantic retrieval, supporting urban planning and cultural heritage protection.
Patent Information
- Application Number
- CN202511564133.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-29
- Publication Date
- 2026-02-10
AI Technical Summary
Existing technologies cannot effectively construct spatiotemporal knowledge graphs of ancient cities, nor can they be compatible with the spatial characteristics and dynamic attribute evolution of ancient cities, resulting in the loss of historical information and failing to support the needs of urban planning, construction, and cultural heritage protection.
Based on the set of place names of the target city, ancient book documents are obtained, sentence segmentation and knowledge extraction are performed, triples are constructed, entity attributes are determined through semantic disambiguation, and a spatiotemporal knowledge graph of ancient Chinese cities is constructed to achieve multi-granular spatiotemporal positioning and semantic retrieval.
It achieves precise spatial positioning from macro-regions to micro-geographic units, improves the accuracy of spatiotemporal information analysis and expression, constructs a unified structured knowledge system, supports deep semantic retrieval and analysis, and provides decision support for urban history research, cultural heritage protection, and planning and construction.
Smart Images

Figure CN121503628A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the technical field of information processing, in particular to a method and device for constructing a time-space knowledge graph of ancient Chinese cities and a storage medium. BACKGROUND
[0002] Ancient Chinese cities contain rich historical information. Constructing a time-space knowledge graph of ancient cities helps to mine the practical experience of ancient people in planning and building cities in terms of time and space, and has important significance for current urban planning and construction and cultural heritage protection. The information of ancient Chinese cities is mainly concentrated in ancient books and documents, such as local chronicles, notes, and city records. However, ancient texts have highly unstructured characteristics and contain a large amount of ambiguous or polysemous time-space descriptions. How to effectively convert them into structured knowledge graphs to support researchers to discover macroscopic laws and deep connections that cannot be revealed by reading the original text is still a technical problem to be solved.
[0003] In the prior art, the construction of a knowledge graph for historical information mainly exists in the following schemes, but none of them can meet the specific needs of ancient city research:
[0004] 1. Knowledge graph construction scheme with historical figures or historical documents as the core: In this scheme, “figures” or “documents” are taken as core entities, mainly used to construct a figure relationship network and track academic inheritance. The technical defect is that cities are only used as the place where events occur or the background of figure activities, and are treated as simple geographic location attributes (such as a coordinate point or a place name string), which cannot model the complex internal structure of cities (such as city walls, government offices, and streets), dynamic attribute evolution (city construction, wars, etc.), and the spatial relationship between various elements.
[0005] 2. Knowledge graph scheme with modern urban planning as the core: In this scheme, the ontology design and analysis logic of the graph serve modern urban management, and its core concepts include land use, traffic network, etc. The technical defect is that its ontology model is significantly different from the spatial form, functional layout, and evolution logic of ancient cities, and cannot be compatible with and express the spatial characteristics of ancient cities (such as central axes, li-fang, city walls, etc.). Directly applying the modern city model will result in the loss of historical information.
[0006] Therefore, in view of the current urban planning and construction and cultural heritage protection, it is more necessary to return to the ancient context and construct a time-space knowledge graph with ancient cities as the core ontology in combination with the characteristics of ancient cities. SUMMARY
[0007] Therefore, the present disclosure proposes a method and device for constructing a time-space knowledge graph of ancient Chinese cities and a storage medium.
[0008] According to an aspect of the present disclosure, a method for constructing a time-space knowledge graph of ancient Chinese cities is provided. The method comprises:
[0009] Based on the city name set of the target city, obtain ancient literature data related to the target city;
[0010] Based on the ancient literature data, perform sentence segmentation processing and knowledge extraction to obtain a plurality of ancient text entries and a plurality of triples corresponding to the plurality of ancient text entries, the triples including entities, attributes of the entities, and relationships between the entities, the types of the entities including time, space, and city, and the entities of the time and space types including a plurality of attributes of different granularities;
[0011] For each triple corresponding to each ancient text entry, perform semantic disambiguation to determine the attribute values of the entities of the time and space types uniquely corresponding to each triple;
[0012] Based on the triples respectively corresponding to the plurality of ancient text entries and the attribute values of the entities of the time and space types uniquely corresponding to the triples, construct a time-space knowledge graph of ancient Chinese cities, the time-space knowledge graph of ancient Chinese cities including structured features capable of supporting semantic retrieval and analysis related to ancient Chinese cities to provide decision support for digital protection of cultural heritage and modern city planning.
[0013] In a possible implementation, the ancient literature data includes first text data and second text data, and based on the city name set of the target city, the ancient literature data related to the target city is obtained, comprising:
[0014] Extract the historical names, aliases, abbreviations, and administrative division transition information of the target city from a historical geography database to determine the city name set of the target city, the city name set of the target city including the ancient name and the modern name of the target city;
[0015] Based on the city name set of the target city, perform keyword matching in a catalog database of ancient books to obtain a book list related to the target city, and take the full-text text corresponding to the book list as the first text data;
[0016] Based on the city name set of the target city, perform keyword matching in a full-text database of ancient books to obtain a title text related to the target city, and take the title text and the context text of the title text as the second text data.
[0017] In a possible implementation, based on the ancient literature data, perform sentence segmentation processing and knowledge extraction to obtain a plurality of ancient text entries and a plurality of triples corresponding to the plurality of ancient text entries, comprising:
[0018] Perform sentence segmentation processing on the ancient literature data using a natural language model to obtain a plurality of ancient text sentences;
[0019] generating an initial ancient text item identified by a unique identifier for each sentence of the ancient text, the initial ancient text item including a book name, a volume number, content, a period, a version, an author, an ancient name and a modern name corresponding to the ancient text;
[0020] correcting the period and the author in the plurality of initial ancient text items to obtain a corrected ancient text item;
[0021] merging the corrected ancient text items that are repeated, retaining the earliest corrected ancient text item corresponding to the period in the repeated corrected ancient text items, to obtain a plurality of ancient text items;
[0022] performing knowledge extraction processing on each ancient text item to obtain a triple corresponding to each ancient text item.
[0023] In a possible implementation, the correcting the period and the author in the plurality of initial ancient text items to obtain a corrected ancient text item includes:
[0024] For any initial ancient text item, in response to a case where there are multiple periods in any initial ancient text item, a layout analysis model is used to determine that the content in the initial ancient text item belongs to a main text area or a note area;
[0025] In a case where the content in the initial ancient text item belongs to the note area, the author corresponding to the initial ancient text item is determined to be a note writer, and the period corresponding to the initial ancient text item is determined to be a corresponding age of the note writer;
[0026] In a case where the content in the initial ancient text item belongs to the main text area, the author corresponding to the initial ancient text item is determined to be an original author, and the period corresponding to the initial ancient text item is determined to be a corresponding age of the original author.
[0027] In a possible implementation, for each triple corresponding to each ancient text item, semantic disambiguation is performed to determine an attribute value of an entity of a time and space type uniquely corresponding to each triple, including:
[0028] For each triple corresponding to each ancient text item, a spatio-temporal attribute candidate library is determined, the spatio-temporal attribute candidate library including all candidate time and space information strings related to the triple;
[0029] A candidate time and space information string is determined from the spatio-temporal attribute candidate library as disambiguated time and space attribute information;
[0030] The disambiguated time and space attribute information corresponding to each triple is standardized to obtain an attribute value of an entity of a time and space type uniquely corresponding to each triple.
[0031] In a possible implementation, for each triple corresponding to each ancient text entry, a candidate library of spatio-temporal attributes is determined, including:
[0032] The string corresponding to the entity of the time and space type in the triple is taken as the first candidate time and space information string in the candidate library of spatio-temporal attributes;
[0033] The string corresponding to the entity of the time and space type appearing in the content of the ancient text entry corresponding to the triple is taken as the second candidate time and space information string in the candidate library of spatio-temporal attributes;
[0034] The third candidate time and space information string in the candidate library of spatio-temporal attributes is determined based on the period and ancient name of the ancient text entry corresponding to the triple;
[0035] A candidate time and space information string is determined from the candidate library of spatio-temporal attributes as the disambiguated time and space attribute information, including:
[0036] In the case where the first candidate time and space information string exists in the candidate library of spatio-temporal attributes, the first candidate time and space information string is taken as the disambiguated time and space attribute information;
[0037] In the case where the first candidate time and space information string does not exist in the candidate library of spatio-temporal attributes but the second candidate time and space information string exists, the disambiguated time and space attribute information is determined based on the second candidate time and space information string;
[0038] In the case where the first candidate time and space information string and the second candidate time and space information string do not exist in the candidate library of spatio-temporal attributes but the third candidate time and space information string exists, the disambiguated time and space attribute information is determined based on the third candidate time and space information string.
[0039] In a possible implementation, the type of the entity further includes a person, an event, a building, a settlement, a geographical entity, and a literature source, and the entity of the literature source type is used to indicate each part in the ancient text entry;
[0040] The relationship between the entities satisfies a constraint condition, and the constraint condition includes that the entity of the city type, the person type, the event type, the building type, the settlement type, the geographical entity type, and the literature source type respectively has an association relationship with the entity of the time type, the entity of the city type, the person type, the event type, the building type, the settlement type, and the geographical entity type respectively has an association relationship with the entity of the space type, and the entity of the person type, the event type, the building type, the settlement type, and the geographical entity type respectively has an association relationship with the entity of the literature source type and the city type.
[0041] In one possible implementation, the attributes of a time-type entity include attributes of time period and time point type. The attributes of the time period type include dynasty, reign title and other time ranges, and the attributes of the time point type include year.
[0042] The attributes of spatial type entities include attributes of spatial surface and spatial point types. The attributes of spatial surface type include high-level administrative regions, unified county administrative regions, county-level administrative regions, and other spatial ranges.
[0043] The attributes of city-type entities include ancient name, category, and modern name;
[0044] The attributes of a person-type entity include name, gender, and category;
[0045] The attributes of an event type entity include content and category;
[0046] The attributes of building-type entities, settlement-type entities, and geographical-type entities include name and category, respectively;
[0047] The attributes of the source type include the title, volume number, content, period, edition, and author of the ancient text entry.
[0048] According to another aspect of this disclosure, a device for constructing a spatiotemporal knowledge graph of ancient Chinese cities is provided. The device includes:
[0049] The acquisition module is used to acquire ancient book and document data related to the target city based on the set of city place names of the target city;
[0050] The sentence segmentation and knowledge extraction module is used to segment and extract knowledge from ancient text data, resulting in multiple ancient text entries and corresponding triples. Each triple includes an entity, the entity's attributes, and the relationship between entities. The types of entities include time, space, and city. Time and space type entities each include multiple attributes of different granularities.
[0051] The semantic disambiguation module is used to perform semantic disambiguation on each triple corresponding to each ancient text entry, and to determine the attribute value of the time and space type entity that uniquely corresponds to each triple.
[0052] The module is used to construct a spatiotemporal knowledge graph of ancient Chinese cities based on the triples corresponding to multiple ancient text entries and the attribute values of the time and space types of entities that uniquely correspond to the triples. The spatiotemporal knowledge graph of ancient Chinese cities includes structured features that can support semantic retrieval and analysis related to ancient Chinese cities, so as to provide decision support for the digital protection of cultural heritage and modern urban planning.
[0053] In one possible implementation, the ancient text data includes first text data and second text data, and the acquisition module is used for:
[0054] Extract historical place names, alternative names, abbreviations, and administrative division changes of the target city from the historical geography database to determine the set of city place names of the target city, which includes the ancient and modern names of the target city.
[0055] Based on the set of city place names of the target city, keyword matching is performed in the ancient book catalog database to obtain the ancient book catalog related to the target city, and the full text of the ancient book catalog is used as the first text data.
[0056] Based on the set of city place names of the target city, keyword matching is performed in the full-text database of ancient books to obtain the hit text related to the target city. The hit text and the context text of the hit text are used as the second text data.
[0057] In one possible implementation, the sentence segmentation and knowledge extraction module is used for:
[0058] Natural language processing was used to segment ancient texts into sentences, resulting in multi-sentence ancient texts.
[0059] For each sentence of ancient text, generate an initial ancient text entry with a unique identifier. The initial ancient text entry includes the book title, volume number, content, period, version, author, ancient name and modern name of the ancient text.
[0060] The period and author in multiple initial ancient book text entries were corrected to obtain the corrected ancient book text entries;
[0061] Duplicate entries in the corrected ancient text entries are merged, and the earliest corrected ancient text entry corresponding to the period of the duplicate entries is retained, resulting in multiple ancient text entries.
[0062] For each ancient text entry, knowledge extraction is performed to obtain the triplet corresponding to each ancient text entry.
[0063] In one possible implementation, the period and author information in multiple initial ancient text entries are corrected to obtain corrected ancient text entries, including:
[0064] For any initial ancient text entry, in response to the existence of multiple periods in any initial ancient text entry, a layout analysis model is used to determine whether the content in the initial ancient text entry belongs to the main text area or the annotation area.
[0065] If the content in the initial ancient text entry belongs to the annotation area, the author corresponding to the initial ancient text entry is identified as the annotator, and the period corresponding to the initial ancient text entry is identified as the era corresponding to the annotator.
[0066] If the content of the initial ancient text entry belongs to the main text area, the author corresponding to the initial ancient text entry is determined as the original author, and the period corresponding to the initial ancient text entry is determined as the era corresponding to the original author.
[0067] In one possible implementation, the semantic disambiguation module is used for:
[0068] For each triple corresponding to each ancient text entry, a spatiotemporal attribute candidate library is determined. The spatiotemporal attribute candidate library includes all candidate time and space information strings related to the triple.
[0069] A candidate time and space information string is selected from the candidate library of time and space attributes as the disambiguated time and space attribute information;
[0070] The disambiguated temporal and spatial attribute information corresponding to each triple is standardized to obtain the attribute values of the entity of the time and space type that uniquely corresponds to each triple.
[0071] In one possible implementation, for each triple corresponding to each ancient text entry, a candidate library of spatiotemporal attributes is determined, including:
[0072] The strings corresponding to the time and space types of entities in the triplet are used as the first candidate time and space information strings in the spatiotemporal attribute candidate library;
[0073] The strings corresponding to other time and space types of entities appearing in the content of the ancient text entries corresponding to the triplet are used as the second candidate time and space information strings in the candidate library of spatiotemporal attributes.
[0074] Based on the period and ancient name of the ancient text entry corresponding to the triple, the third candidate time and space information string in the candidate library of spatiotemporal attributes is determined;
[0075] A candidate time and space information string is determined from the candidate library of time and space attributes as the disambiguated time and space attribute information, including:
[0076] If a first candidate time and space information string exists in the candidate library of spatiotemporal attributes, the first candidate time and space information string shall be used as the disambiguated time and space attribute information.
[0077] If there is no first candidate time and space information string in the spatiotemporal attribute candidate library but there is a second candidate time and space information string, the disambiguated time and space attribute information is determined based on the second candidate time and space information string;
[0078] If there is no first candidate time and space information string or second candidate time and space information string in the spatiotemporal attribute candidate library, but there is a third candidate time and space information string, the disambiguated time and space attribute information is determined based on the third candidate time and space information string.
[0079] In one possible implementation, the types of entities also include people, events, buildings, settlements, geographical entities, and sources of information, with entities of the source of information type used to indicate different parts of ancient text entries;
[0080] The relationships between entities satisfy the following constraints: entities of the city, person, event, building, settlement, geographical entity, and document source type are associated with entities of the time type; entities of the city, person, event, building, settlement, and geographical entity types are associated with entities of the spatial type; and entities of the person, event, building, settlement, and geographical entity types are associated with entities of the document source type and city type.
[0081] In one possible implementation, the attributes of a time-type entity include attributes of time period and time point type. The attributes of the time period type include dynasty, reign title and other time ranges, and the attributes of the time point type include year.
[0082] The attributes of spatial type entities include attributes of spatial surface and spatial point types. The attributes of spatial surface type include high-level administrative regions, unified county administrative regions, county-level administrative regions, and other spatial ranges.
[0083] The attributes of city-type entities include ancient name, category, and modern name;
[0084] The attributes of a person-type entity include name, gender, and category;
[0085] The attributes of an event type entity include content and category;
[0086] The attributes of building-type entities, settlement-type entities, and geographical-type entities include name and category, respectively;
[0087] The attributes of the source type include the title, volume number, content, period, edition, and author of the ancient text entry.
[0088] According to another aspect of this disclosure, an apparatus for constructing a spatiotemporal knowledge graph of ancient Chinese cities is provided, including a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the steps of the above method.
[0089] According to another aspect of this disclosure, a non-volatile computer-readable storage medium is provided, on which a computer program is stored, which, when executed by a processor, implements the steps of the above-described method.
[0090] According to another aspect of this disclosure, a computer program product is provided, including a computer program or a non-volatile computer-readable storage medium carrying the computer program, wherein the computer program, when executed by a processor, implements the steps of the above-described method.
[0091] According to the embodiments of this disclosure, ancient book literature data related to the target city is obtained by using a set of city place names based on the target city. Sentence segmentation and knowledge extraction are performed on the ancient book literature data to obtain multiple ancient book text entries and triples corresponding to the multiple ancient book text entries. The triples include entities, entity attributes, and relationships between entities. The types of entities include time, space, and city. Time and space types of entities each include multiple attributes of different granularities, which can achieve accurate spatial positioning from macro-regions to micro-geographic units. By meeting the time requirements of different scales, the accuracy of knowledge graph in analyzing, interpreting, and expressing spatiotemporal information is significantly improved. By performing semantic disambiguation on each triple corresponding to each ancient text entry, and determining the attribute value of the unique time and space type entity corresponding to each triple, each triple can possess traceable, multi-granular spatiotemporal positioning capabilities. Based on the triples corresponding to multiple ancient text entries and the attribute values of the unique time and space type entities corresponding to each triple, a spatiotemporal knowledge graph of ancient Chinese cities can be constructed. This allows for the construction of a spatiotemporal knowledge graph with ancient cities as the ontology, breaking through the limitations of traditional urban historical data modeling centered on documents or figures. It aggregates diverse knowledge elements such as figures, events, and buildings, originally scattered across different spatiotemporal dimensions and record carriers, into a unified, structured knowledge system through semantic association with core urban entities. The constructed spatiotemporal knowledge graph of ancient Chinese cities can support in-depth semantic retrieval and analysis related to ancient Chinese cities, providing efficient knowledge support and decision-making assistance for fields such as urban history research, cultural heritage protection, and urban planning and construction.
[0092] Other features and aspects of this disclosure will become clear from the following detailed description of exemplary embodiments with reference to the accompanying drawings. Attached Figure Description
[0093] The accompanying drawings, which are included in and form part of this specification, illustrate exemplary embodiments, features, and aspects of this disclosure together with the specification and serve to explain the principles of this disclosure.
[0094] Figure 1 A schematic diagram illustrating an application scenario according to an embodiment of this disclosure is shown.
[0095] Figure 2 A flowchart illustrating a method for constructing a spatiotemporal knowledge graph of ancient Chinese cities according to an embodiment of this disclosure is shown.
[0096] Figure 3 This diagram illustrates the hierarchical structure of a spatiotemporal knowledge graph of ancient Chinese cities according to an embodiment of the present disclosure.
[0097] Figure 4 A structural diagram of an apparatus for constructing a spatiotemporal knowledge graph of ancient Chinese cities according to an embodiment of the present disclosure is shown.
[0098] Figure 5 This is a block diagram illustrating an apparatus 1900 for constructing a spatiotemporal knowledge graph of ancient Chinese cities according to an exemplary embodiment. Detailed Implementation
[0099] Various exemplary embodiments, features, and aspects of this disclosure will now be described in detail with reference to the accompanying drawings. The same reference numerals in the drawings denote elements that have the same or similar functions. Although various aspects of the embodiments are shown in the drawings, they are not necessarily drawn to scale unless specifically indicated otherwise.
[0100] As used herein, the terms “comprising,” “including,” “having,” or variations thereof are open-ended and include one or more of the stated features, integrals, elements, steps, components, or functions, but do not exclude the presence or addition of one or more other features, integrals, elements, steps, components, functions, or groups thereof.
[0101] When an element is referred to as “connected,” “coupled,” “responding,” or a variation thereof relative to another element, it may be directly connected, coupled, or responding to another element, or there may be an intermediate element present.
[0102] Although the terms first, second, third, etc., may be used herein to describe various elements / operations, these elements / operations should not be limited by these terms. These terms are only used to distinguish one element / operation from another. Therefore, without departing from the teachings of the inventive concept, a first element / operation in some embodiments may be referred to as a second element / operation in other embodiments.
[0103] The term “exemplary” as used herein means “serving as an example, embodiment, or illustration.” Any embodiment illustrated herein as “exemplary” is not necessarily to be construed as superior to or better than other embodiments.
[0104] Furthermore, to better illustrate this disclosure, numerous specific details are set forth in the following detailed description. Those skilled in the art will understand that this disclosure can be practiced without certain specific details. In some instances, methods, means, components, and circuits well known to those skilled in the art have not been described in detail in order to highlight the main points of this disclosure.
[0105] Ancient Chinese cities contain a wealth of historical information. Constructing a spatiotemporal knowledge graph of ancient cities can help uncover the practical experience of ancient people in planning and building cities at both temporal and spatial levels, which is of great significance for current urban planning, construction, and cultural heritage protection. Information about ancient Chinese cities is mainly concentrated in ancient texts, such as local gazetteers, anecdotal novels, and gazetteers of cities and towns. However, these texts are highly unstructured and contain a large number of ambiguous or polysemous spatiotemporal descriptions. How to effectively transform them into structured knowledge graphs to support researchers in discovering macroscopic patterns and deep connections that are difficult to reveal by simply reading the original texts remains a pressing technical challenge.
[0106] In existing technologies, the construction of knowledge graphs for historical information mainly involves the following types of schemes, but none of them can meet the specific needs of ancient city research:
[0107] 1. Knowledge graph construction schemes centered on historical figures or documents: These schemes use "figures" or "documents" as core entities, primarily for building networks of relationships between figures and tracing academic lineages. Their technical limitation lies in the fact that cities are merely treated as locations of events or backgrounds for figures' activities, processed as simple geographical attributes (such as a coordinate point or place name string). They fail to model the complex internal structure of cities (such as city walls, government offices, and streets), dynamic attribute evolution (city construction, wars, etc.), and spatial relationships between various elements.
[0108] 2. Knowledge Graph Scheme Based on Modern Urban Planning: In this type of scheme, the ontology design and analysis logic of the graph serve modern urban management. Its core concepts include land use and transportation networks. Its technical drawback lies in the significant differences between its ontology model and the spatial morphology, functional layout, and evolutionary logic of ancient cities. It cannot be compatible with or express the spatial characteristics of ancient cities (such as central axes, neighborhoods, and city walls). Directly applying modern urban models will lead to the loss of historical information.
[0109] Therefore, in order to address current urban planning, construction, and cultural heritage protection, we need a solution that can return to the ancient context, combine the characteristics of ancient cities, and construct a spatiotemporal knowledge graph with ancient cities as the core ontology.
[0110] In view of this, embodiments of this disclosure provide a method, apparatus, and storage medium for constructing a spatiotemporal knowledge graph of ancient Chinese cities. The method of this disclosure acquires ancient text data related to a target city based on a set of city place names. It then performs sentence segmentation and knowledge extraction on the ancient text data to obtain multiple ancient text entries and corresponding triples. Each triple includes an entity, entity attributes, and relationships between entities. Entity types include time, space, and city. Time and space types of entities each include multiple attributes of different granularities, enabling precise spatial positioning from macro-regions to micro-geographical units. By meeting the time requirements at different scales, it significantly improves the accuracy of the knowledge graph's analysis, interpretation, and expression of spatiotemporal information. By performing semantic disambiguation on each triple corresponding to each ancient text entry, and determining the attribute value of the unique time and space type entity corresponding to each triple, each triple can possess traceable, multi-granular spatiotemporal positioning capabilities. Based on the triples corresponding to multiple ancient text entries and the attribute values of the unique time and space type entities corresponding to each triple, a spatiotemporal knowledge graph of ancient Chinese cities can be constructed. This allows for the construction of a spatiotemporal knowledge graph with ancient cities as the ontology, breaking through the limitations of traditional urban historical data modeling centered on documents or figures. It brings together diverse knowledge elements such as figures, events, and buildings, originally scattered across different spatiotemporal dimensions and record carriers, into a unified, structured knowledge system through semantic association with core urban entities. The constructed spatiotemporal knowledge graph of ancient Chinese cities can support in-depth semantic retrieval and analysis related to ancient Chinese cities, providing efficient knowledge support and decision-making assistance for fields such as urban history research, cultural heritage protection, and urban planning and construction.
[0111] Figure 1 The diagram illustrates an application scenario according to an embodiment of this disclosure. The Chinese ancient city spatiotemporal knowledge graph construction system of this disclosure can be deployed on a terminal device or server, such as... Figure 1 As shown, the Chinese ancient city spatiotemporal knowledge graph construction system of this disclosure can retrieve relevant ancient books and documents based on the city place name set of relevant cities, and construct a Chinese ancient city spatiotemporal knowledge graph based on the ancient books and documents with ancient cities as the ontology. The Chinese ancient city spatiotemporal knowledge graph obtained by this disclosure can be used in scenarios such as urban history research, cultural heritage protection, and urban planning and construction, providing reliable data support and auxiliary decision-making capabilities for related fields.
[0112] The terminal devices involved in the embodiments of this disclosure can be any one or more of the following: mobile phones, foldable electronic devices, tablet computers, desktop computers, laptop computers, handheld computers, laptops, ultra-mobile personal computers (UMPCs), netbooks, cellular phones, personal digital assistants (PDAs), and in-vehicle devices. The embodiments of this disclosure do not impose any special limitations on the specific type of terminal device, which can have wired or wireless communication capabilities.
[0113] The server disclosed in this embodiment can be located locally or in the cloud, and can be a physical device or a virtual device, such as a virtual machine or container. It has wireless communication capabilities, which can be configured in the server's chip (system) or other components. The wireless communication capabilities can be implemented through mobile communication technologies such as 2G / 3G / 4G / 5G, as well as Wi-Fi, Bluetooth, frequency modulation (FM), data radio, and satellite communication. It can also communicate via a wired connection to enable interaction with other devices.
[0114] Figure 2 A flowchart illustrating a method for constructing a spatiotemporal knowledge graph of ancient Chinese cities according to an embodiment of this disclosure is shown. Figure 2 As shown, the method may include:
[0115] Step S201: Based on the set of city place names of the target city, obtain ancient book documents related to the target city.
[0116] The target city can be one city or multiple cities. To ensure the comprehensiveness of the search, you can first collect all the names of the target cities at different times.
[0117] In one possible implementation, information on historical place names, alternative names, abbreviations, and administrative division changes of the target city can be extracted from a historical geography database to determine the set of city place names for the target city.
[0118] The historical geography database can include controllable reference materials such as the *Historical Atlas of China* and *Dictionary of Place Names Throughout History*. It can store extracted historical place names, alternative names, abbreviations, and information on administrative division changes of the target city in an institutionalized form, resulting in a set of city place names for the target city. This set of city place names can include both the ancient and modern names of the target city for subsequent retrieval. The ancient names refer to place names or administrative names used by the target city during its historical development; for example, the ancient names of Enshi City could include Shizhou, Shinanfu, and Enshi County.
[0119] Ancient text data can include textual content related to the target city. To ensure comprehensive and specific city-related information is obtained, this disclosure provides a multi-scale strategy for acquiring ancient text data, including whole-book and sub-book-scale acquisition. The whole-book scale is used to collect complete bibliographies of ancient books related to the target city at a macro level, ensuring the completeness of city documents through title retrieval and filtering. The sub-book scale is used to supplement the documents at a micro level, further extracting textual content related to the target city that may have been missed in the whole-book scale through refined retrieval at the chapter, paragraph, or item level, thereby achieving fine-grained improvement of the ancient text data.
[0120] In one possible implementation, the ancient text data includes first text data and second text data. Based on the set of city place names of the target city, ancient text data related to the target city is obtained, including:
[0121] Based on the set of city place names of the target city, keyword matching is performed in the ancient book catalog database to obtain the ancient book catalog related to the target city, and the full text of the ancient book catalog is used as the first text data.
[0122] Based on the set of city place names of the target city, keyword matching is performed in the full-text database of ancient books to obtain the hit text related to the target city. The hit text and the context text of the hit text are used as the second text data.
[0123] For the whole book scale, the ancient book literature data can include the first text data. The ancient book catalog database is a database system that systematically collects, organizes, stores, retrieves, and displays the catalog information related to ancient books (such as book titles, authors, editions, survival status, content summaries, etc.). It can be an existing publicly available database or a dedicated ancient book catalog resource library. When retrieving, the set of city names of the target city can be used as the input condition to perform keyword matching on information fields such as book titles and summaries in the ancient book catalog database to identify ancient book titles related to the target city, so as to cover as many records of different historical periods as possible. To improve the retrieval coverage rate, keyword expansion can also be performed based on the ancient names in the set of city names. For example, by combining common administrative division suffixes (such as "zhou", "fu", "xian", etc.) or common aliases in historical periods, the expanded keywords are added to the above input conditions to expand a richer set of retrieval keywords and avoid missing ancient book literature related to the target city.
[0124] The ancient book titles related to the target city can be ancient books whose titles contain the ancient name of the target city. These ancient books are often directly related to the target city, such as local chronicles, city chronicles, note novels, etc.
[0125] If the full-text data of the ancient book titles is already included in the ancient book catalog database, the corresponding full-text can be directly extracted as the first text data; if the ancient book catalog database only provides bibliographic information, the corresponding ancient book can be scanned to obtain image data, and an optical character recognition (OCR) tool (such as PaddleOCR, etc.) can be used to extract the text content in the image as the first text data. Thus, the ancient book literature data at the whole book scale can be obtained, providing a basis for subsequent information processing and knowledge extraction.
[0126] For the sub-book scale, the ancient book literature data can include the second text data. The second text data includes some texts in the ancient books, which are used to supplement the information that may be missed in the whole book scale retrieval. The ancient book full-text database is a database system that systematically collects, collates, encodes, stores the complete text content of ancient books (rather than just catalog information), and supports full-text retrieval, browsing, analysis, and utilization. It can be an existing publicly available database or a dedicated ancient book full-text digital resource library.
[0127] By utilizing a set of city names of the target city in a full-text database of ancient books, keyword matching is performed on the full text. This involves extracting sentences, paragraphs, or entries related to the target city from the set of city names based on their location within the full text of the ancient book. These are then used as the target text, along with surrounding text fragments (e.g., 40 characters before and after the target city) as the context text. Both the target text and its context text are then used as secondary text data. This approach not only identifies documents where the target city's name is not mentioned in the title but is mentioned in the text, but also reveals more granular city-related information.
[0128] This process can effectively supplement information that may be missed during the full-book scale retrieval, forming a more comprehensive and fine-grained collection of ancient text data, providing more sufficient and high-quality data support for subsequent knowledge extraction and spatiotemporal knowledge graph construction.
[0129] Step S202 involves sentence segmentation and knowledge extraction based on ancient text data to obtain multiple ancient text entries and their corresponding triples.
[0130] Each ancient text entry can correspond to a sentence in the ancient text data (including the first text data and / or the second text data). For each keyword matching result (i.e., ancient text data) at the whole book scale and the sub-book scale, the book title, volume number, content, period, version, and author can be extracted. This information, along with the ancient name and modern name fields, is stored in a structured format in a unified database.
[0131] In step S202, the following can be done:
[0132] The ancient text data is segmented using a natural language model to obtain multiple sentences of ancient texts. Multiple initial ancient text entries, each uniquely identified, are generated for each sentence. The periods and authors in these initial entries are corrected to obtain corrected entries. Duplicate entries in the corrected entries are merged, and the earliest corrected entry from the corresponding period is retained, resulting in multiple ancient text entries. Knowledge extraction is performed on each entry to obtain a corresponding triple.
[0133] Each initial ancient text entry is associated with a sentence from the ancient text data. The natural language processing model can be a pre-trained GujiBERT model, an AI Taiyan ancient Chinese language model, etc. The natural language processing model can be used to segment the full text corresponding to the first text data, as well as the matched text and context text corresponding to the second text data, to obtain each sentence of the ancient text.
[0134] For any given sentence from an ancient text, a unique identifier can be generated, and its corresponding initial ancient text entry can be determined. This initial entry can include various attribute fields stored in the database, such as the book title (BookTitle), volume number (BookVolume), content (BookText), period (BookPeriod), version (BookVersion), author (BookAuthor), ancient name (AncientCityName), and modern name (ModernCityName). Specifically, the book title refers to the name of the ancient text to which the sentence belongs; the volume number refers to the volume number within the ancient text; the content refers to the sentence itself; the period can be the year the text was written or cited; the version refers to the edition information, such as "Qianlong 12th year edition"; and the author refers to the author or commentator of the text. In this way, entries can be stored at the sentence level, constructing a corpus containing multiple initial ancient text entries, thus achieving fine-grained text indexing.
[0135] To ensure data quality, multiple initial ancient text entries can be cleaned. This cleaning process may include correcting the aforementioned periods and authors, as well as removing duplicate entries.
[0136] In one possible implementation, the period and author information in multiple initial ancient text entries are corrected to obtain corrected ancient text entries, including:
[0137] For any initial ancient text entry, in response to the existence of multiple periods within that initial ancient text entry, a layout analysis model is used to determine whether the content of the initial ancient text entry belongs to the main text area or the annotation area. If the content of the initial ancient text entry belongs to the annotation area, the author corresponding to the initial ancient text entry is identified as the annotator, and the period corresponding to the initial ancient text entry is identified as the era corresponding to the annotator. If the content of the initial ancient text entry belongs to the main text area, the author corresponding to the initial ancient text entry is identified as the original author, and the period corresponding to the initial ancient text entry is identified as the era corresponding to the original author.
[0138] Layout analysis models, such as the YOLO (You Only Look Once) based layout analysis model, etc.
[0139] For some ancient text entries, the corresponding ancient text may have multiple authors (typically, early documents have been annotated by scholars throughout history, hence multiple authors, such as "Zuo Zhuan" and "Guanzi"). In this case, the exact period and author information of the text can be automatically determined by the location of the text in the ancient text facsimile and its font size. If the ancient text is confirmed to be either an annotation or the main text, the "Author" field corresponds to the annotator if it is an annotation, and to the original author if it is the main text.
[0140] To illustrate further, the period of the *Guanzi* in the ancient text database is: Eastern Zhou Dynasty, Tang Dynasty; author: Guan Zhong (Eastern Zhou Dynasty) wrote, Fang Xuanling (Tang Dynasty) annotated. This indicates that part of the text of the *Guanzi* was written by Guan Zhong during the Eastern Zhou Dynasty, and part was written by Fang Xuanling during the Tang Dynasty. In ancient texts, the content written by different authors is distinguished by layout and font size. In the case of the *Guanzi*, the original text written by Guan Zhong (Eastern Zhou Dynasty) is in a larger font, while the annotations written by Fang Xuanling (Tang Dynasty) are in a smaller font and appended after the main text. Therefore, a layout analysis model can be used to identify the main text area and the annotation area in the ancient text image, and combined with OCR technology to extract the text separately, thus distinguishing between the main text and the annotations. If it is identified as an annotation, the "Author" field in the ancient text entry corresponds to the annotator, and the "Period" field corresponds to the annotator's era; if it is the main text, the "Author" field in the ancient text entry corresponds to the original author, and the "Period" field corresponds to the original author's era.
[0141] Because ancient books often cite earlier documents during their compilation, there may be highly similar ancient texts among different ancient text entries. To eliminate duplicates and avoid redundancy, we can calculate and estimate the text similarity between texts (for example, by calling the difflib library in Python to calculate text similarity) and pre-set a similarity threshold (for example, 0.5). We can then compare each corrected ancient text entry pairwise to identify duplicate entries whose text similarity of the content part (i.e., the ancient text itself) is greater than the similarity threshold. For any two entries that are determined to be duplicates, we can remove duplicates based on the date of composition of the corresponding document (determined based on the "period" field mentioned above) and retain the earliest corrected ancient text entry from the corresponding period.
[0142] Through the above process, a batch of high-quality ancient text entries that have undergone sentence segmentation, structuring, correction, and deduplication can be obtained, providing a reliable data foundation for subsequent knowledge extraction and spatiotemporal knowledge graph construction.
[0143] In order to construct a spatiotemporal knowledge graph of ancient Chinese cities, this disclosure takes into account the history of ancient urban planning and the geographical paradigm of historical cities, and predefines the core concepts in the spatiotemporal knowledge graph of ancient cities, including entities, the attributes of entities, and the relationships between entities (i.e., predefine which entities can have relationships, i.e. the constraints below), i.e., triples.
[0144] Table 1 illustrates the structure and concept of predefined entities and their attributes according to embodiments of this disclosure.
[0145] Table 1
[0146]
[0147] As shown in Table 1, entity types can include time, space, and city. Time-type entities can be used to represent time information in ancient texts, space-type entities can be used to represent geospatial information, and city-type entities are the core entities of the knowledge graph, used to represent specific ancient cities. The difference between city and space-type entities is that space entities emphasize the objective existence of geographical divisions or natural geographical spaces, such as representing the latitude and longitude of a specific location, while city entities focus on describing the scale and social functional characteristics centered on the city, such as ancient names, categories, and modern names.
[0148] Entities of time and space types can each include multiple attributes of different granularities. For example, the attributes of entities of time and space types can be divided into two categories: point-like and non-point-like. Point-like attributes are finer-grained attributes, which can be accurately located in a smaller range or at a more specific time or space point, thereby improving the spatiotemporal resolution capability.
[0149] The attributes of a time-type entity can include time period and time point type attributes. Time period type attributes can include dynasty, reign title and other time ranges, while time point type attributes can include year.
[0150] The attributes of spatial entities can include attributes of spatial surfaces and spatial points. Attributes of spatial surfaces can include higher-level administrative divisions (such as provinces), county-level administrative divisions, and other spatial extents. Attributes of spatial points can be spatial locations, represented by latitude and longitude. Attributes of city entities can include ancient names, categories, and modern names, used to describe the overall scale and development characteristics of the city.
[0151] In this embodiment, the spatial type attributes take into account multiple granularities, including points (coordinates) and surfaces (regional spatial units such as administrative divisions), to achieve accurate spatial positioning from macro-regions to micro-geographical units. The time type attributes cover point-like time (years) and non-point-like time (time intervals such as dynasties and reign titles), meeting the time requirements of different scales and enabling the knowledge graph to have higher time resolution capabilities.
[0152] Entities can also include at least one of the following: people, events, buildings, settlements, geographical entities, and document sources. Entities of the document source type can be used to indicate various parts of ancient text entries, such as book title, volume number, content, period, edition, and author. Among them, entities of the people, events, buildings, settlements, and geographical entities are all attached to a specific city, that is, they have a relationship with a specific city. For example, an event occurs in a certain city, or a building is located in a certain city.
[0153] The "Person" type entity can be used to represent historical figures mentioned in ancient texts. The attributes of the "Person" type entity can include name, gender, and category. The category can be used to distinguish different identities or roles, such as officials, those who passed the imperial examinations, and local worthies.
[0154] Event type entities can be used to represent historical events related to cities recorded in ancient texts. The attributes of event type entities can include content and category. The category of an event can be distinguished according to the nature of the event, such as city building, disaster, war, etc.
[0155] Building type entities can be used to represent specific buildings mentioned in ancient texts, settlement type entities can be used to represent settlements outside of or belonging to cities, and geographical entities can be used to represent natural or non-natural geographical features related to cities.
[0156] The attributes of building-type entities, settlement-type entities, and geographic-type entities can include name and category, respectively. The building category can be used to represent buildings within the territory of an ancient city, such as government offices, temples, and warehouses. The settlement category can be used to represent the human settlement environment within the territory of an ancient city, such as towns and villages. The geographic entity category can include natural geographic entities (such as mountains and rivers) as well as non-natural geographic entities (such as transportation routes and canals).
[0157] The document source type entity can be used to represent the source information of ancient texts. The attributes of the document source type can include the title, volume number, content, period, version, and author of the ancient text entry. This provides source and contextual support for subsequent knowledge extraction results, thereby ensuring the traceability and reliability of the knowledge graph.
[0158] Of the entities mentioned above, all except those of the document source type can have their attributes extracted from the "content" of the ancient text entries, i.e., the ancient text itself. If these attributes cannot be extracted from the ancient text, they can be determined through the attributes of the document source type entity.
[0159] The knowledge extraction process described above can be achieved by text annotation of relevant ancient texts. Each ancient text entry can be annotated to obtain a triple corresponding to each entry. The triple can include an entity, the entity's attributes, and the relationships between entities.
[0160] Because ancient texts are unstructured, natural language processing tools (such as the HanLP Ancient Chinese Processing Model) can be used to annotate the text corresponding to each entry, extracting entities such as time, place names, and personal names. For example, for the ancient text "Duke Xin Guo Tang He personally experienced the surveying and construction of Zhapu Town in the nineteenth year of Hongwu", natural language processing tools can be used to identify the following entities: personal name "Tang He", official title "Duke Xin Guo", time "the nineteenth year of Hongwu", place name "Zhapu Town", and event "construction of city". Then, entities including time, space, person, and event type, along with their corresponding attributes, are annotated. For time-type entities, the year attribute is annotated as "the nineteenth year of Hongwu", further standardized to the Gregorian calendar year; for space-type entities, the attribute is annotated as "Zhapu Town", further converted to corresponding latitude and longitude coordinates; the name attribute of the task-type entity is "Tang He", and the category attribute of the event-type entity is "construction of city".
[0161] For entities of time and space type, the finest granular attribute value can be taken first. For example, when both "Hongwu era" and "Hongwu nineteenth year" exist in the text, the finer granular "Hongwu nineteenth year" should be taken first as the year attribute value.
[0162] Step S203: Perform semantic disambiguation on each triple corresponding to each ancient text entry to determine the attribute value of the time and space type entity that uniquely corresponds to each triple.
[0163] In one possible implementation, step S203 includes:
[0164] For each triple corresponding to each ancient text entry, a candidate library of spatiotemporal attributes is determined; wherein, the candidate library of spatiotemporal attributes may include all candidate time and space information strings related to the triple.
[0165] In one possible implementation, for each triple corresponding to each ancient text entry, a candidate library of spatiotemporal attributes is determined, including:
[0166] The strings corresponding to the time and space types of entities in the triple are used as the first candidate time and space information strings in the spatiotemporal attribute candidate library; the strings corresponding to other time and space types of entities appearing in the content of the ancient text entries corresponding to the triple are used as the second candidate time and space information strings in the spatiotemporal attribute candidate library; the third candidate time and space information strings are determined based on the period and ancient name of the ancient text entries corresponding to the triple.
[0167] Candidate time and space information strings can include strings associated with time attributes and strings associated with space attributes, respectively. The first candidate time and space information string can be a time or space entity string explicitly contained within the triple itself. The second candidate time and space information string can be other time or space entity strings appearing in the ancient text entry to which it belongs. The third candidate time and space information string can be from the source metadata associated with the ancient text entry, where the time attribute can be inherited from the period (BookPeriod) and the space attribute can be inherited from the ancient city name (AncientCityName).
[0168] In step S203, it is also possible to: determine a candidate time and space information string from the spatiotemporal attribute candidate library as the disambiguated time and space attribute information.
[0169] After obtaining the candidate attributes, a unique value determination process for the spatiotemporal attribute (i.e., disambiguation) is performed for each triple. This process aims to select and determine the unique and most relevant spatiotemporal attribute string for the triple from the candidate spatiotemporal attribute library. This process strictly follows the priority of the candidate attributes and combines syntactic dependency analysis to make automated decisions.
[0170] Among them, it is possible to:
[0171] If a first candidate time and space information string exists in the spatiotemporal attribute candidate library, the first candidate time and space information string is used as the disambiguated time and space attribute information. If a first candidate time and space information string does not exist in the spatiotemporal attribute candidate library but a second candidate time and space information string exists, the disambiguated time and space attribute information is determined based on the second candidate time and space information string. If a first candidate time and space information string does not exist in the spatiotemporal attribute candidate library and a second candidate time and space information string exists but a third candidate time and space information string exists, the disambiguated time and space attribute information is determined based on the third candidate time and space information string.
[0172] First, check if a first candidate time and space information string exists in the candidate library of the triple's spatiotemporal attributes. Since this string originates directly from the triple itself, it has the highest priority and is unique. If a first candidate time and space information string exists, it can be directly selected as the spatiotemporal attribute of the triple, and this step completes the determination of that attribute. If no first candidate time and space information string exists, then check if a second candidate time and space information string exists.
[0173] If no first candidate time and space information string exists, but a unique second candidate time and space information string exists, the second candidate time and space information string is used. Disambiguation can be performed based on syntactic dependency parsing. First, natural language processing tools (such as the HanLP Classical Chinese processing model) can be used to perform dependency parsing on the content (i.e., sentences) of the ancient text entries containing the triples, generating a dependency parsing tree. The dependency parsing tree can structurally represent the grammatical relationships between words in the sentence. The shortest path distance between the node corresponding to the time or space string in each second candidate time and space information string and the node of the core predicate verb of the triple can be calculated within this dependency parsing tree. This distance is quantified by the number of edges required to connect the two nodes. Finally, the candidate entity with the shortest dependency path distance to the core predicate verb node can be selected as the unique spatiotemporal attribute that is semantically closest and most relevant.
[0174] A third candidate time and space information string is selected as a supplement only if neither the first nor the second candidate time and space information string exists in the candidate pool of time and space attributes. That is, the period (BookPeriod) and ancient name (AncientCityName) attributes are retrieved from the metadata of the source documents associated with the ancient text entries, and these are used as the unique time and space attributes of the triples, respectively. It is important to note that all ancient text entries have period (BookPeriod) and ancient name (AncientCityName) attributes, therefore it is possible to assign relevant information to all triples. This allows us to obtain the disambiguated time and space attribute information for each triple.
[0175] In order to obtain the finest spatial and temporal information, in step S203, the disambiguated temporal and spatial attribute information corresponding to each triple can be standardized to obtain the attribute values of the time and space type entities that uniquely correspond to each triple.
[0176] This allows the disambiguated temporal and spatial attributes to be transformed from natural language into structured, computable, multi-granular spatiotemporal data, establishing a unified semantic foundation for the nodes and relationships of knowledge graphs.
[0177] For the attribute values of time-type entities, the year names, eras, sexagenary cycle, and other time expressions in ancient texts can be mapped to standardized time attributes based on a pre-built historical chronology conversion rule library (based on authoritative materials such as the "Chronological Table of Chinese History"), and used as year attribute values.
[0178] If the text contains a year that can be precisely matched (such as "Hongwu 20th year"), it is converted into a unique Gregorian year (1387) while retaining its higher-level granularity information (such as "Hongwu era" "Ming Dynasty").
[0179] If it is not possible to parse down to a specific year, then a coarser-grained time unit (such as the time period attribute value "Hongwu era" or "Ming Dynasty") is retained to form a three-level time structure of "dynasty - reign title - year".
[0180] This multi-granularity representation method can support semantic querying and reasoning of subsequent knowledge graphs at different time levels (such as century, dynasty, and reign title).
[0181] For the attributes of spatial entities, the extracted ancient place name strings can be mapped to standardized historical geographic entities by calling the application programming interface (API) of the Chinese Historical Geographic Information System (CHGIS) to obtain their unique identifier (ID), latitude and longitude coordinates and hierarchical structure.
[0182] When an exact match exists, point information (such as the coordinates of the city center) can be recorded; when the local name refers to a wide or vague area (such as "Northern Sichuan" in the Ming Dynasty), it is recorded as a surface or regional spatial unit, and its superior administrative entity is retained (such as "Northern Sichuan ∈ Sichuan Province" in the Ming Dynasty), so that each spatial entity forms a multi-granularity hierarchy of "point-surface-region-system" to support cross-temporal and spatial alignment and regional-scale statistical analysis.
[0183] The results, after being standardized by time and space attributes, are uniformly converted into a machine-readable structured format and mapped to the attributes of knowledge graph nodes.
[0184] In this way, each triple has traceable, multi-granular spatiotemporal positioning capabilities, realizing a closed-loop mapping of "text-entity-spatiotemporal-knowledge".
[0185] Triads can indicate the spatiotemporal and semantic relationships between multiple entities in an ancient text entry. Spatiotemporal relationships can represent the co-occurrence or dependency of entities in the time and space dimensions, such as an event occurring at a specific time and in a specific space. Semantic relationships can represent the connection between entities in semantic logic or narrative structure, such as a person participating in an event, or a building belonging to a city. For the ancient text "Duke Tang He of Xinguo personally experienced the surveying and construction of Zhapu Town in the nineteenth year of Hongwu" mentioned above, it can be annotated based on the spatiotemporal relationships between multiple entities, such as "the construction event—occurred in—the nineteenth year of Hongwu" (indicating the relationship between the event entity and the time entity) and "the construction event—occurred in—Zhapu Town" (indicating the relationship between the event entity and the city entity), etc.
[0186] The relationships between entities satisfy constraints, which can also be used to indicate the hierarchical structure of the knowledge graph disclosed herein.
[0187] In one possible implementation, the constraints may include: entities of the city, person, event, building, settlement, geographical entity, and document source type have association relationships with entities of the time type; entities of the city, person, event, building, settlement, and geographical entity type have association relationships with entities of the spatial type; and entities of the person, event, building, settlement, and geographical entity type have association relationships with entities of the document source type and city type.
[0188] Figure 3 This diagram illustrates the hierarchical structure of a spatiotemporal knowledge graph of ancient Chinese cities according to an embodiment of the present disclosure. Figure 3As shown, there are different relationships between the nine types of entities, and the attributes included in each of the nine types of entities are consistent with those in Table 1. Specifically, time-type entities are associated with entities of the city, person, event, building, settlement, geographical entity, and document source types, indicating the time when other entities occur (as shown by "when"), such as the time when a city or person exists, the time when an event occurs, and the time when a document is written. Spatial-type entities are associated with entities of the city, person, event, building, settlement, geographical entity, and document source types, indicating the location where other entities occur (as shown by "where"), such as the location of a person, event, building, and the location mentioned in a document. City-type entities are associated with entities of the person, event, building, settlement, and geographical entity types. The relationship between a person and a city-type entity includes birth, death (as shown by "lives at"), governance, and travel; the relationship between an event and a city-type entity is that the event occurs in a city (as shown by "happens in"); and the relationship between a building, settlement, and geographical entity and a city-type entity is that the building, settlement, and geographical entity are located in a city (as shown by "is located"). (in”); entities of the document source type have a relationship with entities of the person, event, building, settlement, and geographical entity types, indicating the source recorded by other types of entities (as shown in “is documented in” in the figure).
[0189] It can be seen that in this embodiment, by combining multi-scale acquisition of ancient books and documents with multi-scale spatiotemporal attribute assignment, a spatiotemporal knowledge graph is constructed with ancient cities as the ontology. This breaks through the limitations of traditional urban historical data modeling and brings together diverse knowledge elements such as people, events, and buildings that were originally scattered in different spatiotemporal dimensions and record carriers into a unified structured knowledge system through semantic association with core urban entities.
[0190] like Figure 3 The specific types of entities, attributes, and relationships, as well as their hierarchical structure, shown in the ancient city spatiotemporal knowledge graph can be analyzed using relevant knowledge modeling tools (such as...). ) to build and implement.
[0191] Step S204: Based on the triples corresponding to multiple ancient text entries and the attribute values of the time and space type entities that uniquely correspond to the triples, construct a spatiotemporal knowledge graph of ancient Chinese cities.
[0192] Each of the aforementioned ancient text entries can be used to extract one or more sets of entities, entity attributes, and relationships between entities. These entity, attribute, and relationship data from multiple ancient text entries can be batch-imported into a knowledge graph database (Neo4j graph database) using relevant knowledge graph construction tools or interfaces (such as the data import tools and programming interfaces provided by Neo4j). During this process, entities and their attributes can be modeled as nodes and node attributes in the graph database, and the relationships between entities can be modeled as edges, thus forming a complete graph structure (as shown above). Figure 3 (Illustration of the hierarchical structure in the text).
[0193] The Chinese Ancient City Spatiotemporal Knowledge Graph includes structured features that support semantic retrieval and analysis related to ancient Chinese cities, providing decision support for the digital protection of cultural heritage and modern urban planning.
[0194] According to the embodiments of this disclosure, ancient book literature data related to the target city is obtained by using a set of city place names based on the target city. Sentence segmentation and knowledge extraction are performed on the ancient book literature data to obtain multiple ancient book text entries and triples corresponding to the multiple ancient book text entries. The triples include entities, entity attributes, and relationships between entities. The types of entities include time, space, and city. Time and space types of entities each include multiple attributes of different granularities, which can achieve accurate spatial positioning from macro-regions to micro-geographic units. By meeting the time requirements of different scales, the accuracy of knowledge graph in analyzing, interpreting, and expressing spatiotemporal information is significantly improved. By performing semantic disambiguation on each triple corresponding to each ancient text entry, and determining the attribute value of the unique time and space type entity corresponding to each triple, each triple can possess traceable, multi-granular spatiotemporal positioning capabilities, achieving a closed-loop mapping of "text-entity-spatiotemporal-knowledge". Based on the triples corresponding to multiple ancient text entries and the attribute values of the unique time and space type entities corresponding to each triple, a spatiotemporal knowledge graph of ancient Chinese cities can be constructed. This allows for the construction of a spatiotemporal knowledge graph with ancient cities as the ontology, breaking through the limitations of traditional urban historical data modeling centered on documents or figures. It brings together diverse knowledge elements such as figures, events, and buildings, originally scattered across different spatiotemporal dimensions and record carriers, into a unified structured knowledge system through semantic association with core urban entities. The constructed spatiotemporal knowledge graph of ancient Chinese cities can support in-depth semantic retrieval and analysis related to ancient Chinese cities, providing efficient knowledge support and decision-making assistance for fields such as urban history research, cultural heritage protection, and urban planning and construction.
[0195] At the application level, the spatiotemporal knowledge graph of ancient Chinese cities is used for semantic retrieval and analysis related to ancient Chinese cities. For example, researchers can construct Cypher queries to retrieve co-occurrence information across multiple historical timescales. Taking the analysis of "the spatiotemporal distribution and dominant players in the construction of coastal county-level cities during the Jiajing period of the Ming Dynasty" as an example, it can return the construction events of coastal county-level cities during this period, the relevant participants, and their spatial distribution. With the help of visualization tools or programming interfaces, researchers can not only intuitively browse this data but also conduct in-depth analysis of the spatiotemporal patterns and evolutionary characteristics of ancient cities in different historical periods.
[0196] The spatiotemporal knowledge graph of ancient Chinese cities constructed in this embodiment provides cross-temporal and multi-scale information retrieval and visualization capabilities for urban history research. Through semantic retrieval and visualization analysis, researchers can use this system to query cross-temporal and multi-scale historical information co-occurrence patterns, revealing the development trajectory, spatial pattern evolution, and complex relationships between ancient cities and historical events and figures. For example, it can retrieve and analyze city-building activities in specific periods to study the patterns of urban construction in different periods.
[0197] For cultural heritage protection, this knowledge graph can comprehensively integrate historical information of ancient cities, accurately record the spatiotemporal evolution of cultural heritage such as ancient buildings and historical settlements, and reveal their interactive relationship with the urban spatial pattern. Through knowledge association analysis, cultural heritage protection agencies can formulate scientific restoration and protection plans, ensuring the authenticity and integrity of historical information and providing a solid data basis for cultural heritage protection.
[0198] For contemporary urban planning and construction, this knowledge graph, through in-depth exploration of the planning concepts and experiences of ancient cities, provides a reference for optimizing modern urban spatial layout, adjusting functional zoning, and protecting historical urban areas. For example, in the context of new urbanization, we can draw on the site selection principles, ecological patterns, and spatial organization methods of ancient cities to provide historical experience and feasible solutions for modern urban and rural planning.
[0199] In summary, the spatiotemporal knowledge graph of ancient Chinese cities proposed in this invention breaks through the limitations of traditional research in the spatiotemporal dimension, forming a systematic and standardized historical knowledge expression system. This graph not only serves historical research and cultural heritage protection but also provides important insights for modern urban planning and decision-making, possessing significant academic value and broad application prospects.
[0200] Figure 4 A structural diagram of an apparatus for constructing a spatiotemporal knowledge graph of ancient Chinese cities according to an embodiment of this disclosure is shown. Figure 4 As shown, the device includes:
[0201] The acquisition module 401 is used to acquire ancient book documents related to the target city based on the set of city place names of the target city;
[0202] The sentence segmentation and knowledge extraction module 402 is used to perform sentence segmentation and knowledge extraction on ancient book documents to obtain multiple ancient book text entries and triples corresponding to multiple ancient book text entries. The triples include entities, entity attributes, and relationships between entities. The types of entities include time, space, and city. Time and space type entities each include multiple attributes of different granularities.
[0203] The semantic disambiguation module 403 is used to perform semantic disambiguation on each triple corresponding to each ancient text entry, and to determine the attribute value of the time and space type entity that uniquely corresponds to each triple.
[0204] Module 404 is used to construct a spatiotemporal knowledge graph of ancient Chinese cities based on the triples corresponding to multiple ancient text entries and the attribute values of the time and space types of entities that uniquely correspond to the triples. The spatiotemporal knowledge graph of ancient Chinese cities includes structured features that can support semantic retrieval and analysis related to ancient Chinese cities, so as to provide decision support for the digital protection of cultural heritage and modern urban planning.
[0205] In one possible implementation, the ancient text data includes first text data and second text data, and the acquisition module 401 is used for:
[0206] Extract historical place names, alternative names, abbreviations, and administrative division changes of the target city from the historical geography database to determine the set of city place names of the target city, which includes the ancient and modern names of the target city.
[0207] Based on the set of city place names of the target city, keyword matching is performed in the ancient book catalog database to obtain the ancient book catalog related to the target city, and the full text of the ancient book catalog is used as the first text data.
[0208] Based on the set of city place names of the target city, keyword matching is performed in the full-text database of ancient books to obtain the hit text related to the target city. The hit text and the context text of the hit text are used as the second text data.
[0209] In one possible implementation, the sentence segmentation and knowledge extraction module 402 is used for:
[0210] Natural language processing was used to segment ancient texts into sentences, resulting in multi-sentence ancient texts.
[0211] For each sentence of ancient text, generate an initial ancient text entry with a unique identifier. The initial ancient text entry includes the book title, volume number, content, period, version, author, ancient name and modern name of the ancient text.
[0212] The period and author in multiple initial ancient book text entries were corrected to obtain the corrected ancient book text entries;
[0213] Duplicate entries in the corrected ancient text entries are merged, and the earliest corrected ancient text entry corresponding to the period of the duplicate entries is retained, resulting in multiple ancient text entries.
[0214] For each ancient text entry, knowledge extraction is performed to obtain the triplet corresponding to each ancient text entry.
[0215] In one possible implementation, the period and author information in multiple initial ancient text entries are corrected to obtain corrected ancient text entries, including:
[0216] For any initial ancient text entry, in response to the existence of multiple periods in any initial ancient text entry, a layout analysis model is used to determine whether the content in the initial ancient text entry belongs to the main text area or the annotation area.
[0217] If the content in the initial ancient text entry belongs to the annotation area, the author corresponding to the initial ancient text entry is identified as the annotator, and the period corresponding to the initial ancient text entry is identified as the era corresponding to the annotator.
[0218] If the content of the initial ancient text entry belongs to the main text area, the author corresponding to the initial ancient text entry is determined as the original author, and the period corresponding to the initial ancient text entry is determined as the era corresponding to the original author.
[0219] In one possible implementation, the semantic disambiguation module 403 is used for:
[0220] For each triple corresponding to each ancient text entry, a spatiotemporal attribute candidate library is determined. The spatiotemporal attribute candidate library includes all candidate time and space information strings related to the triple.
[0221] A candidate time and space information string is selected from the candidate library of time and space attributes as the disambiguated time and space attribute information;
[0222] The disambiguated temporal and spatial attribute information corresponding to each triple is standardized to obtain the attribute values of the entity of the time and space type that uniquely corresponds to each triple.
[0223] In one possible implementation, for each triple corresponding to each ancient text entry, a candidate library of spatiotemporal attributes is determined, including:
[0224] The strings corresponding to the time and space types of entities in the triplet are used as the first candidate time and space information strings in the spatiotemporal attribute candidate library;
[0225] The strings corresponding to other time and space types of entities appearing in the content of the ancient text entries corresponding to the triplet are used as the second candidate time and space information strings in the candidate library of spatiotemporal attributes.
[0226] Based on the period and ancient name of the ancient text entry corresponding to the triple, the third candidate time and space information string in the candidate library of spatiotemporal attributes is determined;
[0227] A candidate time and space information string is determined from the candidate library of time and space attributes as the disambiguated time and space attribute information, including:
[0228] If a first candidate time and space information string exists in the candidate library of spatiotemporal attributes, the first candidate time and space information string shall be used as the disambiguated time and space attribute information.
[0229] If there is no first candidate time and space information string in the spatiotemporal attribute candidate library but there is a second candidate time and space information string, the disambiguated time and space attribute information is determined based on the second candidate time and space information string;
[0230] If there is no first candidate time and space information string or second candidate time and space information string in the spatiotemporal attribute candidate library, but there is a third candidate time and space information string, the disambiguated time and space attribute information is determined based on the third candidate time and space information string.
[0231] In one possible implementation, the types of entities also include people, events, buildings, settlements, geographical entities, and sources of information, with entities of the source of information type used to indicate different parts of ancient text entries;
[0232] The relationships between entities satisfy the following constraints: entities of the city, person, event, building, settlement, geographical entity, and document source type are associated with entities of the time type; entities of the city, person, event, building, settlement, and geographical entity types are associated with entities of the spatial type; and entities of the person, event, building, settlement, and geographical entity types are associated with entities of the document source type and city type.
[0233] In one possible implementation, the attributes of a time-type entity include attributes of time period and time point type. The attributes of the time period type include dynasty, reign title and other time ranges, and the attributes of the time point type include year.
[0234] The attributes of spatial type entities include attributes of spatial surface and spatial point types. The attributes of spatial surface type include high-level administrative regions, unified county administrative regions, county-level administrative regions, and other spatial ranges.
[0235] The attributes of city-type entities include ancient name, category, and modern name;
[0236] The attributes of a person-type entity include name, gender, and category;
[0237] The attributes of an event type entity include content and category;
[0238] The attributes of building-type entities, settlement-type entities, and geographical-type entities include name and category, respectively;
[0239] The attributes of the source type include the title, volume number, content, period, edition, and author of the ancient text entry.
[0240] According to the embodiments of this disclosure, ancient book documents related to the target city are obtained by using a set of city place names of the target city. The ancient book documents are processed to obtain multiple ancient book text entries and triples corresponding to the multiple ancient book text entries. The triples include entities, entity attributes, and relationships between entities. The types of entities include time, space, and city. Entities of time and space types each include multiple attributes of different granularities. This enables precise spatial positioning from macro-regions to micro-geographic units. By meeting the time requirements of different scales, the accuracy of knowledge graphs in analyzing, interpreting, and expressing spatiotemporal information is significantly improved. By performing semantic disambiguation on each triple corresponding to each ancient text entry, and determining the attribute value of the unique time and space type entity corresponding to each triple, each triple can possess traceable, multi-granular spatiotemporal positioning capabilities, achieving a closed-loop mapping of "text-entity-spatiotemporal-knowledge". Based on the triples corresponding to multiple ancient text entries and the attribute values of the unique time and space type entities corresponding to each triple, a spatiotemporal knowledge graph of ancient Chinese cities can be constructed. This allows for the construction of a spatiotemporal knowledge graph with ancient cities as the ontology, breaking through the limitations of traditional urban historical data modeling centered on documents or figures. It brings together diverse knowledge elements such as figures, events, and buildings, originally scattered across different spatiotemporal dimensions and record carriers, into a unified structured knowledge system through semantic association with core urban entities. The constructed spatiotemporal knowledge graph of ancient Chinese cities can support in-depth semantic retrieval and analysis related to ancient Chinese cities, providing efficient knowledge support and decision-making assistance for fields such as urban history research, cultural heritage protection, and urban planning and construction.
[0241] In some embodiments, the functions or modules of the apparatus provided in this disclosure can be used to perform the methods described in the above method embodiments. The specific implementation can be referred to the description of the above method embodiments, and for the sake of brevity, it will not be repeated here.
[0242] This disclosure also provides an apparatus for constructing a spatiotemporal knowledge graph of ancient Chinese cities, including a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the steps of the above method.
[0243] This disclosure also provides a non-volatile computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the steps of the above-described method.
[0244] This disclosure also provides a computer program product, including a computer program or a non-volatile computer-readable storage medium carrying the computer program, wherein the computer program, when executed by a processor, implements the steps of the above method.
[0245] Figure 5This is a block diagram illustrating an apparatus 1900 for constructing a spatiotemporal knowledge graph of ancient Chinese cities, according to an exemplary embodiment. For example, apparatus 1900 may be provided as a server or terminal device. (Refer to...) Figure 5 The apparatus 1900 includes a processing component 1922, which further includes one or more processors, and memory resources represented by memory 1932 for storing instructions, such as application programs, that can be executed by the processing component 1922. The application programs stored in memory 1932 may include one or more modules, each corresponding to a set of instructions. Furthermore, the processing component 1922 is configured to execute instructions to perform the methods described above.
[0246] Device 1900 may also include a power supply component 1926 configured to perform power management of device 1900, a wired or wireless network interface 1950 configured to connect device 1900 to a network, and an input / output interface 1958 (I / O interface). Device 1900 can operate on an operating system, such as Windows Server, stored in memory 1932. TM macOS X TM Unix TM Linux TM FreeBSD TM Or similar.
[0247] In an exemplary embodiment, a non-volatile computer-readable storage medium is also provided, such as a memory 1932 including computer program instructions that can be executed by a processing component 1922 of the device 1900 to perform the above-described method.
[0248] Computer-readable storage media can be tangible devices capable of holding and storing programs / instructions used by instruction execution devices. Computer-readable storage media can be, for example—but not limited to—electrical storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of computer-readable storage media include: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disc read-only memory (CD-ROM), digital multifunction disc (DVD), memory sticks, floppy disks, mechanical encoding devices, such as punch cards or recessed protrusions storing instructions thereon, and any suitable combination of the foregoing. The computer-readable storage media used herein are not to be construed as transient signals themselves, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides or other transmission media (e.g., light pulses through fiber optic cables), or electrical signals transmitted through wires.
[0249] The computer program (or computer-readable program instructions) described herein can be downloaded from a computer-readable storage medium to various computing / processing devices, or downloaded via a network, such as the Internet, local area network, wide area network, and / or wireless network, to an external computer or external storage device. The network may include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards them to the computer-readable storage medium in the respective computing / processing device.
[0250] The computer program (or computer program instructions) used to perform the operations of this disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk, C++, etc., and conventional procedural programming languages such as the "C" language or similar programming languages. The computer-readable program instructions may execute entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, electronic circuitry, such as programmable logic circuitry, field-programmable gate arrays (FPGAs), or programmable logic arrays (PLAs), is personalized by utilizing state information from the computer-readable program instructions to implement various aspects of this disclosure.
[0251] Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0252] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processor of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner; thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0253] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.
[0254] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction containing one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than those shown in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.
[0255] The various embodiments of this disclosure have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is chosen to best explain the principles, practical application, or technical improvements to the embodiments in the market, or to enable others skilled in the art to understand the embodiments disclosed herein.
Claims
1. A method for constructing a spatiotemporal knowledge graph of ancient Chinese cities, characterized in that, The method includes: Based on the set of city place names of the target city, obtain ancient book documents related to the target city; Based on the ancient text data, sentence segmentation and knowledge extraction are performed to obtain multiple ancient text entries and corresponding triples. Each triple includes an entity, the attributes of the entity, and the relationship between the entities. The types of the entities include time, space, and city. The time and space type entities each include multiple attributes of different granularities. Semantic disambiguation is performed on each triple corresponding to each of the ancient text entries to determine the attribute value of the time and space type entity that uniquely corresponds to each triple; Based on the triples corresponding to multiple ancient text entries and the attribute values of the entities of time and space types that the triples uniquely correspond to, a spatiotemporal knowledge graph of ancient Chinese cities is constructed. The spatiotemporal knowledge graph of ancient Chinese cities includes structured features that can support semantic retrieval and analysis related to ancient Chinese cities, so as to provide decision support for the digital protection of cultural heritage and modern urban planning.
2. The method according to claim 1, characterized in that, The ancient text data includes first text data and second text data. Based on the set of city place names of the target city, ancient text data related to the target city is obtained, including: Extract historical place names, alternative names, abbreviations, and administrative division changes of the target city from the historical geography database to determine the set of city place names of the target city, which includes the ancient and modern names of the target city. Based on the set of city place names of the target city, keyword matching is performed in the ancient book catalog database to obtain the ancient book catalog related to the target city, and the full text of the ancient book catalog is used as the first text data. Based on the set of city place names of the target city, keyword matching is performed in the full-text database of ancient books to obtain the hit text related to the target city. The hit text and the context text of the hit text are used as the second text data.
3. The method according to claim 1, characterized in that, Based on the aforementioned ancient text data, sentence segmentation and knowledge extraction are performed to obtain multiple ancient text entries and their corresponding triples, including: The ancient text data was segmented using a natural language model to obtain multiple sentences of ancient text. For each sentence of ancient text, generate an initial ancient text entry identified by a unique identifier. The initial ancient text entry includes the book title, volume number, content, period, version, author, ancient name, and modern name corresponding to the ancient text. The periods and authors in the multiple initial ancient book text entries are corrected to obtain the corrected ancient book text entries; Duplicate entries in the corrected ancient text entries are merged, and the earliest corrected ancient text entry corresponding to the period of the duplicate entries is retained to obtain the plurality of ancient text entries. Each ancient text entry undergoes knowledge extraction processing to obtain a triplet corresponding to each ancient text entry.
4. The method according to claim 3, characterized in that, The periods and authors in multiple initial ancient text entries were corrected to obtain corrected ancient text entries, including: For any initial ancient text entry, in response to the presence of multiple periods in the initial ancient text entry, a layout analysis model is used to determine whether the content in the initial ancient text entry belongs to the main text area or the annotation area. If the content in the initial ancient text entry belongs to the annotation area, the author corresponding to the initial ancient text entry is identified as the annotator, and the period corresponding to the initial ancient text entry is identified as the era corresponding to the annotator. If the content in the initial ancient text entry belongs to the main text area, the author corresponding to the initial ancient text entry is determined as the original author, and the period corresponding to the initial ancient text entry is determined as the era corresponding to the original author.
5. The method according to claim 1, characterized in that, Semantic disambiguation is performed on each triple corresponding to each of the aforementioned ancient text entries to determine the attribute values of the time and space type entities uniquely corresponding to each triple, including: For each triple corresponding to each ancient text entry, a spatiotemporal attribute candidate library is determined, which includes all candidate time and space information strings related to the triple; A candidate time and space information string is determined from the spatiotemporal attribute candidate library as the disambiguated time and space attribute information; The disambiguated time and space attribute information corresponding to each triple is standardized to obtain the attribute values of the time and space type entities that uniquely correspond to each triple.
6. The method according to claim 5, characterized in that, For each triple corresponding to each ancient text entry, a candidate library of spatiotemporal attributes is determined, including: The strings corresponding to the time and space type entities in the triplet are used as the first candidate time and space information strings in the spatiotemporal attribute candidate library; The strings corresponding to other time and space types of entities appearing in the content of the ancient text entries corresponding to the triplet are used as the second candidate time and space information strings in the spatiotemporal attribute candidate library; Based on the period and ancient name of the ancient text entry corresponding to the triple, the third candidate time and space information string in the spatiotemporal attribute candidate library is determined; The step of determining a candidate time and space information string from the spatiotemporal attribute candidate library as the disambiguated time and space attribute information includes: If the first candidate time and space information string exists in the spatiotemporal attribute candidate library, the first candidate time and space information string shall be used as the disambiguated time and space attribute information. If the first candidate time and space information string does not exist in the spatiotemporal attribute candidate library but the second candidate time and space information string exists, the disambiguated time and space attribute information is determined based on the second candidate time and space information string; If the first candidate time and space information string and the second candidate time and space information string do not exist in the spatiotemporal attribute candidate library, but the third candidate time and space information string exists, the disambiguated time and space attribute information is determined based on the third candidate time and space information string.
7. The method according to claim 1, characterized in that, The types of entities also include people, events, buildings, settlements, geographical entities, and sources of documents. Entities of the source of documents type are used to indicate the various parts of the ancient text entries. The relationships between the entities satisfy the following constraints: entities of the city, person, event, building, settlement, geographical entity, and document source type are associated with entities of the time type; entities of the city, person, event, building, settlement, and geographical entity type are associated with entities of the spatial type; and entities of the person, event, building, settlement, and geographical entity type are associated with entities of the document source type and city type.
8. The method according to claim 7, characterized in that, The attributes of the time-type entity include time period and time point type attributes. The attributes of the time period type include dynasty, reign title and other time ranges, and the attributes of the time point type include year. The attributes of the spatial type entity include attributes of spatial surface and spatial point types. The attributes of the spatial surface type include high-level administrative regions, unified county administrative regions, county-level administrative regions, and other spatial ranges. The attributes of the city type entity include ancient name, category, and modern name; The attributes of the person type entity include name, gender, and category; The attributes of the event type entity include content and category; The attributes of the building type entity, settlement type entity, and geographical type entity include name and category, respectively; The attributes of the document source type include the book title, volume number, content, period, version, and author of the ancient book text entry.
9. A device for constructing a spatiotemporal knowledge graph of ancient Chinese cities, characterized in that, The device includes: The acquisition module is used to acquire ancient book and document data related to the target city based on the set of city place names of the target city; The sentence segmentation and knowledge extraction module is used to perform sentence segmentation and knowledge extraction on the ancient book document data to obtain multiple ancient book text entries and triples corresponding to the multiple ancient book text entries. The triples include entities, attributes of the entities, and relationships between the entities. The types of the entities include time, space, and city. The time and space type entities each include multiple attributes of different granularities. The semantic disambiguation module is used to perform semantic disambiguation on each triple corresponding to each of the ancient text entries, and to determine the attribute value of the time and space type entity that uniquely corresponds to each triple; The construction module is used to construct a spatiotemporal knowledge graph of ancient Chinese cities based on the triples corresponding to multiple ancient text entries and the attribute values of the entities of time and space types that the triples uniquely correspond to. The spatiotemporal knowledge graph of ancient Chinese cities includes structured features that can support semantic retrieval and analysis related to ancient Chinese cities, so as to provide decision support for the digital protection of cultural heritage and modern urban planning.
10. A device for constructing a spatiotemporal knowledge graph of ancient Chinese cities, comprising a memory, a processor, and a computer program stored in the memory, characterized in that, The processor executes the computer program to implement the steps of the method according to any one of claims 1 to 8.
11. A non-volatile computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 8.
12. A computer program product comprising a computer program, or a non-volatile computer-readable storage medium carrying a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 8.