Urban multi-department heterogeneous data knowledge graph construction method and system, terminal and storage medium

By constructing a knowledge graph of heterogeneous data in urban multiple departments, determining the ontology model, extracting data information and calculating the semantic similarity of space-time semantic similarity, the semantic islands and co-referential digestion problems of urban multi-department data are solved, and efficient integration of cross-department data and intelligent decision-making support are achieved.

CN120338076AActive Publication Date: 2025-07-18SHENZHEN UNIV

Patent Information

Application Number
CN202510804679.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-17
Publication Date
2025-07-18
Estimated Expiration
2045-06-17

AI Technical Summary

Technical Problem

There are semantic islands and entities that combine the problem of eliminating difficulties in existing urban multi-department heterogeneous data. The existing methods fail to effectively correlate and integrate data from different departments and fields from the perspective of urban comprehensive decision makers.

Method used

By constructing a knowledge graph of heterogeneous data in cities with multi-department, determining the entity objects, description attributes and association relationships of the ontology model, extracting data information using the rule engine, and combining the storage structure of the graph database, calculating the spatiotemporal semantic similarity to fusion and co-referencing problems, realizing the semantic association and entity co-referencing of cross-departmental data.

Benefits of technology

It has realized unified storage of heterogeneous data in cities with multi-departmental departments and efficient integration of cross-departmental data, providing intelligent decision-making support from a global perspective, breaking data barriers, and supporting smart city governance and cross-departmental collaborative analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120338076A_ABST
    Figure CN120338076A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of information processing, and discloses an urban multi-department heterogeneous data knowledge graph construction method and system, a terminal and a storage medium, and the method comprises the steps: determining entity objects, description attributes and association relationships of an urban multi-department heterogeneous data ontology model; according to the entity object, the description attribute and the incidence relation, using a rule engine to extract entity records, attribute information and relation information from the city multi-department heterogeneous data; organizing the entity objects, the entity records, the attribute information and the relation information in combination with a storage structure of a graph database to obtain an initial knowledge graph; and calculating a space-time semantic similarity according to the initial knowledge graph, and fusing different entity records with co-reference problems according to the space-time semantic similarity to obtain a target knowledge graph. According to the method, semantic association, knowledge integration and entity co-reference resolution of cross-department data are realized by constructing the knowledge graph, and intelligent decision support from a global perspective is provided for urban governance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of information processing, and particularly relates to a method, system, terminal, and computer-readable storage medium for constructing a knowledge graph of heterogeneous data of multiple urban departments. Background Art

[0002] With the accelerating promotion of the construction of new smart cities, the amount of data generated by urban operations shows exponential growth, and the data types cover geospatial information, building information models, Internet of Things perception data, and structured and semi-structured data in the business systems of various departments. These data are scattered in different departments such as planning, housing construction, transportation, environment, and public security, forming a complex multi-source heterogeneous data ecosystem. Due to the adoption of independent data standards and storage architectures by each department and the differences in description dimensions in different systems, there are semantic islands among the data of multiple urban departments.

[0003] In current research, knowledge graphs have been widely used to solve the problem of data semantic islands. By converting heterogeneous data into a graph structure, relationships between different entities can be established in the graph structure to achieve cross-domain information fusion and semantic association. However, existing urban knowledge graph research often focuses on vertical application fields and fails to associate and integrate data from different departments and domains from the perspective of urban comprehensive decision-makers. At the same time, for the problem of difficult coreference resolution caused by inconsistent spatio-temporal benchmarks and entity concepts of data in each department, existing methods have not organized and integrated data from the perspective of entity-based thinking and spatio-temporal semantic similarity, and thus cannot achieve a global view and cross-departmental collaboration.

[0004] Therefore, the existing technology still needs to be improved and developed. Summary of the Invention

[0005] The main purpose of the present invention is to provide a method, system, terminal, and computer-readable storage medium for constructing a knowledge graph of heterogeneous data of multiple urban departments, aiming to solve the problems of semantic islands and difficult coreference resolution of existing heterogeneous data of multiple urban departments.

[0006] To achieve the above invention purpose, the present invention provides a method for constructing a knowledge graph of heterogeneous data of multiple urban departments, and the method for constructing a knowledge graph of heterogeneous data of multiple urban departments includes: Determine the entity objects, description attributes, and association relationships of the ontology model of heterogeneous data of multiple urban departments; According to the entity objects, the description attributes, and the association relationships, use a rule engine to extract entity records, attribute information, and relationship information from the heterogeneous data of multiple urban departments; Combine the storage structure of the graph database to organize the entity objects, the entity records, the attribute information, and the relationship information to obtain an initial knowledge graph; Calculate the spatio-temporal semantic similarity according to the initial knowledge graph, and fuse different entity records with co-reference problems according to the spatio-temporal semantic similarity to obtain a target knowledge graph.

[0007] Optionally, the determination of the entity objects, description attributes, and association relationships of the urban multi-department heterogeneous data ontology model specifically includes: Obtain urban multi-department heterogeneous data, determine the key entities of the urban multi-department heterogeneous data ontology model in combination with the urban multi-department heterogeneous data, and classify and grade the key entities to obtain entity types and instance objects, where the entity types and the instance objects constitute the entity objects; Determine the description attributes and association relationships of the urban multi-department heterogeneous data ontology model in combination with the characteristics of the urban multi-department heterogeneous data.

[0008] Optionally, the urban multi-department heterogeneous data includes relational data, GIS data, BIM data, and IoT data; The extraction of entity records, attribute information, and relationship information from urban multi-department heterogeneous data using a rule engine according to the entity objects, the description attributes, and the association relationships specifically includes: According to the entity objects, the description attributes, and the association relationships in the urban multi-department heterogeneous data ontology model, extract the entity records corresponding to the entity objects, the attribute information corresponding to the description attributes, and the relationship information corresponding to the association relationships from the relational data based on the SQL rule engine; According to the entity objects, the description attributes, and the association relationships in the urban multi-department heterogeneous data ontology model, extract the entity records corresponding to the entity objects, the attribute information corresponding to the description attributes, and the relationship information corresponding to the association relationships from the GIS data based on the Geopandas tool; According to the entity objects, the description attributes, and the association relationships in the urban multi-department heterogeneous data ontology model, extract the entity records corresponding to the entity objects, the attribute information corresponding to the description attributes, and the relationship information corresponding to the association relationships from the BIM data based on the Revit API; According to the entity objects, the description attributes, and the association relationships in the urban multi-department heterogeneous data ontology model, extract the entity records corresponding to the entity objects, the attribute information corresponding to the description attributes, and the relationship information corresponding to the association relationships from the IoT data based on the InfluxDBClient engine.

[0009] Optionally, organize the entity objects, entity records, attribute information, and relationship information according to the storage structure of the combined graph database to obtain an initial knowledge graph, specifically including: According to the node-attribute-edge storage structure of the graph database, map the entity types, instance objects, and entity records to concept nodes with concept labels, instance nodes with instance labels, and record nodes with record labels respectively at three levels: the concept layer, the instance layer, and the record layer; Use the key-value pair attributes embedded in the record nodes to represent the attribute information; Use the typed edges between the record nodes to represent the relationship information, and obtain the initial knowledge graph.

[0010] Optionally, calculate the spatio-temporal semantic similarity according to the initial knowledge graph, and fuse different entity records with co-reference problems according to the spatio-temporal semantic similarity to obtain a target knowledge graph, specifically including: Unify the spatio-temporal benchmarks of different entity records corresponding to different instance objects under the same entity type in the initial knowledge graph to obtain different target entity records corresponding to different instance objects under the same entity type; Calculate the time similarity, space similarity, and attribute similarity between the different target entity records respectively; Perform a weighted sum of the time similarity, the space similarity, and the attribute similarity to obtain the spatio-temporal semantic similarity; Compare the spatio-temporal semantic similarity with a similarity threshold. If the spatio-temporal semantic similarity is not less than the similarity threshold, it is considered that the different entity records are different entity records with co-reference problems corresponding to different instance objects under the same entity type; Retain the target nodes in the instance nodes of different instance objects corresponding to different entity records with co-reference problems, and at the same time remove the redundant nodes in the instance nodes of different instance objects corresponding to different entity records with co-reference problems, and incorporate the information of the redundant nodes into the target nodes; Wherein, the target node is the instance node with the largest amount of information or the highest information reliability in the instance nodes of different instance objects corresponding to different entity records with co-reference problems, and the redundant node is the instance node other than the target node in the instance nodes of different instance objects corresponding to different entity records with co-reference problems.

[0011] Optionally, the calculation of the time similarity, space similarity, and attribute similarity between the different target entity records respectively specifically includes: Calculate the time similarity between the different target entity records according to the different time information corresponding to the different target entity records: ; ; ; wherein, represents the th target entity record, represents the th target entity record, represents the time similarity between the target entity record and the target entity record ; represents the length of the overlapping interval between the target entity record and the target entity record ; represents the length of the combined interval between the target entity record and the target entity record ; represents the left endpoint of the time interval of the target entity record , ; represents the right endpoint of the time interval of the target entity record , ; represents the left endpoint of the time interval of the target entity record , ; represents the right endpoint of the time interval of the target entity record , ; represents taking the maximum value, represents taking the minimum value; Calculate the spatial similarity between the different target entity records according to the different spatial information corresponding to the different target entity records: ; ; ; ; wherein, represents the spatial similarity between the target entity record and the target entity record ; represents the spatial representation of the target entity record , represents the target entity record Spatial representation Denote spatial representation And spatial representation The distance similarity between Denote spatial representation And spatial representation The topological similarity between Denote spatial representation And spatial representation The regional similarity between Denote the weight of distance similarity Denote the weight of topological similarity Denote the weight of regional similarity Denote spatial representation And spatial representation The distance between Denote the maximum acceptable distance threshold Denote spatial representation And spatial representation The topological relationship between Denote spatial representation And spatial representation The area of the intersection part between Denote spatial representation And spatial representation The area of the union part between; According to the different target entity records, record the different attribute information under each attribute type, and calculate the attribute similarity between the different target entity records: ; ; ; ; Wherein, Denote the target entity record And the target entity record The attribute similarity between Denote the th attribute type Denote the weight of the th attribute type Denote the target entity record And the target entity record The sub-attribute similarity under the th attribute type between Denote the sub-attribute similarity under the 1st attribute type Denote the target entity record The string instance, represents the target entity record The string instance, represents the string instance The edit distance between the string instance and the string instance represents the string instance The length of, represents the string instance The length of, represents the sub - attribute similarity under the second type of attribute type represents a set score less than 1 and greater than 0.5 represents the target entity record The category of, represents the target entity record The category of, represents the category and the category The relationship between, represents the sub - attribute similarity under the third type of attribute type represents the target entity record The value of, represents the target entity record The value of, represents the maximum possible value threshold.

[0012] Optionally, the weighted sum of the time similarity, the space similarity, and the attribute similarity is calculated to obtain the spatio - temporal semantic similarity, specifically including: According to the pre - set weight coefficients of the time similarity, the space similarity, and the attribute similarity, the weighted sum of the time similarity between different target entity records, the space similarity between different target entity records, and the attribute similarity between different target entity records is calculated to obtain the spatio - temporal semantic similarity between different target entity records: ; Among them, is the spatio - temporal semantic similarity between the target entity record and the target entity record , is the weight coefficient of the time similarity, is the weight coefficient of the space similarity, is the weight coefficient of the attribute similarity.

[0013] To achieve the above - mentioned invention purpose, the present invention also provides a system for constructing a knowledge graph of heterogeneous data of multiple departments in a city. The system for constructing a knowledge graph of heterogeneous data of multiple departments in a city includes: Ontology model construction module: used to determine the entity objects, descriptive attributes, and association relationships of the urban multi-department heterogeneous data ontology model; Knowledge extraction module: used to extract entity records, attribute information, and relationship information from urban multi-department heterogeneous data using a rule engine according to the entity objects, the descriptive attributes, and the association relationships; Initial knowledge graph construction module: used to organize the entity objects, the entity records, the attribute information, and the relationship information in combination with the storage structure of the graph database to obtain an initial knowledge graph; Target knowledge graph construction module: used to calculate the spatio-temporal semantic similarity according to the initial knowledge graph, and fuse different entity records with co-reference problems according to the spatio-temporal semantic similarity to obtain a target knowledge graph.

[0014] To achieve the above invention purposes, the present invention also provides a terminal, the terminal includes: a memory, a processor, and a program for constructing a knowledge graph of urban multi-department heterogeneous data stored on the memory and executable on the processor. When the program for constructing a knowledge graph of urban multi-department heterogeneous data is executed by the processor, the steps of the method for constructing a knowledge graph of urban multi-department heterogeneous data as described above are implemented.

[0015] To achieve the above invention purposes, the present invention also provides a computer-readable storage medium, the computer-readable storage medium stores a program for constructing a knowledge graph of urban multi-department heterogeneous data. When the program for constructing a knowledge graph of urban multi-department heterogeneous data is executed by a processor, the steps of the method for constructing a knowledge graph of urban multi-department heterogeneous data as described above are implemented.

[0016] In the present invention, entity objects, description attributes, and association relationships of an ontology model for heterogeneous data of multiple urban departments are determined; according to the entity objects, the description attributes, and the association relationships, an entity record, attribute information, and relationship information are extracted from the heterogeneous data of multiple urban departments using a rule engine; in combination with the storage structure of a graph database, the entity objects, the entity records, the attribute information, and the relationship information are organized to obtain an initial knowledge graph; the spatio-temporal semantic similarity is calculated based on the initial knowledge graph, and different entity records with co-reference problems are fused according to the spatio-temporal semantic similarity to obtain a target knowledge graph. Through a unified ontology model, a multi-source extraction mechanism, and spatio-temporal semantic fusion technology, from the perspective of entities and spatio-temporal semantic similarity, the data of different departments and fields are organized and fused from the perspective of urban comprehensive decision-makers, realizing semantic association, knowledge integration, and entity co-reference resolution of cross-department data, solving the problems of semantic islands and difficult entity co-reference resolution existing in the existing heterogeneous data of multiple urban departments, and providing intelligent decision-making support with a global perspective for urban governance; at the same time, by constructing a knowledge graph, the unified storage of entities, attributes, and relationships in the heterogeneous data of multiple urban departments is realized, which helps to associate the heterogeneous data of multiple urban departments and break the barriers between the heterogeneous data of multiple urban departments. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 is a flowchart of a preferred embodiment of a method for constructing a knowledge graph of heterogeneous data of multiple urban departments of the present invention; Figure 2 is a schematic diagram of an ontology model of heterogeneous data of multiple urban departments of the present invention; Figure 3 is a schematic diagram of an initial knowledge graph of the present invention; Figure 4 is a schematic diagram of fusing different entity records with co-reference problems of the present invention; Figure 5 is a structural diagram of a preferred embodiment of a system for constructing a knowledge graph of heterogeneous data of multiple urban departments of the present invention; Figure 6 is a structural diagram of a preferred embodiment of a terminal of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0018] To make the objectives, technical solutions, and advantages of the present invention clearer and more definite, the following further describes the present invention in detail with reference to the accompanying drawings and by way of examples. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0019] With the accelerating advancement of the construction of new smart cities, the amount of data generated by urban operations has shown exponential growth. The data types cover geospatial information, building information models, Internet of Things perception data, as well as structured and semi-structured data in the business systems of various departments. These data are scattered in different departments such as planning, housing construction, transportation, environment, and public security, forming a complex multi-source heterogeneous data ecosystem. Due to the adoption of independent data standards and storage architectures by each department and the differences in description dimensions in different systems, there are semantic islands among the data of multiple urban departments. In practical applications, the same entity (such as buildings, transportation facilities, population) appears repeatedly in the data tables of different departments, but there is a lack of effective association between these records, further affecting the efficiency of data sharing and cross-departmental collaborative work.

[0020] In current research, knowledge graphs have been widely used to solve the problem of data semantic islands. By transforming heterogeneous data into a graph structure, it is possible to establish relationships between different entities in the graph structure, realizing cross-domain information fusion and semantic association. For example, urban knowledge graphs can organize data in fields such as transportation, buildings, and public security through a unified knowledge framework, thereby providing support for urban management and decision-making.

[0021] However, existing urban knowledge graph research often focuses on vertical application fields and fails to associate and integrate data from different departments and fields from the perspective of urban comprehensive decision-makers. At the same time, regarding the problem of difficult coreference resolution caused by inconsistent spatio-temporal benchmarks and inconsistent entity concepts in the data of each department, existing methods have not organized and fused the data from the perspective of entity-based thinking and spatio-temporal semantic similarity, and thus cannot achieve a global view and cross-departmental collaboration.

[0022] To solve the above technical problems, the present invention provides a method for constructing a knowledge graph of heterogeneous data of multiple urban departments, which determines the entity objects, description attributes, and association relationships of the ontology model of heterogeneous data of multiple urban departments; according to the entity objects, the description attributes, and the association relationships, uses a rule engine to extract entity records, attribute information, and relationship information from the heterogeneous data of multiple urban departments; combines the storage structure of the graph database to organize the entity objects, the entity records, the attribute information, and the relationship information to obtain an initial knowledge graph; calculates the spatio-temporal semantic similarity according to the initial knowledge graph, and fuses different entity records with co-reference problems according to the spatio-temporal semantic similarity to obtain a target knowledge graph. Through a unified ontology model, a multi-source extraction mechanism, and spatio-temporal semantic fusion technology, the present invention organizes and fuses data from different departments and fields from the perspective of entities and spatio-temporal semantic similarity from the perspective of urban comprehensive decision-makers, realizes semantic association, knowledge integration, and entity co-reference resolution of cross-department data, solves the problems of semantic islands and difficult entity co-reference resolution existing in the existing heterogeneous data of multiple urban departments, and provides intelligent decision-making support from a global perspective for urban governance; at the same time, by constructing a knowledge graph, the unified storage of entities, attributes, and relationships in the heterogeneous data of multiple urban departments is realized, which helps to associate the heterogeneous data of multiple urban departments and break the barriers between the heterogeneous data of multiple urban departments.

[0023] The following further illustrates the application content by describing the embodiments in conjunction with the accompanying drawings.

[0024] A preferred embodiment of the method for constructing a knowledge graph of heterogeneous data of multiple urban departments of the present invention is as Figure 1 shown, and specifically includes: S1. Determine the entity objects, description attributes, and association relationships of the ontology model of heterogeneous data of multiple urban departments.

[0025] In an implementation manner of this embodiment, the determination of the entity objects, description attributes, and association relationships of the ontology model of heterogeneous data of multiple urban departments specifically includes: Obtain the heterogeneous data of multiple urban departments, combine the heterogeneous data of multiple urban departments to determine the key entities of the ontology model of heterogeneous data of multiple urban departments, and classify and grade the key entities to obtain entity types and instance objects, where the entity types and the instance objects constitute the entity objects; Combine the characteristics of the heterogeneous data of multiple urban departments to determine the description attributes and association relationships of the ontology model of heterogeneous data of multiple urban departments; Among them, the entity objects include seven major categories: natural resources and environment, topics and services, transportation infrastructure, urban management and services, population and society, building facilities, and land and planning, as well as sub-categories corresponding to the seven major categories; the description attributes include ID, standard name, index information, department source, data source, absolute location description, relative address description, timestamp, status, geometry, material, and other information; the association relationships include located in, contains, adjacent to, before and after, superior and subordinate, affiliated, dependent, and whole-part.

[0026] Specifically, referring to relevant standards such as "Classification and Codes for Fundamental Geographic Information Elements" and "OGC Geography Markup Language Encoding Standard" (where OGC is the Open Geospatial Consortium, representing the Open Geospatial Alliance), combined with the multi-source heterogeneous characteristics of urban data, a hierarchical semantic framework (i.e., an ontology model) is established, as Figure 2 shown. First, entity types (including instance objects) are constructed. Entities are identifiable objects in the city that have management requirements and specific attributes and functions, including seven major categories: natural resources and environment, topics and services, transportation infrastructure, urban management and services, population and society, building facilities, and land and planning, as well as corresponding middle categories and sub-categories; attribute types are constructed. Attributes describe the necessary and general attributes of urban multi-department entity objects for retrieval and application, including multiple categories and corresponding attributes such as identification, indexing, spatio-temporal, and features; relationship types are constructed. Relationships are used to express the connections between different urban multi-department entity objects for retrieval and application, including multiple categories and corresponding relationships such as space, time, hierarchy, and function. To realize the design of the ontology model for urban multi-department heterogeneous data, the ontology model constructed by the present invention needs to include the following five principles: Principle 1, knowledge related to urban data indexing and retrieval should be able to be transformed into the form of triples, that is, entities, attributes, and relationships; Principle 2, the extracted triple knowledge needs to be able to provide useful information for users to retrieve data; Principle 3, the construction of the ontology model should ensure the accuracy and consistency of semantic expression to avoid affecting data fusion due to different expression methods; Principle 4, when there are ambiguities in the terms used for the same concept, the definition in a more authoritative and practical standard should be selected; Principle 5, for the problem of possible different granularity expressions in data sources, support the use of different records for multi-level representation, allowing multiple records of the same entity object to coexist. By formulating the above principles, the ontology model constructed by the present invention can ensure semantic accuracy while being compatible with different data standards, supporting multi-granularity data expression, and optimizing the data organization method. It should be noted that Figure 2 the attributes herein refer to description attributes, and the relationships refer to association relationships.

[0027] Specifically, the present invention combines the standards related to heterogeneous data of multiple urban departments to determine the key entities included in the ontology model of heterogeneous data of multiple urban departments, and classifies and grades a large number of entities (i.e., key entities) in the city to obtain entity types and instance objects (collectively referred to as entity objects). The urban entity objects are recognizable objects in the city that have management requirements and specific attributes and functions. Then, referring to the designed entity objects and combining the characteristics of heterogeneous data of multiple urban departments, descriptive attributes and association relationships are designed. The descriptive attributes are used to represent the index information and general attributes of the data records corresponding to different urban entity objects, and the association relationships are used to represent the dependency relationships between the data records corresponding to different urban entity objects.

[0028] S2. According to the entity objects, the descriptive attributes, and the association relationships, use a rule engine to extract entity records, attribute information, and relationship information from the heterogeneous data of multiple urban departments.

[0029] In an implementation manner of this embodiment, the heterogeneous data of multiple urban departments includes relational data, GIS (Geographic Information System) data, BIM (Building Information Modeling) data, and IoT (Internet of Things) data. The step of using a rule engine to extract entity records, attribute information, and relationship information from the heterogeneous data of multiple urban departments according to the entity objects, the descriptive attributes, and the association relationships specifically includes: According to the entity objects, the descriptive attributes, and the association relationships in the ontology model of the heterogeneous data of multiple urban departments, use an SQL (Structured Query Language) rule engine to extract the entity records corresponding to the entity objects, the attribute information corresponding to the descriptive attributes, and the relationship information corresponding to the association relationships from the relational data. According to the entity objects, the descriptive attributes, and the association relationships in the ontology model of the heterogeneous data of multiple urban departments, use the Geopandas tool to extract the entity records corresponding to the entity objects, the attribute information corresponding to the descriptive attributes, and the relationship information corresponding to the association relationships from the GIS data. According to the entity objects, the description attributes, and the association relationships in the urban multi-department heterogeneous data ontology model, entity records corresponding to the entity objects, attribute information corresponding to the description attributes, and relationship information corresponding to the association relationships are extracted from the BIM data in Revit format based on the Revit API (Revise instantly Application Programming Interface). According to the entity objects, the description attributes, and the association relationships in the urban multi-department heterogeneous data ontology model, entity records corresponding to the entity objects, attribute information corresponding to the description attributes, and relationship information corresponding to the association relationships are extracted from the IoT data stored in InfluxDB based on the InfluxDBClient engine.

[0030] Specifically, the urban multi-source heterogeneous data targeted by the present invention includes relational data, GIS data, BIM data, and IoT data. A rule engine is used to extract knowledge from the relational database. For example, based on the SQL rule engine, the primary key of the relational data is extracted as the entity record, and the foreign key association is used as the relationship. A rule engine is used to extract knowledge from the GIS database. For example, based on the Geopandas tool to parse the GIS data file, spatial entities and their hierarchical relationships are extracted from the urban GIS scene. A rule engine is used to extract knowledge from the BIM database. For example, by parsing the Revit format file (i.e., the BIM data in Revit format) through the Revit API, room entities and their material attributes are extracted, and the component hierarchical relationship is mapped to an "include" edge. A rule engine is used to extract knowledge from the IoT database. For example, based on the InfluxDBClient engine, sensor node time-series data is extracted from the InfluxDB time-series database (one of the IoT databases), generating entity records with timestamps, and the device topology relationship is converted into a "connect" edge.

[0031] S3. Combine the storage structure of the graph database to organize the entity objects, the entity records, the attribute information, and the relationship information to obtain an initial knowledge graph.

[0032] In an implementation manner of this embodiment, the combining the storage structure of the graph database to organize the entity objects, the entity records, the attribute information, and the relationship information to obtain an initial knowledge graph specifically includes: Combining the node-attribute-edge storage structure of the graph database, the entity types, the instance objects, and the entity records are respectively mapped to concept nodes with concept labels, instance nodes with instance labels, and record nodes with record labels according to three levels: the concept layer, the instance layer, and the record layer. Represent the attribute information using the key-value pair attributes embedded in the record nodes; Represent the relationship information using the typed edges between the record nodes to obtain an initial knowledge graph.

[0033] Specifically, the present invention combines the storage structure of a graph database to realize the organization and association of design concepts (i.e., entity concepts) and extracted knowledge, including: regarding the entity concepts (i.e., entity objects, including entity types and instance objects) determined in step S1 and the entity records extracted in step S3 according to the entity concepts determined in step S1 as entity nodes of a graph model / graph structure / graph database, and distinguishing levels through labels, that is, entities are organized according to three levels: concept, instance, and record. Among them, the concept layer defines the schema of entity types and their attribute relationships; the instance layer represents specific entity objects (i.e., instance objects) and attributes; the record layer stores the original record data (i.e., entity records) from different sources and links them to the instance layer nodes through relationship edges. Store the attribute information extracted in step S3 according to the described attributes determined in step S1 as key-value pair attributes embedded in the graph model in the nodes; organize the relationship information extracted in step S3 according to the association relationships determined in step S1 as typed edge relationships of the graph model. The graph model / graph structure obtained after organization is the initial knowledge graph, as Figure 3 shown. It should be noted that the entity concepts, described attributes, and association relationships determined in step S1 are the entity objects, described attributes, and association relationships in the ontology model / semantic framework. The entity objects, described attributes, and association relationships are equivalent to design concepts, and step S3 is to extract corresponding knowledge according to this design concept. Figure 3 is a schematic diagram of the initial knowledge graph (i.e., a schematic diagram of the organization method of the knowledge graph of heterogeneous data of multiple urban departments). The concept layer expresses the classification and grading situations between different urban entity types and is connected to specific instance objects (such as the resident population 1) in the instance layer on this basis. Each instance object is associated with each corresponding record extracted from heterogeneous data in the record layer, and there are also different relationships between the records; it should be noted that from concept to instance and from instance to record is not a one-to-one mapping, but a one-to-many mapping. Among them, entity types, instance objects, and entity records are collectively referred to as entities.

[0034] S4. Calculate the spatio-temporal semantic similarity according to the initial knowledge graph, and fuse different entity records with co-reference problems according to the spatio-temporal semantic similarity to obtain a target knowledge graph.

[0035] In one implementation of this embodiment, calculating the spatio-temporal semantic similarity according to the initial knowledge graph, and fusing different entity records with co-reference problems according to the spatio-temporal semantic similarity to obtain a target knowledge graph specifically includes: Unify the spatio-temporal benchmarks of different entity records corresponding to different instance objects under the same entity type in the initial knowledge graph to obtain different target entity records corresponding to different instance objects under the same entity type; Calculate the time similarity, space similarity, and attribute similarity between the different target entity records respectively; Perform a weighted sum of the time similarity, the space similarity, and the attribute similarity to obtain the spatio-temporal semantic similarity; Compare the spatio-temporal semantic similarity with a similarity threshold. If the spatio-temporal semantic similarity is not less than the similarity threshold, it is considered that the different entity records are different entity records with co-reference problems corresponding to different instance objects under the same entity type; Retain the target nodes in the instance nodes of different instance objects corresponding to different entity records with co-reference problems, and at the same time remove the redundant nodes in the instance nodes of different instance objects corresponding to different entity records with co-reference problems, and incorporate the information of the redundant nodes into the target nodes; Wherein, the target node is the instance node with the largest amount of information or the highest information reliability in the instance nodes of different instance objects corresponding to different entity records with co-reference problems, and the redundant node is the instance node other than the target node in the instance nodes of different instance objects corresponding to different entity records with co-reference problems.

[0036] In one implementation of this embodiment, the calculating the time similarity, space similarity, and attribute similarity between the different target entity records respectively specifically includes: Calculate the time similarity between the different target entity records according to the different time information corresponding to the different target entity records. The time similarity aims to measure the overlap and consistency degree of the occurrence or existence time of two entities (referring to entity records). Based on Allen interval algebra, the qualitative relationship between two time intervals can be judged first, and then a value between 0 and 1 can be quantified. When the two time intervals are exactly the same or highly overlapped, , such that the value is close to 1; if the two do not intersect or only have a very small overlap, then the value is close to 0, thus effectively distinguishing the differences in time: ; ; ; Among them, represents the th target entity record, represents the th target entity record, represents the time similarity between the target entity record and the target entity record; represents the length of the overlapping interval between the target entity record and the target entity record; represents the length of the combined interval between the target entity record and the target entity record; represents the left endpoint of the time interval , of the target entity record, represents the right endpoint of the time interval , of the target entity record, represents the left endpoint of the time interval , of the target entity record, represents the right endpoint of the time interval , of the target entity record, represents taking the maximum value, represents taking the minimum value; According to the different spatial information corresponding to the different target entity records, calculate the spatial similarity between the different target entity records. The spatial similarity mainly measures the consistency of two entities (referring to entity records) in terms of geographical location, spatial range, and topological relationship: ; Distance similarity is calculated using a linear attenuation function of the following formula. When , the value is 0: ; Topological similarity is strictly corresponding according to the topological relationship between two entities (referring to entity records), and different scores are assigned to different topological relationships. In the present invention, the corresponding result is shown in the following formula: ; Region similarity It often targets surface entities and can be defined as the area overlap rate. When two areas highly overlap, the value approaches 1; if they do not intersect at all, then the value is 0: ; Among them, represents the spatial similarity between the target entity record and the target entity record ; represents the spatial representation of the target entity record ; represents the spatial representation of the target entity record ; represents the distance similarity between the spatial representation and the spatial representation (that is, between the target entity record and the target entity record . The relevant meanings below are the same as here and will not be elaborated further), represents the topological similarity between the spatial representation and the spatial representation ; represents the area similarity between the spatial representation and the spatial representation ; represents the weight of the distance similarity, represents the weight of the topological similarity, represents the weight of the area similarity, , represents the distance between the spatial representation and the spatial representation ; represents the set maximum acceptable distance threshold, represents the topological relationship between the spatial representation and the spatial representation ; represents the area of the intersection between the spatial representation and the spatial representation , that is, represents the area of the intersection of two surface entities, that is, the area of their overlapping region; represents the area of the union between the spatial representation and the spatial representation , that is, represents the area of the union of two surface entities, that is, the total area they cover; Record different attribute information under each attribute type corresponding to the different target entities, and calculate the attribute similarity between the records of the different target entities. The attribute similarity is used to measure the similarity degree between two entities (referring to entity records) in terms of name, category, data source, timestamp, etc. Each attribute can be regarded as a one-dimensional feature. Select corresponding measurement functions for different attribute types (string, numeric, enumeration), and then perform weighted summation: ; For string attributes, the edit distance is used to measure the similarity. Among them, the edit distance refers to the minimum number of single-character edits (insertion, deletion, or replacement) required to convert one string into another string. When two strings are exactly the same, the edit distance is 0, and the string attribute similarity is 1; if the gap is larger, the string attribute similarity is lower: ; For enumeration or category attributes, if the categories of two entities (referring to entity records) are the same, the category attribute similarity is assigned 1; if they are similar categories, the category attribute similarity is assigned a relatively high score; if they are different categories, the category attribute similarity is assigned 0. For example, the category attribute similarity between "school" and "college" can be regarded as 0.8, but the category attribute similarity between "school" and "hospital" is regarded as 0: ; For numeric attributes, the absolute value of the numerical difference between two entities (referring to entity records) can be normalized. For example, normalize the maximum possible numerical range [0, . If the absolute value of the difference does not exceed the tolerable range, the numeric attribute similarity is relatively high: ; Among them, represents the attribute similarity between target entity record and target entity record , represents the th attribute type, represents the weight of the th attribute type, , represents the sub-attribute similarity between target entity record and target entity record under the th attribute type, represents the sub-attribute similarity under the first attribute type (i.e., the string attribute similarity), represents target entity record 's string instance, represents target entity record The string instance represents the string instance and the string instance The edit distance between represents the string instance The length of represents the string instance The length of represents the sub - attribute similarity under the second type of attribute type (i.e., the categorical attribute similarity) represents a set score less than 1 and greater than 0.5, for example, 0.8 can be taken represents the target entity record The category of represents the target entity record The category of represents the category and the category The relationship between represents the sub - attribute similarity under the third type of attribute type (i.e., the numerical attribute similarity) represents the target entity record The value of represents the target entity record The value of represents the maximum possible numerical threshold, that is, the right endpoint of the maximum possible numerical range [0, .

[0037] In an implementation manner of this embodiment, the weighted sum of the time similarity, the space similarity, and the attribute similarity to obtain the spatio - temporal semantic similarity specifically includes: According to the weight coefficients of the time similarity, the space similarity, and the attribute similarity set in advance, the weighted sum of the time similarity between different target entity records, the space similarity between different target entity records, and the attribute similarity between different target entity records is performed to obtain the spatio - temporal semantic similarity between different target entity records: ; Among them, is the spatio - temporal semantic similarity between the target entity record and the target entity record , is the weight coefficient of the time similarity, is the weight coefficient of the space similarity, is the weight coefficient of the attribute similarity, , By setting different weight coefficients through experience, the above formula can comprehensively combine the influences of time, space, and attributes.

[0038] Specifically, based on spatio-temporal semantic similarity, the present invention fuses different entity records with co-reference problems corresponding to different instance objects under the same entity type to generate a unified urban multi-department knowledge graph, including: First, preprocess different entity records, that is, unify the time standard and spatial coordinate system of different data sources, and unify the spatio-temporal reference to the CGCS2000 (China Geodetic Coordinate System 2000) coordinate system and the UTC (Coordinated Universal Time) time standard; then calculate the time similarity, spatial similarity, and attribute similarity between different entity records respectively, and obtain the comprehensive similarity (i.e., spatio-temporal semantic similarity) by comprehensively weighting each similarity. Determine co-referring entities according to the comprehensive similarity, eliminate redundant nodes, merge redundant records, and eliminate co-reference problems.

[0039] Among them, in the actual dataset, it is also necessary to select an appropriate similarity threshold θ in combination with the actual record situation. Only when the spatio-temporal semantic similarity ≥θ, it is considered that these two entities (referring to entity records) represent the same object (referring to instance objects), thereby triggering subsequent fusion operations and reducing the misfusion of irrelevant objects. As Figure 4 shown, for node objects (referring to instance nodes, such as Figure 4 E1 and E2 in Figure 4 ) that exceed the similarity threshold, retain the instance node (such as Figure 4 E1 in Figure 4 ) with richer or more reliable information, and at the same time remove redundant nodes (such as Figure 4 E2 in Figure 4 ), and incorporate the information (such as aliases, sub-node relationships) of the removed nodes into the retained nodes. Here, the removed nodes are redundant nodes, the retained nodes are target nodes, and the sub-node relationships include the record nodes (such as Figure 4 R3 in Figure 4 ) connected by the redundant nodes and the edges between the redundant nodes and the record nodes. It should be noted that Figure 4 C in

[0040] represents the concept node corresponding to the entity type, E1 and E2 represent different instance nodes corresponding to different instance objects, and R1, R2, and R3 represent different record nodes corresponding to different entity records.

[0040] It should be noted that the urban multi-source heterogeneous data knowledge graph provided by the embodiments of the present invention is a city knowledge graph that can be applied to heterogeneous data retrieval. By querying the index information and general data in the knowledge graph, it can achieve querying and indexing of heterogeneous data based on entity objects.

[0041] In summary, the present invention realizes the efficient integration of heterogeneous data of multiple urban departments and the construction of a knowledge graph through the following technical solutions: Based on a hierarchical ontology model, a semantic framework covering seven categories of entities such as natural resources and building facilities is constructed, supporting multi-level mapping to accommodate data granularity differences; Using rule engines (such as SQL, Geopandas, etc.) to achieve automated extraction and graph structure transformation of multi-source heterogeneous data such as relational data, GIS data, BIM data, and IoT data; A spatio-temporal semantic weighting algorithm is proposed to fuse time similarity, space similarity, and attribute similarity, and the spatio-temporal benchmark (UTC time standard, CGCS2000 coordinate system) is unified to eliminate redundancy; Combined with the Neo4j graph database, nodes and edges are organized at three levels of concept, instance, and record, supporting complex cross-domain queries. The present invention can break down departmental data barriers and provide a globally consistent knowledge base for cross-departmental collaborative governance, emergency command, and dynamic data integration.

[0042] The present invention has the following beneficial effects: It realizes cross-departmental data integration, eliminates information silos, and realizes multi-source associated storage of core entity records; It realizes efficient knowledge fusion and reduces the generation of redundant records based on the spatio-temporal semantic similarity algorithm; It provides intelligent retrieval support, and the multi-level model based on the graph database supports data association and query for the same entity object; It is applied to scenarios such as smart city governance, emergency command, and cross-departmental collaborative analysis, providing intelligent decision-making support for urban comprehensive governance.

[0043] In addition, based on the above method for constructing a knowledge graph of heterogeneous data of multiple urban departments, the present invention also provides a system for constructing a knowledge graph of heterogeneous data of multiple urban departments. Among them, a preferred embodiment of the system for constructing a knowledge graph of heterogeneous data of multiple urban departments is as Figure 5 shown, and specifically includes: Ontology model construction module 01: Used to determine the entity objects, description attributes, and association relationships of the ontology model of heterogeneous data of multiple urban departments; Knowledge extraction module 02: Used to extract entity records, attribute information, and relationship information from the heterogeneous data of multiple urban departments using a rule engine according to the entity objects, the description attributes, and the association relationships; Initial knowledge graph construction module 03: Used to organize the entity objects, the entity records, the attribute information, and the relationship information in combination with the storage structure of the graph database to obtain an initial knowledge graph; Target knowledge graph construction module 04: Used to calculate the spatio-temporal semantic similarity according to the initial knowledge graph, and fuse different entity records with co-reference problems according to the spatio-temporal semantic similarity to obtain a target knowledge graph.

[0044] In addition, based on the above method and system for constructing a knowledge graph of heterogeneous data from multiple urban departments, the present invention also correspondingly provides a terminal. In a preferred embodiment of the terminal, as Figure 6 shown, it specifically includes a processor 10, a memory 20, and a display 30. Figure 6 Only some components of the terminal are shown. However, it should be understood that it is not required to implement all the shown components, and more or fewer components can be alternatively implemented. This terminal supports calling the knowledge graph construction and query services through an API interface.

[0045] The memory 20 can be an internal storage unit of the terminal in some embodiments, such as the hard disk or memory of the terminal. The memory 20 can also be an external storage device of the terminal in other embodiments, such as a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, and a Flash Card equipped on the terminal, etc. Further, the memory 20 can also include both the internal storage unit and the external storage device of the terminal. The memory 20 is used to store application software installed on the terminal and various types of data, such as storing the program code of the terminal, etc. The memory 20 can also be used to temporarily store data that has been output or will be output. In one embodiment, a program 40 for constructing a knowledge graph of heterogeneous data from multiple urban departments is stored on the memory 20, and the program 40 for constructing a knowledge graph of heterogeneous data from multiple urban departments can be executed by the processor 10, thereby implementing the steps of the method for constructing a knowledge graph of heterogeneous data from multiple urban departments in this application.

[0046] The processor 10 can be a Central Processing Unit (CPU), a microprocessor, or other data processing chips in some embodiments, and is used to run the program code stored in the memory 20 or process data, such as executing the program 40 for constructing a knowledge graph of heterogeneous data from multiple urban departments, etc.

[0047] The display 30 can be an LED display, a liquid crystal display, a touch liquid crystal display, and an OLED (Organic Light-Emitting Diode) toucher, etc. in some embodiments. The display 30 is used to display information on the terminal and to display a visual user interface.

[0048] In one embodiment, when the processor 10 executes the program 40 for constructing a knowledge graph of heterogeneous data from multiple urban departments in the memory 20, the steps of the method for constructing a knowledge graph of heterogeneous data from multiple urban departments as described above are implemented.

[0049] The present invention also correspondingly provides a computer-readable storage medium. The computer-readable storage medium stores a program for constructing a knowledge graph of heterogeneous data of multiple urban departments. When the program for constructing the knowledge graph of heterogeneous data of multiple urban departments is executed by a processor, the steps of the method for constructing the knowledge graph of heterogeneous data of multiple urban departments as described above are implemented.

[0050] It should be noted that in this text, the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or terminal including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, article or terminal. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of another identical element in the process, method, article or terminal including that element.

[0051] Certainly, those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware (such as a processor, a controller, etc.) through a computer program. The program can be stored in a computer-readable storage medium that can be read by a computer. When the program is executed, it can include the processes of the above method embodiments. The computer-readable storage medium can be a memory, a magnetic disk, an optical disc, etc.

[0052] It should be understood that the application of the present invention is not limited to the above examples. For those of ordinary skill in the art, improvements or transformations can be made according to the above description. All such improvements and transformations should fall within the protection scope of the appended claims of the present invention.

Claims

1. A method for constructing a knowledge graph of heterogeneous data from multiple urban departments, characterized in that, The method for constructing a knowledge graph of heterogeneous data from multiple urban departments includes: Determine the entity objects, descriptive attributes, and association relationships of the ontology model of heterogeneous data from multiple urban departments; According to the entity objects, the descriptive attributes, and the association relationships, use a rule engine to extract entity records, attribute information, and relationship information from the heterogeneous data of multiple urban departments; Combined with the storage structure of the graph database, organize the entity objects, the entity records, the attribute information, and the relationship information to obtain an initial knowledge graph; Calculate the spatio-temporal semantic similarity based on the initial knowledge graph, and fuse different entity records with co-reference problems according to the spatio-temporal semantic similarity to obtain a target knowledge graph.

2. The method for constructing a knowledge graph of heterogeneous data of multiple urban departments according to claim 1, wherein, The determination of the entity objects, descriptive attributes, and association relationships of the ontology model of heterogeneous data from multiple urban departments specifically includes: Obtain the heterogeneous data of multiple urban departments, determine the key entities of the ontology model of heterogeneous data from multiple urban departments in combination with the heterogeneous data of multiple urban departments, and classify and grade the key entities to obtain entity types and instance objects, where the entity types and the instance objects constitute the entity objects; Determine the descriptive attributes and association relationships of the ontology model of heterogeneous data from multiple urban departments in combination with the characteristics of the heterogeneous data of multiple urban departments.

3. The method for constructing a knowledge graph of heterogeneous data of multiple urban departments according to claim 1, characterized in that The heterogeneous data of multiple urban departments includes relational data, GIS data, BIM data, and IoT data; The extraction of entity records, attribute information, and relationship information from the heterogeneous data of multiple urban departments using a rule engine according to the entity objects, the descriptive attributes, and the association relationships specifically includes: According to the entity objects, the descriptive attributes, and the association relationships in the ontology model of heterogeneous data from multiple urban departments, extract the entity records corresponding to the entity objects, the attribute information corresponding to the descriptive attributes, and the relationship information corresponding to the association relationships from the relational data based on the SQL rule engine; According to the entity objects, the descriptive attributes, and the association relationships in the ontology model of heterogeneous data from multiple urban departments, extract the entity records corresponding to the entity objects, the attribute information corresponding to the descriptive attributes, and the relationship information corresponding to the association relationships from the GIS data based on the Geopandas tool; According to the entity objects, the descriptive attributes, and the association relationships in the ontology model of heterogeneous data from multiple urban departments, extract the entity records corresponding to the entity objects, the attribute information corresponding to the descriptive attributes, and the relationship information corresponding to the association relationships from the BIM data based on the Revit API; According to the entity objects, the descriptive attributes, and the association relationships in the ontology model of heterogeneous data from multiple urban departments, extract the entity records corresponding to the entity objects, the attribute information corresponding to the descriptive attributes, and the relationship information corresponding to the association relationships from the IoT data based on the InfluxDBClient engine.

4. The method for constructing an urban multi-department heterogeneous data knowledge graph according to claim 2, characterized in that, The organization of the entity objects, the entity records, the attribute information, and the relationship information in combination with the storage structure of the graph database to obtain an initial knowledge graph specifically includes: Combined with the node-attribute-edge storage structure of the graph database, the entity types, the instance objects, and the entity records are respectively mapped to concept nodes with concept labels, instance nodes with instance labels, and record nodes with record labels according to three levels: the concept layer, the instance layer, and the record layer; Use the key-value pair attributes embedded in the record nodes to represent the attribute information; Use the typed edges between the record nodes to represent the relationship information, obtaining an initial knowledge graph.

5. The method for constructing a knowledge graph of heterogeneous data of multiple urban departments according to claim 4, wherein The calculating the spatio-temporal semantic similarity according to the initial knowledge graph and fusing different entity records with co-reference problems according to the spatio-temporal semantic similarity to obtain a target knowledge graph specifically includes: Unify the spatio-temporal benchmarks of different entity records corresponding to different instance objects under the same entity type in the initial knowledge graph to obtain different target entity records corresponding to different instance objects under the same entity type; Calculate the time similarity, space similarity, and attribute similarity between the different target entity records respectively; Perform a weighted sum of the time similarity, the space similarity, and the attribute similarity to obtain the spatio-temporal semantic similarity; Compare the spatio-temporal semantic similarity with a similarity threshold. If the spatio-temporal semantic similarity is not less than the similarity threshold, it is considered that the different entity records are different entity records with co-reference problems corresponding to different instance objects under the same entity type; Retain the target nodes in the instance nodes of different instance objects corresponding to different entity records with co-reference problems, and at the same time remove the redundant nodes in the instance nodes of different instance objects corresponding to different entity records with co-reference problems, and incorporate the information of the redundant nodes into the target nodes; Among them, the target node is the instance node with the largest amount of information or the highest information reliability in the instance nodes of different instance objects corresponding to different entity records with co-reference problems, and the redundant node is the instance node other than the target node in the instance nodes of different instance objects corresponding to different entity records with co-reference problems.

6. The method for constructing a knowledge graph of heterogeneous data of multiple urban departments according to claim 5, wherein The calculating the time similarity, space similarity, and attribute similarity between the different target entity records respectively specifically includes: Calculate the time similarity between the different target entity records according to the different time information corresponding to the different target entity records; ; ; ; Among them, represents the th target entity record, represents the th target entity record, represents the time similarity between the target entity record and the target entity record represents the length of the overlapping interval between the target entity record and the target entity record represents the length of the combined interval between the target entity record and the target entity record; represents the left endpoint of the time interval , of the target entity record represents the right endpoint of the time interval , of the target entity record represents the left endpoint of the time interval , of the target entity record represents the right endpoint of the time interval , of the target entity record represents taking the maximum value, represents taking the minimum value; Calculate the space similarity between the different target entity records according to the different space information corresponding to the different target entity records; ; ; ; ; Among them, represents the spatial similarity between the target entity record and the target entity record, represents the spatial representation of the target entity record, represents the spatial representation of the target entity record, represents the distance similarity between the spatial representation and the spatial representation, represents the topological similarity between the spatial representation and the spatial representation, represents the regional similarity between the spatial representation and the spatial representation, represents the weight of the distance similarity, represents the weight of the topological similarity, represents the weight of the regional similarity, represents the distance between the spatial representation and the spatial representation, represents the maximum acceptable distance threshold, represents the topological relationship between the spatial representation and the spatial representation, represents the area of the intersection part between the spatial representation and the spatial representation, represents the area of the union part between the spatial representation and the spatial representation; Calculate the attribute similarity between the different target entity records according to the different attribute information under each attribute type corresponding to the different target entity records; ; ; ; ; Among them, represents the attribute similarity between and the target entity record ; represents the th type of attribute; represents the weight of the th type of attribute; represents the sub-attribute similarity between the target entity record and the target entity record under the th type of attribute; represents the sub-attribute similarity under the 1st type of attribute; represents the string instance of the target entity record ; represents the string instance of the target entity record ; represents the edit distance between the string instance and the string instance ; represents the length of the string instance ; represents the length of the string instance ; represents the sub-attribute similarity under the 2nd type of attribute; represents a set score less than 1 and greater than 0.5; represents the category of the target entity record ; represents the category of the target entity record ; represents the relationship between the category and the category ; represents the sub-attribute similarity under the 3rd type of attribute; represents the numerical value of the target entity record ; represents the numerical value of the target entity record ; represents the maximum possible numerical value threshold.

7. The method for constructing a knowledge graph of heterogeneous data of multiple urban departments according to claim 6, wherein, The performing a weighted sum of the time similarity, the space similarity, and the attribute similarity to obtain the spatio-temporal semantic similarity specifically includes: According to the preset weight coefficients of the time similarity, the space similarity, and the attribute similarity, perform a weighted sum of the time similarity between the different target entity records, the space similarity between the different target entity records, and the attribute similarity between the different target entity records to obtain the spatio-temporal semantic similarity between the different target entity records; ; Among them, is the spatio-temporal semantic similarity between the target entity record and the target entity record The spatio-temporal semantic similarity is the weight coefficient of the time similarity is the weight coefficient of the space similarity is the weight coefficient of the attribute similarity 8. A knowledge graph construction system for heterogeneous data of multiple departments in a city, characterized in that, The urban multi-department heterogeneous data knowledge graph construction system includes: Ontology model construction module: used to determine the entity objects, description attributes and association relationships of the urban multi-department heterogeneous data ontology model; Knowledge extraction module: used to extract entity records, attribute information and relationship information from urban multi-department heterogeneous data using a rule engine according to the entity objects, the description attributes and the association relationships; Initial knowledge graph construction module: used to organize the entity objects, the entity records, the attribute information and the relationship information in combination with the storage structure of the graph database to obtain an initial knowledge graph; Target knowledge graph construction module: used to calculate the spatio-temporal semantic similarity according to the initial knowledge graph, and fuse different entity records with co-reference problems according to the spatio-temporal semantic similarity to obtain a target knowledge graph.

9. A terminal, characterized in that, The terminal includes: a memory, a processor, and an urban multi-department heterogeneous data knowledge graph construction program stored on the memory and executable on the processor. When the urban multi-department heterogeneous data knowledge graph construction program is executed by the processor, the steps of the urban multi-department heterogeneous data knowledge graph construction method according to any one of claims 1-7 are implemented.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores an urban multi-department heterogeneous data knowledge graph construction program. When the urban multi-department heterogeneous data knowledge graph construction program is executed by a processor, the steps of the urban multi-department heterogeneous data knowledge graph construction method according to any one of claims 1-7 are implemented.

Citation Information

Patent Citations

  • Space-time multi-mode mixed data processing method, association method and indexing method

    CN113297395A

  • Heterogeneous data fusion indexing system and device based on knowledge graph

    CN116662342A

  • Cross-source data retrieval method based on space-time asset directory, medium and equipment

    CN117312688A

  • Traditional Chinese medicine intelligent query system based on knowledge graph

    CN118395021A

Cited By

  • Heterogeneous data monitoring method and system

    CN120578787A

  • Coding and graph falling management method and system based on building monomers

    CN121031900A

  • A building unit-based code assignment plot management method and system

    CN121031900B

  • BIM-based component-level knowledge graph construction method and system

    CN121660041A

  • Data integration and visualization method of city CIM information platform

    CN121766407A