A method, system, terminal and storage medium for constructing a knowledge graph of heterogeneous data from multiple urban departments

By constructing a knowledge graph of heterogeneous data in cities with multi-department, determining the ontology model, extracting entity records and calculating the semantic similarity of space-time semantic similarity, the semantic islands and co-referentials of urban multi-department data are solved, and the unified storage of data and cross-departmental collaboration is realized, and intelligent decision-making support from a global perspective is provided.

CN120338076BActive Publication Date: 2025-08-22SHENZHEN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510804679.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-17
Publication Date
2025-08-22
Estimated Expiration
2045-06-17

AI Technical Summary

Technical Problem

There are semantic islands and entities that combine the problem of eliminating difficulties in existing urban multi-department heterogeneous data. The existing methods fail to effectively correlate and integrate data from different departments and fields from the perspective of urban comprehensive decision makers.

Method used

By determining the entity objects, description attributes and association relationships of the city's multi-department heterogeneous data ontology model, the rule engine is used to extract entity records and relationship information, the knowledge graph is organized in combination with the storage structure of the graph database, and the spatial and temporal semantic similarity fusion co-referential problems are calculated, so as to realize the semantic association and entity co-referential digestion of cross-departmental data.

Benefits of technology

It realizes unified storage and cross-departmental collaboration of heterogeneous data in urban departments, provides intelligent decision-making support from a global perspective, breaks down data barriers, and improves the efficiency of data sharing and cross-departmental collaborative work.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120338076B_ABST
    Figure CN120338076B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of information processing technology, and discloses a method, system, terminal and storage medium for constructing a knowledge graph of heterogeneous data of multiple departments in a city. The method comprises: determining the entity objects, descriptive attributes and association relationships of the ontology model of heterogeneous data of multiple departments in the city; extracting entity records, attribute information and relationship information from the heterogeneous data of multiple departments in the city using a rule engine according to the entity objects, descriptive attributes and association relationships; organizing the entity objects, entity records, attribute information and relationship information in combination with the storage structure of the graph database to obtain an initial knowledge graph; calculating the spatiotemporal semantic similarity according to the initial knowledge graph, and fusing different entity records with co-reference problems according to the spatiotemporal semantic similarity to obtain a target knowledge graph. The present invention realizes the semantic association, knowledge integration and entity co-reference resolution of cross-departmental data by constructing a knowledge graph, and provides intelligent decision support from a global perspective for urban governance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of information processing technology, and in particular to a method, system, terminal and computer-readable storage medium for constructing a knowledge graph of heterogeneous data from multiple departments in a city. Background Art

[0002] With the accelerated development of new smart cities, the amount of data generated by urban operations is growing exponentially. Data types encompass geospatial information, building information models, IoT sensor data, and structured and semi-structured data from various departmental business systems. This data is dispersed across various departments, including planning, housing and construction, transportation, environment, and public safety, creating a complex, multi-source, heterogeneous data ecosystem. Because each department utilizes independent data standards and storage architectures, and because description dimensions vary across systems, semantic silos exist between data from multiple urban departments.

[0003] In current research, knowledge graphs have been widely used to solve the problem of data semantic silos. By converting heterogeneous data into a graph structure, relationships between different entities can be established within the graph structure, enabling cross-domain information fusion and semantic association. However, existing urban knowledge graph research often focuses on vertical application areas and fails to associate and integrate data from different departments and fields from the perspective of comprehensive urban decision-makers. At the same time, in order to address the difficulties in resolving coreference due to the different spatiotemporal benchmarks of data from various departments and inconsistent entity concepts, existing methods have not yet organized and integrated data from the perspective of entity-based thinking and spatiotemporal semantic similarity, making it impossible to achieve a global view and cross-departmental collaboration.

[0004] Therefore, the existing technology still needs to be improved and developed. Summary of the Invention

[0005] The main purpose of the present invention is to provide a method, system, terminal and computer-readable storage medium for constructing a knowledge graph of heterogeneous data from multiple departments of a city, aiming to solve the problems of semantic islands and difficulty in resolving entity co-references in existing heterogeneous data from multiple departments of a city.

[0006] To achieve the above-mentioned object of the invention, the present invention provides a method for constructing a knowledge graph of heterogeneous data from multiple departments in a city, the method comprising:

[0007] Determine the entity objects, descriptive attributes and association relationships of the heterogeneous data ontology model of multiple departments in the city;

[0008] According to the entity objects, the descriptive attributes and the association relationships, a rule engine is used to extract entity records, attribute information and relationship information from heterogeneous data of multiple departments in the city;

[0009] In combination with the storage structure of the graph database, the entity objects, the entity records, the attribute information and the relationship information are organized to obtain an initial knowledge graph;

[0010] The spatiotemporal semantic similarity is calculated based on the initial knowledge graph, and different entity records with co-reference problems are fused based on the spatiotemporal semantic similarity to obtain a target knowledge graph.

[0011] Optionally, determining the entity objects, descriptive attributes, and association relationships of the city multi-department heterogeneous data ontology model specifically includes:

[0012] Acquire heterogeneous data of multiple urban departments, determine key entities of an ontology model of the heterogeneous data of multiple urban departments in combination with the heterogeneous data of multiple urban departments, and classify and categorize the key entities to obtain entity types and instance objects, wherein the entity types and the instance objects constitute the entity objects;

[0013] In combination with the characteristics of the heterogeneous data of multiple departments in the city, the descriptive attributes and association relationships of the ontology model of the heterogeneous data of multiple departments in the city are determined.

[0014] Optionally, the city multi-department heterogeneous data includes relational data, GIS data, BIM data and IoT data;

[0015] The method of extracting entity records, attribute information, and relationship information from heterogeneous data of multiple departments of a city using a rule engine based on the entity objects, the descriptive attributes, and the association relationships specifically includes:

[0016] According to the entity objects, the descriptive attributes and the association relationships in the city multi-department heterogeneous data ontology model, extracting entity records corresponding to the entity objects, attribute information corresponding to the descriptive attributes and relationship information corresponding to the association relationships from the relational data based on an SQL rule engine;

[0017] Extracting entity records corresponding to the entity objects, attribute information corresponding to the descriptive attributes, and relationship information corresponding to the relationship from the GIS data based on the Geopandas tool according to the entity objects, the descriptive attributes, and the relationship in the urban multi-department heterogeneous data ontology model;

[0018] Extracting entity records corresponding to the entity objects, attribute information corresponding to the descriptive attributes, and relationship information corresponding to the relationship from the BIM data based on the Revit API according to the entity objects, the descriptive attributes, and the relationship in the city multi-department heterogeneous data ontology model;

[0019] According to the entity objects, the descriptive attributes and the association relationships in the city's multi-department heterogeneous data ontology model, the entity records corresponding to the entity objects, the attribute information corresponding to the descriptive attributes and the relationship information corresponding to the association relationships are extracted from the IoT data based on the InfluxDBClient engine.

[0020] Optionally, organizing the entity objects, the entity records, the attribute information, and the relationship information in combination with the storage structure of the graph database to obtain an initial knowledge graph specifically includes:

[0021] Combined with the node-attribute-edge storage structure of the graph database, the entity type, the instance object, and the entity record are mapped to concept nodes with concept labels, instance nodes with instance labels, and record nodes with record labels respectively according to the three levels of concept layer, instance layer, and record layer;

[0022] Representing the attribute information using the key-value pair attributes embedded in the record node;

[0023] The relationship information is represented using typed edges between the record nodes to obtain an initial knowledge graph.

[0024] Optionally, calculating spatiotemporal semantic similarity based on the initial knowledge graph, and fusing different entity records with coreference problems based on the spatiotemporal semantic similarity to obtain a target knowledge graph specifically includes:

[0025] Performing spatiotemporal benchmark unification on different entity records corresponding to different instance objects under the same entity type in the initial knowledge graph to obtain different target entity records corresponding to different instance objects under the same entity type;

[0026] Calculating the temporal similarity, spatial similarity and attribute similarity between the different target entity records respectively;

[0027] Performing weighted summation on the temporal similarity, the spatial similarity, and the attribute similarity to obtain spatiotemporal semantic similarity;

[0028] Comparing the spatiotemporal semantic similarity with a similarity threshold, if the spatiotemporal semantic similarity is not less than the similarity threshold, it is considered that the different entity records are different entity records corresponding to different instance objects under the same entity type and have a coreference problem;

[0029] Retain the target node in the instance nodes of different instance objects corresponding to different entity records with coreference problems, remove redundant nodes in the instance nodes of different instance objects corresponding to different entity records with coreference problems, and merge information of the redundant nodes into the target node;

[0030] Among them, the target node is the instance node with the largest amount of information or the highest information reliability among the instance nodes of different instance objects corresponding to different entity records with coreference problems, and the redundant node is the instance node other than the target node among the instance nodes of different instance objects corresponding to different entity records with coreference problems.

[0031] Optionally, respectively calculating the temporal similarity, spatial similarity, and attribute similarity between the different target entity records specifically includes:

[0032] Calculate the time similarity between the different target entity records based on the different time information corresponding to the different target entity records:

[0033] ;

[0034] ;

[0035] ;

[0036] in, Indicates the target entity records, Indicates the target entity records, Represents the target entity record and the target entity record The time similarity between Represents the target entity record and the target entity record The length of the overlapping interval between Represents the target entity record and the target entity record The length of the joint interval between Represents the target entity record Time interval[ , ]’s left endpoint, Represents the target entity record Time interval[ , ]’s right endpoint, Represents the target entity record Time interval[ , ]’s left endpoint, Represents the target entity record Time interval[ , ]’s right endpoint, Indicates taking the maximum value, Indicates taking the minimum value;

[0037] Calculate the spatial similarity between the different target entity records based on the different spatial information corresponding to the different target entity records:

[0038] ;

[0039] ;

[0040] ;

[0041] ;

[0042] in, Represents the target entity record and the target entity record The spatial similarity between Represents the target entity record The spatial representation of Represents the target entity record The spatial representation of Representation space representation and spatial representation The distance similarity between Representation space representation and spatial representation The topological similarity between Representation space representation and spatial representation The regional similarity between represents the weight of distance similarity, represents the weight of topological similarity, represents the weight of region similarity, Representation space representation and spatial representation The distance between represents the maximum acceptable distance threshold, Representation space representation and spatial representation The topological relationship between Representation space representation and spatial representation The area of ​​the intersection between them, Representation space representation and spatial representation The area of ​​the union between them;

[0043] Calculate the attribute similarity between the different target entity records based on the different attribute information under each attribute type corresponding to the different target entity records:

[0044] ;

[0045] ;

[0046] ;

[0047] ;

[0048] in, Represents the target entity record and the target entity record The attribute similarity between Indicates the Types of attributes, Indicates the The weight of the attribute type, Represents the target entity record and the target entity record Between The similarity of sub-attributes under the attribute type, Indicates the sub-attribute similarity under the first attribute type, Represents the target entity record String instance, Represents the target entity record String instance, Represents a string instance With string instance The edit distance between Represents a string instance length, Represents a string instance length, Indicates the sub-attribute similarity under the second attribute type, Indicates a set score that is less than 1 and greater than 0.5, Represents the target entity record Category, Represents the target entity record Category, Representation category and categories The relationship between Indicates the similarity of sub-attributes under the third attribute type, Represents the target entity record The numerical value of Represents the target entity record The numerical value of Indicates the maximum possible numerical threshold.

[0049] Optionally, performing weighted summation on the temporal similarity, the spatial similarity, and the attribute similarity to obtain spatiotemporal semantic similarity specifically includes:

[0050] According to the preset weight coefficients of temporal similarity, spatial similarity and attribute similarity, the temporal similarity between the different target entity records, the spatial similarity between the different target entity records and the attribute similarity between the different target entity records are weighted and summed to obtain the spatiotemporal semantic similarity between the different target entity records:

[0051] ;

[0052] in, Record for the target entity and the target entity record The spatiotemporal semantic similarity between is the weight coefficient of time similarity, is the weight coefficient of spatial similarity, is the weight coefficient of attribute similarity.

[0053] To achieve the above-mentioned object of the invention, the present invention further provides a system for constructing a knowledge graph of heterogeneous data from multiple departments in a city, the system comprising:

[0054] Ontology model construction module: used to determine the entity objects, description attributes and association relationships of the heterogeneous data ontology model of multiple departments in the city;

[0055] Knowledge extraction module: used to extract entity records, attribute information and relationship information from heterogeneous data of multiple departments of the city using a rule engine based on the entity objects, the descriptive attributes and the association relationships;

[0056] An initial knowledge graph construction module is used to organize the entity objects, entity records, attribute information, and relationship information in combination with the storage structure of the graph database to obtain an initial knowledge graph;

[0057] Target knowledge graph construction module: used to calculate the spatiotemporal semantic similarity based on the initial knowledge graph, and to fuse different entity records with co-reference problems based on the spatiotemporal semantic similarity to obtain the target knowledge graph.

[0058] In order to achieve the above-mentioned purpose of the invention, the present invention also provides a terminal, which includes: a memory, a processor, and a city multi-department heterogeneous data knowledge graph construction program stored in the memory and runnable on the processor. When the city multi-department heterogeneous data knowledge graph construction program is executed by the processor, the steps of the city multi-department heterogeneous data knowledge graph construction method as described above are implemented.

[0059] In order to achieve the above-mentioned purpose of the invention, the present invention also provides a computer-readable storage medium, which stores a city multi-department heterogeneous data knowledge graph construction program, and when the city multi-department heterogeneous data knowledge graph construction program is executed by a processor, it implements the steps of the city multi-department heterogeneous data knowledge graph construction method as described above.

[0060] In the present invention, the entity objects, descriptive attributes and association relationships of the heterogeneous data ontology model of multiple departments in the city are determined; based on the entity objects, the descriptive attributes and the association relationships, a rule engine is used to extract entity records, attribute information and relationship information from the heterogeneous data of multiple departments in the city; combined with the storage structure of the graph database, the entity objects, the entity records, the attribute information and the relationship information are organized to obtain an initial knowledge graph; the spatiotemporal semantic similarity is calculated based on the initial knowledge graph, and different entity records with co-reference problems are merged based on the spatiotemporal semantic similarity to obtain a target knowledge graph. This invention uses a unified ontology model, a multi-source extraction mechanism, and spatiotemporal semantic fusion technology. Based on the idea of ​​entities and the perspective of spatiotemporal semantic similarity, it organizes and integrates data from different departments and fields from the perspective of urban comprehensive decision makers, realizes semantic association, knowledge integration, and entity co-reference resolution of cross-departmental data, solves the problems of semantic islands and entity co-reference resolution in existing heterogeneous data of multiple departments in cities, and provides intelligent decision-making support from a global perspective for urban governance. At the same time, by constructing a knowledge graph, it realizes the unified storage of entities, attributes, and relationships in heterogeneous data of multiple departments in cities, which helps to associate heterogeneous data of multiple departments in cities and break down the barriers between heterogeneous data of multiple departments in cities. BRIEF DESCRIPTION OF THE DRAWINGS

[0061] Figure 1 It is a flowchart of a preferred embodiment of the method for constructing a knowledge graph of heterogeneous data of multiple departments in a city according to the present invention;

[0062] Figure 2 It is a schematic diagram of the city multi-department heterogeneous data ontology model of the present invention;

[0063] Figure 3 is a schematic diagram of the initial knowledge graph of the present invention;

[0064] Figure 4 It is a schematic diagram of the present invention for fusing different entity records that have a coreference problem;

[0065] Figure 5 This is a structural diagram of a preferred embodiment of the city multi-department heterogeneous data knowledge graph construction system of the present invention;

[0066] Figure 6 It is a structural diagram of a preferred embodiment of the terminal of the present invention. DETAILED DESCRIPTION

[0067] In order to make the purpose, technical solutions and advantages of the present invention more clear and distinct, the present invention is further described in detail below with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0068] With the accelerated development of new smart cities, the amount of data generated by urban operations is growing exponentially. Data types encompass geospatial information, building information models, IoT sensor data, and structured and semi-structured data from various departmental business systems. This data is dispersed across various departments, such as planning, housing and construction, transportation, environment, and public safety, creating a complex, multi-source, heterogeneous data ecosystem. Due to the independent data standards and storage architectures used by each department, and the varying descriptive dimensions across different systems, semantic silos exist between data from multiple urban departments. In practice, the same entity (such as a building, transportation facility, or population) appears repeatedly in data tables across different departments, but these records lack effective linkage, further hindering the efficiency of data sharing and cross-departmental collaboration.

[0069] Knowledge graphs are currently being widely used to address the problem of data semantic silos. By transforming heterogeneous data into a graph structure, relationships between different entities can be established within the graph, enabling cross-domain information fusion and semantic association. For example, urban knowledge graphs can organize data from areas such as transportation, construction, and public safety within a unified knowledge framework, thereby supporting urban management and decision-making.

[0070] However, existing urban knowledge graph research often focuses on vertical application areas, failing to connect and integrate data from different departments and fields from the perspective of comprehensive urban decision-makers. Furthermore, to address the difficulties in resolving coreferences caused by inconsistent spatiotemporal benchmarks and inconsistent entity concepts across departments, existing methods have yet to organize and integrate data from an entity-based perspective and spatiotemporal semantic similarity, hindering the realization of a global view and cross-departmental collaboration.

[0071] In order to solve the above technical problems, the present invention provides a method for constructing a knowledge graph of heterogeneous data of multiple departments in a city, determining the entity objects, descriptive attributes and association relationships of the ontology model of heterogeneous data of multiple departments in a city; using a rule engine to extract entity records, attribute information and relationship information from the heterogeneous data of multiple departments in the city according to the entity objects, the descriptive attributes and the association relationships; combining the storage structure of the graph database, organizing the entity objects, the entity records, the attribute information and the relationship information to obtain an initial knowledge graph; calculating the spatiotemporal semantic similarity based on the initial knowledge graph, and merging different entity records with co-reference problems based on the spatiotemporal semantic similarity to obtain a target knowledge graph. This invention uses a unified ontology model, a multi-source extraction mechanism, and spatiotemporal semantic fusion technology. Based on the idea of ​​entities and the perspective of spatiotemporal semantic similarity, it organizes and integrates data from different departments and fields from the perspective of urban comprehensive decision makers, realizes semantic association, knowledge integration, and entity co-reference resolution of cross-departmental data, solves the problems of semantic islands and entity co-reference resolution in existing heterogeneous data of multiple departments in cities, and provides intelligent decision-making support from a global perspective for urban governance. At the same time, by constructing a knowledge graph, it realizes the unified storage of entities, attributes, and relationships in heterogeneous data of multiple departments in cities, which helps to associate heterogeneous data of multiple departments in cities and break down the barriers between heterogeneous data of multiple departments in cities.

[0072] The application content will be further explained below through description of embodiments in conjunction with the accompanying drawings.

[0073] A preferred embodiment of the method for constructing a knowledge graph of heterogeneous data from multiple departments in a city according to the present invention is as follows: Figure 1 As shown, specifically including:

[0074] S1. Determine the entity objects, descriptive attributes and association relationships of the city's multi-department heterogeneous data ontology model.

[0075] In one implementation of this embodiment, determining the entity objects, descriptive attributes, and association relationships of the city multi-department heterogeneous data ontology model specifically includes:

[0076] Acquire the city multi-department heterogeneous data, determine key entities of the city multi-department heterogeneous data ontology model based on the city multi-department heterogeneous data, and classify and categorize the key entities to obtain entity types and instance objects, wherein the entity types and the instance objects constitute the entity objects;

[0077] In combination with the characteristics of the heterogeneous data of multiple departments in the city, the descriptive attributes and association relationships of the ontology model of the heterogeneous data of multiple departments in the city are determined;

[0078] Among them, the entity objects include seven major categories, namely natural resources and environment, special topics and services, transportation infrastructure, urban management and services, population and society, buildings and facilities, and land and planning, as well as seven subcategories corresponding to the seven major categories; the descriptive attributes include ID, standard name, index information, department source, data source, absolute position description, relative address description, timestamp, status, geometry, material and other information; the association relationships include located in, contained, adjacent, before and after, superior and subordinate, subordinate, dependent and whole-part.

[0079] Specifically, referring to the relevant standards such as "Classification and Code of Basic Geographic Information Elements" and "OGC Geography Markup Language Encoding Standard" (OGC stands for Open Geospatial Consortium), combined with the multi-source and heterogeneous characteristics of urban data, a hierarchical semantic framework (i.e., ontology model) is established, such as Figure 2 As shown, first, construct entity types (including instance objects). Entities are identifiable objects in the city that have management needs and have specific attributes and functions, including seven major categories and corresponding middle categories and subcategories: natural resources and environment, topics and services, transportation infrastructure, urban management and services, population and society, buildings and facilities, and land and planning; construct attribute types. Attributes describe the necessary and common attributes of urban multi-department entity objects for retrieval and application, including multiple categories and corresponding attributes such as identification, indexing, time and space, and features; construct relationship types. Relationships are used to express the connections between different urban multi-department entity objects for retrieval and application, including multiple categories and corresponding relationships such as space, time, hierarchy, and function. In order to realize the design of heterogeneous data ontology model for multiple departments in a city, the ontology model constructed by the present invention needs to include the following five principles: Principle 1, knowledge related to urban data index retrieval should be able to be converted into the form of triples, namely entities, attributes and relationships; Principle 2, the extracted triple knowledge needs to be able to provide users with useful information for data retrieval; Principle 3, the construction of the ontology model should ensure the accuracy and consistency of semantic expression to avoid affecting data fusion due to different expressions; Principle 4, when there is ambiguity in the terms used for the same concept, it is necessary to select a definition in a more authoritative standard that meets actual requirements; Principle 5, in view of the fact that data sources may have different granularity expressions, support the use of different records for multi-level representation, and allow the same entity object to have multiple records coexisting. Through the formulation of the above principles, the ontology model constructed by the present invention can be compatible with different data standards while ensuring semantic accuracy, support multi-granularity data expression, and optimize data organization. It should be noted that, Figure 2 The attributes in it refer to descriptive attributes, and the relationships refer to association relationships.

[0080] Specifically, the present invention combines standards related to heterogeneous urban data from multiple departments to identify key entities within the ontology model of this data. It then classifies the vast number of entities (i.e., key entities) within a city into hierarchical categories, resulting in entity types and instance objects (collectively, entity objects). Urban entity objects are identifiable objects within a city that meet management requirements and possess specific attributes and functions. Then, referring to the designed entity objects and combining the characteristics of heterogeneous urban data from multiple departments, descriptive attributes and associations are designed. Descriptive attributes are used to represent index information and common attributes for data records corresponding to different urban entity objects, while associations are used to represent the dependencies between data records corresponding to different urban entity objects.

[0081] S2. Based on the entity objects, the descriptive attributes and the association relationships, a rule engine is used to extract entity records, attribute information and relationship information from heterogeneous data of multiple departments in the city.

[0082] In one implementation of this embodiment, the heterogeneous data of multiple departments in the city includes relational data, GIS (Geographic Information System) data, BIM (Building Information Modeling) data, and IoT (Internet of Things) data.

[0083] The method of extracting entity records, attribute information, and relationship information from heterogeneous data of multiple departments of a city using a rule engine based on the entity objects, the descriptive attributes, and the association relationships specifically includes:

[0084] Extracting, from the relational data, entity records corresponding to the entity objects, attribute information corresponding to the descriptive attributes, and relationship information corresponding to the relationship based on a SQL (Structured Query Language) rule engine according to the entity objects, the descriptive attributes, and the relationship in the city multi-department heterogeneous data ontology model;

[0085] Extracting entity records corresponding to the entity objects, attribute information corresponding to the descriptive attributes, and relationship information corresponding to the relationship from the GIS data based on the Geopandas tool according to the entity objects, the descriptive attributes, and the relationship in the urban multi-department heterogeneous data ontology model;

[0086] Extracting, based on the entity objects, the descriptive attributes, and the association relationships in the city multi-department heterogeneous data ontology model, entity records corresponding to the entity objects, attribute information corresponding to the descriptive attributes, and relationship information corresponding to the association relationships from the BIM data in Revit format based on the Revit API (Revise instantly Application Programming Interface);

[0087] According to the entity objects, the descriptive attributes and the association relationships in the city's multi-department heterogeneous data ontology model, the entity records corresponding to the entity objects, the attribute information corresponding to the descriptive attributes and the relationship information corresponding to the association relationships are extracted from the IoT data stored in InfluxDB based on the InfluxDBClient engine.

[0088] Specifically, the urban multi-source heterogeneous data targeted by the present invention includes relational data, GIS data, BIM data and IoT data. A rule engine is used to extract knowledge from a relational database. For example, based on the SQL rule engine, the primary key of the relational data is extracted as an entity record, and the foreign key association is extracted as a relationship; a rule engine is used to extract knowledge from a GIS database. For example, based on the Geopandas tool, GIS data files are parsed to extract spatial entities and their hierarchical relationships from urban GIS scenes; a rule engine is used to extract knowledge from a BIM database. For example, through the Revit API, Revit format files (i.e., BIM data in Revit format) are parsed to extract room entities and their material properties, and the hierarchical relationships of components are mapped as "inclusion" edges; a rule engine is used to extract knowledge from an IoT database. For example, based on the InfluxDBClient engine, sensor node time series data is extracted from the InfluxDB time series database (a type of IoT database), entity records with timestamps are generated, and device topology relationships are converted into "connection" edges.

[0089] S3. In combination with the storage structure of the graph database, the entity objects, the entity records, the attribute information and the relationship information are organized to obtain an initial knowledge graph.

[0090] In one implementation of this embodiment, the entity objects, entity records, attribute information, and relationship information are organized in combination with the storage structure of the graph database to obtain an initial knowledge graph, specifically including:

[0091] Combined with the node-attribute-edge storage structure of the graph database, the entity type, the instance object, and the entity record are mapped to concept nodes with concept labels, instance nodes with instance labels, and record nodes with record labels respectively according to the three levels of concept layer, instance layer, and record layer;

[0092] Representing the attribute information using the key-value pair attributes embedded in the record node;

[0093] The relationship information is represented using typed edges between the record nodes to obtain an initial knowledge graph.

[0094] Specifically, the present invention combines the storage structure of the graph database to realize the organization and association of design concepts (i.e., entity concepts) and extracted knowledge, including: taking the entity concepts (i.e., entity objects, including entity types and instance objects) determined in step S1 and the entity records extracted in step S3 according to the entity concepts determined in step S1 as entity nodes of the graph model / graph structure / graph database, distinguishing the levels by labels, that is, the entities are organized according to three levels: concepts, instances, and records, wherein the concept layer defines the schema of entity types and their attribute relationships; the instance layer represents specific entity objects (i.e., instance objects) and attributes; the record layer stores original record data (i.e., entity records) from different sources and links to instance layer nodes through relationship edges. The attribute information extracted in step S3 according to the descriptive attributes determined in step S1 is stored in the node as a key-value pair attribute embedded in the graph model; the relationship information extracted in step S3 according to the association relationship determined in step S1 is organized as a typed edge relationship of the graph model. The graph model / graph structure obtained after organization is the initial knowledge graph, such as Figure 3 It should be noted that the entity concepts, descriptive attributes, and association relationships determined in step S1 are the entity objects, descriptive attributes, and association relationships in the ontology model / semantic framework. The entity objects, descriptive attributes, and association relationships are equivalent to the design concepts. Step S3 is to extract the corresponding knowledge according to the design concepts. Figure 3 This is a schematic diagram of the initial knowledge graph (i.e., a diagram of the organization of a knowledge graph for heterogeneous data from multiple departments in a city). The concept layer expresses the hierarchy and classification of different city entity types and, based on this, connects them to specific instance objects (such as resident population 1) at the instance layer. Each instance object is in turn associated with each corresponding record extracted from heterogeneous data at the record layer, and the records also have different relationships. It is important to note that the mapping from concepts to instances and from instances to records is not a one-to-one one, but a one-to-many one. Entity types, instance objects, and entity records are collectively referred to as entities.

[0095] S4. Calculate the spatiotemporal semantic similarity based on the initial knowledge graph, and fuse different entity records with co-reference problems based on the spatiotemporal semantic similarity to obtain a target knowledge graph.

[0096] In one implementation of this embodiment, calculating the spatiotemporal semantic similarity based on the initial knowledge graph and fusing different entity records with coreference problems based on the spatiotemporal semantic similarity to obtain a target knowledge graph specifically includes:

[0097] Performing spatiotemporal benchmark unification on different entity records corresponding to different instance objects under the same entity type in the initial knowledge graph to obtain different target entity records corresponding to different instance objects under the same entity type;

[0098] Calculating the temporal similarity, spatial similarity and attribute similarity between the different target entity records respectively;

[0099] Performing weighted summation on the temporal similarity, the spatial similarity, and the attribute similarity to obtain spatiotemporal semantic similarity;

[0100] Comparing the spatiotemporal semantic similarity with a similarity threshold, if the spatiotemporal semantic similarity is not less than the similarity threshold, it is considered that the different entity records are different entity records corresponding to different instance objects under the same entity type and have a coreference problem;

[0101] Retain the target node in the instance nodes of different instance objects corresponding to different entity records with coreference problems, remove redundant nodes in the instance nodes of different instance objects corresponding to different entity records with coreference problems, and merge information of the redundant nodes into the target node;

[0102] Among them, the target node is the instance node with the largest amount of information or the highest information reliability among the instance nodes of different instance objects corresponding to different entity records with coreference problems, and the redundant node is the instance node other than the target node among the instance nodes of different instance objects corresponding to different entity records with coreference problems.

[0103] In an implementation of this embodiment, respectively calculating the temporal similarity, spatial similarity, and attribute similarity between the different target entity records specifically includes:

[0104] Based on the different time information corresponding to the different target entity records, the time similarity between the different target entity records is calculated. The time similarity is intended to measure the overlap and consistency of the occurrence or existence time of two entities (referring to entity records). Based on Allen interval algebra, the qualitative relationship between the two time intervals can be first determined, and then quantified to obtain a value between 0 and 1. When the two time intervals are exactly the same or highly overlapping, , making The value is close to 1; if the two do not intersect or have only a small overlap, then The value is close to 0, which effectively distinguishes the difference in time:

[0105] ;

[0106] ;

[0107] ;

[0108] in, Indicates the target entity records, Indicates the target entity records, Represents the target entity record and the target entity record The time similarity between Represents the target entity record and the target entity record The length of the overlapping interval between Represents the target entity record and the target entity record The length of the joint interval between Represents the target entity record Time interval[ , ]’s left endpoint, Represents the target entity record Time interval[ , ]’s right endpoint, Represents the target entity record Time interval[ , ]’s left endpoint, Represents the target entity record Time interval[ , ]’s right endpoint, Indicates taking the maximum value, Indicates taking the minimum value;

[0109] Based on the different spatial information corresponding to the different target entity records, the spatial similarity between the different target entity records is calculated. The spatial similarity mainly measures the consistency of two entities (referring to entity records) in terms of geographic location, spatial range, and topological relationship:

[0110] ;

[0111] Distance Similarity The linear attenuation function of the following formula is used for calculation. hour, A value of 0:

[0112] ;

[0113] Topological similarity According to the topological relationship between two entities (referring to entity records), a strict correspondence is performed, and different scores are assigned to different topological relationships. In the present invention, the corresponding result is shown in the following formula:

[0114] ;

[0115] Regional similarity Often for surface entities, it can be defined as the area overlap rate. When two areas highly overlap, The value is close to 1; if the two are completely disjoint, A value of 0:

[0116] ;

[0117] in, Represents the target entity record and the target entity record The spatial similarity between Represents the target entity record The spatial representation of Represents the target entity record The spatial representation of Representation space representation and spatial representation Between (i.e. target entity records and the target entity record The following meanings are the same as here and will not be repeated here. Representation space representation and spatial representation The topological similarity between Representation space representation and spatial representation The regional similarity between represents the weight of distance similarity, represents the weight of topological similarity, represents the weight of region similarity, , Representation space representation and spatial representation The distance between Indicates the maximum acceptable distance threshold set. Representation space representation and spatial representation The topological relationship between Representation space representation and spatial representation The area of ​​the intersection between the two surface entities, that is, the area of ​​the overlapping area between the two; Representation space representation and spatial representation The area of ​​the union between the two surface entities is the area of ​​the union of the two surface entities, that is, the total area covered by the two;

[0118] Based on the different attribute information under each attribute type corresponding to the different target entity records, the attribute similarity between the different target entity records is calculated. Attribute similarity is used to measure the similarity between two entities (referring to entity records) in terms of name, category, data source, timestamp, etc. Each attribute can be regarded as a one-dimensional feature. For different attribute types (string, numeric, enumeration), the corresponding metric function is selected and then weighted summation is performed:

[0119] ;

[0120] For string attributes, the edit distance is used to measure similarity. The edit distance refers to the minimum number of single-character edits (insertion, deletion, or replacement) required to convert one string into another. When two strings are completely identical, the edit distance is 0 and the string attribute similarity is 1. The larger the difference, the lower the string attribute similarity:

[0121] ;

[0122] For enumeration or categorical attributes, if two entities (referring to entity records) have the same category, the categorical attribute similarity is assigned a value of 1; if they are of similar categories, the categorical attribute similarity is assigned a higher score; if they are of different categories, the categorical attribute similarity is assigned a value of 0. For example, the categorical attribute similarity between "school" and "college" can be considered 0.8, but the categorical attribute similarity between "school" and "hospital" is considered 0:

[0123] ;

[0124] For numerical attributes, the absolute value of the difference between the values ​​of two entities (referring to entity records) can be normalized, such as the maximum possible value range [0, ] is normalized. If the absolute value of the difference does not exceed the tolerable range, the similarity of the numerical attribute is high:

[0125] ;

[0126] in, Represents the target entity record and the target entity record The attribute similarity between Indicates the Types of attributes, Indicates the The weight of the attribute type, , Represents the target entity record and the target entity record Between The similarity of sub-attributes under the attribute type, Indicates the sub-attribute similarity under the first attribute type (i.e. string attribute similarity), Represents the target entity record String instance, Represents the target entity record String instance, Represents a string instance With string instance The edit distance between Represents a string instance length, Represents a string instance length, Indicates the sub-attribute similarity under the second attribute type (i.e., category attribute similarity), Indicates a set score that is less than 1 and greater than 0.5, for example, 0.8. Represents the target entity record Category, Represents the target entity record Category, Representation category and categories The relationship between Indicates the sub-attribute similarity under the third attribute type (i.e., numerical attribute similarity). Represents the target entity record The numerical value of Represents the target entity record The numerical value of Indicates the maximum possible numerical threshold, that is, the maximum possible value range [0, ]'s right endpoint.

[0127] In one implementation of this embodiment, performing weighted summation on the temporal similarity, the spatial similarity, and the attribute similarity to obtain the spatiotemporal semantic similarity specifically includes:

[0128] According to the preset weight coefficients of temporal similarity, spatial similarity and attribute similarity, the temporal similarity between the different target entity records, the spatial similarity between the different target entity records and the attribute similarity between the different target entity records are weighted and summed to obtain the spatiotemporal semantic similarity between the different target entity records:

[0129] ;

[0130] in, Record for the target entity and the target entity record The spatiotemporal semantic similarity between is the weight coefficient of time similarity, is the weight coefficient of spatial similarity, is the weight coefficient of attribute similarity, ,By setting different weight coefficients empirically, the above formula can ,combine the effects of time, space and attributes.

[0131] Specifically, based on spatiotemporal semantic similarity, the present invention fuses different entity records corresponding to different instance objects of the same entity type that have coreference problems, and generates a unified urban multi-department knowledge graph, including: first, preprocessing different entity records, that is, unifying the time standards and spatial coordinate systems of different data sources, and unifying the spatiotemporal benchmarks to the CGCS2000 (China Geodetic Coordinate System 2000, 2000 National Geodetic Coordinate System) coordinate system and the UTC (Coordinated Universal Time) time standard; then, respectively calculating the time similarity, spatial similarity and attribute similarity between different entity records, and obtaining the comprehensive similarity (i.e., spatiotemporal semantic similarity) by comprehensively weighting each similarity; determining the coreference entities based on the comprehensive similarity, eliminating redundant nodes, merging redundant records, and eliminating the coreference problem.

[0132] Among them, in the actual data set, it is also necessary to select an appropriate similarity threshold θ based on the actual record situation. ≥θ, the two entities (referring to entity records) are considered to represent the same object (referring to instance objects), thereby triggering subsequent fusion operations and reducing the misfusion of irrelevant objects. Figure 4 As shown, for node objects (referring to instance nodes, such as Figure 4 E1 and E2 in ), retain the instance nodes with richer information or more trustworthy information (such as Figure 4 E1 in ), and remove redundant nodes (such as Figure 4 In E2), the information of the removed node (such as alias, child node relationship) is merged into the retained node. Here, the removed node is the redundant node, the retained node is the target node, and the child node relationship includes the record node connected to the redundant node (such as Figure 4 R3 in ) and the edges between redundant nodes and record nodes. It should be noted that Figure 4 C represents the concept node corresponding to the entity type, E1 and E2 represent different instance nodes corresponding to different instance objects, and R1, R2, and R3 represent different record nodes corresponding to different entity records.

[0133] It should be noted that the urban multi-source heterogeneous data knowledge graph provided in the embodiment of the present invention is an urban knowledge graph that can be applied to heterogeneous data retrieval. By querying the index information and general data in the knowledge graph, it can realize the query and indexing of heterogeneous data based on entity objects.

[0134] In summary, the present invention achieves efficient integration of heterogeneous data from multiple urban departments and the construction of a knowledge graph through the following technical solutions: Based on a hierarchical ontology model, a semantic framework covering seven major entity categories, such as natural resources and building facilities, is constructed, supporting multi-level mapping to accommodate differences in data granularity; rule engines (such as SQL and Geopandas) are utilized to automatically extract and transform heterogeneous data from multiple sources, including relational data, GIS data, BIM data, and IoT data, into graph structures; a spatiotemporal semantic weighting algorithm is proposed to integrate temporal similarity, spatial similarity, and attribute similarity, unifying the spatiotemporal datum (UTC time standard and CGCS2000 coordinate system) to eliminate redundancy; and, in conjunction with the Neo4j graph database, nodes and edges are organized at the concept, instance, and record levels, supporting complex cross-domain queries. This invention can break down departmental data silos and provide a globally consistent knowledge base for cross-departmental collaborative governance, emergency command, and dynamic data integration.

[0135] The present invention has the following beneficial effects: realizing cross-departmental data integration, eliminating information islands, and realizing multi-source associated storage of core entity records; realizing efficient knowledge fusion, and reducing the generation of redundant records based on spatiotemporal semantic similarity algorithms; providing intelligent retrieval support, and supporting data association and query for the same entity object based on a multi-level model of a graph database; and being applied to scenarios such as smart city governance, emergency command, and cross-departmental collaborative analysis, providing intelligent decision-making support for comprehensive urban governance.

[0136] In addition, based on the above-mentioned urban multi-department heterogeneous data knowledge graph construction method, the present invention also provides an urban multi-department heterogeneous data knowledge graph construction system, wherein the preferred embodiment of the urban multi-department heterogeneous data knowledge graph construction system is as follows: Figure 5 As shown, specifically including:

[0137] Ontology model construction module 01: used to determine the entity objects, descriptive attributes and association relationships of the heterogeneous data ontology model of multiple departments in the city;

[0138] Knowledge extraction module 02: for extracting entity records, attribute information and relationship information from heterogeneous data of multiple departments of a city using a rule engine based on the entity objects, the descriptive attributes and the association relationships;

[0139] Initial knowledge graph construction module 03: used to organize the entity objects, the entity records, the attribute information and the relationship information in combination with the storage structure of the graph database to obtain an initial knowledge graph;

[0140] Target knowledge graph construction module 04: used to calculate the spatiotemporal semantic similarity based on the initial knowledge graph, and to fuse different entity records with co-reference problems based on the spatiotemporal semantic similarity to obtain the target knowledge graph.

[0141] In addition, based on the above-mentioned urban multi-department heterogeneous data knowledge graph construction method and system, the present invention also provides a terminal, wherein a preferred embodiment of the terminal is as follows: Figure 6 As shown, it specifically includes a processor 10, a memory 20 and a display 30. Figure 6 Only some components of the terminal are shown, but it should be understood that it is not required to implement all the components shown, and more or fewer components may be implemented instead. The terminal supports calling knowledge graph construction and query services through an API interface.

[0142] In some embodiments, the memory 20 may be an internal storage unit of the terminal, such as a hard disk or memory of the terminal. In other embodiments, the memory 20 may also be an external storage device of the terminal, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, or a flash card equipped on the terminal. Furthermore, the memory 20 may include both an internal storage unit of the terminal and an external storage device. The memory 20 is used to store application software and various types of data installed on the terminal, such as program code of the terminal. The memory 20 may also be used to temporarily store data that has been output or is about to be output. In one embodiment, the memory 20 stores a city multi-department heterogeneous data knowledge graph construction program 40, which can be executed by the processor 10, thereby implementing the steps of the city multi-department heterogeneous data knowledge graph construction method in this application.

[0143] In some embodiments, the processor 10 can be a central processing unit (CPU), a microprocessor, or other data processing chip, used to run program codes or process data stored in the memory 20, such as executing a city multi-department heterogeneous data knowledge graph construction program 40.

[0144] In some embodiments, the display 30 may be an LED display, a liquid crystal display, a touch-sensitive liquid crystal display, an OLED (Organic Light-Emitting Diode) touch screen, etc. The display 30 is used to display information on the terminal and to display a visual user interface.

[0145] In one embodiment, when the processor 10 executes the city multi-department heterogeneous data knowledge graph construction program 40 in the memory 20, the steps of the city multi-department heterogeneous data knowledge graph construction method as described above are implemented.

[0146] The present invention also provides a computer-readable storage medium, wherein the computer-readable storage medium stores a program for constructing a knowledge graph of heterogeneous data of multiple urban departments. When the program for constructing a knowledge graph of heterogeneous data of multiple urban departments is executed by a processor, the steps of the method for constructing a knowledge graph of heterogeneous data of multiple urban departments as described above are implemented.

[0147] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or terminal comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or terminal. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or terminal comprising the element.

[0148] Of course, those skilled in the art will appreciate that all or part of the processes in the above-described method embodiments can be implemented by instructing related hardware (such as a processor, controller, etc.) through a computer program. The program can be stored in a computer-readable storage medium that can be read by a computer. When executed, the program can include the processes in the above-described method embodiments. The computer-readable storage medium can be a memory, a magnetic disk, an optical disk, etc.

[0149] It should be understood that the application of the present invention is not limited to the above examples. For those skilled in the art, improvements or changes can be made based on the above description. All these improvements and changes should fall within the scope of protection of the claims attached to the present invention.

Claims

1. A method for constructing a knowledge graph of heterogeneous data from multiple departments in a city, characterized by: The method for constructing a knowledge graph of heterogeneous data from multiple departments in a city includes: Determine the entity objects, descriptive attributes and association relationships of the heterogeneous data ontology model of multiple departments in the city; According to the entity objects, the descriptive attributes and the association relationships, a rule engine is used to extract entity records, attribute information and relationship information from heterogeneous data of multiple departments in the city; In combination with the storage structure of the graph database, the entity objects, the entity records, the attribute information and the relationship information are organized to obtain an initial knowledge graph; Calculating spatiotemporal semantic similarity based on the initial knowledge graph, and fusing different entity records with co-reference problems based on the spatiotemporal semantic similarity to obtain a target knowledge graph; The step of calculating the spatiotemporal semantic similarity based on the initial knowledge graph and fusing different entity records with coreference problems based on the spatiotemporal semantic similarity to obtain a target knowledge graph specifically includes: Performing spatiotemporal benchmark unification on different entity records corresponding to different instance objects under the same entity type in the initial knowledge graph to obtain different target entity records corresponding to different instance objects under the same entity type; Calculating the temporal similarity, spatial similarity and attribute similarity between the different target entity records respectively; Performing weighted summation on the temporal similarity, the spatial similarity, and the attribute similarity to obtain spatiotemporal semantic similarity; Comparing the spatiotemporal semantic similarity with a similarity threshold, if the spatiotemporal semantic similarity is not less than the similarity threshold, it is considered that the different entity records are different entity records corresponding to different instance objects under the same entity type and have a coreference problem; Retain the target node in the instance nodes of different instance objects corresponding to different entity records with coreference problems, remove redundant nodes in the instance nodes of different instance objects corresponding to different entity records with coreference problems, and merge information of the redundant nodes into the target node; Among them, the target node is the instance node with the largest amount of information or the highest information reliability among the instance nodes of different instance objects corresponding to different entity records with coreference problems, and the redundant node is the instance node other than the target node among the instance nodes of different instance objects corresponding to different entity records with coreference problems.

2. The method for constructing a knowledge graph of heterogeneous urban data from multiple departments according to claim 1 is characterized in that: Determining the entity objects, descriptive attributes, and association relationships of the city multi-department heterogeneous data ontology model specifically includes: Acquire heterogeneous data of multiple urban departments, determine key entities of an ontology model of the heterogeneous data of multiple urban departments in combination with the heterogeneous data of multiple urban departments, and classify and categorize the key entities to obtain entity types and instance objects, wherein the entity types and the instance objects constitute the entity objects; In combination with the characteristics of the heterogeneous data of multiple departments in the city, the descriptive attributes and association relationships of the ontology model of the heterogeneous data of multiple departments in the city are determined.

3. The method for constructing a knowledge graph of heterogeneous urban data from multiple departments according to claim 1 is characterized in that: The heterogeneous data of multiple departments in the city include relational data, GIS data, BIM data and IoT data; The method of extracting entity records, attribute information, and relationship information from heterogeneous data of multiple departments of a city using a rule engine based on the entity objects, the descriptive attributes, and the association relationships specifically includes: According to the entity objects, the descriptive attributes and the association relationships in the city multi-department heterogeneous data ontology model, extracting entity records corresponding to the entity objects, attribute information corresponding to the descriptive attributes and relationship information corresponding to the association relationships from the relational data based on an SQL rule engine; Extracting entity records corresponding to the entity objects, attribute information corresponding to the descriptive attributes, and relationship information corresponding to the relationship from the GIS data based on the Geopandas tool according to the entity objects, the descriptive attributes, and the relationship in the urban multi-department heterogeneous data ontology model; Extracting entity records corresponding to the entity objects, attribute information corresponding to the descriptive attributes, and relationship information corresponding to the relationship from the BIM data based on the Revit API according to the entity objects, the descriptive attributes, and the relationship in the city multi-department heterogeneous data ontology model; According to the entity objects, the descriptive attributes and the association relationships in the city's multi-department heterogeneous data ontology model, the entity records corresponding to the entity objects, the attribute information corresponding to the descriptive attributes and the relationship information corresponding to the association relationships are extracted from the IoT data based on the InfluxDBClient engine.

4. The method for constructing a knowledge graph of heterogeneous urban data from multiple departments according to claim 2 is characterized in that: The storage structure of the graph database is combined to organize the entity objects, the entity records, the attribute information, and the relationship information to obtain an initial knowledge graph, specifically including: Combined with the node-attribute-edge storage structure of the graph database, the entity type, the instance object, and the entity record are mapped to concept nodes with concept labels, instance nodes with instance labels, and record nodes with record labels respectively according to the three levels of concept layer, instance layer, and record layer; Representing the attribute information using the key-value pair attributes embedded in the record node; The relationship information is represented using typed edges between the record nodes to obtain an initial knowledge graph.

5. The method for constructing a knowledge graph of urban multi-department heterogeneous data according to claim 4 is characterized in that: The respectively calculating the temporal similarity, spatial similarity and attribute similarity between the different target entity records specifically includes: Calculate the time similarity between the different target entity records based on the different time information corresponding to the different target entity records: ; ; ; in, Indicates the target entity records, Indicates the target entity records, Represents the target entity record and the target entity record The time similarity between Represents the target entity record and the target entity record The length of the overlapping interval between Represents the target entity record and the target entity record The length of the joint interval between Represents the target entity record Time interval[ , ]’s left endpoint, Represents the target entity record Time interval[ , ]’s right endpoint, Represents the target entity record Time interval[ , ]’s left endpoint, Represents the target entity record Time interval[ , ]’s right endpoint, Indicates taking the maximum value, Indicates taking the minimum value; Calculate the spatial similarity between the different target entity records based on the different spatial information corresponding to the different target entity records: ; ; ; ; in, Represents the target entity record and the target entity record The spatial similarity between Represents the target entity record The spatial representation of Represents the target entity record The spatial representation of Representation space representation and spatial representation The distance similarity between Representation space representation and spatial representation The topological similarity between Representation space representation and spatial representation The regional similarity between represents the weight of distance similarity, represents the weight of topological similarity, represents the weight of region similarity, Representation space representation and spatial representation The distance between represents the maximum acceptable distance threshold, Representation space representation and spatial representation The topological relationship between Representation space representation and spatial representation The area of ​​the intersection between them, Representation space representation and spatial representation The area of ​​the union between them; Calculate the attribute similarity between the different target entity records based on the different attribute information under each attribute type corresponding to the different target entity records: ; ; ; ; in, Represents the target entity record and the target entity record The attribute similarity between Indicates the Types of attributes, Indicates the The weight of the attribute type, Represents the target entity record and the target entity record Between The similarity of sub-attributes under the attribute type, Indicates the sub-attribute similarity under the first attribute type, Represents the target entity record String instance, Represents the target entity record String instance, Represents a string instance With string instance The edit distance between Represents a string instance length, Represents a string instance length, Indicates the sub-attribute similarity under the second attribute type, Indicates a set score that is less than 1 and greater than 0.5, Represents the target entity record Category, Represents the target entity record Category, Representation category and categories The relationship between Indicates the similarity of sub-attributes under the third attribute type, Represents the target entity record The numerical value of Represents the target entity record The numerical value of Indicates the maximum possible numerical threshold.

6. The method for constructing a knowledge graph of urban multi-department heterogeneous data according to claim 5 is characterized in that: The weighted summation of the temporal similarity, the spatial similarity, and the attribute similarity to obtain the spatiotemporal semantic similarity specifically includes: According to the preset weight coefficients of temporal similarity, spatial similarity and attribute similarity, the temporal similarity between the different target entity records, the spatial similarity between the different target entity records and the attribute similarity between the different target entity records are weighted and summed to obtain the spatiotemporal semantic similarity between the different target entity records: ; in, Record for the target entity and the target entity record The spatiotemporal semantic similarity between is the weight coefficient of time similarity, is the weight coefficient of spatial similarity, is the weight coefficient of attribute similarity.

7. A system for constructing a knowledge graph of heterogeneous data from multiple departments in a city, characterized by: The system for constructing a knowledge graph of heterogeneous data of multiple departments in a city is applied to the method for constructing a knowledge graph of heterogeneous data of multiple departments in a city according to any one of claims 1 to 6, and the system for constructing a knowledge graph of heterogeneous data of multiple departments in a city comprises: Ontology model construction module: used to determine the entity objects, description attributes and association relationships of the heterogeneous data ontology model of multiple departments in the city; Knowledge extraction module: used to extract entity records, attribute information and relationship information from heterogeneous data of multiple departments of the city using a rule engine based on the entity objects, the descriptive attributes and the association relationships; An initial knowledge graph construction module is used to organize the entity objects, entity records, attribute information, and relationship information in combination with the storage structure of the graph database to obtain an initial knowledge graph; Target knowledge graph construction module: used to calculate the spatiotemporal semantic similarity based on the initial knowledge graph, and to fuse different entity records with co-reference problems based on the spatiotemporal semantic similarity to obtain the target knowledge graph.

8. A terminal, characterized in that: The terminal includes: a memory, a processor, and a city multi-department heterogeneous data knowledge graph construction program stored in the memory and runnable on the processor. When the city multi-department heterogeneous data knowledge graph construction program is executed by the processor, the steps of the city multi-department heterogeneous data knowledge graph construction method as described in any one of claims 1 to 6 are implemented.

9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a city multi-department heterogeneous data knowledge graph construction program, and when the city multi-department heterogeneous data knowledge graph construction program is executed by a processor, the steps of the city multi-department heterogeneous data knowledge graph construction method as described in any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Space-time multi-mode mixed data processing method, association method and indexing method

    CN113297395A

  • Heterogeneous data fusion indexing system and device based on knowledge graph

    CN116662342A