Multi-source heterogeneous data management method and system based on intelligent atlas
By constructing intelligent graphs and performing standardized processing and dynamic association, the problems of inconsistent formats and static association relationships in the management of multi-source heterogeneous data are solved, achieving efficient data integration and accurate scenario adaptation, and improving the overall performance of data management.
Patent Information
- Application Number
- CN202511115781.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-11
- Publication Date
- 2025-11-21
AI Technical Summary
Existing technologies for managing multi-source heterogeneous data suffer from inconsistent formats, static relationships, rigid tag management, and a lack of full lifecycle optimization mechanisms. This results in low data integration efficiency, insufficient correlation analysis capabilities, and an inability to meet the needs of precise and scenario-based applications.
By constructing an intelligent graph, receiving heterogeneous data sources and performing standardized processing, an intelligent graph of multiple types of entities and their relationships is established, enabling dynamic association and flexible adaptation of the tagging system. Combined with a full-process monitoring and optimization mechanism, a closed-loop management is formed.
It improves the integration efficiency and correlation analysis capabilities of multi-source heterogeneous data, enhances the adaptability of data in multiple fields, supports accurate classification and scenario adaptation, and provides intelligent data management solutions.
Smart Images

Figure CN120994750A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data management technology, and more specifically, to a method and system for managing multi-source heterogeneous data based on intelligent graphs. Background Technology
[0002] In the era of big data, data generated in industries, government affairs, finance, and other fields exhibits characteristics of multi-source and heterogeneity, including structured database tables, semi-structured XML / JSON data, and unstructured text and images. Existing data management technologies suffer from the following shortcomings: First, the large differences in formats among multi-source heterogeneous data, coupled with a lack of unified standardized processing mechanisms, leads to low data integration efficiency and semantic inconsistencies. Second, the relationships between data are mostly statically defined, making it difficult to dynamically adapt to new data and to intuitively represent complex relationships between entities. Third, tagging systems are mostly fixed classifications, unable to dynamically adjust with data content updates or business scenario changes, resulting in poor adaptability of data retrieval and application scenarios. Fourth, the lack of monitoring and optimization mechanisms throughout the data lifecycle makes it difficult to iteratively improve data management performance based on actual application results. These problems restrict the in-depth utilization of multi-source heterogeneous data and fail to meet the needs of precise and scenario-based data applications. Therefore, there is an urgent need for an intelligent data management solution that can achieve unified and standardized processing of multi-source heterogeneous data, dynamic relationship management, flexible adaptation of the tag system, and full-process optimization and iteration. This solution can address the problems of low data integration efficiency, static relationships, rigid tag management, and lack of closed-loop optimization mechanisms in existing technologies, and meet the needs of various fields for in-depth correlation analysis and scenario-based applications of multi-source data. Summary of the Invention
[0003] This invention provides a multi-source heterogeneous data management method and system based on intelligent graphs. It receives at least two heterogeneous data sources through a preset interface and performs standardized processing, solving the problem of inconsistent data formats across multiple sources. Based on standardized data, it constructs an intelligent graph containing various types of entities such as people, houses, and vehicles, as well as relationships such as residence and ownership, achieving visualized and structured management of data relationships. When new data is added, dynamic entity association is achieved through weighted matching of unique identifier attributes and auxiliary description attributes, and the association weights are updated based on time decay coefficients and data importance to ensure the timeliness and accuracy of the graph. Combined with a multi-dimensional tag system covering data categories, application scenarios, and source attributes, it achieves linked updates of tags and the graph through a tag mapping rule base and a tag-entity mapping relationship table, adapting to changes in data content and business scenario switching. Based on graph relationships and the tag system, it retrieves and aggregates data to generate thematic applications, and optimizes the graph weight model and tag mapping rules through full-process monitoring logs and performance indicators, forming a process of "data access - graph construction - dynamic update - tag management - application generation - optimization iteration". The closed-loop management mechanism effectively improves the integration efficiency, correlation analysis capabilities, and application adaptability of multi-source heterogeneous data, providing an intelligent solution for collaborative data management in multiple fields.
[0004] This application provides a method for managing multi-source heterogeneous data based on intelligent graphs, including the following steps: The system receives at least two heterogeneous data sources through a preset interface and performs format standardization processing to obtain standardized data. Based on the standardized data, an intelligent graph containing multiple types of entities and the relationships between entities is constructed. When new standardized data is detected, the system automatically identifies the entity attributes it contains and matches them with the existing entity nodes in the intelligent graph. The standardized data is labeled and dynamically managed based on a pre-defined multi-dimensional labeling system. Based on the entity association relationships and multi-dimensional tag system of the intelligent graph, data retrieval and aggregation are performed to generate thematic data applications that meet business scenarios; Real-time monitoring of the entire process of data access, intelligent map updates, tagging governance and application services, and recording of data operation logs and performance indicators; Based on the logs and performance metrics, optimize the association weight calculation model and multi-dimensional label mapping rules of the intelligent graph.
[0005] In the multi-source heterogeneous data management method based on intelligent graphs described in this application, the heterogeneous data sources include: Structured data, semi-structured data, and unstructured data; Structured data includes relational database table data, semi-structured data includes XML and JSON format data, and unstructured data includes text, images, and log files.
[0006] In the multi-source heterogeneous data management method based on intelligent graphs described in this application, the format standardization process specifically includes: It identifies and processes missing values, outliers, and duplicate data using preset rules; Implement cross-data source field semantic mapping based on a pre-defined data dictionary; Duplicate or low-value data is removed by calculating feature similarity.
[0007] In the multi-source heterogeneous data management method based on intelligent graphs described in this application, the multi-type entities and the relationships between entities include: The various types of entities include people, houses, vehicles, enterprises, and government affairs; The relationships between entities include the residential relationship between a person and a house, the ownership relationship between a person and a vehicle, the employment relationship between a person and a company, the relationship between a person and the handling of government affairs, and the registration relationship between a company and a house.
[0008] In the multi-source heterogeneous data management method based on intelligent graphs described in this application, the construction of the intelligent graph specifically involves: Entity recognition models are used to extract entity and attribute information from standardized data. Semantic relationships between entities are identified through a relation extraction model, and initial weights based on data confidence are assigned to these relationships. A graph database is used to store entities, attributes, and relationships, forming an intelligent graph with a hierarchical structure.
[0009] In the multi-source heterogeneous data management method based on intelligent graphs described in this application, the step of automatically identifying the entity attributes contained in newly added standardized data and matching them with the attributes of existing entity nodes in the intelligent graph includes: Entity attributes are divided into unique identifier attributes and auxiliary descriptive attributes; Set the first matching weight for the unique identifier attribute, set the second matching weight for the auxiliary description attribute, and calculate the overall matching degree; A successful match is determined when the overall matching score exceeds a preset threshold; otherwise, a failed match is determined. After a successful match, the association weights between entity nodes are dynamically updated based on the time decay coefficient of the newly added data and the data importance weight.
[0010] In the multi-source heterogeneous data management method based on intelligent graphs described in this application, the multi-dimensional tagging system includes: Data category tags include source data tags, dictionary-type data tags, scenario-type data tags, scenario application data tags, and theme / topic application data tags; Application scenarios are tagged with tags for population management, traffic control, tourism services, water monitoring, and epidemic prevention and control. Source attribute tags, including tags for carrier source, internet platform source, and government department source.
[0011] In the multi-source heterogeneous data management method based on intelligent graphs described in this application, the dynamic management includes: When the data content is updated, the tag mapping rule base is automatically triggered to update the tag attributes of the corresponding data; When business scenarios change, scenario application tags are automatically added or adjusted based on scenario association rules; After the tags are updated, the tag attributes of the corresponding entity nodes in the smart graph are updated synchronously through the tag-entity mapping relationship table, so as to realize the linkage update between the tag system and the smart graph.
[0012] Secondly, this application provides a multi-source heterogeneous data management system based on intelligent graphs, characterized in that it includes: The data access and standardization module is used to receive at least two heterogeneous data sources through a preset interface, and to perform format standardization processing on the heterogeneous data sources to obtain standardized data. The intelligent graph construction module is used to construct an intelligent graph containing multiple types of entities and the relationships between entities based on the standardized data. The entity matching and update module is used to automatically identify the entity attributes contained in newly added standardized data and match them with the existing entity nodes in the smart graph when the data is detected. The tag management module is used to perform tag mapping and dynamic management of the standardized data based on a preset multi-dimensional tag system; The thematic application generation module is used to perform data retrieval and aggregation based on the entity association relationship and multi-dimensional tag system of the intelligent graph, and generate thematic data applications that meet business scenarios; The monitoring and optimization module is used to monitor the entire process of data access, intelligent graph update, tagging governance and application services in real time, record data operation logs and performance indicators, and optimize the association relationship weight calculation model and multi-dimensional tag mapping rules of the intelligent graph based on the logs and performance indicators.
[0013] The system also includes a memory and a processor. The memory contains a program for a multi-source heterogeneous data management method based on intelligent graphs. When the program for the multi-source heterogeneous data management method based on intelligent graphs is executed by the processor, it performs the following steps: The system receives at least two heterogeneous data sources through a preset interface and performs format standardization processing to obtain standardized data. Based on the standardized data, an intelligent graph containing multiple types of entities and the relationships between entities is constructed. When new standardized data is detected, the system automatically identifies the entity attributes it contains and matches them with the existing entity nodes in the intelligent graph. The standardized data is labeled and dynamically managed based on a pre-defined multi-dimensional labeling system. Based on the entity association relationships and multi-dimensional tag system of the intelligent graph, data retrieval and aggregation are performed to generate thematic data applications that meet business scenarios; Real-time monitoring of the entire process of data access, intelligent map updates, tagging governance and application services, and recording of data operation logs and performance indicators; Based on the logs and performance metrics, optimize the association weight calculation model and multi-dimensional label mapping rules of the intelligent graph.
[0014] As shown above, this invention achieves intelligent management of multi-source heterogeneous data throughout the entire process, from access, standardization, and correlation modeling to dynamic updates, scenario applications, and optimization iterations, by constructing a collaborative management mechanism of intelligent graphs and a multi-dimensional tagging system. Its core lies in using intelligent graphs to intuitively present complex relationships between entities and support dynamic updates, using a multi-dimensional tagging system to achieve accurate data classification and scenario adaptation, and continuously optimizing model parameters through full-process monitoring and performance analysis. This effectively solves pain points in traditional data management such as difficulty in format integration, static correlation relationships, and rigid tag management. This invention not only improves the integration efficiency and correlation analysis capabilities of multi-source heterogeneous data but also enhances the adaptability of data to diverse business scenarios, providing reliable technical support for refined data applications in government, industry, finance, and other fields.
[0015] Other features and advantages of this application will be set forth in the following description and will be apparent in part from the description or may be learned by practicing embodiments of this application. The objectives and other advantages of this application may be realized and obtained by means of the structures particularly pointed out in the written description and the accompanying drawings. Attached Figure Description
[0016] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments of this application will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0017] Figure 1The high-level flowchart of the multi-source heterogeneous data management method based on intelligent graph provided in the embodiments of this application is used to intuitively present the core technical framework and the whole process logic, clearly showing the key links and collaborative relationships of data access, processing, graph construction, updating, tag management, application generation and optimization, and providing a visual top-level reference for technology implementation, deployment and optimization.
[0018] Figure 2 A flowchart illustrating the multi-source heterogeneous data management method based on intelligent graphs provided in this application embodiment; Figure 3 A flowchart illustrating the construction of an intelligent graph based on a multi-source heterogeneous data management method provided in this application embodiment; Figure 4 A flowchart illustrating the dynamic management of a multi-source heterogeneous data management method based on intelligent graphs provided in this application embodiment; Figure 5 This is a structural block diagram of a multi-source heterogeneous data management system based on intelligent graphs provided in an embodiment of this application. Detailed Implementation
[0019] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of the embodiments. The components of the embodiments of this application described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely represents selected embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.
[0020] It should be noted that similar reference numerals and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. Furthermore, in the description of this application, terms such as "first" and "second" are used only to distinguish descriptions and should not be construed as indicating or implying relative importance. It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of related data must comply with the relevant laws, regulations, and standards of the relevant countries and regions.
[0021] Please refer to Figure 1 , Figure 1This is a high-level flowchart of a multi-source heterogeneous data management method based on intelligent graphs in some embodiments of this application. The high-level flowchart forms a closed loop around "intelligent management of the entire lifecycle of multi-source heterogeneous data": First, multiple types of heterogeneous data are accessed through preset interfaces, and after standardization processing (including data cleaning, field mapping, and redundancy removal), data in a unified format is obtained; based on the standardized data, an intelligent graph is constructed, extracting entities such as people and houses, as well as relationships such as residence and ownership, and storing them as a hierarchical structure; when new data is added, the associated graph nodes are matched by weighted matching of entity attributes, and the associated weights are dynamically updated; combined with a multi-dimensional tag system, data tag mapping is realized, and tags are automatically adjusted and synchronized to graph entities as data is updated or scenarios change; aggregated data is retrieved using graph relationships and the tag system to generate scenario-based thematic applications; finally, through full-process monitoring and recording of logs and performance indicators, the graph weight model and tag mapping rules are continuously optimized, forming a complete technical link of "access - standardization - graph construction - dynamic update - tag management - application generation - optimization iteration," improving data management efficiency and scenario adaptability.
[0022] Please refer to Figure 2 , Figure 2 This is a flowchart of a multi-source heterogeneous data management method based on intelligent graphs in some embodiments of this application.
[0023] The first aspect of this invention discloses a multi-source heterogeneous data management method based on intelligent graphs, which is used in terminal devices such as computers and mobile terminals. This multi-source heterogeneous data management method based on intelligent graphs includes the following steps: S201. Receive at least two heterogeneous data sources through a preset interface and perform format standardization processing to obtain standardized data; S202. Based on the standardized data, construct an intelligent graph containing multiple types of entities and the relationships between entities; S203. When new standardized data is detected, automatically identify the entity attributes it contains and match them with the existing entity nodes in the intelligent graph. S204. Based on a preset multi-dimensional labeling system, perform label mapping and dynamic management on the standardized data; S205. Based on the entity association relationship and multi-dimensional tag system of the intelligent graph, perform data retrieval and aggregation to generate thematic data applications that meet business scenarios; S206. Real-time monitoring of the entire process of data access, intelligent map updates, labeling governance and application services, recording data operation logs and performance indicators; S207. Based on the logs and performance metrics, optimize the association weight calculation model and multi-dimensional label mapping rules of the intelligent graph.
[0024] The system receives at least two heterogeneous data sources with different data structures or types through a preset interface, performs format standardization processing on them, and converts unstructured, semi-structured, or structurally significantly different data into standardized data in a unified format, laying the foundation for subsequent data processing. Based on the standardized data, an intelligent graph containing multiple types of entities (such as people, events, and objects) and the relationships between entities (such as subordination, interaction, and causality) is constructed to present the inherent logic between data in a structured manner. When new standardized data is detected, the system automatically identifies the entity attribute information contained in the data and performs attribute matching with existing entity nodes in the intelligent graph to achieve dynamic updates and association strengthening of entity information. At the same time, based on a preset multi-dimensional tag system (covering attributes, categories, scenarios, etc.), the standardized data is tagged and mapped dynamically. The management mechanism ensures accurate correspondence and real-time adjustment between tags and data features. Relying on the entity relationships and multi-dimensional tag system of the intelligent graph, it performs precise data retrieval and aggregation analysis to generate thematic data applications that meet specific business scenario needs, thereby enhancing the business value of the data. During this process, it monitors the entire process in real time, including data access efficiency, intelligent graph update frequency, tagging accuracy, and application service response speed. It also records detailed data operation logs (such as data source, processing time, and operator) and performance indicators (such as processing time and accuracy). Finally, based on the recorded log data and performance indicators, it continuously optimizes the intelligent graph's relationship weight calculation model (to improve association accuracy) and the mapping rules for multi-dimensional tags (to enhance tag adaptability), forming a closed-loop optimization mechanism for data processing and application, ensuring the efficient and stable operation of the entire system.
[0025] According to an embodiment of the present invention, the heterogeneous data source includes: Structured data, semi-structured data, and unstructured data; Structured data includes relational database table data, semi-structured data includes XML and JSON format data, and unstructured data includes text, images, and log files.
[0026] The heterogeneous data sources encompass various types of data with different data organization structures and representations, specifically including three major categories: structured data, semi-structured data, and unstructured data. Structured data refers to data with strict format definitions and fixed data structures, typically represented by relational database table data. This type of data is usually stored in a standardized row-column format, with clear relationships between data items and easy querying and processing. Semi-structured data is a data type between structured and unstructured data. It has some structural information but does not strictly follow a fixed schema definition, specifically including XML and JSON format data. This type of data organizes information through tags or key-value pairs and has a certain degree of self-description. Unstructured data refers to data without a predefined data structure. Its content format is flexible and diverse, specifically including text data (such as documents and reports), image data (such as pictures and image files), and log files (such as system operation logs and operation record logs). This type of data usually requires specific parsing and processing to extract useful information. After the aforementioned heterogeneous data sources are connected to the system through a preset interface, they are transformed into standardized data in a unified format through format standardization processing. This provides comprehensive data input for subsequent intelligent map construction, entity attribute matching, and data application, forming a close connection with the entire data processing process in the overall technical solution, ensuring the integrity and consistency of data from access to application.
[0027] According to an embodiment of the present invention, the format standardization process specifically includes: It identifies and processes missing values, outliers, and duplicate data using preset rules; Implement cross-data source field semantic mapping based on a pre-defined data dictionary; Duplicate or low-value data is removed by calculating feature similarity.
[0028] For heterogeneous data sources such as structured, semi-structured, and unstructured data, the system first preprocesses the data using preset cleaning rules. This accurately identifies and addresses missing values (e.g., by supplementing missing information through interpolation or default value filling), outliers (e.g., by filtering and correcting data that deviates from the normal range through statistical analysis or business rules), and duplicate data (e.g., by identifying and retaining unique valid data through field comparison), ensuring data integrity and accuracy. Based on this, a preset data dictionary (containing semantic definitions, type specifications, and mapping relationships for fields from each data source) is used to achieve cross-data source field semantic mapping. This unifies and transforms fields from different data sources that have different expressions but the same semantics, eliminating semantic ambiguity caused by differences in field naming or format, and standardizing the data structure. Simultaneously, feature similarity calculation methods (e.g., algorithms based on text similarity and numerical distance) are used to further filter the processed data, removing data with high duplication or low value for business analysis, reducing data volume and improving subsequent processing efficiency. Through the above series of operations, heterogeneous data sources are transformed into standardized data with unified format, consistent semantics, and reliable quality. This provides a high-quality data foundation for subsequent intelligent graph construction (such as entity recognition and relationship extraction), entity attribute matching, and multi-dimensional label mapping. It also forms an organic connection with the entire process of data access, graph construction, and application services in the overall technical solution, ensuring the continuity and effectiveness of the data processing chain.
[0029] According to an embodiment of the present invention, the multiple types of entities and the relationships between entities include: The various types of entities include people, houses, vehicles, enterprises, and government affairs; The relationships between entities include the residential relationship between a person and a house, the ownership relationship between a person and a vehicle, the employment relationship between a person and a company, the relationship between a person and the handling of government affairs, and the registration relationship between a company and a house.
[0030] The various types of entities and their interrelationships are the core components of the intelligent graph. The various types of entities cover key objects with significant analytical value in business scenarios, specifically including people (such as individual natural persons), houses (such as residential and commercial buildings), vehicles (such as motor vehicles), enterprises (such as various legal entities or business entities), and government affairs (such as administrative approvals and public services). The interrelationships between entities are formed based on the business interactions and attribute characteristics of the entities, specifically including the residential relationship between people and houses due to actual residence, the ownership relationship between people and vehicles due to ownership, the appointment relationship between people and enterprises due to job assignment, the processing relationship between people and government affairs due to application, and the registration relationship between enterprises and houses due to registered address. These diverse entities are interconnected through the aforementioned relationships. After entity recognition models extract entity and attribute information, and relationship extraction models identify semantic relationships and assign initial weights, they are stored in a graph database to form a hierarchical intelligent graph. This graph accurately reflects the inherent connections between various entities in the business scenario, providing structured relational support for subsequent steps such as entity attribute matching, data retrieval and aggregation, and the generation of thematic data applications based on the intelligent graph. It is closely integrated with the data standardization processing, dynamic updates of the intelligent graph, and full-process monitoring and optimization in the overall technical solution, ensuring the integrity of the technical system and its business adaptability.
[0031] Please refer to Figure 3 , Figure 3 This is a flowchart illustrating the construction of an intelligent graph based on a multi-source heterogeneous data management method according to some embodiments of this application. According to embodiments of the present invention, the construction of the intelligent graph specifically involves: S301. Extract entity and attribute information from standardized data using an entity recognition model; S302. Identify semantic relationships between entities through a relation extraction model and assign initial weights to the relationships based on data confidence. S303. Use a graph database to store entities, attributes and relationships to form an intelligent graph with a hierarchical structure.
[0032] In the construction process based on standardized data, the first step is to use an entity recognition model to deeply analyze the standardized data, accurately extracting various entities (such as people, events, and objects) and their corresponding attribute information (such as entity features, parameters, and states), providing core data support for the construction of nodes in the intelligent graph. Subsequently, a relation extraction model is used to perform semantic-level association analysis on the extracted entities, identifying various semantic relationships such as subordination, interaction, and causation between entities. Each relationship is then assigned an initial weight based on the reliability of the data itself (i.e., data confidence), quantifying this process. The reliability of the relationships is assessed. Finally, a graph database is used as the storage medium to structure and store the extracted entities, entity attributes, and relationships between entities with initial weights. By sorting and organizing the entity hierarchy and relationship paths, an intelligent graph with a clear hierarchical structure is formed. This graph can intuitively and accurately present the inherent logic and relationship strength between entities in the data, providing a solid structured data foundation for subsequent operations such as entity attribute matching and data retrieval aggregation. It is organically connected with the data standardization processing, dynamic updates, and closed-loop optimization in the overall technical solution, ensuring the consistency and integrity of the technical system.
[0033] According to an embodiment of the present invention, when newly standardized data is detected, automatically identifying the entity attributes it contains and matching them with existing entity nodes in the intelligent graph includes: Entity attributes are divided into unique identifier attributes and auxiliary descriptive attributes; Set the first matching weight for the unique identifier attribute, set the second matching weight for the auxiliary description attribute, and calculate the overall matching degree; A successful match is determined when the overall matching score exceeds a preset threshold; otherwise, a failed match is determined. After a successful match, the association weights between entity nodes are dynamically updated based on the time decay coefficient of the newly added data and the data importance weight.
[0034] The process of automatically identifying entity attributes and matching them with existing entity nodes in the intelligent graph when newly standardized data is detected is as follows: First, the entity attributes extracted from the newly standardized data are classified into unique identifier attributes (such as a person's ID number or a house's property ownership certificate number, which can uniquely identify the entity) and auxiliary descriptive attributes (such as a person's age or a house's area, which are used to supplement the entity's characteristics). Differentiated matching weights are set for the two types of attributes. A first matching weight (relatively high) is set for unique identifier attributes to highlight their core role in entity matching, and a second matching weight (relatively low) is set for auxiliary descriptive attributes to supplement entity matching. The comprehensive matching degree is calculated based on the actual matching situation of the two types of attributes. The system presets a matching threshold. When the comprehensive matching degree exceeds the preset threshold, a decision is made... If a newly added entity attribute successfully matches an existing entity node in the intelligent graph, it is identified as the same entity. If the overall matching degree does not reach a preset threshold, the match is deemed a failure, and the entity can be included in the intelligent graph as a new entity node. After a successful match, the association weights between corresponding entity nodes in the intelligent graph are dynamically updated based on the time decay coefficient of the newly added data (which decreases in weight as the data is generated, reflecting the impact of data timeliness) and the data importance weight (set according to the value of the data to business analysis, such as higher weight for key attribute data). This ensures that the association weights reflect the latest status and value of the data in real time. This process is closely integrated with the construction of the intelligent graph (such as initial weight assignment) and the dynamic update mechanism, ensuring the accuracy and timeliness of entity associations in the intelligent graph. This provides reliable relational data support for subsequent data retrieval, aggregation, and the generation of thematic applications, ensuring the consistency and effectiveness of the overall technical solution.
[0035] According to an embodiment of the present invention, the multi-dimensional tagging system includes: Data category tags include source data tags, dictionary-type data tags, scenario-type data tags, scenario application data tags, and theme / topic application data tags; Application scenarios are tagged with tags for population management, traffic control, tourism services, water monitoring, and epidemic prevention and control. Source attribute tags, including tags for carrier source, internet platform source, and government department source.
[0036] The multi-dimensional tagging system is the core framework for tagging and dynamically managing standardized data. It constructs a tagging system from multiple dimensions, including the essential characteristics of the data, business application scenarios, and data source channels. Specifically, it includes data category tags, scenario application tags, and source attribute tags. Data category tags are used to identify the inherent type attributes of data, specifically covering source data tags (marking the basic type of the original data), dictionary-type data tags (marking dictionary-type reference data used for data standardization), scenario tag-type data tags (marking basic tag data related to business scenarios), scenario application data tags (marking data directly serving scenario applications), and thematic application data. The system comprises several tags: Tags (marking data for specific themes); Scenario application tags (associating data with corresponding business scenarios, including population management tags for population-related business scenarios, traffic management tags for transportation-related business scenarios, tourism service tags for tourism industry business scenarios, water monitoring tags for water management business scenarios, and epidemic prevention and control tags for epidemic prevention and control business scenarios); and Source attribute tags (tracing the data acquisition channels, including operator source tags (marking data from telecommunications operators), internet platform source tags (marking data from internet service platforms), and government department source tags (marking data from government departments). This multi-dimensional tagging system accurately maps standardized data with tags and dynamically adjusts the correspondence between tags and data in real time through a dynamic management mechanism. It provides multi-dimensional tag filtering conditions for entity association analysis and data retrieval aggregation based on intelligent graphs. It is closely integrated with entity attribute matching, intelligent graph updates, and the generation of thematic data applications, ensuring that data tags accurately reflect data characteristics and business needs, improving the targeting and effectiveness of data applications, and guaranteeing the consistency and integrity of the overall technical solution.
[0037] Please refer to Figure 4 , Figure 4 This is a flowchart illustrating the dynamic management of a multi-source heterogeneous data management method based on intelligent graphs according to some embodiments of this application. According to embodiments of the present invention, the dynamic management includes: S401. When the data content is updated, the tag mapping rule base is automatically triggered to update the tag attributes of the corresponding data. S402. When business scenarios change, automatically add or adjust scenario application tags based on scenario association rules; S403. After the label is updated, the label attributes of the corresponding entity nodes in the smart graph are updated synchronously through the label-entity mapping relationship table, so as to realize the linkage update between the label system and the smart graph.
[0038] Specifically, when the content of standardized data is updated (such as changes in entity attributes or modifications to data field values), the system automatically triggers a preset tag mapping rule library. Based on the updated data characteristics, it re-matches the tag mapping rules and synchronously updates the tag attributes of the corresponding data to ensure consistency between tags and data content. When a business scenario changes (such as switching from a population management scenario to a traffic management scenario), based on preset scenario association rules (covering the adaptation relationship of tags under different scenarios), the system automatically supplements (adds tags required for new scenarios) or adjusts (removes redundant tags from old scenarios) the scenario application tags, enabling the tag system to quickly respond to the changing needs of business scenarios. After the tags are updated, the system synchronizes the updated tag attribute information to the corresponding entity nodes in the intelligent graph through a preset tag-entity mapping relationship table (recording the corresponding association between tags and entity nodes in the intelligent graph), achieving real-time updates of entity node tag attributes and thus achieving a linkage update effect between the tag system and the intelligent graph. This dynamic management mechanism maintains the accuracy and timeliness of tag attributes by responding to data changes and scene switching in real time. It strengthens the collaborative relationship between the multi-dimensional tag system and the intelligent graph, and forms a closed loop with data standardization processing, entity attribute matching and thematic data application generation, further ensuring the consistency and business adaptability of data tags and intelligent graphs in the overall technical solution.
[0039] According to an embodiment of the present invention, a cross-system data collaborative governance step is further included, specifically: It connects to external heterogeneous systems through a pre-defined cross-domain interface protocol and receives standardized data from these systems. The format standardization processing module is used to uniformly transform external data, and the entity recognition and relation extraction model is used to extract the entities and relationships in the external data. By employing a weighted matching rule between unique identifier attributes and auxiliary description attributes, external entities and their relationships are matched with the intelligent graph across systems to generate a cross-system fusion graph. External data is labeled using a multi-dimensional labeling system, and internal and external data labels are updated collaboratively through a dynamic management mechanism.
[0040] First, a pre-defined cross-domain interface protocol is used to establish a connection with external heterogeneous systems (such as other business platforms, data middleware, etc.) to specifically receive standardized data that has been processed by external systems, ensuring the standardization and compatibility of cross-system data access. Next, the system's original format standardization processing module is reused to further unify and convert the received external data, ensuring that the external data format is completely consistent with the system's standardized data format. Then, entity information (such as personnel, enterprises, etc.) and inter-entity relationships (such as employment associations, registration associations, etc.) are extracted from the converted external data using entity recognition and relationship extraction models, providing foundational data for cross-system data fusion. Finally, the core rules for entity attribute matching from the original technical solution are adopted. This involves setting a first matching weight for unique identifier attributes and a second matching weight for auxiliary description attributes, then calculating a comprehensive matching degree. This process performs cross-system entity matching between externally extracted entities and their relationships and existing entity nodes in the system's intelligent graph. The matching results are used to integrate internal and external entity relationship information, generating a fused graph containing cross-system data. Finally, based on the system's pre-set multi-dimensional tagging system (covering data category tags, scenario application tags, and source attribute tags), external data is tagged and mapped. A dynamic tag management mechanism (including data update-triggered tag updates and scenario-switched tag adjustments) ensures coordinated updates of internal and external data tags, maintaining consistency between cross-system data tags and the system's tagging system. This step, by reusing the original system's core modules such as format standardization, entity recognition, attribute matching, and tag management, achieves effective cross-system data fusion without altering the technical features of the original claims. This provides richer data support for cross-domain business analysis and thematic data applications, ensuring the scalability and consistency of the technical solution.
[0041] Please refer to Figure 5 , Figure 5 This is a structural block diagram of a multi-source heterogeneous data management system based on intelligent graphs provided in an embodiment of this application.
[0042] A second aspect of this invention also discloses a multi-source heterogeneous data management system based on intelligent graphs, comprising: The data access and standardization module 501 is used to receive at least two heterogeneous data sources through a preset interface, and to perform format standardization processing on the heterogeneous data sources to obtain standardized data. The intelligent graph construction module 502 is used to construct an intelligent graph containing multiple types of entities and the relationships between entities based on the standardized data. The entity matching and update module 503 is used to automatically identify the entity attributes contained in newly added standardized data when the data is detected, and to match the attributes with existing entity nodes in the intelligent graph. The tag management module 504 is used to perform tag mapping and dynamic management of the standardized data based on a preset multi-dimensional tag system; Thematic application generation module 505 is used to perform data retrieval and aggregation based on the entity association relationship and multi-dimensional tag system of the intelligent graph, and generate thematic data applications that meet business scenarios. The monitoring and optimization module 506 is used to monitor the entire process of data access, intelligent graph update, tagging governance and application services in real time, record data operation logs and performance indicators, and optimize the association relationship weight calculation model and multi-dimensional tag mapping rules of the intelligent graph based on the logs and performance indicators.
[0043] The system also includes a memory and a processor. The memory includes a program for a multi-source heterogeneous data management method based on intelligent graphs. When the program for the multi-source heterogeneous data management method based on intelligent graphs is executed by the processor, it implements the steps of the multi-source heterogeneous data management method based on intelligent graphs as described in any one of the first aspects.
[0044] This invention discloses a multi-source heterogeneous data management method and system based on intelligent graphs. It constructs a complete data processing system encompassing data access, processing, graph construction, tag management, application generation, and optimization. The system receives heterogeneous data sources such as structured, semi-structured, and unstructured data through a preset interface. After format standardization processing, including missing value handling, field semantic mapping, and duplicate data removal, unified data is obtained. Based on this, entities and attributes are extracted through entity recognition, semantic associations are identified through relationship extraction and assigned initial weights, and then stored in a graph database to form an intelligent graph containing multiple types of entities such as people and houses, as well as relationships such as residence and employment. When new data is detected, entity attribute matching is performed according to the weight matching rules of unique identifiers and auxiliary descriptive attributes. If successful, the association weights are updated based on time decay coefficients and data importance weights. Simultaneously, tag mapping is performed using a multi-dimensional tag system covering data categories, application scenarios, and source attributes. Dynamic management is achieved through data update triggering rule bases, scene switching adjusting tags, and linkage updates with the intelligent graph. Based on graph associations and the tag system, aggregated data is retrieved to generate thematic applications, and the entire process is monitored in real time, recording logs and indicators to optimize the graph weight model and tag mapping rules. Building upon this foundation, a cross-system data collaborative governance mechanism has been expanded. By connecting to external standardized data through cross-domain interfaces, the original system processing modules are reused to complete the transformation, entity extraction and matching, generating a cross-system fusion map and realizing the collaborative updating of internal and external labels. All aspects of the entire solution are closely connected, ensuring the stability of core processes while possessing good scalability, and providing comprehensive support for data applications in multiple scenarios.
[0045] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods, such as: multiple units or components can be combined, or integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the various components shown or discussed can be through some interfaces, and the indirect coupling or communication connection between devices or units can be electrical, mechanical, or other forms.
[0046] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units. They may be located in one place or distributed across multiple network units. Some or all of the units may be selected to achieve the purpose of this embodiment according to actual needs.
[0047] In addition, in the various embodiments of the present invention, each functional unit can be integrated into one processing unit, or each unit can be a separate unit, or two or more units can be integrated into one unit; the integrated unit can be implemented in hardware or in the form of hardware plus software functional units.
[0048] Those skilled in the art will understand that all or part of the steps of the above method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a readable storage medium. When the program is executed, it performs the steps of the above method embodiments. The aforementioned storage medium includes various media that can store program code, such as mobile storage devices, read-only memory, random access memory, magnetic disks, or optical disks.
[0049] Alternatively, if the integrated units of this invention are implemented as software functional modules and sold or used as independent products, they can also be stored in a readable storage medium. Based on this understanding, the technical solutions of the embodiments of this invention, or the parts that contribute to the prior art, can be embodied in the form of a software product. This software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the methods described in the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as mobile storage devices, ROM, RAM, magnetic disks, or optical disks.
Claims
1. A multi-source heterogeneous data management method based on intelligent graphs, characterized in that, Includes the following steps: The system receives at least two heterogeneous data sources through a preset interface and performs format standardization processing to obtain standardized data. Based on the standardized data, an intelligent graph containing multiple types of entities and the relationships between entities is constructed. When new standardized data is detected, the system automatically identifies the entity attributes it contains and matches them with the existing entity nodes in the intelligent graph. The standardized data is labeled and dynamically managed based on a pre-defined multi-dimensional labeling system. Based on the entity association relationships and multi-dimensional tag system of the intelligent graph, data retrieval and aggregation are performed to generate thematic data applications that meet business scenarios; Real-time monitoring of the entire process of data access, intelligent map updates, tagging governance and application services, and recording of data operation logs and performance indicators; Based on the logs and performance metrics, optimize the association weight calculation model and multi-dimensional label mapping rules of the intelligent graph.
2. The multi-source heterogeneous data management method based on intelligent graphs according to claim 1, characterized in that, The heterogeneous data sources include: Structured data, semi-structured data, and unstructured data; Structured data includes relational database table data, semi-structured data includes XML and JSON format data, and unstructured data includes text, images, and log files.
3. The multi-source heterogeneous data management method based on intelligent graphs according to claim 1, characterized in that, The format standardization process is specifically as follows: It identifies and processes missing values, outliers, and duplicate data using preset rules; Implement cross-data source field semantic mapping based on a pre-defined data dictionary; Duplicate or low-value data is removed by calculating feature similarity.
4. The multi-source heterogeneous data management method based on intelligent graphs according to claim 1, characterized in that, The various types of entities and the relationships between them include: The various types of entities include people, houses, vehicles, enterprises, and government affairs; The relationships between entities include the residential relationship between a person and a house, the ownership relationship between a person and a vehicle, the employment relationship between a person and a company, the relationship between a person and the handling of government affairs, and the registration relationship between a company and a house.
5. The multi-source heterogeneous data management method based on intelligent graphs according to claim 1, characterized in that, The construction of the intelligent graph is specifically as follows: Entity recognition models are used to extract entity and attribute information from standardized data. Semantic relationships between entities are identified through a relation extraction model, and initial weights based on data confidence are assigned to these relationships. A graph database is used to store entities, attributes, and relationships, forming an intelligent graph with a hierarchical structure.
6. The multi-source heterogeneous data management method based on intelligent graphs according to claim 1, characterized in that, When newly standardized data is detected, the system automatically identifies the entity attributes it contains and matches them with existing entity nodes in the intelligent graph, including: Entity attributes are divided into unique identifier attributes and auxiliary descriptive attributes; Set the first matching weight for the unique identifier attribute, set the second matching weight for the auxiliary description attribute, and calculate the overall matching degree; A successful match is determined when the overall matching score exceeds a preset threshold; otherwise, a failed match is determined. After a successful match, the association weights between entity nodes are dynamically updated based on the time decay coefficient of the newly added data and the data importance weight.
7. The multi-source heterogeneous data management method based on intelligent graphs according to claim 1, characterized in that, The multi-dimensional tagging system includes: Data category tags include source data tags, dictionary-type data tags, scenario-type data tags, scenario application data tags, and theme / topic application data tags; Application scenarios are tagged with tags for population management, traffic control, tourism services, water monitoring, and epidemic prevention and control. Source attribute tags, including tags for carrier source, internet platform source, and government department source.
8. The multi-source heterogeneous data management method based on intelligent graphs according to claim 1, characterized in that, The dynamic management includes: When the data content is updated, the tag mapping rule base is automatically triggered to update the tag attributes of the corresponding data; When business scenarios change, scenario application tags are automatically added or adjusted based on scenario association rules; After the tags are updated, the tag attributes of the corresponding entity nodes in the smart graph are updated synchronously through the tag-entity mapping relationship table, so as to realize the linkage update between the tag system and the smart graph.
9. A multi-source heterogeneous data management system based on intelligent graphs, characterized in that: include: The data access and standardization module is used to receive at least two heterogeneous data sources through a preset interface, and to perform format standardization processing on the heterogeneous data sources to obtain standardized data. The intelligent graph construction module is used to construct an intelligent graph containing multiple types of entities and the relationships between entities based on the standardized data. The entity matching and update module is used to automatically identify the entity attributes contained in newly added standardized data and match them with the existing entity nodes in the smart graph when the data is detected. The tag management module is used to perform tag mapping and dynamic management of the standardized data based on a preset multi-dimensional tag system; The thematic application generation module is used to perform data retrieval and aggregation based on the entity association relationship and multi-dimensional tag system of the intelligent graph, and generate thematic data applications that meet business scenarios; The monitoring and optimization module is used to monitor the entire process of data access, intelligent graph update, tagging governance and application services in real time, record data operation logs and performance indicators, and optimize the association relationship weight calculation model and multi-dimensional tag mapping rules of the intelligent graph based on the logs and performance indicators.
10. A multi-source heterogeneous data management system based on intelligent graphs, characterized in that: The system also includes a memory and a processor. The memory contains a program for a multi-source heterogeneous data management method based on intelligent graphs. When the program for the multi-source heterogeneous data management method based on intelligent graphs is executed by the processor, it performs the following steps: The system receives at least two heterogeneous data sources through a preset interface and performs format standardization processing to obtain standardized data. Based on the standardized data, an intelligent graph containing multiple types of entities and the relationships between entities is constructed. When new standardized data is detected, the system automatically identifies the entity attributes it contains and matches them with the existing entity nodes in the intelligent graph. The standardized data is labeled and dynamically managed based on a pre-defined multi-dimensional labeling system. Based on the entity association relationships and multi-dimensional tag system of the intelligent graph, data retrieval and aggregation are performed to generate thematic data applications that meet business scenarios; Real-time monitoring of the entire process of data access, intelligent map updates, tagging governance and application services, and recording of data operation logs and performance indicators; Based on the logs and performance metrics, optimize the association weight calculation model and multi-dimensional label mapping rules of the intelligent graph.
Citation Information
Cited By
Multi-dimensional integration and intelligent management method for silkworm germplasm resources
CN121660397A
Data processing method for enterprise digital intelligence quality management
CN121836471A