Knowledge Graph-Based Intelligent Data Resource Entry System
The knowledge graph-based intelligent data resource entry system solves the problem of the inability to assess the value and relationships of data resources, and realizes intelligent management and efficient storage of data resources.
Patent Information
- Application Number
- CN202511087820.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-05
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2045-08-05
AI Technical Summary
Existing technologies cannot accurately assess the value and relationships of data resources during storage, leading to low management efficiency.
An intelligent data resource entry system based on knowledge graphs is adopted. The system acquires target data resources through a data acquisition unit, constructs a basic knowledge graph through a basic graph construction unit, optimizes relationship distances through a relationship correction and optimization unit, and evaluates and enters the data into the table to perform resource fitness analysis and identify and store the data.
It enables the accurate identification and quantification of relationships between data resources, transforming into an intelligent management model based on value assessment, thereby improving the efficiency and accuracy of data resource management.
Smart Images

Figure CN120596486B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of knowledge graphs, and more particularly to an intelligent data resource table entry system based on knowledge graphs. Background Technology
[0002] With the rapid development of information technology and the widespread application of data-driven decision-making, various industries have generated massive amounts of data resources, which possess significant value and application potential. Currently, data resources are generally stored using traditional file storage or relational database storage. These storage methods merely preserve data and lack the ability to mine and analyze its intrinsic value. They fail to understand the value of different data resources and the relationships between them, resulting in a lack of effective data resource identification and consequently, inefficient subsequent data resource management. Summary of the Invention
[0003] This invention addresses the technical problem of low data resource management efficiency caused by the inability to accurately assess data value and relationships during data resource storage in existing technologies, by providing an intelligent data resource table entry system based on knowledge graphs.
[0004] The technical solution of the present invention to solve the above-mentioned technical problems is as follows:
[0005] This invention provides an intelligent data resource table entry system based on knowledge graphs, comprising: a data acquisition unit for acquiring target data resources to be entered into the table, wherein the target data resources include multiple sets of data, each set of data including main data and several related data, and there are several association categories between the main data and the several related data; a basic graph construction unit for acquiring multiple relational attributes of the main data between the multiple sets of data, configuring multiple basic relational distances, and constructing a basic knowledge graph; a relation correction and optimization unit for analyzing the similarity of several related data between each set of data, correcting the basic relational distances, obtaining multiple relational distances, and adjusting the basic knowledge graph to obtain a knowledge graph; and an entry evaluation unit for analyzing the resource fitness of the target data resources based on the scale of the knowledge graph and the data volume of the multiple sets of data, identifying the knowledge graph, and storing it in the table.
[0006] Optionally, the execution steps of the data acquisition unit include: collecting multiple sets of data to be entered into the table, wherein each set of data includes master data and several related data, and there are several related categories between the master data and the several related data; integrating multiple sets of data to obtain target data resources.
[0007] Optionally, the execution steps of the basic knowledge graph construction unit include: obtaining multiple relational attributes of master data among multiple sets of data; allocating and configuring multiple basic relational distances according to the multiple relational attributes; and connecting multiple sets of data, master data within each set of data, and several related data according to the multiple sets of data, multiple basic relational distances, multiple relational attributes, and several association categories to obtain a basic knowledge graph, wherein multiple relational attributes and multiple basic relational distances are used to connect the master data within multiple sets of data, and several association categories are used to connect the master data within each set of data and several related data.
[0008] Optionally, the execution steps of the basic map construction unit include: collecting a set of sample relationship attributes based on the data resource entry records within a historical time period; calculating the ratio of the average occurrence frequency of the same sample relationship attribute to the occurrence frequency of each sample relationship attribute to obtain multiple principal distance correction factors; obtaining a preset relationship distance; and using multiple principal distance correction factors to correct and calculate the preset relationship distance to obtain multiple basic relationship distances.
[0009] Optionally, the execution steps of the relationship correction and optimization unit include: acquiring multiple connected data groups within the basic knowledge graph, each data group including two sets of data; analyzing the similarity between the two sets of data within each data group, correcting the basic relationship distance, and obtaining multiple relationship distances; and using the multiple relationship distances to adjust the basic knowledge graph to obtain a knowledge graph.
[0010] Optionally, the execution steps of the relationship correction and optimization unit further include: calculating the similarity of several related data in two groups of data within each data group, and calculating the mean to obtain multiple similarities; calculating the ratio of each similarity to the mean of multiple similarities as multiple association distance correction factors; and using the multiple association distance correction factors to perform correction calculations on multiple basic relationship distances to obtain multiple relationship distances.
[0011] Optionally, the execution steps of the evaluation table entry unit include: statistically calculating the total length of relation distances within the knowledge graph as knowledge graph scale information; calculating the resource fitness of the target data resource based on the knowledge graph scale information and the number of multiple sets of data; and using the resource fitness to identify the knowledge graph and store it in a table.
[0012] The beneficial effects of this invention are:
[0013] The data acquisition unit acquires target data resources to be entered into the table. These resources consist of multiple sets of data, each containing master data and several related data points. Several association categories exist between the master data and the related data, providing a data foundation for subsequent association analysis and value assessment. The basic knowledge graph construction unit acquires multiple relational attributes between the master data sets, configures multiple basic relational distances, and constructs a basic knowledge graph. This establishes a preliminary data association framework based on the relational attributes between the master data, and quantifies the degree of association between data points using basic relational distances. The relational correction and optimization unit analyzes the similarity of several related data points between each set, corrects the basic relational distances, obtains multiple relational distances, adjusts the basic knowledge graph to obtain a knowledge graph, and further optimizes the relational distances through similarity analysis of related data, improving the accuracy of the knowledge graph and the precision of data value assessment. The evaluation and table entry unit analyzes the resource suitability of the target data resources based on the scale of the knowledge graph and the data volume of the multiple sets. The knowledge graph is then labeled and stored in the table, thereby quantifying and evaluating the overall value of the data resources and achieving intelligent labeling and efficient table entry management of data resources.
[0014] The above technical solutions enable accurate identification and quantification of relationships between data resources. Through knowledge graph construction and relationship distance optimization, precise assessment of data resource value is achieved. Simultaneously, resource fitness calculation and identification mechanisms transform data resource management from a traditional, simple storage model to an intelligent management model based on value assessment, enabling accurate quantitative assessment and efficient table entry management of data resource value. Attached Figure Description
[0015] Figure 1 A schematic diagram of the structure of the knowledge graph-based intelligent data resource table entry system provided by the present invention;
[0016] Figure 2 This is a schematic diagram illustrating the process of obtaining a knowledge graph provided by the present invention.
[0017] In the attached diagram, the components represented by each number are as follows:
[0018] Data acquisition unit 11, basic map construction unit 12, relationship correction and optimization unit 13, evaluation and table entry unit 14. Detailed Implementation
[0019] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0020] In the description of this invention, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the stated features. In the description of this invention, "a plurality of" means two or more, unless otherwise explicitly specified.
[0021] In the description of this invention, the term "for example" is used to mean "used as an example, illustration, or description." Any embodiment described as "for example" in this invention is not necessarily to be construed as being more preferred or advantageous than other embodiments. The following description is provided to enable any person skilled in the art to make and use the invention. Details are set forth in the following description for purposes of explanation. It should be understood that those skilled in the art will recognize that the invention can be made without using these specific details. In other instances, well-known structures and processes will not be described in detail to avoid obscuring the description of the invention with unnecessary detail. Therefore, the invention is not intended to be limited to the embodiments shown, but is consistent with the broadest scope of the principles and features disclosed herein.
[0022] like Figure 1 As shown, this embodiment of the invention provides an intelligent data resource table entry system based on knowledge graphs, including a data acquisition unit 11, a basic graph construction unit 12, a relationship correction and optimization unit 13, and an evaluation table entry unit 14.
[0023] The data acquisition unit 11 is used to acquire the target data resources to be entered into the table. The target data resources include multiple sets of data, each set of data including main data and several related data, and there are several association categories between the main data and the several related data.
[0024] Specifically, the data acquisition unit 11 is responsible for acquiring the target data resources to be entered into the table. For example, the data acquisition unit 11 acquires the target data resources from various data sources through a data acquisition interface. These target data resources refer to the original data set that needs to be structured and used for knowledge graph construction.
[0025] The target data resource contains multiple sets of data, each forming a relatively independent data unit. The data structure of the target data resource exhibits a specific hierarchical characteristic. Within each set of data, the data is further divided into two levels: primary data and related data. Primary data refers to the core identifying information of the set, possessing uniqueness and representativeness. Related data refers to auxiliary information associated with the primary data, numbering in the quantity, and including specific business data related to the primary data. The relationship between primary data and related data is not a simple containment relationship, but rather a semantic relationship established through several association categories. These association categories are classification labels describing the semantic relationship between primary data and related data, laying the foundation for the edge relationships in subsequent knowledge graph construction.
[0026] Taking the equipment operation and maintenance data management of a smart manufacturing enterprise as an example, the data acquisition unit 11 acquires equipment operation and maintenance data resources from multiple data sources, such as the enterprise's equipment management system, maintenance record system, and sensor monitoring system, as target data resources. These resources include multiple sets of equipment operation and maintenance data. The first set of data contains master data for "CNC-A001 CNC machine tool," and associated data includes fault codes "E001," "E005," and "E012," maintenance plans such as "replace spindle bearing," "adjust transmission system," and "calibrate coordinate system," fault parameters such as "temperature 75℃," "vibration 0.3mm," and "speed deviation 5%," and maintenance durations such as "4 hours," "2.5 hours," and "1 hour." The second set of data contains master data for "CNC-A002 CNC machine tool," and associated data includes fault codes "E003," "E001," and "E008," maintenance plans such as "adjust tool parameters," "replace spindle bearing," and "clean cooling system," fault parameters such as "vibration 0.5mm," "temperature 78℃," and "pressure abnormality," and maintenance durations such as "2 hours," "4 hours," and "3 hours." The main data for the third set of data is "CNC-B001 type CNC machine tool". The associated data includes fault codes "E001" and "E015", repair solutions "replace spindle bearing" and "replace hydraulic pump", fault parameters "temperature 78℃" and "insufficient hydraulic pressure", and repair time "3.5 hours" and "5 hours".
[0027] Semantic relationships are established between these master data and related data through association categories such as "fault information", "maintenance plan", "operating parameters" and "maintenance records".
[0028] Through this structured data organization method, the data acquisition unit 11 can provide standardized data input for subsequent knowledge graph construction, ensuring that the semantic integrity and relevance of the data are effectively maintained.
[0029] The basic knowledge graph construction unit 12 is used to obtain multiple relational attributes of the master data between multiple sets of data, configure multiple basic relational distances, and construct a basic knowledge graph.
[0030] Specifically, the basic knowledge graph construction unit 12 is responsible for constructing an initial knowledge graph structure based on the target data resources provided by the data acquisition unit 11. Specifically, the basic knowledge graph construction unit 12 analyzes multiple relational attributes between master data sets, quantifies the strength of these relational attributes, and constructs a basic knowledge graph with a spatial topological structure accordingly.
[0031] First, the basic knowledge graph construction unit 12 acquires multiple relational attributes of the master data across multiple sets of data. These relational attributes refer to the semantic association features between the master data of different sets of data. These attributes reflect semantic relationships such as similarity, relevance, or hierarchy between the master data, providing a semantic foundation for subsequent distance calculations. Based on the acquired relational attributes, the basic knowledge graph construction unit 12 further configures multiple basic relational distances. Basic relational distances are quantified spatial distance values representing the strength of association between different master data nodes in the knowledge graph. The configuration process of basic relational distances is based on the semantic weights and frequency of occurrence of relational attributes, using a mathematical calculation model to transform abstract semantic relationships into quantifiable distance parameters. Subsequently, based on the acquired relational attributes and configured basic relational distances, the basic knowledge graph construction unit 12 constructs a basic knowledge graph. The basic knowledge graph is a graph structure with master data as nodes, relational attributes as edges, and basic relational distances as edge weights, simultaneously containing the association relationships between master data and related data within each set of data.
[0032] In the aforementioned application scenario of equipment operation and maintenance data management, the basic knowledge graph construction unit 12 analyzes multiple relationship attributes between various equipment models, including relationship attributes such as "same model," "same series," and "first-level similar models." Based on the frequency and importance of these relationship attributes, corresponding basic relationship distances are configured. For example, the basic relationship distance corresponding to the "same model" relationship attribute is relatively small, the basic relationship distance corresponding to the "same series" relationship attribute is moderate, and the basic relationship distance corresponding to the "first-level similar models" relationship attribute is relatively large. The final constructed basic knowledge graph uses each equipment model as a node, relationship attributes as connecting edges, and basic relationship distances as edge weights. It also includes the connection relationships between each equipment node and its associated data such as fault information, maintenance plans, operating parameters, and maintenance records.
[0033] Through this graph structure construction method, the basic graph construction unit 12 can transform the original multiple sets of data into a structured knowledge graph representation, providing a basic graph structure framework for subsequent relationship optimization and resource evaluation.
[0034] The relation correction and optimization unit 13 is used to analyze the similarity of several related data between each group of data, correct the basic relation distance, obtain multiple relation distances, and adjust the basic knowledge graph to obtain a knowledge graph.
[0035] Specifically, the relation correction and optimization unit 13 is responsible for accurately correcting and optimizing the basic knowledge graph based on the similarity analysis of associated data. Specifically, the relation correction and optimization unit 13 analyzes the similarity characteristics of associated data among multiple sets of data, identifies the diversity and value differences of the data, and dynamically adjusts the distance of basic relations accordingly, thereby obtaining a final knowledge graph that more accurately reflects the value of the data.
[0036] First, the relationship correction and optimization unit 13 analyzes the similarity of several related data points between each group of data. The similarity of related data refers to the quantified value of the degree of similarity between corresponding related data points in different groups of data. This similarity reflects the repetitiveness and diversity characteristics of the data. Higher similarity indicates that the related data are more similar and the data diversity is poorer, thus affecting the overall value assessment of the data. Then, based on the analysis results of the related data similarity, the relationship correction and optimization unit 13 corrects the basic relationship distance. The correction process incorporates the similarity information of the related data into the adjustment of the basic relationship distance. When the similarity of the related data is high, it indicates poor data diversity and low data value, and the corresponding relationship distance is increased; when the similarity of the related data is low, it indicates good data diversity and high data value, and the corresponding relationship distance is decreased, thus forming multiple corrected relationship distances.
[0037] By applying corrected relation distances, relation correction optimization unit 13 structurally adjusts the basic knowledge graph to obtain the final knowledge graph. The final knowledge graph, while maintaining the original node and edge relationships, more accurately reflects the true strength of the correlations and value differences between data points through optimized relation distances.
[0038] In the aforementioned application scenario of equipment operation and maintenance data management, the relationship correction and optimization unit 13 analyzes the similarity of related data between different groups of data. For example, the three groups of data, "CNC-A001 CNC machine tool," "CNC-A002 CNC machine tool," and "CNC-B001 CNC machine tool," all contain the fault code "E001," the repair plan "replace spindle bearing," and temperature-related fault parameters ("temperature 75℃," "temperature 78℃," and "temperature 78℃," respectively), indicating that these related data have high similarity and poor data diversity. Based on this analysis result, the relationship correction and optimization unit 13 increases the distance between these equipment nodes to reflect the decrease in their data value.
[0039] Through this similarity-based dynamic correction mechanism, the relationship correction optimization unit 13 can transform the static basic knowledge graph into an optimized knowledge graph that accurately reflects the distribution of data value, providing a more accurate graph structure foundation for subsequent resource assessment.
[0040] The evaluation and table entry unit 14 is used to analyze the resource fitness of the target data resource based on the scale of the knowledge graph and the amount of data of the multiple sets of data, and to identify and store the knowledge graph in a table.
[0041] Specifically, the evaluation and input unit 14 is responsible for comprehensively evaluating the optimized knowledge graph, quantifying the value level of the target data resources, and realizing the tagged storage of the knowledge graph. Specifically, the evaluation and input unit 14 calculates a quantitative indicator reflecting the overall value of the data resources by comprehensively analyzing the structural and data scale characteristics of the knowledge graph, and then uses this indicator to tag the knowledge graph and store it in a structured manner.
[0042] First, evaluation unit 14 performs a comprehensive analysis based on the knowledge graph scale and the data volume of multiple sets of data. The knowledge graph scale refers to the total distance between all relations in the knowledge graph, reflecting the overall complexity of the graph structure and the density of connections between data. The data volume of multiple sets of data refers to the number of data sets contained in the target data resource, reflecting the size of the data resource. The knowledge graph scale and the data volume of multiple sets of data together constitute the basic dimensions for data resource value assessment. Then, based on the analysis of the knowledge graph scale and data volume, evaluation unit 14 calculates the resource fitness of the target data resource. Resource fitness is a comprehensive quantitative indicator that measures the overall value of a data resource. It is calculated as the reciprocal of the ratio of the knowledge graph scale to the data volume. When the knowledge graph scale is small and the data volume is large, it indicates that the data are closely connected, the data value is high, and the resource fitness is high; conversely, when the knowledge graph scale is large and the data volume is small, it indicates that the data are loosely connected, the data value is low, and the resource fitness is low.
[0043] After obtaining the resource fitness, evaluation and table entry unit 14 identifies the knowledge graph and stores it in the table. The identification process refers to assigning corresponding value level labels to the knowledge graph according to the numerical range of the resource fitness, realizing a visual representation of the data resource value; table entry and storage refers to storing the identified knowledge graph in the system database in a structured data format for easy subsequent retrieval, analysis and application.
[0044] In the aforementioned application scenario of equipment operation and maintenance data management, the total distance of all relationships in the optimized knowledge graph, as evaluated by the input unit 14, is used as the knowledge graph scale. The number of equipment operation and maintenance data groups included is used as the data volume, and the reciprocal of the ratio of the two is calculated to obtain the resource fitness. For example, if the knowledge graph scale is 100 distance units and the data volume is 20 groups of equipment data, then the resource fitness is the reciprocal of 20 / 100 = 0.2, which is 5.0. Based on this resource fitness value, the input unit 14 assigns a corresponding value level identifier to the knowledge graph and stores it in the equipment operation and maintenance data resource library, realizing the quantitative management and visual display of the value of equipment operation and maintenance data.
[0045] Through this quantitative assessment and identification storage mechanism, the assessment entry unit 14 can transform complex knowledge graphs into data resources with clear value identifiers, providing a structured and visualized data value reference for data asset management, thereby achieving accurate quantitative assessment and efficient entry management of data resource value.
[0046] Furthermore, the execution steps of the data acquisition unit 11 include:
[0047] Collect multiple sets of data to be entered into the table. Each set of data includes master data and several related data. There are several association categories between the master data and the related data.
[0048] Integrate multiple sets of data to obtain the target data resources.
[0049] In one feasible implementation, the data acquisition unit 11 first collects raw data from multiple data sources through a data acquisition interface, forming multiple sets of data to be entered into the table. Each set of data has a unified data structure, including two levels: master data and several related data. The master data serves as the core identifier of the set and has a unique characteristic; the related data provides detailed content as auxiliary information related to the master data. Semantic relationships are established between the master data and the related data through several association categories, which define the specific association types and semantic meanings between the data.
[0050] Then, the data acquisition unit 11 performs unified formatting and structured integration on the collected data sets, eliminating format differences between different data sources and ensuring the consistency and integrity of the data structure. Through integration processing, the multiple sets of data to be entered into the table are organized into a unified data set, forming the target data resource. This target data resource maintains the independence and correlation of each set of data, providing standardized data input for subsequent knowledge graph construction.
[0051] In the aforementioned application scenario of equipment operation and maintenance data management, the data acquisition unit 11 first collects operation and maintenance data of each piece of equipment from different data sources such as the equipment management system, maintenance record system, and sensor monitoring system. This forms multiple sets of data with "CNC-A001 CNC machine tool," "CNC-A002 CNC machine tool," and "CNC-B001 CNC machine tool" as the main data. Each set of data contains corresponding fault codes, maintenance plans, fault parameters, maintenance duration, and other related data. Then, the data acquisition unit 11 unifies the format and integrates the structure of these multiple sets of data from different systems, ultimately forming a target data resource containing complete equipment operation and maintenance information.
[0052] Through the data acquisition process, the data acquisition unit 11 can effectively extract and integrate the required data resources from a complex multi-source data environment, providing a data foundation for subsequent knowledge graph construction and data value assessment.
[0053] Furthermore, the execution steps of the basic map construction unit 12 include:
[0054] Retrieve multiple relational attributes of master data among multiple sets of data;
[0055] Assign and configure multiple basic relationship distances based on multiple relationship attributes;
[0056] Based on multiple sets of data, multiple basic relation distances, multiple relation attributes, and several association categories, a basic knowledge graph is obtained by connecting multiple sets of data and the master data and several related data within each set of data. Specifically, multiple relation attributes and multiple basic relation distances are used to connect the master data within multiple sets of data, and several association categories are used to connect the master data and several related data within each set of data.
[0057] In a preferred embodiment, firstly, the basic graph construction unit 12 identifies and extracts various relational features between the master data of different groups of data through semantic analysis. Relationship attributes are feature labels that describe the semantic associations between master data, reflecting multi-dimensional relational features such as similarity, hierarchy, and categorization between master data. Multiple relationship attributes constitute a complete descriptive system of relationships between master data, providing a semantic basis for subsequent distance quantization.
[0058] Then, based on the semantic weight and statistical features of each relation attribute, the basic graph construction unit 12 assigns a corresponding distance value to each relation attribute, resulting in multiple basic relation distances. The basic relation distance is a quantitative representation of a relation attribute in the knowledge graph space, and the magnitude of the distance value reflects the association strength of the corresponding relation attribute. The higher the frequency of a relation attribute, the more common it is, and the smaller the corresponding basic relation distance; conversely, the lower the frequency of a relation attribute, the less common and less valuable it is, and the larger the corresponding basic relation distance.
[0059] Subsequently, based on multiple sets of data, multiple basic relation distances, multiple relation attributes, and several association categories, the multiple sets of data, as well as the main data and several associated data within each set, are connected to obtain a basic knowledge graph. Specifically, the basic graph construction unit 12 constructs two levels of connections: the first level uses multiple relation attributes and multiple basic relation distances to connect the main data within multiple sets of data, forming an association network between main data nodes; the second level uses several association categories to connect the main data and several associated data within each set of data, forming a hierarchical structure of main data and associated data. Through these two levels of connections, the basic graph construction unit 12 constructs a complete basic knowledge graph.
[0060] In the aforementioned application scenario of equipment operation and maintenance data management, the basic knowledge graph construction unit 12 first analyzes the relationship attributes between master data such as "CNC-A001 CNC machine tool," "CNC-A002 CNC machine tool," and "CNC-B001 CNC machine tool," identifying multiple relationship attributes such as "same model," "same series," and "first-level similar models." Then, based on the importance and frequency of these relationship attributes, a smaller basic relationship distance is assigned to the "same model" relationship attribute, a medium basic relationship distance is assigned to the "same series" relationship attribute, and a larger basic relationship distance is assigned to the "first-level similar models" relationship attribute. Finally, the basic knowledge graph construction unit 12 uses these relationship attributes and basic relationship distances to connect each equipment model node, and uses association categories such as "fault information," "maintenance plan," "operating parameters," and "maintenance records" to connect each equipment node with its corresponding associated data nodes, ultimately constructing a basic knowledge graph containing inter-equipment relationships and internal equipment information.
[0061] Through this hierarchical graph construction process, the basic graph construction unit 12 can transform the original multiple sets of data into a knowledge graph representation with a clear structure and well-defined relationships, laying the foundation for subsequent relationship optimization processing.
[0062] Furthermore, such as Figure 2 As shown, the execution steps of the basic map construction unit 12 include:
[0063] Based on the data resource records entered into the table within a historical time period, collect a set of sample relationship attributes;
[0064] Calculate the ratio of the average frequency of occurrence of the same sample relation attribute to the frequency of occurrence of each sample relation attribute to obtain multiple main distance correction factors;
[0065] Get the preset relationship distance;
[0066] Multiple primary distance correction factors are used to correct and calculate the preset relationship distances, thereby obtaining multiple basic relationship distances.
[0067] In a preferred embodiment, the basic graph construction unit 12 achieves precise configuration of basic relationship distances through statistical analysis methods based on historical data. Specifically, the basic graph construction unit 12 calculates the principal distance correction factor of relationship attributes based on the statistical characteristics of historical data resources, and corrects the preset relationship distance accordingly, thereby obtaining the basic relationship distances that reflect the true strength of the relationships.
[0068] First, the basic map construction unit 12 accesses the historical database to retrieve and analyze data resource entries that have been stored in tables within a historical time period. Various relational attribute information is extracted from these data resource entries to form a sample relational attribute set. This sample relational attribute set contains all relational attribute types that appeared during the historical process and their corresponding statistical data, providing a data foundation for subsequent frequency analysis and value assessment.
[0069] Then, the basic graph construction unit 12 statistically analyzes the frequency of occurrence of each relation attribute in the sample relation attribute set, and then calculates the average frequency of occurrence of all relation attributes of the same type. By calculating the ratio of the average frequency of occurrence to the frequency of occurrence of each specific relation attribute, the principal distance correction factor for that relation attribute is obtained. The principal distance correction factor reflects the scarcity and importance of a specific relation attribute relative to the average level of similar attributes; a larger ratio indicates a scarcer relation attribute and lower data value.
[0070] Next, the basic graph construction unit 12 reads the preset relation distance from the configuration parameters. The preset relation distance is the baseline distance value set during initialization, serving as the basis parameter for subsequent correction calculations. Subsequently, the basic graph construction unit 12 uses the principal distance correction factor corresponding to each relation attribute as a correction coefficient, multiplying it by the preset relation distance to obtain the corrected basic relation distance. When the principal distance correction factor is high, it indicates that the relation attribute is relatively scarce, occurs infrequently, and has low data value, resulting in a relatively large basic relation distance; when the principal distance correction factor is low, it indicates that the relation attribute is relatively common, occurs frequently, and has high data value, resulting in a relatively small basic relation distance.
[0071] In the aforementioned application scenario of equipment operation and maintenance data management, the basic graph construction unit 12 collects a set of sample relationship attributes from historical equipment operation and maintenance data entry records, including relationship attributes such as "same model," "same series," and "first-level similar models" and their historical occurrence frequency. For example, statistical analysis reveals that the "same model" relationship attribute occurs more frequently, while the "first-level similar models" relationship attribute occurs less frequently. The ratio of the average occurrence frequency to the occurrence frequency of each relationship attribute is calculated, revealing that the primary distance correction factor for the "same model" relationship attribute is lower, while the primary distance correction factor for the "first-level similar models" relationship attribute is higher. Finally, these primary distance correction factors are used to correct the preset relationship distances, assigning a larger basic relationship distance to the "same model" relationship attribute and a smaller basic relationship distance to the "first-level similar models" relationship attribute.
[0072] By configuring distances based on historical data statistics, the basic graph construction unit 12 can dynamically adjust the distances of basic relationships according to the actual importance and scarcity of relationship attributes, ensuring that the distance parameters in the knowledge graph can accurately reflect the strength of the real relationships between data.
[0073] Furthermore, the execution steps of the relationship correction and optimization unit 13 include:
[0074] Obtain multiple interconnected data groups within the aforementioned basic knowledge graph, with each data group containing two sets of data;
[0075] Analyze the similarity between two data sets within each data set, correct the basic relation distance, and obtain multiple relation distances;
[0076] The knowledge graph is obtained by adjusting the basic knowledge graph using the multiple relationship distances.
[0077] In a preferred embodiment, firstly, the relationship correction and optimization unit 13 identifies pairs of primary data nodes directly connected by relational attributes in the basic knowledge graph through graph traversal, and groups each pair of connected primary data nodes and their corresponding data content into a data group. The data group is the basic unit for similarity analysis; each data group contains two complete sets of data, including their respective primary data and several corresponding related data. In this way, the relationship correction and optimization unit 13 transforms the connection relationships in the graph structure into a set of data groups suitable for similarity calculation.
[0078] Then, the relationship correction and optimization unit 13 analyzes the similarity of several related data points between two data groups within each data group. By comparing the content, type, and numerical characteristics of the related data in the two data groups, a similarity index of the related data is calculated. Based on the similarity analysis results, the relationship correction and optimization unit 13 performs correction calculations on the basic relationship distance corresponding to that data group: when the similarity of the related data is high, it indicates poor data diversity and low data value, and the corrected relationship distance increases; when the similarity of the related data is low, it indicates good data diversity and high data value, and the corrected relationship distance decreases. By processing all data groups, multiple corrected relationship distances are obtained.
[0079] Subsequently, the relation correction and optimization unit 13 applies the corrected relation distances to the corresponding edges of the basic knowledge graph, replacing the original basic relation distances. Through this distance update operation, the structure of the basic knowledge graph is optimized and adjusted, forming the final knowledge graph. The final knowledge graph, while maintaining the original node and edge topological relationships, more accurately reflects the true value relationships and correlation strength between data through the optimized relation distances.
[0080] In the aforementioned application scenario of equipment operation and maintenance data management, the relationship correction and optimization unit 13 first identifies connected pairs of equipment nodes in the basic knowledge graph, such as data groups like ("CNC-A001 CNC machine tool", "CNC-A002 CNC machine tool") and ("CNC-A001 CNC machine tool", "CNC-B001 CNC machine tool"). Then, it analyzes the similarity of the associated data between the two sets of equipment data within each data group, finding that "CNC-A001 CNC machine tool" shares the fault code "E001" and the maintenance plan "replace the spindle bearing" with both "CNC-A002 CNC machine tool" and "CNC-B001 CNC machine tool," indicating a high degree of similarity. Based on this similarity analysis, the relationship correction and optimization unit 13 increases the relationship distance between these equipment nodes to reflect the decrease in their data value. Finally, the corrected relationship distance is applied to update the basic knowledge graph, obtaining a knowledge graph that accurately reflects the distribution of equipment data value.
[0081] Through a dynamic correction process based on similarity analysis, the relationship correction optimization unit 13 can transform the static basic knowledge graph into an optimized knowledge graph that accurately reflects the differences in data value, providing a more accurate graph structure foundation for subsequent resource value assessment.
[0082] Furthermore, the execution steps of the relationship correction and optimization unit 13 also include:
[0083] Calculate the similarity of several related data points between two data groups within each data group, and calculate the mean to obtain multiple similarity scores;
[0084] Calculate the ratio of each similarity to the mean of multiple similarities, and use it as a multiple association distance correction factor;
[0085] The multiple correlation distance correction factors are used to correct and calculate the distances of multiple basic relationships, thereby obtaining multiple relationship distances.
[0086] In a preferred embodiment, the relationship correction and optimization unit 13 achieves precise correction of the basic relationship distance through refined similarity calculation and value assessment. Specifically, the relationship correction and optimization unit 13 transforms the similarity features of the associated data into quantified value indicators through multi-level calculations, and adjusts the basic relationship distance accordingly.
[0087] First, the relationship correction and optimization unit 13 calculates the similarity value between corresponding related data in each of the two data groups within each data group. For example, it compares the similarity of fault codes, maintenance plans, and fault parameters in the two data groups. After the calculation is completed, the mean of the similarity of all related data within the data group is calculated to obtain the comprehensive similarity of the data group. By performing this operation on all data groups, multiple similarity values representing the degree of similarity of related data in different data groups are obtained. Then, the relationship correction and optimization unit 13 calculates the overall mean of the similarity of all data groups, and then calculates the ratio of the similarity of each data group to the overall mean to obtain multiple association distance correction factors. The association distance correction factor reflects the degree of deviation of the similarity of a specific data group relative to the overall similarity level. When the ratio is greater than 1, it indicates that the similarity of the data group is higher than the average level, the data diversity is poor, and the data value is low; when the ratio is less than 1, it indicates that the similarity of the data group is lower than the average level, the data diversity is good, and the data value is high.
[0088] Subsequently, the relationship correction and optimization unit 13 uses the association distance correction factor corresponding to each data group as a correction coefficient, multiplying it by the basic relationship distance corresponding to that data group to obtain the corresponding relationship distance. When the association distance correction factor is high, it indicates high data similarity and poor diversity, so the basic relationship distance is increased for correction; when the association distance correction factor is low, it indicates low data similarity and good diversity, so the basic relationship distance is decreased for correction. Through this correction calculation, multiple relationship distances reflecting the differences in the true value of the data are obtained.
[0089] In the aforementioned application scenario of equipment operation and maintenance data management, the relationship correction and optimization unit 13 calculates the similarity of fault codes, maintenance plans, fault parameters, and maintenance time for each data group ("CNC-A001 CNC machine tool", "CNC-A002 CNC machine tool"). Then, it calculates the mean of these similarities to obtain the similarity of the data group. By performing the same operation on all equipment data groups, multiple similarities are obtained. The mean of all similarities is calculated, and the ratio of the similarity of each data group to the mean of all similarities is calculated to obtain the association distance correction factor for each data group. Subsequently, these association distance correction factors are used to correct the corresponding basic relationship distances, assigning larger relationship distances to equipment data groups with higher similarity and smaller relationship distances to equipment data groups with lower similarity.
[0090] Through similarity analysis and value calculation mechanisms, the relationship correction and optimization unit 13 can more accurately identify and quantify the value differences between data, achieve precise correction of the basic relationship distance, and thus construct a knowledge graph that truly reflects the distribution of data resource value.
[0091] Furthermore, the execution steps of the evaluation input unit 14 include:
[0092] The total length of relation distances within the knowledge graph is statistically calculated and used as the knowledge graph scale information.
[0093] Based on the knowledge graph scale information and the number of multiple sets of data, the resource fitness of the target data resource is calculated.
[0094] The knowledge graph is identified and stored in a table using the resource fitness.
[0095] In a preferred embodiment, firstly, the evaluation input unit 14 traverses the graph, accessing all edge connections in the knowledge graph, obtaining the relation distance value corresponding to each edge, and summing them to obtain the total relation distance length. The total relation distance length reflects the overall structural complexity of the knowledge graph and the density of relationships between data, serving as scale information for the knowledge graph. When the total relation distance length is small, it indicates low similarity and high diversity among data, indicating high data value; when the total relation distance length is large, it indicates high similarity and low diversity among data, indicating low data value.
[0096] Subsequently, the number of multiple data sets contained in the target data resource in Table 14 is evaluated. Then, the ratio of knowledge graph scale information to the number of multiple data sets is calculated, and the reciprocal of this ratio is taken as the resource fitness. Resource fitness is a comprehensive quantitative indicator that measures the overall value of data resources. Its calculation formula is: Resource Fitness = Number of Data Sets / Knowledge Graph Scale Information. The smaller the ratio, the smaller the total relation distance and the larger the data volume, indicating high data value and sufficient quantity, and a higher resource fitness. Conversely, the larger the ratio, the larger the total relation distance and the smaller the data volume, indicating low data value and insufficient quantity, and a lower resource fitness.
[0097] Next, the evaluation and entry unit 14 assigns corresponding value level labels to the knowledge graph based on the calculated resource fitness. Using a preset value level threshold range, knowledge graphs with different resource fitness ranges are labeled as high-value, medium-value, low-value, etc. After completing the value labeling, the evaluation and entry unit 14 stores the labeled knowledge graphs in a structured data format in the database, achieving standardized entry and management of the knowledge graphs.
[0098] In the aforementioned application scenario of equipment operation and maintenance data management, the evaluation table unit 14 calculates the sum of the relationship distances between all equipment nodes in the optimized knowledge graph. For example, the relationship distance between "CNC-A001 CNC machine tool" and "CNC-A002 CNC machine tool" is 8 units, the relationship distance between "CNC-A001 CNC machine tool" and "CNC-B001 CNC machine tool" is 5 units, and the relationship distance between "CNC-A002 CNC machine tool" and "CNC-B001 CNC machine tool" is 2 units, totaling 15 distance units, which is used as the knowledge graph scale information. Simultaneously, the number of included equipment data groups is counted as 3 groups. The ratio of the knowledge graph scale information to the number of data groups is calculated to be 15 / 3 = 5. The resource fitness is the reciprocal of this ratio, i.e., 1 / 5 = 0.2. Based on the resource fitness value of 0.2, according to the preset value level threshold range (e.g., resource fitness ≥ 0.5 is high value, 0.2 ≤ resource fitness < 0.5 is medium value, and resource fitness < 0.2 is low value), the evaluation entry unit 14 assigns a medium value level label to the knowledge graph and stores it in the data resource library to realize the standardized entry management of the knowledge graph.
[0099] Through quantitative evaluation and identification storage mechanisms, the evaluation entry unit 14 can transform complex knowledge graphs into data resources with clear value identifiers, realizing accurate quantitative evaluation and efficient entry management of data resource value, and providing a structured and visualized data value reference for data asset management.
[0100] It should be noted that, in the above embodiments, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant description of other embodiments.
[0101] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0102] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0103] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0104] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0105] Although preferred embodiments of the invention have been described, those skilled in the art, once they have learned the basic inventive concept, can make other changes and modifications to these embodiments.
[0106] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of this invention and its equivalents, this invention also intends to include these modifications and variations.
Claims
1. A knowledge graph-based intelligent data resource table entry system, characterized in that, The system includes: The data acquisition unit is used to acquire target data resources to be entered into the table. The target data resources include multiple sets of data, which are multiple sets of equipment operation and maintenance data acquired from multiple data sources such as the enterprise's equipment management system, maintenance record system, and sensor monitoring system. Each set of data includes master data and several related data. There are several association categories between the master data and the several related data. The several association categories include fault information, maintenance plan, operating parameters, and maintenance records. Semantic relationships are established between the master data and the related data through the several association categories. The basic knowledge graph construction unit is used to obtain multiple relational attributes of the master data between multiple sets of data, configure multiple basic relational distances, and construct a basic knowledge graph. The relation correction and optimization unit is used to analyze the similarity of several related data between each group of data, correct the basic relation distances, obtain multiple relation distances, and adjust the basic knowledge graph to obtain a knowledge graph, including: Obtain multiple interconnected data groups within the aforementioned basic knowledge graph, with each data group containing two sets of data; Analyze the similarity between two data sets within each data set, correct the basic relation distance, and obtain multiple relation distances, including: Calculate the similarity of several related data points between two data groups within each data group, and calculate the mean to obtain multiple similarity scores; Calculate the ratio of each similarity to the mean of multiple similarities, and use it as multiple association distance correction factors; Using the aforementioned multiple association distance correction factors, the distances of multiple basic relationships are corrected and calculated to obtain multiple relationship distances; The knowledge graph is obtained by adjusting the basic knowledge graph using the multiple relation distances. The evaluation and table entry unit is used to analyze the resource fitness of the target data resource based on the scale of the knowledge graph and the data volume of the multiple sets of data, identify the knowledge graph, and store it in a table, including: The total length of relation distances within the knowledge graph is statistically calculated and used as the knowledge graph scale information. Based on the knowledge graph scale information and the number of multiple sets of data, the resource fitness of the target data resource is calculated. The knowledge graph is identified and stored in a table using the resource fitness.
2. The knowledge graph-based intelligent data resource table entry system according to claim 1, characterized in that, The execution steps of the data acquisition unit include: Collect multiple sets of data to be entered into the table. Each set of data includes master data and several related data. There are several association categories between the master data and the related data. Integrate multiple sets of data to obtain the target data resources.
3. The intelligent data resource entry system based on knowledge graphs according to claim 1, characterized in that, The execution steps of the basic map construction unit include: Retrieve multiple relational attributes of master data among multiple sets of data; Assign and configure multiple basic relationship distances based on multiple relationship attributes; Based on multiple sets of data, multiple basic relation distances, multiple relation attributes, and several association categories, a basic knowledge graph is obtained by connecting multiple sets of data and the master data and several related data within each set of data. Specifically, multiple relation attributes and multiple basic relation distances are used to connect the master data within multiple sets of data, and several association categories are used to connect the master data and several related data within each set of data.
4. The knowledge graph-based intelligent data resource table entry system according to claim 1, characterized in that, The execution steps of the basic atlas construction unit include: Based on the data resource records entered into the table within a historical time period, collect a set of sample relationship attributes; Calculate the ratio of the average frequency of occurrence of the same sample relation attribute to the frequency of occurrence of each sample relation attribute to obtain multiple main distance correction factors; Get the preset relationship distance; Multiple primary distance correction factors are used to correct and calculate the preset relationship distances, thereby obtaining multiple basic relationship distances.
Citation Information
Patent Citations
Data management method, system and device based on knowledge graph and medium
CN112685405A
Knowledge graph optimization method and device suitable for network security situation awareness data
CN120090833A