Intelligent data resource table entering system based on knowledge graph

Through the intelligent data resource entry system based on knowledge graph, the problems of value assessment and relationship identification in data resource storage are solved, and intelligent management and efficient storage of data resources are realized.

CN120596486AActive Publication Date: 2025-09-05BEIJING XINLIU DATA TECH CO LTD +1
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202511087820.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-05
Publication Date
2025-09-05
Estimated Expiration
2045-08-05

AI Technical Summary

Technical Problem

The data resource storage method in existing technologies lacks the mining and analysis of the intrinsic value of data, and is unable to accurately evaluate the value and correlation of data, resulting in low management efficiency.

Method used

An intelligent data resource entry system based on knowledge graph is adopted. The target data resources are acquired through the data acquisition unit, the basic graph construction unit constructs the basic knowledge graph, the relationship correction optimization unit optimizes the relationship distance, and the evaluation entry unit performs resource fitness evaluation and identification storage.

Benefits of technology

It has achieved accurate identification and quantification of the relationships between data resources, transformed into an intelligent management model based on value assessment, and improved the efficiency and accuracy of data resource management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120596486A_ABST
    Figure CN120596486A_ABST
Patent Text Reader

Abstract

The invention provides an intelligent data resource table entering system based on a knowledge graph, and belongs to the field of knowledge graphs. Comprising a data acquisition unit, a basic atlas construction unit, a relation correction optimization unit and an evaluation table entry unit. Wherein the data acquisition unit is used for acquiring a target data resource to be entered into a table; the basic graph construction unit is used for constructing a basic knowledge graph; the relation correction optimization unit is used for adjusting the basic knowledge graph to obtain the knowledge graph; and the evaluation table entry unit is used for analyzing and obtaining the resource fitness of the target data resource according to the scale of the knowledge graph and the data volume of the multiple groups of data, and identifying the knowledge graph and storing the knowledge graph in a table. According to the method and the device, the technical problem of low data resource management efficiency caused by the fact that the data value and the association relationship cannot be accurately evaluated during data resource storage in the prior art is solved, and the technical effect of realizing accurate quantitative evaluation and efficient table entry management of the data resource value by constructing the knowledge graph and calculating the resource fitness is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of knowledge graphs, and in particular to an intelligent data resource entry system based on knowledge graphs. Background Art

[0002] The rapid development of information technology and the widespread application of data-driven decision-making have generated massive amounts of data resources across all industries. These data resources possess significant value and application potential. Currently, data resources are commonly stored using traditional file storage or relational databases. This storage method simply stores data, lacking the ability to mine and analyze the data's intrinsic value. This inability to understand the value of different data resources and the relationships between them leads to a lack of effective data resource identification, resulting in inefficient subsequent data resource management. Summary of the Invention

[0003] The present invention aims to solve the technical problem in the prior art that data value and association relationships cannot be accurately assessed when data resources are stored, resulting in low efficiency in data resource management, and provides an intelligent data resource entry system based on knowledge graph to solve the problem.

[0004] The technical solution of the present invention to solve the above technical problems is as follows: The present invention provides an intelligent data resource entry system based on a knowledge graph, comprising: a data acquisition unit, for acquiring target data resources to be entered into a table, wherein the target data resources include multiple groups of data, each group of data includes primary data and several related data, and there are several related categories between the primary data and the several related data; a basic graph construction unit, for acquiring multiple relationship attributes of the primary data between the multiple groups of data, configuring multiple basic relationship distances, and constructing a basic knowledge graph; a relationship correction and optimization unit, for analyzing the similarity of several related data between each group of data, correcting the basic relationship distance, obtaining multiple relationship distances, and adjusting the basic knowledge graph to obtain a knowledge graph; an evaluation entry unit, for analyzing and obtaining the resource adaptability of the target data resource based on the scale of the knowledge graph and the data volume of the multiple groups of data, and marking the knowledge graph and incorporating it into a table for storage.

[0005] Optionally, the execution steps of the data acquisition unit include: collecting multiple groups of data to be entered into a table, wherein each group of data includes main data and several related data, and there are several association categories between the main data and the several related data; integrating the multiple groups of data to obtain target data resources.

[0006] Optionally, the execution steps of the basic graph construction unit include: obtaining multiple relationship attributes of the main data between multiple groups of data; allocating and configuring multiple basic relationship distances based on the multiple relationship attributes; connecting multiple groups of data and the main data and several related data in each group of data based on the multiple groups of data, multiple basic relationship distances, multiple relationship attributes and several related categories to obtain a basic knowledge graph, wherein multiple relationship attributes and multiple basic relationship distances are used to connect the main data in the multiple groups of data, and several related categories are used to connect the main data and several related data in each group of data.

[0007] Optionally, the execution steps of the basic graph construction unit include: collecting a set of sample relationship attributes based on the data resource entry records within the historical time; calculating the ratio of the average occurrence frequency of the same sample relationship attribute to the occurrence frequency of each sample relationship attribute to obtain multiple main distance correction factors; obtaining a preset relationship distance; using multiple main distance correction factors to correct and calculate the preset relationship distance respectively to obtain multiple basic relationship distances.

[0008] Optionally, the execution steps of the relationship correction optimization unit include: obtaining multiple data groups connected in the basic knowledge graph, each data group including two groups of data; analyzing the similarity between the two groups of data in each data group, correcting the basic relationship distance, and obtaining multiple relationship distances; using the multiple relationship distances, adjusting the basic knowledge graph to obtain a knowledge graph.

[0009] Optionally, the execution steps of the relationship correction optimization unit also include: calculating the similarity of several related data in two groups of data in each data group, and calculating the average to obtain multiple similarities; calculating the ratio of each similarity to the average of multiple similarities as multiple related distance correction factors; using the multiple related distance correction factors to correct and calculate multiple basic relationship distances respectively to obtain multiple relationship distances.

[0010] Optionally, the execution steps of the evaluation table entry unit include: statistically calculating the total length of the relationship distance within the knowledge graph as the knowledge graph scale information; calculating the resource fitness of the target data resource based on the knowledge graph scale information and the number of multiple groups of data; and using the resource fitness to identify the knowledge graph and incorporate it into the table for storage.

[0011] The beneficial effects of the present invention are: The data acquisition unit acquires the target data resource to be entered into a table. The target data resource includes multiple data sets, each of which includes master data and multiple associated data. Multiple association categories exist between the master data and the multiple associated data, providing a data foundation for subsequent association analysis and value assessment. The basic graph construction unit acquires multiple relationship attributes of the master data between the multiple data sets, configures multiple basic relationship distances, and constructs a basic knowledge graph. This establishes a preliminary data association framework based on the relationship attributes between the master data, and quantifies the degree of association between the data using the basic relationship distances. The relationship correction and optimization unit analyzes the similarity of multiple associated data between each data set, corrects the basic relationship distances to obtain multiple relationship distances, and adjusts the basic knowledge graph to obtain a knowledge graph. The relationship distances are further optimized through similarity analysis of the associated data, improving the accuracy of the knowledge graph and the precision of data value assessment. The entry evaluation unit analyzes the resource adaptability of the target data resource based on the scale of the knowledge graph and the data volume of the multiple data sets. The knowledge graph is then identified and stored in a table, thereby quantifying and assessing the overall value of the data resource, achieving intelligent identification and efficient entry management of data resources.

[0012] The above technical solution accurately identifies and quantifies the relationships between data resources. By building a knowledge graph and optimizing relationship distances, it enables precise assessment of the value of data resources. Furthermore, through resource fitness calculation and identification mechanisms, data resource management shifts from a traditional simple storage model to an intelligent management model based on value assessment, enabling accurate quantitative assessment of data resource value and efficient table entry management. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] Figure 1 A schematic diagram of the structure of the intelligent data resource entry system based on knowledge graph provided by the present invention; Figure 2 A schematic diagram of the process of obtaining a knowledge graph provided by the present invention.

[0014] In the accompanying drawings, the components represented by the reference numerals are as follows: Data acquisition unit 11, basic map construction unit 12, relationship correction and optimization unit 13, evaluation and table entry unit 14. DETAILED DESCRIPTION

[0015] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without making any creative efforts shall fall within the scope of protection of the present invention.

[0016] In the description of the present invention, the terms "first" and "second" are used for descriptive purposes only and should not be understood to indicate or imply relative importance or implicitly specify the number of the technical features indicated. Therefore, a feature specified as "first" or "second" may explicitly or implicitly include one or more of the specified features. In the description of the present invention, "plurality" means two or more, unless otherwise specifically defined.

[0017] In the description of the present invention, the term "for example" is used to mean "used as an example, illustration or illustration". Any embodiment of the present invention described as "for example" is not necessarily to be construed as being more preferred or advantageous than other embodiments. The following description is given to enable any person skilled in the art to implement and use the present invention. In the following description, details are listed for the purpose of explanation. It should be understood that a person of ordinary skill in the art can recognize that the present invention can be implemented without using these specific details. In other examples, well-known structures and processes are not elaborated in detail to avoid obscuring the description of the present invention with unnecessary details. Therefore, the present invention is not intended to be limited to the embodiments shown, but is consistent with the widest scope consistent with the principles and features disclosed herein.

[0018] like Figure 1 As shown, an embodiment of the present invention provides an intelligent data resource entry system based on a knowledge graph, including a data acquisition unit 11, a basic graph construction unit 12, a relationship correction and optimization unit 13, and an evaluation entry unit 14.

[0019] The data acquisition unit 11 is used to acquire target data resources to be entered into a table, wherein the target data resources include multiple groups of data, each group of data includes main data and several associated data, and there are several association categories between the main data and the several associated data.

[0020] Specifically, the data acquisition unit 11 is responsible for acquiring target data resources to be entered into the table. For example, the data acquisition unit 11 acquires target data resources from various data sources through a data acquisition interface. The target data resources refer to the original data set that needs to be structured stored and the knowledge graph constructed.

[0021] The target data resource contains multiple sets of data, each of which constitutes a relatively independent data unit. The data structure of the target data resource has specific hierarchical characteristics. Within each set of data, the data is further divided into two levels: primary data and associated data. Primary data refers to the core identification information of that set of data, which is unique and representative. Associated data refers to auxiliary information associated with the primary data, which is a number of pieces and contains specific business data related to the primary data. The relationship between primary data and associated data is not a simple inclusion relationship, but rather a semantic association relationship established through several association categories. Association categories refer to classification labels that describe the semantic relationship between primary data and associated data. These association categories lay the foundation for the edge relationships in the subsequent knowledge graph construction.

[0022] Taking the equipment operation and maintenance data management of a smart manufacturing enterprise as an example, the data acquisition unit 11 acquires equipment operation and maintenance data resources as target data resources from multiple data sources, such as the enterprise's equipment management system, maintenance record system, and sensor monitoring system. These data resources include multiple sets of equipment operation and maintenance data. The first set of data includes the master data for "CNC-A001 CNC machine tool," and the associated data includes fault codes "E001," "E005," and "E012," maintenance plans "replace spindle bearing," "adjust transmission system," and "calibrate coordinate system," fault parameters "temperature 75°C," "vibration 0.3mm," and "speed deviation 5%," and maintenance times of "4 hours," "2.5 hours," and "1 hour." The second set of data includes the master data for "CNC-A002 CNC machine tool," and the associated data includes fault codes "E003," "E001," and "E008," maintenance plans "adjust tool parameters," "replace spindle bearing," and "clean cooling system," fault parameters "vibration 0.5mm," "temperature 78°C," and "abnormal pressure," and maintenance times of "2 hours," "4 hours," and "3 hours." The main data of the third group of data is "CNC-B001 CNC machine tool", and the related data includes fault codes "E001" and "E015", maintenance plans "replace spindle bearing" and "replace hydraulic pump", fault parameters "temperature 78°C" and "insufficient hydraulic pressure", and maintenance time "3.5 hours" and "5 hours".

[0023] Semantic relationships are established between these master data and associated data through associated categories such as "fault information", "maintenance plan", "operating parameters", and "maintenance records".

[0024] Through this structured data organization method, the data acquisition unit 11 can provide standardized data input for subsequent knowledge graph construction, ensuring that the semantic integrity and relevance of the data are effectively maintained.

[0025] The basic graph construction unit 12 is used to obtain multiple relationship attributes of master data between multiple groups of data, configure multiple basic relationship distances, and construct a basic knowledge graph.

[0026] Specifically, the basic graph construction unit 12 is responsible for constructing an initial knowledge graph structure based on the target data resources provided by the data acquisition unit 11. Specifically, the basic graph construction unit 12 analyzes multiple relationship attributes of the master data between multiple groups of data, quantifies the strength of these relationship attributes, and constructs a basic knowledge graph with a spatial topological structure based on this.

[0027] First, the basic graph construction unit 12 obtains multiple relationship attributes of primary data between multiple data sets. Relationship attributes refer to semantic association features between primary data from different data sets. These relationship attributes reflect semantic relationships such as similarity, correlation, or hierarchy between the primary data, providing a semantic basis for subsequent distance calculations. Based on the acquired relationship attributes, the basic graph construction unit 12 further configures multiple basic relationship distances. Basic relationship distances are spatial distance quantization values ​​that represent the strength of associations between different primary data nodes in the knowledge graph. The configuration of basic relationship distances is based on the semantic weights and occurrence frequencies of relationship attributes, converting abstract semantic relationships into quantifiable distance parameters through a mathematical calculation model. Subsequently, based on the acquired relationship attributes and the configured basic relationship distances, the basic graph construction unit 12 constructs a basic knowledge graph. The basic knowledge graph is a graph structure with primary data as nodes, relationship attributes as edges, and basic relationship distances as edge weights. It also contains the association relationships between primary data and associated data within each data set.

[0028] In the above-mentioned application scenario of equipment operation and maintenance data management, the basic graph construction unit 12 analyzes multiple relationship attributes between each equipment model, including relationship attributes such as "same model", "same series", and "first-level similar models". Based on the frequency of occurrence and importance of these relationship attributes, the corresponding basic relationship distance is configured. For example, the basic relationship distance corresponding to the "same model" relationship attribute is small, the basic relationship distance corresponding to the "same series" relationship attribute is moderate, and the basic relationship distance corresponding to the "first-level similar models" relationship attribute is large. The basic knowledge graph finally constructed uses each equipment model as a node, the relationship attribute as the connecting edge, and the basic relationship distance as the edge weight. It also includes the connection relationship between each equipment node and its related data such as fault information, maintenance plan, operating parameters, and maintenance records.

[0029] Through this graph structure construction method, the basic graph construction unit 12 can convert the original multiple groups of data into a structured knowledge graph representation, providing a basic graph structure framework for subsequent relationship optimization and resource evaluation.

[0030] The relationship correction optimization unit 13 is used to analyze the similarity of several related data between each group of data, correct the basic relationship distance, obtain multiple relationship distances, and adjust the basic knowledge graph to obtain a knowledge graph.

[0031] Specifically, the relationship correction and optimization unit 13 is responsible for accurately correcting and optimizing the basic knowledge graph based on similarity analysis of related data. Specifically, the relationship correction and optimization unit 13 analyzes the similarity characteristics of the related data between multiple groups of data, identifies the diversity and value differences of the data, and dynamically adjusts the basic relationship distance accordingly, thereby obtaining a final knowledge graph that more accurately reflects the value of the data.

[0032] First, the relationship correction and optimization unit 13 analyzes the similarity of several pieces of associated data within each data set. The similarity of associated data refers to a quantitative measure of the degree of similarity between corresponding associated data within different data sets. This similarity reflects the data's repetitive and diverse characteristics. Higher similarity indicates closer associated data and lower data diversity, which impacts the overall data value assessment. Then, based on the analysis results of the associated data similarity, the relationship correction and optimization unit 13 performs correction processing on the basic relationship distance. This correction process incorporates the associated data similarity information into the adjustment of the basic relationship distance. When the associated data similarity is high, it indicates low data diversity and low data value, and the corresponding relationship distance is increased. When the associated data similarity is low, it indicates good data diversity and high data value, and the corresponding relationship distance is reduced, thereby forming multiple corrected relationship distances.

[0033] By applying the corrected multiple relationship distances, the relationship correction optimization unit 13 adjusts the structure of the basic knowledge graph to obtain a final knowledge graph. The final knowledge graph, while maintaining the original node and edge relationships, more accurately reflects the true association strength and value differences between data through the optimized relationship distances.

[0034] In the aforementioned equipment operation and maintenance data management scenario, the relationship correction and optimization unit 13 analyzes the similarity of associated data between each set of data. For example, the data sets for "CNC-A001 CNC Machine Tool," "CNC-A002 CNC Machine Tool," and "CNC-B001 CNC Machine Tool" all contain the fault code "E001," the maintenance plan "Replace Spindle Bearing," and temperature-related fault parameters ("Temperature 75°C," "Temperature 78°C," and "Temperature 78°C," respectively). This indicates that these associated data have high similarity and low data diversity. Based on this analysis, the relationship correction and optimization unit 13 increases the relationship distance between these equipment nodes to reflect the reduced data value.

[0035] Through this similarity-based dynamic correction mechanism, the relationship correction optimization unit 13 can transform the static basic knowledge graph into an optimized knowledge graph that accurately reflects the data value distribution, providing a more accurate graph structure foundation for subsequent resource evaluation.

[0036] The evaluation and table entry unit 14 is used to analyze and obtain the resource adaptability of the target data resource based on the scale of the knowledge graph and the data volume of the multiple groups of data, and to identify the knowledge graph and store it in a table.

[0037] Specifically, the evaluation table entry unit 14 is responsible for comprehensively evaluating the optimized knowledge graph, quantifying the value level of the target data resource, and implementing the identified storage of the knowledge graph. Specifically, the evaluation table entry unit 14 calculates a quantitative indicator reflecting the overall value of the data resource by comprehensively analyzing the structural characteristics and data scale characteristics of the knowledge graph. Based on this indicator, the knowledge graph is value-identified and structured.

[0038] First, the evaluation table entry unit 14 performs a comprehensive analysis based on the scale of the knowledge graph and the data volume of the multiple data sets. The scale of the knowledge graph refers to the total length of all relationship distances within the knowledge graph, reflecting the overall complexity of the graph structure and the density of connections between data sets. The data volume of the multiple data sets refers to the number of data sets contained in the target data resource, reflecting the scale of the data resource. The scale of the knowledge graph and the data volume of the multiple data sets together constitute the fundamental dimensions for assessing the value of data resources. Then, based on the analysis of the knowledge graph scale and data volume, the evaluation table entry unit 14 calculates the resource fitness of the target data resource. Resource fitness is a comprehensive quantitative indicator that measures the overall value of a data resource. It is calculated as the inverse of the ratio of the knowledge graph scale to the data volume. When the knowledge graph scale is small and the data volume is large, it indicates close connections between data, high data value, and high resource fitness. Conversely, when the knowledge graph scale is large and the data volume is small, it indicates loose connections between data, low data value, and low resource fitness.

[0039] After obtaining resource fitness, the evaluation and table entry unit 14 identifies the knowledge graph and stores it in a table. Identification involves assigning corresponding value-level labels to the knowledge graph based on the numerical range of resource fitness, thereby visually representing the value of the data resource. Table entry involves storing the identified knowledge graph in a structured data format in the system database to facilitate subsequent retrieval, analysis, and application.

[0040] In the application scenario of the above-mentioned equipment operation and maintenance data management, the evaluation table entry unit 14 counts the total length of all relationship distances in the optimized knowledge graph as the knowledge graph scale, counts the number of equipment operation and maintenance data groups contained as the data volume, and calculates the inverse of the ratio of the two to obtain resource fitness. For example, if the knowledge graph scale is 100 distance units and the data volume is 20 groups of equipment data, then the resource fitness is the inverse of 20 / 100=0.2, that is, 5.0. Based on this resource fitness value, the evaluation table entry unit 14 assigns a corresponding value level identifier to the knowledge graph and stores it in the equipment operation and maintenance data resource library to achieve quantitative management and visual display of the value of equipment operation and maintenance data.

[0041] Through this quantitative evaluation and identification storage mechanism, the evaluation entry unit 14 can transform complex knowledge graphs into data resources with clear value identification, providing a structured and visual data value reference for data asset management, thereby achieving accurate quantitative evaluation of data resource value and efficient entry management.

[0042] Furthermore, the data acquisition unit 11 executes the following steps: Collect multiple sets of data to be entered into a table, wherein each set of data includes master data and several related data, and there are several association categories between the master data and the several related data; Integrate multiple sets of data to obtain target data resources.

[0043] In one feasible implementation, the data acquisition unit 11 first collects raw data from multiple data sources through a data acquisition interface, forming multiple sets of data to be entered into a table. Each set of data has a unified data structure, consisting of two levels: master data and multiple sets of associated data. The master data serves as the core identifier of the data set and possesses unique characteristics; the multiple sets of associated data serve as auxiliary information related to the master data, providing detailed content. Semantic associations are established between the master data and the multiple sets of associated data through multiple association categories. These association categories define the specific association type and semantic meaning between the data.

[0044] The data acquisition unit 11 then performs unified formatting and structural integration on the multiple sets of collected data, eliminating format differences between different data sources and ensuring the consistency and integrity of the data structure. Through this integration process, the multiple sets of data to be entered into the table are organized into a unified data set, forming the target data resource. This target data resource maintains the independence and relevance of each set of data, providing standardized data input for the subsequent construction of the knowledge graph.

[0045] In the aforementioned application scenario of equipment operation and maintenance data management, data acquisition unit 11 first collects operation and maintenance data for each device from various data sources, such as the equipment management system, maintenance record system, and sensor monitoring system. This data is then collected into multiple sets of data, with "CNC-A001 CNC machine tool," "CNC-A002 CNC machine tool," and "CNC-B001 CNC machine tool" as primary data. Each set of data contains relevant data, such as the corresponding fault code, maintenance plan, fault parameters, and maintenance duration. Data acquisition unit 11 then unifies the format and structures of these multiple sets of data from different systems, ultimately forming a target data resource containing complete equipment operation and maintenance information.

[0046] Through the data acquisition process, the data acquisition unit 11 can effectively extract and integrate the required data resources from a complex multi-source data environment, providing a data basis for subsequent knowledge graph construction and data value assessment.

[0047] Furthermore, the execution steps of the basic graph construction unit 12 include: Get multiple relationship attributes of master data between multiple groups of data; According to multiple relationship attributes, multiple basic relationship distances are allocated and configured; According to multiple groups of data, multiple basic relationship distances, multiple relationship attributes and several associated categories, multiple groups of data and the main data and several associated data in each group of data are connected to obtain a basic knowledge graph, wherein multiple relationship attributes and multiple basic relationship distances are used to connect the main data in multiple groups of data, and several associated categories are used to connect the main data and several associated data in each group of data.

[0048] In a preferred embodiment, the basic graph construction unit 12 first uses semantic analysis to identify and extract various relationship features between master data from different groups of data. Relationship attributes are feature labels that describe the semantic connections between master data, reflecting multi-dimensional relationship characteristics such as similarity, hierarchy, and categorization. Multiple relationship attributes constitute a complete description system for the relationships between master data, providing a semantic foundation for subsequent distance quantification.

[0049] Then, the basic graph construction unit 12 assigns a corresponding distance value to each relationship attribute based on the semantic weight and statistical characteristics of each relationship attribute, thereby obtaining a plurality of basic relationship distances. The basic relationship distance is a quantitative representation of the relationship attribute in the knowledge graph space, and the size of the distance value reflects the strength of the association of the corresponding relationship attribute. The higher the frequency of occurrence of the relationship attribute, the more common the relationship attribute is, and the smaller the corresponding basic relationship distance is; the lower the frequency of occurrence of the relationship attribute, the less common the relationship attribute is, the lower its value is, and the larger the corresponding basic relationship distance is.

[0050] Subsequently, based on multiple data sets, multiple basic relationship distances, multiple relationship attributes, and several association categories, the multiple data sets, as well as the primary data and several associated data within each data set, are connected to obtain a basic knowledge graph. Specifically, the basic graph construction unit 12 constructs two levels of connection relationships: the first level uses multiple relationship attributes and multiple basic relationship distances to connect the primary data within the multiple data sets, forming an association network between the primary data nodes; the second level uses several association categories to connect the primary data and several associated data within each data set, forming a hierarchical structure of primary data and associated data. Through these two levels of connection, the basic graph construction unit 12 constructs a complete basic knowledge graph.

[0051] In the above-mentioned application scenario of equipment operation and maintenance data management, the basic graph construction unit 12 first analyzes the relationship attributes between the main data such as "CNC-A001 CNC machine tool", "CNC-A002 CNC machine tool", and "CNC-B001 CNC machine tool", and identifies multiple relationship attributes such as "same model", "same series", and "first-level similar model". Then, based on the importance and frequency of occurrence of these relationship attributes, a smaller basic relationship distance is assigned to the "same model" relationship attribute, a medium basic relationship distance is assigned to the "same series" relationship attribute, and a larger basic relationship distance is assigned to the "first-level similar model" relationship attribute. Finally, the basic graph construction unit 12 uses these relationship attributes and basic relationship distances to connect the equipment model nodes, and uses association categories such as "fault information", "maintenance plan", "operating parameters", and "maintenance records" to connect each equipment node with its corresponding associated data node, and finally constructs a basic knowledge graph containing the relationship between devices and the internal information of the equipment.

[0052] Through this hierarchical graph construction process, the basic graph construction unit 12 can convert the original multiple groups of data into a knowledge graph representation with clear structure and clear relationships, laying the graph structure foundation for subsequent relationship optimization processing.

[0053] Further, such as Figure 2 As shown, the execution steps of the basic graph construction unit 12 include: Collect sample relationship attribute sets based on the data resource entry records within the historical time; Calculate the ratio of the average occurrence frequency of the same sample relationship attribute to the occurrence frequency of each sample relationship attribute to obtain multiple principal distance correction factors; Get the preset relationship distance; A plurality of main distance correction factors are used to respectively correct and calculate the preset relationship distances to obtain a plurality of basic relationship distances.

[0054] In a preferred embodiment, the basic graph construction unit 12 accurately configures the basic relationship distances using a statistical analysis method based on historical data. Specifically, the basic graph construction unit 12 calculates the principal distance correction factor of the relationship attribute based on the statistical characteristics of the historical data resources and adjusts the preset relationship distance accordingly, thereby obtaining a basic relationship distance that reflects the true relationship strength.

[0055] First, the basic graph construction unit 12 accesses the historical database to retrieve and analyze the data resource entry records that have been stored in the table during the historical time period. Various relationship attribute information is extracted from these data resource entry records to form a sample relationship attribute set. The sample relationship attribute set contains all relationship attribute types that appeared during the historical operation process and their corresponding statistical data, providing a data foundation for subsequent frequency analysis and value assessment.

[0056] Next, the basic graph construction unit 12 counts the frequency of occurrence of each relationship attribute in the sample relationship attribute set and calculates the average frequency of occurrence of all relationship attributes of the same type. By calculating the ratio of the average frequency of occurrence to the frequency of occurrence of each specific relationship attribute, the principal distance correction factor for that relationship attribute is obtained. The principal distance correction factor reflects the scarcity and importance of a specific relationship attribute relative to the average level of similar attributes. A larger ratio indicates a more scarce relationship attribute and a lower data value.

[0057] Next, the basic map construction unit 12 reads the preset relationship distance from the configuration parameters. The preset relationship distance is the reference distance value set during initialization, which serves as the basic parameter for subsequent correction calculations. Subsequently, the basic map construction unit 12 uses the main distance correction factor corresponding to each relationship attribute as a correction coefficient, multiplies it by the preset relationship distance, and obtains the corrected basic relationship distance. When the main distance correction factor is high, it indicates that the relationship attribute is relatively scarce, the frequency of occurrence is low, the data value is low, and the corresponding basic relationship distance is relatively large; when the main distance correction factor is low, it indicates that the relationship attribute is relatively common, the frequency of occurrence is high, the data value is high, and the corresponding basic relationship distance is relatively small.

[0058] In the above-mentioned application scenario of equipment operation and maintenance data management, the basic graph construction unit 12 collects a set of sample relationship attributes from the historical equipment operation and maintenance data entry records, including relationship attributes such as "same model", "same series", "first-level similar model" and their historical occurrence times. For example, through statistical analysis, it is found that the "same model" relationship attribute has a higher frequency of occurrence, and the "first-level similar model" relationship attribute has a lower frequency of occurrence. By calculating the ratio of the average occurrence frequency to the occurrence frequency of each relationship attribute, it is found that the main distance correction factor of the "same model" relationship attribute is lower, and the main distance correction factor of the "first-level similar model" relationship attribute is higher. Finally, these main distance correction factors are used to correct the preset relationship distance, assigning a larger basic relationship distance to the "same model" relationship attribute, and assigning a smaller basic relationship distance to the "first-level similar model" relationship attribute.

[0059] Through distance configuration based on historical data statistics, the basic graph construction unit 12 can dynamically adjust the basic relationship distance according to the actual importance and scarcity of relationship attributes, ensuring that the distance parameters in the knowledge graph can accurately reflect the true relationship strength between data.

[0060] Furthermore, the execution steps of the relationship correction optimization unit 13 include: Acquire multiple data groups connected in the basic knowledge graph, each data group including two groups of data; Analyze the similarity between two groups of data in each data group, correct the basic relationship distance, and obtain multiple relationship distances; The multiple relationship distances are used to adjust the basic knowledge graph to obtain a knowledge graph.

[0061] In a preferred embodiment, the relationship correction and optimization unit 13 first traverses the graph to identify pairs of primary data nodes directly connected by relationship attributes in the basic knowledge graph. Each pair of connected primary data nodes and their corresponding data content are combined into a data group. A data group is the basic unit for similarity analysis. Each data group contains two complete sets of data, including the respective primary data and the corresponding number of associated data. In this way, the relationship correction and optimization unit 13 converts the connection relationships in the graph structure into a collection of data groups that can be used for similarity calculation.

[0062] Next, the relationship correction and optimization unit 13 analyzes the similarity of several associated data points between two data groups within each data group. By comparing the content, type, and value of the associated data in the two data groups, it calculates a similarity index for the associated data. Based on the similarity analysis results, the relationship correction and optimization unit 13 performs a correction calculation on the basic relationship distance corresponding to the data group. When the associated data similarity is high, it indicates poor data diversity and low data value, and the corrected relationship distance increases. When the associated data similarity is low, it indicates good data diversity and high data value, and the corrected relationship distance decreases. By processing all data groups, multiple corrected relationship distances are obtained.

[0063] Subsequently, the relationship correction and optimization unit 13 applies the corrected multiple relationship distances to the corresponding edges of the basic knowledge graph, replacing the original basic relationship distances. Through this distance update operation, the structure of the basic knowledge graph is optimized and adjusted to form the final knowledge graph. The final knowledge graph, while maintaining the original node and edge topology, more accurately reflects the true value relationship and association strength between data through the optimized relationship distances.

[0064] In the above-mentioned application scenario of equipment operation and maintenance data management, the relationship correction optimization unit 13 first identifies the connected equipment node pairs in the basic knowledge graph, such as ("CNC-A001 CNC machine tool", "CNC-A002 CNC machine tool"), ("CNC-A001 CNC machine tool", "CNC-B001 CNC machine tool") and other data groups. Then, the similarity of the associated data between the two groups of equipment data in each data group is analyzed, and it is found that "CNC-A001 CNC machine tool" and "CNC-A002 CNC machine tool" and "CNC-B001 CNC machine tool" all share the fault code "E001" and the maintenance plan "replace the spindle bearing", and the similarity is relatively high. Based on this similarity analysis, the relationship correction optimization unit 13 increases the relationship distance between these equipment nodes to reflect the reduction in their data value. Finally, the corrected relationship distance is applied to update the basic knowledge graph to obtain a knowledge graph that can accurately reflect the value distribution of equipment data.

[0065] Through a dynamic correction process based on similarity analysis, the relationship correction optimization unit 13 can transform the static basic knowledge graph into an optimized knowledge graph that accurately reflects the difference in data value, providing a more accurate graph structure foundation for subsequent resource value evaluation.

[0066] Furthermore, the execution steps of the relationship correction optimization unit 13 also include: Calculate the similarity of several related data in two groups of data in each data group, and calculate the mean to obtain multiple similarities; Calculate the ratio of each similarity to the mean of multiple similarities as multiple correlation distance correction factors; The multiple correlation distance correction factors are used to perform correction calculations on the multiple basic relationship distances respectively to obtain multiple relationship distances.

[0067] In a preferred embodiment, the relationship correction and optimization unit 13 achieves precise correction of the basic relationship distance through detailed similarity calculation and value evaluation. Specifically, the relationship correction and optimization unit 13 uses multi-level calculations to convert the similarity characteristics of the associated data into quantitative value indicators, and accurately adjusts the basic relationship distance accordingly.

[0068] First, for each data group, the relationship correction and optimization unit 13 calculates the similarity between each corresponding associated data point within the two data groups. For example, it compares the similarity of fault codes, repair solutions, and fault parameters between the two data groups. After the calculations are completed, the similarities of all associated data points within the data group are averaged to obtain the overall similarity of the data group. By performing this operation on all data groups, multiple similarities are obtained, representing the degree of similarity between the associated data points of different data groups. Next, the relationship correction and optimization unit 13 calculates the overall mean of the similarities of all data groups and then calculates the ratio of each data group's similarity to the overall mean to obtain multiple association distance correction factors. The association distance correction factor reflects the degree of deviation of the similarity of a particular data group from the overall similarity level. When the ratio is greater than 1, it indicates that the similarity of the data group is above average, the data diversity is poor, and the data value is low. When the ratio is less than 1, it indicates that the similarity of the data group is below average, the data diversity is good, and the data value is high.

[0069] Subsequently, the relationship correction optimization unit 13 uses the correlation distance correction factor corresponding to each data group as a correction coefficient and multiplies it by the base relationship distance corresponding to that data group to obtain the corresponding relationship distance. When the correlation distance correction factor is high, it indicates high data similarity and low diversity, and the base relationship distance is corrected to increase. When the correlation distance correction factor is low, it indicates low data similarity and high diversity, and the base relationship distance is corrected to decrease. Through this correction calculation, multiple relationship distances reflecting the actual value differences of the data are obtained.

[0070] In the application scenario of the above-mentioned equipment operation and maintenance data management, the relationship correction optimization unit 13 calculates the fault code similarity, maintenance plan similarity, fault parameter similarity, and maintenance time similarity for the data groups ("CNC-A001 CNC machine tool", "CNC-A002 CNC machine tool"), and then calculates the mean of these similarities to obtain the similarity of the data group. By performing the same operation on all equipment data groups, multiple similarities are obtained. The mean of all similarities is calculated, and the ratio of the similarity of each data group to the mean of all similarities is calculated to obtain the association distance correction factor of each data group. Afterwards, these association distance correction factors are used to correct the corresponding basic relationship distances, assigning larger relationship distances to equipment data groups with higher similarity, and assigning smaller relationship distances to equipment data groups with lower similarity.

[0071] Through similarity analysis and value calculation mechanism, the relationship correction optimization unit 13 can more accurately identify and quantify the value differences between data, achieve precise correction of basic relationship distances, and thus construct a knowledge graph that truly reflects the value distribution of data resources.

[0072] Furthermore, the steps of executing the evaluation table entry unit 14 include: Statistically calculating the total length of the relationship distances within the knowledge graph as knowledge graph scale information; Calculate the resource adaptability of the target data resource based on the knowledge graph scale information and the number of multiple groups of data; The resource adaptability is used to identify the knowledge graph and incorporate it into table storage.

[0073] In a preferred embodiment, the evaluation table entry unit 14 first accesses all edge connections in the knowledge graph through graph traversal, obtains the relationship distance value corresponding to each edge, and performs cumulative summation to obtain the total length of the relationship distance. The total length of the relationship distance reflects the overall structural complexity of the knowledge graph and the density of the associations between data, serving as the scale information of the knowledge graph. When the total length of the relationship distance is small, it indicates that the similarity between the data is low, the diversity is good, and the data value is high; when the total length of the relationship distance is large, it indicates that the similarity between the data is high, the diversity is poor, and the data value is low.

[0074] Subsequently, the evaluation table entry unit 14 counts the number of data groups contained in the target data resource, and then calculates the ratio of the knowledge graph scale information to the number of data groups. The inverse of this ratio is used as the resource fitness. Resource fitness is a comprehensive quantitative indicator that measures the overall value of a data resource. Its calculation formula is: Resource fitness = Number of data groups / Knowledge graph scale information. A smaller ratio indicates a smaller total length of relationship distances and a larger amount of data, with high data value and sufficient quantity, thus increasing resource fitness. A larger ratio indicates a larger total length of relationship distances and a smaller amount of data, with low data value and insufficient quantity, thus decreasing resource fitness.

[0075] Next, the evaluation and entry unit 14 assigns a corresponding value level identifier to the knowledge graph based on the calculated resource adaptability. Using preset value level thresholds, knowledge graphs within different resource adaptability ranges are labeled as high value, medium value, low value, and other different levels. After completing the value identification, the evaluation and entry unit 14 stores the identified knowledge graph in a structured data format in the database, achieving standardized entry management of the knowledge graph.

[0076] In the aforementioned application scenario of equipment operation and maintenance data management, the evaluation table entry unit 14 counts the sum of the relationship distances between all equipment nodes in the optimized knowledge graph. For example, the relationship distance between "CNC-A001 CNC machine tool" and "CNC-A002 CNC machine tool" is 8 units, the relationship distance between "CNC-A001 CNC machine tool" and "CNC-B001 CNC machine tool" is 5 units, and the relationship distance between "CNC-A002 CNC machine tool" and "CNC-B001 CNC machine tool" is 2 units. The cumulative total is 15 distance units, which serves as the knowledge graph scale information. At the same time, the number of device data groups included is counted as 3 groups. The ratio of the knowledge graph scale information to the number of data groups is calculated to be 15 / 3=5, and the resource adaptability is the inverse of this ratio, that is, 1 / 5=0.2. Based on the resource fitness value of 0.2, according to the preset value level threshold range (for example: resource fitness ≥ 0.5 is high value, 0.2 ≤ resource fitness < 0.5 is medium value, and resource fitness < 0.2 is low value), the evaluation entry unit 14 assigns a medium value level identifier to the knowledge graph and stores it in the data resource library to realize standardized entry management of the knowledge graph.

[0077] Through the quantitative evaluation and identification storage mechanism, the evaluation entry unit 14 can transform complex knowledge graphs into data resources with clear value identification, realize accurate quantitative evaluation of data resource value and efficient entry management, and provide structured and visual data value reference for data asset management.

[0078] It should be noted that, in the above embodiments, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant description of other embodiments.

[0079] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0080] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded computer, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0081] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0082] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 The steps for the function specified in one or more boxes.

[0083] Although preferred embodiments of the present invention have been described, additional changes and modifications to these embodiments may occur to those skilled in the art once the basic inventive concepts become known.

[0084] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the present invention and its equivalents, the present invention is also intended to include these modifications and variations.

Claims

1. The intelligent data resource entry system based on knowledge graph is characterized by: The system comprises: A data acquisition unit is used to acquire a target data resource to be entered into a table, wherein the target data resource includes multiple groups of data, each group of data includes primary data and a plurality of associated data, and there are a plurality of association categories between the primary data and the plurality of associated data; The basic graph construction unit is used to obtain multiple relationship attributes of master data between multiple groups of data, configure multiple basic relationship distances, and build a basic knowledge graph; A relationship correction and optimization unit is used to analyze the similarity of several related data between each group of data, correct the basic relationship distance, obtain multiple relationship distances, and adjust the basic knowledge graph to obtain a knowledge graph; The evaluation and table entry unit is used to analyze and obtain the resource adaptability of the target data resource based on the scale of the knowledge graph and the data volume of the multiple groups of data, and to identify the knowledge graph and incorporate it into the table for storage.

2. The intelligent data resource entry system based on knowledge graph according to claim 1 is characterized in that: The execution steps of the data acquisition unit include: Collect multiple sets of data to be entered into a table, wherein each set of data includes master data and several related data, and there are several association categories between the master data and the several related data; Integrate multiple sets of data to obtain target data resources.

3. The intelligent data resource entry system based on knowledge graph according to claim 1 is characterized in that: The execution steps of the basic graph construction unit include: Get multiple relationship attributes of master data between multiple groups of data; According to multiple relationship attributes, multiple basic relationship distances are allocated and configured; According to multiple groups of data, multiple basic relationship distances, multiple relationship attributes and several associated categories, multiple groups of data and the main data and several associated data in each group of data are connected to obtain a basic knowledge graph, wherein multiple relationship attributes and multiple basic relationship distances are used to connect the main data in multiple groups of data, and several associated categories are used to connect the main data and several associated data in each group of data.

4. The intelligent data resource entry system based on knowledge graph according to claim 1 is characterized in that: The execution steps of the basic graph construction unit include: Collect sample relationship attribute sets based on the data resource entry records within the historical time; Calculate the ratio of the average occurrence frequency of the same sample relationship attribute to the occurrence frequency of each sample relationship attribute to obtain multiple principal distance correction factors; Get the preset relationship distance; A plurality of main distance correction factors are used to respectively correct and calculate the preset relationship distances to obtain a plurality of basic relationship distances.

5. The intelligent data resource entry system based on knowledge graph according to claim 1 is characterized in that: The execution steps of the relationship correction optimization unit include: Acquire multiple data groups connected in the basic knowledge graph, each data group including two groups of data; Analyze the similarity between two groups of data in each data group, correct the basic relationship distance, and obtain multiple relationship distances; The multiple relationship distances are used to adjust the basic knowledge graph to obtain a knowledge graph.

6. The intelligent data resource entry system based on knowledge graph according to claim 5 is characterized in that: The execution steps of the relationship correction optimization unit also include: Calculate the similarity of several related data in two groups of data in each data group, and calculate the mean to obtain multiple similarities; Calculate the ratio of each similarity to the mean of multiple similarities as multiple correlation distance correction factors; The multiple correlation distance correction factors are used to perform correction calculations on the multiple basic relationship distances respectively to obtain multiple relationship distances.

7. The intelligent data resource entry system based on knowledge graph according to claim 1 is characterized in that: The execution steps of the evaluation table entry unit include: Statistically calculating the total length of the relationship distances within the knowledge graph as knowledge graph scale information; Calculate the resource adaptability of the target data resource based on the knowledge graph scale information and the number of multiple groups of data; The resource adaptability is used to identify the knowledge graph and incorporate it into table storage.

Citation Information

Patent Citations

  • Data management method, system and device based on knowledge graph and medium

    CN112685405A

  • Standard information management method and system based on knowledge graph

    CN119179789A

  • Knowledge graph optimization method and device suitable for network security situation awareness data

    CN120090833A

  • Knowledge data storage method, device, computer apparatus, and storage medium

    WO2020143326A1