Knowledge graph construction method based on data assets

Through the knowledge graph construction method based on data assets, the shortcomings of the existing technology in data preprocessing and knowledge inference are solved, automatic identification and cleaning of data, fusion of multimodal data, and interpretability of knowledge graphs are realized, and the quality and credibility of knowledge graphs are improved.

CN120106200AInactive Publication Date: 2025-06-06MAILEFENG (XIAMEN) E-COMMERCE CO LTD

Patent Information

Application Number
CN202510570328.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-06
Publication Date
2025-06-06
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The prior art has shortcomings in data preprocessing and knowledge inference, especially in the automatic identification and cleaning of multi-source heterogeneous data, the fusion of multi-modal data, the dynamic and real-time update of knowledge graphs, and the interpretability.

Method used

The knowledge graph construction method based on data assets is adopted, and different types of data assets are automatically identified and cleaned through adaptive data preprocessing technology, deep learning is used to fusion of multimodal data, and an inference engine based on semantic understanding is built, and the missing information in the knowledge graph is inferred through graph neural networks, and the update process is recorded using blockchain technology.

Benefits of technology

It realizes automatic identification and cleaning of data, improves data quality and diversity, enhances the semantic understanding and interpretability of the knowledge graph, and ensures the update transparency and security of the knowledge graph.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120106200A_ABST
    Figure CN120106200A_ABST
Patent Text Reader

Abstract

The invention discloses a knowledge graph construction method based on data assets, and relates to the technical field of artificial intelligence and knowledge management, and the method comprises the steps: collecting data assets, employing an adaptive data preprocessing technology, automatically recognizing and cleaning different types of data assets, and carrying out the fusion of multi-modal data through deep learning. Constructing an inference engine based on semantic understanding, analyzing data context by using a natural language processing technology, deducing an implicit relationship between data, and enhancing a relationship between entities in the atlas; the method comprises the following steps: deducing missing information in a knowledge graph through knowledge reasoning and a completion algorithm of a graph neural network, generating an interpretable knowledge graph, increasing the credibility of a user to the knowledge graph by recording a derivation process of a node and a relationship between nodes, and recording an updating process of the knowledge graph by utilizing a block chain technology. According to the method, the overall richness and accuracy of the knowledge graph are improved, the credibility and use convenience of the user to the knowledge graph are enhanced, and the information integrity and accuracy of knowledge are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence and knowledge management technology, and in particular to a method for constructing a knowledge graph based on data assets. Background Art

[0002] With the rapid development of information technology, the exponential growth of data resources has led to the widespread application of knowledge graphs. As a structured way to store and represent knowledge, knowledge graphs have gradually become a core component in fields such as information retrieval, artificial intelligence, and recommendation systems. Traditional knowledge graph construction methods often rely on artificial rules and prior knowledge, which limits their adaptability and scalability. With the rapid development of deep learning and natural language processing technologies, automated knowledge graph construction methods have emerged. These methods can extract implicit knowledge through automatic analysis of massive data, and then construct rich knowledge graphs, thereby improving the level and depth of knowledge expression. However, although the current knowledge graph construction technology has made certain progress, it still faces many challenges such as data heterogeneity, data noise, missing information, and interpretability.

[0003] Existing related technologies have certain deficiencies in data preprocessing and knowledge reasoning. For example, in the data preprocessing stage, the automatic identification and cleaning of multi-source heterogeneous data still rely heavily on manual rules, lacking flexibility and intelligence. In addition, in the process of multimodal data fusion, existing methods cannot fully consider the semantic relationship between data, which often leads to the one-sidedness and incompleteness of information in the knowledge graph. In terms of knowledge reasoning, traditional methods are usually limited to static relationship deduction based on graphs, lacking the ability to update dynamically and in real time. Through the topological characteristics of graph neural networks, the data reasoning and completion of knowledge graphs can be further optimized. However, these methods often ignore the problem of interpretability, resulting in insufficient trust in knowledge graphs by users during use. Therefore, in response to the above problems, our invention proposes a knowledge graph construction method based on data assets. Summary of the invention

[0004] In view of the above-mentioned existing problems, the present invention solves the following technical problems: how to effectively collect, identify and clean various types of data assets, and integrate them into a unified knowledge framework to ensure the quality and availability of the data; how to analyze the contextual relationship of the data through advanced reasoning engines and natural language processing technology to derive the implicit relationship and knowledge between the data; how to reason and complete the missing information in the knowledge graph through graph neural network, and make the generated knowledge graph interpretable to enhance the user's trust in the knowledge graph; how to use blockchain technology to record the update process of the knowledge graph to ensure the traceability and non-tamperability of each change, thereby improving the transparency and security of data updates.

[0005] In order to solve the above technical problems, a knowledge graph construction method based on data assets is proposed, including:

[0006] Collect data assets, use adaptive data preprocessing technology to automatically identify and clean different types of data assets, and fuse multimodal data through deep learning; build a reasoning engine based on semantic understanding, use natural language processing technology to analyze data context, deduce implicit relationships between data, and enhance the relationship between entities in the knowledge graph; deduce missing information in the knowledge graph through the knowledge reasoning and completion algorithm of the graph neural network, generate an interpretable knowledge graph, increase user trust in the knowledge graph by recording the derivation process of nodes and relationships between nodes, and use blockchain technology to record the update process of the knowledge graph.

[0007] As a preferred solution of the method for constructing a knowledge graph based on data assets described in the present invention, the data assets include text data, image data, audio data, sensor and Internet of Things data, and structured data obtained through data crawlers and API calls.

[0008] As a preferred solution of the method for constructing a knowledge graph based on data assets described in the present invention, the automatic identification and cleaning of different types of data assets includes performing metadata analysis on the collected data assets, including file format, data format, data structure, extracting representative features, and formulating data source identification standards, including data type standards, content feature standards, and structural complexity standards:

[0009] The data type standard is to define preset types of different data assets according to the original format of the data assets and user requirements, and mark the received data as a potential error when the format does not match;

[0010] The content feature standard is to set a keyword density threshold and a language mode preset according to user needs. When the keyword density is greater than 3% and the set language mode is met, it is marked as high applicability;

[0011] The structural complexity standard is to judge the complexity of the data source based on the number of fields and the depth of the hierarchy of the statistical data assets. When the complexity index exceeds 3, it is marked as a high structure type.

[0012] Identify all potential errors for filtering, and prioritize the extraction of highly applicable and highly structured data assets using recursive depth-first search and use all identified data assets as graph nodes.

[0013] As a preferred solution of the method for constructing a knowledge graph based on data assets described in the present invention, the automatic identification and cleaning of different types of data assets also includes converting the extracted data assets into a unified UTF-8 encoding, using regular expressions to remove HTML tags, special characters, URLs and irrelevant information, using natural language processing tools for word segmentation, removing stop words in the text and performing stem extraction or word form restoration to obtain the final word or phrase as a graph node, converting the processed data assets into a feature vector representation through an improved Word2Vec method, fusing multimodal data through deep learning, and fusing different types of data assets: ;

[0014] in, is the feature vector of word or phrase j, is the fused feature vector, is the normalization factor, is the total number of different types of data assets, is the variable index, is the weight of each type of data asset, is a nonlinear transformation function used to process each type of data asset.

[0015] As a preferred solution of the method for constructing a knowledge graph based on data assets described in the present invention, the construction of an inference engine based on semantic understanding includes analyzing the contextual relationship in text data through natural language processing technology, and calculating the similarity between contexts based on the fused feature vectors: ;

[0016] in, is the context similarity value, ∈ is the smoothing parameter to prevent division by zero error; is the relationship vector between the current word or phrase and the remaining words in the context;

[0017] when When it is greater than 0.75, it indicates that there is similarity between the contexts of the two words, and an edge is added between the current two nodes, and a relationship type label is added to the current edge;

[0018] Create a reasoning engine based on semantic understanding, and build a semantic model in combination with an improved graph neural network to establish a relationship mapping between entities ,in, is the relational mapping function, is the mapped relationship vector, which derives the implicit relationship and enhances the relationship definition between entities in the knowledge graph: ;

[0019] in, is the enhanced relationship representation, is the gain factor, used to enhance the definition of the relationship; The relational outputs are normalized so that the results can be interpreted as probability distributions.

[0020] As a preferred solution of the method for constructing a knowledge graph based on data assets described in the present invention, the generating of an interpretable knowledge graph includes deducing the missing information in the knowledge graph through the knowledge reasoning and completion algorithm of the graph neural network, setting all nodes in the knowledge graph as a set U, and edges as a set B, and using an implicit feature representation matrix H, wherein each node is represented as , and use the message passing mechanism of graph neural networks to establish relationships between nodes: ;

[0021] in, For the Node in layer The feature representation of represents the new feature of the node after a message propagation in the graph neural network; is the activation function, For Node The set of neighbor nodes, that is, the set of nodes directly connected to node i; For the The weight matrix of the layer is responsible for converting the features of neighboring nodes into feature representations of the current layer; Neighbor node In the The feature representation of the layer, For the The bias term of the layer;

[0022] Infer missing nodes and edges through feature relationships: ;

[0023] in, The new feature vector to be completed is expected, through the derived new information; A function computed over the current feature representation matrix H that represents how features are aggregated and affect the inference results.

[0024] As a preferred solution of the method for constructing a knowledge graph based on data assets described in the present invention, wherein: the generating of an interpretable knowledge graph further includes: using a graphical tool to present the nodes and edges of the knowledge graph in a graphical form, adding user-readable semantic labels to each node and edge, and recording the missing nodes and edges derived by completion. After each derivation or completion, the system automatically generates a record including a timestamp, operation type, and the status of the nodes and edges before and after the change;

[0025] Records will be stored in a structured form in SON format. Each record will be packaged into a block, and hashes will be generated using the encryption characteristics of the blockchain. Through the consensus mechanism, all nodes in the network will share the updated version, forming an unalterable chain storage.

[0026] When a user queries the knowledge graph, the update history of each node and edge is confirmed by querying the blockchain to generate an interpretable knowledge graph.

[0027] Another purpose of the present invention is to provide a knowledge graph construction system based on data assets, which aims to effectively improve the construction efficiency and information quality of the knowledge graph and strengthen the user's trust in it. This method combines blockchain technology to achieve transparency and immutability of the knowledge graph update process, providing new ideas and solutions for the management and application of digital assets.

[0028] As a preferred solution of the knowledge graph construction system based on data assets described in the present invention, it includes a data processing module, a graph derivation module, and a graph generation module;

[0029] The data processing module collects data assets, uses adaptive data preprocessing technology, automatically identifies and cleans different types of data assets, and integrates multimodal data through deep learning;

[0030] The graph derivation module builds a reasoning engine based on semantic understanding, uses natural language processing technology to analyze data context, derives implicit relationships between data, and enhances the relationships between entities in the knowledge graph;

[0031] The graph generation module derives the missing information in the knowledge graph through the knowledge reasoning and completion algorithm of the graph neural network, generates an interpretable knowledge graph, increases the user's trust in the knowledge graph by recording the derivation process of nodes and relationships between nodes, and uses blockchain technology to record the update process of the knowledge graph.

[0032] Beneficial effects of the present invention: The present invention realizes automatic identification and cleaning of data, significantly improves the quality of data, and lays a solid foundation for the subsequent construction of knowledge graphs. Specifically, through multimodal data fusion, the diversity of data sources is broadened, so that the knowledge graph can more comprehensively reflect the information of the real world, thereby improving the practicality and accuracy of the graph.

[0033] Building a reasoning engine based on semantic understanding, analyzing the data context through natural language processing technology, not only derives the implicit relationship between data, but also enhances the relationship between entities in the knowledge graph. The beneficial effect of this step is that it improves the semantic understanding ability of the knowledge graph, allowing the graph to more accurately reflect the complex relationship between entities and provide users with more in-depth knowledge services.

[0034] Through the knowledge reasoning and completion algorithm of the graph neural network, the missing information in the knowledge graph is derived, and the update process is recorded using blockchain technology. The beneficial effects of this step are reflected in: on the one hand, an interpretable knowledge graph is generated, which increases the user's trust in the knowledge graph; on the other hand, the application of blockchain technology ensures the transparency and non-tamperability of the knowledge graph update process, providing users with reliable knowledge traceability guarantees.

[0035] To sum up, not only the comprehensive collection and efficient processing of data assets are achieved, but also a knowledge graph with rich semantics, accurate relationships and strong interpretability is constructed, ultimately achieving the beneficial effect of improving the quality of knowledge services and user trust. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0037] Figure 1 An overall flow chart of a method for constructing a knowledge graph based on data assets provided for one embodiment of the present invention.

[0038] Figure 2 A system solution module diagram of a knowledge graph construction system based on data assets provided for one embodiment of the present invention. DETAILED DESCRIPTION

[0039] In order to make the above-mentioned purposes, features and advantages of the present invention more obvious and easy to understand, the specific implementation methods of the present invention are described in detail below in conjunction with the drawings of the specification. Obviously, the examples described are part of the embodiments of the present invention, but not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary persons in the art without creative work should fall within the scope of protection of the present invention.

[0040] In the following description, many specific details are set forth to facilitate a full understanding of the present invention, but the present invention may also be implemented in other ways different from those described herein, and those skilled in the art may make similar generalizations without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.

[0041] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure or characteristic that may be included in at least one implementation of the present invention. The term "in one embodiment" that appears in different places in this specification does not necessarily refer to the same embodiment, nor is it an embodiment that is mutually exclusive with other embodiments, either individually or selectively.

[0042] The present invention is described in detail with reference to schematic diagrams. When describing the embodiments of the present invention, for the sake of convenience, the cross-sectional diagrams showing the device structure will not be partially enlarged according to the general scale, and the schematic diagrams are only examples, which should not limit the scope of protection of the present invention. In addition, in actual production, the three-dimensional dimensions of length, width and depth should be included.

[0043] At the same time, in the description of the present invention, it should be noted that the directions or positional relationships indicated by the terms "upper, lower, inner and outer" are based on the directions or positional relationships shown in the drawings, which are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific direction, be constructed and operated in a specific direction, and therefore cannot be understood as limiting the present invention. In addition, the terms "first, second or third" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance.

[0044] In the present invention, unless otherwise clearly specified and limited, the terms "install, connect, connect" should be understood in a broad sense, for example: it can be a fixed connection, a detachable connection or an integral connection; it can also be a mechanical connection, an electrical connection or a direct connection, or it can be indirectly connected through an intermediate medium, or it can be the internal communication of two components. For ordinary technicians in this field, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.

[0045] Example 1, reference Figure 1 , which is the first embodiment of the present invention, and provides a method for constructing a knowledge graph based on data assets, including:

[0046] S1: Collect data assets, use adaptive data preprocessing technology, automatically identify and clean different types of data assets, and fuse multimodal data through deep learning.

[0047] Furthermore, data crawlers and API calls are used to obtain text data, image data, audio data, sensor and IoT data, and structured data;

[0048] The text data includes, but is not limited to, scientific papers and patent documents containing technical background, research results and innovations; user opinions obtained from social media, product evaluation websites, etc. to help understand user needs and market trends; news reports and blog articles covering industry dynamics, technology trends and market analysis.

[0049] The image data includes, but is not limited to, drawings and design drawings including product designs, engineering drawings, prototype drawings, etc., which can intuitively demonstrate technical implementation; brand and product images help identify and classify the visual features of different products or services and are widely used in market analysis.

[0050] The audio data includes, but is not limited to, recordings of product introductions and customer interviews, from which important information and user feedback can be extracted through natural language processing technology; recordings of lectures and seminars, from which knowledge and insights can be extracted from speeches by industry experts;

[0051] The sensor and IoT data include, but are not limited to, device monitoring data including real-time data collected by sensors, such as temperature, humidity, location, etc., which can analyze the usage status and performance of the product; user behavior data including data collected based on smart devices and applications, such as usage time, function access frequency, etc.

[0052] The structured data includes, but is not limited to, databases and spreadsheets including market research data, sales data, financial data, etc., to facilitate quantitative analysis and trend forecasting; industry standards and Benchmark data to help build an industry knowledge graph and provide comparison with the market average.

[0053] It should also be noted that metadata analysis is performed on the collected data assets, including file formats, data formats (such as JSON, XML), and data structures (such as tables and image pixels), representative features are extracted, and data source identification standards are formulated, including data type standards, content feature standards, and structural complexity standards: the data type standards are based on the original format of the data assets, and the preset types of different data assets are defined according to user needs. When the received data format does not match, it is marked as a potential error;

[0054] The content feature standard is to set a keyword density threshold and a language pattern (such as common phrases or sentence structures) preset according to user needs. When the keyword density is greater than 3% and the set language pattern is satisfied, it is marked as highly applicable;

[0055] Specifically, in image data, by detecting the contour complexity and the proportion of the main color, if both exceed 85%, it is determined as a high-quality image with high applicability.

[0056] The structural complexity standard is to count the number of fields and the hierarchical depth of the data asset to judge the complexity of the data source: ;

[0057] Among them, is the number of fields, is the hierarchical depth, is the complexity index;

[0058] When the complexity index exceeds 3, it is marked as highly structural; identify and filter out all potential errors, and at the same time, for data assets with high applicability and high structure, use recursive depth-first search for priority extraction and all identified data assets.

[0059] It should be noted that the extracted data assets are converted to a unified UTF-8 encoding, use regular expressions to remove HTML tags, special characters, website addresses and irrelevant information, use natural language processing tools (such as NLTK or spaCy) for word segmentation, split the text into words or phrases, and at the same time perform part-of-speech tagging. After removing the stop words (such as meaningless words like "of", "is", "in", etc.) in the text, perform stemming (such as converting "running" to "run") or lemmatization (such as converting "better" to "good") to obtain the final words or phrases as graph nodes. Convert the processed data assets into feature vector representations through an improved Word2Vec method, and perform multi-modal data fusion through deep learning to fuse different types of data assets: ;

[0060] Among them, is the feature vector of word or phrase j, is the word index within the context window, ranging from t to t + k; is the word frequency of word or phrase j in the text, is the maximum value of all word frequencies within the context window, represents the semantic information of word or phrase j;

[0061] Perform multi-modal data fusion through deep learning to fuse different types of data assets: ;

[0062] in, is the fused feature vector, is the normalization factor, is the total number of different types of data assets, is the variable index, is the weight of each type of data asset, is a nonlinear transformation function used to process each type of data asset.

[0063] The present invention realizes automatic identification and cleaning of different types of data by collecting data assets and adopting adaptive data preprocessing technology. Multimodal data fusion is carried out through deep learning, and this process ensures the diversity and integrity of data. Its role is to improve data quality, enhance the accuracy of subsequent analysis, and ultimately achieve a more accurate and detailed knowledge graph construction effect. By using data crawlers, API calls and other methods to obtain text, audio, image and sensor data, the types of data sources are broadened, more comprehensive information capture is achieved, and the richness of knowledge graphs is improved, which can better serve the needs of users in different fields. The data is unified into UTF-8 encoding, and regular expressions and natural language processing tools are used to deeply analyze the text, which improves the systematization and standardization of data processing. At the same time, the effective fusion of multimodal data is achieved through deep learning, which provides more detailed and rich semantic relationships for the knowledge graph, enhancing the comprehensiveness and practicality of the graph.

[0064] S2: Build an inference engine based on semantic understanding, use natural language processing technology to analyze data context, deduce implicit relationships between data, and enhance the relationships between entities in the knowledge graph.

[0065] Furthermore, we use natural language processing technology to analyze the contextual relationship in the text data, and calculate the similarity between the contexts based on the fused feature vectors: ;

[0066] in, is the context similarity value, ∈ is the smoothing parameter to prevent division by zero error; is the relationship vector between the current word or phrase and the remaining words in the context;

[0067] when When it is greater than 0.75, it indicates that there is similarity between the contexts of the two words, and an edge is added between the current two nodes, and a relationship type label is added to the current edge;

[0068] It should also be noted that the creation of a reasoning engine based on semantic understanding, combined with an improved graph neural network to build a semantic model and establish a relationship mapping between entities ,in, is the relational mapping function, is the mapped relationship vector, which derives the implicit relationship and enhances the relationship definition between entities in the knowledge graph: ;

[0069] in, is the enhanced relationship representation, is the gain factor, used to enhance the definition of the relationship; The relational outputs are normalized so that the results can be interpreted as probability distributions.

[0070] Detailed metadata analysis includes the development of data source identification standards, which enables the effective identification and filtering of potential erroneous data during the data cleaning process, thereby ensuring the high quality of basic data. Recursive depth-first search is used to prioritize the extraction of highly applicable and highly structured data, which can quickly build more accurate graph nodes and improve the efficiency and accuracy of graph construction. Natural language processing is used to analyze contextual relationships, and the relationships between entities in the knowledge graph are enhanced based on similarity calculations. This provides users with more accurate relationship mapping and information extraction, improves the efficiency of knowledge retrieval and application, and ultimately achieves a scientific supplement to the content of the knowledge graph.

[0071] S3: Use the knowledge reasoning and completion algorithm of the graph neural network to derive the missing information in the knowledge graph and generate an interpretable knowledge graph. By recording the derivation process of nodes and the relationships between nodes, the user's trust in the knowledge graph is increased, and the update process of the knowledge graph is recorded using blockchain technology.

[0072] Furthermore, the missing information in the knowledge graph is derived through the knowledge reasoning and completion algorithm of the graph neural network. All nodes in the knowledge graph are set as set U, and edges are set as set B. The implicit feature representation matrix H is used, where each node is represented as , and use the message passing mechanism of graph neural networks to establish relationships between nodes: ;

[0073] in, For the Node in layer The feature representation of represents the new feature of the node after a message propagation in the graph neural network; is the activation function, For Node The set of neighbor nodes, that is, the set of nodes directly connected to node i; For the The weight matrix of the layer is responsible for converting the features of neighboring nodes into feature representations of the current layer; Neighbor node In the The feature representation of the layer, For the The bias term of the layer;

[0074] Infer missing nodes and edges through feature relationships: ;

[0075] in, The new feature vector to be completed is expected, through the derived new information; A function computed over the current feature representation matrix H that represents how features are aggregated and affect the inference results.

[0076] It should be noted that the nodes and edges of the knowledge graph are presented in a graphical form using a graphical tool, user-readable semantic labels are added to each node and edge, and missing nodes and edges derived by completion are recorded. After each derivation or completion, the system automatically generates a record containing the timestamp, operation type, and the status of the nodes and edges before and after the change;

[0077] Records will be stored in a structured form in SON format. Each record will be packaged into a block, and hashes will be generated using the encryption characteristics of the blockchain. Through the consensus mechanism, all nodes in the network will share the updated version, forming an unalterable chain storage.

[0078] When a user queries the knowledge graph, the update history of each node and edge is confirmed by querying the blockchain to generate an interpretable knowledge graph.

[0079] The derivation and completion of missing information through graph neural networks provides interpretability for knowledge graphs. Such completion not only improves the completeness of graph information, but also enhances users’ trust in the accuracy of content, facilitating better decision support.

[0080] Embodiment 2, the second embodiment of the present invention, is different from Embodiment 1 in that:

[0081] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions for a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, etc., which can store program codes.

[0082] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as an ordered list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by an instruction execution system, device or apparatus (such as a computer-based system, a system including a processor, or other system that can fetch instructions from an instruction execution system, device or apparatus and execute instructions), or in conjunction with such instruction execution systems, devices or apparatuses. For the purposes of this specification, "computer-readable medium" can be any device that can contain, store, communicate, propagate or transmit a program for use by an instruction execution system, device or apparatus, or in conjunction with such instruction execution systems, devices or apparatuses.

[0083] More specific examples of computer-readable media (a non-exhaustive list) include the following: an electrical connection with one or more wires (electronic device), a portable computer disk case (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable and programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disk read-only memory (CDROM). In addition, the computer-readable medium may even be a paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, deciphering or, if necessary, processing in another suitable manner, and then stored in a computer memory.

[0084] It should be understood that the various parts of the present invention can be implemented by hardware, software, firmware or a combination thereof. In the above-mentioned embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, it can be implemented by any one of the following technologies known in the art or their combination: a discrete logic circuit having a logic gate circuit for implementing a logic function for a data signal, a dedicated integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.

[0085] Example 3, reference Figure 2 , which is the third embodiment of the present invention, and provides a knowledge graph construction system based on data assets, including a data processing module, a graph derivation module, and a graph generation module;

[0086] The data processing module collects data assets, uses adaptive data preprocessing technology to automatically identify and clean different types of data assets, and integrates multimodal data through deep learning;

[0087] The graph inference module builds an inference engine based on semantic understanding, uses natural language processing technology to analyze data context, infers implicit relationships between data, and enhances the relationships between entities in the knowledge graph;

[0088] The graph generation module derives the missing information in the knowledge graph through the knowledge reasoning and completion algorithm of the graph neural network to generate an interpretable knowledge graph. By recording the derivation process of nodes and the relationships between nodes, it increases the user's trust in the knowledge graph and uses blockchain technology to record the update process of the knowledge graph.

[0089] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.

Claims

1. A method for constructing a knowledge graph based on data assets, characterized in that: include, Collect data assets, use adaptive data preprocessing technology to automatically identify and clean different types of data assets, and integrate multimodal data through deep learning; Build a reasoning engine based on semantic understanding, use natural language processing technology to analyze data context, deduce implicit relationships between data, and enhance the relationship between entities in the knowledge graph; The missing information in the knowledge graph is derived through the knowledge reasoning and completion algorithm of the graph neural network to generate an explainable knowledge graph. By recording the derivation process of nodes and the relationships between nodes, the user's trust in the knowledge graph is increased, and the blockchain technology is used to record the update process of the knowledge graph.

2. The method for constructing a knowledge graph based on data assets according to claim 1, characterized in that: The data assets include text data, image data, audio data, sensor and IoT data, and structured data obtained through data crawlers and API calls.

3. The method for constructing a knowledge graph based on data assets according to claim 2, characterized in that: The automatic identification and cleaning of different types of data assets includes performing metadata analysis on the collected data assets, including file format, data format, and data structure, extracting representative features, and formulating data source identification standards, including data type standards, content feature standards, and structural complexity standards: The data type standard is to define preset types of different data assets according to the original format of the data assets and user needs, and mark the received data as a potential error when the format does not match; The content feature standard is to set a keyword density threshold and a language mode preset according to user needs. When the keyword density is greater than 3% and the set language mode is met, it is marked as high applicability; The structural complexity standard is to judge the complexity of the data source based on the number of fields and the depth of the hierarchy of the statistical data assets. When the complexity index exceeds 3, it is marked as a high structure type. Identify all potential errors for filtering, and prioritize the extraction of highly applicable and highly structured data assets using recursive depth-first search.

4. The method for constructing a knowledge graph based on data assets according to claim 3, characterized in that: The automatic identification and cleaning of different types of data assets also includes converting the extracted data assets into a unified UTF-8 encoding, using regular expressions to remove HTML tags, special characters, URLs and irrelevant information, using natural language processing tools to perform word segmentation, splitting the text into words or phrases, and performing part-of-speech tagging at the same time, removing stop words in the text and performing stem extraction or word form restoration to obtain the final word or phrase as a graph node, converting the processed data assets into a feature vector representation through an improved Word2Vec method, and fusing multimodal data through deep learning to fuse different types of data assets: ; in, is the feature vector of word or phrase j, is the fused feature vector, is the normalization factor, is the total number of different types of data assets, is the variable index, is the weight of each type of data asset, is a nonlinear transformation function used to process each type of data asset.

5. The method for constructing a knowledge graph based on data assets according to claim 4, characterized in that: The construction of the inference engine based on semantic understanding includes analyzing the contextual relationship in the text data through natural language processing technology, and calculating the similarity between the contexts based on the fused feature vectors: ; in, is the context similarity value, ∈ is the smoothing parameter to prevent division by zero error; is the relationship vector between the current word or phrase and the remaining words in the context; when When it is greater than 0.75, it indicates that there is similarity between the contexts of the two words, and an edge is added between the current two nodes, and a relationship type label is added to the current edge; Create a reasoning engine based on semantic understanding, and build a semantic model in combination with an improved graph neural network to establish a relationship mapping between entities ,in, is the relational mapping function, is the mapped relationship vector, which infers the implicit relationship and enhances the relationship definition between entities in the graph: ; in, is the enhanced relationship representation, is the gain factor, used to enhance the definition of the relationship; The relational outputs are normalized so that the results can be interpreted as probability distributions.

6. The method for constructing a knowledge graph based on data assets according to claim 5, characterized in that: The generating of the explainable knowledge graph includes deducing the missing information in the knowledge graph through the knowledge reasoning and completion algorithm of the graph neural network, setting all nodes in the knowledge graph as a set U, and edges as a set B, and using an implicit feature representation matrix H, where each node is represented as , and use the message passing mechanism of graph neural networks to establish relationships between nodes: ; in, For the Node in layer The feature representation of represents the new feature of the node after a message propagation in the graph neural network; is the activation function, For Node The set of neighbor nodes, that is, the set of nodes directly connected to node i; For the The weight matrix of the layer is responsible for converting the features of neighboring nodes into feature representations of the current layer; Neighbor node In the The feature representation of the layer, For the The bias term of the layer; Infer missing nodes and edges through feature relationships: ; in, The new feature vector to be completed is expected, through the derived new information; A function computed over the current feature representation matrix H that represents how features are aggregated and affect the inference results.

7. The method for constructing a knowledge graph based on data assets according to claim 6, characterized in that: Generating an interpretable knowledge graph also includes using a graphical tool to present the nodes and edges of the knowledge graph in a graphical form, adding user-readable semantic labels to each node and edge, and recording the missing nodes and edges derived by completion. After each derivation or completion, the system automatically generates a record including a timestamp, an operation type, and the states of the nodes and edges before and after the change. Records will be stored in a structured form in SON format. Each record will be packaged into a block, and hashes will be generated using the encryption characteristics of the blockchain. Through the consensus mechanism, all nodes in the network will share the updated version, forming an unalterable chain storage. When a user queries the knowledge graph, the update history of each node and edge is confirmed by querying the blockchain to generate an interpretable knowledge graph.

8. A system for constructing a knowledge graph based on data assets, applied to a method for constructing a knowledge graph based on data assets as claimed in any one of claims 1 to 7, characterized in that: It includes a data processing module, a graph derivation module, and a graph generation module; The data processing module collects data assets, uses adaptive data preprocessing technology, automatically identifies and cleans different types of data assets, and integrates multimodal data through deep learning; The graph derivation module builds a reasoning engine based on semantic understanding, uses natural language processing technology to analyze data context, derives implicit relationships between data, and enhances the relationships between entities in the knowledge graph; The graph generation module derives the missing information in the knowledge graph through the knowledge reasoning and completion algorithm of the graph neural network, generates an interpretable knowledge graph, increases the user's trust in the knowledge graph by recording the derivation process of nodes and relationships between nodes, and uses blockchain technology to record the update process of the knowledge graph.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method for constructing a knowledge graph based on data assets described in any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method for constructing a knowledge graph based on data assets described in any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Deep reading method and system based on knowledge graph

    CN118193748A

  • Railway engineering design knowledge base construction method based on cloud platform

    CN118886336A

  • Mirror image type industry linkage engine system and method based on knowledge graph

    CN119537864A

  • Recommending content using multimodal memory embeddings

    WO2024220281A1

Cited By

  • Knowledge graph construction method and device based on material reserve information

    CN120764653A

  • Multi-reactor nuclear power station knowledge management intelligent platform and method based on block chain technology

    CN121524291A