Industrial chain knowledge graph construction method and system based on deep learning

Through deep learning-based methods, the problems of low efficiency and insufficient accuracy of industrial chain knowledge graph construction in the existing technology are solved, and more efficient and accurate industrial chain knowledge graph construction is achieved, which is suitable for industrial chain analysis and decision-making support in different industries.

CN119940515AActive Publication Date: 2025-05-06GONGXIN HUMANISTIC (BEIJING) MANAGEMENT CONSULTING CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202411872464.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-18
Publication Date
2025-05-06
Estimated Expiration
2044-12-18

AI Technical Summary

Technical Problem

The existing technology is inefficient and has limited accuracy when building the knowledge graph of the industrial chain, and there are problems with overfitting and professional term understanding of steps such as data collection, processing and relationship extraction.

Method used

Using a deep learning-based method, the cloud data acquisition program is deployed to obtain data from multiple data sources, random forests and word embedding are used to combine deep learning models for data classification and screening, combined with ELECTRA model and BM-25F scores for entity word filtering, and GCNE model is used for entity recognition and relationship extraction, and finally build an industrial chain knowledge graph.

Benefits of technology

It improves the efficiency, accuracy and completeness of the construction of the industrial chain knowledge graph, can better assist industrial chain analysis, prediction and decision-making, and has good scalability, and is suitable for industrial chains related to different industries.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119940515A_ABST
    Figure CN119940515A_ABST
Patent Text Reader

Abstract

The invention discloses an industrial chain knowledge graph construction method and system based on deep learning, a storage medium and electronic equipment, and the method comprises the steps: deploying a cloud data collection program, obtaining industrial chain data from multiple data sources, and carrying out the preprocessing; classifying, screening and classifying the preprocessed data; an ELCTRA model is combined with a BM-25F score to filter entity words and mark the entity words; constructing a graph structure, and performing entity recognition and relation extraction by using a GCNE model; storing the obtained triple in a graph database to construct a knowledge graph; during user retrieval, according to the query requirement of the knowledge graph, a proper index is created, knowledge graph data is matched, and related data is pushed to the user, and the method improves the construction efficiency, accuracy and integrity of the industrial chain knowledge graph, can assist industrial chain analysis, prediction and decision making, has good expansibility, and is suitable for industrial chains of different industries.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of knowledge graph technology, and in particular to a method and system for constructing an industrial chain knowledge graph based on deep learning. Background Art

[0002] With the rapid development of information technology, knowledge graphs are increasingly used in various fields, especially in industrial chain analysis. The industrial chain graph is a visualization tool for structured data systems built using artificial intelligence algorithms and other technologies. It is used to describe the various links in the industry and their interrelationships. The industrial chain knowledge graph plays a huge role. In corporate strategic planning, it can help companies clarify their positioning and competitive situation, optimize supply chains and expand their businesses, and assist in industrial planning; in the field of academic research, it can provide analytical tools and research perspectives for industrial economics, regional economics, and technological innovation research.

[0003] The industrial chain knowledge map divides the industrial chain into different links or nodes. Taking the electronic industrial chain as an example, it can be divided into links or nodes such as chip design, wafer manufacturing, packaging and testing, component production, electronic product assembly, sales and service, and end users. Among them, chip design companies are responsible for the research and development of chip functions and architectures; wafer manufacturing companies convert designs into actual chip wafers; packaging and testing companies package and quality test chips; component manufacturers produce various components required for electronic equipment; electronic product assembly companies assemble chips and components into complete electronic products; sales and service links include product promotion, sales, and after-sales service; and end users are individuals or companies that use electronic products. These different links or nodes are interconnected and together constitute a complete industrial chain. Through the industrial chain map, the relationship and interaction between each link can be clearly seen. In this way, the industrial chain map can clearly show the relationship between the upstream and downstream of the industrial chain, as well as the degree of dependence and mutual influence between different links.

[0004] In the prior art, building an industrial chain knowledge graph involves multiple links, which requires a lot of computing resources and time. In addition, in terms of data collection, data processing, entity recognition and relationship extraction during the construction process, the accuracy of traditional methods is limited, and deep learning models also have problems with overfitting and professional terminology understanding. The traditional industrial chain knowledge graph is inefficient and error-prone, and cannot meet actual needs. The present invention aims to solve these problems and build a better industrial chain knowledge graph.

[0005] To this end, the present invention provides a method and system for constructing an industrial chain knowledge graph based on deep learning. Summary of the invention

[0006] In order to make up for the shortcomings of the existing technology, the present invention provides a method for constructing an industrial chain knowledge graph based on deep learning, which can improve the efficiency, accuracy and completeness of the construction of the industrial chain knowledge graph, assist in industrial chain analysis, prediction and decision-making, has good scalability, and is suitable for related industrial chains in different industries.

[0007] In order to achieve the above objectives, in a first aspect, an embodiment of the present invention provides a method for constructing an industrial chain knowledge graph based on deep learning, characterized in that it includes the following steps:

[0008] S1. Deploy a cloud data collection program, wherein the cloud data collection program is configured to obtain industrial chain data from multiple data sources and clean and pre-process the data;

[0009] S2. Classify and filter the preprocessed data. For structured data, use the random forest algorithm to classify and filter. For unstructured data, use word embedding combined with a deep learning model to classify and filter the stored data to obtain text data related to the industrial chain.

[0010] S3, filter the text data for entity words, use the ELECTRA model as a classifier, and combine the BM-25F score to filter out important entity words, remove non-word words, obtain entity words, and annotate them;

[0011] S4, construct graph structure representation, create entity nodes, create and initialize edges, and use GCNE model for entity recognition and relationship extraction;

[0012] S5. Based on the results of entity recognition and relationship extraction, the entity, attribute and relationship triples required for the knowledge graph are obtained, and the triples are stored in the graph database to construct the industrial chain knowledge graph;

[0013] S6. When users search for industrial chain information, create appropriate indexes based on the query requirements of the knowledge graph, match the knowledge graph data, and push relevant data to users.

[0014] In one embodiment, for step S1, rdd conversion is used for unstructured data in the collected data, DataFrame conversion is used for structured data, and after the data is aggregated using an aggregation function, the processed data is saved to cloud storage.

[0015] In one embodiment, for step S3, the BM-25F score of each field is calculated, the text embedding of each field is generated using the ELECTRA model, and other features that are helpful in identifying entity words are extracted; a deep learning model is constructed, the BM-25F score, text embedding and other features are used as input, the Transformer model is used for feature fusion and processing, and the probability of whether the output word is an important entity word is determined.

[0016] In one embodiment, for step S3, entity word labeling specifically includes: combining the evaluation of industry experts in the industrial chain, labeling the extracted entity words into six categories: company, expert, industry, product, technology, and application, and automatically labeling in BIOES format.

[0017] In one embodiment, for step S4, constructing a graph structure representation specifically includes identifying entities in the text based on word segmentation results and industry chain domain knowledge, creating a node for each entity, and assigning an initial feature vector representation to each entity node. The initial feature vector is created based on basic attribute information of the entity, and the attribute information is extracted from the text or obtained through other external data sources. According to the relationship description between entities in the text, edges connecting the entity nodes are created.

[0018] In one embodiment, step S4 specifically includes: inputting the constructed graph structure data into the GCNE model, determining the set of neighbor nodes at each node according to the structure of the graph, and for each neighbor node, calculating the message from the neighbor node to the current node, wherein the message is calculated based on the current feature vector of the neighbor node and the features of the edges connecting them, and then aggregating the messages from all neighbor nodes; updating the state of the current node using the aggregated neighbor node information and edge information; using a linear layer to map the updated node feature vector to an entity label space; constructing a conditional random field (CRF) model, whose input is the representation of the node in the entity label space output by the linear layer, using the CRF model to predict the entity label of each node, selecting all possible node representations from the graph structure, and splicing them together to form a representation of an entity pair; then mapping the representation of the entity pair to the relationship label space, and predicting the relationship label between the entity pairs through the Softmax function.

[0019] In a second aspect, the present invention provides an industrial chain knowledge graph construction system based on deep learning, characterized in that it includes:

[0020] Cloud data collection device, used to deploy cloud data collection programs, obtain industrial chain data from multiple data sources, and clean and pre-process the data;

[0021] Data classification and screening device, used to classify and screen the pre-processed data, where the random forest algorithm is used to classify and screen the structured data, and the word embedding combined with the deep learning model is used to classify and screen the stored data for unstructured data, so as to obtain text data related to the industrial chain;

[0022] The entity word processing device is used to filter entity words in text data, use the ELECTRA model as a classifier, and combine the BM-25F score to screen out important entity words, remove non-word words, obtain entity words, and mark them;

[0023] Graph structure construction and analysis device, used to construct graph structure representation, create entity nodes, and create and initialize edges, using GCNE (Graph Convolutional Networks with Edge Features) model for entity recognition (NER) and relationship extraction (RE);

[0024] A knowledge graph construction device is used to obtain the entity, attribute and relationship triples required for the knowledge graph based on the results of entity recognition and relationship extraction, and store the triples in a graph database, thereby constructing an industrial chain knowledge graph;

[0025] The data push device is used to match the knowledge graph data according to the keyword fields and the industry to which the user's company belongs when the user searches for industrial chain information, and push the most relevant data to the user.

[0026] In a third aspect, an embodiment of the present invention further provides an electronic device, comprising a processor and a memory, wherein the memory stores computer-executable instructions that can be executed by the processor, and the processor executes the computer-executable instructions to implement any one of the methods provided in the first aspect.

[0027] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium, wherein the computer-readable storage medium stores computer-executable instructions. When the computer-executable instructions are called and executed by a processor, the computer-executable instructions prompt the processor to implement any one of the methods provided in the first aspect.

[0028] In a fifth aspect, an embodiment of the present invention further provides a computer program product, comprising a computer program stored on a computer-readable storage medium, and when the computer program is executed by a processor, it implements any method provided in the first aspect.

[0029] The present invention proposes a method for constructing an industrial chain knowledge graph based on deep learning, including: deploying a cloud data acquisition program, obtaining industrial chain data from multiple data sources and preprocessing; classifying and screening the preprocessed data, using a random forest algorithm for structured data, and using word embedding combined with a deep learning model for unstructured data; filtering entity words and annotating them using the ELECTRA model combined with the BM-25F score; constructing a graph structure, using the GCNE model for entity recognition and relationship extraction; storing the resulting triples in a graph database to construct a knowledge graph; and matching and pushing data by keywords and the industry to which they belong when users search. This method improves the efficiency, accuracy and completeness of the construction of the industrial chain knowledge graph, can assist in industrial chain analysis, prediction and decision-making, has good scalability, and is suitable for related industrial chains in different industries.

[0030] The technical solution of the present invention is further described in detail below through the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] The above and other purposes, features and advantages of the present invention will become more apparent by describing the embodiments of the present invention in more detail in conjunction with the accompanying drawings. The accompanying drawings are used to provide a further understanding of the embodiments of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings, the same reference numerals generally represent the same components or steps.

[0032] Figure 1 It is a flowchart of a method provided by an exemplary embodiment of the present invention.

[0033] Figure 2 It is a schematic diagram of the structure of a device provided by an exemplary embodiment of the present invention. DETAILED DESCRIPTION

[0034] Below, the exemplary embodiments according to the present invention will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments of the present invention, and it should be understood that the present invention is not limited to the exemplary embodiments described here.

[0035] It should be noted that the relative arrangement of components and steps, the numerical expressions and numerical values ​​set forth in these embodiments do not limit the scope of the present invention unless specifically stated otherwise.

[0036] Those skilled in the art can understand that the terms "first" and "second" in the embodiments of the present invention are only used to distinguish different steps, devices or modules, etc., and neither represent any specific technical meaning nor indicate the necessary logical order between them.

[0037] It should also be understood that, in the embodiments of the present invention, “plurality” may refer to two or more than two, and “at least one” may refer to one, two or more than two.

[0038] It should also be understood that any component, data or structure mentioned in the embodiments of the present invention can generally be understood as one or more, unless explicitly limited or otherwise indicated in the context.

[0039] To facilitate understanding of this embodiment, firstly, a method for constructing an industrial chain knowledge graph based on deep learning disclosed in an embodiment of the present invention is described in detail, as shown in the attached figure. Figure 1 As shown, a method for constructing an industrial chain knowledge graph based on deep learning includes the following steps:

[0040] S1. Deploy a cloud data collection program, which is configured to obtain industrial chain data from multiple data sources, clean and pre-process the data, and save the processed data to cloud storage.

[0041] In the embodiment of the present application, before classifying the industrial chain, it is necessary to collect text data related to the industrial chain, including corporate reports, industry research reports, news articles, technical documents, etc. These text data should cover all aspects of the industrial chain, such as raw material supply, production and processing, sales and circulation, technology research and development, etc. The cloud data collection tool can be implemented based on the Scrapy framework of Python, for example.

[0042] In the above step S1, rdd conversion is used for unstructured data in the collected data, and DataFrame data structure is used for conversion. After the data is aggregated using aggregation functions (such as groupby, reduce, aggregate functions), the processed data is saved to cloud storage.

[0043] When processing the collected data, for unstructured data, first load it as RDD (Resilient Distributed Dataset), then perform conversion operations such as word splitting on text data, and then use aggregation functions for aggregation. For structured data, load it as DataFrame data structure, perform conversion operations (such as selecting columns, adding new columns, etc.), and then use aggregation functions for grouping and aggregation. After determining the path and format of cloud storage (such as the path of services such as Amazon S3, Google CloudStorage, and data formats such as CSV and Parquet), save the processed RDD data of unstructured data as a text file, and save the processed DataFrame of structured data as a file of a specified format (such as CSV) to cloud storage. In actual operation, adjustments and optimizations can be made according to data characteristics and business needs.

[0044] Furthermore, the data is preprocessed to remove noise information in the text, such as redundant punctuation marks, special characters, numbers that have little effect on entity recognition and relationship extraction, etc. For example, multiple consecutive spaces are replaced with a single space, and some irrelevant advertising information or web page tags are deleted.

[0045] Furthermore, the data is segmented: Use appropriate segmentation tools to segment the cleaned text. For Chinese text, you can use tools such as Jieba segmentation; for English text, you can use the segmenter in NLTK. At the same time, according to the characteristics of the professional terms of the industrial chain, build a custom dictionary, and segment some specific industry terms as a whole to avoid incorrect segmentation. For example, terms such as "industrial chain", "supply chain", and "raw material supplier" should be identified as a whole. Considering the characteristics of the industrial chain data text, some professional terms need to be specially processed. For example, if there are some abbreviations or compound words for a specific industry, they can be processed as a whole to avoid incorrect segmentation. You can build a custom dictionary, add these professional terms to the dictionary, use the segmentation tool to identify them, and save the processed data to the target location.

[0046] S2. Classify and filter the preprocessed data. For structured data, use the random forest algorithm to classify and filter. For unstructured data, use word embedding combined with a deep learning model to classify and filter the stored data to obtain text data related to the industrial chain.

[0047] In the embodiment of the present application, for structured data: select a suitable traditional machine learning algorithm, such as random forest or XGBoost, according to the data characteristics and task objectives, and divide the data set into a training set and a test set (for example, in a ratio of 7:3 or 8:2). Use the training set to train the selected model. For random forest, some key parameters need to be set, such as the number of trees, maximum depth, etc.; for XGBoost, parameters such as the learning rate, maximum depth of the tree, and minimum weight of leaf nodes need to be adjusted to optimize model performance. In the industrial chain data, different types of structured data can be accurately identified and distinguished, such as enterprise data of different scales, industries, and development stages, as well as product data of different quality levels and market positioning, and valuable structured data information can be extracted. After classification and screening by the random forest algorithm, the data quality and availability are improved, making the retained data more representative and accurate, providing a more reliable foundation for subsequent analysis and modeling. On the other hand, the random forest algorithm can automatically mine hidden patterns and relationships in structured data, such as discovering the relationship between enterprise characteristics and enterprise development, product characteristics and market demand, which helps to deeply understand the operation mechanism of the industrial chain. At the same time, by screening key data, it reduces the data dimension and complexity, reduces the data volume, reduces the difficulty of subsequent processing and analysis, and improves computing efficiency.

[0048] For unstructured data: for example, text data, use word embedding combined with deep learning models (such as CNN, LSTM) or pre-trained language models (such as BERT) to perform preliminary text classification and screening on the stored data. In this case, use word embedding technology such as Word2Vec to map words in the text to a low-dimensional vector space. When training the word embedding model, select appropriate parameters such as vector dimension, window size, minimum word frequency, etc. according to the scale and characteristics of the text data. Select an appropriate deep learning model combined with word embedding for text classification and screening according to task requirements. If you are processing shorter texts and focusing on local features, you can choose CNN (convolutional neural network); if you are processing longer texts and need to consider the order information of the text, you can choose LSTM (long short-term memory network). You can also choose a pre-trained language model such as BERT, which has a strong language understanding ability. Make full use of unstructured data. Unstructured text data contains rich industrial chain information. Through the above methods, effective classification and screening can be carried out to mine and utilize this information to supplement and improve the industrial chain knowledge map.

[0049] The industrial chain data set includes the following features: enterprise-related features, such as enterprise scale (for example, represented by the number of employees, asset size, etc.), years of establishment, industry category, geographical location, etc. Product-related features, such as product type, product quality level, market share, sales price range, etc. Technology-related features, such as technological advancement indicators (such as the number of patents, the proportion of R&D investment, etc.), technology application fields, etc. Supply relationship characteristics, such as the number of major suppliers, supply stability indicators (such as the number of supply interruptions, supply lead time fluctuations, etc.). Sales channel characteristics, such as the proportion of online sales, the number of offline dealers, the proportion of exports, etc. Classification prediction results, that is, predicting the category to which the new industrial chain data point belongs. For example, enterprises can be divided into different development stage categories (start-up, growth, maturity, decline), or products can be divided into high demand, medium demand, low demand categories, etc.

[0050] S3, filter the text data for entity words, use the ELECTRA model as a classifier, and combine the BM-25F score to filter out important entity words, remove non-word words, obtain entity words, and annotate them;

[0051] Among them, the BM-25F (Best Match 25F) is an information retrieval algorithm, which is mainly used for text retrieval and sorting tasks. When calculating the relevance between documents and query content, BM-25F will assign different weights according to the characteristics of different fields, and comprehensively consider factors such as word frequency, document length, and inverse document frequency. In the embodiment of the present application, for step S3, the BM-25F score of each field is calculated, and the ELECTRA (Efficiently Learning an Encoder that Classifies Token Replacements Accurately) model is used to generate text embedding for each field, and other features that help identify entity words are extracted; a deep learning model is constructed, and the BM-25F score, text embedding and other features are used as input, and the Transformer model is used for feature fusion and processing, and the output is the probability of whether the word is an important entity word. The algorithm is used to further screen out important entity words and reduce the influence of non-word vocabulary. BM-25F can more accurately evaluate the relevance of documents to specific queries, thereby playing an important role in tasks such as text classification and screening of important entity words.

[0052] Specifically, the optional method of the above screening is, for example, a pre-labeled dataset for the industrial chain, which contains the title, body and summary of the document, as well as the important entity words in each document; for example, data = [{"title":"Automobile Company A Develops New Technology","body":"Automobile Company A announced the development of a new technology to improve production efficiency.","summary":"New Technology Improves Production Efficiency","entities":["Automobile Company A","New Technology"]};

[0053] Calculate the BM-25F score for each field (title, text, abstract); use the ELECTRA model to generate text embeddings for each field, and extract other features that help identify entity words, such as word frequency, word position, etc. Build a deep learning model where the input layer: takes the BM-25F score, text embedding, and other features as input, the middle layer: uses the Transformer model for feature fusion and processing, and for the output layer: outputs the probability of whether the word is an important entity word. Through the above steps, combine the BM-25F score and the ELECTRA model to build a deep learning classifier for screening important entity words. Using the BM-25F's ability to score different fields and the ELECTRA model's powerful semantic understanding ability, the accuracy and efficiency of entity word screening can be effectively improved.

[0054] Further, for step S3, the entity word annotation specifically includes: combining the evaluation of industry experts in the industrial chain, annotating the extracted entity words into six categories: company, expert, industry, product, technology, and application. BIOES is a annotation method for named entity recognition (NER). BIOES respectively stands for Begin (entity start), Inside (entity inside), Outside (non-entity part), End (entity end), Single (single entity). The present invention also uses the BIOES format for automatic annotation. The text after word segmentation is tagged with parts of speech to better understand the text structure and semantics. An automatic annotation program can be written in a programming language (such as Python). Read the extracted entity word list and the category information to which they belong, wherein the category information can be predetermined by manual judgment or reference to expert opinions. Then, according to the BIOES annotation rules, each entity word is annotated, and the program loops through each word in the entity word and performs corresponding annotations according to its position and category.

[0055] Furthermore, we can also use the part-of-speech tagger in NLTK or some Chinese part-of-speech tagging tools (such as HanLP) to implement BIOES tagging using these toolkits. By properly configuring the toolkit, it can adapt to the specific requirements of entity word tagging in related industries, such as setting the correct category labels and tagging rules.

[0056] S4. Build a graph structure representation, create entity nodes, and create and initialize edges. Use the GCNE (Graph Convolutional Networks with Edge Features) model for entity recognition (NER) and relationship extraction (RE);

[0057] In an embodiment of the present application, step S4 specifically includes: inputting the constructed graph structure data into the GCNE model, at each node, determining its neighbor node set according to the structure of the graph, for each neighbor node, calculating the message transmitted from the neighbor node to the current node, wherein the message is calculated based on the current feature vector of the neighbor node and the features of the edges connecting them, and then aggregating the messages from all neighbor nodes; using the aggregated neighbor node information and edge information to update the state of the current node; using a linear layer to map the updated node feature vector to the entity label space; constructing a conditional random field (CRF) model, whose input is the representation of the node in the entity label space output by the linear layer, using the CRF model to predict the entity label of each node, selecting all possible node representations from the graph structure and concatenating them to form an entity pair representation; then mapping the entity pair representation to the relationship label space, and predicting the relationship label between the entity pairs through the Softmax function.

[0058] The following are the detailed implementation steps for using the GCNE model to perform entity recognition and relationship extraction on industrial chain text data to construct an industrial chain knowledge graph: S401. Construct a graph structure representation, identify entities in the text based on the word segmentation results and industrial chain domain knowledge, and create a node for each entity. Entities may include industrial chain-related enterprises, products, technologies, raw materials, markets, regulations, etc. For example, from the text "A metal processing enterprise uses advanced smelting technology to produce high-quality metal products, and the products are mainly sold to the domestic market", entities such as "metal processing enterprises", "smelting technology", "metal products", and "domestic market" can be identified, and corresponding nodes can be created. An initial feature vector representation is assigned to each entity node. The initial feature vector is based on some basic attribute information of the entity. For example, for industrial chain-related enterprise entities, it can include the encoding of information such as enterprise name, location, enterprise scale, establishment time, and industry; for industrial chain-related product entities, it can include the encoding of information such as product name, product category, product function, applicable field, quality standard, etc.; for industrial chain-related technology entities, it can include the encoding of information such as technology name, technology field, technology research and development time, and application effect. The above attribute information can be extracted from the text or obtained through other external data sources. According to the relationship description between entities in the text, create edges connecting entity nodes. For example, in the above text, there is an "adoption" relationship between "metal processing enterprises" and "smelting technology", a "production" relationship between "metal enterprises" and "metal products", and a "sale" relationship between "metal products" and "domestic market". Therefore, the corresponding edges can be created, and the types of edges can be marked as "adoption", "production", "sale", etc. The feature vector of the edge can include the encoding of information such as the strength, time span, direction, and legality of the relationship. In the initial stage, if this information is not clear in the text, some default values ​​can be assigned, or reasonable guesses can be made based on domain knowledge. For example, for the "production" relationship, where production is the core business relationship, the relationship strength can be set to a higher value, and the time span can be evaluated based on the production cycle of the enterprise or the update cycle of the product.

[0059] S402. Apply the GCNE model for entity recognition and relationship extraction, input the constructed graph structure data into the GCNE model, and at each node, determine the set of neighbor nodes according to the graph structure. For each neighbor node, calculate the message transmitted from the neighbor node to the current node. The message is calculated based on the current feature vector of the neighbor node and the features of the edge connecting them. For example, if the feature of the edge contains the strength information of the relationship, then the message calculation can multiply the feature vector of the neighbor node by the relationship strength weight. Then, aggregate the messages from all neighbor nodes. The aggregation operation can be performed in a variety of ways, such as summation, averaging, weighted summation, etc. For example, weighted summation can assign different weights according to the closeness of the connection between the neighbor node and the current node (such as the weight of the edge or other related metrics), so that more important neighbor nodes have a greater impact on the aggregation results. Use the aggregated neighbor node information and edge information to update the state of the current node. The update method can use the nonlinear function ReLU function to process the fused information, and then perform weighted summation with the original feature vector of the current node to obtain the updated node feature vector. Use a linear layer to map the updated node feature vector to the entity label space. For example, if the entity label space includes labels such as enterprises, products, technologies, and raw materials related to the industrial chain, the updated node feature vector is mapped to this space through the linear layer to obtain the representation of each node in the entity label space. For CRF prediction of entity labels, a conditional random field (CRF) model is constructed, and its input is the representation of the node in the entity label space output by the linear layer. The CRF model defines a transition probability matrix between adjacent labels, which reflects the possibility of transition between different entity labels. For example, an enterprise entity label is more likely to be followed by a cooperative enterprise entity label or a product entity label, and is less likely to be directly followed by a technical detail entity label. The entity label of each node is predicted using the CRF model. In the prediction process, not only the representation of the node itself in the entity label space is considered, but also the labels of adjacent nodes and the transition probability between them. By maximizing the label probability of the entire sequence (based on the energy function or probability calculation formula of CRF), the label assignment of the entire sequence is optimized, making the label assignment more reasonable and logical. The node representations are spliced ​​together in the graph structure to form the representation of entity pairs. The representation of the entity pair is then mapped to the relationship label space, and the relationship label between the entity pairs is predicted by the Softmax function. For example, for the pair of entities “metal processing enterprise” and “smelting technology”, their node representations are concatenated and mapped to the relationship label space, and the relationship between them is predicted to be an “adoption” relationship.

[0060] S5. Based on the results of entity recognition and relationship extraction, the entity, attribute and relationship triples required for the knowledge graph are obtained, and the triples are stored in the graph database to construct the industrial chain knowledge graph;

[0061] In an embodiment of the present application, entities and relationships are stored in a graph database, and entity and relationship information is stored in the graph database based on the results of entity recognition and relationship extraction. According to the aforementioned method, nodes are created in the graph database, and their attributes are set. For example, "metal processing enterprise" is stored as a node, and its attributes include the name of the enterprise related to the industrial chain, the region where it is located, the size of the enterprise, the time of establishment, the industry to which it belongs, etc.; the "adoption" relationship is stored as an edge, and its attributes include the relationship strength, time span, direction, etc. If relevant entities and relationships already exist in the graph database, they can be updated according to the new recognition and extraction results.

[0062] S6. When users search for industrial chain information, create appropriate indexes based on the query requirements of the knowledge graph, match the knowledge graph data, and push relevant data to users.

[0063] In the embodiment of the present application, a suitable visualization tool, such as Cytoscape, Gephi, etc., is selected to display the knowledge graph in a graphical manner. Through the visualization tool, the knowledge graph can be adjusted in layout, color and size of nodes and edges can be set to highlight important entities. For example, the enterprise node can be set to different sizes according to its size, and the relationship edge can be set to different colors according to its strength. Through the visualization display, the distribution of entities and relationships in the industrial chain can be more intuitively understood.

[0064] Provide convenience for further analysis. The present invention can provide corresponding application programs according to the application requirements of the industrial chain data text. For example, if the purpose is to provide cooperation suggestions for enterprises, an enterprise cooperation suggestion system can be provided to recommend potential cooperation partners to enterprises by querying the industrial chain-related enterprise entities and relationships in the knowledge graph. If the purpose is to analyze the structure and dynamic changes of the industrial chain, an industrial chain analysis system can be provided to monitor and analyze the dynamic changes of entities and relationships in the knowledge graph, and provide a report on the structure and dynamic changes of the industrial chain. Create appropriate indexes according to the query requirements of the knowledge graph. For example, if you often need to query enterprises in a certain area, you can create an index on the regional attributes of the enterprise node to improve query efficiency.

[0065] Second, as attached Figure 2 As shown, the embodiment of the present invention also provides an industrial chain knowledge graph construction system 200 based on deep learning, including: a cloud data acquisition device 201, which is used to deploy a cloud data acquisition program, obtain industrial chain data from multiple data sources, and clean and pre-process the data;

[0066] Data classification and screening device 202, used to classify and screen the pre-processed data, wherein for structured data, a random forest algorithm is used for classification and screening, and for unstructured data, a word embedding combined with a deep learning model is used to classify and screen the stored data to obtain text data related to the industrial chain;

[0067] The entity word processing device 203 is used to filter the text data for entity words, use the ELECTRA model as a classifier, and combine the BM-25F score to filter out important entity words, remove non-word words, obtain entity words, and mark them;

[0068] Graph structure construction and analysis device 204, used to construct graph structure representation, create entity nodes, and create and initialize edges, using GCNE (Graph Convolutional Networks with Edge Features) model for entity recognition (NER) and relationship extraction (RE);

[0069] The knowledge graph construction device 205 is used to obtain the entity, attribute and relationship triples required for the knowledge graph according to the results of entity recognition and relationship extraction, and store the triples in the graph database, so as to construct the industrial chain knowledge graph;

[0070] The data push device 206 is used to match the knowledge graph data according to the keyword field and the industry to which the user's company belongs when the user searches for industrial chain information, and push the most relevant data to the user.

[0071] In a third aspect, an embodiment of the present invention further provides an electronic device, comprising a processor and a memory, wherein the memory stores computer-executable instructions that can be executed by the processor, and the processor executes the computer-executable instructions to implement any one of the methods provided in the first aspect.

[0072] In a fourth aspect, the present invention provides a data processing system based on a knowledge graph, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program and implements the data processing method based on the knowledge graph as described above.

[0073] In addition to the above-mentioned methods and devices, the embodiments of this document may also be a computer program product, which includes computer program instructions, which, when executed by a processor, enable the processor to execute the steps of the methods according to various embodiments of this document described in the above-mentioned "Exemplary Methods" section of this specification.

[0074] The basic principles of the present disclosure are described above in conjunction with specific embodiments. However, it should be noted that the advantages, strengths, effects, etc. mentioned in the present disclosure are only examples and not limitations, and it cannot be considered that these advantages, strengths, effects, etc. are required by each embodiment of the present disclosure. In addition, the specific details disclosed above are only for the purpose of illustration and ease of understanding, and are not limitations. The above details do not limit the present disclosure to the necessity of adopting the above specific details to be implemented.

[0075] Each embodiment in this specification is described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the embodiments can be referred to each other. For the system embodiment, since it basically corresponds to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.

[0076] The block diagrams of the devices, apparatuses, equipment, and systems involved in this disclosure are only illustrative examples and are not intended to require or imply that they must be connected, arranged, and configured in the manner shown in the block diagrams. As will be appreciated by those skilled in the art, these devices, apparatuses, equipment, and systems can be connected, arranged, and configured in any manner. Words such as "including," "comprising," "having," and the like are open words, referring to "including but not limited to," and can be used interchangeably therewith. The words "or" and "and" used herein refer to the words "and / or," and can be used interchangeably therewith, unless the context clearly indicates otherwise. The word "such as" used herein refers to the phrase "such as but not limited to," and can be used interchangeably therewith.

[0077] The method and apparatus of the present disclosure may be implemented in many ways. For example, the method and apparatus of the present disclosure may be implemented by software, hardware, firmware, or any combination of software, hardware, and firmware. The above order of steps for the method is for illustration only, and the steps of the method of the present disclosure are not limited to the order specifically described above, unless otherwise specifically stated. In addition, in some embodiments, the present disclosure may also be implemented as a program recorded in a recording medium, which includes machine-readable instructions for implementing the method according to the present disclosure. Therefore, the present disclosure also covers a recording medium storing a program for executing the method according to the present disclosure.

[0078] It should also be noted that in the apparatus, equipment and method of the present disclosure, each component or each step can be decomposed and / or recombined. These decompositions and / or recombinations should be regarded as equivalent schemes of the present disclosure. The above description of the disclosed aspects is provided to enable any technician in the field to make or use the present disclosure. Various modifications to these aspects are very obvious to those skilled in the art, and the general principles defined herein can be applied to other aspects without departing from the scope of the present disclosure. Therefore, the present disclosure is not intended to be limited to the aspects shown here, but to the widest scope consistent with the principles and novel features disclosed herein.

[0079] The above description has been given for the purpose of illustration and description. In addition, this description is not intended to limit the embodiments of the present disclosure to the forms disclosed herein. Although multiple example aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, changes, additions and sub-combinations thereof.

Claims

1. A method for constructing an industrial chain knowledge graph based on deep learning, characterized in that: The following steps are involved: S1. Deploy a cloud data collection program, wherein the cloud data collection program is configured to obtain industrial chain data from multiple data sources and clean and pre-process the data; S2. Classify and filter the preprocessed data. For structured data, use the random forest algorithm to classify and filter. For unstructured data, use word embedding combined with a deep learning model to classify and filter the stored data to obtain text data related to the industrial chain. S3, filter the text data for entity words, use the ELECTRA model as a classifier, and combine the BM-25F score to filter out important entity words, remove non-word words, obtain entity words, and annotate them; S4, construct graph structure representation, create entity nodes, create and initialize edges, and use GCNE model for entity recognition and relationship extraction; S5. Based on the results of entity recognition and relationship extraction, the entity, attribute and relationship triples required for the knowledge graph are obtained, and the triples are stored in the graph database to construct the industrial chain knowledge graph; S6. When users search for industrial chain information, create appropriate indexes based on the query requirements of the knowledge graph, match the knowledge graph data, and push relevant data to users.

2. The method according to claim 1, characterized in that: For step S1, rdd conversion is used for unstructured data in the collected data, and DataFrame conversion is used for structured data. After the data is aggregated using an aggregation function, the processed data is saved to cloud storage.

3. The method according to claim 1, characterized in that: For step S3, the BM-25F score of each field is calculated, the text embedding of each field is generated using the ELECTRA model, and other features that help identify entity words are extracted; a deep learning model is built, the BM-25F score and text embedding are used as input, the Transformer model is used for feature fusion and processing, and the probability of whether the output word is an important entity word is constructed.

4. The method according to claim 1, characterized in that: For step S3, entity word labeling specifically includes: combining the evaluation results of industry experts in the industrial chain, labeling the extracted entity words into six categories: company, expert, industry, product, technology, and application, and automatically labeling them in BIOES format.

5. The method according to claim 1, characterized in that: For step S4, constructing a graph structure representation specifically includes identifying entities in the text based on word segmentation results and industry chain domain knowledge, creating a node for each entity, and assigning an initial feature vector representation to each entity node. The initial feature vector is created based on basic attribute information of the entity, and the attribute information is extracted from the text or obtained through other external data sources. According to the relationship description between entities in the text, edges connecting entity nodes are created.

6. The method according to claim 4, characterized in that: Step S4 specifically includes: inputting the constructed graph structure data into the GCNE model, determining the set of neighbor nodes at each node according to the structure of the graph, calculating the message from the neighbor node to the current node for each neighbor node, wherein the message is calculated based on the current feature vector of the neighbor node and the feature of the edge connecting the node, and then aggregating the messages from all neighbor nodes; updating the state of the current node using the aggregated neighbor node information and edge information; mapping the updated node feature vector to the entity label space using a linear layer; constructing a conditional random field (CRF) model, whose input is the representation of the node in the entity label space output by the linear layer, using the CRF model to predict the entity label of each node, selecting all possible node representations from the graph structure, and splicing them together to form the representation of the entity pair; then mapping the representation of the entity pair to the relationship label space, and predicting the relationship label between the entity pairs through the Softmax function.

7. A deep learning-based industrial chain knowledge graph construction system, characterized in that: include: Cloud data collection device, used to deploy cloud data collection programs, obtain industrial chain data from multiple data sources, and clean and pre-process the data; Data classification and screening device, used to classify and screen the pre-processed data, where the random forest algorithm is used to classify and screen the structured data, and the word embedding combined with the deep learning model is used to classify and screen the stored data for unstructured data, so as to obtain text data related to the industrial chain; The entity word processing device is used to filter entity words in text data, use the ELECTRA model as a classifier, and combine the BM-25F score to screen out important entity words, remove non-word words, obtain entity words, and mark them; Graph structure construction and analysis device, used to construct graph structure representation, create entity nodes, and create and initialize edges, using GCNE (Graph Convolutional Networks with Edge Features) model for entity recognition (NER) and relationship extraction (RE); A knowledge graph construction device is used to obtain the entity, attribute and relationship triples required for the knowledge graph based on the results of entity recognition and relationship extraction, and store the triples in a graph database, thereby constructing an industrial chain knowledge graph; The data push device is used to create appropriate indexes, match knowledge graph data, and push relevant data to users based on the query requirements of the knowledge graph when users retrieve industrial chain information.

8. An electronic device, characterized in that: The electronic device comprises: a memory and a processor, wherein the memory and the processor are coupled; the memory stores program instructions, and when the program instructions are executed by the processor, the electronic device executes the method according to any one of claims 1 to 7.

9. A computer-readable storage medium, characterized in that: The method comprises a computer program, which, when executed on an electronic device, enables the electronic device to execute the method according to any one of claims 1 to 7.

10. A computer program product, comprising a computer program stored on a computer-readable storage medium, wherein when the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Pre-training dual attention neural network semantic inference dialogue retrieval method and system, retrieval equipment and storage medium

    CN113535918A

  • Text enhancement method and device, electronic equipment and storage medium

    CN113822047A

  • Energy industry knowledge graph construction method and device based on multi-source heterogeneous data fusion technology

    CN117313849A

  • NLP-based recommender system for efficient analysis of trouble tickets

    US20240241902A1