Knowledge graph construction method and device, equipment and medium
By constructing a structured knowledge entity base and dependency syntax tree for PLC patent data, the problem of unified processing of multi-source heterogeneous patent data was solved, realizing accurate and automated construction and management of PLC patent data, and improving the patent analysis capabilities in the PLC field.
Patent Information
- Application Number
- CN202510968423.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-14
- Publication Date
- 2025-11-28
AI Technical Summary
Existing knowledge graph construction methods struggle to unify the processing of complex and semantically diverse multi-source heterogeneous patent big data, resulting in difficulties and low accuracy in constructing knowledge graphs for PLC patent data, and consequently, the effectiveness of patent analysis cannot be guaranteed.
By structuring the patent data in the patent specification, and using entity recognition models and disambiguation and alignment with a structured knowledge entity database, a PLC patent data knowledge graph is constructed. This includes entity annotation, disambiguation and alignment processing.
It enables accurate and automated construction and management of PLC patent data, improves the accuracy of patent analysis and data processing efficiency in the PLC field, and supports the dynamic updating and integration of PLC patent data.
Smart Images

Figure CN121031733A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of knowledge graph construction technology, specifically to a knowledge graph construction method, apparatus, device, and medium. Background Technology
[0002] As the primary productive force, science and technology directly determine the quality and efficiency of economic development. Technological innovation can optimize industrial structure and improve total factor productivity. Therefore, in recent years, the country has vigorously promoted scientific and technological innovation, attached importance to intellectual property rights and the transformation of research results, and the number of patent applications has risen sharply, with various patent database platforms filled with massive amounts of patent big data.
[0003] Because patent big data contains a wealth of technical solutions and knowledge, it is possible to extract creatively valuable technical information through analysis. For example, it can reveal technological evolution trends in existing technical fields, identify emerging fields, and pinpoint the relationships between technical knowledge. However, since patent documents are written in natural language, they not only have diverse textual expressions, but also differ in terminology and industry standards across different technical fields. Therefore, it is difficult to form a unified structured representation and semantic association for patent data generated from such diverse patent documents, posing significant challenges to the mining and analysis of patent information. Summary of the Invention
[0004] To address one of the aforementioned technical deficiencies, this application provides a knowledge graph construction method, apparatus, device, and medium.
[0005] In a first aspect, embodiments of this application provide a method for constructing a knowledge graph based on PLC patent data, comprising: obtaining first patent data from multiple patent databases, the first patent data including structured patent data and unstructured patent data; performing standardization processing on the structured patent data and the unstructured patent data respectively to obtain initial standardized patent data; performing cross-data source semantic mapping and unified encoding processing on the initial standardized patent data to obtain a standardized patent dataset with global entity identifiers; performing entity annotation, disambiguation, and alignment processing on the text data in the standardized patent dataset based on an entity recognition model; generating a structured knowledge entity library based on the processing results of the standardized patent dataset and the entity recognition model; constructing a dependency syntax tree corresponding to each text sentence in the text data based on the structured knowledge entity library; determining target entity triples based on the central word of each text sentence and the corresponding dependency syntax tree; and constructing a PLC patent data knowledge graph based on the target entity triples corresponding to all dependency syntax trees.
[0006] In one optional embodiment, the structured patent data and the unstructured patent data are respectively standardized to obtain initial standardized patent data, including: extracting and standardizing fields from the structured patent data to obtain patent basic data with a unified format; performing text recognition on the unstructured patent data and converting it into structured patent text data; and using the patent basic data and the patent text data as the initial standardized patent data.
[0007] In one optional embodiment, the process of performing text recognition on the unstructured patent data and converting it into structured patent text data includes: performing natural language recognition and format conversion on the text data in the unstructured patent data to obtain first text data; performing character extraction and format conversion on the image and text data in the unstructured patent data to obtain second text data; and using the first text data and the second text data as structured patent text data.
[0008] In an optional embodiment, the method further includes: converting the patent text data into text vectors using a first embedding model; and determining redundant patent data and performing deduplication processing by calculating vector similarity.
[0009] In one optional embodiment, the initial standardized patent data is subjected to cross-data source semantic mapping and unified encoding processing to obtain a standardized patent dataset with global entity identifiers. This includes: based on a semantic mapping function and a standardization function, mapping patent data from different patent databases that correspond to the same semantics in the initial standardized patent data to the same entity identifiers to obtain a standardized patent dataset with global entity identifiers.
[0010] In one optional embodiment, entity annotation, disambiguation, and alignment processing of text data in the standardized patent dataset based on an entity recognition model includes: extracting text data from the standardized patent dataset; segmenting the text data by sentence and word segmentation to obtain a text sequence including multiple words; converting each word in the text sequence into a word vector through a second embedding model; performing entity recognition and annotation on each word vector through a bidirectional long short-term memory network model to obtain the entity label score corresponding to each word; and correcting the entity label score corresponding to each word based on a conditional random field model to determine the entity label of the corresponding word.
[0011] In one optional embodiment, generating a structured knowledge entity library based on the processing results of the standardized patent dataset and the entity recognition model includes: determining the global entity identifier and basic patent data associated with the word segmentation labeled with entity tags, as entity attributes of the corresponding word segmentation, based on the processing results of the standardized patent dataset and the entity recognition model; and generating a structured knowledge entity library based on the entity tags and entity attributes corresponding to each word segmentation.
[0012] In one optional embodiment, constructing a dependency syntax tree corresponding to each text sentence in the text data based on the structured knowledge entity library includes: performing syntactic analysis on the text data to determine the syntactic structure of each text sentence; wherein the syntactic structure includes each word segment in the corresponding text sentence, its corresponding part of speech, and the dependency relationships between each word segment; determining candidate entity triples corresponding to each text sentence based on the structured knowledge entity library and the syntactic structure of each text sentence; and constructing a corresponding dependency syntax tree based on the candidate entity triples corresponding to each text sentence.
[0013] In one optional embodiment, determining the candidate entity triples corresponding to each text sentence based on the structured knowledge entity library and the syntactic structure of each text sentence includes: constructing a PLC terminology ontology library based on the structured knowledge entity library and PLC standard terminology; and determining the candidate entity triples corresponding to each text sentence based on the PLC terminology ontology library and the syntactic structure of each text sentence.
[0014] In one optional embodiment, determining the target entity triple based on the headword of each text sentence and the corresponding dependency syntax tree includes: performing a shortest path traversal on the corresponding dependency syntax tree based on the headword of each text sentence to determine the target triple corresponding to the corresponding dependency syntax tree.
[0015] In an optional embodiment, the method further includes: periodically obtaining second patent data from multiple patent databases; filtering the second patent data based on the PLC terminology ontology to determine newly added entities and / or newly added entity dependencies; and updating the PLC patent data knowledge graph according to the newly added entities and / or newly added entity dependencies.
[0016] In an optional embodiment, the second patent data corresponds to a timestamp, and the method further includes: determining the entity change frequency and / or entity dependency change characteristics of the PLC patent data knowledge graph based on the timestamps corresponding to the newly added entities and / or the newly added entity dependencies, for use in PLC domain analysis.
[0017] Secondly, embodiments of this application provide a knowledge graph construction device based on PLC patent data, comprising: an acquisition module for acquiring first patent data from multiple patent databases, the first patent data including structured patent data and unstructured patent data; a first processing module for standardizing the structured patent data and the unstructured patent data respectively to obtain initial standardized patent data; a second processing module for performing cross-data source semantic mapping and unified encoding processing on the initial standardized patent data to obtain a standardized patent dataset with global entity identifiers; a third processing module for performing entity annotation, disambiguation, and alignment processing on text data in the standardized patent dataset based on an entity recognition model; a generation module for generating a structured knowledge entity library based on the processing results of the standardized patent dataset and the entity recognition model; a first construction module for constructing a dependency syntax tree corresponding to each text sentence in the text data based on the structured knowledge entity library; a determination module for determining target entity triples based on the central word of each text sentence and the corresponding dependency syntax tree; and a second construction module for constructing a PLC patent data knowledge graph based on the target entity triples corresponding to all dependency syntax trees.
[0018] Thirdly, embodiments of this application provide an electronic device, including a processor and a memory, wherein when the processor executes a computer program, it is used to implement the knowledge graph construction method based on PLC patent data as described in the first aspect.
[0019] Fourthly, embodiments of this application provide a computer program product, including a computer program / instruction, which, when executed by a processor, implements the knowledge graph construction method based on PLC patent data as described in the first aspect.
[0020] In summary, the knowledge graph construction method based on PLC patent data provided in this application acquires multi-source heterogeneous patent big data from different patent databases, performs format standardization processing on the structured and unstructured patent data, and performs entity recognition, disambiguation, and alignment processing on the unstructured patent data to construct a structured knowledge entity library with a unified structure. Based on this structured PLC standard terminology, an enhanced semantic PLC terminology ontology library is constructed. When extracting entity dependency relationships from patent data based on dependency syntax trees, entities and entity dependency relationships adapted to the PLC domain can be extracted from unstructured text data, transforming unstructured text data into a structured knowledge graph. The entire process can be completed accurately and automatically, and the extracted entities and entity dependency relationships have a higher adaptability to the PLC domain.
[0021] In addition, the knowledge graph construction method based on PLC patent data provided in this application can also periodically acquire new patent data and identify incremental entities and / or entity dependencies in the patent data based on timestamps. Furthermore, by comparing the incremental entities and / or entity dependencies with the knowledge graph of the original PLC patent data at the structural and semantic levels, the incremental entities and / or entity dependencies that meet the requirements are integrated into the knowledge graph of the PLC patent data, realizing the dynamic fusion and updating of the knowledge graph of PLC patent data, which greatly facilitates analysis in the PLC field. Attached Figure Description
[0022] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:
[0023] Figure 1 A flowchart illustrating a knowledge graph construction method based on PLC patent data, provided for embodiments of this application;
[0024] Figure 2 This application provides a schematic diagram of a process for processing patent data.
[0025] Figure 3 A schematic diagram of a knowledge graph structure provided in an embodiment of this application;
[0026] Figure 4 This is a schematic diagram of another knowledge graph structure provided in an embodiment of this application;
[0027] Figure 5 A schematic diagram of a knowledge graph construction device based on PLC patent data provided in this application embodiment;
[0028] Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0029] To make the technical solutions and advantages of the embodiments of this application clearer, the exemplary embodiments of this application will be described in further detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not an exhaustive list of all embodiments. It should be noted that, unless otherwise specified, the embodiments and features in the embodiments of this application can be combined with each other.
[0030] In recent years, with the exponential growth in the number of patents, a massive, complex, and semantically diverse multi-source heterogeneous patent big data has been formed. Since knowledge graphs, as structured semantic networks, can effectively organize and manage complex knowledge within a domain, constructing knowledge graphs based on patent big data is beneficial for comprehensive and systematic patent analysis. For example, knowledge graphs can be used to study technological evolution trends in related fields, discover emerging fields, and mine technological knowledge and knowledge relationships, etc.
[0031] With the rapid development of intelligent manufacturing and industry automation, Programmable Logic Controllers (PLCs) are being used as core control units across various industries. Applying PLCs to the construction and management of patent data knowledge graphs can improve data processing efficiency. However, traditional knowledge graph construction methods struggle to uniformly process complex, semantically diverse, and multi-source heterogeneous patent big data. Consequently, PLCs face difficulties in achieving unified automated control in the construction and management of patent data knowledge graphs, resulting in difficulties in knowledge graph construction, low accuracy, and consequently, uncertainty in the accuracy of patent analysis based on knowledge graphs.
[0032] To address the aforementioned issues, this application provides a method for constructing a knowledge graph based on PLC patent data, aiming to improve the accuracy of knowledge graph construction based on PLC patent big data in the domain internet environment and reduce the difficulty of knowledge mining and data analysis.
[0033] The knowledge graph construction method based on PLC patent data provided in the embodiments of this application will now be described in conjunction with the accompanying drawings.
[0034] Figure 1 A flowchart of a knowledge graph construction method based on PLC patent data provided for embodiments of this application is shown below. Figure 1 As shown, the method includes:
[0035] S101. Obtain first patent data from multiple patent databases, the first patent data including structured patent data and unstructured patent data;
[0036] S102. Standardize the structured patent data and unstructured patent data respectively to obtain initial standardized patent data;
[0037] S103. Perform cross-data source semantic mapping and unified encoding processing on the initial standardized patent data to obtain a standardized patent dataset with global entity identifiers.
[0038] S104. Perform entity annotation, disambiguation and alignment processing on text data in the standardized patent dataset based on the entity recognition model;
[0039] S105. Generate a structured knowledge entity database based on the processing results of the standardized patent dataset and the entity recognition model.
[0040] S106. Based on the structured knowledge entity base, construct the dependency syntax tree corresponding to each text sentence in the text data;
[0041] S107. Based on the central word of each text sentence and the corresponding dependency syntax tree, determine the target entity triples;
[0042] S108. Construct a PLC patent data knowledge graph based on the target entity triples corresponding to all dependency syntax trees.
[0043] In this embodiment, patent data is divided into structured patent data and unstructured patent data. Structured patent data refers to data uniquely associated with patent documents, such as application number, application date, classification number, applicant, inventor, etc. Unstructured patent data refers to data used to embody the technical content of patent documents, such as abstracts, claims, descriptions, and drawings. Based on this, to construct a knowledge graph based on PLC patent data, patent data is first obtained from multiple patent databases, and the structured and unstructured patent data included are identified. Optionally, to distinguish it from newly added patent data used to update the knowledge graph, the patent data used to create the knowledge graph is referred to as the first patent data, and correspondingly, the patent data used to update the knowledge graph is referred to as the second patent data. The types of multiple patent databases are not limited; they can optionally be authoritative official databases, such as, but not limited to, the China National Intellectual Property Administration (CNIPA), Web of Science, Google Patents, etc. The specific database type is determined according to actual needs and will not be detailed here.
[0044] In practical applications, since different patent databases have their own exclusive data storage methods, in order to perform unified data processing, after obtaining the first patent data, the structured patent data and unstructured patent data are standardized separately to obtain the initial standardized patent data. The specific processing method will be described in detail in subsequent embodiments.
[0045] Furthermore, because the technical content in patent documents is written in natural language, the same object can be described with different textual content in different patent documents. For example, the proper noun "PLC" may be described as "PLC controller" in one patent document and "programmable controller" in another. Similarly, textual content containing the manufacturer "Siemens" may be described as "Siemens AG" in one patent document and "Siemens Ltd." in another, and so on. In these examples, although different patent documents use different expressions to describe the same object, the semantic meaning of the same object is the same in different patent documents. Therefore, when constructing a knowledge graph, it should be treated as the same entity.
[0046] Based on this, to ensure the semantic uniqueness of each entity in the constructed knowledge graph, in this embodiment, after obtaining the initial standardized patent data, cross-data source semantic mapping and unified encoding processing are further performed on the initial standardized patent data to obtain a standardized patent dataset with global entity identifiers. Specifically, objects described with different text content in the standardized patent dataset, if corresponding to the same semantics, are mapped to the same entity identifier. This entity identifier is the entity identifier of the corresponding text content in the knowledge graph, used for retrieving the corresponding entity.
[0047] In addition to creating a unified global entity identifier for text content with the same semantics but different expressions, this application embodiment also performs entity annotation, disambiguation, and alignment processing on the text data in the standardized patent dataset based on an entity recognition model to transform unstructured text content into structured semantic annotations for accurate entity boundary and type localization. This is done to label each word segment in the corresponding text data with an entity tag. Based on this, different text content with the same semantics can be labeled with the same type of entity tag. For example, in a patent document containing the text content "PLC controller based on edge computing," after the above processing, "edge computing" can be labeled as "technology," and "PLC controller" as "equipment." Similarly, in another patent document containing the text content "a control method based on PLC," after the above processing, "PLC" can be labeled as "equipment," and "control method" as "method." Thus, even if two patent documents express the proper noun "PLC" differently, they are both labeled with the same entity type, ensuring the accuracy of entity types in the constructed knowledge graph.
[0048] Based on the above, after creating global entity identifiers for structured and unstructured patent data, and performing structured semantic annotation on unstructured patent data, a structured knowledge entity library can be generated based on the processing results of the standardized patent dataset and the entity recognition model. This library can then be used for subsequent entity recognition and relation extraction. The structured knowledge entity library includes all entity tags obtained from the above processing, as well as global entity identifiers and basic patent data associated with the text data labeled with these entity tags, such as application numbers, IPC classification numbers, source documents, etc.
[0049] Furthermore, after obtaining the structured knowledge entity base, in order to determine the dependency relationships between entities, in this embodiment, a corresponding dependency syntax tree is first constructed based on each text sentence in the text data of the standardized patent dataset. Then, based on the structured knowledge entity base and the dependency syntax tree corresponding to each text sentence, the target entity triples of each dependency syntax tree are determined. The target entity triples include a head entity, a tail entity, and the dependency relationship between them. For example, for the target entity triple (PLC controller, control, system), "PLC controller" is the head entity, "system" is the tail entity, and "control" is the dependency relationship between them, meaning "system" is controlled by "PLC controller". Based on this, a PLC patent data knowledge graph can be constructed using a graph structure based on the target entity triples corresponding to all dependency syntax trees.
[0050] In the above embodiments, the implementation principle of the main method steps of the knowledge graph construction method based on PLC patent data provided in this application embodiment has been explained. Below, through specific embodiments, the detailed implementation of each of the above method steps will be illustrated by example.
[0051] For step S101, this application embodiment does not limit the specific method of obtaining the first patent data from multiple patent databases. Optionally, it can be obtained through the open interfaces provided by each patent database, or it can be obtained through web crawling technology. The specific method is determined according to actual needs.
[0052] For step S102, the storage format of structured patent data such as application number, application date, classification number, applicant, and inventor may differ in different patent databases. Therefore, in this embodiment, after obtaining the first patent data and determining the structured and unstructured patent data within it, the structured patent data undergoes field extraction and normalization processing to obtain patent basic data with a unified format. The specific method for field extraction and normalization is not limited. For example, an Extensible Markup Language (XML) parser can be used to extract and normalize the structured patent data; alternatively, a lightweight data exchange format, such as a JavaScript Object Notation (JSON) parser, can be used; or regular expressions can be used for conversion, etc.
[0053] Of course, the above is only an illustrative example and is not limited to this in actual application. This step aims to transform structured patent data from different patent databases into a unified data format. The specific implementation method can be flexibly selected according to actual needs. As long as the method meets the actual needs, it can be used for implementation, which will not be described in detail here.
[0054] Accordingly, for unstructured patent data such as specification abstracts, claims, specifications, and drawings, which contain a large amount of text, text recognition is performed on this unstructured data to convert it into a machine-processable data format, transforming it into structured patent text data. Based on this, the patent base data obtained from processing structured patent data and the patent text data obtained from processing unstructured data are used together as initial standardized patent data. This provides a data foundation for subsequent entity annotation and knowledge graph construction, facilitating data identification and processing.
[0055] In practical applications, since the abstract, claims, and specification are written in natural language, when performing text recognition and format conversion on this part of the unstructured patent data, the corresponding text data can be subjected to natural language recognition and format conversion to obtain the first text data. As for the drawings in the specification, although they include text content, in patent databases, the drawings are usually stored as graphic objects. Therefore, for this part of the unstructured patent data, the corresponding graphic data can be extracted and format converted to obtain the second text data. Based on this, the first and second text data are combined as structured patent text data.
[0056] In this embodiment, the specific method for performing natural language recognition and format conversion on the text data is not limited. Optionally, natural language processing tools can be used to perform text recognition and format conversion on the corresponding text data. Furthermore, the specific type of tool used is not limited, for example, including but not limited to Jieba, HanLP, THU LexicalAnalyzer for Chinese (THULAC), Natural Language Toolkit (NLTK), Stanford Core Natural Language Processing (Stanford CoreNLP), etc. The specific tool selected can be chosen according to actual needs, and will not be detailed here.
[0057] Correspondingly, there are no restrictions on the specific methods for character extraction and format conversion of the image and text data. Optionally, Optical Character Recognition (OCR) tools can be used to extract characters and convert formats from the corresponding image and text data. Furthermore, there are no restrictions on the specific type of tool used, such as, but not limited to, Tesseract OCR, EasyOCR, Umi-OCR, etc. The choice of which tool to use depends on the actual needs and will not be detailed here.
[0058] Figure 2 The diagram illustrates the standardization and formatting processes performed on structured and unstructured data, respectively. Figure 2 This is only one option; please refer to the above explanation for details, which will not be elaborated here.
[0059] In this embodiment, step S102 can be considered as a preprocessing operation before entity annotation and knowledge graph construction. Optionally, in addition to the above formatting operation, since the first patent data obtained from different patent databases may include redundant data with the same content, for example, different patent databases may store patent data corresponding to the same patent document, the obtained first patent data can also be deduplicated.
[0060] Optionally, the patent text data can be converted into text vectors using a first embedding model. Then, redundant patent data can be identified and deduplicated by calculating vector similarity. For example, the first embedding model can be a document vector model (Doc2Vec). Based on this, the patent text data corresponding to each first patent data can be input into Doc2Vec to obtain the corresponding text vectors. Further, the cosine similarity between each text vector is calculated using the following formula (1). Based on the comparison between the corresponding cosine similarity and a preset similarity threshold, it is determined whether the two correspond to the same patent document; where v A and v B Each represents a different text vector, sim(v) A ,v B The ) represents the pre-defined similarity between the two text vectors. Furthermore, if two text vectors with a cosine similarity greater than or equal to the preset similarity threshold correspond to the same patent document, only one copy of the first patent data and the initial standardized patent data after the above formatting process are retained to remove redundant patent data.
[0061]
[0062] In practical applications, because different patent databases may define different field types and store different field contents, there may be slight differences in the content of the same patent document stored in different patent databases. Alternatively, during the character extraction and format conversion of the aforementioned image and text data, differences in the layout of images and text in different patent databases may result in different layouts of the accompanying drawings in the specification of the same patent document in different patent databases, causing the OCR tool to fail to recognize some of the image and text data provided by some patent databases or to be unable to recognize some content. Therefore, in this embodiment of the application, missing data fields or text content can also be supplemented.
[0063] In one optional approach, after identifying the first patent data corresponding to the same patent document obtained from different patent databases, the text recognition results and image recognition results of this first patent data can be compared and contrasted to fill in missing text content. In another optional approach, historical patent documents within the same technical field can be combined to infer field and text patterns to fill in missing fields or text content. Based on this, when removing redundant patent data, the most complete text can be retained, and the remaining patent data can be deleted.
[0064] Of course, the preprocessing operations in the above embodiments are only illustrative examples and are not limited to these in actual applications. For example, they may also include compliance verification, invalid data filtering, data compression, etc. The specific operation content can be determined according to actual needs and will not be described in detail here.
[0065] Based on the above, after obtaining the initial standardized patent data that meets the requirements, entity annotation and knowledge graph construction can be performed based on the initial standardized patent data.
[0066] For step S103, before entity annotation and knowledge graph construction, to ensure the uniqueness of entities within the knowledge graph, this embodiment performs cross-data source semantic mapping and unified encoding processing on the initial standardized patent data to obtain a standardized patent dataset with global entity identifiers. Optionally, referring to the following formula (2), based on the semantic mapping function and standardization function, patent data from different patent databases in the initial standardized patent data that correspond to the same semantics are mapped to the same entity identifier to obtain a standardized patent dataset with global entity identifiers. Wherein, e i For patent data from different patent databases that correspond to the same semantics, such as "PLC", "PLC controller", "programmable controller", f s f is a semantic mapping function. n For the standardized function, EntityID has multiple e i A common global entity identifier.
[0067] EntityID = f n (f s (e i ),i=1,2,…,n (2)
[0068] For step S104, in order to perform entity annotation, disambiguation, and alignment processing on the text data in the standardized patent dataset based on the entity recognition model, text data can first be extracted from the standardized patent dataset, and text recognition can be performed on the corresponding text data to identify the segmentation boundaries, for example, using a period as the segmentation boundary. Then, based on the identified segmentation boundaries, the text data is segmented by sentence and word-segmented to obtain a text sequence including multiple words. Further, each word in the text sequence is converted into a word vector through a second embedding model, and then entity recognition and annotation are performed on each word vector through a bidirectional long short-term memory network model (BiLSTM) to obtain the entity label score corresponding to each word, and the entity label score corresponding to each word is corrected based on a conditional random field model (CRF) to determine the entity label of the corresponding word.
[0069] Optionally, the second embedding model can be a Bidirectional Encoder Representations from Transformers (BERT) model. Assume the text sequence obtained after segmentation and word segmentation is X = {x1, x2, ..., x...}. n}, where x i Each word is a segment. Based on this, inputting the text sequence X into the BERT model allows each segmented word x to be segmented. i Convert to the corresponding text vector Furthermore, entity annotation of the text sequence X using a BiLSTM model can capture the contextual dependency information of each word in the patent text, enhancing the ability to recognize long sentences and complex expressions, and obtaining an independent score for the tag to which each word belongs. Further, by incorporating a CRF layer and introducing a transition matrix, the tag sequence Y = {y1, y2, ..., y...} is defined. n The total score function (3) is combined with the conditional probability function (4) that maximizes the sequence label. The label sequence corresponding to each word is corrected by minimizing the negative log-likelihood function (5), thereby maximizing the conditional probability of the label sequence and ensuring the accuracy of entity recognition.
[0070]
[0071] L=-log(P(Y|X)) (5)
[0072] in, For label y i-1 To tag y i The CRF transition matrix, The i-th word is assigned the label y. i The score.
[0073] In step S105, after the above processing, each word in the text sequence can be labeled with its corresponding entity tag. For example, "PLC controller" is labeled as the "equipment" entity, and "Schneider Electric" is labeled as the "organization" entity. Since these text sequences are determined based on unstructured patent data, and the corresponding unstructured data has a direct correlation with the structured data corresponding to the same patent document (e.g., application number, IPC classification number, inventor, applicant, etc.), and each word in the text sequence is mapped to a global entity identifier, the global entity identifier and basic patent data associated with the labeled entity tags can be determined based on the processing results of the standardized patent dataset and entity recognition model. These can be used as the entity attributes of the labeled words. Furthermore, a structured knowledge entity base can be generated based on the entity tags and entity attributes corresponding to each word.
[0074] In this embodiment of the application, since the structured knowledge entity library includes all the entity tags obtained from the above processing and the basic patent data associated with the corresponding entity tags, that is, the structured patent data after standardization, the structured knowledge entity library can not only determine the type of each entity and the dependencies between them, but also determine the patent information associated with each entity, providing a search basis for the construction and updating of the knowledge graph when determining entity nodes.
[0075] For step S106, when constructing the dependency syntax tree corresponding to each text sentence in the text data based on the structured knowledge entity library, syntactic analysis can be performed on the text data first to determine the syntactic structure of each text sentence. The syntactic structure includes each word segment in the corresponding text sentence, its corresponding part of speech, and the dependency relationships between the word segments. Then, based on the structured knowledge entity library and the syntactic structure of each text sentence, the determined syntactic structure of each text sentence can be mapped to the dependency relationships of each entity in the structured knowledge entity library, and candidate entity triples corresponding to each text sentence can be determined. Based on this, the corresponding dependency syntax tree can be constructed according to the candidate entity triples corresponding to each text sentence.
[0076] Optionally, a period can be used as the end boundary of each text sentence. Based on this, when performing syntactic analysis on text data, the corresponding text content can be identified first, and the positions of each period can be determined. Then, based on each period, the corresponding text content is segmented into multiple text sentences, forming a syntactic structure analysis dataset. Further, all text sentences in the syntactic structure analysis dataset are traversed, and the tokens labeled as entities in each text sentence are analyzed sentence by sentence to determine their corresponding parts of speech. Then, based on each token and its corresponding part of speech, the corresponding entity labels and entity dependency relations are matched in a structured knowledge entity base to form candidate entity triples for each text sentence.
[0077] For example, given the text sentence "The PLC controller communicates with the remote module via Ethernet," text recognition and syntactic analysis can identify "PLC controller" and "remote module" as labeled entities. Based on word segmentation such as "via," "Ethernet," and "communication," the dependency relationship between these two entity events can be determined to be "communication relationship," with the corresponding communication medium being "Ethernet." If the entity label corresponding to "PLC controller" and "remote module" in the structured knowledge entity base is "device," then the corresponding candidate entity triple is (device, communication relationship, device). Thus, in the dependency syntactic tree constructed based on this candidate entity triple, it can be determined that there is a communication relationship between the two devices.
[0078] Further, optionally, in addition to mapping the entity label types and dependency relationships of each node in the dependency syntax tree through a structured knowledge entity base, to improve the semantic matching between the dependency syntax tree and the PLC domain, in this embodiment, when determining the candidate entity triples corresponding to each text sentence, a PLC terminology ontology can be constructed first based on the structured knowledge entity base and PLC standard terminology. Then, based on the PLC terminology ontology and the syntactic structure of each text sentence, the candidate entity triples corresponding to each text sentence are determined. The specific content of the PLC standard terminology is not limited, but optionally includes, but is not limited to, standardized and high-frequency terms covering key areas such as PLC controller architecture, communication protocols, execution units, control logic, and application scenarios.
[0079] Thus, the PLC terminology ontology constructed based on the structured knowledge entity base and PLC standard terminology not only ensures accurate and unique labeling of each entity, but also allows for adaptive adjustment of entity labels to suit the characteristics of the PLC domain. Based on this, the syntactic structure of each text sentence is mapped to entity triples in the PLC terminology ontology, which can serve as candidate entity triples for each text sentence, giving the corresponding candidate entity triples stronger PLC semantic characteristics.
[0080] For step S107, when determining the target entity triples based on the headword and corresponding dependency syntax tree of each text sentence, the headword and other dependent words governed by the headword of each text sentence can be determined first based on the syntactic structure, such as the grammatical structure "subject-verb-object". Then, based on the node positions of each headword and its governed dependent words in the dependency syntax tree, the linear distance from each dependent word node to the headword node, i.e., the dependency distance, is calculated. Redundant syntactic dependency relations with large distances are deleted until the dependency distances from all dependent word nodes to the headword node are equal, thus obtaining the target dependency syntax tree. The triples in the target dependency syntax tree can be represented as follows: Where h i , t i For entities, r syn This is a syntactic dependency relation.
[0081] Based on this, shortest path traversal of the target dependency parse tree yields the corresponding target triples. Optionally, to ensure semantic consistency between the entity recognition and entity dependency extraction stages and avoid semantic bias caused by model differences, the semantic embedding capability of BiLSTM-CRF can be reused to assign local feature vectors to the text words corresponding to each entity label. Optionally, h = (x1, x2, ..., x...) can be used. N This means that, in this way, we can select from candidate entity triples: (h, r synExtracting the shortest dependency path P from ,t)∈T h→t Arrange all nodes on the path in order to form a path matrix: P = [h1; ...; h L ] T Where L is the path length. Based on this, by using a convolutional neural network to perform multi-window convolution and pooling operations on the path matrix, global semantic features can be obtained. Then, the path vector and entity vector are concatenated: v path ⊕h i ⊕h j The data is then input into the Softmax classifier, which can predict entity dependency relationships to obtain the target entity triples.
[0082] For step S108, based on the above, after obtaining the target entity triples corresponding to each dependency syntax tree, the dependency relationships between entity nodes can be extracted from the target entity triples to construct a PLC patent data knowledge graph based on the target entity triples corresponding to all dependency syntax trees. Figure 3 A schematic diagram of the structure of a PLC patent data knowledge graph constructed through the above process is shown, as follows: Figure 3 As shown, each node (such as equipment, control strategy, application system, etc.) in the PLC patent data knowledge graph is not only unique, but the types of each entity and the dependencies between entities (such as "composition relationship", "applicability relationship", "technology dependence", etc.) also conform to the PLC field specifications, which is conducive to the management and application of PLC patent data.
[0083] In this application, in addition to constructing a PLC patent data knowledge graph based on the first patent data, second patent data can also be periodically obtained from multiple patent databases through open interfaces or web crawling technology for each patent data. The second patent data can be patent data corresponding to recently granted or published patent documents. Furthermore, after obtaining the second patent data, it can undergo format standardization, deduplication, supplementation of missing fields and content, entity recognition, disambiguation, and alignment processing. Specific processing procedures can be found in the description of the first patent data in the aforementioned embodiments, and will not be repeated here.
[0084] Optionally, after the above processing, the second patent data can be filtered based on the PLC terminology ontology to determine whether there are any newly added entities and / or newly added entity dependencies. If newly added entities and / or newly added entity dependencies are confirmed, to avoid directly integrating the corresponding newly added entities and / or newly added entity dependencies into the PLC patent data knowledge graph and introducing a large number of noisy nodes, leading to semantic confusion in the graph, in this embodiment of the application, a candidate structured knowledge entity library can first be constructed based on the corresponding newly added entities and / or newly added entity dependencies, and the newly added entities and / or newly added entity dependencies can be stored in the candidate structured knowledge entity library. Based on this, after multiple rounds of semantic consistency detection and frequency statistics, the verified newly added entities and / or newly added entity dependencies are then integrated into the PLC patent data knowledge graph to update the PLC patent data knowledge graph according to the corresponding newly added entities and / or newly added entity dependencies.
[0085] In this embodiment, the specific method for semantic consistency detection and frequency statistics of newly added entities and / or newly added entity dependencies is not limited. In one optional method, for two entities associated with newly added entity dependencies, the existing PLC patent data knowledge graph can be searched to see if there are one or more alternative paths between them; based on the semantic representation of edges and nodes, the semantics of the alternative paths and the newly added entity dependencies are compared (similarity between the semantic aggregation vector of the alternative paths and the vector of the newly added entity dependencies); if the semantics of the alternative paths and the newly added entity dependencies are highly similar, the newly added entity dependencies are determined to be redundant; redundant relationships are filtered to avoid duplicate addition.
[0086] In another alternative approach, the newly added entity dependency relations can be encoded using a pre-trained language model to obtain their semantic vector representations. A set of semantic vectors containing existing entity dependency relations in the PLC patent data knowledge graph can be constructed as a reference library. The semantic similarity between the newly added entity dependency relations and each entity dependency relation in the candidate structured knowledge entity library can be determined based on cosine similarity. If the similarity is lower than a preset threshold, the corresponding entity dependency relation is judged to be semantically abnormal and will not be introduced. If the similarity reaches or exceeds the threshold, it is considered semantically consistent and can proceed to the next step of verification. Based on the logical rules (transitivity, symmetry, anti-relationship, etc.) defined in the PLC patent data knowledge graph, it can be determined whether the newly added entity dependency relations violate the basic logical structure of the PLC patent data knowledge graph, ensuring that they are consistent with the semantics and logic between entities in the PLC patent data knowledge graph.
[0087] Furthermore, for newly added entity dependencies that fail to match any existing entity dependency relationships in the PLC patent data knowledge graph in the semantic consistency detection, they are temporarily stored in the PLC patent data knowledge graph as potential entity dependencies for continuous observation and verification. The verification is further aided by combining the frequency of occurrence of the entity dependency relationship, relationship clustering, and logical reasoning verification. When it is determined that the entity dependency relationship exhibits semantic stability and logical rationality, it is then integrated into the PLC patent data knowledge graph.
[0088] Optionally, the second patent data obtained in this application embodiment all correspond to timestamps. Based on this, the entity change frequency and / or entity dependency change characteristics of the PLC patent data knowledge graph can be determined based on the timestamps corresponding to newly added entities and / or newly added entity dependencies, for use in PLC domain analysis. Optionally, a time window Δt can be set, and the second patent data obtained within the corresponding time window can be represented as: S Δt ={(d i ,t i )|t i ∈[t-Δt,t]}, where t i For patent publication timestamps, d i The second patent data obtained is used as a basis for modeling the frequency of change and relational structure characteristics of newly added entities and / or newly added entity dependencies within a time window, in order to identify entity evolution trends in the PLC field. Optionally, the rate of change of the frequency of occurrence of entity h within the time window Δt can be defined as:
[0089]
[0090] Where f(h,t) is the number of times entity h appears in the sliding window [t-Δt,t], and deg(h,t) is the degree centrality of entity h in the knowledge graph. When EI(h,t) is greater than the preset threshold, entity h can be determined to be an emerging entity in the PLC field.
[0091] Further, optionally, a time decay factor can be used to assign weights to the newly added entities and / or their dependencies to identify the importance of the corresponding newly added entities and / or their dependencies. For example, the weight of each newly added entity and / or its dependency can be expressed as: Where λ is the time decay factor, e is the newly added entity and / or the newly added entity dependency, and w(t) is the corresponding weight. Optionally, Figure 4 This illustrates a schematic diagram of a knowledge graph that updates dynamically over time, such as... Figure 4As shown, based on the above weights, the display pattern of the decay of the strength of new entities and / or new entity dependencies over time can be simulated, making the expression of PLC patent data knowledge graph more consistent with dynamic knowledge flow.
[0092] It should be noted that the embodiments in this application are not limited to... Figure 1 The execution order of each method step in the flowchart shown is not specified. S101, S102, etc., are only used to distinguish different method steps and do not limit the execution order. Optionally, the above method steps can be flexibly adjusted according to different actual processing needs, which will be detailed in this step.
[0093] In summary, the knowledge graph construction method based on PLC patent data provided in this application acquires multi-source heterogeneous patent big data from different patent databases, performs format standardization processing on the structured and unstructured patent data, and performs entity recognition, disambiguation, and alignment processing on the unstructured patent data to construct a structured knowledge entity library with a unified structure. Based on this structured PLC standard terminology, an enhanced semantic PLC terminology ontology library is constructed. When extracting entity dependency relationships from patent data based on dependency syntax trees, entities and entity dependency relationships adapted to the PLC domain can be extracted from unstructured text data, transforming unstructured text data into a structured knowledge graph. The entire process can be completed accurately and automatically, and the extracted entities and entity dependency relationships have a higher adaptability to the PLC domain.
[0094] In addition, the knowledge graph construction method based on PLC patent data provided in this application can also periodically acquire new patent data and identify incremental entities and / or entity dependencies in the patent data based on timestamps. Furthermore, by comparing the incremental entities and / or entity dependencies with the knowledge graph of the original PLC patent data at the structural and semantic levels, the incremental entities and / or entity dependencies that meet the requirements are integrated into the knowledge graph of the PLC patent data, realizing the dynamic fusion and updating of the knowledge graph of PLC patent data, which greatly facilitates analysis in the PLC field.
[0095] Based on the above, this application also provides a knowledge graph construction device based on PLC patent data. Figure 5 A schematic diagram of the structure of a knowledge graph construction device.
[0096] like Figure 5 As shown, the knowledge graph construction device 500 includes an acquisition module 501, a first processing module 502, a second processing module 503, a third processing module 504, a generation module 505, a first construction module 506, a determination module 507, and a second construction module 508, wherein:
[0097] The acquisition module 501 is used to acquire first patent data from multiple patent databases, including structured patent data and unstructured patent data; the first processing module 502 is used to standardize the structured patent data and unstructured patent data respectively to obtain initial standardized patent data; the second processing module 503 is used to perform cross-data source semantic mapping and unified encoding processing on the initial standardized patent data to obtain a standardized patent dataset with global entity identifiers; the third processing module 504 is used to perform entity annotation, disambiguation and alignment processing on the text data in the standardized patent dataset based on an entity recognition model; the generation module 505 is used to generate a structured knowledge entity library based on the processing results of the standardized patent dataset and the entity recognition model; the first construction module 506 is used to construct the dependency syntax tree corresponding to each text sentence in the text data based on the structured knowledge entity library; the determination module 507 is used to determine the target entity triples based on the central word of each text sentence and the corresponding dependency syntax tree; the second construction module 508 is used to construct a PLC patent data knowledge graph based on the target entity triples corresponding to all dependency syntax trees.
[0098] In one optional embodiment, the first processing module 502 performs standardization processing on the structured patent data and the unstructured patent data respectively to obtain initial standardized patent data, which is used for: extracting and standardizing fields from the structured patent data to obtain patent basic data with a unified format; performing text recognition on the unstructured patent data and converting it into structured patent text data; and using the patent basic data and the patent text data as the initial standardized patent data.
[0099] In an optional embodiment, the first processing module 502 performs text recognition on unstructured patent data and converts it into structured patent text data, for the following purposes: performing natural language recognition and format conversion on the text data in the unstructured patent data to obtain first text data; performing character extraction and format conversion on the graphic data in the unstructured patent data to obtain second text data; and using the first text data and the second text data as structured patent text data.
[0100] In an optional embodiment, the first processing module 502 is further configured to: convert patent text data into text vectors through a first embedding model; and determine redundant patent data and perform deduplication processing through vector similarity calculation.
[0101] In an optional embodiment, the second processing module 503 performs cross-data source semantic mapping and unified encoding processing on the initial standardized patent data to obtain a standardized patent dataset with global entity identifiers. This is used to: based on the semantic mapping function and the standardization function, map patent data from different patent databases in the initial standardized patent data that correspond to the same semantics to the same entity identifiers, thereby obtaining a standardized patent dataset with global entity identifiers.
[0102] In an optional embodiment, the third processing module 504 performs entity annotation, disambiguation, and alignment processing on the text data in the standardized patent dataset based on an entity recognition model. This is used to: extract text data from the standardized patent dataset; segment the text data by sentence and word segmentation to obtain a text sequence including multiple words; convert each word in the text sequence into a word vector through a second embedding model; perform entity recognition and annotation on each word vector through a bidirectional long short-term memory network model to obtain the entity label score corresponding to each word; and correct the entity label score corresponding to each word based on a conditional random field model to determine the entity label of the corresponding word.
[0103] In an optional embodiment, the generation module 505 generates a structured knowledge entity library based on the processing results of the standardized patent dataset and the entity recognition model. This library is used to: determine the global entity identifier and basic patent data associated with the word segmentation labeled with entity tags, based on the processing results of the standardized patent dataset and the entity recognition model, as the entity attributes of the corresponding word segmentation; and generate a structured knowledge entity library based on the entity tags and entity attributes corresponding to each word segmentation.
[0104] In an optional embodiment, the first construction module 506 constructs a dependency syntax tree corresponding to each text sentence in the text data based on a structured knowledge entity library, for the following purposes: performing syntactic analysis on the text data to determine the syntactic structure of each text sentence; wherein the syntactic structure includes each word segment in the corresponding text sentence, the corresponding part of speech, and the dependency relationship between each word segment; determining the candidate entity triples corresponding to each text sentence based on the structured knowledge entity library and the syntactic structure of each text sentence; and constructing the corresponding dependency syntax tree based on the candidate entity triples corresponding to each text sentence.
[0105] In an optional embodiment, the first construction module 506 determines the candidate entity triples corresponding to each text sentence based on the structured knowledge entity library and the syntactic structure of each text sentence, for the purpose of: constructing a PLC terminology ontology library based on the structured knowledge entity library and PLC standard terminology; and determining the candidate entity triples corresponding to each text sentence based on the PLC terminology ontology library and the syntactic structure of each text sentence.
[0106] In an optional embodiment, the determining module 507 determines the target entity triple based on the headword of each text sentence and the corresponding dependency syntax tree, and is used to: determine the target triple corresponding to the corresponding dependency syntax tree by performing shortest path traversal on the headword of each text sentence.
[0107] In an optional embodiment, the acquisition module 501 is further configured to: periodically acquire second patent data from multiple patent databases; filter the second patent data based on the PLC terminology ontology to determine newly added entities and / or newly added entity dependencies; and update the PLC patent data knowledge graph based on the newly added entities and / or newly added entity dependencies.
[0108] In an optional embodiment, the second patent data corresponds to a timestamp, and the determining module 507 is further configured to: determine the entity change frequency and / or entity dependency change characteristics of the PLC patent data knowledge graph based on the timestamps corresponding to the newly added entities and / or the newly added entity dependencies, for use in PLC domain analysis.
[0109] It should be noted that the specific functions and implementation processes of each module in the knowledge graph construction device based on PLC patent data can be found in the description of the corresponding embodiment in the knowledge graph construction method based on PLC patent data, and will not be repeated here.
[0110] Based on the above, this application also provides an electronic device. Figure 6 A schematic diagram of the structure of an electronic device, such as Figure 6 As shown, the electronic device 600 includes: a processor 601 and a memory 602 storing a computer program; wherein the processor 601 and the memory 602 may be one or more.
[0111] Memory 602 is primarily used to store computer programs, which can be executed by processor 601, causing processor 601 to control the electronic device to perform corresponding functions, actions, or tasks. In addition to storing computer programs, memory 602 can also be configured to store various other data to support operation on the electronic device. Examples of this data include instructions for any application program or method used to operate on the electronic device.
[0112] The memory 602 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk or optical disk.
[0113] In this embodiment, the implementation of processor 601 is not limited; it can be, for example, but not limited to, a CPU, GPU, or MCU. Processor 601 can be considered a control system for an electronic device, capable of executing computer programs stored in memory 602 to control the electronic device to perform corresponding functions, actions, or tasks. It is worth noting that, depending on the implementation of the electronic device and the context in which it operates, the required functions, actions, or tasks will differ; correspondingly, the computer programs stored in memory 602 will also differ, and processor 601 can control the electronic device to perform different functions and complete different actions or tasks by executing different computer programs.
[0114] In some alternative embodiments, such as Figure 6 As shown, electronic devices may also include other components such as displays, power supply components, and communication components. Figure 6 The diagram only shows some components and does not mean that the electronic device includes only these components. Figure 6 The components shown are not exhaustive; electronic devices may include other components to meet different application needs. For example, in cases where voice interaction is required, the electronic device may also include an audio component. The specific components that an electronic device may include depend on its product form and are not limited here.
[0115] In this embodiment of the application, when the processor 601 executes the computer program in the memory 602, it is used to implement the above-mentioned knowledge graph construction method based on PLC patent data.
[0116] It should be noted that the specific functions of the processor in the above-mentioned electronic device can be found in the description of the corresponding embodiment in the knowledge graph construction method based on PLC patent data, and will not be repeated here.
[0117] Accordingly, this application also provides a computer program product, including a computer program / instructions, which, when executed by a processor, implement the steps that can be executed by an electronic device in the above method embodiments.
[0118] The communication components in the above embodiments are configured to facilitate wired or wireless communication between the device housing the communication component and other devices. The device housing the communication component can access wireless networks based on communication standards, such as WiFi, 2G, 3G, 4G / LTE, 5G, or combinations thereof. In one exemplary embodiment, the communication component receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In one exemplary embodiment, the communication component further includes a Near Field Communication (NFC) module to facilitate short-range communication. For example, the NFC module may be implemented based on Radio Frequency Identification (RFID), Infrared Data Association (IrDA) technology, Ultra-Wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.
[0119] The display in the above embodiments includes a screen, which may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touchscreen to receive input signals from a user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can sense not only the boundaries of the touch or swipe action, but also the duration and pressure associated with the touch or swipe operation.
[0120] The power supply component in the above embodiments provides power to various components of the device in which the power supply component is located. The power supply component may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to the device in which the power supply component is located.
[0121] The audio component in the above embodiments can be configured to output and / or input audio signals. For example, the audio component includes a microphone (MIC) configured to receive external audio signals when the device containing the audio component is in an operating mode, such as call mode, recording mode, and voice recognition mode. The received audio signals can be further stored in memory or transmitted via a communication component. In some embodiments, the audio component also includes a speaker for outputting audio signals.
[0122] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0123] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0124] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0125] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0126] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0127] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.
[0128] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.
[0129] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0130] The above description is merely an embodiment of this application and is not intended to limit this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principle of this application should be included within the scope of the claims of this application.
Claims
1. A method for constructing a knowledge graph based on PLC patent data, characterized in that, include: First patent data is obtained from multiple patent databases, including structured patent data and unstructured patent data; The structured patent data and the unstructured patent data are standardized respectively to obtain initial standardized patent data; The initial standardized patent data is subjected to cross-data source semantic mapping and unified encoding processing to obtain a standardized patent dataset with global entity identifiers; Based on the entity recognition model, the text data in the standardized patent dataset is subjected to entity annotation, disambiguation and alignment processing; Based on the processing results of the standardized patent dataset and the entity recognition model, a structured knowledge entity database is generated. Based on the structured knowledge entity base, construct the dependency syntax tree corresponding to each text sentence in the text data; Based on the central word of each text sentence and its corresponding dependency syntax tree, the target entity triple is determined; Construct a PLC patent data knowledge graph based on the target entity triples corresponding to all dependency syntax trees.
2. The method according to claim 1, characterized in that, The structured patent data and the unstructured patent data are standardized respectively to obtain initial standardized patent data, including: The structured patent data is subjected to field extraction and standardization to obtain patent basic data with a unified format; The unstructured patent data is subjected to text recognition and converted into structured patent text data; The patent basic data and the patent text data are used as initial standardized patent data; The step of performing text recognition on the unstructured patent data and converting it into structured patent text data includes: Natural language processing and format conversion are performed on the text data in the unstructured patent data to obtain the first text data; Character extraction and format conversion are performed on the graphic data in the unstructured patent data to obtain the second text data; The first text data and the second text data are used as structured patent text data.
3. The method according to claim 2, characterized in that, Based on an entity recognition model, the text data in the standardized patent dataset is subjected to entity annotation, disambiguation, and alignment processing, including: Extract text data from the standardized patent dataset; The text data is segmented into sentences and words to obtain a text sequence containing multiple words; Each word segment in the text sequence is converted into a word vector using a second embedding model; Each word vector is identified and labeled using a bidirectional long short-term memory network model to obtain the entity label score corresponding to each word segmentation. The entity label score corresponding to each word is corrected based on the conditional random field model, and the entity label of the corresponding word is determined.
4. The method according to claim 3, characterized in that, Based on the processing results of the standardized patent dataset and the entity recognition model, a structured knowledge entity library is generated, including: Based on the processing results of the standardized patent dataset and the entity recognition model, the global entity identifier and basic patent data associated with the word segmentation of the labeled entity are determined as the entity attributes of the corresponding word segmentation. A structured knowledge entity library is generated based on the entity tags and entity attributes corresponding to each word segment.
5. The method according to claim 4, characterized in that, Based on the structured knowledge entity base, construct the dependency syntax tree corresponding to each text sentence in the text data, including: The text data is subjected to syntactic analysis to determine the syntactic structure of each text sentence; wherein, the syntactic structure includes each word segment in the corresponding text sentence, its corresponding part of speech, and the dependency relationships between each word segment; Based on the structured knowledge entity base and the syntactic structure of each text sentence, determine the candidate entity triples corresponding to each text sentence; Construct the corresponding dependency syntax tree based on the candidate entity triples corresponding to each text sentence; The step of determining the candidate entity triplet corresponding to each text sentence based on the structured knowledge entity base and the syntactic structure of each text sentence includes: Based on the structured knowledge entity base and PLC standard terminology, construct a PLC terminology ontology base; Based on the PLC terminology ontology and the syntactic structure of each text sentence, candidate entity triples corresponding to each text sentence are determined.
6. The method according to claim 5, characterized in that, Based on the central words of each text sentence and the corresponding dependency syntax tree, the target entity triples are determined, including: Based on the central word of each text sentence and the shortest path traversal of the corresponding dependency syntax tree, the target triplet corresponding to the corresponding dependency syntax tree is determined.
7. The method according to claim 6, characterized in that, Also includes: Second patent data is periodically retrieved from multiple patent databases, and the second patent data is timestamped. The second patent data is filtered based on the PLC terminology ontology to determine new entities and / or new entity dependencies. The PLC patent data knowledge graph is updated based on the newly added entities and / or the newly added entity dependencies; Based on the timestamps corresponding to the newly added entities and / or the new entity dependencies, the entity change frequency and / or entity dependency change characteristics of the PLC patent data knowledge graph are determined for use in PLC domain analysis.
8. A knowledge graph construction device based on PLC patent data, characterized in that, include: The acquisition module is used to acquire first patent data from multiple patent databases, wherein the first patent data includes structured patent data and unstructured patent data; The first processing module is used to standardize the structured patent data and the unstructured patent data respectively to obtain initial standardized patent data. The second processing module is used to perform cross-data source semantic mapping and unified encoding processing on the initial standardized patent data to obtain a standardized patent dataset with global entity identifiers. The third processing module is used to perform entity annotation, disambiguation and alignment processing on the text data in the standardized patent dataset based on the entity recognition model. The generation module is used to generate a structured knowledge entity library based on the processing results of the standardized patent dataset and the entity recognition model. The first construction module is used to construct a dependency syntax tree corresponding to each text sentence in the text data based on the structured knowledge entity library; The determination module is used to determine the target entity triples based on the headword of each text sentence and the corresponding dependency syntax tree; The second construction module is used to construct a PLC patent data knowledge graph based on the target entity triples corresponding to all dependency syntax trees.
9. An electronic device, characterized in that, include: A processor and a memory, wherein, when the processor executes a computer program, the processor is configured to implement the method as described in any one of claims 1-7.
10. A computer program product comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, the method as described in any one of claims 1-7 is implemented.