Method, device and equipment for extracting patent entity relationship, medium and product

By building a corpus and using a multi-module collaborative method of sequence coding, graph convolution network and relationship prediction module, the problem of insufficient efficiency and accuracy of unstructured patent text knowledge extraction in the field of CCUS is solved, and deep feature mining and efficient knowledge extraction are achieved.

CN120218067APending Publication Date: 2025-06-27CHINA THREE GORGES CORPORATION
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510257961.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-05
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The prior art has not in-depth feature extraction and narrow dimensions in entity relationship extraction, resulting in insufficient efficiency and accuracy of unstructured patent text knowledge extraction in the field of CCUS.

Method used

Preprocess the target patent text and build a corpus, including marked and classified text entities and semantic relationships. Then, based on the corpus, semantic information is obtained through the sequence encoding module, and input the graph convolution network encoding module based on the label attention mechanism to obtain the structure dependence information, and finally input the relationship prediction module to obtain the subject-object relationship.

Benefits of technology

Deeply explore hidden features, improve the efficiency and accuracy of unstructured text knowledge extraction, enhance the mining and utilization of semantic information, and provide support for the intelligent construction of patents in the CCUS field.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120218067A_ABST
    Figure CN120218067A_ABST
Patent Text Reader

Abstract

The invention relates to a patent entity relationship extraction method and device, equipment, a medium and a product, in particular to the technical field of natural language processing. Comprising the steps that a target patent text is obtained and preprocessed, a corpus is constructed, and the corpus comprises marked and classified text entities and semantic relations between the text entities; based on a corpus, semantic information is obtained through a sequence coding module; inputting the semantic information into a graph convolutional network coding module based on a label attention mechanism to obtain structure dependence information output by the graph convolutional network coding module based on the label attention mechanism; and inputting the structure dependence information into a relation prediction module to obtain a subject-object relation output by the relation prediction module. The corpus is constructed through preprocessing, hidden features can be deeply mined through multi-module collaboration, and the extraction efficiency and accuracy of unstructured text knowledge are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of natural language processing, and in particular, to a method, apparatus, device, medium, and product for extracting patent entity relationships. Background Art

[0002] In the critical period of coping with climate change, the technology of carbon capture, utilization, and storage (CCUS) is the core means to achieve carbon emission reduction and promote sustainable development. As a knowledge treasure of CCUS technology innovation, the unstructured text in patents contains a large amount of key information. The number of patent applications in the CCUS field in China has been continuously rising. How to mine technical information and associations from these massive texts has become an important direction of patent analysis.

[0003] Traditional patent analysis mainly relies on manual methods, which are time-consuming, inefficient, and costly. The development of computer technology has enabled the automation of massive data processing. In the field of text information processing, entity relationship extraction technology is a key foundation of natural language processing and artificial intelligence, supporting numerous downstream tasks. From early machine learning to current deep learning, named entity recognition technology has been continuously evolving, and deep learning methods have become the mainstream due to their high accuracy.

[0004] However, currently, when extracting entity relationships, the hidden features of sentences cannot be deeply mined, and there is insufficient attention at the feature level, resulting in incomplete extracted features, which severely restricts the efficient and accurate extraction and utilization of knowledge in unstructured patent texts. Therefore, there is an urgent need for new methods to improve the extraction efficiency and accuracy, provide support for constructing a high-quality patent knowledge graph in the CCUS field, realize the effective association of unstructured patent text data, and thus promote the intelligent process of patents in the CCUS field. Summary of the Invention

[0005] In order to solve the above technical problems or at least partially solve the above technical problems, this application provides a method, apparatus, device, medium, and product for extracting patent entity relationships, which can solve the problem that the feature extraction in the existing technology for entity relationship extraction is not in-depth and the dimension is narrow, resulting in insufficient extraction efficiency and accuracy of knowledge in unstructured patent texts in the CCUS field.

[0006] To achieve the above object, the technical solutions provided by the embodiments of this application are as follows:

[0007] In a first aspect, the present application provides a method for extracting patent entity relationships, including: obtaining a target patent text for preprocessing and constructing a corpus, where the corpus includes text entities that have been marked and classified and the semantic relationships between the text entities; based on the corpus, obtaining semantic information through a sequence encoding module; inputting the semantic information into a graph convolutional network encoding module based on a label attention mechanism to obtain the structural dependency information output by the graph convolutional network encoding module based on the label attention mechanism; and inputting the structural dependency information into a relationship prediction module to obtain the subject-object relationship output by the relationship prediction module.

[0008] As an optional implementation provided in an embodiment of the present application, the sequence encoding module includes a MacBERT pre-training model, a global word representation vector GloVe model, and a bidirectional long short-term memory network BiLSTM model; obtaining semantic information through the sequence encoding module based on the corpus includes: based on the corpus, using the MacBERT pre-training model to perform word embedding on the text entities to obtain word labels; using the GloVe model to perform part-of-speech tagging embedding on the part-of-speech information of the text entities to obtain part-of-speech labels, and performing named entity recognition embedding on the relevant information of the text entities to obtain entity recognition labels; concatenating the word labels, part-of-speech labels, and entity recognition labels to obtain a concatenated matrix; and inputting the concatenated matrix into the BiLSTM model to obtain the semantic information output by the BiLSTM model.

[0009] As an optional implementation provided in an embodiment of the present application, inputting the semantic information into a graph convolutional network encoding module based on a label attention mechanism to obtain the structural dependency information output by the graph convolutional network encoding module based on the label attention mechanism includes: the graph convolutional network encoding module based on the label attention mechanism performs the following operations: constructing a dependency syntax analysis tree based on the semantic information; calculating a weighted adjacency matrix between node i and all its neighbor nodes j based on the similarity and distance information between the nodes in the dependency syntax analysis tree; based on the weighted adjacency matrix, performing a weighted sum of the structural dependency information of all neighbor nodes j in the previous layer, then normalizing by dividing by the degree of node i in the dependency syntax analysis tree, and then adding a bias vector, and obtaining the structural dependency information of node i in the current layer through an activation function.

[0010] As an optional implementation provided in an embodiment of the present application, calculating the weighted adjacency matrix between node i and all its neighbor nodes j based on the similarity and distance information between the nodes in the dependency syntax analysis tree includes calculating the weighted adjacency matrix according to the following formula:

[0011]

[0012] where e ij is the attention weight, e ij =softmax(s ij, d ij ), where s ij is the node similarity between node i and node j; d ij is the reciprocal of the distance between node i and node j.

[0013] As an alternative implementation provided by the embodiments of the present application, based on the weighted adjacency matrix, the upper-layer structure dependence information of all neighbor nodes j is weighted and summed, and then normalized by dividing by the degree of node i in the dependency parsing tree, and then a bias vector is added, and the structure dependence information of node i in the current layer is obtained through an activation function, including calculating the structure dependence information of node i in the current layer according to the following formula:

[0014]

[0015] In the formula, l is the current layer, σ is the activation function, n is the number of all neighbor nodes j, is the weighted adjacency matrix between node i in the current layer and all its neighbor nodes j, W (l) is the weight matrix, is the upper-layer structure dependence information of all neighbor nodes j, d i is the degree of node i in the dependency parsing tree, b (l) is the bias vector.

[0016] As an alternative implementation provided by the embodiments of the present application, the relationship prediction module includes a fully connected neural network and a global correspondence module; the structure dependence information is input into the relationship prediction module to obtain the subject-object relationship output by the relationship prediction module, including: using the fully connected neural network to perform sequence labeling on the structure dependence information to obtain subject-object pairs; calculating the relationship probability scores of the subject-object pairs through the global correspondence module; when the relationship probability scores of the subject-object pairs are greater than a preset threshold, outputting the entity relationship of the subject-object pairs.

[0017] In a second aspect, the present application provides an extraction device for patent entity relationships, the device includes:

[0018] A corpus construction component, configured to obtain a target patent text for preprocessing and construct a corpus, the corpus including text entities that have been labeled and classified and semantic relationships between the text entities;

[0019] A sequence encoding component, configured to obtain semantic information based on the corpus through a sequence encoding module;

[0020] A graph convolutional network encoding component, configured to input the semantic information into a graph convolutional network encoding module based on a label attention mechanism to obtain structure dependence information output by the graph convolutional network encoding module based on the label attention mechanism;

[0021] The relationship prediction component is used to input the structural dependency information into the relationship prediction module to obtain the subject-object relationship output by the relationship prediction module.

[0022] In the third aspect, the present application provides an electronic device comprising: a processor, a memory, and a computer program stored in the memory and executable on the processor, wherein the computer program, when executed by the processor, implements the method for extracting patent entity relationships as described in the first aspect or any one of its optional embodiments.

[0023] In a fourth aspect, the present application provides a computer-readable storage medium, comprising: a computer program is stored on the computer-readable storage medium, and when the computer program is executed by a processor, the method for extracting patent entity relationships as described in the first aspect or any one of its optional embodiments is implemented.

[0024] In the fifth aspect, the present application provides a computer program product, including: the computer program product includes a computer program, when the computer program runs on a computer, the computer implements the method for extracting patent entity relationships as described in the first aspect or any one of its optional embodiments.

[0025] Compared with the prior art, the technical solution provided by the embodiments of the present application has the following advantages:

[0026] The embodiment of the present application provides a method, device, equipment, medium and product for extracting patent entity relationships, wherein the method first obtains the target patent text for preprocessing and constructs a corpus, which includes text entities that have been marked and classified and semantic relationships between text entities; then, semantic information is obtained through a sequence encoding module based on the corpus, and then the semantic information is input into a graph convolutional network encoding module based on a label attention mechanism to obtain structural dependency information output by the graph convolutional network encoding module based on a label attention mechanism, and the structural dependency information is further input into a relationship prediction module to obtain a subject-object relationship output by the relationship prediction module. In this way, the present application constructs a corpus through preprocessing, and through the collaboration of multiple modules, it can deeply mine hidden features, improve the efficiency and accuracy of extracting unstructured text knowledge, enhance the mining and utilization of semantic information, provide strong support for the intelligent construction of patents in the CCUS field and realize text data association. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.

[0028] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the accompanying drawings required in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0029] Figure 1A It is a schematic flowchart of a method for extracting patent entity relationships provided by an embodiment of the present application;

[0030] Figure 1B It is a schematic flowchart of multi-module collaboration provided by an embodiment of the present application;

[0031] Figure 2 It is a schematic structural diagram of an apparatus for extracting patent entity relationships provided by an embodiment of the present application;

[0032] Figure 3 It is a schematic structural diagram of an electronic device according to an embodiment of the present application. Detailed implementation manners

[0033] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the technical terms required in the description of the embodiments or the prior art.

[0034] CCUS refers to the general term for a series of technologies and processes that separate carbon dioxide from industrial processes, energy utilization, or the atmosphere, capture, transport, and utilize or store it to achieve carbon dioxide emission reduction or resource utilization. Carbon capture refers to the process of separating and enriching carbon dioxide from industrial emission sources (such as thermal power plants, steel mills, cement plants, etc.) or the atmosphere using physical, chemical, or biological methods. Carbon utilization is the process of converting the captured carbon dioxide into valuable products or energy through chemical, biological, or physical means, such as converting carbon dioxide into chemicals, fuels, building materials, etc., to achieve the resource utilization of carbon dioxide and reduce its emission into the atmosphere. Carbon sequestration is to transport the captured carbon dioxide to a suitable geological or marine environment for long-term storage, isolating it from the atmosphere, so as to achieve the purpose of reducing the concentration of carbon dioxide in the atmosphere. Common sequestration methods include geological sequestration and marine sequestration, etc.

[0035] BIO tagging is a method used to tag entities in text in natural language processing (NLP) tasks. Each word in the text is tagged with three labels, B, I, and O, to represent the relationship between the word and the named entity. Among them, "B" indicates the start position of the named entity, "I" indicates the position inside the named entity, and "O" indicates that the word does not belong to any named entity. This tagging method can structure the entity information in the text, making it easier for computers to understand and process the text semantics, and providing basic support for subsequent natural language processing tasks.

[0036] The MacBERT pre-trained model is a pre-trained language model improved based on the Bidirectional Encoder Representations from Transformers (BERT). Like BERT, MacBERT also adopts the encoder architecture of the Transformer, consisting of multiple stacked Transformer blocks, which can process text sequences in parallel, effectively capture the long-range dependencies in the text, and perform in-depth semantic representation learning on the text. The main innovation of MacBERT lies in the adoption of the Macaron Mask strategy. During the pre-training process, this strategy performs a more refined masking operation on the tokens in the text, considering not only the conventional random masking but also introducing some additional masking methods, enabling the model to better learn the semantic information of the text and enhancing the model's understanding and generalization ability of language knowledge.

[0037] The Bidirectional Long Short-Term Memory (BiLSTM) model is a deep learning model used to process sequence data. The BiLSTM model consists of two LSTM layers in opposite directions. One is the forward LSTM layer, which processes the input data in the normal order of the sequence, reading information from the beginning to the end of the sequence; the other is the backward LSTM layer, which processes the input data in the reverse order of the sequence, reading information from the end to the beginning of the sequence. These two directions of LSTM layers can capture the context information in different directions of the input sequence, and then the outputs of these two directions are combined, usually through simple concatenation or summation operations, to obtain the final output result.

[0038] The LayerNorm function is a technique used in deep learning to normalize the inputs of each layer of a neural network.

[0039] Global Vectors for Word Representation (GloVe) is an unsupervised learning algorithm for word vector representation. It aims to map words into a low-dimensional vector space to capture the semantic and syntactic relationships between words. The GloVe model learns word vectors based on the global word co-occurrence matrix. It utilizes the co-occurrence information of words in the corpus and obtains the low-dimensional vector representation of each word by performing operations such as decomposing the co-occurrence matrix. Specifically, it believes that the semantic similarity between two words can be reflected by the frequency of their co-occurrence in the corpus. The higher the co-occurrence frequency, the more similar the semantics of the two words, and the closer their distance in the vector space.

[0040] The Graph Convolutional Network encoding module is constructed based on the Graph Convolutional Network, which is a neural network specifically designed for processing graph data. Graph data consists of nodes (vertices) and edges connecting the nodes, and can be used to represent various complex relationships. The core function of this module is to encode the node features in the graph. By performing convolutional operations on the graph structure, it aggregates the information of nodes and their neighbor nodes to learn more expressive node representations.

[0041] The Dependency Parse Tree is a tree-shaped data structure used in the field of natural language processing to clearly display the syntactic structure of a sentence. Each word in the sentence corresponds to a node in the tree, and the nodes are connected by edges. These edges represent the dependency relationships between words and have a clear direction, pointing from the dependent word to the governor. The entire tree structure presents a hierarchical relationship, reflecting the syntactic hierarchy and semantic dependency relationships among the components of the sentence.

[0042] A virtual edge is an edge that does not actually exist in the original data structure but is artificially created according to specific rules, algorithms, or requirements. It is mainly used in the data processing process to better represent and mine the potential relationships, dependency information, or semantic connections between data elements, providing richer structural information for subsequent analysis and model processing. In dependency parsing, virtual edges can be used to connect words that do not have a direct dependency relationship in traditional grammar analysis but have important connections in semantic understanding, helping to more comprehensively analyze the syntactic structure of a sentence.

[0043] The Label Attention Mechanism (LAM) is a mechanism that can automatically learn the importance weights of different positions or elements for specific labels or targets when processing sequence data or structured data. It calculates the degree of association between the input data and the label and assigns an attention weight to each data element, allowing the model to pay more attention to important information related to the label, thereby improving the model's ability to predict or understand the label.

[0044] In the context of the global response to climate change, carbon capture, utilization and storage (CCUS) technology has attracted widespread attention from the international community as a core means to achieve carbon reduction targets and promote sustainable development. Patents, as key knowledge carriers of scientific and technological innovation results, record in detail many scientific and technological details of CCUS technology. The number of patent applications in the CCUS field in my country has shown a significant growth trend year by year. In this context, how to accurately mine key technical information and the inherent correlation between technologies from massive and complex unstructured patent texts has become a research focus and hotspot in the current field of patent analysis.

[0045] In the past, patent analysis was mainly completed by manual operations. This traditional method has many disadvantages, such as lengthy analysis cycles, low work efficiency, and high manpower and time costs. With the rapid development of modern science and technology, using computers to store, process, and analyze massive amounts of data has become a mainstream trend in the industry. In the field of text information processing, entity relationship extraction technology plays a key role as an important basic technology in the fields of natural language processing and artificial intelligence. It can effectively solve the problem of identifying and extracting specific words in texts, and provide strong support for many natural language processing downstream tasks such as syntactic analysis, text classification, automatic question answering, and knowledge graph construction. In the early days, machine learning technology laid the foundation for named entity recognition technology, and now, deep learning-based recognition methods have become the mainstream technology in this field with their higher accuracy.

[0046] However, the current entity relationship extraction technology has obvious deficiencies in the feature extraction link. It fails to fully explore the hidden features in the deep level of the sentence, and focuses on a relatively narrow feature dimension, which makes it difficult for the extracted features to fully reflect the text information. This defect seriously restricts the efficient and accurate extraction of unstructured or semi-structured patent text knowledge, which in turn affects the effective use of technical texts. Therefore, there is an urgent need for a new technical solution to improve the efficiency and accuracy of knowledge extraction. By using knowledge graph technology, the details and semantic information of patent texts are deeply mined and clearly presented, providing a solid theoretical and technical guarantee for the construction of high-quality and efficient patent knowledge graphs in the CCUS field, realizing the organic connection of unstructured patent text data, and fully promoting the process of patent intelligence construction in the CCUS field.

[0047] In order to solve some or all of the technical problems existing in the related art, the embodiment of the present application provides a method, device, equipment, medium and product for extracting patent entity relationships, wherein the method first obtains the target patent text for preprocessing and constructs a corpus, which includes text entities that have been marked and classified and semantic relationships between text entities; then, semantic information is obtained through a sequence encoding module based on the corpus, and then the semantic information is input into a graph convolutional network encoding module based on a label attention mechanism to obtain structural dependency information output by the graph convolutional network encoding module based on a label attention mechanism, and the structural dependency information is further input into a relationship prediction module to obtain a subject-object relationship output by the relationship prediction module. In this way, the present application constructs a corpus through preprocessing, and through the collaboration of multiple modules, it can deeply mine hidden features, improve the efficiency and accuracy of unstructured text knowledge extraction, enhance the mining and utilization of semantic information, and provide strong support for the intelligent construction of patents in the CCUS field and realize text data association.

[0048] In order to more clearly understand the above-mentioned purposes, features and advantages of the present application, the scheme of the present application will be further described below. It should be noted that the embodiments of the present application and the features in the embodiments can be combined with each other without conflict.

[0049] In the following description, many specific details are set forth to facilitate a full understanding of the present application, but the present application may also be implemented in other ways different from those described herein; obviously, the embodiments in the specification are only part of the embodiments of the present application, rather than all of the embodiments.

[0050] A method for extracting patent entity relationships provided in the embodiments of the present application can be implemented by an extraction device or electronic device for patent entity relationships, and the electronic device includes but is not limited to a personal computer, a laptop computer, a tablet computer, a smart phone, etc. The operating system of the electronic device may include Android, a mobile operating system (iOS) developed by Apple, an operating system (Windows) developed by Microsoft Corporation of the United States, etc., and the embodiments of the present application are not limited to this. The electronic device can be operated alone to implement the present application, or it can be connected to a network and implemented by interactive operations with other computer devices in the network. Among them, the network in which the electronic device is located includes but is not limited to the Internet, a wide area network, a metropolitan area network, a local area network, a virtual private network (VPN) network, etc.

[0051] It should be noted that the protection scope of the method for extracting patent entity relationships described in the embodiment of the present application is not limited to the execution order of the steps listed in the present embodiment. All solutions implemented by adding, reducing or replacing steps in the prior art based on the principles of the present application are included in the protection scope of the present application.

[0052] As Figure 1A shown Figure 1A is a schematic flowchart of a method for extracting patent entity relationships according to an embodiment of the present application. This method can be executed by an extraction device for patent entity relationships, where the device can be implemented by software and / or hardware and is generally integrated in an electronic device. The method mainly includes the following steps S101 to S104:

[0053] S101. Obtain the target patent text for preprocessing and construct a corpus.

[0054] In some embodiments, patents in the patent database are used as data sources, and according to the target retrieval subject terms, the target patent text is retrieved. Patent texts related to the CCUS field. Exemplarily, the target retrieval subject terms are CCUS + carbon capture + carbon utilization + carbon sequestration. The target patent text includes information such as the patent title, abstract, applicant, application date, legal status, etc.

[0055] In some embodiments, preprocessing is performed on the target patent text (especially the abstract), including but not limited to: data cleaning, corpus classification, boundary clarification, and BIO annotation.

[0056] Among them, data cleaning includes: importing the target patent text into a data processing environment, using technologies such as hash algorithms and data fingerprints to identify and delete duplicate records in the dataset; dealing with missing values in the data by methods such as deleting records containing missing values, filling with mean / median / mode, and predicting and filling based on machine learning algorithms; removing noises such as irrelevant characters, garbled codes, and special symbols in the data; performing standardization processing on the data, such as unifying the text format and converting the data into a specific data type; and, by setting data quality indicators and rules, checking whether the data meets the requirements.

[0057] Corpus classification includes: collecting various types of corpus data and performing preprocessing operations such as cleaning and word segmentation to prepare for subsequent classification; extracting features from the preprocessed corpus and converting the text into a vector form that can be processed by a computer; selecting a suitable classification algorithm according to the characteristics of the corpus and the task requirements; using the labeled corpus data as a training set, inputting the extracted features into the selected classification model for training, and adjusting the parameters of the model so that it can learn the feature patterns of different categories of corpus; using the test set to evaluate the trained model, calculating evaluation indicators to measure the classification performance of the model; optimizing and adjusting the model according to the evaluation results; applying the trained model to the actual corpus classification task, classifying and predicting new unlabeled corpus, and monitoring and maintaining according to the actual application situation.

[0058] Boundary clarification includes: removing noise, standardizing formats, and segmenting text to make the text easier to process; extracting features such as character level, word level, and syntactic structure to provide a basis for boundary identification; determining the boundaries of the text based on the extracted features; selecting evaluation indicators to measure the effectiveness of boundary clarification, adjusting the model or method based on the evaluation results, and conducting manual review and correction when necessary.

[0059] BIO annotation includes: collecting text data that needs to be annotated and performing necessary preprocessing; determining the label system for BIO annotation according to specific tasks and needs, clarifying the meaning and scope of each label, B represents the beginning of the entity, I represents the middle and end of the entity, and O represents the non-entity part; manually or using annotation tools to annotate the text according to the annotation scheme, marking each word or character with a corresponding BIO label; performing a quality check on the annotated data, checking whether the annotation is accurate and complete, and whether it complies with the annotation specifications through manual sampling, automatic inspection, etc.; making corrections and adjustments in a timely manner for problems found during the inspection to ensure the quality of the annotated data; storing the annotated data and establishing a suitable data management mechanism to facilitate subsequent data use, query and maintenance.

[0060] The corpus (CCUS_dataset) is constructed using the preprocessed target patent text. The corpus contains labeled and classified text entities and semantic relationships between text entities. Specifically, various text entities are identified from the preprocessed target patent text, and the semantic relationships between text entities are extracted; then the semantic relationships between text entities are merged and cleaned, and then stored in the corpus according to the preset format and structure.

[0061] The above embodiment obtains the target patent text for preprocessing and constructs a corpus, which covers marked and classified text entities and semantic relationships. This provides a rich and valuable data foundation for subsequent model training. In the process of obtaining semantic information through a sequence encoding module based on the corpus, the model can learn more comprehensive and in-depth semantic features from a large amount of diverse text data. Compared with the traditional entity relationship extraction technology that only focuses on surface features, this method can dig out potential semantic associations and hidden features in the text, such as some indirectly expressed technical relationships, implicit causal relationships, etc., thereby comprehensively improving the breadth and depth of feature extraction. The target patent text is preprocessed and a corpus is constructed. This process can convert the original unstructured or semi-structured patent text into an ordered data form that can be understood by the model. In the subsequent processing flow, the model learns and predicts based on the corpus, avoiding the disordered processing of the original patent text and greatly improving the processing efficiency. For example, when facing a large amount of patent text, there is no need to analyze the text structure from scratch every time, but directly use the organized information in the corpus to quickly locate and extract key knowledge.

[0062] S102. Obtain semantic information through a sequence encoding module based on an entity relationship corpus.

[0063] In some embodiments, the sequence encoding module includes a MacBERT pre-trained model, a global word representation vector GloVe model, and a bidirectional long short-term memory network BiLSTM model. When performing step S102, based on the corpus, the MacBERT pre-trained model is used to perform word embedding on the text entity to obtain word tags; the GloVe model is used to perform part-of-speech tagging embedding on the part-of-speech information of the text entity to obtain part-of-speech tags, and perform named entity recognition embedding on the relevant information of the text entity to obtain entity recognition tags; the word tags, part-of-speech tags, and entity recognition tags are concatenated to obtain a concatenated matrix; the concatenated matrix is input into the BiLSTM model to obtain the semantic information output by the BiLSTM model.

[0064] Specifically, first, based on the corpus, the MacBERT pre-trained model is used to perform word embedding on the text entity to obtain word tags. The input sequence is x = [x1, x2, …, x n , where x i is the i-th word of the sequence. The MacBERT pre-trained model is used to transform the input sequence x into an input encoding e = [e1, e2, …, e n , where e i represents the embedding encoding of the i-th word. Then, the sequence representation is obtained through the forward pass calculation of the MacBERT pre-trained model, that is, h = [h1, h2, …, hn]. Furthermore, the LayerNorm function is used for normalization processing to obtain the final word tag (word).

[0065] The normalization processing using the LayerNorm function is shown in the following formula (1):

[0066]

[0067] In formula (1), α i is a normalization weight, as shown in formula (2):

[0068]

[0069] In formula (2), e c represents a learnable vector, which is used as a query vector to calculate the attention weight of each word.

[0070] Then, the GloVe model is used to perform part-of-speech tagging embedding on the part-of-speech information of the text entity to obtain part-of-speech tags (pos), and perform named entity recognition embedding on the relevant information of the text entity to obtain entity recognition tags (ner). As shown in the following formulas (3) and (4):

[0071]

[0072] In formula (3), represents the pos embedding of the i-th word. In formula (4), represents the ner embedding of the i-th word.

[0073] Concatenate the word tag, part-of-speech tag, and entity recognition tag to obtain the concatenated matrix h (0) ∈R n×m , as shown in the following formulas (5) and (6):

[0074]

[0075] m = d word + d pos + d ner (6)

[0076] where d word , d pos , d ner represent the dimensions of the word tag, part-of-speech tag, and entity recognition tag respectively.

[0077] Since the information contained in this matrix is incomplete, the BiLSTM layer is combined to read the sequence in both the forward and backward directions, and the state vector is used to capture the context information to obtain the encoded representation h t (0) of the entire sequence, that is, the semantic information, which is used as the input for the subsequent model.

[0078] The above-mentioned embodiment extracts semantic information from the corpus by the sequence encoding module, which can deeply understand the meaning of each word and sentence in the patent text, as well as the semantic relationship between them, providing a basis for subsequent analysis. For example, for a patent text describing a complex technical principle, this module can accurately grasp the logical relationship between each step in the technical process, the functional connection between different technical components, and other semantic information, rather than just staying at the surface word recognition.

[0079] S103. Input the semantic information into the graph convolutional network encoding module based on the label attention mechanism to obtain the structural dependency information output by the graph convolutional network encoding module based on the label attention mechanism.

[0080] The graph convolutional network encoding module based on the label attention mechanism combines the dependency parsing tree and the graph convolutional network to obtain a more comprehensive context representation by fusing the syntactic structure of the sentence and the relationship information between entities.

[0081] In some embodiments, the graph convolutional network encoding module based on the label attention mechanism performs the following operations: constructing a dependency syntactic analysis tree based on semantic information; calculating a weighted adjacency matrix between node i and all its neighbor nodes j based on the similarity and distance information between the nodes in the dependency syntactic analysis tree; based on the weighted adjacency matrix, performing a weighted sum of the upper-layer structure dependency information of all neighbor nodes j, then normalizing by dividing by the degree of node i in the dependency syntactic analysis tree, and then adding a bias vector, and obtaining the structure dependency information of node i in the current layer through an activation function.

[0082] Among them, the dependency syntactic analysis tree shows the grammatical and semantic relationships between text segments. Each node represents a word, and the edge represents a dependency relationship. By adding virtual edges to construct LAM, it is used to obtain the dependency relationships of the k-order neighborhood. For the weighted adjacency matrix of the weight values in LAM, the similarity and distance between nodes are used to represent.

[0083] Weighted adjacency matrix It is used to describe the connection weight between node i and node j in the dependency syntactic analysis tree, and it plays a role in controlling the intensity of information transmission between nodes in the calculation of the l-th layer. As shown in formula (7):

[0084]

[0085] In formula (7), e ij is the attention weight, e ij = softmax(s ij , d ij ), where s ij is the node similarity between node i and node j in the dependency syntactic analysis tree, x i and x j are the feature vectors of node i and node j in the dependency syntactic analysis tree respectively; d ij is the reciprocal of the distance between node i and node j, e is the Euler number, and d is the distance between node i and node j. The softmax operation is used to combine the two for normalization to obtain the attention weight e ij . N i represents the set of neighbor nodes of node i. Combining the node similarity and distance as the weight value, and combining it with the logical adjacency matrix to obtain the weighted adjacency matrix, so as to ensure that the information between nodes can be fully captured and the differences between different nodes are retained. Among them, the logical adjacency matrix is usually a binary matrix representing whether there is a connection between nodes (the elements are usually 0 or 1, 1 means there is a connection, and 0 means there is no connection).

[0086] The constructed weighted adjacency matrix is initially obtained from the initialized adjacency matrix, and all element values of the initialized adjacency matrix are set to 0. After calculating the weights by the attention mechanism, finally, the logical adjacency matrix is convolved with the node features to obtain the final representation: the sentence representation based on specific relationships {q1, q2, q3, …, q n}. {q1, q2, q3, …, q n} can be regarded as the new feature representation obtained after each part of the sentence (corresponding to the nodes in the graph) considers the specific relationships between nodes reflected by the weighted adjacency matrix and the logical adjacency matrix and participates in the convolution of the node features.

[0087] Based on the weighted adjacency matrix, the weighted sum of the upper-layer structural dependency information of all neighbor nodes j is calculated, and then it is normalized by dividing by the degree of node i in the dependency parse tree, and then the bias vector is added. The structural dependency information of node i in the current layer is obtained through the activation function, as shown in formula (8):

[0088]

[0089] Among them, l is the current layer, σ represents the activation function, which introduces non-linearity into the model so that the model can learn more complex functional relationships. n is the number of all neighbor nodes j. is the weighted adjacency matrix between node i in the current layer l and all its neighbor nodes j. W (l) is the weight matrix, which is used to perform a linear transformation on the hidden state of node j in the (l - 1)-th layer, adjust the feature representation of the information, so as to better participate in the calculation of the hidden state of node i in the current layer. d i is the degree of node i in the dependency parse tree, that is, the number of edges directly connected to node i; b (l) is the bias vector.

[0090] The above embodiment uses the graph convolutional network encoding module based on the label attention mechanism, comprehensively considering semantic information and the node relationships in the graph structure. The label attention mechanism enables the model to dynamically focus on the information corresponding to different labels (i.e., different types of entities and relationships) when processing data, and is no longer limited to the narrow feature level. For example, when processing complex technical descriptions in patent texts, it can simultaneously focus on the features corresponding to various types of entities such as technical terms, related materials, application scenarios, etc. and their relationships, effectively solving the problem of too narrow feature focus in traditional methods.

[0091] S104. Input the structural dependency information into the relationship prediction module to obtain the subject-object relationship output by the relationship prediction module.

[0092] In some embodiments, the relationship prediction module includes a fully connected neural network and a global correspondence module; when performing step S104, the fully connected neural network is used to perform sequence labeling on the structural dependency information to obtain subject-object pairs; the global correspondence module calculates the relationship probability scores of the subject-object pairs; when the relationship probability score of a subject-object pair is greater than a preset threshold, the entity relationship of the subject-object pair is output.

[0093] In the sentence, there are phenomena of intersection, overlap, and nested structures between the subject and the object. To solve this problem of subject-object overlap, the present application first uses two sequence labeling operations to separately extract the subject (sub) and the object (obj).

[0094] The labeling operation formulas are shown in formulas (9) and (10):

[0095]

[0096] Among them, W sub , W obj are the trainable weights of the subject and the object respectively, h i ∈R d×1 is the encoded representation of the i-th word, u j ∈R d×1 is the j-th relationship representation in the trainable embedding matrix U∈R d×nr ( nr is the size of the full relationship set).

[0097] In the embodiments of the present application, to ensure the stability of the model, a fully connected neural network is used to implement the sequence labeling task of the subject and the object. After sequence labeling, all possible subjects and objects in the sentence relationship are obtained.

[0098] However, the initially identified entity pairs may have problems of errors or redundancy. Therefore, to further improve the accuracy of entity relationship triple extraction, the global correspondence module can be used to determine the correct subject and object pairs among them. Specifically, the learning of the global correspondence matrix in the global correspondence module exists independently of the relationship and can be performed in parallel with the prediction of potential relationships. For all possible entity pairs obtained, the global correspondence module calculates the probability score P of each pair of subject and object. When the P value is greater than the preset threshold λ2, the entity relationship pair is retained, otherwise it is filtered out, as shown in formula (11):

[0099]

[0100] Among them, σ is the Sigmoid function; W g is the trainable weight; are the encoded representations of the i-th and j-th tokens in the input sentence respectively, representing potential subject and object pairs; b gDenote the bias vector.

[0101] As Figure 1B shown in the above embodiments, the multi-module collaboration is utilized. The sequence encoding module is used to obtain semantic information, providing a basis for subsequent analysis. The graph convolutional network encoding module based on the label attention mechanism further mines the structural dependency information, considering the structural relationships between entities in the text. The relationship prediction module finally outputs the subject-object relationship. This way of multi-module collaboration significantly improves the accuracy of knowledge extraction by gradually refining and deepening the understanding of the text. Taking the extraction of technology application relationships from patent texts as an example, through the collaborative effect of each module, the technology subject (such as a certain invention) and the application object (such as a specific application field) can be accurately identified, reducing the situation of incorrect extraction.

[0102] In summary, the embodiment of the present application provides a method for extracting patent entity relationships. The method first obtains the target patent text for preprocessing and constructs a corpus, which includes text entities that have been labeled and classified and the semantic relationships between the text entities. Then, based on the corpus, semantic information is obtained through the sequence encoding module, and then the semantic information is input into the graph convolutional network encoding module based on the label attention mechanism to obtain the structural dependency information output by the graph convolutional network encoding module based on the label attention mechanism. Further, the structural dependency information is input into the relationship prediction module to obtain the subject-object relationship output by the relationship prediction module. In this way, through preprocessing to construct a corpus and through multi-module collaboration, the present application can deeply mine hidden features, improve the efficiency and accuracy of unstructured text knowledge extraction, enhance the mining and utilization of semantic information, provide strong support for the intelligent construction of patents in the CCUS field, and realize the association of text data.

[0103] The method for extracting patent entity relationships provides a key theoretical and technical basis for constructing a high-quality and high-efficiency patent knowledge graph. Utilizing the structural dependency information output by the graph convolutional network encoding module based on the label attention mechanism and combining it with the subject-object relationship obtained by the relationship prediction module helps to construct a patent knowledge graph. The knowledge graph can present the details and semantic information in the patent text in an intuitive graph structure form, integrating the entities and relationships scattered throughout the text to achieve the efficient utilization of the semantic information in the patent text. For example, in the patent analysis of the CCUS field, through the knowledge graph, the relationships between different carbon capture technologies, related equipment, application scenarios, etc. can be clearly shown, facilitating researchers to quickly understand the overall technology in this field.

[0104] Through a series of operations such as preprocessing unstructured patent texts, extracting entity relationships, and constructing knowledge graphs, the effective association between unstructured patent text data is achieved. This enables the originally isolated patent text information to be interconnected, forming an organic knowledge network. In the field of CCUS, it helps to integrate the technical information scattered in various patents, promote cross-patent technology integration and innovation, accelerate the process of patent intelligence construction in this field, and drive the further development and application of CCUS technology.

[0105] As Figure 2 shown, Figure 2 FIG. is a schematic structural diagram of an extraction device for patent entity relationships provided by an embodiment of the present application. The device includes:

[0106] A corpus construction component 201, configured to obtain a target patent text for preprocessing and construct a corpus, where the corpus includes text entities that have been labeled and classified and the semantic relationships between the text entities;

[0107] A sequence encoding component 202, configured to obtain semantic information through a sequence encoding module based on the corpus;

[0108] A graph convolutional network encoding component 203, configured to input the semantic information into a graph convolutional network encoding module based on a label attention mechanism to obtain the structural dependency information output by the graph convolutional network encoding module based on the label attention mechanism;

[0109] A relationship prediction component 204, configured to input the structural dependency information into a relationship prediction module to obtain the subject-object relationship output by the relationship prediction module.

[0110] As an optional implementation manner provided by an embodiment of the present application, the sequence encoding module includes a MacBERT pre-trained model, a global word representation vector GloVe model, and a bidirectional long short-term memory network BiLSTM model; the sequence encoding component 202 is specifically configured to: based on the corpus, perform word embedding on the text entity using the MacBERT pre-trained model to obtain word labels; perform part-of-speech tagging embedding on the part-of-speech information of the text entity using the GloVe model to obtain part-of-speech labels, and perform named entity recognition embedding on the relevant information of the text entity to obtain entity recognition labels; splice the word labels, part-of-speech labels, and entity recognition labels to obtain a splicing matrix; input the splicing matrix into the BiLSTM model to obtain the semantic information output by the BiLSTM model.

[0111] As an alternative implementation provided by the embodiments of the present application, the graph convolutional network encoding component 203 is specifically configured to: the graph convolutional network encoding module based on the label attention mechanism performs the following operations: construct a dependency syntactic analysis tree based on semantic information; calculate the weighted adjacency matrix between node i and all its neighbor nodes j based on the similarity and distance information between the nodes in the dependency syntactic analysis tree; based on the weighted adjacency matrix, perform a weighted sum of the upper-layer structure dependency information of all neighbor nodes j, then normalize it by dividing by the degree of node i in the dependency syntactic analysis tree, and then add a bias vector, and obtain the structure dependency information of node i in the current layer through an activation function.

[0112] As an alternative implementation provided by the embodiments of the present application, the graph convolutional network encoding component 203 calculates the weighted adjacency matrix between node i and all its neighbor nodes j based on the similarity and distance information between the nodes in the dependency syntactic analysis tree, and is specifically configured to calculate the weighted adjacency matrix according to the following formula:

[0113]

[0114] In the formula, e ij is the attention weight, e ij = softmax(s ij , d ij ), where s ij is the node similarity between node i and node j; d ij is the reciprocal of the distance between node i and node j.

[0115] As an alternative implementation provided by the embodiments of the present application, the graph convolutional network encoding component 203 performs a weighted sum of the upper-layer structure dependency information of all neighbor nodes j based on the weighted adjacency matrix, then normalizes it by dividing by the degree of node i in the dependency syntactic analysis tree, and then adds a bias vector, and obtains the structure dependency information of node i in the current layer through an activation function, and is specifically configured to calculate the structure dependency information of node i in the current layer according to the following formula:

[0116]

[0117] In the formula, l is the current layer, σ is the activation function, n is the number of all neighbor nodes j, is the weighted adjacency matrix between node i in the current layer and all its neighbor nodes j, W (l) is the weight matrix, is the upper-layer structure dependency information of all neighbor nodes j, d i is the degree of node i in the dependency syntactic analysis tree, b (l) is the bias vector.

[0118] As an alternative implementation provided by the embodiments of the present application, the relationship prediction module includes a fully connected neural network and a global correspondence module; the relationship prediction component 204 is specifically configured to: perform sequence labeling on the structural dependency information by using the fully connected neural network to obtain subject-object pairs; calculate the relationship probability scores of the subject-object pairs through the global correspondence module; and output the entity relationship of the subject-object pairs when the relationship probability scores of the subject-object pairs are greater than a preset threshold.

[0119] For the specific limitations on the device for extracting patent entity relationships, reference can be made to the limitations on the method for extracting patent entity relationships in the foregoing text, which will not be elaborated herein. Each module in the above device for extracting patent entity relationships can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor of the computer device in hardware form or independent thereof, or stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to the above modules.

[0120] In one embodiment, the present application provides an electronic device, which can be a terminal, and its internal structure diagram can be as Figure 3 shown. The electronic device includes a processor, a memory, a communication interface, a display screen, and an input device connected through a system bus. Among them, the processor of the electronic device is used to provide computing and control capabilities. The memory of the electronic device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of the electronic device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be implemented through WIFI, a carrier network, near field communication (NFC), or other technologies. When the computer program is executed by the processor, it implements a method for detecting lags. The display screen of the electronic device can be a liquid crystal display screen or an electronic ink display screen, and the input device of the electronic device can be a touch layer covered on the display screen, or a button, a trackball, or a touchpad provided on the housing of the electronic device, or an external keyboard, a touchpad, or a mouse, etc.

[0121] Those skilled in the art can understand that Figure 3 the structure shown in

[0122] In one embodiment, the device for extracting patent entity relationships provided by the present application can be implemented in the form of a computer program, and the computer program can be in Figure 3The memory of the electronic device may store various program modules constituting the extraction device of the patent entity relationship, for example, Figure 2 The corpus construction component 201, the sequence encoding component 202, the graph convolution network encoding component 203 and the relationship prediction component 204 are shown. The computer program composed of various program modules enables the processor to execute the steps of the patent entity relationship extraction method of each embodiment of the present application described in this specification.

[0123] For example, Figure 3 The electronic device shown can be Figure 2 The corpus construction component 201 in the patent entity relationship extraction device shown obtains the target patent text for preprocessing and constructs a corpus, which includes text entities that have been marked and classified and semantic relationships between text entities; the electronic device can execute based on the corpus through the sequence encoding component 202, and obtain semantic information through the sequence encoding module; the electronic device can execute through the graph convolutional network encoding component 203 to input the semantic information into the graph convolutional network encoding module based on the label attention mechanism, and obtain the structural dependency information output by the graph convolutional network encoding module based on the label attention mechanism; the electronic device can execute through the relationship prediction component 204 to input the structural dependency information into the relationship prediction module, and obtain the subject-object relationship output by the relationship prediction module.

[0124] In one embodiment, the present application provides an electronic device, including a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:

[0125] The target patent text is obtained for preprocessing, and a corpus is constructed. The corpus includes marked and classified text entities and semantic relations between text entities; based on the corpus, semantic information is obtained through a sequence encoding module; the semantic information is input into a graph convolutional network encoding module based on a label attention mechanism to obtain structural dependency information output by the graph convolutional network encoding module based on a label attention mechanism; the structural dependency information is input into a relationship prediction module to obtain a subject-object relationship output by the relationship prediction module.

[0126] In one embodiment, when the computer program executes the computer program, the following steps are further implemented: The sequence encoding module includes a MacBERT pre-trained model, a global word representation vector GloVe model, and a bidirectional long short-term memory network BiLSTM model; Based on the corpus, semantic information is obtained through the sequence encoding module, including: Based on the corpus, the MacBERT pre-trained model is used to perform word embedding on text entities to obtain word tags; The GloVe model is used to perform part-of-speech tagging embedding on the part-of-speech information of text entities to obtain part-of-speech tags, and to perform named entity recognition embedding on the relevant information of text entities to obtain entity recognition tags; The word tags, part-of-speech tags, and entity recognition tags are concatenated to obtain a concatenated matrix; The concatenated matrix is input into the BiLSTM model to obtain the semantic information output by the BiLSTM model.

[0127] In one embodiment, when the computer program executes the computer program, the following steps are further implemented: The semantic information is input into a graph convolutional network encoding module based on a label attention mechanism to obtain the structural dependency information output by the graph convolutional network encoding module based on the label attention mechanism, including: The graph convolutional network encoding module based on the label attention mechanism performs the following operations: A dependency parsing tree is constructed based on the semantic information; Based on the similarity and distance information between nodes in the dependency parsing tree, a weighted adjacency matrix between node i and all its neighbor nodes j is calculated; Based on the weighted adjacency matrix, the structural dependency information of the previous layer of all neighbor nodes j is weighted and summed, and then normalized by dividing by the degree of node i in the dependency parsing tree, and then a bias vector is added, and the structural dependency information of node i in the current layer is obtained through an activation function.

[0128] In one embodiment, when the computer program executes the computer program, the following steps are further implemented: Based on the similarity and distance information between nodes in the dependency parsing tree, a weighted adjacency matrix between node i and all its neighbor nodes j is calculated, including calculating the weighted adjacency matrix according to the following formula:

[0129]

[0130] In the formula, e ij is the attention weight, e ij = softmax(s ij , d ij ), where s ij is the node similarity between node i and node j; d ij is the reciprocal of the distance between node i and node j.

[0131] In one embodiment, when the computer program executes the computer program, the following steps are further implemented: based on the weighted adjacency matrix, weighted summing is performed on the previous layer structural dependency information of all neighbor nodes j, and then normalizing by dividing by the degree of node i in the dependency syntactic analysis tree, and then adding the bias vector, and obtaining the structural dependency information of node i in the current layer through the activation function, including calculating the structural dependency information of node i in the current layer according to the following formula:

[0132]

[0133] Where l is the current layer, σ is the activation function, n is the number of all neighbor nodes j, is the weighted adjacency matrix between node i and all its neighbor nodes j in the current layer, W (l) is the weight matrix, is the upper layer structure dependency information of all neighbor nodes j, d i is the degree of node i in the dependency parse tree, b (l) is the bias vector.

[0134] In one embodiment, the computer program further implements the following steps when executing the computer program: the relationship prediction module includes a fully connected neural network and a global correspondence module; the structural dependency information is input into the relationship prediction module to obtain the subject-object relationship output by the relationship prediction module, including: using the fully connected neural network to sequence the structural dependency information to obtain the subject-object pair; calculating the relationship probability score of the subject-object pair through the global correspondence module; when the relationship probability score of the subject-object pair is greater than a preset threshold, outputting the entity relationship of the subject-object pair.

[0135] When the processor in the electronic device provided by the present application executes a computer program, it first obtains the target patent text for preprocessing and constructs a corpus, which includes text entities that have been marked and classified and semantic relationships between text entities; then, semantic information is obtained through a sequence encoding module based on the corpus, and then the semantic information is input into a graph convolutional network encoding module based on a label attention mechanism to obtain structural dependency information output by the graph convolutional network encoding module based on a label attention mechanism, and the structural dependency information is further input into a relationship prediction module to obtain a subject-object relationship output by the relationship prediction module. In this way, the present application constructs a corpus through preprocessing, and through the collaboration of multiple modules, it can deeply mine hidden features, improve the efficiency and accuracy of unstructured text knowledge extraction, enhance the mining and utilization of semantic information, provide strong support for the intelligent construction of patents in the CCUS field and realize text data association.

[0136] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored, and when the computer program is executed by the computer program, the following steps are implemented:

[0137] Obtain the target patent text for preprocessing and construct a corpus, which includes text entities that have been labeled and classified and the semantic relationships between the text entities; based on the corpus, obtain semantic information through a sequence encoding module; input the semantic information into a graph convolutional network encoding module based on a label attention mechanism to obtain the structural dependency information output by the graph convolutional network encoding module based on the label attention mechanism; input the structural dependency information into a relationship prediction module to obtain the subject-object relationship output by the relationship prediction module.

[0138] In one embodiment, when the computer program executes the computer program, the following steps are further implemented: The sequence encoding module includes a MacBERT pre-trained model, a global word representation vector GloVe model, and a bidirectional long short-term memory network BiLSTM model; based on the corpus, obtain semantic information through the sequence encoding module, including: based on the corpus, use the MacBERT pre-trained model to perform word embedding on the text entity to obtain word labels; use the GloVe model to perform part-of-speech tagging embedding on the part-of-speech information of the text entity to obtain part-of-speech labels, and perform named entity recognition embedding on the relevant information of the text entity to obtain entity recognition labels; splice the word labels, part-of-speech labels, and entity recognition labels to obtain a splicing matrix; input the splicing matrix into the BiLSTM model to obtain the semantic information output by the BiLSTM model.

[0139] In one embodiment, when the computer program executes the computer program, the following steps are further implemented: Input the semantic information into a graph convolutional network encoding module based on a label attention mechanism to obtain the structural dependency information output by the graph convolutional network encoding module based on the label attention mechanism, including: The graph convolutional network encoding module based on the label attention mechanism performs the following operations: Construct a dependency parsing tree based on the semantic information; calculate the weighted adjacency matrix between node i and all its neighbor nodes j based on the similarity and distance information between the nodes in the dependency parsing tree; based on the weighted adjacency matrix, perform a weighted sum of the structural dependency information of all neighbor nodes j in the previous layer, then normalize it by dividing by the degree of node i in the dependency parsing tree, and then add a bias vector, and obtain the structural dependency information of node i in the current layer through an activation function.

[0140] In one embodiment, when the computer program executes the computer program, the following steps are further implemented: Calculate the weighted adjacency matrix between node i and all its neighbor nodes j based on the similarity and distance information between the nodes in the dependency parsing tree, including calculating the weighted adjacency matrix according to the following formula:

[0141]

[0142] where, e ij is the attention weight, e ij = softmax(s ij,d ij ), where s ij is the node similarity between node i and node j; d ij is the reciprocal of the distance between node i and node j.

[0143] In one embodiment, when the computer program is executed, the computer program further implements the following steps: Based on the weighted adjacency matrix, perform a weighted sum of the upper-layer structural dependency information of all neighbor nodes j, and then normalize it by dividing by the degree of node i in the dependency parsing tree, and then add the bias vector, and obtain the structural dependency information of node i in the current layer through the activation function, including calculating the structural dependency information of node i in the current layer according to the following formula:

[0144]

[0145] In the formula, l is the current layer, σ is the activation function, n is the number of all neighbor nodes j, is the weighted adjacency matrix between node i in the current layer and all its neighbor nodes j, W (l) is the weight matrix, is the upper-layer structural dependency information of all neighbor nodes j, d i is the degree of node i in the dependency parsing tree, b (l) is the bias vector.

[0146] In one embodiment, when the computer program is executed, the computer program further implements the following steps: The relationship prediction module includes a fully connected neural network and a global correspondence module; input the structural dependency information into the relationship prediction module to obtain the subject-object relationship output by the relationship prediction module, including: using the fully connected neural network to perform sequence labeling on the structural dependency information to obtain the subject-object pair; calculating the relationship probability score of the subject-object pair through the global correspondence module; when the relationship probability score of the subject-object pair is greater than the preset threshold, output the entity relationship of the subject-object pair.

[0147] When the computer program in the computer-readable storage medium provided by this application executes the computer program, it first obtains the target patent text for preprocessing and constructs a corpus. This corpus includes text entities that have been marked and classified, as well as the semantic relationships between the text entities. Then, based on the corpus, semantic information is obtained through the sequence encoding module. Subsequently, the semantic information is input into the graph convolutional network encoding module based on the label attention mechanism to obtain the structural dependency information output by the graph convolutional network encoding module based on the label attention mechanism. Further, the structural dependency information is input into the relationship prediction module to obtain the subject-object relationship output by the relationship prediction module. In this way, through preprocessing to construct a corpus and the collaboration of multiple modules, this application can deeply mine hidden features, improve the efficiency and accuracy of unstructured text knowledge extraction, enhance the mining and utilization of semantic information, provide strong support for the intelligent construction of patents in the CCUS field, and realize the association of text data.

[0148] Those skilled in the art should understand that the embodiments of this application can be provided as a method, a system, or a computer program product. Therefore, this application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, this application can take the form of a computer program product implemented on one or more computer-usable storage media that contain computer-usable program code.

[0149] In several embodiments provided by this application, it should be understood that the disclosed device and method can also be implemented in other ways. The device embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the accompanying drawings show the possible architectures, functions, and operations of the device, method, and computer program product according to multiple embodiments of this application. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code, and the module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, as well as the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.

[0150] In this application, the processor may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0151] In this application, the memory may include non-permanent memory in a computer-readable medium, in the form of random access memory (RAM) and / or non-volatile memory such as read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.

[0152] In this application, the computer-readable medium includes permanent and non-permanent, removable and non-removable storage media. The storage media can implement information storage by any method or technology, and the information can be computer-readable instructions, data structures, program modules, or other data. Examples of the computer's storage media include, but are not limited to, Phase Change Memory (PRAM), Static Random Access Memory (SRAM), Dynamic Random Access Memory (DRAM), other types of Random Access Memory (RAM), Read-Only Memory (ROM), Electrically Erasable Programmable Read-Only Memory (EEPROM), flash memory or other memory technologies, Compact Disc Read-Only Memory (CD-ROM), Digital Versatile Disc (DVD) or other optical storage, magnetic cassette tapes, disk storage or other magnetic storage devices, or any other non-transmission media that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transitory media such as modulated data signals and carrier waves.

[0153] It should be noted that in this article, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprising", "including" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.

[0154] The above are only specific embodiments of the present application, enabling those skilled in the art to understand or implement the present application. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to these embodiments herein, but rather to the broadest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for extracting patent entity relations, characterized in that: include: Obtaining the target patent text for preprocessing and constructing a corpus, wherein the corpus includes marked and classified text entities and semantic relations between the text entities; Based on the corpus, semantic information is obtained through a sequence encoding module; Inputting the semantic information into a graph convolutional network encoding module based on a label attention mechanism to obtain structural dependency information output by the graph convolutional network encoding module based on a label attention mechanism; The structural dependency information is input into a relationship prediction module to obtain a subject-object relationship output by the relationship prediction module.

2. The method according to claim 1, characterized in that The sequence encoding module includes a MacBERT pre-trained model, a global word representation vector GloVe model, and a bidirectional long short-term memory network BiLSTM model; The obtaining of semantic information based on the corpus through a sequence encoding module includes: Based on the corpus, the MacBERT pre-trained model is used to embed the text entities to obtain word labels; The GloVe model is used to perform part-of-speech tagging and embedding on the part-of-speech information of the text entity to obtain a part-of-speech tag, and the text entity related information is embedded in named entity recognition to obtain an entity recognition tag; Concatenating the word tags, the part-of-speech tags, and the entity recognition tags to obtain a concatenated matrix; The concatenated matrix is ​​input into the BiLSTM model to obtain the semantic information output by the BiLSTM model.

3. The method according to claim 1, characterized in that The step of inputting the semantic information into a graph convolutional network coding module based on a label attention mechanism to obtain structural dependency information output by the graph convolutional network coding module based on a label attention mechanism includes: The graph convolutional network encoding module based on the label attention mechanism performs the following operations: Constructing a dependency syntactic analysis tree based on the semantic information; Based on the similarity and distance information between each node in the dependency syntactic analysis tree, a weighted adjacency matrix between node i and all its neighboring nodes j is calculated; Based on the weighted adjacency matrix, the weighted sum of the previous layer structural dependency information of all neighbor nodes j is performed, and then normalized by dividing it by the degree of the node i in the dependency syntactic analysis tree, and then adding the bias vector, and obtaining the structural dependency information of the node i in the current layer through the activation function.

4. The method according to claim 3, characterized in that The step of calculating the weighted adjacency matrix between node i and all its neighboring nodes j based on the similarity and distance information between the nodes in the dependency syntactic analysis tree includes calculating the weighted adjacency matrix according to the following formula: In the formula, e ij is the attention weight, e ij =softmax(s ij ,d ij ), where s ij is the node similarity between node i and node j; d ij It is the inverse of the distance between node i and node j.

5. The method according to claim 3, characterized in that: Based on the weighted adjacency matrix, weighted sum is performed on the previous layer structural dependency information of all neighbor nodes j, and then normalized by dividing it by the degree of the node i in the dependency syntactic analysis tree, and then adding the bias vector, and obtaining the structural dependency information of the node i in the current layer through the activation function, including calculating the structural dependency information of the node i in the current layer according to the following formula: Where l is the current layer, σ is the activation function, n is the number of all neighbor nodes j, is the weighted adjacency matrix between node i and all its neighbor nodes j in the current layer, W (l) is the weight matrix, is the upper layer structure dependency information of all neighbor nodes j, d i is the degree of node i in the dependency parse tree, b (l) is the bias vector.

6. The method according to claim 1, characterized in that The relationship prediction module includes a fully connected neural network and a global correspondence module; The step of inputting the structural dependency information into a relationship prediction module to obtain a subject-object relationship output by the relationship prediction module includes: Using the fully connected neural network to sequence-label the structure-dependent information to obtain subject-object pairs; Calculating the relationship probability score of the subject-object pair by the global correspondence module; When the relationship probability score of the subject-object pair is greater than a preset threshold, the entity relationship of the subject-object pair is output.

7. A patent entity relationship extraction device, characterized in that: include: A corpus construction component is used to obtain the target patent text for preprocessing and construct a corpus, wherein the corpus includes marked and classified text entities and semantic relations between the text entities; A sequence encoding component, used for obtaining semantic information through a sequence encoding module based on the corpus; A graph convolutional network coding component, used for inputting the semantic information into a graph convolutional network coding module based on a label attention mechanism, and obtaining the structural dependency information output by the graph convolutional network coding module based on the label attention mechanism; The relationship prediction component is used to input the structural dependency information into the relationship prediction module to obtain the subject-object relationship output by the relationship prediction module.

8. An electronic device, characterized in that: include: A processor, a memory, and a computer program stored in the memory and executable on the processor, wherein the computer program, when executed by the processor, implements a method for extracting patent entity relationships as described in any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that: include: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method for extracting patent entity relationships as described in any one of claims 1 to 6 is implemented.

10. A computer program product, characterized in that include: The computer program product includes a computer program, and when the computer program is run on a computer, the computer is enabled to implement the method for extracting patent entity relationships as described in any one of claims 1 to 6.

Citation Information

Cited By

  • Text extraction method and system, medium and terminal

    CN121524367A