Knowledge extraction method and system based on data enhancement and priority constraint, terminal and storage medium

By performing data enhancement and priority constraints on the initial text data, and using pre-trained language models and non-autoregressive models to train knowledge extraction models, the problems of data sparsity and insufficient diversity of relation instances are solved, the accuracy of knowledge extraction is improved, and accurate extraction of triple data is achieved.

CN120654788AActive Publication Date: 2025-09-16SHENZHEN UNIV
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510516881.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-23
Publication Date
2025-09-16
Estimated Expiration
2045-04-23

AI Technical Summary

Technical Problem

In the existing technology, due to data sparsity and insufficient diversity of triple knowledge relationship instances, the accuracy of the knowledge extraction model is insufficient and cannot meet user needs.

Method used

Through data enhancement and priority constraint methods, including data enhancement processing of initial text data, using pre-trained language models and non-autoregressive models for text encoding and model training, and using entity self-reference enhancement strategy and Lagrange multiplier method to optimize the loss function, the knowledge extraction accuracy of the model is improved.

Benefits of technology

It effectively alleviates the problems of data sparsity and insufficient diversity of relationship instances, improves the accuracy of the knowledge extraction model, and realizes the accurate extraction of triple data in the current text data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120654788A_ABST
    Figure CN120654788A_ABST
Patent Text Reader

Abstract

The invention discloses a knowledge extraction method and system based on data enhancement and priority constraint, a terminal and a storage medium, and the method comprises the steps: obtaining initial text data, and carrying out data enhancement processing on the initial text data to obtain enhanced text data; determining a pre-training language model and a non-autoregression model, performing text coding processing on the enhanced text data through the pre-training language model, and performing model training through the non-autoregression model to obtain a knowledge extraction model; and obtaining current text data, and performing triple knowledge extraction processing on the current text data through the knowledge extraction model to obtain a target triple knowledge extraction result. According to the method, data enhancement processing is carried out on the initial text data, the problems of data sparsity and insufficient diversity of triple knowledge relationship examples are effectively relieved, the non-autoregression model is trained by enhancing the text data, and the knowledge extraction precision of the knowledge extraction model is effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of natural language processing technology, and in particular to a knowledge extraction method, system, terminal and computer-readable storage medium based on data enhancement and priority constraints. Background Art

[0002] Knowledge extraction is the foundation of building a knowledge graph. It requires extracting the required triples of knowledge from various types of data. Knowledge extraction involves large amounts of data and generally requires automated extraction through algorithmic models. The accuracy of the model algorithm determines the effectiveness of knowledge extraction.

[0003] However, due to data sparsity and insufficient diversity of triple knowledge relationship instances in the existing technology, the knowledge extraction accuracy of the knowledge extraction model is insufficient and cannot meet users' needs for knowledge extraction. Summary of the Invention

[0004] The main purpose of the present invention is to provide a knowledge extraction method, system, terminal and computer-readable storage medium based on data enhancement and priority constraints, aiming to solve the problem in the existing technology that due to data sparsity and insufficient diversity of triple knowledge relationship instances, the knowledge extraction accuracy of the knowledge extraction model is insufficient and cannot meet the user's knowledge extraction needs.

[0005] To achieve the above object, the present invention provides a knowledge extraction method based on data enhancement and priority constraints, the knowledge extraction method based on data enhancement and priority constraints comprising the following steps:

[0006] Acquiring initial text data, and performing data enhancement processing on the initial text data to obtain enhanced text data;

[0007] Determine a pre-trained language model and a non-autoregressive model, perform text encoding processing on the enhanced text data using the pre-trained language model, and perform model training using the non-autoregressive model to obtain a knowledge extraction model;

[0008] Current text data is acquired, and triple knowledge extraction processing is performed on the current text data using the knowledge extraction model to obtain a target triple knowledge extraction result.

[0009] Optionally, the method for knowledge extraction based on data enhancement and priority constraints, wherein the step of obtaining initial text data and performing data enhancement processing on the initial text data to obtain enhanced text data, specifically includes:

[0010] Acquire triple data from initial text data, and acquire relational properties of the triple data;

[0011] Performing data enhancement processing on the triple data according to the relationship properties to obtain initial enhanced text data;

[0012] An entity self-referential enhancement strategy is adopted to generate a self-referential triple corresponding to the triple data, and enhanced text data is obtained according to the self-referential triple and the initial enhanced text data.

[0013] Optionally, in the knowledge extraction method based on data enhancement and priority constraints, the relationship properties include symmetric relationships, non-directional relationships, and asymmetric relationships;

[0014] The performing data enhancement processing on the triple data according to the relationship properties to obtain initial enhanced text data specifically includes:

[0015] If the relationship property is the symmetric relationship or the non-directional relationship, obtaining a symmetric triple of the triple data;

[0016] If the relationship property is the asymmetric relationship, obtaining a reversible triple of the triple data;

[0017] The symmetric triples and the reversible triples are added to the triple data of the text data to obtain initial enhanced text data.

[0018] Optionally, the method for knowledge extraction based on data enhancement and priority constraints, wherein determining a pre-trained language model and a non-autoregressive model, performing text encoding processing on the enhanced text data using the pre-trained language model, and performing model training using the non-autoregressive model to obtain a knowledge extraction model, specifically includes:

[0019] Determining a pre-trained language model, inputting the enhanced text data into the pre-trained language model, and performing text encoding processing on the enhanced text data through an encoder in the pre-trained language model to obtain context embedded data;

[0020] Determine a non-autoregressive model, and perform context feature extraction and weight dynamic adjustment according to the context embedding data through a non-autoregressive decoder in the non-autoregressive model to obtain dynamic attention gating;

[0021] Performing loss calculation according to the dynamic attention gating through the fully connected layer in the non-autoregressive model to obtain subject loss and object loss, and performing weighted optimization processing and normalization processing on the subject loss and the object loss to obtain a target loss and an initial knowledge extraction model;

[0022] The initial knowledge extraction model is iteratively trained according to the target loss and the enhanced text data. When the number of iterations reaches a preset iteration threshold, the training is completed and a knowledge extraction model is obtained.

[0023] Optionally, the method for knowledge extraction based on data enhancement and priority constraints, wherein determining a non-autoregressive model, performing context feature extraction processing and dynamic weight adjustment processing according to the context embedding data by a non-autoregressive decoder in the non-autoregressive model, to obtain dynamic attention gating, specifically includes:

[0024] Inputting the context embedding data into a non-autoregressive decoder in the non-autoregressive model to obtain a preset number of query embedding data, and calculating the dependency relationship between the triples in the enhanced text data based on the query embedding data;

[0025] Context features are calculated based on the context embedding data and the query embedding data, and the dependency relationships between triples in the enhanced text data and the context features are dynamically adjusted in weight to obtain dynamic attention gating.

[0026] Optionally, the knowledge extraction method based on data augmentation and priority constraints, wherein the loss calculation is performed according to the dynamic attention gating through the fully connected layer in the non-autoregressive model to obtain the subject loss and the object loss, specifically includes:

[0027] Inputting the dynamic attention gate into the fully connected layer in the non-autoregressive model, mapping the dynamic attention gate to the relationship category space through the fully connected layer to obtain the relationship prediction probability;

[0028] Calculating the subject start position prediction probability, the subject end position prediction probability, the object start position prediction probability, and the object end position prediction probability according to the relationship prediction probability;

[0029] The subject loss is calculated based on the subject start position prediction probability and the subject end position prediction probability, and the object loss is calculated based on the object start position prediction probability and the object end position prediction probability.

[0030] Optionally, the knowledge extraction method based on data enhancement and priority constraints, wherein the subject loss and the object loss are subjected to weighted optimization and normalization processing to obtain the target loss and the initial knowledge extraction model, specifically includes:

[0031] The subject loss and the object loss are weightedly optimized using a priority constraint optimization method of a Lagrange multiplier method to obtain an optimal loss of the subject position and an optimal loss of the object position;

[0032] A rescaling strategy is used to normalize the optimal loss of the subject position, the optimal loss of the object position, the subject loss, and the object loss to obtain a target loss. The initial training of the non-autoregressive model is completed to obtain an initial knowledge extraction model.

[0033] In addition, to achieve the above-mentioned purpose, the present invention further provides a knowledge extraction system based on data enhancement and priority constraints, wherein the knowledge extraction system based on data enhancement and priority constraints comprises:

[0034] A data enhancement processing module is used to obtain initial text data and perform data enhancement processing on the initial text data to obtain enhanced text data;

[0035] A model training module is used to determine a pre-trained language model and a non-autoregressive model, perform text encoding processing on the enhanced text data using the pre-trained language model, and perform model training using the non-autoregressive model to obtain a knowledge extraction model;

[0036] The triplet knowledge extraction module is used to obtain current text data, perform triplet knowledge extraction processing on the current text data through the knowledge extraction model, and obtain target triplet knowledge extraction results.

[0037] In addition, to achieve the above-mentioned purpose, the present invention also provides a terminal, wherein the terminal includes: a memory, a processor, and a knowledge extraction program based on data enhancement and priority constraints stored on the memory and runnable on the processor, wherein the knowledge extraction program based on data enhancement and priority constraints implements the steps of the knowledge extraction method based on data enhancement and priority constraints as described above when executed by the processor.

[0038] In addition, to achieve the above-mentioned purpose, the present invention also provides a computer-readable storage medium, wherein the computer-readable storage medium stores a knowledge extraction program based on data enhancement and priority constraints, and when the knowledge extraction program based on data enhancement and priority constraints is executed by a processor, the steps of the knowledge extraction method based on data enhancement and priority constraints as described above are implemented.

[0039] In the present invention, initial text data is obtained, and data enhancement processing is performed on the initial text data to obtain enhanced text data; a pre-trained language model and a non-autoregressive model are determined, and text encoding processing is performed on the enhanced text data through the pre-trained language model, and model training is performed through the non-autoregressive model to obtain a knowledge extraction model; current text data is obtained, and triple knowledge extraction processing is performed on the current text data through the knowledge extraction model to obtain a target triple knowledge extraction result. By performing data enhancement on the initial text data, the present invention can effectively alleviate the problems of data sparsity and insufficient diversity of triple data relationship instances. At the same time, by using the enhanced text data after data enhancement processing to train the non-autoregressive model, the knowledge extraction accuracy of the model can be effectively improved. The knowledge extraction model obtained after training can achieve accurate extraction of triple data in the current text data. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] Figure 1 It is a flow chart of a preferred embodiment of the knowledge extraction method based on data enhancement and priority constraints of the present invention;

[0041] Figure 2 This is a schematic diagram of the overall process of a preferred embodiment of the knowledge extraction method based on data enhancement and priority constraints of the present invention;

[0042] Figure 3 1 is a structural diagram of a preferred embodiment of the knowledge extraction system based on data enhancement and priority constraints of the present invention;

[0043] Figure 4 FIG. 4 is a structural diagram of a preferred embodiment of the terminal of the present invention. DETAILED DESCRIPTION

[0044] In order to make the purpose, technical solutions and advantages of the present invention more clear and distinct, the present invention is further described in detail below with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0045] Knowledge extraction is the foundation for building a knowledge graph, and it requires extracting the required triple knowledge from various types of data. Knowledge extraction involves large amounts of data, and generally requires automated extraction through algorithmic models. The accuracy of the model algorithm determines the effectiveness of knowledge extraction. Entity-relationship joint extraction refers to the direct extraction of triple relationships with the structure of (subject, relationship, object) from text data. The triples involved are often distributed in different sentences, and the model algorithm needs to have the ability to understand and process long texts. In the existing technology, there are many methods for entity-relationship joint extraction, but the following problems still exist: 1. Data sparsity and insufficient diversity of relationship instances: The entity relationship instances in existing data sets are limited, and the diversity of relationship instances is not fully reflected, which restricts the generalization ability of the model; 2. Entity relationship optimization conflict: In the entity-relationship joint extraction framework, the optimization goals of entity recognition and relationship classification conflict. Low-priority tasks (such as relationship classification) will interfere with high-priority tasks (such as entity recognition), thereby reducing the overall performance.

[0046] To solve the above problems, the present invention proposes an improved entity relationship joint extraction method, which alleviates the data sparsity problem through symmetry and entity self-reference data enhancement methods, utilizes priority constraint joint optimization, and dynamically adjusts the optimization weights through Lagrange multipliers to alleviate the optimization conflict between entity and relationship recognition, thereby improving the accuracy of entity relationship joint extraction.

[0047] The knowledge extraction method based on data enhancement and priority constraints described in the preferred embodiment of the present invention is as follows: Figure 1 As shown, the knowledge extraction method based on data enhancement and priority constraints includes the following steps:

[0048] Step S10: Acquire initial text data, and perform data enhancement processing on the initial text data to obtain enhanced text data.

[0049] The present invention is a knowledge extraction task, the purpose of which is to extract all triples (subject, relationship, object) from text data. First, the present invention divides the triples in the text data into three categories according to the nature of the relationship.

[0050] Specifically, triple data in the initial text data is obtained, and relational properties of the triple data are obtained; the relational properties include symmetric relations, non-directional relations, and asymmetric relations.

[0051] like Figure 2As shown, after obtaining the initial text data, the present invention first needs to perform data enhancement processing on the initial text data. The data enhancement processing process in the present invention includes data enhancement based on relationship properties and data enhancement based on entity self-reference. Among them, the data enhancement process based on relationship properties is: according to the characteristics of the relationship, the initial text data is divided into three categories: 1. Symmetric relationship: The relationship between the triple data satisfies the symmetry, that is, if (e s ,r,e o ) is true (where e is the abbreviation of entity, r is the abbreviation of relation, S is the abbreviation of subject, O is the abbreviation of object, e s is the subject in the entity, e o is the object in the entity), then (e o ,r,e s ) is also true, for example, the relations "genre" and "sister city". "True" means: for example, if the triple (Guangzhou, adjacent, Guangdong) is indeed adjacent to the two cities, then the new triple (Guangdong, adjacent, Guangzhou) is also correct. This means that the relation r satisfies symmetry, so the new triple can be added to the original triple data. 2. Undirected relations: The relation r∈R is undirected (where R is the set of all relations), for example, "member of". 3. Asymmetric relations: The relation r∈R has a clear directionality, for example, "follows" and "followed by".

[0052] If the relationship property is the symmetrical relationship or the undirected relationship, then the symmetrical triple of the triple data is obtained; if the relationship property is the asymmetrical relationship, then the reversible triple of the triple data is obtained; the symmetrical triple and the reversible triple are added to the triple data of the text data to obtain initial enhanced text data.

[0053] For the above three types of relationships, the present invention proposes the following strategies: 1. If the triple (e s ,r,e o ) is a symmetric relationship or a non-directional relationship, then a new triple (e o ,r,e s ). 2. If the triple (e s ,r,e o ) is an asymmetric relationship, and there is a reversible relationship r' in the preset relationship set, then add the triple (e o , r′, e s). Among them, reversible relations: For example, the triple (Xiao Ming, father, Xiao Wang), and at the same time, the reversible relationship son of the father relationship exists in the preset relationship set, then the new triple (Xiao Wang, son, Xiao Ming) is also correct, which means that the relationship r satisfies the reversible relationship, so the new triple can be added to the original triple data.

[0054] An entity self-referential enhancement strategy is adopted to generate a self-referential triple corresponding to the triple data, and enhanced text data is obtained according to the self-referential triple and the initial enhanced text data.

[0055] Data enhancement based on entity self-reference: Entities may have some implicit association with themselves in semantics. In order to capture this feature, the present invention designs an entity self-reference enhancement strategy. Specifically, for each entity e i ∈E (E is the abbreviation of entity set Entity), generate the corresponding self-referential triple (e i ,NR,e i ), where "NR" stands for No Relation, indicating that there is no clear explicit semantic association between the entity and itself (e.g., Guangdong, No Relation (NR), Guangdong), but implicit contextual relevance (the entity self-referential method here is to generate such triples in order to enhance entity recognition and enhance the diversity of relationship recognition).

[0056] Step S20: determine a pre-trained language model and a non-autoregressive model, perform text encoding processing on the enhanced text data using the pre-trained language model, and perform model training using the non-autoregressive model to obtain a knowledge extraction model.

[0057] The pre-trained language model in the present invention preferably adopts a model such as BERT (Bidirectional Encoder Representations from Transformers, a deep learning model based on Transformer).

[0058] Specifically, a pre-trained language model is determined, the enhanced text data is input into the pre-trained language model, and the enhanced text data is subjected to text encoding processing by an encoder in the pre-trained language model to obtain context embedded data.

[0059] like Figure 2 As shown, after data enhancement is performed on the text data, the obtained enhanced text data is input into the pre-trained language model, and the enhanced text data is subjected to text encoding processing by the pre-trained language model.

[0060] For a document D of length l (i.e., the enhanced text data in the present invention), it is expressed as Among them, xt Represents the word at position t. This paper uses PLM (Pretrained Language Model) as the encoder to obtain the document context embedding H and attention weight A through PLM:

[0061] H, A = PLM(D);

[0062] Where H∈R l×d , A∈R h×l×l , d is the encoding dimension of PLM, and h represents the number of attention heads.

[0063] like Figure 2 As shown, the context embedding data is input into the non-autoregressive decoder in the non-autoregressive model to obtain a preset number of query embedding data, and the dependency relationship between the triplets in the enhanced text data is calculated based on the query embedding data; context features are calculated based on the context embedding data and the query embedding data, and the dependency relationship between the triplets in the enhanced text data and the context features are dynamically adjusted in weight to obtain dynamic attention gating.

[0064] The non-autoregressive decoder in this invention consists of several identical Transformer blocks, each of which contains submodules such as multi-head self-attention, multi-head cross-attention, and feedforward network.

[0065] The non-autoregressive decoder processes the contextual embedding data as follows:

[0066] 1. Query Embedding: The non-autoregressive decoder first generates a fixed number of query embeddings Q∈R N×d , N is the number of decoding layers, and each query embedding corresponds to a potential triple. These query embeddings are generated by the following formula:

[0067] Q = Proj(Concat(Q0,H));

[0068] Where H represents the output of the encoder (also known as the contextual embedded data in this invention), Concat(·) is the concatenation operation, and Proj(·) is the linear transformation of the conditional projection. Q0 is an initialization matrix initialized with a Gaussian distribution with variance 0 and mean 1.

[0069] 2. Decoder Layer: The decoder consists of N stacked decoding layers. Each decoding layer contains the following modules: a. Self-attention module: used to capture the interdependence between triplets. Its calculation formula is:

[0070]

[0071] Among them, A self is the output of the self-attention module, which represents the dependency between triplets. is the learnable parameter of the self-attention module, d k is the scaling factor of attention, and T is the matrix transpose. b. Cross attention module: It is used to integrate the context information of the input document. Its calculation formula is:

[0072]

[0073] Among them, A cross c. Dynamic attention gating mechanism: It is used to dynamically adjust the weights of cross attention and self-attention to enhance the ability to capture key information. The gating calculation method is:

[0074] A output =σ(W gate )·A cross +(1-σ(W gate ))·A self ;

[0075] Among them, σ is the Sigmoid activation function, W gate is a learnable parameter. The output of the decoder A output Relation classification and entity location prediction for triples.

[0076] The dynamic attention gate is input into the fully connected layer in the non-autoregressive model, and the dynamic attention gate is mapped to the relationship category space through the fully connected layer to obtain the relationship prediction probability; the subject starting position prediction probability, the subject ending position prediction probability, the object starting position prediction probability and the object ending position prediction probability are calculated according to the relationship prediction probability; the subject loss is calculated according to the subject starting position prediction probability and the subject ending position prediction probability, and the object loss is calculated according to the object starting position prediction probability and the object ending position prediction probability.

[0077] Furthermore, the present invention classifies and predicts the position of the enhanced text data based on the dynamic attention gating: 1. Relation classification: mapping to the relationship category space through the fully connected layer:

[0078] P R =Softmax(QW R );

[0079] in, represents the relationship prediction probability, C r is the number of relation categories, W R2. Entity Recognition: Predict the starting and ending positions of the subject and object through the position score function:

[0080]

[0081] in, They represent the probability predictions of the subject’s starting position, the subject’s ending position, the object’s starting position, and the object’s ending position, respectively. and represents a learnable parameter.

[0082] Furthermore, the loss calculation in the present invention includes classification loss calculation and joint optimization and priority constraints, wherein the classification loss processing process is as follows: for the classification task of entities and relationships, cross-entropy loss is used to measure the accuracy of classification.

[0083] Among them, the goal of relation classification is to predict the relation category in the triple, and its loss is defined as:

[0084]

[0085] Among them, L r is the cross entropy loss between entities and relations, R represents the total relation set, which contains multiple relations, and r refers to a relation in the R set. The goal of entity position classification is to predict the subject and object positions, and the subject loss L s and object loss L o They are defined as:

[0086]

[0087] Here, S is the subject, O is the object, start indicates the starting position, and end indicates the position. Since both S and O are entities, entity prediction requires both the entity's starting position prediction and the entity's ending position prediction, as the entity needs to be found.

[0088] The priority constrained optimization method of the Lagrange multiplier method is used to perform weighted optimization processing on the subject loss and the object loss to obtain the optimal loss of the subject position and the optimal loss of the object position; the optimal loss of the subject position, the optimal loss of the object position, the subject loss and the object loss are normalized using a rescaling strategy to obtain the target loss, the initial training of the non-autoregressive model is completed, and the initial knowledge extraction model is obtained.

[0089] The process of joint optimization and priority constraints in this invention is as follows: To balance the tasks of entity location classification and relationship classification, this invention adopts a priority-constrained optimization method based on the Lagrange multiplier method. Subject location classification and object location classification are set as high-priority objectives, and relationship classification is set as a low-priority objective. By introducing Lagrange multipliers, the losses of each task are weighted and optimized, thereby achieving the indestructibility of the priority objectives. The optimization objective is defined as:

[0090]

[0091] in, and are the optimal losses for the subject position and the object position, λ s and λ o is a Lagrangian operator used to constrain the optimization of high-priority objectives.

[0092] In order to maintain the stability of high-priority targets, the present invention dynamically updates the Lagrangian operator in each round of training:

[0093]

[0094] Among them, α is the learning rate, λ max It is the upper limit of the Lagrangian operator, and clip is used to prevent training instability caused by excessive updates.

[0095] In order to further improve the stability of the training process, the present invention adopts a rescaling strategy to normalize the loss function. The final loss function form is:

[0096]

[0097] This rescaling method effectively prevents the gradient explosion phenomenon caused by the rapid growth of the Lagrange multiplier, thereby making the optimization process more stable. Through the above method, the present invention achieves the maximum effect of the relationship classification task while ensuring that the performance of the subject and object position classification task does not decrease.

[0098] The initial knowledge extraction model is iteratively trained according to the target loss and the enhanced text data. When the number of iterations reaches a preset iteration threshold, the training is completed and a knowledge extraction model is obtained.

[0099] After obtaining the initial knowledge extraction model, the initial knowledge extraction model needs to be iteratively trained to correct the model parameters. The specific implementation process is as follows: using text enhancement data, the text enhancement data is repeatedly input as training data into the updated initial knowledge extraction model, and the model parameters are corrected each time using the calculated loss to realize the training of the model, and finally a trained model (that is, the knowledge extraction model in the present invention) is obtained.

[0100] Step S30: Acquire current text data, perform triple knowledge extraction processing on the current text data through the knowledge extraction model, and obtain target triple knowledge extraction results.

[0101] When it is necessary to extract triple data from the current text data, the current text data is input into the trained knowledge extraction model, and the current text data is processed for triple knowledge extraction through the knowledge extraction model, so that the triples extracted from the previous text data are obtained in the calculation results of the knowledge extraction model.

[0102] Beneficial effects of the invention: The present invention proposes a document-level entity relationship joint extraction method, which aims to solve the challenges of data sparsity, relationship instance diversity and entity relationship optimization conflicts faced in the existing technology by introducing a data enhancement strategy based on relationship properties and a priority constraint optimization method based on the Lagrange multiplier method. The knowledge extraction method proposed in the present invention can effectively improve the accuracy of entity relationship joint extraction.

[0103] Furthermore, if Figure 3 As shown, based on the above-mentioned knowledge extraction method based on data enhancement and priority constraints, the present invention also provides a knowledge extraction system based on data enhancement and priority constraints, wherein the knowledge extraction system based on data enhancement and priority constraints includes:

[0104] The data enhancement processing module 51 is used to obtain initial text data and perform data enhancement processing on the initial text data to obtain enhanced text data;

[0105] A model training module 52 is used to determine a pre-trained language model and a non-autoregressive model, perform text encoding processing on the enhanced text data using the pre-trained language model, and perform model training using the non-autoregressive model to obtain a knowledge extraction model;

[0106] The triplet knowledge extraction module 53 is used to obtain current text data, perform triplet knowledge extraction processing on the current text data through the knowledge extraction model, and obtain target triplet knowledge extraction results.

[0107] Furthermore, if Figure 4 As shown, based on the above-mentioned knowledge extraction method and system based on data enhancement and priority constraints, the present invention also provides a terminal, which includes a processor 10, a memory 20 and a display 30. Figure 4 Only some of the components of the terminal are shown, but it should be understood that implementation of all of the shown components is not required, and more or fewer components may be implemented instead.

[0108] In some embodiments, the memory 20 may be an internal storage unit of the terminal, such as a hard disk or memory of the terminal. In other embodiments, the memory 20 may also be an external storage device of the terminal, such as a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), etc. equipped on the terminal. Furthermore, the memory 20 may also include both an internal storage unit of the terminal and an external storage device. The memory 20 is used to store application software and various types of data installed on the terminal, such as the program code of the installation terminal. The memory 20 may also be used to temporarily store data that has been output or is to be output. In one embodiment, a knowledge extraction program 40 based on data enhancement and priority constraints is stored on the memory 20, and the knowledge extraction program 40 based on data enhancement and priority constraints can be executed by the processor 10, thereby realizing the knowledge extraction method based on data enhancement and priority constraints in the present application.

[0109] In some embodiments, the processor 10 may be a central processing unit (CPU), a microprocessor, or other data processing chip, used to run the program code or process data stored in the memory 20, such as executing the knowledge extraction method based on data enhancement and priority constraints.

[0110] In some embodiments, the display 30 may be an LED display, a liquid crystal display, a touch-sensitive liquid crystal display, or an OLED (Organic Light-Emitting Diode) touchscreen. The display 30 is used to display information on the terminal and to display a visual user interface. The terminals communicate with each other via a system bus.

[0111] In one embodiment, when the processor 10 executes the knowledge extraction program 40 based on data enhancement and priority constraints in the memory 20 , the steps of the knowledge extraction method based on data enhancement and priority constraints described above are implemented.

[0112] The present invention also provides a computer-readable storage medium, wherein the computer-readable storage medium stores a knowledge extraction program based on data enhancement and priority constraints, and when the knowledge extraction program based on data enhancement and priority constraints is executed by a processor, the steps of the knowledge extraction method based on data enhancement and priority constraints as described above are implemented.

[0113] In summary, the present invention provides a knowledge extraction method, system, and terminal based on data enhancement and priority constraints. The method includes: obtaining initial text data, and performing data enhancement processing on the initial text data to obtain enhanced text data; determining a pre-trained language model and a non-autoregressive model, performing text encoding processing on the enhanced text data through the pre-trained language model, and performing model training through the non-autoregressive model to obtain a knowledge extraction model; obtaining current text data, and performing triple knowledge extraction processing on the current text data through the knowledge extraction model to obtain a target triple knowledge extraction result. By performing data enhancement on the initial text data, the present invention can effectively alleviate the problems of data sparsity and insufficient diversity of triple data relationship instances. At the same time, by using the enhanced text data after data enhancement processing to train the non-autoregressive model, the knowledge extraction accuracy of the model can be effectively improved. The knowledge extraction model obtained after training can achieve accurate extraction of triple data in the current text data.

[0114] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or terminal comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or terminal. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or terminal comprising the element.

[0115] Of course, those skilled in the art will appreciate that all or part of the processes in the above-described method embodiments can be implemented by instructing related hardware (such as a processor, controller, etc.) through a computer program. The program can be stored in a computer-readable storage medium that can be read by a computer. When the program is executed, it can include the processes in the above-described method embodiments. The computer-readable storage medium can be a memory, a magnetic disk, an optical disk, etc.

[0116] It should be understood that the application of the present invention is not limited to the above examples. For those skilled in the art, improvements or changes can be made based on the above description. All these improvements and changes should fall within the scope of protection of the claims attached to the present invention.

Claims

1. A knowledge extraction method based on data enhancement and priority constraints, characterized in that: The knowledge extraction method based on data enhancement and priority constraints includes: Acquiring initial text data, and performing data enhancement processing on the initial text data to obtain enhanced text data; Determine a pre-trained language model and a non-autoregressive model, perform text encoding processing on the enhanced text data using the pre-trained language model, and perform model training using the non-autoregressive model to obtain a knowledge extraction model; Current text data is acquired, and triple knowledge extraction processing is performed on the current text data using the knowledge extraction model to obtain a target triple knowledge extraction result.

2. The knowledge extraction method based on data enhancement and priority constraints according to claim 1 is characterized in that: The obtaining of initial text data and performing data enhancement processing on the initial text data to obtain enhanced text data specifically includes: Acquire triple data from initial text data, and acquire relational properties of the triple data; Performing data enhancement processing on the triple data according to the relationship properties to obtain initial enhanced text data; An entity self-referential enhancement strategy is adopted to generate a self-referential triple corresponding to the triple data, and enhanced text data is obtained according to the self-referential triple and the initial enhanced text data.

3. The knowledge extraction method based on data enhancement and priority constraints according to claim 2 is characterized in that: The relationship properties include symmetrical relationship, non-directional relationship and asymmetrical relationship; The performing data enhancement processing on the triple data according to the relationship properties to obtain initial enhanced text data specifically includes: If the relationship property is the symmetric relationship or the non-directional relationship, obtaining a symmetric triple of the triple data; If the relationship property is the asymmetric relationship, obtaining a reversible triple of the triple data; The symmetric triples and the reversible triples are added to the triple data of the text data to obtain initial enhanced text data.

4. The knowledge extraction method based on data enhancement and priority constraints according to claim 1, characterized in that: The determining of a pre-trained language model and a non-autoregressive model, performing text encoding processing on the enhanced text data using the pre-trained language model, and performing model training using the non-autoregressive model to obtain a knowledge extraction model specifically includes: Determining a pre-trained language model, inputting the enhanced text data into the pre-trained language model, and performing text encoding processing on the enhanced text data through an encoder in the pre-trained language model to obtain context embedded data; Determining a non-autoregressive model, and performing context feature extraction and weight dynamic adjustment according to the context embedding data by a non-autoregressive decoder in the non-autoregressive model to obtain dynamic attention gating; Performing loss calculation according to the dynamic attention gating through the fully connected layer in the non-autoregressive model to obtain subject loss and object loss, and performing weighted optimization processing and normalization processing on the subject loss and the object loss to obtain a target loss and an initial knowledge extraction model; The initial knowledge extraction model is iteratively trained according to the target loss and the enhanced text data. When the number of iterations reaches a preset iteration threshold, the training is completed and a knowledge extraction model is obtained.

5. The knowledge extraction method based on data enhancement and priority constraints according to claim 4, characterized in that: The non-autoregressive decoder in the non-autoregressive model performs context feature extraction processing and weight dynamic adjustment processing according to the context embedding data to obtain dynamic attention gating, specifically including: Inputting the context embedding data into a non-autoregressive decoder in the non-autoregressive model to obtain a preset number of query embedding data, and calculating the dependency relationship between the triples in the enhanced text data based on the query embedding data; Context features are calculated based on the context embedding data and the query embedding data, and the dependency relationships between triples in the enhanced text data and the context features are dynamically adjusted in weight to obtain dynamic attention gating.

6. The knowledge extraction method based on data enhancement and priority constraints according to claim 4 is characterized in that: The loss calculation is performed by the fully connected layer in the non-autoregressive model according to the dynamic attention gating to obtain the subject loss and the object loss, specifically including: Inputting the dynamic attention gate into the fully connected layer in the non-autoregressive model, mapping the dynamic attention gate to the relationship category space through the fully connected layer to obtain the relationship prediction probability; Calculating the subject start position prediction probability, the subject end position prediction probability, the object start position prediction probability, and the object end position prediction probability according to the relationship prediction probability; The subject loss is calculated based on the subject start position prediction probability and the subject end position prediction probability, and the object loss is calculated based on the object start position prediction probability and the object end position prediction probability.

7. The knowledge extraction method based on data enhancement and priority constraints according to claim 4 is characterized in that: The weighted optimization processing and normalization processing of the subject loss and the object loss to obtain the target loss and the initial knowledge extraction model specifically include: The subject loss and the object loss are weightedly optimized using a priority constraint optimization method of a Lagrange multiplier method to obtain an optimal loss of the subject position and an optimal loss of the object position; A rescaling strategy is used to normalize the optimal loss of the subject position, the optimal loss of the object position, the subject loss, and the object loss to obtain a target loss. The initial training of the non-autoregressive model is completed to obtain an initial knowledge extraction model.

8. A knowledge extraction system based on data enhancement and priority constraints, characterized in that: The knowledge extraction system based on data enhancement and priority constraints includes: A data enhancement processing module is used to obtain initial text data and perform data enhancement processing on the initial text data to obtain enhanced text data; A model training module is used to determine a pre-trained language model and a non-autoregressive model, perform text encoding processing on the enhanced text data using the pre-trained language model, and perform model training using the non-autoregressive model to obtain a knowledge extraction model; The triplet knowledge extraction module is used to obtain current text data, perform triplet knowledge extraction processing on the current text data through the knowledge extraction model, and obtain target triplet knowledge extraction results.

9. A terminal, characterized in that: The terminal includes: a memory, a processor, and a knowledge extraction program based on data enhancement and priority constraints stored in the memory and runnable on the processor. When the knowledge extraction program based on data enhancement and priority constraints is executed by the processor, the steps of the knowledge extraction method based on data enhancement and priority constraints as described in any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a knowledge extraction program based on data enhancement and priority constraints. When the knowledge extraction program based on data enhancement and priority constraints is executed by a processor, the steps of the knowledge extraction method based on data enhancement and priority constraints as described in any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Triple extraction model training method, triple extraction method, device and equipment

    CN115033717A

  • Knowledge extraction method and system fusing pre-training language model

    CN117521802A

  • Knowledge graph link prediction method based on heuristic information and graph neural network

    CN118036726A

  • Triple information extraction method, apparatus, and device, and computer-readable storage medium

    WO2022116417A1