Continuous learning data space relation extraction method and system based on parameter regularization and dynamic memory
By using BERT pre-trained language model and parameter regularization technology, combined with dynamic memory management, the problem of unreasonable encoder layer parameter changes and sample selection in continuous learning relationship extraction is solved, and the accuracy of relationship prediction in the construction of knowledge graph data space is improved.
Patent Information
- Application Number
- CN202510469216.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-15
- Publication Date
- 2025-07-22
AI Technical Summary
The existing continuous learning relationship extraction method has failed to effectively solve the problem of embedding space representation deviation caused by the parameter changes of the encoder layer and the unreasonable typical sample selection strategy caused by the change of the encoder layer, resulting in the forgetting of old tasks and reducing the prediction accuracy.
The BERT pre-trained language model is used as an encoder, combining parameter regularization and dynamic memory technology, and limit parameter changes through cross entropy loss and custom parameter regularization loss functions, optimize sample memory allocation and data enhancement, and enrich sample set diversity.
Effectively maintaining old task knowledge, improving the accuracy of relationship prediction, and is suitable for continuous learning relationship extraction in the construction of knowledge graph data space.
Smart Images

Figure CN120353594A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of knowledge graphs, and specifically, relates to a method and system for extracting continuous learning data space relationships based on parameter regularization and dynamic memory. Background Art
[0002] As an infrastructure for data circulation and utilization, the data space plays an important role in data storage, data integration, data analysis and application, and can support the integration of multi-source heterogeneous data and promote cross-domain and cross-modal application and innovation of data. The knowledge graph can better integrate multi-source heterogeneous data and construct a more manageable data space by building an entity relationship network between data. The knowledge graph itself plays an important role in the field of natural language processing. In the process of building a knowledge graph, relationship extraction is a key task. Relationship extraction originated in the era of information explosion, when people needed to obtain useful information from a large amount of text data. The purpose of relationship extraction is to extract the relationship between entities in the text, which helps to build the edges of the knowledge graph, and can help to automatically obtain the relationship between entities from a large amount of text, greatly improving the efficiency of information processing.
[0003] However, traditional relationship extraction methods are usually based on supervised learning and require a large amount of labeled training data as support. However, this method has problems such as long training time consuming computing resources and the model being unable to adapt to newly emerging relationships. Therefore, researchers have begun to focus on continuous learning relationship extraction, aiming to solve the limitations of traditional relationship extraction methods. Continuous learning research aims to learn from continuous data to continuously update and improve the current model. The main challenge of continuous learning is to avoid catastrophically forgetting old knowledge when learning new tasks.
[0004] In 2019, the concept of continuous learning relation extraction was proposed, and the EA-EMR model was proposed. It simply replays the data stored in memory and optimizes the output of the encoder layer through an embedding alignment model to limit the changes in the previous relation class embeddings. On the basis of the original model, the EMAR model proposed in 2020 adds a method based on episodic memory activation and reconsolidation. It adopts a three-stage training method. During the second-stage training, it replays the data in memory and new task data simultaneously. By activating the previously learned relation knowledge, it introduces it into the current learning task. The RP-CRE model proposed in 2021 finely represents the sample embeddings by weighted fusion of different relation prototypes into the sample embeddings through the attention mechanism, improving the accuracy of sample relation classification. The proposed adversarial class data augmentation strategy (ACA) in 2022 analyzes the root cause of forgetting in continuous relation extraction to obtain more training data. The Cdec model proposed in 2023 maintains the balance between model plasticity and stability by freezing some important weight parameters of the classifier during the model training stage.
[0005] Since none of the above models consider the problem of embedding space representation deviation caused by the continuous change of encoder layer parameters and the unreasonable selection strategy of typical samples in continuous tasks. Therefore, the above models will forget more old task knowledge in continuous tasks with multiple iterations, reducing the prediction accuracy. Summary of the Invention
[0006] Aiming at the deficiencies of the existing continuous learning relation extraction methods, that is, the problem of embedding space representation deviation caused by the continuous change of encoder layer parameters and the unreasonable selection strategy of typical samples in continuous tasks, the present invention proposes a continuous learning data space relation extraction method and system based on parameter regularization and dynamic memory, using the BERT pre-trained language model as the encoder, model parameter regularization, dynamic memory construction, and memory sample data augmentation technology to limit the change of encoder layer parameters, optimize the selection strategy of typical samples, and enrich the diversity of the sample memory set.
[0007] The present invention is realized through the following technical solutions: A continuous learning data space relation extraction method based on parameter regularization and dynamic memory: The method specifically includes the following steps:
[0008] Step 1: Obtain a sentence-level relation extraction data set and preprocess the data set;
[0009] Step 2: Construct a continuous learning data space relation extraction model based on parameter regularization and dynamic memory;
[0010] Step 3: Use the training set data divided in Step 1 to train the continuous learning data spatial relationship extraction model based on parameter regularization and dynamic memory constructed in Step 2;
[0011] Step 4: Use the test set data divided in Step 1 to perform relationship extraction on the continuously learning data spatial relationship extraction model trained in Step 3 to obtain the predicted entity relationship categories.
[0012] Further, in Step 1,
[0013] Step 1.1: Preprocess the data set. For each sentence sample, label a pair of entities as the relationship prediction target;
[0014] Step 1.2: Divide the data set into multiple sub-data sets, each sub-data set containing a training set and a test set for multiple task training.
[0015] Further, in Step 2,
[0016] Step 2.1: Use the BERT encoder to obtain the embedding representation of the sample entity pairs included in the current task;
[0017] Step 2.2: Pass the embedding representation of the sample entity pairs included in the current task through a multi-layer perceptron to obtain a preliminary relationship prediction result;
[0018] Step 2.3: Input the prediction result into the cross-entropy loss function and the custom parameter regularization loss function to reduce the deviation between the prediction result and the true result;
[0019] Step 2.4: Define a method for constructing a sample memory set, output the embedding representation of the sample entity pairs included in the current task after preliminary training, and select typical samples of the current task through the dynamic memory method;
[0020] Step 2.5: Define a data augmentation method for symmetric and asymmetric relationships in the sample memory set. For symmetric relationship classes, exchange the positions of the entity pairs, and for asymmetric relationship classes, replace the entity pairs between samples;
[0021] Step 2.6: Define a method for calculating the parameter regularization loss, calculate the gradient changes of different parameters of the model under the current task, obtain the importance weights of each parameter, and use the sum of the products of the changes in all parameter values and the corresponding weights as the parameter regularization loss.
[0022] Further, in Step 2.5,
[0023] The relationships are divided into symmetric relationships and asymmetric relationships;
[0024] The symmetric relationship means that the exchange of the positions of the head entity and the tail entity of the corresponding relationship instance does not affect its correctness, and the asymmetric relationship means that the positions of the head entity and the tail entity of the corresponding relationship cannot be randomly interchanged;
[0025] For the symmetric relationship, exchange the head entity and the tail entity of the relationship corresponding sample to generate a new relationship sample;
[0026] For the asymmetric relationship, randomly select two instances from the samples corresponding to the relationship, and replace the head entity and the tail entity of the first instance with the head entity and the tail entity of the second instance to generate a new relationship sample;
[0027] Through step 2.5, the in-memory sample size is tripled.
[0028] Furthermore, in step 2.6,
[0029] Approximate the importance of the parameter by the sensitivity of the model output to the change of the parameter; that is, when a certain parameter in the model changes, the more significant the output change of the same sample, the more important the parameter is for the same sample.
[0030] Furthermore, in step 3,
[0031] Step 3.1: Input the training set of the primary tasks divided in step 1 into the continuous learning data space relationship extraction model based on parameter regularization and dynamic memory constructed in step 2 to obtain the relationship prediction result;
[0032] Step 3.2: Input the relationship prediction result obtained in step 3.1 into the cross-entropy loss function and the custom parameter regularization loss;
[0033] Step 3.3: Minimize the joint loss function including the cross-entropy loss function and the custom loss function in step 2.6 to preliminarily train the model;
[0034] Step 3.4: Output the embedding representation of the sample entity pairs included in the current task after preliminary training. Select typical samples of the current task through the dynamic memory method and store them in the sample memory to construct a sample memory set;
[0035] Step 3.5: Perform data augmentation on the symmetric relationship and the asymmetric relationship for all samples in the sample memory set to form an augmented sample memory data set;
[0036] Step 3.6: Input the sample entities included in the augmented sample data set into the model trained in step 3.3 to obtain the replay relationship prediction result;
[0037] Step 3.7: Input the prediction result in step 3.6 into the cross-entropy loss function to calculate the cross-entropy loss, and minimize the cross-entropy loss function to retrain the model;
[0038] Step 3.8: Save the output model after training in Step 3.7, and obtain the parameter regularization weights according to the samples included in the current task and the parameter gradient calculated by the model.
[0039] Step 3.9: Repeat the above Steps 3.1 - 3.8 until the training of all sub - datasets is completed.
[0040] Further, in Step 4,
[0041] Step 4.1: Input the test - set data divided in Step 1 into the continuous learning data - space relationship extraction model based on parameter regularization and dynamic memory trained in Step 3.
[0042] Step 4.2: Obtain the relationship prediction result.
[0043] A continuous learning data - space relationship extraction system based on parameter regularization and dynamic memory
[0044] The relationship extraction system includes a pre - processing module, a relationship extraction model construction module, a training module, and a prediction module;
[0045] The pre - processing module obtains a sentence - level relationship extraction data set, pre - processes the data set, and divides it into multiple sub - data sets, each sub - data set containing a training set and a test set;
[0046] The relationship extraction model construction module constructs a continuous learning data - space relationship extraction model based on parameter regularization and dynamic memory;
[0047] The training module uses the training - set data divided by the pre - processing module to train the continuous learning data - space relationship extraction model constructed by the relationship extraction model construction module;
[0048] The prediction module uses the test - set data divided by the pre - processing module to perform relationship extraction on the continuously learning data - space relationship extraction model trained by the training module, and obtains the predicted entity - to - entity relationship category.
[0049] An electronic device includes a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, the steps of the above - mentioned method are implemented.
[0050] A computer - readable storage medium is used to store computer instructions, and when the computer instructions are executed by a processor, the steps of the above - mentioned method are implemented.
[0051] Advantages of the present invention
[0052] The present invention encodes the text information of samples by using the pre-trained language model BERT. At the same time, in order to limit the parameter changes in multiple consecutive tasks, the changes in the sample embedding representation are maintained by calculating the parameter importance weights and restricting their changes in the loss function. And by using the method of dynamic memory, the memory space is allocated more reasonably, and the test results of the previous task are used as the basis for memory allocation of the relationships in the current task. Relationships with lower accuracy will be allocated more memory sample numbers and their corresponding memory spaces. Finally, data augmentation based on rules is performed on the sample set, enriching the diversity of the memory sample set.
[0053] The method of the present invention can better maintain and recover the knowledge of old tasks, improve the accuracy of relationship prediction in continuous learning relationship extraction, and is applicable to the continuous learning relationship extraction technology in the field of knowledge graph data space construction. Brief Description of the Drawings
[0054] Figure 1 is a schematic flowchart of the method of the present invention;
[0055] Figure 2 is a model of the continuous learning data space relationship extraction method based on parameter regularization and dynamic memory of the present invention;
[0056] Figure 3 is a schematic diagram of the construction of dynamic memory of the present invention. Detailed Embodiments
[0057] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0058] The experimental methods used in the following embodiments are all conventional methods unless otherwise specified. The materials, reagents, methods, and instruments used, unless otherwise specified, are all conventional materials, reagents, methods, and instruments in this field, and those skilled in the art can obtain them through commercial channels.
[0059] As Figure 1 shown, the present invention proposes a continuous learning data space relationship extraction method based on parameter regularization and dynamic memory:
[0060] Step 1: Obtain a sentence-level relationship extraction data set, preprocess the data set, and divide it into multiple sub-data sets, each sub-data set containing a training set and a test set;
[0061] Step 1.1: Obtain the sentence-level relation extraction dataset. For each sentence sample, the present invention labels a pair of entities therein as the relation prediction target;
[0062] Step 1.2: Divide the dataset into multiple sub-datasets, each of which contains a training set and a test set for multiple task trainings.
[0063] Step 2: Construct a continuous learning data space relation extraction model based on parameter regularization and dynamic memory;
[0064] Step 2.1: Use the BERT encoder to obtain the embedding representations of the entity pairs included in the current task;
[0065] The pre-trained language model BERT is used as the encoder. Specifically, for sample i, the special tokens in the sentence are used to represent the relationship between entities. For the input sample with a pair of entity information already labeled, the start and end positions of the head entity and the tail entity are marked with [E11] / E
[12] and [E21] / E
[22] respectively. The encoder processes the entire sentence, and the hidden layer representations of [E11] and [E21] are used as the relationship representations between the sentence entities. The representation of sample i is as follows:
[0066] h i =LayerNorm(WDropout([h 11 ;h 21 )+b)
[0067] where are the hidden layer representations of the start tokens of the head entity and the tail entity respectively. (where d is the dimension of the encoder hidden layer) and are the weight and bias matrices respectively. h i is the relation representation of sample i.
[0068] Step 2.2: Pass the embedding representations of the entity pairs included in the current task through a multi-layer perceptron to obtain the preliminary relation prediction result;
[0069] Specifically, for the current task T k , before the task starts training, the model first saves the weight of the old task classification network and adds a classification network for the new task data In the first stage of training, the classifier classifies the relation corresponding to the new task to obtain the probability distribution P(x i ; θ) as follows:
[0070] P(x i ; θ)=softmax([W old h i ; Wnew h i +b c )
[0071] where [W old h i ; W new h i represents the concatenation of the outputs of the old task network and the new task network, and
[0072] Step 2.3: Input the prediction result into the cross-entropy loss function and the custom parameter regularization loss function in Step 2.6 to reduce the deviation between the prediction result and the true result.
[0073] When the new task T k appears, train the model on the new task data . As Figure 2 shown, the sample first obtains the relationship representation between corresponding entities through the encoder, and then obtains the final classification distribution through the classifier. In the first-stage learning, the loss function of T k consists of two parts: cross-entropy loss and parameter regularization loss.
[0074]
[0075] where P(y i |x i ) is the score of the relevant sample output by the classifier layer, and (x i , y i ) is the sample from .
[0076]
[0077] where represents the parameters of the old encoder network (determined by the encoder of the previous task in the task sequence T k-1 ), and Ω k-1 represents the parameter regularization weight of the k-1 task.
[0078] Step 2.4: Define a method for constructing a sample memory set, output the embedding representation of the sample entity pairs included in the current task after preliminary training, and select typical samples of the current task through the dynamic memory method;
[0079] To retain the knowledge learned from previous tasks, representative samples are selected from the training samples and stored in the memory so that they can be directly accessed during the entire task learning process. Specifically, as Figure 3 shown, a fixed memory size m is set before the start of the first task. At the kth task, equal space is allocated for all samples of the current task Therefore, after each new task training is completed, when selecting a representative subset from the new task dataset, it is necessary to delete a part of the samples from the existing fixed memory. The basis for deletion is the test accuracy P of the model in the previous task. k-1 , the samples with high accuracy are most likely to be deleted. The accuracy is used as the additional weight of the samples, and the samples already in the memory are randomly deleted according to the full weight probability. All relationship categories with an accuracy threshold higher than α may be deleted. At the same time, a minimum deletion limit is set, and the number of samples of each relationship category in the memory shall not be less than β. Finally, dynamic allocation of memory space can be achieved.
[0080] Step 2.5: Define the data augmentation methods for symmetric and asymmetric relationships in the sample memory set. For symmetric relationship classes, swap the positions of entity pairs, and for asymmetric relationship classes, replace entity pairs between samples.
[0081] Relationships are divided into symmetric and asymmetric relationships. Symmetric relationships refer to those where the exchange of the positions of the head entity and the tail entity of the corresponding relationship instance does not affect its correctness, such as the "spouse" relationship type. Asymmetric relationships refer to those where the positions of the head entity and the tail entity of the corresponding relationship cannot be randomly interchanged. The specific memory enhancement strategies are as Figure 2 shown. For symmetric relationships, swap the head entity and the tail entity of the sample corresponding to relationship r to generate a new relationship sample For asymmetric relationships, randomly select two instances and from the samples corresponding to relationship r. The head entity and the tail entity of sample are replaced by the head entity and the tail entity of instance to generate a new relationship sample Through the above strategies, the size of the memory samples can be tripled.
[0082] Step 2.6: Define the calculation method of parameter regularization loss, calculate the gradient changes of different parameters of the model in the current task, obtain the importance weights of each parameter, and take the sum of the products of the changes in all parameter values and the corresponding weights as the parameter regularization loss.
[0083] The importance of a parameter is approximated by the sensitivity of the model output to the change of that parameter. Specifically, when a certain parameter in the model changes, the more significant the change in the output of the same sample, the more important the parameter is for the same sample. Formally, define the model f(θ) after the current task training, and given the instance x i , the model output is f(x i ; θ). Applying a small perturbation δ k ∈θ to the parameter θ k will cause a change in the model output, which can be expressed as:
[0084]
[0085] where is the parameter θ k on the sample x i gradient, representing the importance of the parameters of the sample.
[0086] Based on the above formula, the parameter gradient of the model can measure the importance of the sample parameters, that is, a small perturbation of the parameters will affect the degree of the representation output corresponding to the instance x i Then, the importance weight Ω of the parameter θ of the task is obtained by accumulating the gradients of all instances in the task k : k
[0087]
[0088] where N is the total number of task instances at that time. The parameters with smaller calculated weights will not have much impact on the output, indicating that the parameter carries less old task knowledge. On the contrary, the parameters with larger weights retain more old task knowledge.
[0089] Step 3: Use the training set data divided in Step 1 to train the continuous learning data space relationship extraction model based on parameter regularization and dynamic memory constructed in Step 2;
[0090] As Figure 2 shown, Step 3 includes:
[0091] Step 3.1: Input the training set of the one-time task divided in Step 1 into the continuous learning data space relationship extraction model based on parameter regularization and dynamic memory constructed in Step 2 to obtain the relationship prediction result;
[0092] Step 3.2: Input the relationship prediction result obtained in Step 3.1 into the cross-entropy loss function and the custom parameter regularization loss;
[0093] Step 3.3: Minimize the joint loss function including the cross-entropy loss function and the custom loss function in Step 2.6 to preliminarily train the model.
[0094] Step 3.4: For the embedding representation output of the sample entity pairs included in the current task after preliminary training, select typical samples of the current task through the dynamic memory method and store them in the sample memory to construct a sample memory set;
[0095] Step 3.5: Perform symmetric relationship and asymmetric relationship data augmentation on all samples in the sample memory set to form an augmented sample memory data set.
[0096] Step 3.6: Input the sample entities included in the augmented sample data set into the model trained in Step 3.3 to obtain the replay relationship prediction result;
[0097] Step 3.7: Input the prediction result of Step 3.6 into the cross-entropy loss function to calculate the cross-entropy loss, and retrain the model by minimizing the cross-entropy loss function;
[0098] Step 3.8: Save the output model after training in Step 3.7, and obtain the parameter regularization weights based on the samples included in the current task and the parameter gradient calculated by the model;
[0099] Step 3.9: Repeat the above Steps 3.1 - 3.8 until the training of all sub-datasets is completed.
[0100] Design the loss function as a joint loss function:
[0101]
[0102] where λ is a hyperparameter used to smooth the plasticity and stability of the model. By continuously optimizing the loss function, all parameters of the model are updated.
[0103] Step 4: Use the test set data divided in Step 1 to perform relationship extraction on the continuously learning data space relationship extraction model based on parameter regularization and dynamic memory trained in Step 3, and obtain the predicted entity relationship categories.
[0104] Step 4.1: Input the test set data divided in Step 1 into the continuously learning data space relationship extraction model based on parameter regularization and dynamic memory trained in Step 3;
[0105] Step 4.2: Obtain the relationship prediction result.
[0106] To sum up, in order to prevent the embedded representation of sample entity pairs from deviating severely, the present invention adds parameter changes as a penalty to the loss function to restrict the changes of some important parameters in the encoder part.
[0107] To effectively alleviate the problem of sample overfitting caused by repeated replay of a small number of samples, the relationship samples are divided into symmetric and asymmetric relationships and subjected to a certain degree of data augmentation respectively to enrich the diversity of the memory sample set.
[0108] A continuous learning data space relationship extraction system based on parameter regularization and dynamic memory,
[0109] The relationship extraction system includes a preprocessing module, a relationship extraction model construction module, a training module, and a prediction module;
[0110] The preprocessing module obtains a sentence-level relationship extraction data set, preprocesses the data set, and divides it into multiple sub-datasets, each sub-dataset containing a training set and a test set;
[0111] The relationship extraction model construction module constructs a continuous learning data space relationship extraction model based on parameter regularization and dynamic memory;
[0112] The training module uses the training set data divided by the preprocessing module to train the continuous learning data space relationship extraction model constructed by the relationship extraction model construction module based on parameter regularization and dynamic memory;
[0113] The prediction module uses the test set data divided by the preprocessing module to perform relationship extraction on the continuously learning data space relationship extraction model trained by the training module, and obtains the predicted entity relationship categories.
[0114] An electronic device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the above method are implemented.
[0115] A computer-readable storage medium is used to store computer instructions, and when the computer instructions are executed by a processor, the steps of the above method are implemented.
[0116] The memory in the embodiments of the present application may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read only memory (ROM), a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchlink DRAM (SLDRAM), and direct rambus RAM (DR RAM). It should be noted that the memory for the method described in the present invention is intended to include, but not be limited to, these and any other suitable types of memory.
[0117] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired means such as coaxial cable, optical fiber, digital subscriber line (DSL), or wireless means such as infrared, wireless, microwave, etc. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server, data center, etc. that contains one or more integrated available media. The available media can be magnetic media such as floppy disks, hard disks, magnetic tapes, optical media such as high-density digital video discs (DVDs), or semiconductor media such as solid state discs (SSDs), etc.
[0118] In the implementation process, each step of the above method can be completed by the integrated logic circuit of the hardware in the processor or the instructions in the form of software. The steps of the method disclosed in the embodiments of the present application can be directly embodied as being executed and completed by the hardware processor, or executed and completed by a combination of the hardware and software modules in the processor. The software module can be located in a mature storage medium in the art such as random access memory, flash memory, read-only memory, programmable read-only memory, or electrically erasable programmable memory, registers, etc. This storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps of the above method. To avoid repetition, it will not be described in detail here.
[0119] It should be noted that the processor in the embodiments of the present application may be an integrated circuit chip with signal processing capabilities. In the implementation process, the steps of the above method embodiments can be completed by the integrated logic circuit in the hardware of the processor or instructions in the form of software. The above-mentioned processor may be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor, or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as being executed and completed by a hardware decoding processor, or executed and completed by a combination of hardware and software modules in the decoding processor. The software module may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. The storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps of the above method.
[0120] The above has introduced in detail a method and system for extracting the relationship of continuous learning data space based on parameter regularization and dynamic memory proposed by the present invention, and has elaborated on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present invention.
Claims
1. A method for extracting continuous learning data space relationships based on parameter regularization and dynamic memory, characterized in that: The method specifically includes the following steps: Step 1: Obtain a sentence-level relationship extraction dataset and preprocess the dataset; Step 2: Construct a continuous learning data space relationship extraction model based on parameter regularization and dynamic memory; Step 3: Use the training set data divided in Step 1 to train the continuous learning data space relationship extraction model constructed in Step 2; Step 4: Use the test set data divided in Step 1 to perform relationship extraction on the continuously learning data space relationship extraction model trained in Step 3 to obtain the predicted entity relationship categories.
2. The method according to claim 1, wherein: In Step 1, Step 1.1: Preprocess the dataset. For each sentence sample, mark a pair of entities as the relationship prediction target; Step 1.2: Divide the dataset into multiple sub-datasets, each sub-dataset containing a training set and a test set for multiple task training.
3. The method according to claim 2, wherein: In Step 2, Step 2.1: Use the BERT encoder to obtain the embedding representation of the sample entity pairs included in the current task; Step 2.2: Pass the embedding representation of the sample entity pairs included in the current task through a multi-layer perceptron to obtain a preliminary relationship prediction result; Step 2.3: Input the prediction result into the cross-entropy loss function and the custom parameter regularization loss function to reduce the deviation between the prediction result and the true result; Step 2.4: Define a method for constructing a sample memory set, output the embedding representation of the sample entity pairs included in the current task after preliminary training, and select typical samples of the current task through the dynamic memory method; Step 2.5: Define a data augmentation method for symmetric and asymmetric relationships in the sample memory set. For symmetric relationship classes, exchange the positions of the entity pairs, and for asymmetric relationship classes, replace the entity pairs between samples; Step 2.6: Define a method for calculating the parameter regularization loss, calculate the gradient changes of different parameters of the model under the current task, obtain the importance weights of each parameter, and take the sum of the products of all parameter value changes and the corresponding weights as the parameter regularization loss.
4. The method according to claim 3, wherein: In Step 2.5, The relationships are divided into symmetric relationships and asymmetric relationships; Among them, a symmetric relationship means that the exchange of the positions of the head entity and the tail entity of the corresponding relationship instance does not affect its correctness, and an asymmetric relationship means that the positions of the head entity and the tail entity of the corresponding relationship cannot be interchanged randomly; For symmetric relationships, exchange the head entity and the tail entity of the relationship corresponding sample to generate a new relationship sample; For asymmetric relationships, randomly select two instances from the samples corresponding to the relationship, and replace the head entity and the tail entity of the first instance with the head entity and the tail entity of the second instance to generate a new relationship sample; Through Step 2.5, the memory sample size is tripled.
5. The method according to claim 4, wherein: In Step 2.6, Approximate the importance of the parameter by the sensitivity of the model output to the change of the parameter; that is, when a certain parameter in the model changes, the more significant the output change of the same sample, the more important the parameter is for the same sample.
6. The method according to claim 5, characterized in that: In Step 3, Step 3.1: Input the training set of the primary tasks divided in Step 1 into the continuous learning data space relationship extraction model based on parameter regularization and dynamic memory constructed in Step 2 to obtain relationship prediction results; Step 3.2: Input the relationship prediction results obtained in Step 3.1 into the cross-entropy loss function and the custom parameter regularization loss; Step 3.3: Minimize the combined loss function that includes the cross-entropy loss function and the custom loss function in Step 2.6 to preliminarily train the model; Step 3.4: Output the embedded representations of the sample entity pairs included in the current task after preliminary training, select typical samples of the current task through the dynamic memory method, and store them in the sample memory to construct a sample memory set; Step 3.5: Perform symmetric and asymmetric relationship data augmentation on all samples in the sample memory set to form an augmented sample memory data set; Step 3.6: Input the sample entities included in the augmented sample data set into the model trained in Step 3.3 to obtain replay relationship prediction results; Step 3.7: Input the prediction results in Step 3.6 into the cross-entropy loss function to calculate the cross-entropy loss, and minimize the cross-entropy loss function to retrain the model; Step 3.8: Save the output model after training in Step 3.7, calculate the parameter gradients based on the samples included in the current task and the model, and obtain the parameter regularization weights; Step 3.9: Repeat the above Steps 3.1 - 3.8 until the training of all sub-data sets is completed.
7. The method according to claim 6, wherein: In Step 4, Step 4.1: Input the test set data divided in Step 1 into the continuous learning data space relationship extraction model based on parameter regularization and dynamic memory trained in Step 3; Step 4.2: Obtain relationship prediction results.
8. A relationship extraction system for performing the continuous learning data space relationship extraction method based on parameter regularization and dynamic memory according to any one of claims 1 to 7, characterized in that: The relationship extraction system includes a preprocessing module, a relationship extraction model construction module, a training module, and a prediction module; The preprocessing module obtains a sentence-level relationship extraction data set and preprocesses the data set; The relationship extraction model construction module constructs a continuous learning data space relationship extraction model based on parameter regularization and dynamic memory; The training module uses the training set data divided by the preprocessing module to train the continuous learning data space relationship extraction model constructed by the relationship extraction model construction module; The prediction module uses the test set data divided by the preprocessing module to perform relationship extraction on the continuous learning data space relationship extraction model trained by the training module to obtain the predicted entity relationship categories.
9. An electronic device, comprising a memory and a processor, the memory storing a computer program, characterized in that, The processor implements the steps of the method according to any one of claims 1 to 7 when executing the computer program.
10. A computer-readable storage medium for storing computer instructions, characterized in that, The computer instructions implement the steps of the method according to any one of claims 1 to 7 when executed by the processor.