Material supply chain named entity recognition model training method and system

By imitating sentences and entity words, high-quality training samples of the power material supply chain are generated, which solves the data scarcity problem of the named entity recognition model and improves the recognition accuracy and training efficiency of the model.

CN120354854BActive Publication Date: 2025-09-16STATE GRID DIGITAL TECHNOLOGY HOLDING CO LTD +2
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510813641.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-18
Publication Date
2025-09-16
Estimated Expiration
2045-06-18

AI Technical Summary

Technical Problem

In the field of electric power material supply chain, named entity recognition models lack high-quality labeled data, resulting in insufficient recognition accuracy. In addition, large language models have hallucination problems in information extraction, making them difficult to apply directly.

Method used

By imitating sentences and entity words, a large number of training samples are generated. Sentences and entity words are imitated using a large language model, a set of imitated sentences is constructed, and entity types are labeled to improve the efficiency and accuracy of model training.

Benefits of technology

It achieves the rapid generation of high-quality labeled data in small sample scenarios, improves the recognition accuracy of named entity recognition models, solves the problem of data scarcity, and reduces the need for manual labeling.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120354854B_ABST
    Figure CN120354854B_ABST
Patent Text Reader

Abstract

The present invention provides a method and system for training a named entity recognition model for a material supply chain. The implementation scheme is as follows: obtaining training samples for an entity recognition model for an electric power material supply chain, wherein the training samples include sentence samples and the entity types of each entity word sample in the sentence samples; based on the imitation requirement that the entity types of each entity word in the imitation sentence are the same as the entity word samples corresponding to the sentence samples, sentence imitation is performed on the sentence samples to obtain at least one target imitation sentence; based on each target imitation sentence and the entity types of the entity word samples corresponding to each entity word in each target imitation sentence, each first imitation sample is determined; and based on each first imitation sample, the entity recognition model is trained. Thus, a large number of labeled samples can be quickly generated when there are relatively few samples, thereby improving the training accuracy of the model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a material supply chain named entity recognition model training method and system. Background Art

[0002] At present, in the field of electric power material supply chain, the Chinese named entity recognition task faces the dilemma of scarcity of high-quality annotated data. As a result, the existing named entity recognition model based on supervised learning is not accurate enough in this field and cannot meet the demand for accurate information extraction in electric power material supply chain management.

[0003] Furthermore, while large language models possess powerful linguistic knowledge, they suffer from the "hallucination problem" in named entity recognition, making them difficult to directly apply to information extraction tasks in the same field. The hallucination problem refers to the phenomenon where large language models generate seemingly plausible content that is actually false or fabricated.

[0004] Therefore, how to generate high-quality labeled data in the field of power material supply chain to improve the accuracy of named entity recognition models is a technical problem that needs to be solved in this field. Summary of the Invention

[0005] The present invention provides a material supply chain named entity recognition model training method and system, which can solve at least one of the above technical problems.

[0006] According to one aspect of the present invention, a method for training a named entity recognition model for a material supply chain is provided, comprising:

[0007] Obtaining training samples of an entity recognition model for a power material supply chain text, wherein the training samples include sentence samples and entity types of each entity word sample in the sentence samples;

[0008] Based on the imitation requirement that each entity word in the imitation sentence has the same entity type as the corresponding entity word sample in the sentence sample, sentence imitation is performed on the sentence sample to obtain multiple target imitation sentences;

[0009] Determine each first imitation sample based on each target imitation sentence and the entity type of the entity word sample corresponding to each entity word in each target imitation sentence;

[0010] The entity recognition model is trained based on each of the first imitation samples.

[0011] In one embodiment, the imitation requirement that each entity word in the imitation sentence has the same entity type as the corresponding entity word sample in the sentence sample is performed on the sentence sample to obtain at least one target imitation sentence, including:

[0012] Based on the imitation requirement that each entity word in the imitation sentence is the same as the corresponding entity word sample in the sentence sample, sentence imitation is performed on the sentence sample to obtain at least one first imitation sentence;

[0013] Based on the imitation requirement that the entity type of the imitated entity word is the same as that of the corresponding entity word sample in the sentence sample, and the entity type of each entity word sample, entity word imitation is performed on each entity word sample to obtain an imitation entity word set for each entity word sample;

[0014] Entity words in the imitated entity word set of each entity word sample are used to replace the entity words in each first imitated sentence that are the same as the entity word sample, so as to obtain multiple target imitated sentences.

[0015] In one embodiment, the imitation requirement that each entity word in the imitation sentence is the same as the corresponding entity word sample in the sentence sample is performed on the sentence sample to obtain at least one first imitation sentence, including:

[0016] Classifying the entity triples corresponding to each entity word sample in the sentence sample according to the entity type to obtain entity subsets corresponding to each entity type, wherein the entity triples include the position of the first character and the position of the last character of the corresponding entity word sample in the sentence sample, as well as the entity type of the entity word sample;

[0017] Based on the entity subsets corresponding to the entity types, construct a plurality of mapping relationships from the entity types to the entity subsets, wherein each mapping relationship corresponds to one entity type;

[0018] Generate sentence imitation instructions based on the imitation requirement that each entity word in the imitation sentence is the same as the corresponding entity word sample in the sentence sample, and use the mapping relationship corresponding to each entity type as the entity word prompt word of each entity type;

[0019] The sentence imitation instruction is input into a first large language model to obtain a plurality of first imitation sentences output by the first large language model.

[0020] In one embodiment, the sentence imitation instruction also includes a first imitation task, which is used to indicate the identity role of the first large language model, and to instruct the first large language model to imitate and generate N imitation sentences based on the imitation requirements and the identity role, with reference to the sentence sample, where N is a positive integer greater than or equal to 1.

[0021] In one embodiment, the imitation requirement that the entity word of the imitation entity word is the same as the entity type of the corresponding entity word sample in the sentence sample, and the entity type of each entity word sample, performing entity word imitation on each entity word sample to obtain the imitation entity word set of each entity word sample includes:

[0022] Constructing a writing imitation prompt word for the entity word sample in such a manner that the entity type of the entity word sample in the sentence sample corresponds to the mapping relationship;

[0023] Generate entity word imitation instructions based on the imitation requirement that the entity word to be imitated has the same entity type as the corresponding entity word sample in the sentence sample, and the imitation prompt word of the entity word sample;

[0024] The entity word imitation instruction is input into the second largest language model to obtain the imitation entity word set of the entity word sample output by the second largest language model.

[0025] In one embodiment, the entity word imitation instruction also includes a second imitation task, which is used to instruct the second large language model to imitate and generate M imitation entity words based on the imitation requirements and with reference to the entity word samples, where M is a positive integer greater than or equal to 1.

[0026] In one embodiment, the method of replacing the entity words in the entity word set of each entity word sample with the entity words identical to the entity word sample in each first imitation sentence to obtain multiple target imitation sentences includes:

[0027] In each of the first imitated sentences, the imitated sentences whose entity words are not included in the sentence samples are removed to obtain a first imitated sentence set;

[0028] removing, from the first imitation sentence set, imitation sentences whose similarity to the sentence sample does not meet a preset similarity condition, to obtain a second imitation sentence set;

[0029] Using each imitation entity word in the imitation entity word set corresponding to each entity word sample, replacing the entity word identical to the entity word sample in each imitation sentence in the second imitation sentence set to obtain a third imitation sentence set;

[0030] The plurality of target imitation sentences are determined based on the union of the second imitation sentence set and the third imitation sentence set.

[0031] In one embodiment, determining each first imitation sample based on each target imitation sentence and the entity type of the entity word sample corresponding to each entity word in each target imitation sentence includes:

[0032] Based on the requirement of semantic coherence, the target imitation sentence is continued to be written to obtain a continuation sentence of the target imitation sentence;

[0033] Determining the entity type of each entity word in the target imitation sentence based on the entity type of the entity word sample corresponding to each entity word in the target imitation sentence;

[0034] The first imitation sample is determined based on the target imitation sentence, a continuation sentence of the target imitation sentence, and the entity type of each entity word in the target imitation sentence.

[0035] In one embodiment, the training of the entity recognition model based on each of the first imitation writing samples includes:

[0036] splicing the target imitation sentence and the continuation sentence of the target imitation sentence in the first imitation sample to obtain a spliced ​​sentence;

[0037] Inputting the concatenated sentence into an encoder in the entity recognition model to obtain an encoding vector of the concatenated sentence output by the encoder;

[0038] Extracting the encoding vector of the target imitation sentence from the encoding vectors of the concatenated sentences based on the position information of the target imitation sentence in the concatenated sentences;

[0039] Inputting the encoding vector of the target imitation sentence into the decoder in the entity recognition model, and obtaining the predicted entity type of each entity word in the target imitation sentence output by the decoder;

[0040] Determining a loss function based on the entity type of each entity word in the target imitation sentence included in the first imitation sample and the predicted entity type of each entity word in the target imitation sentence;

[0041] Based on the loss function, model parameters of the entity recognition model are adjusted.

[0042] According to another aspect of the present invention, a material supply chain named entity recognition model training device is provided, comprising:

[0043] A training sample acquisition module is used to obtain training samples of an entity recognition model for an electric power material supply chain, wherein the training samples include sentence samples and entity types of each entity word sample in the sentence samples;

[0044] A sentence imitation module is configured to perform sentence imitation on the sentence sample based on the imitation requirement that the entity types of the entity words in the imitation sentence are the same as those of the corresponding entity word samples in the sentence sample, thereby obtaining at least one target imitation sentence;

[0045] An imitation sample determination module, configured to determine each first imitation sample based on each target imitation sentence and the entity type of the entity word sample corresponding to each entity word in each target imitation sentence;

[0046] A model training module is used to train the entity recognition model based on each of the first imitation samples.

[0047] In one embodiment, the sentence imitation module includes:

[0048] A sentence imitation unit, configured to perform sentence imitation on the sentence sample based on the imitation requirement that each entity word in the imitation sentence is the same as the corresponding entity word sample in the sentence sample, to obtain at least one first imitation sentence;

[0049] An entity word imitation unit is configured to perform entity word imitation on each entity word sample based on the imitation requirement that the entity type of the imitated entity word is the same as that of the corresponding entity word sample in the sentence sample and the entity type of each entity word sample, thereby obtaining an imitation entity word set for each entity word sample;

[0050] The entity word replacement unit is used to replace the entity words in the imitation entity word set of each entity word sample with the entity words in each first imitation sentence that are the same as the entity word sample, so as to obtain multiple target imitation sentences.

[0051] In one embodiment, the sentence imitation unit is specifically used to:

[0052] Classifying the entity triples corresponding to each entity word sample in the sentence sample according to the entity type to obtain entity subsets corresponding to each entity type, wherein the entity triples include the position of the first character and the position of the last character of the corresponding entity word sample in the sentence sample, as well as the entity type of the entity word sample;

[0053] Based on the entity subsets corresponding to the entity types, construct a plurality of mapping relationships from the entity types to the entity subsets, wherein each mapping relationship corresponds to one entity type;

[0054] Generate sentence imitation instructions based on the imitation requirement that each entity word in the imitation sentence is the same as the corresponding entity word sample in the sentence sample, and use the mapping relationship corresponding to each entity type as the entity word prompt word of each entity type;

[0055] The sentence imitation instruction is input into a first large language model to obtain a plurality of first imitation sentences output by the first large language model.

[0056] In one embodiment, the sentence imitation instruction also includes a first imitation task, which is used to indicate the identity role of the first large language model, and to instruct the first large language model to imitate and generate N imitation sentences based on the imitation requirements and the identity role, with reference to the sentence sample, where N is a positive integer greater than or equal to 1.

[0057] In one embodiment, the entity word imitation unit includes:

[0058] Constructing a writing imitation prompt word for the entity word sample in such a manner that the entity type of the entity word sample in the sentence sample corresponds to the mapping relationship;

[0059] Generate entity word imitation instructions based on the imitation requirement that the entity word to be imitated has the same entity type as the corresponding entity word sample in the sentence sample, and the imitation prompt word of the entity word sample;

[0060] The entity word imitation instruction is input into the second largest language model to obtain the imitation entity word set of the entity word sample output by the second largest language model.

[0061] In one embodiment, the entity word imitation instruction also includes a second imitation task, which is used to instruct the second large language model to imitate and generate M imitation entity words based on the imitation requirements and with reference to the entity word samples, where M is a positive integer greater than or equal to 1.

[0062] In one embodiment, the entity word replacement unit is specifically used to:

[0063] From the plurality of first imitative sentences, remove the imitative sentences whose entity words are not included in the sentence samples to obtain a first imitative sentence set;

[0064] removing, from the first imitation sentence set, imitation sentences whose similarity to the sentence sample does not meet a preset similarity condition, to obtain a second imitation sentence set;

[0065] Using each imitation entity word in the imitation entity word set corresponding to each entity word sample, replacing the entity word identical to the entity word sample in each imitation sentence in the second imitation sentence set to obtain a third imitation sentence set;

[0066] The plurality of target imitation sentences are determined based on the union of the second imitation sentence set and the third imitation sentence set.

[0067] In one embodiment, the imitation writing sample determination module includes:

[0068] A sentence continuation unit, configured to continue writing the target imitation sentence based on a semantically coherent continuation requirement to obtain a continuation sentence of the target imitation sentence;

[0069] An entity type determining unit, configured to determine the entity type of each entity word in the target imitation sentence based on the entity type of the entity word sample corresponding to each entity word in the target imitation sentence;

[0070] The imitation sample construction unit is used to determine the first imitation sample based on the target imitation sentence, the continuation sentence of the target imitation sentence, and the entity type of each entity word in the target imitation sentence.

[0071] In one embodiment, the model training module includes:

[0072] a splicing unit, configured to splice the target imitation sentence and the continuation sentence of the target imitation sentence in the first imitation sample to obtain a spliced ​​sentence;

[0073] an encoding unit, configured to input the concatenated sentence into an encoder in the entity recognition model, and obtain an encoding vector of the concatenated sentence output by the encoder;

[0074] A coding vector extraction unit, configured to extract the coding vector of the target imitation sentence from the coding vectors of the concatenated sentence based on position information of the target imitation sentence in the concatenated sentence;

[0075] A decoding unit, configured to input the encoding vector of the target imitation sentence into a decoder in the entity recognition model, and obtain the predicted entity type of each entity word in the target imitation sentence output by the decoder;

[0076] A loss function determining unit, configured to determine a loss function based on the entity type of each entity word in the target imitation sentence included in the first imitation sample and the predicted entity type of each entity word in the target imitation sentence;

[0077] A model parameter adjustment unit is used to adjust the model parameters of the entity recognition model based on the loss function.

[0078] According to another aspect of the present invention, a material supply chain named entity recognition model training system is provided, comprising: at least one processor, and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the processor obtains the instructions from the memory and executes the instructions, so that the at least one processor can execute the material supply chain named entity recognition model training method described in any embodiment of the present invention.

[0079] According to another aspect of the present invention, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are provided to a computer to instruct the computer to execute the material supply chain named entity recognition model training method described in any embodiment of the present invention.

[0080] Using the technical solution of the present invention, a training sample for an entity recognition model of a power material supply chain text is obtained, wherein the training sample includes a sentence sample and the entity type of each entity word sample in the sentence sample. Based on the imitation requirement that each entity word in the imitation sentence has the same entity type as the corresponding entity word sample in the sentence sample, sentence imitation is performed on the sentence sample to obtain multiple target imitation sentences. In this way, multiple imitation sentences can be obtained in which the entity word types are the same as the entity word types in the sentence sample. Based on each target imitation sentence and the entity type of the entity word sample corresponding to each entity word in each target imitation sentence, each first imitation sample is determined. In this way, since the entity types are known, the entity word types in the imitation sentences can be annotated without manually or using a model to identify the type of the entity words in the imitation sentences. Thus, a large number of imitation samples can be quickly generated and annotated. Subsequently, the entity recognition model is trained based on these large number of imitation samples to improve the recognition accuracy of the entity recognition model.

[0081] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present invention, nor is it intended to limit the scope of the present invention. Other features of the present invention will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0082] The accompanying drawings are provided for a better understanding of the present invention and do not constitute a limitation of the present invention.

[0083] Figure 1 This is a flow chart of a method for training a named entity recognition model for a material supply chain according to an embodiment of the present invention;

[0084] Figure 2 This is a structural block diagram of a material supply chain named entity recognition model training device according to an embodiment of the present invention;

[0085] Figure 3 is a block diagram of an electronic device for implementing the method according to an embodiment of the present invention. DETAILED DESCRIPTION

[0086] The following description of exemplary embodiments of the present invention is made in conjunction with the accompanying drawings, and various details of the embodiments of the present invention are included to facilitate understanding. These details should be considered as merely exemplary. Therefore, it should be appreciated by those skilled in the art that various changes and modifications may be made to the embodiments described herein without departing from the scope of the present invention. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0087] Figure 1 The present invention is a flowchart of a method for training a named entity recognition model for a material supply chain according to an embodiment of the present invention.

[0088] like Figure 1 As shown, the material supply chain named entity recognition model training method may include:

[0089] S110, obtaining training samples of an entity recognition model for the power material supply chain, wherein the training samples include sentence samples and entity types of each entity word sample in the sentence samples;

[0090] S120, based on the imitation requirement that each entity word in the imitation sentence has the same entity type as the corresponding entity word sample in the sentence sample, performing sentence imitation on the sentence sample to obtain at least one target imitation sentence;

[0091] S130, determining each first imitation sample based on each target imitation sentence and the entity type of the entity word sample corresponding to each entity word in each target imitation sentence;

[0092] S140: Training an entity recognition model based on each first imitation writing sample.

[0093] In an embodiment of the present invention, when the amount of text samples collected in the power material supply chain is small, that is, when a small sample occurs, the sample construction process of steps S110 to S140 of the embodiment of the present invention can be performed on the existing training samples to obtain multiple first imitation samples. The first imitation samples are added as training samples to the existing training sample set, thereby obtaining a large number of training samples. Using a large number of training samples to train the entity recognition model can improve the entity recognition model's recognition accuracy for entity types of entity words in text sentences in the power material supply chain.

[0094] For example, multiple texts related to the power material supply chain can be collected from multiple platforms, and multiple training samples can be obtained by filtering, sentence segmentation, and entity word tagging of the texts.

[0095] For example, a sentence sample can be a single sentence or a paragraph including multiple sentences. A sentence sample includes multiple entity word samples. An entity word sample can include one or more words.

[0096] For example, for the original sample set The training samples in , based on this, new training samples are constructed to address the problem of shortage of training samples in small sample scenarios. represents a sentence sample, Indicates the entity type of each entity word sample. It may also include multiple entity triples, each entity triple corresponds to an entity word sample, and the entity triple includes the position of the first word of the corresponding entity word sample in the sentence sample and the position of the last word in the sentence sample, as well as the entity type of the entity word sample.

[0097] For example, , Represents entity word samples The entity triplet of Represents an entity sample The entity triplet of Represented as training samples The entity type collection, Entity word samples The entity type, For the first entity sample Entity type. , Represents entity word samples The first letter of In sentence samples location, Represents entity word samples The last word in In the sentence sample location. , Represents entity word samples The first letter of In the sentence sample location, Represents entity word samples The last word in In the sentence sample location.

[0098] Exemplarily, based on the imitation requirement that each entity word in the imitation sentence has the same entity type as the corresponding entity word sample in the sentence sample, an imitation instruction is determined, the imitation instruction is input into the large language model, and at least one target imitation sentence output by the large language model is obtained.

[0099] The at least one target imitation sentence may include one or more target imitation sentences.

[0100] Among them, the large language model can be GPT (Generative Pre-trained Transformer) or other natural language models, such as the Deep Seek model.

[0101] For example, the imitation instruction is:

[0102] {Imitation requirements: The entity types of each entity word in the imitation sentence are the same as those of the corresponding entity word sample in the sentence sample;

[0103] Sample sentence: }.

[0104] For example, the imitation instruction is:

[0105] {Imitation requirements: The entity types of each entity word in the imitation sentence are the same as those of the corresponding entity word sample in the sentence sample;

[0106] Sample sentence: ;

[0107] Entity word prompt: }.

[0108] For example, the imitation requirement that each entity word in the imitation sentence has the same entity type as the corresponding entity word sample in the sentence sample can be divided into multiple sub-imitation requirements and some operations to implement the imitation requirement.

[0109] For example, the first imitation sample may include a target imitation sentence and the entity type of each entity word in the target imitation sentence. The entity type of each entity word in the target imitation sentence may be similar to the above Perform triple construction.

[0110] Exemplarily, the target imitation sentence in the first imitation sample is input into the entity recognition model, and the predicted entity type of each entity word in the target imitation sentence output by the entity recognition model is obtained. Based on the difference between the entity type of each entity word in the target imitation sentence in the first imitation sample and the predicted entity type of each entity word in the target imitation sentence output by the entity recognition model, a loss function is determined, and the model parameters of the entity recognition model are adjusted using the loss function. In this way, the model training is continuously stopped until the model accuracy reaches the preset accuracy requirement or the number of model training times reaches a preset training number threshold.

[0111] According to the above embodiment, a training sample for an entity recognition model of a power material supply chain text is obtained, wherein the training sample includes a sentence sample and the entity type of each entity word sample in the sentence sample; based on the imitation requirement that each entity word in the imitation sentence has the same entity type as the corresponding entity word sample in the sentence sample, sentence imitation is performed on the sentence sample to obtain multiple target imitation sentences. In this way, multiple imitation sentences can be obtained in which the entity word type is the same as the entity word type in the sentence sample. Based on each of the target imitation sentences and the entity type of the entity word sample corresponding to each entity word in each of the target imitation sentences, each first imitation sample is determined. In this way, since the entity type is known, it is not necessary to manually or use a model to identify the type of the entity words in the imitation sentence, and the type of the entity words in the imitation sentence can be annotated. Thus, a large number of imitation samples can be quickly generated and annotated. Subsequently, the entity recognition model is trained based on these large number of imitation samples to improve the recognition accuracy of the entity recognition model.

[0112] In one embodiment, based on the imitation requirement that each entity word in the imitated sentence has the same entity type as the corresponding entity word sample in the sentence sample, sentence imitation is performed on the sentence sample to obtain at least one target imitated sentence, including: based on the imitation requirement that each entity word in the imitated sentence has the same entity type as the corresponding entity word sample in the sentence sample, sentence imitation is performed on the sentence sample to obtain at least one first imitated sentence; based on the imitation requirement that the entity type of the imitated entity word is the same as the corresponding entity word sample in the sentence sample and the entity type of each entity word sample, entity word imitation is performed on each entity word sample to obtain an imitated entity word set of each entity word sample; using the entity words in the imitated entity word set of each entity word sample to replace the entity words in each first imitated sentence that are the same as the entity word sample to obtain multiple target imitated sentences.

[0113] It can be understood that the imitation requirement that each entity word in the imitation sentence is the same as the entity type of the corresponding entity word sample in the sentence sample is divided into two sub-imitation requirements and a replacement operation. That is, the two sub-imitation requirements are: the imitation requirement that each entity word in the imitation sentence is the same as the corresponding entity word sample in the sentence sample, and the imitation requirement that the entity type of the imitation entity word is the same as the entity word sample in the sentence sample. The replacement operation is: using the entity words in the imitation entity word set of each entity word sample to replace the entity words in each first imitation sentence that are the same as the entity word sample.

[0114] Exemplarily, based on the imitation requirements that each entity word in the imitated sentence is identical to the corresponding entity word sample in the sentence sample, and that the imitated sentence has the same semantics as the sentence sample, a sentence imitation instruction is determined, the sentence imitation instruction is input into the large language model, the large language model imitates the sentence sample according to the sentence imitation instruction, and the large language model outputs at least one first imitated sentence.

[0115] For example, the sentence imitation instruction is:

[0116] {Imitation requirements: Each entity word in the imitation sentence is the same as the corresponding entity word sample in the sentence sample, and the imitation sentence has the same semantics as the sentence sample;

[0117] Sample sentence: ;

[0118] Entity word prompt: }.

[0119] Exemplarily, based on the imitation requirement that the entity type of the imitated entity word is the same as that of the corresponding entity word sample in the sentence sample, an entity word imitation instruction is determined, the entity word imitation instruction is input into the large language model, and the large language model imitates the entity word of the sentence sample according to the entity word imitation instruction. The large language model outputs multiple imitated entity words of the entity word sample, that is, the imitated entity word set.

[0120] For example, the entity word imitation instruction is:

[0121] {Imitation requirement: the entity type of the imitated entity word is the same as the corresponding entity word sample in the sentence sample;

[0122] Sample sentence: ;

[0123] Entity word samples Imitation prompt words: }.

[0124] Exemplarily, the entity words in the entity word set of the imitation entity word samples of each entity word sample are used to directly replace the entity words in each first imitation sentence that are the same as the entity word sample, to obtain multiple target imitation sentences, and each target imitation sentence is different from each other. Alternatively, after filtering out the sentences that do not meet the requirements in the multiple first imitation sentences, the entity words that are the same as the entity word sample in each filtered first imitation sentence are replaced to obtain multiple target imitation sentences. Alternatively, after filtering out the sentences that do not meet the requirements in the multiple first imitation sentences, the entity words that are the same as the entity word sample in each filtered first imitation sentence are replaced to obtain multiple second imitation sentences, and the filtered first imitation sentences are combined with the multiple second imitation sentences to obtain multiple target imitation sentences.

[0125] According to the above implementation, sentence imitation is first performed on the sentence sample based on the imitation requirement of unchanged entity words, and at the same time, entity word imitation is performed on each entity word sample based on the imitation requirement of unchanged entity word type to obtain the imitated entity words of each entity word sample. Then, based on each imitated entity word of the entity word sample, the entity words that are the same as the entity word sample in each first imitated sentence obtained by imitation are replaced. In this way, a large number of imitated sentences can be obtained, and the imitated sentences have the same semantics as the sentence samples and the entity word type remains unchanged. Subsequently, when constructing the imitated samples, it is not necessary to manually mark the entity type of the imitated sentences or perform model recognition, and the entity type of each entity word in the imitated sentences in the imitated samples can be determined, thereby improving the efficiency of sample construction.

[0126] In one embodiment, based on the imitation requirement that each entity word in the imitated sentence is the same as the corresponding entity word sample in the sentence sample, sentence imitation is performed on the sentence sample to obtain at least one first imitated sentence, including: classifying the entity triples corresponding to each entity word sample in the sentence sample according to the entity type to obtain an entity subset corresponding to each entity type, wherein the entity triples include the position of the first character of the corresponding entity word sample in the sentence sample and the position of the last character in the sentence sample, as well as the entity type of the entity word sample; based on the entity subset corresponding to each entity type, constructing multiple mapping relationships from the entity type to the entity subset, wherein each mapping relationship corresponds to an entity type; based on the imitation requirement that each entity word in the imitated sentence is the same as the corresponding entity word sample in the sentence sample, and using the mapping relationship corresponding to each entity type as the entity word prompt word for each entity type, generating a sentence imitation instruction; inputting the sentence imitation instruction into a first large language model to obtain multiple first imitated sentences output by the first large language model.

[0127] It is understandable that an entity type corresponds to an entity subset. For example, for entity type , the corresponding entity subset can be recorded as: ,in, Entity Type Entity word samples, Entity word samples The entity triplet of Represents entity word samples The first letter of the sentence sample The position in Entity word samples The last word in the sentence sample The position in.

[0128] It should be noted that the entity word sample Can include multiple different entity words ( When the value is different, the corresponding entity words are different), but the entity types of these entity words are .

[0129] For example, for the entity type , and its corresponding mapping relationship is: .

[0130] For example, the sentence imitation instruction may be:

[0131] {Imitation requirements: The imitation sentence has the same semantics as the sentence sample, and each entity word in the imitation sentence has the same entity word sample as the corresponding entity word sample in the sentence sample;

[0132] Sample sentence: ;

[0133] Entity word prompt: }.

[0134] in, Represents entity type The mapping relationship Entity subset , Represents entity type The mapping relationship Entity subset .

[0135] In one embodiment, the sentence imitation instruction also includes a first imitation task, which is used to indicate the identity role of the first large language model, and to instruct the first large language model to imitate and generate N imitation sentences based on the imitation requirements and identity role, with reference to sentence samples, where N is a positive integer greater than or equal to 1.

[0136] For example, the sentence imitation instruction may be:

[0137] {Imitation task: Your role is a Chinese teacher. Your task is to imitate sentences based on sample sentences and generate 10 imitation sentences;

[0138] Imitation requirements: The imitation sentence has the same semantics as the sample sentence, and each entity word in the imitation sentence has the same entity word sample as the corresponding entity word sample in the sample sentence;

[0139] Sample sentence: ;

[0140] Entity word prompt: }.

[0141] For example, by using a preset template and filling in the specific imitation task, imitation requirements, sentence samples and prompt words into the template, corresponding sentence imitation instructions can be obtained.

[0142] Exemplarily, a sentence imitation instruction is input into the first large language model, the first large language model generates a plurality of first imitation sentences according to the sentence imitation instruction, and the first large language model outputs a plurality of first imitation sentences.

[0143] According to the above-mentioned implementation, based on the imitation requirement that each entity word in the imitated sentence is the same as the corresponding entity word sample in the sentence sample, sentence imitation is performed on the sentence sample to generate multiple imitated sentences. In this way, the entity type of the entity words in these imitated sentences can be determined based on the corresponding entity word sample in the sentence sample, without the need for manual labeling, thereby improving the construction efficiency of the training samples.

[0144] In one embodiment, based on the imitation requirement that the entity type of the imitated entity word is the same as that of the corresponding entity word sample in the sentence sample, and the entity type of each entity word sample, entity word imitation is performed on each entity word sample to obtain an imitated entity word set for each entity word sample, including: constructing an imitation prompt word for the entity word sample in a manner that the entity type of the entity word sample is a corresponding mapping relationship in the sentence sample; generating an entity word imitation instruction based on the imitation requirement that the entity type of the imitated entity word is the same as that of the corresponding entity word sample in the sentence sample, and the imitation prompt word of the entity word sample; inputting the entity word imitation instruction into the second largest language model to obtain an imitated entity word set for the entity word sample output by the second largest language model.

[0145] For example, the imitation prompt words can be: In the entity word sample Is an entity type is Entity, where Alternatively, the entity prompt word can be: In the entity word sample The entity type is ,in, .

[0146] For example, the entity word imitation instruction may be:

[0147] {Imitation requirements: The entity type of the imitated entity word is the same as the corresponding entity word sample in the sentence sample;

[0148] Imitation prompt words: in the sentence sample In the entity word sample Is an entity type is Entity}.

[0149] In one embodiment, the entity word imitation instruction also includes a second imitation task, which is used to instruct the second largest language model to imitate and generate M imitation entity words based on the imitation requirements and with reference to the entity word samples, where M is a positive integer greater than or equal to 1.

[0150] For example, the entity word imitation instruction may be:

[0151] {Imitation requirements: The entity type of the imitated entity word is the same as the corresponding entity word sample in the sentence sample;

[0152] Imitation prompt words: in the sentence sample In the entity word sample Is an entity type is Entity, where ;

[0153] Imitation task: Targeting entity word samples , please generate M imitation entity words}.

[0154] For example, the imitation writing requirement can be integrated into the imitation writing task. For example, the entity word imitation writing instruction can be:

[0155] {Imitation prompt words: in the sentence sample In the entity word sample Is an entity type is Entity, where ;

[0156] Imitation task: Targeting entity word samples , please generate M entity types and entity word samples Same imitation entity word}.

[0157] For example, by using a preset template, the imitation writing task, imitation writing requirements, sentence samples and prompt words of each entity word sample are filled into the template, and the entity word imitation writing instruction of each entity word sample can be obtained.

[0158] Exemplarily, the entity word imitation instruction of the entity word sample is input into the second largest language model, the second largest language model generates multiple imitation entity words of the entity word sample according to the entity word imitation instruction, and the second largest language model outputs the imitation entity word set of the entity word sample.

[0159] Illustratively, the first largest language model may be the same as or different from the second largest language model.

[0160] According to the above embodiment, based on the imitation requirement that the entity type of the imitated entity word is the same as that of the corresponding entity word sample in the sentence sample, and the entity type of each entity word sample, entity word imitation is performed on each entity word sample to obtain an imitated entity word set of each entity word sample. In this way, the entity type of each entity word in the imitated entity word set of the entity word sample is the same as that of the entity word sample. Subsequently, for the first imitated sentence obtained by the above-mentioned imitation, the entity word samples of the same type can be replaced by the entity words in the imitated entity word set to generate multiple second imitated sentences with the same semantics and unchanged entity types.

[0161] In one embodiment, entity words in the imitated entity word set of each entity word sample are used to replace the entity words in each first imitated sentence that are identical to the entity word sample to obtain multiple target imitated sentences, including: among multiple first imitated sentences, imitated sentences whose entity words are not included in the sentence sample are removed to obtain a first imitated sentence set; in the first imitated sentence set, imitated sentences whose similarity with the sentence sample does not meet the preset similarity conditions are removed to obtain a second imitated sentence set; using each imitated entity word in the imitated entity word set corresponding to each entity word sample, entity words in each imitated sentence in the second imitated sentence set that are identical to the entity word sample are replaced to obtain a third imitated sentence set; based on the union of the second imitated sentence set and the third imitated sentence set, multiple target imitated sentences are determined.

[0162] In this example, for the training samples , for sentence samples , imitate and get multiple first imitated sentences , can be represented by a set of imitative sentences: First, the imitation sentences that do not meet the requirements are filtered out from the collection. The filtering process can include rule filtering, semantic similarity filtering and semantic coherence similarity filtering. Then, sentence samples are used to Entity word samples in The imitation entity word set , replace entity words on each imitated sentence in the filtered imitated sentence set to obtain multiple second imitated sentences, merge the multiple second imitated sentences into the filtered imitated sentence set and perform union to obtain a set including multiple target imitated sentences.

[0163] Exemplarily, the multiple imitation sentences obtained above may be filtered using any one of the following methods or multiple methods simultaneously.

[0164] For example, the rule filtering of the imitation sentences may be as follows: for the first imitation sentence , if there is an entity word Not included in the sentence sample , then filter out , get the set of imitation sentences after rule filtering . Then we have: , , .

[0165] For example, the process of filtering the semantic similarity of the imitation sentences can be as follows:

[0166] If you imitate the sentence Compared with the original sentence The greater the change in , the greater the diversity of the newly constructed samples. Therefore, in order to measure the degree of change in the imitated sentence, it is necessary to calculate the semantic similarity between the imitated sentence and the sample sentence.

[0167] For length Sentence sample With length The first imitation sentence , respectively input these two into the pre-trained language model BERT to obtain sentence samples The encoding vector of and the first imitation sentence The encoding vector of Then, by performing average pooling on these two encoding vectors, we get the sentence sample and the first imitation sentence Vector representation of and .

[0168] For example, taking the sentence sample , which is represented by the vector The calculation process can be as follows:

[0169] ;

[0170] ;

[0171] in, Represents a sentence sample Chinese characters The encoding vector of Represents a sentence sample Chinese characters The encoding vector of Represents the function corresponding to the pre-trained language model BERT.

[0172] For sentence samples and the first imitation sentence Perform inner product operation on the vector representation of and the first imitation sentence In some examples, the similarity score can be normalized to interval.

[0173] Therefore, the plurality of first imitation sentences may be filtered according to the similarity scores, for example, imitation sentences with similarity scores less than a preset threshold may be removed.

[0174] For example, the process of filtering the semantic coherence similarity of the imitation sentences can be as follows:

[0175] The perplexity absolute difference (PerplexityAbsoluteDifference) can be used to measure the semantic coherence similarity between sentences. Perplexity is an indicator used to evaluate the degree of semantic coherence of sentences. The more semantically coherent a sentence is, the lower its perplexity.

[0176] ;

[0177] = ;

[0178] in, Represents a sentence sample The probability of Represents a sentence sample The product of the probability of each word appearing together, that is, the sentence sample The probability of Represents a sentence sample Chinese subtitles The probability of occurrence, Represents a sentence sample Chinese subtitles and The joint occurrence probability of Represents a sentence sample Chinese subtitles ,Character ,···,Character The joint occurrence probability of .

[0179] in, Represents a sentence sample degree of confusion.

[0180] Therefore, from the above two formulas, we can see that the sentence sample The greater the probability, the smaller the perplexity, indicating that the sentence sample The higher the degree of semantic coherence.

[0181] For the first imitation sentence The perplexity can also be similar to the sentence sample The perplexity is calculated by the method of calculating , so we will not give examples one by one here.

[0182] If the first imitation sentence With sentence samples The more consistent the writing style is, the smaller the absolute difference in confusion between the two is. With sentence samples The absolute difference of confusion is calculated as follows:

[0183] ;

[0184] in, Indicates the first imitation sentence With sentence samples The confusion level is absolutely bad. Indicates a very small error value.

[0185] For the imitation sentence collection ,like , then directly keep the imitation sentence Otherwise, calculate the first imitation sentences one by one. With sentence samples The confusion is absolutely bad. If the first imitation sentence With sentence samples The absolute difference of confusion is recorded as , then the first imitation sentence With sentence samples The normalized absolute difference of perplexity is as follows:

[0186] .

[0187] Among them, if The smaller it is, the better the first imitation sentence is. Compared with the original sentence sample Have similar writing styles.

[0188] Therefore, the imitated sentences whose absolute difference in normalized perplexity with the sentence sample is greater than a preset threshold are removed from the plurality of first imitated sentences, so as to filter the imitated sentences based on semantic coherence similarity.

[0189] For example, the balanced process of filtering the semantic similarity and semantic coherence similarity of the imitation sentence is as follows:

[0190] Original sentence sample and the first imitation sentence The similarity is , the standardized perplexity absolute difference is .like The smaller it is, the better it is at imitating the sentence. Compared with the original sentence There are more changes to increase the diversity of samples; if The smaller it is, the better it is at imitating the sentence. Compared with the original sentence Having a similar writing style can avoid introducing more noise.

[0191] For example, the balanced filtering index of the first imitation sentence is:

[0192] ;

[0193] in, Represents the balanced filtering index of the first imitated sentence.

[0194] The imitation sentence set composed of the above multiple first imitation sentences Each imitation sentence in Value, if the number of retention is , then select Keep The smallest value imitated sentences, and remove other imitated sentences to get the filtered imitated sentence set .

[0195] The following introduces the filtered imitation sentence set The process of replacing entity words in each imitation sentence is as follows:

[0196] For training samples , sentence sample The first imitation sentence It must include The entities corresponding to all entity types in . For sentence samples Entity word samples in , and its imitation entity word set is ,from Randomly select a similar entity , for the first imitation sentence Samples of Chinese and entity words The same entity words are replaced. In this way, multiple new imitation sentences can be obtained. Then, the first imitation sentence after filtering is The imitation sentences after replacing entity words are all merged into the same set, obtaining a set including multiple target imitation sentences.

[0197] For example, for the sentence sample "This is the material procurement department of XXX1 company", the first imitation sentence is "This is the material procurement department of XXX1 company". If the imitation entity word of the entity word "XXX1 company" is "XXX2 company" and the imitation entity word of "material procurement department" is "material supply department", then these two imitation entity words are used to replace the corresponding entity words in the first imitation sentence. After replacing the entity words, the target imitation sentence obtained is "material supply department of XX2 company".

[0198] In actual application, imitate the sentence based on the target The entity types of the entity word samples corresponding to each entity word in the target imitation sentence are used as the entity types of each entity word in the target imitation sentence. Combining the positions of the first and last characters of each entity word in the target imitation sentence and the entity types of each entity word, the target imitation sentence can be obtained. Entity collection In this way, the new imitation sample can be .

[0199] According to the above embodiment, multiple first imitation sentences are filtered, and entity words in the imitation entity word set of each entity word sample are used to replace the entity words in each first imitation sentence obtained after filtering that are identical to the entity word sample. Then, each filtered first imitation sentence and each first imitation sentence after entity word replacement are merged into the same set to obtain multiple target imitation sentences. In this way, a larger number of target imitation sentences with a higher similarity to the original sentence samples can be obtained.

[0200] In one embodiment, each first imitation sample is determined based on each target imitation sentence and the entity type of the entity word sample corresponding to each entity word in each target imitation sentence, including: based on the continuation requirement of semantic coherence, the target imitation sentence is continued to obtain a continued sentence of the target imitation sentence; based on the entity type of the entity word sample corresponding to each entity word in the target imitation sentence, the entity type of each entity word in the target imitation sentence is determined; based on the target imitation sentence, the continued sentence of the target imitation sentence, and the entity type of each entity word in the target imitation sentence, the first imitation sample is determined.

[0201] For example, a continuation instruction may be determined based on the semantically coherent continuation requirement and the target imitation sentence, and the continuation instruction may be input into the large language model to obtain a continuation sentence of the target imitation sentence output by the large language model.

[0202] For example, when the context of a sentence is insufficient, it is easy to cause ambiguity in named entity recognition. For example, for the sentence "UHV equipment", due to the lack of relevant description, it may be mistakenly recognized as a similar noun in other fields. However, this example continues it through a large language model, which can introduce the prior knowledge of the large language model and remove the ambiguity problem of entity recognition. For example, if the sentence "UHV equipment" is continued as "UHV equipment is a key facility in the power system", then "UHV equipment" can only be classified by the named entity recognition model as a material equipment entity in the field of the power material supply chain. In this way, applying the continued sentence of the original sentence to the training process of the named entity recognition model can improve the recognition accuracy of the model.

[0203] For example, the continue instruction may be:

[0204] {To the sentence To continue writing, just output the continued sentence}.

[0205] According to the above embodiment, by continuing the imitation sentence and subsequently applying the continued sentence to the training process of the named entity recognition model, the recognition accuracy of the named entity recognition model can be improved.

[0206] In one embodiment, the large language model used in the above example can be ChatGPT (gpt-3.5-turbo-0613). The temperature parameter in the large language model is used to control the diversity of text generated by the large model and is set to 0.5. The presence penalty parameter in the large language model is used to control the repetitiveness of text generated by the large model and is set to 0. For sentence imitation, entity word imitation, and sentence continuation, the maximum number of tokens in the text generated by the large language model is set to 2000, 300, and 120, respectively.

[0207] In one embodiment, an entity recognition model is trained based on each first imitation sample, including: splicing a target imitation sentence and a continuation sentence of the target imitation sentence in the first imitation sample to obtain a spliced ​​sentence; inputting the spliced ​​sentence into an encoder in the entity recognition model to obtain an encoding vector of the spliced ​​sentence output by the encoder; extracting the encoding vector of the target imitation sentence from the encoding vector of the spliced ​​sentence based on the position information of the target imitation sentence in the spliced ​​sentence; inputting the encoding vector of the target imitation sentence into a decoder in the entity recognition model to obtain the predicted entity type of each entity word in the target imitation sentence output by the decoder; determining a loss function based on the entity type of each entity word in the target imitation sentence included in the first imitation sample and the predicted entity type of each entity word in the target imitation sentence; and adjusting the model parameters of the entity recognition model based on the loss function.

[0208] During the training process, for the original training sample set , after the above sample construction, a new training sample set is obtained Then, The semantic expansion of the imitation sentences in all the imitation samples in the training set is obtained For new samples , the original sentence Continuing the sentence Perform splicing to obtain a spliced ​​sentence .in, Indicates a special delimiter.

[0209] For example, a masked language modeling (MLM) encoding model is used as the encoder to concatenate the sentences Encode, if the length of the original sentence T is , continue the sentence The length is ,but The embedding representation of .from Extract the corresponding original sentence The part of the encoding vector embedded representation .

[0210] For example, a sequence annotation method is used to perform the named entity recognition task. If the named entity sequence decoding layer is , new sample Central Plains Sentences The sequence labeling probability is , the original sentence is obtained by computing the decoding layer Each character in The predicted probability of annotation for:

[0211] .

[0212] For example, the model is trained by maximum likelihood estimation, and the objective loss function of the training is:

[0213] .

[0214] In practice, the learning rate for encoding layer parameters can be set to 1e-5, and the learning rate for decoding layer parameters can be set to 1e-3. The learning algorithm is AdamW, with a weight decay rate of 0.01. A linear learning rate warmup is used, with a warmup rate of 0.1. The maximum sentence length for each dataset is uniformly set to 350, the training batch size is 2, and the number of training epochs is 30.

[0215] According to the above embodiment, by splicing the imitation sentence and its continuation sentence and inputting them into the encoder, the encoder can combine the entire sentence context to encode the sentence, thus obtaining a more accurate encoding result. Then, a decoder is used to decode the encoding vector corresponding to the imitation sentence in the encoding result. Because the encoding is combined with the content of the continuation sentence, the decoding result of the encoding vector of the imitation sentence is more accurate, thus achieving higher recognition accuracy of the trained named entity recognition model.

[0216] Figure 2 It is a structural block diagram of a material supply chain named entity recognition model training device according to an embodiment of the present invention.

[0217] like Figure 2 As shown, the material supply chain named entity recognition model training device includes:

[0218] A training sample acquisition module 210 is configured to acquire training samples for an entity recognition model of an electric power material supply chain, wherein the training samples include sentence samples and entity types of entity word samples in the sentence samples;

[0219] A sentence imitation module 220 is configured to perform sentence imitation on the sentence sample based on the imitation requirement that the entity types of the entity words in the imitation sentence are the same as those of the corresponding entity word samples in the sentence sample, to obtain at least one target imitation sentence;

[0220] The imitation sample determination module 230 is configured to determine each first imitation sample based on each target imitation sentence and the entity type of the entity word sample corresponding to each entity word in each target imitation sentence;

[0221] The model training module 240 is used to train the entity recognition model based on each of the first imitation samples.

[0222] In one embodiment, the sentence imitation module includes:

[0223] A sentence imitation unit, configured to perform sentence imitation on the sentence sample based on the imitation requirement that each entity word in the imitation sentence is the same as the corresponding entity word sample in the sentence sample, to obtain at least one first imitation sentence;

[0224] An entity word imitation unit is configured to perform entity word imitation on each entity word sample based on the imitation requirement that the entity type of the imitated entity word is the same as that of the corresponding entity word sample in the sentence sample and the entity type of each entity word sample, thereby obtaining an imitation entity word set for each entity word sample;

[0225] The entity word replacement unit is used to replace the entity words in the imitation entity word set of each entity word sample with the entity words in each first imitation sentence that are the same as the entity word sample, so as to obtain multiple target imitation sentences.

[0226] In one embodiment, the sentence imitation unit is specifically used to:

[0227] Classifying the entity triples corresponding to each entity word sample in the sentence sample according to the entity type to obtain entity subsets corresponding to each entity type, wherein the entity triples include the position of the first character and the position of the last character of the corresponding entity word sample in the sentence sample, as well as the entity type of the entity word sample;

[0228] Based on the entity subsets corresponding to the entity types, construct a plurality of mapping relationships from the entity types to the entity subsets, wherein each mapping relationship corresponds to one entity type;

[0229] Generate sentence imitation instructions based on the imitation requirement that each entity word in the imitation sentence is the same as the corresponding entity word sample in the sentence sample, and use the mapping relationship corresponding to each entity type as the entity word prompt word of each entity type;

[0230] The sentence imitation instruction is input into a first large language model to obtain a plurality of first imitation sentences output by the first large language model.

[0231] In one embodiment, the sentence imitation instruction also includes a first imitation task, which is used to indicate the identity role of the first large language model, and to instruct the first large language model to imitate and generate N imitation sentences based on the imitation requirements and the identity role, with reference to the sentence sample, where N is a positive integer greater than or equal to 1.

[0232] In one embodiment, the entity word imitation unit includes:

[0233] Constructing a writing imitation prompt word for the entity word sample in such a manner that the entity type of the entity word sample in the sentence sample corresponds to the mapping relationship;

[0234] Generate entity word imitation instructions based on the imitation requirement that the entity word to be imitated has the same entity type as the corresponding entity word sample in the sentence sample, and the imitation prompt word of the entity word sample;

[0235] The entity word imitation instruction is input into the second largest language model to obtain the imitation entity word set of the entity word sample output by the second largest language model.

[0236] In one embodiment, the entity word imitation instruction also includes a second imitation task, which is used to instruct the second large language model to imitate and generate M imitation entity words based on the imitation requirements and with reference to the entity word samples, where M is a positive integer greater than or equal to 1.

[0237] In one embodiment, the entity word replacement unit is specifically used to:

[0238] From the plurality of first imitative sentences, remove the imitative sentences whose entity words are not included in the sentence samples to obtain a first imitative sentence set;

[0239] removing, from the first imitation sentence set, imitation sentences whose similarity to the sentence sample does not meet a preset similarity condition, to obtain a second imitation sentence set;

[0240] Using each imitation entity word in the imitation entity word set corresponding to each entity word sample, replacing the entity word identical to the entity word sample in each imitation sentence in the second imitation sentence set to obtain a third imitation sentence set;

[0241] The plurality of target imitation sentences are determined based on the union of the second imitation sentence set and the third imitation sentence set.

[0242] In one embodiment, the imitation writing sample determination module includes:

[0243] A sentence continuation unit, configured to continue writing the target imitation sentence based on a semantically coherent continuation requirement to obtain a continuation sentence of the target imitation sentence;

[0244] An entity type determining unit, configured to determine the entity type of each entity word in the target imitation sentence based on the entity type of the entity word sample corresponding to each entity word in the target imitation sentence;

[0245] The imitation sample construction unit is used to determine the first imitation sample based on the target imitation sentence, the continuation sentence of the target imitation sentence, and the entity type of each entity word in the target imitation sentence.

[0246] In one embodiment, the model training module includes:

[0247] a splicing unit, configured to splice the target imitation sentence and the continuation sentence of the target imitation sentence in the first imitation sample to obtain a spliced ​​sentence;

[0248] an encoding unit, configured to input the concatenated sentence into an encoder in the entity recognition model, and obtain an encoding vector of the concatenated sentence output by the encoder;

[0249] A coding vector extraction unit, configured to extract the coding vector of the target imitation sentence from the coding vectors of the concatenated sentence based on position information of the target imitation sentence in the concatenated sentence;

[0250] A decoding unit, configured to input the encoding vector of the target imitation sentence into a decoder in the entity recognition model, and obtain the predicted entity type of each entity word in the target imitation sentence output by the decoder;

[0251] A loss function determining unit, configured to determine a loss function based on the entity type of each entity word in the target imitation sentence included in the first imitation sample and the predicted entity type of each entity word in the target imitation sentence;

[0252] A model parameter adjustment unit is used to adjust the model parameters of the entity recognition model based on the loss function.

[0253] For the description of specific functions and examples of each module and submodule of the system in the embodiment of the present invention, please refer to the relevant description of the corresponding steps in the above method embodiment, which will not be repeated here.

[0254] In the technical solution of the present invention, the acquisition, storage and application of user personal information involved are in compliance with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0255] According to an embodiment of the present invention, the present invention further provides a system and a readable storage medium.

[0256] For example, an embodiment of the present invention provides a material supply chain named entity recognition model training system, comprising: at least one processor, and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the processor retrieves and executes the instructions from the memory, so that the at least one processor can execute the material supply chain named entity recognition model training method described in any embodiment of the present invention. This system can be applied to electronic devices.

[0257] Exemplarily, an embodiment of the present invention provides a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are provided to a computer to instruct the computer to execute the material supply chain named entity recognition model training method described in any embodiment of the present invention.

[0258] Figure 3A schematic block diagram of an example electronic device 800 that can be used to implement embodiments of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or claimed herein.

[0259] like Figure 3 As shown, electronic device 800 includes a computing unit 801, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 802 or a computer program loaded from a storage unit 808 into a random access memory (RAM) 803. RAM 803 may also store various programs and data required for the operation of electronic device 800. Computing unit 801, ROM 802, and RAM 803 are interconnected via a bus 804. An input / output (I / O) interface 805 is also connected to bus 804.

[0260] Multiple components in the electronic device 800 are connected to the I / O interface 805, including an input unit 806, such as a keyboard, a mouse, etc.; an output unit 807, such as various types of displays, speakers, etc.; a storage unit 808, such as a magnetic disk, an optical disk, etc.; and a communication unit 809, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 809 allows the electronic device 800 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0261] The computing unit 801 can be any general-purpose and / or specialized processing component with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 801 performs the various methods and processes described above, such as the material supply chain named entity recognition model training method. For example, in some embodiments, the material supply chain named entity recognition model training method can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 800 via the ROM 802 and / or the communication unit 809. When the computer program is loaded into the RAM 803 and executed by the computing unit 801, one or more steps of the material supply chain named entity recognition model training method described above can be performed. Alternatively, in other embodiments, the computing unit 801 may be configured to execute the material supply chain named entity recognition model training method in any other appropriate manner (for example, by means of firmware).

[0262] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system comprising at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0263] The program code for implementing the method of the present invention can be written in any combination of one or more programming languages. Such program code can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0264] In the context of the present invention, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or apparatus. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of machine-readable storage media may include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), optical fibers, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0265] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0266] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.

[0267] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact through a communication network. The client-server relationship arises through computer programs running on the respective computers and having a client-server relationship with each other. The server may be a cloud server, a server in a distributed system, or a server integrated with a blockchain.

[0268] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in the present invention can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in the present invention can be achieved. This is not limited herein.

[0269] The above specific embodiments do not limit the scope of protection of the present invention. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the principles of the present invention are intended to be included within the scope of protection of the present invention.

Claims

1. A material supply chain named entity recognition model training method, characterized in that: include: Acquire training samples of an entity recognition model for an electric power material supply chain, wherein the training samples include sentence samples and entity types of each entity word sample in the sentence samples; Based on the imitation requirement that the entity types of each entity word in the imitated sentence are the same as those of the corresponding entity word sample in the sentence sample, sentence imitation is performed on the sentence sample to obtain at least one target imitated sentence, including: classifying the entity triples corresponding to each entity word sample in the sentence sample according to the entity type to obtain entity subsets corresponding to each entity type, wherein the entity triples include the position of the first word in the corresponding entity word sample in the sentence sample and the position of the last word in the sentence sample, as well as the entity type of the entity word sample; based on the entity subsets corresponding to each entity type, multiple mapping relationships from entity type to entity subset are constructed, wherein each mapping relationship corresponds to an entity type; based on the entity subsets corresponding to each entity word in the imitated sentence and the entity type, multiple mapping relationships are constructed from entity type to entity subset, wherein each mapping relationship corresponds to an entity type; based on the entity triples corresponding to the entity word in the imitated sentence and the entity type, multiple mapping relationships are constructed from entity type to entity subset, wherein each mapping relationship corresponds to an entity type. The method comprises the following steps: generating a sentence imitation instruction based on the imitation requirement that the entity word samples corresponding to the sentence samples are the same, and using the mapping relationship corresponding to each entity type as the entity word prompt word of each entity type; inputting the sentence imitation instruction into a first large language model to obtain a plurality of first imitation sentences output by the first large language model; performing entity word imitation on each entity word sample based on the imitation requirement that the entity type of the imitated entity word is the same as that of the corresponding entity word sample in the sentence sample, and the entity type of each entity word sample, to obtain an imitation entity word set of each entity word sample; using the entity words in the imitation entity word set of each entity word sample to replace the entity words in each first imitation sentence that are the same as the entity word sample, to obtain a plurality of target imitation sentences; Determine each first imitation sample based on each target imitation sentence and the entity type of the entity word sample corresponding to each entity word in each target imitation sentence; The entity recognition model is trained based on each of the first imitation samples.

2. The method according to claim 1, characterized in that The sentence imitation instruction also includes a first imitation task, which is used to indicate the identity role of the first large language model, and to instruct the first large language model to imitate and generate N imitation sentences based on the imitation requirements and the identity role, with reference to the sentence sample, where N is a positive integer greater than or equal to 1.

3. The method according to claim 1, characterized in that The imitation requirement based on the entity type of the imitation entity word is the same as that of the corresponding entity word sample in the sentence sample, and the entity type of each entity word sample, performing entity word imitation on each entity word sample to obtain an imitation entity word set of each entity word sample, includes: Constructing a writing imitation prompt word for the entity word sample in such a manner that the entity type of the entity word sample in the sentence sample corresponds to the mapping relationship; Generate entity word imitation instructions based on the imitation requirement that the entity word to be imitated has the same entity type as the corresponding entity word sample in the sentence sample, and the imitation prompt word of the entity word sample; The entity word imitation instruction is input into the second largest language model to obtain the imitation entity word set of the entity word sample output by the second largest language model.

4. The method according to claim 3, characterized in that The entity word imitation instruction also includes a second imitation task, which is used to instruct the second large language model to imitate and generate M imitation entity words based on the imitation requirements and with reference to the entity word samples, where M is a positive integer greater than or equal to 1.

5. The method according to claim 1, characterized in that The method of using entity words in the entity word set of each entity word sample to replace the entity words in each first imitation sentence that are identical to the entity word sample to obtain multiple target imitation sentences includes: In each of the first imitated sentences, the imitated sentences whose entity words are not included in the sentence samples are removed to obtain a first imitated sentence set; removing, from the first imitation sentence set, imitation sentences whose similarity to the sentence sample does not meet a preset similarity condition, to obtain a second imitation sentence set; Using each imitation entity word in the imitation entity word set corresponding to each entity word sample, replacing the entity word identical to the entity word sample in each imitation sentence in the second imitation sentence set to obtain a third imitation sentence set; The plurality of target imitation sentences are determined based on the union of the second imitation sentence set and the third imitation sentence set.

6. The method according to claim 1, wherein The determining of each first imitation sample based on each target imitation sentence and the entity type of the entity word sample corresponding to each entity word in each target imitation sentence includes: Based on the requirement of semantic coherence, the target imitation sentence is continued to be written to obtain a continuation sentence of the target imitation sentence; Determining the entity type of each entity word in the target imitation sentence based on the entity type of the entity word sample corresponding to each entity word in the target imitation sentence; The first imitation sample is determined based on the target imitation sentence, a continuation sentence of the target imitation sentence, and the entity type of each entity word in the target imitation sentence.

7. The method according to claim 6, characterized in that The training of the entity recognition model based on each of the first imitation writing samples includes: splicing the target imitation sentence and the continuation sentence of the target imitation sentence in the first imitation sample to obtain a spliced ​​sentence; Inputting the concatenated sentence into an encoder in the entity recognition model to obtain an encoding vector of the concatenated sentence output by the encoder; Extracting the encoding vector of the target imitation sentence from the encoding vectors of the concatenated sentences based on the position information of the target imitation sentence in the concatenated sentences; Inputting the encoding vector of the target imitation sentence into the decoder in the entity recognition model, and obtaining the predicted entity type of each entity word in the target imitation sentence output by the decoder; Determining a loss function based on the entity type of each entity word in the target imitation sentence included in the first imitation sample and the predicted entity type of each entity word in the target imitation sentence; Based on the loss function, model parameters of the entity recognition model are adjusted.

8. A material supply chain named entity recognition model training device, characterized in that: include: A training sample acquisition module is used to obtain training samples of an entity recognition model for an electric power material supply chain, wherein the training samples include sentence samples and entity types of each entity word sample in the sentence samples; A sentence imitation module is configured to perform sentence imitation on the sentence sample based on the imitation requirement that the entity types of the entity words in the imitation sentence are the same as those of the corresponding entity word samples in the sentence sample, thereby obtaining at least one target imitation sentence; An imitation sample determination module, configured to determine each first imitation sample based on each target imitation sentence and the entity type of the entity word sample corresponding to each entity word in each target imitation sentence; A model training module, configured to train the entity recognition model based on each of the first imitation writing samples; Wherein, the sentence imitation module includes: A sentence imitation unit, configured to perform sentence imitation on the sentence sample based on the imitation requirement that each entity word in the imitation sentence is the same as the corresponding entity word sample in the sentence sample, to obtain at least one first imitation sentence; An entity word imitation unit is configured to perform entity word imitation on each entity word sample based on the imitation requirement that the entity type of the imitated entity word is the same as that of the corresponding entity word sample in the sentence sample and the entity type of each entity word sample, thereby obtaining an imitation entity word set for each entity word sample; An entity word replacement unit, configured to replace the entity words in each of the first imitation sentences that are identical to the entity word samples with entity words in the imitation entity word set of each of the entity word samples, to obtain a plurality of target imitation sentences; The sentence imitation unit is specifically used for: Classifying the entity triples corresponding to each entity word sample in the sentence sample according to the entity type to obtain entity subsets corresponding to each entity type, wherein the entity triples include the position of the first character and the position of the last character of the corresponding entity word sample in the sentence sample, as well as the entity type of the entity word sample; Based on the entity subsets corresponding to the entity types, construct a plurality of mapping relationships from the entity types to the entity subsets, wherein each mapping relationship corresponds to one entity type; Generate sentence imitation instructions based on the imitation requirement that each entity word in the imitation sentence is the same as the corresponding entity word sample in the sentence sample, and use the mapping relationship corresponding to each entity type as the entity word prompt word of each entity type; The sentence imitation instruction is input into a first large language model to obtain a plurality of first imitation sentences output by the first large language model.

9. A material supply chain named entity recognition model training system, characterized by: include: at least one processor, and a memory communicatively coupled to the at least one processor; In which, the memory stores instructions that can be executed by the at least one processor, and the processor obtains the instructions from the memory and executes the instructions, so that the at least one processor can execute the material supply chain named entity recognition model training method described in any one of claims 1-7.

10. A non-transitory computer-readable storage medium storing computer instructions, characterized in that: The computer instructions are used to be provided to a computer to instruct the computer to execute the material supply chain named entity recognition model training method described in any one of claims 1-7.

Citation Information

Patent Citations

  • Named entity recognition neural structure in reading understanding form

    CN114648113A

  • Target recognition model sample collecting and training method for small sample problem

    CN119964062A