Material supply chain named entity recognition model training method and system

Through imitation sentence technology, the imitation sentences are generated and the model is trained, which solves the problem that the naming entity recognition model in the power supply chain lacks high-quality labeled data, improves the recognition accuracy of the model, and realizes efficient training in small samples.

CN120354854AActive Publication Date: 2025-07-22STATE GRID DIGITAL TECHNOLOGY HOLDING CO LTD +2
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510813641.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-18
Publication Date
2025-07-22
Estimated Expiration
2045-06-18

AI Technical Summary

Technical Problem

In the field of power supply chain, the naming entity recognition model lacks high-quality labeling data, resulting in insufficient recognition accuracy. The large language model has hallucinations problems in the naming entity recognition task and is difficult to directly apply.

Method used

Through imitation sentence technology, imitation sentences with the same type of entity word in the original sentence are generated, and a named entity recognition model is trained based on these imitation sentences, and sentence imitation and physical word imitation are used to use large language models to construct imitation samples to improve the number and quality of training samples.

Benefits of technology

Quickly generate a large number of labeled samples in the case of small samples, which improves the recognition accuracy of the named entity recognition model, solves the problem of scarcity of high-quality labeled data, and enhances the recognition accuracy of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120354854A_ABST
    Figure CN120354854A_ABST
Patent Text Reader

Abstract

The invention provides a material supply chain named entity recognition model training method and system. According to the implementation scheme, training samples of the entity recognition model of the electric power material supply chain are obtained, and the training samples comprise sentence samples and entity types of all entity word samples in the sentence samples; performing sentence imitation on the sentence sample based on an imitation requirement that each entity word in the imitation sentence is the same as the entity type of the corresponding entity word sample in the sentence sample to obtain at least one target imitation sentence; based on each target imitation sentence and an entity type of an entity word sample corresponding to each entity word in each target imitation sentence, respectively determining each first imitation sample; and training the entity recognition model based on each first imitation writing sample. Therefore, a large number of labeled samples can be quickly generated under the condition that the samples are few, and the training precision of the model is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and particularly to a method and system for training a named entity recognition model for a material supply chain. Background Art

[0002] Currently, in the field of power material supply chain, the task of Chinese named entity recognition faces the dilemma of scarce high-quality labeled data, resulting in insufficient accuracy of existing supervised learning-based named entity recognition models in this field and being unable to meet the demand for precise information extraction in power material supply chain management.

[0003] In addition, although large language models have powerful language knowledge, there is a "hallucination problem" in the named entity recognition task and it is difficult to be directly applied to the information extraction task in this field. Among them, the hallucination problem refers to the phenomenon that large language models generate content that seems reasonable but is actually incorrect or fictional.

[0004] Therefore, how to generate high-quality labeled data in the field of power material supply chain to improve the accuracy of the named entity recognition model is a technical problem to be solved in this field. Summary of the Invention

[0005] The present invention provides a method and system for training a named entity recognition model for a material supply chain, which can solve at least one of the above technical problems.

[0006] According to one aspect of the present invention, there is provided a method for training a named entity recognition model for a material supply chain, including: Obtaining training samples of an entity recognition model for power material supply chain texts, wherein the training samples include sentence samples and entity types of each entity word sample in the sentence samples; Performing sentence imitation on the sentence samples according to the imitation requirement that each entity word in the imitated sentence has the same entity type as the corresponding entity word sample in the sentence sample, to obtain a plurality of target imitated sentences; Based on each of the target imitated sentences and the entity types of each entity word corresponding to each entity word sample in each of the target imitated sentences, respectively determining each first imitation sample; Training the entity recognition model based on each of the first imitation samples.

[0007] In one implementation manner, the performing sentence imitation on the sentence samples according to the imitation requirement that each entity word in the imitated sentence has the same entity type as the corresponding entity word sample in the sentence sample, to obtain at least one target imitated sentence, includes: Based on the requirement of sentence imitation that each entity word in the imitated sentence is the same as the corresponding entity word sample in the sentence sample, perform sentence imitation on the sentence sample to obtain at least one first imitated sentence; Based on the requirement of entity word imitation that the entity type of the imitated entity word is the same as the corresponding entity word sample in the sentence sample, and the entity types of each of the entity word samples, perform entity word imitation on each of the entity word samples to obtain an imitated entity word set for each of the entity word samples; Use the entity words in the imitated entity word sets of each of the entity word samples to replace the entity words in each of the first imitated sentences that are the same as the entity word samples, to obtain multiple target imitated sentences.

[0008] In one implementation, the performing sentence imitation on the sentence sample based on the requirement of sentence imitation that each entity word in the imitated sentence is the same as the corresponding entity word sample in the sentence sample to obtain at least one first imitated sentence includes: Classify the entity triples corresponding to each entity word sample in the sentence sample according to the entity type to obtain entity subsets corresponding to each of the entity types, where the entity triple includes the position of the first character of the corresponding entity word sample in the sentence sample and the position of the last character of the corresponding entity word sample in the sentence sample, as well as the entity type of the entity word sample; Based on the entity subsets corresponding to each of the entity types, construct multiple mapping relationships from the entity type to the entity subset, where each mapping relationship corresponds to one entity type; Based on the requirement of sentence imitation that each entity word in the imitated sentence is the same as the corresponding entity word sample in the sentence sample, and using the mapping relationships corresponding to each of the entity types as the entity word prompt words for each of the entity types, generate a sentence imitation instruction; Input the sentence imitation instruction into the first large language model to obtain multiple first imitated sentences output by the first large language model.

[0009] In one implementation, the sentence imitation instruction further includes a first imitation task, and the first imitation task is used to indicate the identity role of the first large language model, and is used to indicate that the first large language model imitates and generates N imitated sentences according to the imitation requirement and the identity role, with reference to the sentence sample, where N is a positive integer greater than or equal to 1.

[0010] In one implementation, the performing entity word imitation on each of the entity word samples based on the requirement of entity word imitation that the entity type of the imitated entity word is the same as the corresponding entity word sample in the sentence sample, and the entity types of each of the entity word samples to obtain an imitated entity word set for each of the entity word samples includes: Construct the imitation prompt words of the entity word sample in such a way that the entity type of the entity word sample in the sentence sample corresponds to the mapping relationship. Generate an entity word imitation instruction based on the imitation requirement that the entity type of the imitated entity word is the same as that of the corresponding entity word sample in the sentence sample, and the imitation prompt words of the entity word sample. Input the entity word imitation instruction into the second large language model to obtain the set of imitated entity words of the entity word sample output by the second large language model.

[0011] In one implementation, the entity word imitation instruction further includes a second imitation task, which is used to instruct the second large language model to imitate and generate M imitated entity words according to the imitation requirement with reference to the entity word sample, where M is a positive integer greater than or equal to 1.

[0012] In one implementation, the step of obtaining multiple target imitation sentences by replacing the entity words in each of the first imitation sentences that are the same as the entity word sample with the entity words in the set of imitated entity words of each entity word sample includes: In each of the first imitation sentences, remove the imitation sentences whose entity words are not included in the sentence sample to obtain a first set of imitation sentences. In the first set of imitation sentences, remove the imitation sentences whose similarity to the sentence sample does not meet the preset similarity condition to obtain a second set of imitation sentences. Use each of the imitated entity words in the set of imitated entity words corresponding to each entity word sample to replace the entity words in each of the imitation sentences in the second set of imitation sentences that are the same as the entity word sample to obtain a third set of imitation sentences. Based on the union of the second set of imitation sentences and the third set of imitation sentences, determine the multiple target imitation sentences.

[0013] In one implementation, the step of respectively determining each first imitation sample based on each of the target imitation sentences and the entity type of each entity word in each of the target imitation sentences corresponding to the entity word sample includes: Continue writing the target imitation sentence based on the requirement of semantic coherence to obtain the continued writing sentence of the target imitation sentence. Based on the entity type of each entity word in the target imitation sentence corresponding to the entity word sample, determine the entity type of each entity word in the target imitation sentence. Based on the target imitation sentence, the continued writing sentence of the target imitation sentence, and the entity type of each entity word in the target imitation sentence, determine the first imitation sample.

[0014] In one implementation, training the entity recognition model based on each of the first imitation samples includes: Concatenate the target imitation sentence and the continued sentence of the target imitation sentence in the first imitation sample to obtain a concatenated sentence; Input the concatenated sentence into the encoder in the entity recognition model to obtain the encoded vector of the concatenated sentence output by the encoder; Based on the position information of the target imitation sentence in the concatenated sentence, extract the encoded vector of the target imitation sentence from the encoded vector of the concatenated sentence; Input the encoded vector of the target imitation sentence into the decoder in the entity recognition model to obtain the predicted entity types of each entity word in the target imitation sentence output by the decoder; Determine a loss function based on the entity types of each entity word in the target imitation sentence included in the first imitation sample and the predicted entity types of each entity word in the target imitation sentence; Adjust the model parameters of the entity recognition model based on the loss function.

[0015] According to another aspect of the present invention, there is provided a training device for a material supply chain named entity recognition model, including: A training sample acquisition module, configured to acquire training samples of an entity recognition model for an electric power material supply chain, where the training samples include sentence samples and the entity types of each entity word sample in the sentence samples; A sentence imitation module, configured to perform sentence imitation on the sentence samples based on the imitation requirement that each entity word in the imitation sentence has the same entity type as the corresponding entity word sample in the sentence samples, to obtain at least one target imitation sentence; An imitation sample determination module, configured to respectively determine each first imitation sample based on each of the target imitation sentences and the entity types of each entity word corresponding to each entity word in each of the target imitation sentences; A model training module, configured to train the entity recognition model based on each of the first imitation samples.

[0016] In one implementation, the sentence imitation module includes: A sentence imitation unit, configured to perform sentence imitation on the sentence samples based on the imitation requirement that each entity word in the imitation sentence is the same as the corresponding entity word sample in the sentence samples, to obtain at least one first imitation sentence; An entity word imitation unit for performing entity word imitation on each of the entity word samples based on the imitation requirement that the imitation entity words are of the same entity type as the corresponding entity word samples in the sentence sample, and the entity types of each of the entity word samples, to obtain the imitation entity word sets of each of the entity word samples; An entity word replacement unit for using the entity words in the imitation entity word sets of each of the entity word samples to replace the entity words in each of the first imitation sentences that are the same as the entity word samples, to obtain a plurality of target imitation sentences.

[0017] In one implementation manner, the sentence imitation unit is specifically configured to: Classify the entity triples corresponding to each entity word sample in the sentence sample according to the entity type, to obtain entity subsets corresponding to each of the entity types, where the entity triple includes the position of the first character of the corresponding entity word sample in the sentence sample and the position of the last character of the corresponding entity word sample in the sentence sample, and the entity type of the entity word sample; Based on the entity subsets corresponding to each of the entity types, construct a plurality of mapping relationships from the entity type to the entity subset, where each mapping relationship corresponds to an entity type; Based on the imitation requirement that each entity word in the imitation sentence is the same as the corresponding entity word sample in the sentence sample, and using the mapping relationships corresponding to each of the entity types as the entity word prompt words for each of the entity types, generate a sentence imitation instruction; Input the sentence imitation instruction into a first large language model to obtain a plurality of first imitation sentences output by the first large language model.

[0018] In one implementation manner, the sentence imitation instruction further includes a first imitation task, and the first imitation task is used to indicate the identity role of the first large language model, and is used to indicate that the first large language model, according to the imitation requirement and the identity role, refers to the sentence sample and imitates and generates N imitation sentences, where N is a positive integer greater than or equal to 1.

[0019] In one implementation manner, the entity word imitation unit includes: Construct the imitation prompt word of the entity word sample in such a way that in the sentence sample, the entity type of the entity word sample is the corresponding mapping relationship; Based on the imitation requirement that the imitation entity words are of the same entity type as the corresponding entity word samples in the sentence sample, and the imitation prompt word of the entity word sample, generate an entity word imitation instruction; Input the entity word imitation instruction into a second large language model to obtain the imitation entity word set of the entity word sample output by the second large language model.

[0020] In one implementation, the entity word imitation instruction further includes a second imitation task, which is used to instruct the second large language model to imitate and generate M imitation entity words according to the imitation requirements with reference to the entity word samples, where M is a positive integer greater than or equal to 1.

[0021] In one implementation, the entity word replacement unit is specifically configured to: In the multiple first imitation sentences, remove the imitation sentences whose entity words are not included in the sentence samples to obtain a first set of imitation sentences; In the first set of imitation sentences, remove the imitation sentences whose similarity with the sentence samples does not meet the preset similarity condition to obtain a second set of imitation sentences; Use each imitation entity word in the imitation entity word set corresponding to each entity word sample to replace the entity word identical to the entity word sample in each imitation sentence in the second set of imitation sentences to obtain a third set of imitation sentences; Based on the union of the second set of imitation sentences and the third set of imitation sentences, determine the multiple target imitation sentences.

[0022] In one implementation manner, the imitation sample determination module includes: A sentence continuation unit, which is used to continue writing the target imitation sentence based on the requirement of semantic coherence to obtain a continued sentence of the target imitation sentence; An entity type determination unit, which is used to determine the entity type of each entity word in the target imitation sentence based on the entity type of the entity word sample corresponding to each entity word in the target imitation sentence; An imitation sample construction unit, which is used to determine the first imitation sample based on the target imitation sentence, the continued sentence of the target imitation sentence, and the entity type of each entity word in the target imitation sentence.

[0023] In one implementation, the model training module includes: A splicing unit, which is used to splice the target imitation sentence and the continued sentence of the target imitation sentence in the first imitation sample to obtain a spliced sentence; An encoding unit, which is used to input the spliced sentence into the encoder in the entity recognition model to obtain the encoded vector of the spliced sentence output by the encoder; An encoded vector extraction unit, which is used to extract the encoded vector of the target imitation sentence from the encoded vector of the spliced sentence based on the position information of the target imitation sentence in the spliced sentence; A decoding unit, configured to input the encoded vector of the target imitated sentence into a decoder in the entity recognition model, and obtain predicted entity types of each entity word in the target imitated sentence output by the decoder; A loss function determination unit, configured to determine a loss function based on entity types of each entity word in the target imitated sentence included in the first imitated sample and predicted entity types of each entity word in the target imitated sentence; A model parameter adjustment unit, configured to adjust model parameters of the entity recognition model based on the loss function.

[0024] According to another aspect of the present invention, there is provided a material supply chain named entity recognition model training system, including: at least one processor, and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the processor obtains the instructions from the memory and executes the instructions, so that the at least one processor can execute the material supply chain named entity recognition model training method according to any embodiment of the present invention.

[0025] According to another aspect of the present invention, there is provided a non-transitory computer-readable storage medium storing computer instructions, where the computer instructions are used to provide to a computer to instruct the computer to execute the material supply chain named entity recognition model training method according to any embodiment of the present invention.

[0026] By adopting the technical solution of the present invention, training samples of an entity recognition model for power material supply chain texts are obtained, where the training samples include sentence samples and entity types of each entity word sample in the sentence samples; based on the imitation requirement that entity types of each entity word in the imitated sentence are the same as those of the corresponding entity word samples in the sentence samples, sentence imitation is performed on the sentence samples to obtain a plurality of target imitated sentences. In this way, a plurality of imitated sentences with entity types of entity words being the same as those of entity words in the sentence samples can be obtained. Based on each of the target imitated sentences and entity types of each entity word sample corresponding to each entity word in each of the target imitated sentences, respective first imitated samples are determined. In this way, since the entity types are known, it is not necessary to manually or use a model to identify the types of entity words in the imitated sentences, and the types of entity words in the imitated sentences can be labeled. Thus, a large number of imitated samples can be quickly generated and labeled. Subsequently, based on these large number of imitated samples, the entity recognition model is trained, which can improve the recognition accuracy of the entity recognition model.

[0027] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present invention, nor is it used to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] The drawings are used to better understand the solution and do not constitute a limitation to the present invention. Among them: Figure 1 is a flowchart of a method for training a named entity recognition model of a material supply chain according to an embodiment of the present invention; Figure 2 is a structural block diagram of a device for training a named entity recognition model of a material supply chain according to an embodiment of the present invention; Figure 3 is a block diagram of an electronic device for implementing the method of the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0029] The following describes exemplary embodiments of the present invention with reference to the drawings. Various details of the embodiments of the present invention are included to facilitate understanding and should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope of the present invention. Similarly, descriptions of well-known functions and structures are omitted below for clarity and conciseness.

[0030] Figure 1 is a flowchart of a method for training a named entity recognition model of a material supply chain according to an embodiment of the present invention.

[0031] As Figure 1 shown, the method for training a named entity recognition model of a material supply chain may include: S110, obtaining training samples of a named entity recognition model for an electric power material supply chain, where the training samples include sentence samples and entity types of each entity word sample in the sentence samples; S120, performing sentence imitation on the sentence samples according to the imitation requirement that each entity word in the imitated sentence has the same entity type as the corresponding entity word sample in the sentence samples, to obtain at least one target imitated sentence; S130, respectively determining each first imitation sample based on each target imitated sentence and the entity types of each entity word corresponding to each entity word sample in each target imitated sentence; S140, training the named entity recognition model based on each first imitation sample.

[0032] In an embodiment of the present invention, when the amount of text samples collected in the power material supply chain is small, that is, in the case of small samples, at this time, the existing training samples can be subjected to the sample construction process of steps S110 to S140 of the embodiment of the present invention to obtain a plurality of first imitated samples. The first imitated samples are used as training samples and added to the existing training sample set. In this way, a large number of training samples can be obtained. By using a large number of training samples to train the entity recognition model, the recognition accuracy of the entity type of the entity words in the text sentences in the power material supply chain can be improved.

[0033] Exemplarily, a plurality of texts related to the power material supply chain can be collected from multiple platforms. Multiple training samples can be obtained by filtering, clause splitting, entity word annotation, etc. of the texts.

[0034] Exemplarily, a sentence sample can be a single sentence or a paragraph including multiple sentences. The sentence sample includes a plurality of entity word samples. The entity word sample can include one or more words.

[0035] For example, for the training samples in the original sample set , new training samples are constructed based on this to address the problem of shortage of training samples in the small sample scenario. Among them, represents the sentence sample, represents the entity type of each entity word sample. It can also include a plurality of entity triples. Each entity triple corresponds to an entity word sample. The entity triple includes the position of the first character of the corresponding entity word sample in the sentence sample and the position of the last character in the sentence sample, as well as the entity type of the entity word sample.

[0036] For example, , represents the entity triple of the entity word sample , represents the entity triple of the entity sample , represents the set of entity types for the training sample , is the entity type of the entity word sample , is the entity type of the th entity sample . Among them, , represents the first character in the entity word sample in the sentence sample position, represents the last character in the entity word sample at the position of the sentence sample . , represents the first character in the entity word sample at the position of the sentence sample . represents the last character in the entity word sample at the position of the sentence sample .

[0037] Exemplarily, based on the requirement of the imitation sentence that each entity word has the same entity type as the corresponding entity word sample in the sentence sample, an imitation instruction is determined, and the imitation instruction is input into the large language model to obtain at least one target imitation sentence output by the large language model.

[0038] Among them, the at least one target imitation sentence may include one or more target imitation sentences.

[0039] Among them, the large language model may be GPT (Generative Pre-trained Transformer), or other natural language models, for example, the Deep Seek model.

[0040] For example, the imitation instruction is:[[]] {Imitation requirement: Each entity word in the imitation sentence has the same entity type as the corresponding entity word sample in the sentence sample; Sentence sample: }.

[0041] Another example, the imitation instruction is:[[]] {Imitation requirement: Each entity word in the imitation sentence has the same entity type as the corresponding entity word sample in the sentence sample; Sentence sample: ; Entity word prompt: }.

[0042] Exemplarily, the requirement of the imitation sentence that each entity word has the same entity type as the corresponding entity word sample in the sentence sample can be divided into multiple sub-imitation requirements and some operations to implement the imitation requirement.

[0043] Exemplarily, the first imitation sample may include the target imitation sentence and the entity type of each entity word in the target imitation sentence. For the entity type of each entity word in the target imitation sentence, triple construction can be carried out similarly to the above .

[0044] Exemplarily, input the target imitation sentence in the first imitation sample into the entity recognition model to obtain the predicted entity types of each entity word in the target imitation sentence output by the entity recognition model. Based on the difference between the entity types of each entity word in the target imitation sentence in the first imitation sample and the predicted entity types of each entity word in the target imitation sentence output by the entity recognition model, determine the loss function, and use this loss function to adjust the model parameters of the entity recognition model. In this way, continuously train the model until the model accuracy reaches the preset accuracy requirement or the model training times reach the preset training times threshold before stopping the training of the model.

[0045] According to the above embodiment, obtain the training samples of the entity recognition model for the power material supply chain text, where the training samples include sentence samples and the entity types of each entity word sample in the sentence sample; based on the imitation requirement that the entity types of each entity word in the imitation sentence are the same as those of the corresponding entity word samples in the sentence sample, perform sentence imitation on the sentence sample to obtain multiple target imitation sentences. In this way, multiple imitation sentences with the same entity types as the entity words in the sentence sample can be imitated. Based on each of the target imitation sentences and the entity types of each entity word sample corresponding to each entity word in each of the target imitation sentences, respectively determine each first imitation sample. In this way, since the entity types are known, it is not necessary to manually or use a model to identify the types of entity words in the imitation sentence, and the types of entity words in the imitation sentence can be labeled. Thus, a large number of imitation samples can be quickly generated and labeled. Subsequently, based on these large numbers of imitation samples, train the entity recognition model to improve the recognition accuracy of the entity recognition model.

[0046] In one embodiment, based on the imitation requirement that the entity types of each entity word in the imitation sentence are the same as those of the corresponding entity word samples in the sentence sample, performing sentence imitation on the sentence sample to obtain at least one target imitation sentence includes: based on the imitation requirement that each entity word in the imitation sentence is the same as the corresponding entity word sample in the sentence sample, performing sentence imitation on the sentence sample to obtain at least one first imitation sentence; based on the imitation requirement that the entity types of the imitation entity words are the same as those of the corresponding entity word samples in the sentence sample and the entity types of each entity word sample, performing entity word imitation on each entity word sample to obtain the imitation entity word sets of each entity word sample; using the entity words in the imitation entity word sets of each entity word sample to replace the entity words in each first imitation sentence that are the same as the entity word samples to obtain multiple target imitation sentences.

[0047] Understandably, the requirement that each entity word in the imitated sentence has the same entity type as the corresponding entity word sample in the sentence sample is divided into two sub-imitating requirements and a replacement operation. That is, the two sub-imitating requirements are: the requirement that each entity word in the imitated sentence is the same as the corresponding entity word sample in the sentence sample, and the requirement that the imitated entity word has the same entity type as the corresponding entity word sample in the sentence sample. The replacement operation is: using the entity words in the imitated entity word set of each entity word sample to replace the entity words in each first imitated sentence that are the same as the entity word sample.

[0048] Exemplarily, based on the requirement that each entity word in the imitated sentence is the same as the corresponding entity word sample in the sentence sample, and the requirement that the imitated sentence has the same semantics as the sentence sample, a sentence imitation instruction is determined. The sentence imitation instruction is input into the large language model, and the large language model performs sentence imitation on the sentence sample according to the sentence imitation instruction, and the large language model outputs at least one first imitated sentence.

[0049] For example, the sentence imitation instruction is: {Imitation requirement: Each entity word in the imitated sentence is the same as the corresponding entity word sample in the sentence sample, and the imitated sentence has the same semantics as the sentence sample; Sentence sample: ; Entity word prompt: }.

[0050] Exemplarily, based on the requirement that the imitated entity word has the same entity type as the corresponding entity word sample in the sentence sample, an entity word imitation instruction is determined. The entity word imitation instruction is input into the large language model, and the large language model performs entity word imitation on the sentence sample according to the entity word imitation instruction, and the large language model outputs multiple imitated entity words of the entity word sample, that is, the imitated entity word set.

[0051] For example, the entity word imitation instruction is: {Imitation requirement: The imitated entity word has the same entity type as the corresponding entity word sample in the sentence sample; Sentence sample: ; Entity word sample Imitation prompt: }.

[0052] Exemplarily, using the entity words in the set of imitated entity words of each entity word sample, directly replace the entity words that are the same as the entity word sample in each first imitated sentence to obtain a plurality of target imitated sentences, and each target imitated sentence is different from each other. Or, first filter the sentences that do not meet the requirements among the plurality of first imitated sentences, and then replace the entity words that are the same as the entity word sample in each filtered first imitated sentence to obtain a plurality of target imitated sentences. Or, first filter the sentences that do not meet the requirements among the plurality of first imitated sentences, and then replace the entity words that are the same as the entity word sample in each filtered first imitated sentence to obtain a plurality of second imitated sentences, and take the union of each filtered first imitated sentence and the plurality of second imitated sentences to obtain a plurality of target imitated sentences.

[0053] According to the above embodiments, first, based on the imitation requirement of unchanged entity words, perform sentence imitation on the sentence sample, and at the same time, based on the imitation requirement of unchanged entity word types, perform entity word imitation on each entity word sample to obtain the imitated entity words of each entity word sample. Then, based on each imitated entity word of the entity word sample, replace the entity word that is the same as the entity word sample in each first imitated sentence obtained by imitation. In this way, a large number of imitated sentences can be obtained, and the imitated sentences have the same semantics as the sentence sample and the entity word types remain unchanged. Subsequently, when constructing the imitation sample, it is not necessary to perform manual annotation or model recognition of the entity types of the entity words in the imitated sentences, and the entity types of each entity word in the imitated sentences in the imitation sample can be determined, improving the sample construction efficiency.

[0054] In one embodiment, based on the imitation requirement that each entity word in the imitated sentence is the same as the corresponding entity word sample in the sentence sample, perform sentence imitation on the sentence sample to obtain at least one first imitated sentence, including: classifying the entity triples corresponding to each entity word sample in the sentence sample according to the entity type to obtain entity subsets corresponding to each entity type, where the entity triple includes the position of the first character of the corresponding entity word sample in the sentence sample and the position of the last character in the sentence sample, as well as the entity type of the entity word sample; based on the entity subsets corresponding to each entity type, construct a plurality of mapping relationships from the entity type to the entity subset, where each mapping relationship corresponds to an entity type; based on the imitation requirement that each entity word in the imitated sentence is the same as the corresponding entity word sample in the sentence sample, and using the mapping relationships corresponding to each entity type as the entity word prompt words for each entity type, generate a sentence imitation instruction; input the sentence imitation instruction into the first large language model to obtain a plurality of first imitated sentences output by the first large language model.

[0055] It can be understood that one entity type corresponds to one entity subset. For example, for the entity type , the corresponding entity subset can be denoted as: , where, For entity type of entity word samples, is the entity word sample of the entity triple, indicating the position of the first character of the entity word sample in the sentence sample . The position of the last character of the entity word sample in the sentence sample .

[0056] It should be noted that the entity word sample can include multiple different entity words ( when the values are different, the corresponding entity words are different), but the entity types of these entity words are all .

[0057] Exemplarily, for the entity type , its corresponding mapping relationship is: .

[0058] Exemplarily, the sentence imitation instruction can be: {Imitation requirement: The imitated sentence has the same semantics as the sentence sample, and each entity word in the imitated sentence is the same as the corresponding entity word sample in the sentence sample; Sentence sample: ; Entity word prompt: }.

[0059] Among them, represents the mapping relationship of the entity type as the entity subset , , represents the mapping relationship of the entity type as the entity subset .

[0060] In one implementation, the sentence imitation instruction further includes a first imitation task, and the first imitation task is used to indicate the identity role of the first large language model, and is used to indicate that the first large language model generates N imitation sentences according to the imitation requirements and the identity role, referring to the sentence sample, where N is a positive integer greater than or equal to 1.

[0061] Exemplarily, the sentence imitation instruction can be: {Imitation task: Your identity role is a Chinese teacher, and your task is to imitate sentences referring to the sentence sample and generate 10 imitation sentences; ​Imitation requirements: The imitated sentence has the same semantics as the sentence sample, and each entity word in the imitated sentence is the same as the corresponding entity word sample in the sentence sample; Sentence sample: ; Entity word prompt: }}.

[0062] Exemplarily, by using a pre-set template and filling in the specific imitation task, imitation requirements, sentence sample, and prompt into the template, the corresponding sentence imitation instruction can be obtained.

[0063] Exemplarily, inputting the sentence imitation instruction into the first large language model, the first large language model generates multiple first imitated sentences according to the sentence imitation instruction, and the first large language model outputs multiple first imitated sentences.

[0064] According to the above implementation manner, based on the imitation requirement that each entity word in the imitated sentence is the same as the corresponding entity word sample in the sentence sample, sentence imitation is performed on the sentence sample, and multiple imitated sentences can be generated. In this way, the entity types of the entity words in these imitated sentences can be determined according to the corresponding entity word samples in the sentence sample, without manual annotation, improving the construction efficiency of training samples.

[0065] In one implementation manner, based on the imitation requirement that the entity types of the imitated entity words are the same as the corresponding entity word samples in the sentence sample, and the entity types of each entity word sample, entity word imitation is performed on each entity word sample to obtain an imitated entity word set for each entity word sample, including: constructing an imitation prompt for the entity word sample in the manner of a corresponding mapping relationship of the entity type of the entity word sample in the sentence sample; generating an entity word imitation instruction based on the imitation requirement that the entity types of the imitated entity words are the same as the corresponding entity word samples in the sentence sample and the imitation prompt of the entity word sample; inputting the entity word imitation instruction into the second large language model to obtain the imitated entity word set of the entity word sample output by the second large language model.

[0066] Exemplarily, the imitation prompt can be: In the sentence sample where the entity word sample is an entity with an entity type of , where . Or, the entity prompt can be: In the sentence sample where the entity type of the entity word sample is , where .

[0067] Exemplarily, the entity word imitation instruction can be: {Imitation requirement: The imitated entity words have the same entity type as the corresponding entity word samples in the sentence samples; Imitation prompt: In the sentence sample among them, the entity word sample is an entity with an entity type of }.

[0068] In one implementation, the entity word imitation instruction further includes a second imitation task, and the second imitation task is used to instruct the second large language model to generate M imitated entity words according to the imitation requirements with reference to the entity word samples, where M is a positive integer greater than or equal to 1.

[0069] Exemplarily, the entity word imitation instruction can be: {Imitation requirement: The imitated entity words have the same entity type as the corresponding entity word samples in the sentence samples; Imitation prompt: In the sentence sample among them, the entity word sample is an entity with an entity type of , where ; Imitation task: For the entity word sample , please generate M imitated entity words

[0070] Exemplarily, the imitation requirement can be incorporated into the imitation task. For example, the entity word imitation instruction can be: {Imitation prompt: In the sentence sample among them, the entity word sample is an entity with an entity type of , where ; Imitation task: For the entity word sample , please generate M imitated entity words with the same entity type as the entity word sample }.

[0071] Exemplarily, for a pre-set template, by filling in the imitation tasks, imitation requirements, sentence samples, and prompts of each entity word sample into the template, the entity word imitation instructions for each entity word sample can be obtained.

[0072] Exemplarily, inputting the entity word imitation instruction of the entity word sample into the second large language model, the second large language model generates multiple imitated entity words of the entity word sample according to the entity word imitation instruction, and the second large language model outputs the set of imitated entity words of the entity word sample.

[0073] Exemplarily, the first large language model can be the same as or different from the second large language model.

[0074] According to the above embodiments, based on the requirement that the entity types of the imitated entity words are the same as those of the corresponding entity word samples in the sentence samples, and the entity types of each entity word sample, entity word imitation is performed on each entity word sample to obtain an imitated entity word set for each entity word sample. In this way, the entity types of each entity word in the imitated entity word set of the entity word sample are the same as those of the entity word sample. Subsequently, for the first imitated sentence obtained above, the entity words in the imitated entity word set can be used to replace the entity word samples of the same type, generating multiple second imitated sentences with the same semantics and unchanged entity types.

[0075] In one embodiment, the entity words in the imitated entity word set of each entity word sample are used to replace the entity words in each first imitated sentence that are the same as the entity word sample, obtaining multiple target imitated sentences, including: among the multiple first imitated sentences, removing the imitated sentences whose entity words are not included in the sentence sample to obtain a first imitated sentence set; in the first imitated sentence set, removing the imitated sentences whose similarity to the sentence sample does not meet the preset similarity condition to obtain a second imitated sentence set; using each imitated entity word in the imitated entity word set corresponding to each entity word sample to replace the entity words in each imitated sentence in the second imitated sentence set that are the same as the entity word sample to obtain a third imitated sentence set; based on the union of the second imitated sentence set and the third imitated sentence set, determining multiple target imitated sentences.

[0076] In this example, for the training sample , for the sentence sample , multiple first imitated sentences are obtained by imitation , which can be represented by an imitated sentence set: . First, filter out the imitated sentences that do not meet the requirements in this set. Among them, the filtering process can include rule filtering, semantic similarity filtering, and semantic coherence similarity filtering. Then, using the imitated entity word set of the entity word sample in the sentence sample , entity word replacement is performed on each imitated sentence in the filtered imitated sentence set to obtain multiple second imitated sentences, and the multiple second imitated sentences are merged into the filtered imitated sentence set for union to obtain a set including multiple target imitated sentences.

[0077] Exemplarily, the following process of any one way or multiple ways can be simultaneously performed on the multiple imitated sentences obtained above for filtering.

[0078] Exemplarily, the rule filtering for the imitated sentence can be: for the first imitated sentence , if there is an entity word that is not included in the sentence sample , then filter it out , the set of imitated sentences after rule filtering is obtained . Then there are: , , .

[0079] Exemplarily, the process of semantic similarity filtering for imitated sentences can be as follows: If the imitated sentence has a greater change compared to the original sentence , the greater the diversity of the newly constructed sample. Therefore, in order to measure the degree of change of the imitated sentence, it is necessary to calculate the semantic similarity between the imitated sentence and the sample sentence.

[0080] For a sentence sample of length and the first imitated sentence of length , the encoding vectors of the sentence sample and the encoding vectors of the first imitated sentence are obtained by inputting these two into the pre-trained language model BERT respectively. Then, by performing average pooling on these two encoding vectors, the vector representations of the sentence sample and the first imitated sentence are obtained along with and .

[0081] Exemplarily, for the sentence sample , the calculation process of its obtained vector representation can be as the following formula: ; ; where represents the encoding vector of the character in the sentence sample , represents the encoding vector of the character in the sentence sample , represents the function corresponding to the pre-trained language model BERT.

[0082] Performing an inner product operation on the vector representations of the sentence sample and the first imitated sentence can obtain the similarity score between the sentence sample and the first imitated sentence . In some examples, the similarity score can be normalized to the interval through the sigmoid function.

[0083] Therefore, the above-mentioned multiple first imitated sentences can be filtered based on the above similarity scores. For example, the imitated sentences with similarity scores less than a preset threshold can be removed.

[0084] Exemplarily, the process of filtering the semantic coherence similarity of the imitated sentences can be as follows: The semantic coherence similarity between sentences can be measured by the Perplexity Absolute Difference. Perplexity is an indicator used to evaluate the semantic coherence degree of a sentence. The more coherent the semantics of a sentence is, the smaller the perplexity of the sentence is.

[0085] ; = ; Among them, represents the probability of the sentence sample , represents the product of the probabilities of the joint occurrence of each character in the sentence sample , that is, the probability of the sentence sample , represents the occurrence probability of the character in the sentence sample , represents the joint occurrence probability of the character and the character in the sentence sample , represents the joint occurrence probability of the characters , , ···, in the sentence sample .

[0086] Among them, represents the perplexity of the sentence sample .

[0087] Therefore, from the above two formulas, it can be seen that the greater the probability of the sentence sample , the smaller the perplexity, indicating that the semantic coherence degree of the sentence sample is higher.

[0088] For the perplexity of the first imitated sentence , it can also be calculated in a similar way to the perplexity of the sentence sample , so no further examples will be given here.

[0089] If the first imitated sentence and the sentence sample The more consistent the writing style is, the smaller the absolute difference in perplexity between the two. Imitate the sentence and the sentence sample The calculation process of the absolute difference in perplexity is as follows: ; where represents the absolute difference in perplexity between the first imitated sentence and the sentence sample , and represents a very small error value.

[0090] For the set of imitated sentences , if , then directly retain the imitated sentence . Otherwise, calculate the absolute difference in perplexity between each first imitated sentence and the sentence sample in turn. If the absolute difference in perplexity between the first imitated sentence and the sentence sample is denoted as , then the normalized absolute difference in perplexity between the first imitated sentence and the sentence sample is as follows: .

[0091] where, if is smaller, it means that the first imitated sentence has a more similar writing style compared to the original sentence sample .

[0092] Therefore, among the above multiple first imitated sentences, remove the imitated sentences with a normalized absolute difference in perplexity greater than the preset threshold from the sentence sample to achieve the filtering of the semantic coherence similarity of the imitated sentences.

[0093] Exemplarily, the balancing process of filtering the semantic similarity and semantic coherence similarity of the imitated sentences is as follows: The similarity between the original sentence sample and the first imitated sentence is , and the normalized absolute difference in perplexity is . If is smaller, it means that the imitated sentence has more changes compared to the original sentence , which can increase the diversity of the samples; if is smaller, it means that the imitated sentence has a more similar writing style compared to the original sentence , which can avoid the introduction of more noise.

[0094] Exemplarily, the balanced filtering index of the first imitated sentence is: ; where represents the balanced filtering index of the first imitated sentence.

[0095] For the set of imitated sentences composed of the above multiple first imitated sentences calculate the value for each imitated sentence in it. If the retention number is , then select and retain the imitated sentences with the smallest value, and remove the other imitated sentences to obtain the filtered set of imitated sentences .

[0096] The following introduces the process of entity word replacement for each imitated sentence in the filtered set of imitated sentences as follows: For the training sample , the first imitated sentence of the sentence sample must contain all the entities corresponding to all entity types in . For the entity word sample in the sentence sample , its set of imitated entity words is . Randomly select a similar entity from and replace the entity word in the first imitated sentence that is the same as the entity word sample . In this way, multiple new imitated sentences can be obtained. Then, incorporate both the filtered first imitated sentences and the imitated sentences after entity word replacement into the same set to obtain a set including multiple target imitated sentences. For example, for the sentence sample "This is the material procurement department of XXX1 Company", its first imitated sentence is "Here is the material procurement department of XXX1 Company". If the imitated entity word for the entity word "XXX1 Company" is "XXX2 Company" and the imitated entity word for "material procurement department" is "material supply department", then use these two imitated entity words to replace the corresponding entity words in the first imitated sentence. The target imitated sentence obtained after replacing the entity words is "The material supply department of XX2 Company".

[0097] In actual application, using the target imitated sentence

[0098] as The entity types of the entity word samples corresponding to the respective entity words in it are used as the entity types of the respective entity words in the target imitated sentence. Combining the positions of the first and last characters of the respective entity words in the target imitated sentence and the entity types of the respective entity words, the target imitated sentence can be obtained entity set . In this way, the new imitated samples can be .

[0099] According to the above implementation manner, filter multiple first imitated sentences, use the entity words in the imitated entity word set of each entity word sample to replace the entity words in each first imitated sentence obtained after filtering that are the same as the entity word sample, and then incorporate each filtered first imitated sentence and each first imitated sentence after replacing the entity words into the same set to obtain multiple target imitated sentences. In this way, target imitated sentences with higher similarity and larger quantity to the original sentence samples can be obtained

[0100] In one implementation manner, based on each target imitated sentence and the entity types of the entity word samples corresponding to the respective entity words in each target imitated sentence, determine each first imitated sample respectively, including: based on the requirement of continued writing for semantic coherence, continue writing the target imitated sentence to obtain the continued writing sentence of the target imitated sentence; based on the entity types of the entity word samples corresponding to the respective entity words in the target imitated sentence, determine the entity types of the respective entity words in the target imitated sentence; based on the target imitated sentence, the continued writing sentence of the target imitated sentence, and the entity types of the respective entity words in the target imitated sentence, determine the first imitated sample

[0101] Exemplarily, a continued writing instruction can be determined based on the requirement of continued writing for semantic coherence and the target imitated sentence, and the continued writing instruction is input into the large language model to obtain the continued writing sentence of the target imitated sentence output by the large language model

[0102] Exemplarily, when the context content of the sentence is insufficient, it is easy to cause the problem of ambiguity in named entity recognition. For example, for the sentence "UHV equipment", due to the lack of relevant descriptions, it may be misrecognized as a similar noun in other fields. However, in this example, by using the large language model to continue writing it, the prior knowledge of the large language model can be introduced to remove the ambiguity problem of entity recognition. For example, continuing to write the sentence "UHV equipment" as "UHV equipment is a key facility in the power system", then "UHV equipment" can only be classified as a material equipment entity in the field of the power material supply chain by the named entity recognition model. In this way, applying the continued writing sentence of the original sentence to the training process of the named entity recognition model can improve the recognition accuracy of the model

[0103] Exemplarily, the continued writing instruction can be {For the sentence Continue writing. Just output the continued sentence.

[0104] According to the above embodiments, by continuing to write the imitated sentences and then applying the continued sentences to the training process of the named entity recognition model, the recognition accuracy of the named entity recognition model can be improved.

[0105] In one embodiment, the large language model involved in the above example can be ChatGPT (gpt-3.5-turbo-0613). Among them, the temperature parameter in the large language model is used to control the diversity of the text generated by the large model and is set to 0.5. The presence penalty parameter in the large language model is used to control the repetition degree of the text generated by the large model and is set to 0. For the processes of sentence imitation, entity word imitation, and sentence continuation, the maximum number of tokens for the generated text of the large language model is set to 2000, 300, and 120 respectively.

[0106] In one embodiment, training the entity recognition model based on each first imitation sample includes: concatenating the target imitated sentence and the continued sentence of the target imitated sentence in the first imitation sample to obtain a concatenated sentence; inputting the concatenated sentence into the encoder in the entity recognition model to obtain the encoded vector of the concatenated sentence output by the encoder; extracting the encoded vector of the target imitated sentence from the encoded vector of the concatenated sentence based on the position information of the target imitated sentence in the concatenated sentence; inputting the encoded vector of the target imitated sentence into the decoder in the entity recognition model to obtain the predicted entity types of each entity word in the target imitated sentence output by the decoder; determining the loss function based on the entity types of each entity word in the target imitated sentence included in the first imitation sample and the predicted entity types of each entity word in the target imitated sentence; and adjusting the model parameters of the entity recognition model based on the loss function.

[0107] During the training process, for the original training sample set , after the above sample construction, a new training sample set is obtained. Then, the semantic expansion of the imitated sentences in all the imitation samples in is performed to obtain the training set . For the new sample , the original sentence is concatenated with the continued sentence to obtain the concatenated sentence . Here,

[0108] Exemplarily, a masked language modeling type encoding model (MLM) is used as the encoder for the concatenated sentence Encode. If the length of the original sentence T is , continue to write the sentence with a length of , then 's embedding representation is . Extract from the part of the encoding vector embedding representation corresponding to the original sentence . .

[0109] Exemplarily, a sequence labeling method is used to perform the named entity recognition task. If the named entity sequence decoding layer is , for the new sample , the sequence labeling probability of the original sentence in it is . The labeling prediction probability of each character in the original sentence calculated through the decoding layer is: .

[0110] Exemplarily, the model is trained by maximum likelihood estimation, and the target loss function for training is: .

[0111] In practical applications, the learning rate of the encoding layer parameters can be set to 1e-5, the learning rate of the decoding layer parameters can be set to 1e-3, the learning algorithm is AdamW, and the weight decay rate is 0.01. Sampling linear learning rate warm-up is used, and the warmup rate is 0.1. The maximum sentence length of each dataset is unified to 350, the training batch size is 2, and the number of training epochs is 30.

[0112] According to the above embodiments, by splicing the imitated sentence and its continued sentence and inputting them into the encoder, the encoder can encode by combining the understanding of the entire sentence context, so that the obtained encoding result is more accurate. Then, the decoder is used to decode the encoding vector corresponding to the imitated sentence part in the encoding result. Since the content of the continued sentence is combined for understanding during encoding, the decoding result of this part of the encoding vector of the imitated sentence is more accurate. In this way, the recognition accuracy of the named entity recognition model obtained by training is higher.

[0113] Figure 2 is the structural block diagram of the named entity recognition model training device for the material supply chain according to an embodiment of the present invention.

[0114] As Figure 2 shown, the named entity recognition model training device for the material supply chain includes: A training sample acquisition module 210, configured to acquire training samples for an entity recognition model of an electric power material supply chain, where the training samples include sentence samples and entity types of respective entity word samples in the sentence samples; A sentence imitation module 220, configured to perform sentence imitation on the sentence samples based on an imitation requirement that entity types of respective entity words in an imitation sentence are the same as those of corresponding entity word samples in the sentence samples, so as to obtain at least one target imitation sentence; An imitation sample determination module 230, configured to respectively determine respective first imitation samples based on the respective target imitation sentences and entity types of respective entity words corresponding to the respective entity words in the respective target imitation sentences; A model training module 240, configured to train the entity recognition model based on the respective first imitation samples.

[0115] In an implementation manner, the sentence imitation module includes: A sentence imitation unit, configured to perform sentence imitation on the sentence samples based on an imitation requirement that respective entity words in an imitation sentence are the same as corresponding entity word samples in the sentence samples, so as to obtain at least one first imitation sentence; An entity word imitation unit, configured to perform entity word imitation on the respective entity word samples based on an imitation requirement that an imitation entity word is the same as the entity type of a corresponding entity word sample in the sentence samples and entity types of the respective entity word samples, so as to obtain imitation entity word sets of the respective entity word samples; An entity word replacement unit, configured to use entity words in the imitation entity word sets of the respective entity word samples to replace entity words that are the same as the entity word samples in the respective first imitation sentences, so as to obtain a plurality of target imitation sentences.

[0116] In an implementation manner, the sentence imitation unit is specifically configured to: Classify entity triples corresponding to respective entity word samples in the sentence samples according to entity types, so as to obtain entity subsets corresponding to the respective entity types, where the entity triples include positions of the first character and the last character of a corresponding entity word sample in the sentence sample and the entity type of the entity word sample; Construct a plurality of mapping relationships from entity types to entity subsets based on the entity subsets corresponding to the respective entity types, where each mapping relationship corresponds to one entity type; Generate a sentence imitation instruction based on an imitation requirement that respective entity words in an imitation sentence are the same as corresponding entity word samples in the sentence samples and using the mapping relationships corresponding to the respective entity types as entity word prompt words for the respective entity types; Input the sentence imitation instruction into the first large language model to obtain multiple first imitation sentences output by the first large language model.

[0117] In one implementation, the sentence imitation instruction further includes a first imitation task, which is used to indicate the identity role of the first large language model and to instruct the first large language model to generate N imitation sentences by referring to the sentence sample according to the imitation requirements and the identity role, where N is a positive integer greater than or equal to 1.

[0118] In one implementation, the entity word imitation unit includes: Construct an imitation prompt word for the entity word sample in such a way that the entity type of the entity word sample in the sentence sample is the corresponding mapping relationship. Based on the imitation requirement that the imitation entity word has the same entity type as the corresponding entity word sample in the sentence sample, and the imitation prompt word of the entity word sample, generate an entity word imitation instruction. Input the entity word imitation instruction into the second large language model to obtain the set of imitation entity words of the entity word sample output by the second large language model.

[0119] In one implementation, the entity word imitation instruction further includes a second imitation task, which is used to instruct the second large language model to generate M imitation entity words by referring to the entity word sample according to the imitation requirements, where M is a positive integer greater than or equal to 1.

[0120] In one implementation, the entity word replacement unit is specifically used for: In the multiple first imitation sentences, remove the imitation sentences whose entity words are not included in the sentence sample to obtain a first set of imitation sentences. In the first set of imitation sentences, remove the imitation sentences whose similarity to the sentence sample does not meet the preset similarity condition to obtain a second set of imitation sentences. Use each imitation entity word in the set of imitation entity words corresponding to each entity word sample to replace the entity word in each imitation sentence in the second set of imitation sentences that is the same as the entity word sample to obtain a third set of imitation sentences. Based on the union of the second set of imitation sentences and the third set of imitation sentences, determine the multiple target imitation sentences.

[0121] In one implementation, the imitation sample determination module includes: A sentence continuation unit, which is used to continue the target imitation sentence based on the requirement of semantic coherence to obtain the continued sentence of the target imitation sentence. An entity type determination unit, configured to determine the entity types of each entity word in the target sentence for imitation based on the entity types of the entity word samples corresponding to each entity word in the target sentence for imitation; An imitation sample construction unit, configured to determine the first imitation sample based on the target sentence for imitation, the continued sentence of the target sentence for imitation, and the entity types of each entity word in the target sentence for imitation.

[0122] In one implementation manner, the model training module includes: A splicing unit, configured to splice the target sentence for imitation and the continued sentence of the target sentence for imitation in the first imitation sample to obtain a spliced sentence; An encoding unit, configured to input the spliced sentence into an encoder in the entity recognition model to obtain an encoded vector of the spliced sentence output by the encoder; An encoded vector extraction unit, configured to extract the encoded vector of the target sentence for imitation from the encoded vector of the spliced sentence based on the position information of the target sentence for imitation in the spliced sentence; A decoding unit, configured to input the encoded vector of the target sentence for imitation into a decoder in the entity recognition model to obtain the predicted entity types of each entity word in the target sentence for imitation output by the decoder; A loss function determination unit, configured to determine a loss function based on the entity types of each entity word in the target sentence for imitation included in the first imitation sample and the predicted entity types of each entity word in the target sentence for imitation; A model parameter adjustment unit, configured to adjust the model parameters of the entity recognition model based on the loss function.

[0123] For the specific functions and examples of each module and sub-module of the system according to the embodiments of the present invention, reference may be made to the relevant descriptions of the corresponding steps in the foregoing method embodiments, which will not be elaborated herein.

[0124] In the technical solution of the present invention, the acquisition, storage, and application of user personal information involved all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.

[0125] According to the embodiments of the present invention, the present invention also provides a system and a readable storage medium.

[0126] Exemplarily, an embodiment of the present invention provides a training system for a material supply chain named entity recognition model, including: at least one processor, and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the processor obtains the instructions from the memory and executes the instructions, so that the at least one processor can execute the material supply chain named entity recognition model training method according to any embodiment of the present invention. This system can be applied to an electronic device.

[0127] Exemplarily, an embodiment of the present invention provides a non-transitory computer-readable storage medium storing computer instructions, and the computer instructions are used to provide to a computer to instruct the computer to execute the material supply chain named entity recognition model training method according to any embodiment of the present invention.

[0128] Figure 3 FIG. shows a schematic block diagram of an exemplary electronic device 800 that can be used to implement the embodiments of the present invention. The electronic device is intended to represent various forms of digital computers, such as, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, a personal digital assistant, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or claimed herein.

[0129] As Figure 3 shown, the electronic device 800 includes a computing unit 801, which can execute various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 802 or a computer program loaded from a storage unit 808 into a random access memory (RAM) 803. In the RAM 803, various programs and data required for the operation of the electronic device 800 can also be stored. The computing unit 801, the ROM 802, and the RAM 803 are connected to each other through a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.

[0130] A plurality of components in the electronic device 800 are connected to the I / O interface 805, including: an input unit 806, such as a keyboard, a mouse, etc.; an output unit 807, such as various types of displays, speakers, etc.; a storage unit 808, such as a magnetic disk, an optical disk, etc.; and a communication unit 809, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 809 allows the electronic device 800 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0131] The computing unit 801 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 801 executes the various methods and processes described above, such as the method for training the material supply chain named entity recognition model. For example, in some embodiments, the method for training the material supply chain named entity recognition model can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as the storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 800 via the ROM 802 and / or the communication unit 809. When the computer program is loaded into the RAM 803 and executed by the computing unit 801, one or more steps of the method for training the material supply chain named entity recognition model described above can be executed. Alternatively, in other embodiments, the computing unit 801 can be configured to execute the method for training the material supply chain named entity recognition model in any other suitable manner (e.g., by means of firmware).

[0132] Various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs, which can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a dedicated or general-purpose programmable processor, and can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0133] The program code for implementing the method of the present invention can be written in any combination of one or more programming languages. These program codes can be provided to the processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing devices, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The program code can be executed entirely on the machine, partially on the machine, executed partially on the machine and partially on a remote machine as an independent software package, or executed entirely on a remote machine or server.

[0134] In the context of the present invention, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. The machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0135] In order to provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, speech input, or tactile input).

[0136] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.

[0137] A computer system may include a client and a server. The client and the server are generally far from each other and usually interact via a communication network. The relationship between the client and the server is generated by computer programs running on respective computers and having a client-server relationship with each other. The server may be a cloud server, a server of a distributed system, or a server incorporating a blockchain.

[0138] It should be understood that various forms of the processes shown above can be used, steps can be reordered, added or deleted. For example, the steps recited in the present invention can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in the present invention can be achieved, and no limitations are imposed herein.

[0139] The above specific embodiments do not constitute a limitation on the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the principles of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for training a named entity recognition model for a material supply chain, characterized in that Including: Obtain training samples for the entity recognition model of the power material supply chain, where the training samples include sentence samples and the entity types of each entity word sample in the sentence samples; Based on the requirement of sentence imitation that each entity word in the imitated sentence has the same entity type as the corresponding entity word sample in the sentence sample, perform sentence imitation on the sentence sample to obtain at least one target imitated sentence; Based on each of the target imitated sentences and the entity types of the corresponding entity word samples of each entity word in each of the target imitated sentences, respectively determine each first imitation sample; Based on each of the first imitation samples, train the entity recognition model.

2. The method according to claim 1, wherein The performing sentence imitation on the sentence sample based on the requirement of sentence imitation that each entity word in the imitated sentence has the same entity type as the corresponding entity word sample in the sentence sample to obtain at least one target imitated sentence includes: Based on the requirement of sentence imitation that each entity word in the imitated sentence is the same as the corresponding entity word sample in the sentence sample, perform sentence imitation on the sentence sample to obtain at least one first imitated sentence; Based on the requirement of entity word imitation that the imitated entity word has the same entity type as the corresponding entity word sample in the sentence sample and the entity types of each of the entity word samples, perform entity word imitation on each of the entity word samples to obtain an imitation entity word set for each of the entity word samples; Use the entity words in the imitation entity word sets of each of the entity word samples to replace the entity words in each of the first imitated sentences that are the same as the entity word samples to obtain multiple target imitated sentences.

3. The method according to claim 2, wherein The performing sentence imitation on the sentence sample based on the requirement of sentence imitation that each entity word in the imitated sentence is the same as the corresponding entity word sample in the sentence sample to obtain at least one first imitated sentence includes: Classify the entity triples corresponding to each entity word sample in the sentence sample according to the entity type to obtain entity subsets corresponding to each of the entity types, where the entity triple includes the position of the first character of the corresponding entity word sample in the sentence sample, the position of the last character of the corresponding entity word sample in the sentence sample, and the entity type of the entity word sample; Based on the entity subsets corresponding to each of the entity types, construct multiple mapping relationships from the entity type to the entity subset, where each mapping relationship corresponds to an entity type; Based on the requirement of sentence imitation that each entity word in the imitated sentence is the same as the corresponding entity word sample in the sentence sample and using the mapping relationships corresponding to each of the entity types as the entity word prompt words for each of the entity types, generate a sentence imitation instruction; Input the sentence imitation instruction into the first large language model to obtain multiple first imitated sentences output by the first large language model.

4. The method according to claim 3, characterized in that The sentence imitation instruction further includes a first imitation task, and the first imitation task is used to indicate the identity role of the first large language model and is used to indicate that the first large language model generates N imitated sentences by referring to the sentence sample according to the imitation requirement and the identity role, where N is a positive integer greater than or equal to 1.

5. The method according to claim 3, wherein Per the imitation requirements that the entity types of the imitated entity words are the same as those of the corresponding entity word samples in the sentence samples, and based on the entity types of each of the entity word samples, perform entity word imitation on each of the entity word samples to obtain an imitated entity word set for each of the entity word samples, including: Construct an imitation prompt for the entity word sample in such a way that the entity type of the entity word sample in the sentence sample is the corresponding mapping relationship; Generate an entity word imitation instruction based on the imitation requirements that the entity types of the imitated entity words are the same as those of the corresponding entity word samples in the sentence samples, and the imitation prompt for the entity word sample; Input the entity word imitation instruction into the second large language model to obtain the imitated entity word set for the entity word sample output by the second large language model.

6. The method according to claim 5, characterized in that The entity word imitation instruction further includes a second imitation task, which is used to instruct the second large language model to imitate and generate M imitated entity words according to the imitation requirements with reference to the entity word sample, where M is a positive integer greater than or equal to 1.

7. The method according to claim 2, wherein Replacing the entity words identical to the entity word samples in each of the first imitation sentences with the entity words in the imitated entity word sets for each of the entity word samples to obtain a plurality of target imitation sentences, including: In each of the first imitation sentences, remove the imitation sentences whose entity words are not included in the sentence sample to obtain a first imitation sentence set; In the first imitation sentence set, remove the imitation sentences whose similarity to the sentence sample does not meet the preset similarity condition to obtain a second imitation sentence set; Replace the entity words identical to the entity word samples in each of the imitation sentences in the second imitation sentence set with each of the imitated entity words in the imitated entity word sets corresponding to each of the entity word samples to obtain a third imitation sentence set; Based on the union of the second imitation sentence set and the third imitation sentence set, determine the plurality of target imitation sentences.

8. The method according to claim 1, wherein Based on each of the target imitation sentences and the entity types of each of the entity words corresponding to each of the entity word samples in each of the target imitation sentences, respectively determine each of the first imitation samples, including: Perform continuation writing on the target imitation sentence based on the requirement of semantic coherence to obtain a continuation sentence of the target imitation sentence; Based on the entity types of each of the entity words corresponding to each of the entity word samples in the target imitation sentence, determine the entity types of each of the entity words in the target imitation sentence; Based on the target imitation sentence, the continuation sentence of the target imitation sentence, and the entity types of each of the entity words in the target imitation sentence, determine the first imitation sample.

9. The method according to claim 8, wherein Based on each of the first imitation samples, train the entity recognition model, including: Concatenate the target imitation sentence and the continuation sentence of the target imitation sentence in the first imitation sample to obtain a concatenated sentence; Input the concatenated sentence into the encoder in the entity recognition model to obtain the encoded vector of the concatenated sentence output by the encoder; Extract the encoded vector of the target sentence for imitation from the encoded vector of the spliced sentence based on the position information of the target sentence for imitation in the spliced sentence; Input the encoded vector of the target sentence for imitation into the decoder in the entity recognition model to obtain the predicted entity types of each entity word in the target sentence for imitation output by the decoder; Determine the loss function based on the entity types of each entity word in the target sentence for imitation included in the first imitation sample and the predicted entity types of each entity word in the target sentence for imitation; Adjust the model parameters of the entity recognition model based on the loss function.

10. An apparatus for training a named entity recognition model for a material supply chain, characterized in that, It includes: A training sample acquisition module for acquiring training samples of an entity recognition model for the power material supply chain, where the training samples include sentence samples and the entity types of each entity word sample in the sentence samples; A sentence imitation module for imitating the sentence samples based on the imitation requirement that the entity types of each entity word in the sentence for imitation are the same as the corresponding entity word samples in the sentence samples to obtain at least one target sentence for imitation; An imitation sample determination module for respectively determining each first imitation sample based on each target sentence for imitation and the entity types of each entity word sample corresponding to each entity word in each target sentence for imitation; A model training module for training the entity recognition model based on each first imitation sample.

11. A training system for a named entity recognition model of a material supply chain, characterized in that, It includes: At least one processor and a memory communicatively connected to the at least one processor; Wherein, the memory stores instructions executable by the at least one processor, and the processor obtains the instructions from the memory and executes the instructions so that the at least one processor can execute the method for training a named entity recognition model for a material supply chain according to any one of claims 1-9.

12. A non-transitory computer-readable storage medium storing computer instructions, characterized in that, The computer instructions are used to provide to a computer to instruct the computer to execute the method for training a named entity recognition model for a material supply chain according to any one of claims 1-9.

Citation Information

Patent Citations

  • Sentence entity completion method in combination with triplet knowledge base

    CN108563637A

  • Named entity recognition neural structure in reading understanding form

    CN114648113A

  • Prompt word expanding and writing method and device, storage medium and electronic equipment

    CN117573913A

  • Target recognition model sample collecting and training method for small sample problem

    CN119964062A

  • Weighting dictionary entities for language understanding models

    US20150262078A1