Text processing model training method, system and equipment and storage medium

By building a lightweight text processing model and using the annotation training set training method of the large language model, the problem of complex steps and high cost in the cross-domain article deconstruction scenarios of large language models is solved, and text processing efficiency and key information extraction efficiency are improved.

CN120256548APending Publication Date: 2025-07-04SF TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202311865614.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-30
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

In the cross-domain article deconstruction scenario of large language models, the existing technology has large parameters and complicated steps, which leads to the problem of recalling appropriate few-shot samples for imitation learning for each pending case, which is high in use.

Method used

By using vector encoding processing based on text samples, a set of cases to be marked is constructed, the similarity is calculated and the set number of reference cases are obtained for annotation is obtained. A large language model is used to generate an annotation training set, and a lightweight text processing model is trained, which eliminates the recall steps of similar few-shot samples.

Benefits of technology

It realizes the improvement of text processing efficiency, saves the cost of large language model interface calls, builds a lightweight text processing model, and improves the efficiency of text key information extraction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120256548A_ABST
    Figure CN120256548A_ABST
Patent Text Reader

Abstract

The invention discloses a text processing model training method, system and device and a storage medium, and the method comprises the steps: carrying out the vector coding processing based on a first number of text samples, and constructing a to-be-labeled case set; calculating the similarity between each to-be-labeled task in the to-be-labeled case set and each reference case in a preset reference case set; further, for any to-be-labeled task, obtaining a set number of reference cases of the to-be-labeled task according to a sequence of the similarity from large to small; calling a large language model to label each to-be-labeled task by using a corresponding set number of reference cases to obtain a labeled training set; and training based on the annotation training set to obtain a lightweight text processing model. According to the method, the training set is constructed through the large language model with the large parameter quantity, then the lightweight text processing model is trained, the step of similar few-sample learning sample recall is omitted, and the large language model interface calling cost is saved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of natural language processing, and particularly to a method, system, device, and storage medium for training a text processing model. Background Art

[0002] A large language model (LLM) is a natural language processing model based on deep learning technology. It can generate natural language text and requires a large amount of training data and computing resources to be trained. An LLM can learn language rules by analyzing a vast amount of text to understand the information in the text, so as to accurately judge the logical relationships and connections between sentences, paragraphs, or chapters. Based on this, an LLM can be widely applied to text processing scenarios such as article deconstruction.

[0003] For cross-domain article deconstruction scenarios, the few-shot learning ability of the large language model can usually be utilized. By placing reference cases in the prompt words, correct recognition / deconstruction / disassembly results can be obtained by imitating the reference cases.

[0004] To improve performance, some conventional large language models are usually set with a relatively long solution link and a large number of parameters. However, in the article deconstruction method using few-shot examples, it is necessary to recall appropriate few-shot examples for imitation learning for each case to be processed, resulting in complicated steps and high usage costs. Summary of the Invention

[0005] The main object of the present invention is to provide a method, system, device, and storage medium for training a text processing model, which solves the problem of low efficiency in text processing and extracting key information of text through a large language model. By using a large language model with a large number of parameters to annotate and construct a training set, a lightweight text processing model is trained, eliminating the step of recalling similar few-shot examples and saving the cost of calling the large language model interface.

[0006] To achieve the above object, the embodiments of the present application provide the following technical solutions:

[0007] According to the first aspect of the embodiments of the present application, a method for training a text processing model is provided, and the method includes:

[0008] Performing vector encoding processing on a first number of text samples to construct a set of cases to be annotated;

[0009] Calculate the similarity between each to-be-annotated task in the to-be-annotated case set and each reference case in the pre-set reference case set; the reference case set is obtained after vector encoding processing based on the pre-annotated second number of reference texts; the first number is greater than the second number; the reference case set includes at least one type of reference case;

[0010] For any to-be-annotated task, obtain a set number of reference cases for the to-be-annotated task in descending order of similarity;

[0011] Call a large language model to annotate each to-be-annotated task with the corresponding set number of reference cases to obtain an annotation training set;

[0012] Train a lightweight text processing model based on the annotation training set.

[0013] Optionally, perform vector encoding processing on the first number of text samples to construct a to-be-annotated case set, including:

[0014] Split the first number of text samples according to the first set rule to obtain a first text segment group;

[0015] Perform vector encoding processing on each first text segment in the first text segment group to obtain the text vectors of each first text segment, which are used as the to-be-annotated case set.

[0016] Optionally, calling a large language model to annotate each to-be-annotated task with the corresponding set number of reference cases includes:

[0017] For any to-be-annotated task, based on the pre-set prompt word library, construct a prompt word for the to-be-annotated task according to the set number of reference cases corresponding to the to-be-annotated task, and the prompt word library includes descriptive words representing the business types corresponding to the set number of reference cases;

[0018] Call a large language model to generate annotation information for the to-be-annotated task according to the prompt word of the to-be-annotated task;

[0019] Annotate the to-be-annotated task according to the annotation information.

[0020] Optionally, the method further includes:

[0021] Receive a to-be-processed text task, and the to-be-processed text task carries the to-be-processed text;

[0022] Split the to-be-processed text according to the third set rule to obtain a number of target segments;

[0023] Input the number of target segments into the lightweight text processing model to obtain the processed text result.

[0024] According to the second aspect of the embodiments of the present application, a text processing model training system is provided. The system includes:

[0025] A to-be-annotated case determination module, configured to perform vector encoding processing on a first number of text samples to construct a to-be-annotated case set;

[0026] A similarity calculation module, configured to calculate the similarity between each to-be-annotated task in the to-be-annotated case set and each reference case in a preset reference case set; the reference case set is obtained by performing vector encoding processing on a second number of pre-annotated reference texts; the first number is greater than the second number; the reference case set includes at least one type of reference case;

[0027] A reference case determination module, configured to, for any to-be-annotated task, obtain a set number of reference cases of the to-be-annotated task in descending order of similarity;

[0028] An annotation module, configured to call a large language model to annotate each to-be-annotated task with the corresponding set number of reference cases to obtain an annotation training set;

[0029] A training module, configured to train a lightweight text processing model based on the annotation training set.

[0030] According to the third aspect of the embodiments of the present application, an electronic device is provided, including: a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor runs the computer program, the method of the first aspect is implemented.

[0031] According to the fourth aspect of the embodiments of the present application, a computer-readable storage medium is provided, on which computer-readable instructions are stored. The computer-readable instructions can be executed by a processor to implement the method of the first aspect.

[0032] In summary, the embodiments of the present application provide a method, system, device, and storage medium for training a text processing model. First, vector encoding processing is performed on a first number of text samples to construct a set of cases to be annotated; the similarity between each task to be annotated in the set of cases to be annotated and each reference case in a preset reference case set is calculated; wherein, the reference case set is obtained after vector encoding processing on a second number of pre-annotated reference texts; the first number is greater than the second number; further, for any task to be annotated, a set number of reference cases of the task to be annotated are obtained in descending order of similarity; then, a large language model is called to annotate each task to be annotated with the corresponding set number of reference cases to obtain an annotation training set; a lightweight text processing model is trained based on the annotation training set. The problem of low efficiency in text processing and extracting key text information is solved by the large language model. By using a large language model with a large number of parameters to annotate and construct a training set, and then training a lightweight text processing model, the step of recalling similar few-shot examples is omitted, and the cost of calling the large language model interface is saved. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on the structures shown in these drawings.

[0034] The structures, ratios, sizes, etc. illustrated in this specification are only used to cooperate with the content disclosed in this specification for those who are familiar with this technology to understand and read, and are not used to limit the limited conditions under which the present invention can be implemented. Therefore, they do not have any technical essence. Any modification of the structure, change of the proportional relationship, or adjustment of the size, without affecting the effects that the present invention can produce and the purposes that can be achieved, should still fall within the scope covered by the technical content disclosed in the present invention.

[0035] Figure 1 It is a schematic flowchart of the method for training a text processing model provided by the embodiments of the present application;

[0036] Figure 2 It is a schematic overall flowchart of the training and application of the text processing model provided by the embodiments of the present application;

[0037] Figure 3 It is a block diagram of the text processing model training system provided by the embodiments of the present application;

[0038] Figure 4 It shows a schematic structural diagram of an electronic device provided by the embodiments of the present application;

[0039] Figure 5 The figure shows a schematic diagram of a computer-readable storage medium provided by an embodiment of the present application.

[0040] The realization of the purpose of the present invention, its functional characteristics and advantages will be further described in conjunction with embodiments with reference to the accompanying drawings. Specific embodiments

[0041] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0042] It should be noted that all directional indications (such as up, down, left, right, front, back...) in the embodiments of the present invention are only used to explain the relative position relationship and movement conditions between components in a specific posture (as shown in the accompanying drawings). If the specific posture changes, the directional indications will also change accordingly.

[0043] In addition, the descriptions such as "first" and "second" in the present invention are only for descriptive purposes, and cannot be understood as indicating or implying their relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one such feature. In the description of the present invention, "a plurality" means at least two, such as two, three, etc., unless otherwise specifically defined.

[0044] In the present invention, unless otherwise clearly defined and limited, the terms "connection", "fixation", etc. should be understood in a broad sense. For example, "fixation" may be a fixed connection, a detachable connection, or integrated; it may be a mechanical connection or an electrical connection; it may be directly connected or indirectly connected through an intermediate medium, and it may be the communication inside two components or the interaction relationship between two components, unless otherwise clearly limited. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific situations.

[0045] In addition, the technical solutions between various embodiments of the present invention can be combined with each other, but it must be based on the fact that those of ordinary skill in the art can achieve it. When the combination of technical solutions results in contradictions or cannot be achieved, it should be considered that such a combination of technical solutions does not exist and is not within the protection scope required by the present invention.

[0046] Figure 1 The figure shows a text processing model training method provided by an embodiment of the present application. The method includes:

[0047] Step 101: Perform vector encoding processing based on the first number of text samples to construct a set of cases to be annotated;

[0048] Step 102: Calculate the similarity between each task to be annotated in the set of cases to be annotated and each reference case in the preset reference case set; the reference case set is obtained after performing vector encoding processing on the second number of pre-annotated reference texts; the first number is greater than the second number; the reference case set includes at least one type of reference case;

[0049] Step 103: For any task to be annotated, obtain a set number of reference cases for the task to be annotated in descending order of similarity;

[0050] Step 104: Call a large language model to annotate each task to be annotated using the corresponding set number of reference cases to obtain an annotation training set;

[0051] Step 105: Train a lightweight text processing model based on the annotation training set.

[0052] In a possible implementation manner, before step 101, it further includes: splitting the second number of reference texts according to a second set rule to obtain a second text segment group; respectively performing information annotation on the second text segment group to obtain a set of annotations for each second text segment; respectively performing vector encoding processing on the second text segment group to obtain a set of text vectors for each second text segment; classifying the second text segment group, the set of annotations for each second text segment, and the set of text vectors to obtain a reference case set.

[0053] By performing annotation and vector encoding processing according to the user's own needs to form a reference case set for subsequent dynamic few-shot use of the large model. Use a small number of samples to construct learning examples for the large language model, and utilize the few-shot learning ability of the large language model to annotate a large number of bids, thereby constructing training samples.

[0054] In a possible implementation manner, in step 101, performing vector encoding processing based on the first number of text samples to construct a set of cases to be annotated includes:

[0055] Splitting the first number of text samples according to a first set rule to obtain a first text segment group; respectively performing vector encoding processing on each first text segment in the first text segment group to obtain the text vector of each first text segment as the set of cases to be annotated.

[0056] In a possible implementation, the first setting rule, the second setting rule, and the third setting rule may be the same or different, and are set according to actual applications. The setting rule includes splitting according to the semantic structure of the article. For example, the program identifies the chapter and section directory markers in the document. Since the content semantics under each chapter and section directory are consistent, it is split according to the chapter and section directory.

[0057] In a possible implementation, the method of vector encoding processing may be to use a general embedding model, including but not limited to SimCSE, Text2Vec, etc.

[0058] In a possible implementation, the above-mentioned large language model is called to label each task to be labeled with a corresponding set number of reference cases, specifically including:

[0059] For any task to be labeled, based on a preset prompt word library, a prompt word for the task to be labeled is constructed according to the set number of reference cases corresponding to the task to be labeled. The prompt word library includes descriptive words representing the business types corresponding to the set number of reference cases; the large language model is called to generate labeling information for the task to be labeled according to the prompt word of the task to be labeled; the task to be labeled is labeled according to the labeling information.

[0060] A prompt "prompt" is constructed using the recalled set number of reference cases, and the few-shot ability of the large language model is used to label the task to be labeled.

[0061] In a possible implementation, in step 102, the similarity between each task to be labeled in the set of tasks to be labeled and each reference case in the preset reference case set is calculated, including:

[0062] For each task to be labeled in the set of tasks to be labeled, a parameter for calculating the similarity between the text vector of the task to be labeled and the text vector of each reference case in the reference case set is calculated. The parameter of the similarity is the Euclidean distance or the cosine distance.

[0063] In a possible implementation, when the parameter of the similarity is the Euclidean distance, the set number of reference cases of the task to be labeled is obtained in ascending order of the Euclidean distance.

[0064] In a possible implementation, when the parameter of the similarity is the cosine distance, the set number of reference cases of the task to be labeled is obtained in descending order of the cosine distance.

[0065] In a possible implementation, the method further includes:

[0066] Receive the text task to be processed, where the text task to be processed carries the text to be processed; split the text to be processed according to the third set rule to obtain a number of target segments; input the number of target segments into the lightweight text processing model to obtain the processed text result.

[0067] In the application stage of the text processing model, use the trained lightweight text processing model to parse the text to be processed. In the above training process of the text processing model, the data of the training set is labeled through the powerful reading and comprehension ability of the large language model, and then the lightweight large language model (text processing model) is used to train the learned ability. The knowledge is solidified in the trained text processing model, saving the step of recalling similar few-shot examples and also saving the cost of calling the large language model interface during the training of the traditional model.

[0068] Figure 2 Shows the overall flow diagram of the training and application of the text processing model provided by the embodiments of the present application. Specifically, it may include the following stages:

[0069] The first stage: Construct a training set. Specifically, it includes several aspects: constructing a reference case set, constructing a case set to be labeled, and labeling the case set to be labeled based on the large language model.

[0070] The first aspect: Construct a reference case set.

[0071] Step 1: Split a small amount of reference text according to the set rule to obtain a text segment group; for example, the output structure: <semantic block-1, semantic block-2,..., semantic block-n>;

[0072] Step 2: Perform information annotation on the text segment group respectively to obtain a set of annotations for each text segment;

[0073] Step 3: Perform vector encoding processing on the text segment group respectively to obtain a set of text vectors for each text segment;

[0074] Step 4: Classify the text segment group, the set of annotations for each text segment, and the set of text vectors to obtain a reference case set. The output structure is as follows: <semantic block-1: [annotation information, vector encoding], semantic block-2: [annotation information, vector encoding],..., semantic block-n: [annotation information, vector encoding]>.

[0075] Example: Semantic block-1: The bidder must be an independent legal person or other organization registered within the territory of the People's Republic of China, with a business license, express delivery business license, and road transport license issued by the administrative department for industry and commerce, capable of independently bearing civil liability, and having the ability to engage in this project. Annotation information: License requirements: Business license, express delivery business license, road transport license.

[0076] In a possible implementation, the setting rules include splitting according to the semantic structure of the article. For example, the program identifies the chapter and section markers in the document. Since the content under each chapter and section has consistent semantics, it is split according to the chapter and section markers.

[0077] In a possible implementation, the method of vector encoding processing can be to use a general embedding model including but not limited to simcse, text2vec, etc.

[0078] By annotating according to the user's own needs and performing vector encoding processing, a reference case set is formed for subsequent dynamic few-shot use by the large model. Use a small number of samples to construct the learning examples of the large language model, and utilize the few-shot learning ability of the large language model to annotate a large number of tender documents, thereby constructing training samples.

[0079] Second aspect: Construct a set of cases to be annotated.

[0080] Step 1: Split a large number of text samples according to the set rules to obtain a group of text segments;

[0081] Step 2: Perform vector encoding processing on each text segment in the group of text segments respectively to obtain the text vectors of each text segment, which are used as the set of cases to be annotated. The output structure is as follows: <semantic block-1: [vector encoding], semantic block-2: [vector encoding],...>. Semantic block = text.

[0082] In a possible implementation, the setting rules include splitting according to the semantic structure of the article. For example, the program identifies the chapter and section markers in the document. Since the content under each chapter and section has consistent semantics, it is split according to the chapter and section markers.

[0083] In a possible implementation, the method of vector encoding processing can be to use a general embedding model including but not limited to simcse, text2vec, etc.

[0084] Third aspect: Annotate the set of cases to be annotated based on the large language model.

[0085] Step 1: For each annotation task in the set of cases to be annotated, calculate the parameter of the similarity between the text vector of the annotation task and the text vectors of each reference case in the reference case set. The parameter of the similarity is the Euclidean distance or the cosine distance.

[0086] Step 2: Obtain the set number of reference cases of the annotation task in the order of similarity from large to small, such as the topk reference cases.

[0087] In a possible implementation, when the parameter of similarity is the Euclidean distance, a set number of reference cases for the task to be labeled are obtained in ascending order of the Euclidean distance.

[0088] In a possible implementation, when the parameter of similarity is the cosine distance, a set number of reference cases for the task to be labeled are obtained in descending order of the cosine distance.

[0089] Step 3: Call the large language model to label each task to be labeled using the corresponding set number of reference cases, and obtain a labeled training set; the output structure is as follows: <semantic block - 1: [labeling information, vector encoding], semantic block - 2: [labeling information, vector encoding],..., semantic block - n: [labeling information, vector encoding]>.

[0090] Specifically, for any task to be labeled, based on a pre - set prompt library, a prompt "prompt" for the task to be labeled is constructed according to the set number of reference cases corresponding to the task to be labeled. The prompt library includes descriptive words representing the business types corresponding to the set number of reference cases; call the large language model to generate labeling information for the task to be labeled according to the prompt "prompt" of the task to be labeled; label the task to be labeled according to the labeling information.

[0091] The following is an example of the Prompt template in the prompt library:

[0092] "You are a document disassembling assistant. Please extract key information from the input text according to the given examples.

[0093] <example>

[0094] Semantic block - 1: [label key information]

[0095] Semantic block - 2: [label key information]

[0096] Semantic block - 3: [label key information]

[0097] Semantic block - 4: [label key information]

[0098] Semantic block - 5: [label key information]".

[0099] Construct a prompt using the set number of recalled reference cases, and use the few-shot ability of the large language model to annotate the task to be annotated. Utilize the few-shot ability of the large language model (the ability to learn by imitating a small number of samples), and incorporate the key information extraction examples that have been manually annotated into the prompt, enabling the large model to imitate the examples, so that key information can be extracted from new texts in the application.

[0100] The second stage: Train a lightweight text processing model based on the training set.

[0101] In the above first stage, a large language model with a large number of parameters (for example, greater than 100B) is used. Based on the labeled training set composed of a large amount of obtained labeled data, at this time, a large language model with a small number of parameters (such as 7B, 13B) is used for training, and a lightweight text processing model is trained based on the labeled training set.

[0102] The third stage: Application of the lightweight text processing model.

[0103] Import the text to be processed, split the text to be processed into multiple semantic document blocks according to the semantic structure, respectively use the trained lightweight text processing model to perform inference on the semantic document blocks, generate inference content, and finally integrate and output the processing result.

[0104] In summary, the embodiment of the present application provides a method for training a text processing model. First, perform vector encoding processing on the text samples of the first number to construct a set of cases to be annotated; calculate the similarity between each task to be annotated in the set of cases to be annotated and each reference case in the pre-set reference case set; among them, the reference case set is obtained after vector encoding processing based on the second number of pre-annotated reference texts; the first number is greater than the second number; further, for any task to be annotated, obtain the set number of reference cases of the task to be annotated in descending order of similarity; then call the large language model to annotate each task to be annotated using the corresponding set number of reference cases to obtain a labeled training set; train a lightweight text processing model based on the labeled training set. Solve the problem of low efficiency in text processing and key information extraction of texts through the large language model. By using a large language model with a large number of parameters to annotate and construct a training set, and then training a lightweight text processing model, the step of recalling similar few-shot examples is omitted, saving the cost of calling the large language model interface.

[0105] Based on the same technical concept, the embodiment of the present application also provides a text processing model training system, as Figure 3 shown, the system includes:

[0106] The module 301 for determining cases to be annotated is used to perform vector encoding processing on the text samples of the first number to construct a set of cases to be annotated;

[0107] A similarity calculation module 302 is configured to calculate the similarity between each to-be-annotated task in the to-be-annotated case set and each reference case in the preset reference case set; the reference case set is obtained after vector encoding processing based on a second number of pre-annotated reference texts; the first number is greater than the second number; the reference case set includes at least one type of reference case;

[0108] A reference case determination module 303 is configured to, for any to-be-annotated task, obtain a set number of reference cases of the to-be-annotated task in descending order of similarity;

[0109] An annotation module 304 is configured to call a large language model to annotate each to-be-annotated task with the corresponding set number of reference cases to obtain an annotation training set;

[0110] A training module 305 is configured to train a lightweight text processing model based on the annotation training set.

[0111] An embodiment of the present application also provides an electronic device corresponding to the method provided in the foregoing embodiment. Please refer to Figure 4 , which shows a schematic diagram of an electronic device provided in some embodiments of the present application. The electronic device 20 may include: a processor 200, a memory 201, a bus 202, and a communication interface 203. The processor 200, the communication interface 203, and the memory 201 are connected through the bus 202; a computer program that can run on the processor 200 is stored in the memory 201, and when the processor 200 runs the computer program, it executes the method provided in any of the foregoing embodiments of the present application.

[0112] Among them, the memory 201 may include a high-speed random access memory (RAM: Random Access Memory), and may also include a non-volatile memory, such as at least one disk memory. The communication connection between the system network element and at least one other network element is realized through at least one physical port 203 (which may be wired or wireless), and the Internet, wide area network, local area network, metropolitan area network, etc. can be used.

[0113] The bus 202 may be an ISA bus, a PCI bus, or an EISA bus, etc. The bus may be divided into an address bus, a data bus, a control bus, etc. Among them, the memory 201 is used to store a program, and after the processor 200 receives an execution instruction, it executes the program. The method disclosed in any of the foregoing embodiments of the present application can be applied to the processor 200 or implemented by the processor 200.

[0114] The processor 200 may be an integrated circuit chip with the ability to process signals. In the implementation process, each step of the above method can be completed by the integrated logic circuit of the hardware in the processor 200 or the instructions in the form of software. The above-mentioned processor 200 may be a general-purpose processor, including a central processing unit (CPU for short), a network processor (NP for short), etc.; it may also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as being executed and completed by the hardware decoding processor, or executed and completed by a combination of the hardware and software modules in the decoding processor. The software module may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory 201, and the processor 200 reads the information in the memory 201 and combines its hardware to complete the steps of the above method.

[0115] The electronic device provided in the embodiments of the present application and the method provided in the embodiments of the present application are based on the same inventive concept and have the same beneficial effects as the method adopted, run or implemented by it.

[0116] The embodiments of the present application also provide a computer-readable storage medium corresponding to the method provided in the foregoing embodiments. Please refer to Figure 5 , which shows that the computer-readable storage medium is an optical disc 30, on which a computer program (i.e., a program product) is stored. When the computer program is run by the processor, it will execute the method provided in any of the foregoing embodiments.

[0117] It should be noted that examples of the computer-readable storage medium may also include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other optical and magnetic storage media, which will not be elaborated here one by one.

[0118] The computer-readable storage medium provided in the above embodiments of the present application and the method provided in the embodiments of the present application are based on the same inventive concept and have the same beneficial effects as the method adopted, run or implemented by the application program stored therein.

[0119] It should be noted that:

[0120] The algorithms and displays provided herein are not inherently related to any particular computer, virtual apparatus, or other device. Various general-purpose apparatuses may also be used in conjunction with the teachings provided herein. The structure required to construct such apparatuses will be apparent from the above description. In addition, the present application is not directed to any particular programming language. It should be understood that the content of the present application described herein can be implemented using various programming languages, and the description of a particular language above is for the purpose of disclosing the best mode of the present application.

[0121] In the specification provided herein, a number of specific details are set forth. However, it can be understood that embodiments of the present application may be practiced without these specific details. In some instances, well-known methods, structures, and techniques have not been shown in detail so as not to obscure the understanding of this specification.

[0122] Similarly, it should be understood that, in order to streamline the present application and assist in understanding one or more of the various inventive aspects, in the above description of the exemplary embodiments of the present application, the various features of the present application are sometimes grouped together into a single embodiment, figure, or description thereof. However, the disclosed method should not be construed as reflecting an intention that the claimed present application requires more features than are expressly recited in each claim. Rather, as reflected in the following claims, the inventive aspects lie in less than all the features of the single foregoing disclosed embodiment. Thus, the claims following the detailed description are hereby expressly incorporated into the detailed description, with each claim standing on its own as a separate embodiment of the present application.

[0123] Those skilled in the art will appreciate that the modules in the devices in the embodiments can be adaptively changed and disposed in one or more devices different from the embodiments. The modules or units or components in the embodiments can be combined into one module or unit or component, and in addition, they can be divided into multiple sub-modules or sub-units or sub-components. Except that at least some of such features and / or processes or units are mutually exclusive, any combination can be used to combine all the features disclosed in this specification (including the accompanying claims, abstract, and drawings) and all the processes or units of any method or device so disclosed. Unless otherwise expressly stated, each feature disclosed in this specification (including the accompanying claims, abstract, and drawings) can be replaced by an alternative feature that provides the same, equivalent, or similar purpose.

[0124] In addition, those skilled in the art can understand that although some embodiments described herein include certain features included in other embodiments rather than other features, the combination of features of different embodiments means that it is within the scope of this application and forms different embodiments. For example, in the following claims, any one of the claimed embodiments can be used in any combination.

[0125] Each component embodiment of the present application can be implemented in hardware, or in software modules running on one or more processors, or in a combination thereof. Those skilled in the art should understand that a microprocessor or a digital signal processor (DSP) can be used in practice to implement some or all of the functions of some or all of the components in the virtual machine creation device according to the embodiments of the present application. The present application can also be implemented as a device or device program (such as a computer program and a computer program product) for executing part or all of the methods described herein. Such a program for implementing the present application can be stored on a computer-readable medium, or can be in the form of one or more signals. Such signals can be downloaded from an Internet website, or provided on a carrier signal, or provided in any other form.

[0126] It should be noted that the above embodiments illustrate rather than limit the present application, and those skilled in the art can design alternative embodiments without departing from the scope of the appended claims. In the claims, any reference signs placed between parentheses shall not be construed as limiting the claim. The word "comprising" does not exclude the presence of elements or steps not listed in the claim. The word "a" or "an" preceding an element does not exclude the presence of a plurality of such elements. The present application can be implemented by means of hardware including several different elements and by means of a suitably programmed computer. In the unit claims listing several devices, several of these devices can be embodied by the same item of hardware. The use of the words first, second, and third, etc. does not denote any order. These words can be interpreted as names.

[0127] As described above, the above are only the preferred specific embodiments of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present application should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claimed claims.

[0128] The above are only the preferred embodiments of the present invention, and do not thereby limit the patent scope of the present invention. Any equivalent structural transformation made under the concept of the present invention by using the content of the specification and drawings of the present invention, or any direct / indirect application in other related technical fields shall be included within the patent protection scope of the present invention.

Claims

1. A method for training a text processing model, characterized in that, The method includes: Performing vector encoding processing based on a first number of text samples to construct a set of cases to be labeled; Calculating the similarity between each task to be labeled in the set of cases to be labeled and each reference case in a preset reference case set; the reference case set is obtained by performing vector encoding processing on a second number of pre-labeled reference texts; the first number is greater than the second number; the reference case set includes at least one type of reference case; For any task to be labeled, obtaining a set number of reference cases for the task to be labeled in descending order of similarity; Invoking a large language model to label each task to be labeled using the corresponding set number of reference cases to obtain a labeled training set; Training a lightweight text processing model based on the labeled training set.

2. The method according to claim 1, characterized in that, The performing vector encoding processing based on a first number of text samples to construct a set of cases to be labeled includes: Splitting the first number of text samples according to a first set rule to obtain a first text segment group; Performing vector encoding processing on each first text segment in the first text segment group respectively to obtain text vectors of each first text segment, which are used as the set of cases to be labeled.

3. The method according to claim 2, characterized in that, The invoking a large language model to label each task to be labeled using the corresponding set number of reference cases includes: For any task to be labeled, based on a preset prompt word library, constructing a prompt word for the task to be labeled according to the corresponding set number of reference cases for the task to be labeled, where the prompt word library includes description words representing the business types corresponding to the set number of reference cases; Invoking the large language model to generate annotation information for the task to be labeled according to the prompt word of the task to be labeled; Annotating the task to be labeled according to the annotation information.

4. The method according to claim 1, wherein Before performing vector encoding processing based on a first number of text samples to construct a set of cases to be labeled, it further includes: Splitting the second number of reference texts according to a second set rule to obtain a second text segment group; Performing information annotation on each second text segment group respectively to obtain a set of annotations for each second text segment; Performing vector encoding processing on each second text segment group respectively to obtain a set of text vectors of each second text segment; Classifying the second text segment group, the set of annotations for each second text segment, and the set of text vectors to obtain the reference case set.

5. The method according to claim 4, wherein The calculating the similarity between each task to be labeled in the set of tasks to be labeled and each reference case in a preset reference case set includes: For each task to be labeled in the set of cases to be labeled, calculating a parameter of the similarity between the text vector of the task to be labeled and the text vector of each reference case in the reference case set, where the parameter of the similarity is the Euclidean distance or the cosine distance.

6. The method according to claim 5, wherein When the parameter of the similarity is the Euclidean distance, the obtaining a set number of reference cases for the task to be labeled in descending order of similarity includes: Obtaining a set number of reference cases for the task to be labeled in ascending order of the Euclidean distance; When the parameter of the similarity is the cosine distance, obtaining the set number of reference cases of the task to be labeled in the order from the largest to the smallest similarity includes: Obtaining the set number of reference cases of the task to be labeled in the order from the largest to the smallest cosine distance.

7. The method according to any one of claims 1 to 6, characterized in that, The method further includes: Receiving a text task to be processed, where the text task to be processed carries the text to be processed; Splitting the text to be processed according to the third set rule to obtain a number of target segments; Inputting the number of target segments into the lightweight text processing model to obtain a processed text result.

8. A text processing model training system, characterized in that, The system includes: A to-be-labeled case determination module, configured to perform vector encoding processing based on a first number of text samples to construct a set of to-be-labeled cases; A similarity calculation module, configured to calculate the similarity between each to-be-labeled task in the set of to-be-labeled cases and each reference case in a preset reference case set; the reference case set is obtained after vector encoding processing based on a second number of pre-labeled reference texts; the first number is greater than the second number; the reference case set includes at least one type of reference case; A reference case determination module, configured to, for any to-be-labeled task, obtain the set number of reference cases of the to-be-labeled task in the order from the largest to the smallest similarity; A labeling module, configured to call a large language model to label each to-be-labeled task with the corresponding set number of reference cases to obtain a labeled training set; A training module, configured to train a lightweight text processing model based on the labeled training set.

9. An electronic device, comprising: A memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the processor executes the computer program to implement the method according to any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, A computer-readable instruction is stored thereon, and the computer-readable instruction can be executed by the processor to implement the method according to any one of claims 1-7.

Citation Information

Cited By

  • Mechanical arm control method and system based on large model, terminal and medium

    CN121223778A

  • Large model training-oriented data annotation method and device, electronic equipment and storage medium

    CN121580017A