Text processing method, system and equipment and storage medium
By splitting the target document and classifying semantic feature of large language models, building a training set for specific fields, solving the problems of high cost and limited training effects of building a specific field embedding model, and achieving efficient knowledge recall and model generalization.
Patent Information
- Application Number
- CN202311870567.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-30
- Publication Date
- 2025-07-04
AI Technical Summary
Building a domain-specific embedding model is costly and has limited training effects. The existing methods are time-consuming and labor-intensive and difficult to cover all domain-specific knowledge, resulting in insufficient generalization capabilities of the model.
By setting rules for the target document, using a large language model to classify text fragments based on semantic features and prompt words, a training set for specific fields, including primary classification and secondary classification, text fragments that do not belong to the initial category are selected, and positive and negative sample pairs are constructed for training.
It quickly improves the training effect and application performance of specific domain models, and improves the knowledge recall ability and generalization ability of the model.
Smart Images

Figure CN120256622A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of natural language processing, and specifically to text processing methods, systems, devices, and storage media. Background Art
[0002] In the process of knowledge question answering based on an external knowledge base, knowledge recall is a very important link. However, knowledge recall highly depends on the effect of the annotation (embedding) model. There are significant differences between general embedding models and domain-specific embedding models in many aspects. General embedding models are trained on a large amount of general text data, so they can capture the general characteristics of language, while domain-specific embedding models are trained on text data in a specific domain, so they can capture the specific semantics and context of that domain. Generally speaking, domain-specific embedding models are usually superior to general embedding models in domain-specific tasks because they can more accurately capture the semantics and context of that domain. However, the cost of building a domain-specific embedding model is relatively high, requiring a large amount of domain-specific text data and training time.
[0003] Currently, there are many methods for building a domain-specific embedding model, such as using manually collected training data, unsupervised training, etc. However, the method of manually collecting training data is not only time-consuming and laborious, but also difficult to cover all domain-specific knowledge, which may lead to weak generalization ability of the model. Although the method of unsupervised training can obtain training data on a large scale, due to the lack of effective annotation information, the training effect of the model may be affected. Therefore, it is necessary to find a more effective method that can overcome the above problems and quickly improve the training effect and application performance of the domain-specific embedding model. Summary of the Invention
[0004] The main purpose of the present invention is to provide a text processing method, system, device, and storage medium, which use a large language model to construct a training set for training a domain-specific model and solve the problem of knowledge recall of the large language model.
[0005] To achieve the above purpose, the embodiments of the present application provide the following technical solutions:
[0006] According to the first aspect of the embodiments of the present application, a text processing method is provided, and the method includes:
[0007] Split the target document based on a set rule to obtain a number of text fragments;
[0008] Perform a first classification on the number of text fragments according to the semantic features of each text fragment to obtain a set of first text type groups;
[0009] Call a large language model to perform secondary classification on the set of first text type groups based on semantic features and prompt words, to obtain a set of second text type groups, where at least one second text type group contains text segments different from those in at least one first text type group;
[0010] Construct a training set based on the set of second text type groups.
[0011] Optionally, call a large language model to perform secondary classification on the set of first text type groups based on semantic features and prompt words, to obtain a set of second text type groups, including:
[0012] Call the large language model to screen out the text segments to be processed that do not belong to the corresponding set of first text type groups from several first text type groups in the set of first text type groups;
[0013] Call the large language model to classify the text segments to be processed, to obtain the set of second text type groups.
[0014] Optionally, perform primary classification on several text segments according to the semantic features of each text segment, to obtain a set of first text type groups, including:
[0015] Extract the semantic features of each text segment to obtain the text vector corresponding to each text segment;
[0016] Cluster the text vectors corresponding to each text segment to obtain a set of first text type groups.
[0017] Optionally, the category definitions of the set of first text type groups match the category definitions of the set of second text type groups;
[0018] Call the large language model to screen out the text segments to be processed that do not belong to the corresponding set of first text type groups from several first text type groups in the set of first text type groups, including:
[0019] For any first text type group in the set of first text type groups, call the large language model to calculate the distance between the text vector corresponding to each text segment in the first text type group and the center of the set of first text type groups;
[0020] Obtain a set number of text segments of the first text type group in the order from the nearest to the farthest distance;
[0021] Construct a first prompt word according to the set number of text segments of the first text type group;
[0022] Input the first prompt word into the large language model to obtain the category definition of the first text type group;
[0023] Traverse each text segment in the first text type group, and construct a second prompt word according to the currently traversed text segment, a set number of text segments, and the category definition.
[0024] Input the second prompt word into the large language model to obtain several text segments of the second text type group belonging to the category definition, and determine the remaining several text segments in the second text type group as the text segments to be processed that do not belong to the second text type group.
[0025] Optionally, call the large language model to classify the text segments to be processed into the second text type group set, including:
[0026] Traverse the text segments to be processed, and construct a third prompt word according to the currently traversed text segment to be processed and the category definition of the second text type group set.
[0027] Input the third prompt word into the large language model to obtain the category definition corresponding to the text segments to be processed.
[0028] Classify the text segments to be processed into the second text type group set according to the corresponding category definition.
[0029] Optionally, extract the semantic features of each text segment to obtain the text vector corresponding to each text segment, including:
[0030] Call the general annotation model to extract the semantic features of each text segment to obtain each text segment and the text vector corresponding to each text segment.
[0031] According to the second aspect of the embodiments of the present application, a text processing system is provided, and the system includes:
[0032] A splitting module for splitting the target document based on a set rule to obtain several text segments.
[0033] A primary classification module for performing a primary classification on several text segments according to the semantic features of each text segment to obtain a first text type group set.
[0034] A secondary classification module for calling the large language model to perform a secondary classification on the first text type group set based on the semantic features and prompt words to obtain a second text type group set, where at least one second text type group has different text segments from at least one first text type group.
[0035] A training set construction module for constructing a training set based on the second text type group set.
[0036] According to the third aspect of the embodiments of the present application, an electronic device is provided, including: a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor runs the computer program, it is configured to implement the method of the first aspect above.
[0037] According to the fourth aspect of the embodiments of the present application, a computer-readable storage medium is provided, on which computer-readable instructions are stored. The computer-readable instructions can be executed by a processor to implement the method of the first aspect above.
[0038] In summary, the embodiments of the present application provide a text processing method, system, device, and storage medium. By splitting a target document based on set rules, a number of text fragments are obtained; according to the semantic features of each text fragment, the number of text fragments is classified once to obtain a first set of text type groups; further, a large language model is called to perform a secondary classification on the first set of text type groups based on the semantic features and prompt words to obtain a second set of text type groups; at least one second text type group contains different text fragments from at least one first text type group; and a training set is further constructed based on the second set of text type groups. Using a large language model to construct a training set for training a specific domain model solves the problem of knowledge recall of the large language model. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on the structures shown in these drawings.
[0040] The structures, ratios, sizes, etc. illustrated in this specification are only used to cooperate with the content disclosed in the specification for those who are familiar with this technology to understand and read, and are not used to limit the limited conditions under which the present invention can be implemented. Therefore, they do not have technical substance. Any modification of the structure, change of the proportional relationship, or adjustment of the size, without affecting the effects that the present invention can produce and the purposes that can be achieved, should still fall within the scope covered by the technical content disclosed in the present invention.
[0041] Figure 1 It is a flowchart of the text processing method provided by the embodiments of the present application;
[0042] Figure 2 It is a schematic diagram of the model construction process provided by the embodiments of the present application;
[0043] Figure 3 It is a block diagram of the text processing system provided by the embodiments of the present application;
[0044] Figure 4 The structural schematic diagram of an electronic device provided by an embodiment of the present application is shown;
[0045] Figure 5 The schematic diagram of a computer-readable storage medium provided by an embodiment of the present application is shown.
[0046] The realization of the object, functional features and advantages of the present invention will be further described with reference to the embodiments and the accompanying drawings. Detailed implementation manners
[0047] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without making creative efforts shall fall within the protection scope of the present invention.
[0048] It should be noted that all the directional indications (such as up, down, left, right, front, back...) in the embodiments of the present invention are only used to explain the relative position relationship and movement conditions between components in a specific posture (as shown in the accompanying drawings). If the specific posture changes, the directional indications will also change accordingly.
[0049] In addition, the descriptions such as "first" and "second" in the present invention are only for descriptive purposes, and cannot be understood as indicating or implying their relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of such features. In the description of the present invention, "a plurality" means at least two, such as two, three, etc., unless otherwise specifically defined.
[0050] In the present invention, unless otherwise clearly defined and limited, the terms "connection", "fixation", etc. shall be understood in a broad sense. For example, "fixation" may be a fixed connection, a detachable connection, or integrated; it may be a mechanical connection or an electrical connection; it may be directly connected or indirectly connected through an intermediate medium, and it may be the internal connection of two components or the interaction relationship between two components, unless otherwise clearly limited. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific situations.
[0051] In addition, the technical solutions between the various embodiments of the present invention can be combined with each other, but it must be based on the fact that those of ordinary skill in the art can implement it. When the combination of technical solutions is contradictory or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the protection scope required by the present invention.
[0052] Figure 1 The figure shows a text processing method provided by an embodiment of the present application. The method includes:
[0053] Step 101: Split the target document based on a set rule to obtain a number of text fragments;
[0054] Step 102: Perform a first classification on the number of text fragments according to the semantic features of each text fragment to obtain a set of first text type groups;
[0055] Step 103: Call a large language model to perform a second classification on the set of first text type groups based on the semantic features and prompt words to obtain a set of second text type groups. At least one second text type group contains different text fragments from at least one first text type group;
[0056] Step 104: Construct a training set based on the set of second text type groups.
[0057] In a possible implementation manner, the set rule in step 101 may be to split according to line breaks, full stops, question marks, exclamation marks, and / or semicolons.
[0058] In a possible implementation manner, it further includes: obtaining a text processing model according to the training set.
[0059] In a possible implementation manner, in step 102, performing a first classification on the number of text fragments according to the semantic features of each text fragment to obtain a set of first text type groups includes: extracting the semantic features of each text fragment to obtain text vectors corresponding to each text fragment; clustering the text vectors corresponding to each text fragment to obtain a set of first text type groups.
[0060] In a possible implementation manner, the above first classification method includes a clustering algorithm, such as k-means, etc.
[0061] In a possible implementation manner, extracting the semantic features of each text fragment to obtain text vectors corresponding to each text fragment includes:
[0062] Calling a general annotation model to extract the semantic features of each text fragment to obtain the text fragments of each text fragment and the text vectors corresponding to each text fragment.
[0063] In a possible implementation manner, the above general annotation model includes OpenAI, ada2, etc.
[0064] The classification accuracy of the clustering algorithm may not be as good as that of the large language model. The sentences under each category may not necessarily all belong to the same category. Therefore, the large language model is used for secondary confirmation.
[0065] In a possible implementation, in step 103, a large language model is called to perform secondary classification on the first text type group set based on semantic features and prompt words to obtain a second text type group set, including:
[0066] The large language model is called to screen out text fragments to be processed that do not belong to the corresponding first text type group set from several first text type groups in the first text type group set; the large language model is called to classify the text fragments to be processed to obtain a second text type group set.
[0067] In a possible implementation, the above large language model includes OpenAI, GPT4, etc.
[0068] In a possible implementation, the category definitions of the first text type group set match the category definitions of the second text type group set; the large language model is called to screen out text fragments to be processed that do not belong to the corresponding first text type group set from several first text type groups in the first text type group set, including:
[0069] For any first text type group in the first text type group set, the large language model is called to calculate the distance between the text vector corresponding to each text fragment in the first text type group and the center of the first text type group set; obtain a set number of text fragments of the first text type group in the order from near to far according to the distance; construct a first prompt word based on the set number of text fragments of the first text type group; input the first prompt word into the large language model to obtain the category definition of the first text type group; traverse each text fragment in the first text type group, and construct a second prompt word based on the currently traversed text fragment, the set number of text fragments, and the category definition; input the second prompt word into the large language model to obtain several text fragments of the second text type group belonging to the category definition, and determine the remaining several text fragments in the second text type group as text fragments to be processed that do not belong to the second text type group.
[0070] In a possible implementation, calling the large language model to classify the text fragments to be processed into the second text type group set includes:
[0071] Traverse the text fragments to be processed, and construct a third prompt word based on the currently traversed text fragment to be processed and the category definition of the second text type group set; input the third prompt word into the large language model to obtain the category definition corresponding to the text fragment to be processed; classify the text fragment to be processed into the second text type group set according to the corresponding category definition.
[0072] In a possible implementation, constructing a training set based on the second text type group set includes:
[0073] Construct positive and negative sample pairs according to the set of the second text type groups. The positive sample in the positive and negative sample pairs is a text segment in the text type group that matches the model to be trained, and the negative sample in the positive and negative sample pairs is a text segment in the text type group that does not match the model to be trained.
[0074] The above method utilizes the powerful reading comprehension ability of a large language model with a large number of parameters to quickly construct a training set in a specific domain, and solves the problem of recalling professional domain knowledge in the external knowledge base of the large language model.
[0075] Figure 2 The schematic diagram of the model construction process provided by the embodiment of the present application is shown, which specifically includes the following stages:
[0076] The first stage: Split a large number of unlabeled professional domain documents according to the sentence dimension to obtain a number of text segments.
[0077] Among them, splitting according to the sentence dimension can be splitting by line break, period, question mark, exclamation mark, and / or semicolon. For example, "Please rewrite question q 10 times and then return. One json per line (no line breaks), a total of 10 lines." will be split into two sentences "Please rewrite question q 10 times and then return." and "One json per line (no line breaks), a total of 10 lines."
[0078] The output form is <sentence1, sentence2, sentence3,...>.
[0079] The second stage: Extract sentence vectors from a number of text segments by using a general annotation embedding model.
[0080] In a possible implementation manner, the above general annotation model includes OpenAI, ada2, etc.
[0081] Extract the semantic features of each text segment to obtain the text vector corresponding to each text segment; the output form is <[sentence1, embedding1], [sentence2, embedding2], [sentence3, embedding3],...>.
[0082] The third stage: Cluster the text vectors corresponding to each text segment to obtain a set of the first text type groups, and the set of the first text type groups is a number of clustering clusters.
[0083] In a possible implementation manner, the above one-time classification method includes a clustering algorithm, such as k-means, etc.
[0084] Cluster the text vectors corresponding to each text segment so that sentences with similar semantics are clustered together to obtain several clustering clusters. The output format is <cluster1:[[sentence1, embedding1], [sentence2, embedding2]], cluster2:[[sentence3, embedding3], [sentence4, embedding4]],...>.
[0085] The fourth stage: Call the large language model LLM to perform secondary classification on several clustering clusters based on semantic features and prompt words to obtain clustering category definitions and undefined samples.
[0086] Call the large language model to perform secondary classification on several first text type groups in the first text type group set to obtain clustering category definitions: The output format is <cluster1, cluster2, cluster3,...>; Remove the sentences that do not belong to the clustering category and cache the sentence as <undefined sentence>. The specific steps are as follows:
[0087] Step 1: For each clustering cluster, call the large language model to calculate the distance between each text segment vector in the clustering cluster and the center of the clustering cluster; for example, if the clustering cluster is N clusters, for each clustering cluster, calculate the distance between each sample in the clustering cluster and the center of the cluster.
[0088] Step 2: Construct the first prompt word from a set number of text segments in the order of the distances in the clustering cluster from near to far; sort by distance from near to far, take out the M samples closest to the cluster center, and construct the first prompt word prompt1 together according to the prompt1 template. The prompt1 template is to summarize the category of the text.
[0089] Step 3: Input the first prompt word into the large language model to obtain the category definition of the clustering cluster;
[0090] Step 4: Traverse each text segment in the clustering cluster, and construct the second prompt word according to the currently traversed text segment, a set number of text segments, and the category definition of the clustering cluster; after obtaining the
category definition
[0091] Step 5: Input the second prompt into the large language model to obtain several text segments belonging to the cluster, and determine the remaining several text segments in the cluster as the text segments to be processed that do not belong to the cluster.
[0092] Fifth stage: Call the large language model LLM to annotate the undefined samples. Construct a prompt to classify <undefined sentences>, and finally all sentences can belong to one category.
[0093] Traverse the text segments to be processed, and construct the third prompt according to the currently traversed text segment to be processed and the category definitions of several clusters; input the third prompt into the large language model to obtain the category definition corresponding to the text segment to be processed, so as to complete the classification of the text segment to be processed into several clusters. Reclassify the undefined sentences, and finally the obtained clusters remain unchanged.
[0094] According to the prompt3 template, traverse and classify the text of the undefined category, and bring in "N category definitions" + "classify a single undefined text" in the prompt3 template to classify all undefined samples. The prompt3 template is "Classify the input text according to the given category definitions".
[0095] The output format is <cluster1:[[sentence1,embedding1],[sentence2,embedding2]],cluster2:[[sentence3,embedding3],[sentence4,embedding4]],...>.
[0096] Sixth stage: Construct a training set based on the clustering categories.
[0097] Construct positive and negative sample pairs. The positive sample in the positive and negative sample pairs is the text segment in the text type group that matches the model to be trained, and the negative sample in the positive and negative sample pairs is the text segment in the text type group that does not match the model to be trained. The output format is <[sentence1,sentence2(positive),sentence3(negative)],[sentence3,sentence4(positive),sentence1(negative)]>.
[0098] Seventh stage: Train an embedding model in the professional field according to the training set, such as a text processing model.
[0099] Through the powerful natural language understanding ability of the large language model, annotate on a large scale of unannotated data, and then train on the annotated data to improve the generalization ability and training effect of the model.
[0100] In summary, the embodiments of the present application provide a text processing method. By splitting the target document based on set rules, a number of text fragments are obtained; according to the semantic features of each text fragment, the number of text fragments is classified once to obtain a set of first text type groups; further, a large language model is called to perform a secondary classification on the set of first text type groups based on the semantic features and prompt words to obtain a set of second text type groups; at least one second text type group contains different text fragments from at least one first text type group; and a training set is further constructed based on the set of second text type groups. Using a large language model to construct a training set for training a specific domain model solves the problem of knowledge recall of the large language model.
[0101] Based on the same technical concept, the embodiments of the present application also provide a text processing system, which includes:
[0102] A splitting module 301, configured to split the target document based on set rules to obtain a number of text fragments;
[0103] A primary classification module 302, configured to perform a primary classification on the number of text fragments according to the semantic features of each text fragment to obtain a set of first text type groups;
[0104] A secondary classification module 303, configured to call a large language model to perform a secondary classification on the set of first text type groups based on the semantic features and prompt words to obtain a set of second text type groups, where at least one second text type group contains different text fragments from at least one first text type group;
[0105] A training set construction module 304, configured to construct a training set based on the set of second text type groups.
[0106] The embodiments of the present application also provide an electronic device corresponding to the method provided in the foregoing embodiments. Please refer to Figure 4 , which shows a schematic diagram of an electronic device provided in some embodiments of the present application. The electronic device 20 may include: a processor 200, a memory 201, a bus 202, and a communication interface 203. The processor 200, the communication interface 203, and the memory 201 are connected through the bus 202; a computer program that can run on the processor 200 is stored in the memory 201, and when the processor 200 runs the computer program, it executes the method provided in any of the foregoing embodiments of the present application.
[0107] Among them, the memory 201 may include high-speed random access memory (RAM), and may also include non-volatile memory, such as at least one disk memory. The communication connection between this system network element and at least one other network element is realized through at least one physical port 203 (which can be wired or wireless), and the Internet, wide area network, local area network, metropolitan area network, etc. can be used.
[0108] The bus 202 can be an ISA bus, a PCI bus, an EISA bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. Among them, the memory 201 is used to store programs. After receiving an execution instruction, the processor 200 executes the program. Any implementation manner of the method disclosed in any implementation manner of the foregoing embodiments of the present application can be applied to the processor 200 or implemented by the processor 200.
[0109] The processor 200 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the integrated logic circuit in the hardware of the processor 200 or the instructions in the form of software. The above-mentioned processor 200 can be a general-purpose processor, including a central processing unit (CPU for short), a network processor (NP for short), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as being executed by the hardware decoding processor, or executed by the combination of the hardware and software modules in the decoding processor. The software module can be located in a mature storage medium in the art, such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory 201, and the processor 200 reads the information in the memory 201 and combines its hardware to complete the steps of the above method.
[0110] The electronic device provided in the embodiments of the present application and the method provided in the embodiments of the present application are based on the same inventive concept and have the same beneficial effects as the method adopted, run, or implemented by it.
[0111] The embodiments of the present application also provide a computer-readable storage medium corresponding to the method provided in the foregoing embodiments. Please refer to Figure 5, which shows that the computer-readable storage medium is an optical disc 30, on which a computer program (i.e., program product) is stored. When the computer program is run by a processor, it will execute the method provided by any of the foregoing embodiments.
[0112] It should be noted that examples of the computer-readable storage medium may also include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory, or other optical and magnetic storage media, which will not be elaborated here one by one.
[0113] The computer-readable storage medium provided by the above embodiments of the present application and the method provided by the embodiments of the present application are based on the same inventive concept and have the same beneficial effects as the methods adopted, run, or implemented by the application programs stored therein.
[0114] It should be noted that:
[0115] The algorithms and displays provided herein are not inherently related to any particular computer, virtual device, or other equipment. Various general-purpose devices can also be used in conjunction with the teachings herein. The structure required to construct such devices is obvious from the above description. In addition, the present application is not directed to any particular programming language. It should be understood that the content of the present application described herein can be implemented using various programming languages, and the description of the specific language above is to disclose the best implementation mode of the present application.
[0116] In the specification provided herein, a large number of specific details are set forth. However, it can be understood that the embodiments of the present application can be practiced without these specific details. In some instances, well-known methods, structures, and technologies are not shown in detail so as not to obscure the understanding of this specification.
[0117] Similarly, it should be understood that, in order to streamline the present application and assist in understanding one or more of the various inventive aspects, in the foregoing description of the exemplary embodiments of the present application, the various features of the present application are sometimes grouped together into a single embodiment, figure, or description thereof. However, the disclosed method should not be construed as reflecting the intention that the claimed present application requires more features than are expressly recited in each claim. Rather, as reflected in the following claims, the inventive aspects lie in less than all the features of the single foregoing disclosed embodiment. Thus, the claims following the detailed description are hereby expressly incorporated into the detailed description, where each claim stands on its own as a separate embodiment of the present application.
[0118] Those skilled in the art can understand that the modules in the devices in the embodiments can be adaptively changed and set in one or more devices different from the embodiments. The modules or units or components in the embodiments can be combined into one module or unit or component, and in addition, they can be divided into multiple sub-modules or sub-units or sub-components. Except that at least some of such features and / or processes or units are mutually exclusive, any combination can be adopted to combine all the features disclosed in this specification (including the accompanying claims, abstract, and drawings) and all the processes or units of any method or device so disclosed. Unless otherwise explicitly stated, each feature disclosed in this specification (including the accompanying claims, abstract, and drawings) can be replaced by an alternative feature that provides the same, equivalent, or similar purpose.
[0119] In addition, those skilled in the art can understand that although some of the embodiments described herein include certain features included in other embodiments rather than other features, the combination of the features of different embodiments means that it is within the scope of this application and forms different embodiments. For example, in the following claims, any one of the claimed embodiments can be used in any combination.
[0120] Each component embodiment of the present application can be implemented in hardware, or in software modules running on one or more processors, or in a combination thereof. Those skilled in the art should understand that a microprocessor or a digital signal processor (DSP) can be used in practice to implement some or all of the functions of some or all of the components in the virtual machine creation device according to the embodiments of the present application. The present application can also be implemented as a device or device program (such as a computer program and a computer program product) for executing part or all of the methods described herein. Such a program implementing the present application can be stored on a computer-readable medium, or can be in the form of one or more signals. Such signals can be downloaded from an Internet website, or provided on a carrier signal, or provided in any other form.
[0121] It should be noted that the above embodiments illustrate the present application rather than limit the present application, and those skilled in the art can design alternative embodiments without departing from the scope of the appended claims. In the claims, any reference signs placed between parentheses shall not be construed as limiting the claim. The word "comprising" does not exclude the presence of elements or steps not listed in the claim. The word "a" or "an" preceding an element does not exclude the presence of a plurality of such elements. The present application can be implemented by means of hardware including several different elements and by means of a suitably programmed computer. In the unit claims listing several devices, several of these devices may be embodied by the same item of hardware. The use of the words first, second, and third, etc. does not denote any order. These words can be interpreted as names.
[0122] As described above, the above are only the preferred specific embodiments of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present application should be covered by the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the said claims.
[0123] The above are only the preferred embodiments of the present invention, and do not limit the patent scope of the present invention accordingly. All equivalent structural transformations made under the concept of the present invention by using the content of the specification and drawings of the present invention, or directly / indirectly applied to other related technical fields are included in the patent protection scope of the present invention.
Claims
1. A text processing method, characterized in that The method includes: Splitting the target document based on set rules to obtain a number of text fragments; Performing a first classification on the number of text fragments according to the semantic features of each text fragment to obtain a set of first text type groups; Invoking a large language model to perform a second classification on the set of first text type groups based on semantic features and prompt words to obtain a set of second text type groups, where at least one second text type group contains different text fragments from at least one first text type group; Constructing a training set based on the set of second text type groups.
2. The method according to claim 1, wherein It further includes: Training a text processing model according to the training set.
3. The method according to claim 1, characterized in that The step of invoking the large language model to perform a second classification on the set of first text type groups based on semantic features and prompt words to obtain a set of second text type groups includes: Invoking the large language model to screen out the text fragments to be processed that do not belong to the corresponding set of first text type groups from a number of first text type groups in the set of first text type groups; Invoking the large language model to classify the text fragments to be processed to obtain the set of second text type groups.
4. The method according to claim 1, wherein The step of performing a first classification on the number of text fragments according to the semantic features of each text fragment to obtain a set of first text type groups includes: Extracting the semantic features of each text fragment to obtain the text vectors corresponding to each text fragment; Clustering the text vectors corresponding to each text fragment to obtain the set of first text type groups.
5. The method according to claim 3 or 4, characterized in that, The category definitions of the set of first text type groups match the category definitions of the set of second text type groups; The step of invoking the large language model to perform a second classification on a number of first text type groups in the set of first text type groups and screening out the text fragments to be processed that do not belong to the corresponding set of first text type groups includes: For any first text type group in the set of first text type groups, invoking the large language model to calculate the distance between the text vector corresponding to each text fragment in the first text type group and the center of the set of first text type groups; Obtaining a set number of text fragments of the first text type group in the order from the nearest to the farthest distance; Constructing a first prompt word according to the set number of text fragments of the first text type group; Inputting the first prompt word into the large language model to obtain the category definition of the first text type group; Traversing each text fragment in the first text type group, and constructing a second prompt word according to the currently traversed text fragment, the set number of text fragments, and the category definition; Inputting the second prompt word into the large language model to obtain a number of text fragments of the second text type group belonging to the category definition, and determining the remaining number of text fragments in the second text type group as the text fragments to be processed that do not belong to the second text type group.
6. The method according to claim 5, wherein The step of invoking the large language model to classify the text fragments to be processed into the set of second text type groups includes: Traversing the text fragments to be processed, and constructing a third prompt word according to the currently traversed text fragment to be processed and the category definition of the set of second text type groups; Input the third prompt into the large language model to obtain the category definition corresponding to the text segment to be processed; Classify the text segment to be processed into the set of the second text type groups according to the corresponding category definition.
7. The method according to claim 4, wherein The extraction of the semantic features of each text segment to obtain the text vector corresponding to each text segment includes: Call a general annotation model to extract the semantic features of each text segment to obtain the text segments of each text segment and the text vector corresponding to each text segment.
8. The method according to claim 1, wherein The construction of the training set based on the set of the second text type groups includes: Construct positive and negative sample pairs according to the set of the second text type groups. The positive sample in the positive and negative sample pairs is the text segment in the text type group that matches the model to be trained, and the negative sample in the positive and negative sample pairs is the text segment in the text type group that does not match the model to be trained.
9. A text processing system, characterized in that, The system includes: A splitting module for splitting the target document based on a set rule to obtain a number of text segments; A primary classification module for performing primary classification on the number of text segments according to the semantic features of each text segment to obtain a set of the first text type groups; A secondary classification module for calling a large language model to perform secondary classification on the set of the first text type groups based on the semantic features and prompts to obtain a set of the second text type groups, where at least one second text type group contains different text segments from at least one first text type group; A training set construction module for constructing a training set based on the set of the second text type groups.
10. An electronic device, comprising: A memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the processor executes the computer program to implement the method according to any one of claims 1-8.
11. A computer-readable storage medium, characterized in that, A computer-readable instruction is stored thereon, and the computer-readable instruction can be executed by the processor to implement the method according to any one of claims 1-8.