Text processing model training method and device and storage medium

By using the sample generation model to generate training data containing thinking paths, the initial text processing model is trained, which solves the problem of low output accuracy of model by manually labeling training samples, and achieves higher accuracy of reply information and logical reasoning capabilities.

CN120012942AActive Publication Date: 2025-05-16ALIBABA (CHINA) CO LTD

Patent Information

Application Number
CN202510464871.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-14
Publication Date
2025-05-16
Estimated Expiration
2045-04-14

AI Technical Summary

Technical Problem

In the prior art, manual annotated training samples are used to train the initial text processing model, resulting in low accuracy of the model output reply information.

Method used

By obtaining the sample generation model, the model is used to generate training sample data containing thinking paths, and using this data to train the initial text processing model to obtain the target text processing model. This model is used to retrieve problem information and output reply information based on the processing results.

Benefits of technology

It significantly improves the accuracy of the model's output reply information, enhances the model's logical reasoning and tool usage capabilities, and can better deal with complex problems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120012942A_ABST
    Figure CN120012942A_ABST
Patent Text Reader

Abstract

The invention discloses a text processing model training method and device and a storage medium, and relates to the field of artificial intelligence, and the method comprises the steps: obtaining a sample generation model which is used for generating training sample data containing a thinking path; adopting a sample generation model to generate training sample data containing a thinking path; and training the initial text processing model based on the training sample data to obtain a target text processing model. Through the method and the device, the problem that the accuracy of reply information output by the obtained model is relatively low due to the fact that the initial text processing model is trained by adopting a manually labeled training sample in the related technology is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence, and more specifically, to a training method, device and storage medium for a text processing model. Background Art

[0002] Advances in the field of natural language processing, especially the rise of large-scale language models, have revolutionized machine understanding and generation of human language. However, despite their outstanding performance in many natural language processing tasks, these models still show certain limitations when dealing with tasks that require deep logical reasoning and the use of external tools. The root cause of this phenomenon can be traced back to the inadequacy of the model's training data and training methods.

[0003] Traditional training data relies on manual annotation. In order to improve the logical reasoning ability of the model, researchers tend to use data samples containing step-by-step reasoning processes. This step-by-step reasoning process aims to help the model learn how to gradually expand its thinking and conduct methodical logical analysis, which is extremely critical to improving the model's logical reasoning and problem-solving skills. However, there are obvious disadvantages in relying on manual annotation to construct such data samples. Manual annotation not only consumes a lot of time, but also requires professional knowledge from the annotator. This requirement is particularly strict for complex logical reasoning tasks, resulting in extremely high costs for the data preparation process. In addition, manual annotation is difficult to quickly adapt to new fields and new types of problems, limiting the model's ability to adapt to a variety of scenarios. In addition, the efficiency and scale of manual annotation are difficult to meet the urgent need for large amounts of data in the pre-training stage of large-scale language models. The above problems directly restrict the performance of the model in logical reasoning tasks. The initial text processing model may perform poorly in reasoning in deep understanding, cross-document information retrieval, and combining multi-domain knowledge. Especially when dealing with complex problems, the accuracy and logical coherence of its answers are often unsatisfactory.

[0004] Currently, no effective solution has been proposed to the problem that the initial text processing model is trained by using manually annotated training samples in related technologies, resulting in low accuracy of the response information output by the obtained model. Summary of the invention

[0005] The main purpose of the present application is to provide a training method, device and storage medium for a text processing model, so as to solve the problem in the related art that the initial text processing model is trained using manually annotated training samples, resulting in low accuracy of the response information output by the obtained model.

[0006] In order to achieve the above-mentioned purpose, according to one aspect of the present application, a method for training a text processing model is provided. The method comprises: obtaining a sample generation model, wherein the sample generation model is used to generate training sample data containing a thinking path; using the sample generation model to generate training sample data containing a thinking path; training an initial text processing model based on the training sample data to obtain a target text processing model, wherein the target text processing model is used to perform retrieval enhancement processing on question information, and output reply information based on the processing result.

[0007] Furthermore, the sample generation model is obtained through the following steps: extracting paragraph samples from a text sample set; identifying keywords in the paragraph samples, and using a preset method to mask at least one keyword in the paragraph samples to obtain a target paragraph; determining an initial training sample based on the target paragraph; and training the sample model based on the initial training sample to obtain a sample generation model.

[0008] Furthermore, determining the initial training sample based on the target paragraph includes: determining the prompt word corresponding to the target paragraph; inputting the target paragraph and the prompt word into the initial text processing model, and outputting a first thinking path for the target paragraph and a first reply information for the target paragraph; when the contents of the first reply information and the paragraph sample are the same, taking the target paragraph, the prompt word corresponding to the target paragraph, the first thinking path and the first reply information as the initial training sample.

[0009] Furthermore, the initial training sample includes a target paragraph, a prompt word corresponding to the target paragraph, a first thinking path and a first reply information. The sample model is trained based on the initial training sample to obtain a sample generation model including: inputting the target paragraph and the prompt word corresponding to the target paragraph into the sample model, outputting a second thinking path for the target paragraph and a second reply information for the target paragraph; calculating a first loss value of the first thinking path and the second thinking path based on a first loss function, and calculating a second loss value of the first reply information and the second reply information based on the first loss function; updating the parameters of the sample model based on the first loss value and the second loss value, and repeatedly executing the steps of inputting the target paragraph and the prompt word corresponding to the target paragraph into the updated sample model, and calculating the loss value between the thinking path output by the updated sample model and the first thinking path, and the loss value between the reply information output by the updated sample model and the first reply information based on the first loss function, until the calculated loss value meets the first preset condition, and using the updated sample model as the sample generation model.

[0010] Furthermore, using the sample generation model to generate training sample data containing a thinking path includes: using the sample generation model to generate preset training samples, wherein the preset training samples include preset prompt words, preset text samples, processed preset text samples, preset thinking paths and preset reply information; when the preset text samples and the preset reply information are consistent, the preset prompt words, processed preset text samples, preset thinking paths and preset reply information are used as training sample data.

[0011] Furthermore, the training sample data includes preset prompt words, processed preset text samples, preset thinking paths and preset reply information. Training the initial text processing model based on the training sample data includes: inputting the preset prompt words and the processed preset text samples into the initial text processing model, outputting a third thinking path for the processed preset text samples and a third reply information for the processed preset text samples; calculating a third loss value of the preset thinking path and the third thinking path based on the second loss function, and calculating a fourth loss value of the preset reply information and the third reply information based on the second loss function; updating the parameters of the initial text processing model based on the third loss value and the fourth loss value, and repeatedly executing the steps of inputting the preset prompt words and the processed preset text samples into the updated initial text processing model, calculating the loss value between the thinking path output by the updated initial text processing model and the preset thinking path, and the loss value between the reply information output by the updated initial text processing model and the preset reply information, until the calculated loss value meets the second preset condition, and taking the updated initial text processing model as the target text processing model.

[0012] In order to achieve the above purpose, according to another aspect of the present application, an information processing method based on a target text processing model is provided. The method includes: receiving question information; inputting the question information into the above target text processing model; and outputting reply information corresponding to the question information through the target text processing model.

[0013] Furthermore, outputting reply information corresponding to the question information through the target text processing model includes: generating a target thinking chain corresponding to the question information through the target text processing model; obtaining requirement description information based on the target thinking chain and the question information through the target text processing model, wherein the requirement description information is used to indicate the information that needs to be supplemented in the target thinking chain; calling an external information source to obtain target requirement information corresponding to the requirement description information; filling the target requirement information into the target thinking chain to obtain the target thinking path; and generating reply information based on the target thinking path.

[0014] In order to achieve the above-mentioned purpose, according to another aspect of the present application, a training device for a text processing model is also provided. The device comprises: an acquisition unit, used to acquire a sample generation model, wherein the sample generation model is used to generate training sample data containing a thinking path; a generation unit, used to use the sample generation model to generate training sample data containing a thinking path; a first training unit, used to train the initial text processing model based on the training sample data to obtain a target text processing model, wherein the target text processing model is used to perform retrieval enhancement processing on question information, and output reply information based on the processing result.

[0015] In order to achieve the above-mentioned purpose, according to another aspect of the present application, an information processing device based on a target text processing model is provided. The device comprises: a receiving unit for receiving question information; an input unit for inputting the question information into the above-mentioned target text processing model; and an output unit for outputting reply information corresponding to the question information through the target text processing model.

[0016] According to another aspect of the present application, a computer-readable storage medium is provided, the computer-readable storage medium including a stored program, wherein when the program is running, the device where the computer-readable storage medium is located is controlled to execute any training method for a text processing model.

[0017] According to another aspect of the present application, an electronic device is provided, comprising: one or more processors, a memory, and one or more programs, wherein the one or more programs are stored in the memory and are configured to be executed by the one or more processors, and the one or more programs include a training method for executing any one of the text processing models.

[0018] According to another aspect of the present application, a computer program product is provided, comprising computer instructions, which, when executed by a processor, implement the steps of any one of the above-mentioned methods for training a text processing model.

[0019] In an embodiment of the present application, a sample generation model is obtained, wherein the sample generation model is used to generate training sample data containing a thinking path; the sample generation model is used to generate training sample data containing a thinking path; the initial text processing model is trained based on the training sample data to obtain a target text processing model, wherein the target text processing model is used to perform retrieval enhancement processing on question information and output reply information based on the processing result. Generating training data through the sample production model significantly reduces the dependence on manual labeling, thereby solving the technical problem that the initial text processing model is trained using manually annotated training samples, resulting in low accuracy of the reply information output by the obtained model.

[0020] In this application, the initial text processing model is trained using training sample data containing thinking paths because these sample data provide the steps and logical reasoning process for problem solving. During the training process, the initial text processing model gradually masters the method of solving the problem by learning these steps and reasoning processes. This not only enhances the model's understanding of the logical structure of the problem, but also improves its ability to apply different tools and strategies. Therefore, the target text processing model obtained through this training performs retrieval enhancement processing when facing new problems, can better perform logical reasoning and tool use, and generate content that is more in line with logic and actual needs, thereby improving the accuracy of the model output reply information. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] The drawings constituting a part of the present application are used to provide a further understanding of the present application. The illustrative embodiments and descriptions of the present application are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0022] Figure 1 A hardware structure block diagram of a computer terminal for implementing a training method for a text processing model is shown;

[0023] Figure 2 is a flowchart of a training method for a text processing model provided in an embodiment of the present application;

[0024] Figure 3 is a schematic diagram of the training process of the text processing model provided according to an embodiment of the present application;

[0025] Figure 4 is a flowchart of an information processing method based on a target text processing model provided in an embodiment of the present application;

[0026] Figure 5 is a schematic diagram of an information processing device based on a target text processing model provided according to an embodiment of the present application;

[0027] Figure 6 is a schematic diagram of a training device for a text processing model provided in an embodiment of the present application;

[0028] Figure 7 It is a structural block diagram of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION

[0029] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present application.

[0030] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0031] It should be noted that the relevant information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for display, data for analysis, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. For example, an interface is set up between this system and the relevant users or organizations. Before obtaining relevant information, it is necessary to send an acquisition request to the aforementioned user or organization through the interface, and obtain relevant information after receiving the consent information fed back by the aforementioned user or organization.

[0032] Example 1

[0033] According to an embodiment of the present application, an embodiment of a training method for a text processing model is also provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0034] The method embodiment provided in the first embodiment of the present application can be executed in a mobile terminal, a computer terminal or a similar computing device. Figure 1 FIG. 1 shows a hardware structure block diagram of a computer terminal (or mobile device) for implementing a training method for a text processing model. Figure 1As shown, the computer terminal 10 (or mobile device) may include one or more (102a, 102b, ..., 102n are used to illustrate) processors 102 (the processor 102 may include but is not limited to a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 104 for storing data, and a transmission device 106 for communication functions. In addition, it may also include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of the BUS bus), a network interface, a power supply and / or a camera. It can be understood by those skilled in the art that Figure 1 The structure shown is only for illustration and does not limit the structure of the above electronic device. Figure 1 More or fewer components as shown, or with Figure 1 Different configurations are shown.

[0035] It should be noted that the one or more processors 102 and / or other data processing circuits described above may generally be referred to herein as "data processing circuits". The data processing circuits may be embodied in whole or in part as software, hardware, firmware, or any other combination thereof. In addition, the data processing circuit may be a single independent processing module, or may be incorporated in whole or in part into any of the other components in the computer terminal 10 (or mobile device). As involved in the embodiments of the present application, the data processing circuit acts as a processor control (e.g., selection of a variable resistor terminal path connected to an interface).

[0036] The memory 104 can be used to store software programs and modules of application software, such as program instructions / data storage devices corresponding to the methods in the embodiments of the present application. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, that is, implementing the above methods. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some examples, the memory 104 may further include a memory remotely arranged relative to the processor 102, and these remote memories may be connected to the computer terminal 10 via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0037] The transmission device 106 is used to receive or send data via a network. The specific example of the above network may include a wireless network provided by a communication provider of the computer terminal 10. In one example, the transmission device 106 includes a network adapter (Network Interface Controller, NIC), which can be connected to other network devices through a base station so as to communicate with the Internet. In one example, the transmission device 106 can be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.

[0038] The display may be, for example, a touch screen liquid crystal display (LCD) that enables a user to interact with a user interface of the computer terminal 10 (or mobile device).

[0039] In the above operating environment, this application also provides Figure 2 The training method of the text processing model shown. Figure 2 It is a flowchart of the training method of the text processing model according to Example 1 of the present application.

[0040] Step S201, obtaining a sample generation model, wherein the sample generation model is used to generate training sample data containing a thinking path.

[0041] The above-mentioned thinking path can refer to the logical flow, decision-making process or a series of steps when solving a problem or completing a task. In machine learning, such samples can help the model understand how to achieve the final goal or result through a series of logical steps. Through training samples containing thinking paths, the model can learn how to extract useful information from data through unsupervised learning and imitate the way humans think, thereby making more reasonable decisions on complex problems.

[0042] Step S202: using a sample generation model to generate training sample data including a thinking path.

[0043] The training sample data generated by the sample generation model is training sample data that has not been manually annotated. For example, the training sample data generated by the sample generation model includes processed paragraphs, prompt words, thinking paths, and final answers. For example, the original text is "The frequency of use of smart speakers in the home is mainly affected by their functional diversity and user-friendliness. With the continuous advancement of technology, the accuracy of voice recognition of smart speakers has increased year by year, which allows users to interact with them more naturally and smoothly. In addition, the improvement in accuracy also means that smart speakers can understand user intentions more accurately and provide more personalized services, thereby attracting more users to use them as home entertainment and life assistance devices.", and the processed paragraph is "The frequency of use of smart speakers in the home is mainly affected by their functions... and users.... With the continuous advancement of technology, the accuracy of... of smart speakers has increased year by year, which allows users to interact with them more naturally and smoothly. In addition, the improvement of... also means that smart speakers can understand user intentions more accurately and provide more personalized services, thereby attracting more users to use them as home entertainment and life assistance devices.", and the prompt word is "Combined with the text content, supplement the content of... in the text,... represents several deleted keywords", for example, the thinking path includes the following steps:

[0044] 1. Determine the definition and importance of smart speaker functional diversity and user-friendliness;

[0045] 2. Analyze the evolution of speech recognition technology and identify key technological breakthroughs;

[0046] 3. Explore the specific impact of improved speech recognition accuracy on user interaction experience;

[0047] 4. How improved evaluation accuracy can promote the popularity and use of smart speakers in homes;

[0048] 5. Study the smart speaker market report to understand the relationship between accuracy, user satisfaction, and usage frequency;

[0049] 6. Simulate the use cases of smart speakers in home scenarios at different accuracy levels and predict the changing trend of usage frequency.

[0050] The final answer is "Increasing voice recognition accuracy will significantly increase the frequency of smart speaker usage in the home. Improved accuracy means smoother user interaction and reduces the frequency of erroneous commands, thereby increasing user satisfaction and dependence. In addition, more accurate voice recognition enables smart speakers to better understand users' personalized needs and provide customized services, further enhancing their appeal in the home environment. According to market trends, as the technology matures, the frequency of smart speaker usage is positively correlated with voice recognition accuracy. It is expected that improved accuracy will directly lead to an increase in usage frequency, especially in high-frequency usage scenarios such as daily inquiries, home entertainment, and smart home control."

[0051] Step S203, training the initial text processing model based on the training sample data to obtain a target text processing model, wherein the target text processing model is used to perform retrieval enhancement processing on the question information and output reply information based on the processing result.

[0052] Optionally, the initial text processing model is trained by unsupervised learning based on the training sample data in step S202, because these sample data provide the steps and logical reasoning process for problem solving. The above-mentioned initial text processing model can be a large language model or other models, which are not limited in this application. During the training process, the initial text processing model gradually masters the method of solving the problem by learning these steps and reasoning processes. This not only enhances the model's understanding of the logical structure of the problem, but also improves its ability to apply different tools and strategies. Therefore, the target text processing model obtained through this training performs retrieval enhancement processing when facing new problems, can better perform logical reasoning and tool use, and generate content that is more in line with logic and actual needs. Thereby improving the accuracy of the model output reply information.

[0053] In order to train a sample generation model that can generate a large amount of training data without relying on manual annotation, optionally, in the training method of the text processing model provided in the embodiment of the present application, the sample generation model is obtained by the following steps:

[0054] The first step is to extract paragraph samples from the text sample set.

[0055] Optionally, the above text sample set can be an encyclopedia, professional literature, news reports, etc. By extracting paragraphs from texts in different fields, the training samples can cover a wide range of text knowledge, which helps the model learn cross-disciplinary understanding and expression capabilities. By randomly extracting paragraph samples, the diversity of training data can be ensured, avoiding over-reliance on a certain type of text and causing model bias.

[0056] For example, a sample paragraph about "quantum computing" reads "Quantum computing is a type of computing that uses quantum mechanical phenomena, such as superposition and entanglement, to process information. Unlike traditional binary bits, quantum computers use quantum bits (qubits) as the basic unit of information, which can be in multiple states at the same time, greatly accelerating the solution of certain computing tasks."

[0057] The second step is to identify keywords in the paragraph sample and use a preset method to mask at least one keyword in the paragraph sample to obtain the target paragraph.

[0058] Optionally, the above preset method can select words with a high frequency of occurrence (for example, greater than a preset frequency) in the paragraph sample as keywords. Masking them can prompt the sample generation model to learn key information in the context. The preset method can also use natural language processing technology (such as part-of-speech tagging or named entity recognition) to identify nouns, proper nouns, numbers, etc. These words often carry specific information. Masking them can train the model's ability to understand and infer entity information. The preset method can also simply randomly select a certain proportion of words for masking. This method maintains data diversity while training the model's ability to cope with random missing information.

[0059] For example, after masking the keywords in the sample paragraph, the content of the target paragraph is "Quantum computing is a computing method that uses...phenomena, such as... and... to process information. Unlike traditional..., quantum computers use... as the basic unit of information and can be in multiple states at the same time, greatly accelerating the solution of certain... ", and the... in the content represents the masked keywords.

[0060] The third step is to determine the initial training samples based on the target paragraph.

[0061] For example, based on the target paragraph, the prompt word "Please explain the concept of quantum computing and complete the missing key information." is designed. According to the above prompt word, the target text processing model generates a thinking path, and the final answer is "Quantum computing is a computing method that uses quantum mechanical phenomena, such as superposition and entanglement, to process information. Unlike traditional binary bits, quantum computers use quantum bits (qubits) as the basic unit of information and can be in multiple states at the same time, thereby greatly accelerating the solution of certain computing tasks." Check whether the output answer is consistent with the content of the original sample paragraph. If the content is consistent, determine the target paragraph, prompt word, thinking path and final answer as the initial training sample.

[0062] The fourth step is to train the sample model based on the initial training samples to obtain a sample generation model.

[0063] For example, the sample model is trained using the initial training samples generated in the second and third steps to obtain a sample generation model. The sample generation model will be used to generate more training data containing thinking paths for subsequent text processing model training.

[0064] In summary, the sample generation model constructed through these steps can significantly improve the efficiency of generating training data and reduce the dependence on manually labeled training data.

[0065] In order to obtain an initial training sample with high data quality, optionally, in the training method of the text processing model provided in the embodiment of the present application, determining the initial training sample based on the target paragraph includes:

[0066] The first step is to determine the prompt words corresponding to the target paragraph.

[0067] Optionally, the target paragraph refers to a paragraph obtained by masking at least one keyword in the paragraph sample using a preset method. A prompt word that guides the initial text processing model to generate a thinking path needs to be designed for the target paragraph with masked keywords. The prompt word can enable the initial text processing model to understand the relevance of the missing keywords and the required task type. The designed prompt word can require the initial text processing model to complete the information, and can also require the initial text processing model to understand the logical relationship between the information. The prompt word helps the initial text processing model learn how to build a coherent thinking path and reasoning chain.

[0068] For example, the target paragraph is "On the moon, scientists used ...... technology to find new evidence that ...... exists on the moon's surface. This discovery may change our understanding of the moon's origin .......", and ...... represents the masked keywords. The prompt words designed for this target paragraph are "Please add the name of the technology mentioned in the paragraph, the specific content of the discovery, and the impact of this discovery on the theory of the moon's origin."

[0069] In the second step, the target paragraph and the prompt words are input into the initial text processing model, and the first thinking path for the target paragraph and the first response information for the target paragraph are output.

[0070] Optionally, the above-mentioned initial text processing model can be an existing, mature language model, or a large language model pre-trained for generating thinking paths. The target paragraph and prompt words determined in the first step are input into the initial text processing model. The task of the initial text processing model is to generate thinking paths and response information based on the prompt words to complete the masked keywords. The goal of this step is to generate preliminary thinking path samples through the existing initial text processing model to provide a learning benchmark for subsequent training. It should be noted that the initial text processing model and the sample model are different in functional positioning and processing goals. The sample generation model is obtained based on the sample model training, and the initial text processing model is used to output the first thinking path and the first response information for the target paragraph.

[0071] For example, input the target paragraph and prompt words of the first step above into the initial text processing model, and the thinking path generated by the initial text processing model is "I need to identify what technology is used for lunar exploration. I will search recent aerospace news and scientific journals. Then, I need to determine what was discovered through this technology, which may involve the geological structure of the lunar surface. Finally, I need to analyze the potential impact of this discovery on theories of the origin of the moon, which may point to new theories or evidence of the formation of the moon." The first response information is: "On the moon, scientists used laser spectroscopy analysis technology to discover new evidence of the existence of water ice on the lunar surface. This discovery may change our theory of the origin of the moon and support the hypothesis that the moon may have been formed by fragments of the earth after a huge impact."

[0072] In the third step, when the contents of the first reply information and the paragraph sample are the same, the target paragraph, the prompt words corresponding to the target paragraph, the first thinking path and the first reply information are used as initial training samples.

[0073] Optionally, by verifying the consistency of the generated response information with the original paragraph, it is ensured that only the correct thinking path and response information are used for training, thereby improving the high quality of the initial training samples and avoiding the negative impact of inaccurate or irrelevant thinking paths on model training.

[0074] For example, after verifying that the content of the response information and paragraph samples generated in the second step is the same, the target paragraph "On the moon, scientists used... technology to discover new evidence that... exists on the surface of the moon. This discovery may change our understanding of the origin of the moon...", ... represents the masked keywords. The prompt words corresponding to the target paragraph "Please supplement the name of the technology mentioned in the paragraph, the specific content of the discovery, and the impact of this discovery on the theory of the origin of the moon.", the first thinking path "I need to identify what technology is used for lunar exploration. I will search recent aerospace news and scientific journals. Then, I need to determine what was discovered through this technology, which may involve the geological structure of the lunar surface. Finally, I need to analyze the potential impact of this discovery on the theory of the origin of the moon, which may point to new theories or evidence of the formation of the moon." and the first response information "On the moon, scientists used laser spectroscopy analysis technology to discover new evidence that water ice exists on the surface of the moon. This discovery may change our theory of the origin of the moon and support the hypothesis that the moon may be formed by fragments of the earth after a huge impact." as initial training samples.

[0075] In summary, through the above steps, the data quality of the initial training samples can be ensured, high-quality training materials can be provided for the initial text processing model, and the capabilities of the target text processing model in information completion, logical reasoning, and tool use can be improved.

[0076] In order to obtain a sample generation model that can provide high-quality training data for the training of the initial text processing model, optionally, in the text processing model training method provided in the embodiment of the present application, the initial training sample includes a target paragraph, a prompt word corresponding to the target paragraph, a first thinking path and a first reply information, and the sample model is trained based on the initial training sample to obtain the sample generation model including:

[0077] In the first step, the target paragraph and the prompt words corresponding to the target paragraph are input into the sample model, and the second thinking path for the target paragraph and the second response information for the target paragraph are output.

[0078] Optionally, at the beginning of the training process, the sample model can be a pre-trained initial model, but its capabilities in retrieval enhancement and thinking path generation have not been optimized by the training method of the text processing model provided for the embodiment of the present application. By inputting the target paragraph and prompt words into this initial model, the second thinking path and the second reply information generated by it are obtained, and then compared with the output of the sample model, which is used as a benchmark to guide the updating of parameters and the optimization of the model. As the training process proceeds, the performance of the sample model will gradually improve, and eventually, after meeting specific preset conditions, it will evolve into a sample generation model with the ability to generate high-quality thinking paths and reply information.

[0079] For example, the initial training samples determined are: target paragraph "On the moon, scientists used... technology to find new evidence that... exists on the moon's surface. This discovery may change our understanding of the moon's origin...", ... represents the masked keywords. The prompt words corresponding to the target paragraph are "Please add the name of the technology mentioned in the paragraph, the specific content of the discovery, and the impact of this discovery on the theory of the moon's origin.", the first thinking path is "I need to identify what technology is used for lunar exploration. I will search recent aerospace news and scientific journals. Then, I need to determine what was discovered through this technology, which may involve the geological structure of the moon's surface. Finally, I need to analyze the potential impact of this discovery on the theory of the moon's origin, which may point to new theories or evidence of the moon's formation." and the first response information is "On the moon, scientists used laser spectroscopy analysis technology to find new evidence that water ice exists on the moon's surface. This discovery may change our theory of the moon's origin and support the hypothesis that the moon may have been formed by fragments of the earth after a huge impact.". Input this target paragraph and the prompt words corresponding to the target paragraph into the sample model, and output the second thinking path for the target paragraph: "I need to guess the technology used for lunar exploration. It may be some advanced technology, but I am not sure. Then, I have to speculate what this technology discovered, which may be related to a feature on the lunar surface. Finally, I try to understand the impact of this discovery on theories of the origin of the moon. This may involve some scientific theory, but I don’t know which one it is." The second response information for the target paragraph is "On the moon, scientists used high-tech technology to discover new evidence that there is some unidentified substance on the surface of the moon. This discovery may have an impact on theories of the origin of the moon, but I don’t know the specific impact."

[0080] In the second step, the first loss values ​​of the first thinking path and the second thinking path based on the first loss function are calculated, and the second loss values ​​of the first reply information and the second reply information based on the first loss function are calculated.

[0081] Optionally, the first loss function can be used to measure the difference between the second thinking path and the second reply information, and the first thinking path and the first reply information. The first loss function can include an evaluation of the logical consistency of the thinking path, the accuracy of keyword completion, and the consistency between the reply information and the target paragraph. The first loss value provides feedback on the quality of the content generated by the model in the current state and guides subsequent parameter updates.

[0082] For example, through the first loss function, we can calculate the structural difference between the first thinking path and the second thinking path in the example of the first step, the accuracy of keyword completion, and the consistency of the first reply information and the second reply information. Assume that the first loss function includes the following aspects: logical consistency (i.e., the rationality of the thinking path), the accuracy of keyword completion (i.e., whether all the masked keywords are correctly completed), and the accuracy of the reply information (i.e., whether the reply information matches the content of the target paragraph). Through the calculation of these indicators, the first loss value of 0.2 and the second loss value of 0.18 were obtained.

[0083] The third step is to update the parameters of the sample model based on the first loss value and the second loss value, and repeatedly input the target paragraph and the prompt word corresponding to the target paragraph into the updated sample model, and calculate the loss value between the thinking path output by the updated sample model and the first thinking path, and the loss value between the reply information output by the updated sample model and the first reply information based on the first loss function, until the calculated loss value meets the first preset condition, and the updated sample model is used as the sample generation model.

[0084] Optionally, the first loss value and the second loss value can be used to guide the parameter update of the sample model, and the model parameters can be adjusted through the back propagation algorithm to reduce the two loss values ​​until the preset optimization criteria are reached. This process may require multiple iterations, and each iteration will re-execute the first and second steps based on the model state updated in the previous iteration to gradually improve the model's ability to generate thinking paths and reply information until the first loss value meets the preset optimization conditions, at which point the sample model evolves into a sample generation model.

[0085] For example, the first loss value is 0.2, and the second loss value is 0.18. The back propagation algorithm is used to update the parameters and optimize the sample model. The input, output, and loss calculation process is repeated until both loss values ​​drop to a preset threshold, such as 0.01, indicating that the thinking path and response information generated by the model are highly consistent with the output of the sample model. At this point, the updated sample model is formally determined as the sample generation model, and can then be used to generate high-quality thinking path data on a large scale for further pre-training.

[0086] In summary, the sample generation model obtained by training through the above steps provides rich and high-quality training data for the subsequent training of the initial text processing model, thereby promoting the overall performance improvement of the target text processing model, especially in terms of tool usage ability and logical reasoning.

[0087] In order to obtain more accurate generated training sample data, optionally, in the training method of the text processing model provided in the embodiment of the present application, using the sample generation model to generate training sample data containing the thinking path includes:

[0088] In the first step, a sample generation model is used to generate preset training samples, wherein the preset training samples include preset prompt words, preset text samples, processed preset text samples, preset thinking paths and preset response information.

[0089] For example, the preset training samples generated by the sample generation model obtained after training include: the preset text sample "The distance between the Earth and Mars changes as they revolve around the sun. The closest distance is about 55 million kilometers, and the farthest distance is about 401 million kilometers. This change cycle is about 15 months, and every time this happens, scientists will take this opportunity to launch a probe.", the processed preset text sample "The distance between the Earth and Mars changes as they revolve around the sun. The closest distance is about..., and the farthest distance is about.... This change cycle is about..., and every time this happens, scientists will take this opportunity to launch a probe.", ... represents the covered content. Preset prompt words: "Please supplement the range of distance changes between the Earth and Mars, that is, the closest and farthest distances, and the length of this change cycle.", preset thinking path: "I need to determine the closest and farthest distances between the Earth and Mars, as well as the cycle of distance changes. I can obtain this information by querying astronomical data." and preset reply information: "The distance between the Earth and Mars changes as they orbit the sun. The closest distance is about 55 million kilometers, and the farthest distance is about 401 million kilometers. This change cycle is about 15 months. Every time this happens, scientists will take this opportunity to launch a probe."

[0090] In the second step, when the preset text sample and the preset response information are consistent, the preset prompt words, the processed preset text sample, the preset thinking path and the preset response information are used as training sample data.

[0091] Optionally, after obtaining the preset response information, you need to verify whether it correctly fills in the blanks in the processed preset text sample. You can check whether the preset text sample and the preset response information are consistent. In the above example, verify whether the completed "55 million kilometers", "401 million kilometers" and "15 months" are correct. When the preset text sample and the preset response information are consistent, the preset prompt words, the processed preset text sample, the preset thinking path and the preset response information can be used as training sample data.

[0092] In summary, through the above steps, effective and identical preset training samples are used for training the initial text processing model, so that the obtained target text processing model can correctly learn how to reason and answer questions in the absence of information, thereby improving the reasoning ability of the target text processing model.

[0093] In order to train a target text processing model with stronger logical reasoning ability, optionally, in the text processing model training method provided in the embodiment of the present application, the training sample data includes preset prompt words, processed preset text samples, preset thinking paths and preset reply information, and training the initial text processing model based on the training sample data includes:

[0094] In the first step, the preset prompt words and the processed preset text samples are input into the initial text processing model, and a third thinking path for the processed preset text samples and third response information for the processed preset text samples are output.

[0095] For example, the processed preset text sample "The distance between the Earth and Mars changes as they orbit the sun, and the closest distance is about ......, and the farthest distance can reach about ....... This change cycle is about ......, and every time this happens, scientists will use this opportunity to launch a probe." and the preset prompt "Please supplement the range of the distance change between the Earth and Mars, that is, the closest and farthest distances, and the length of this change cycle." are input into the initial text processing model, and the initial text processing model generates a third thinking path and a third response information based on its current learning state. The generated third thinking path is "I will try to recall information about the orbital cycle of the Earth and Mars, and then infer the closest and farthest distances based on this cycle." The third response information is: "The distance between the Earth and Mars changes as they orbit the sun, but I need to further search for the specific closest and farthest distances and cycles."

[0096] The second step is to calculate the third loss value of the preset thinking path and the third thinking path based on the second loss function, and calculate the fourth loss value of the preset reply information and the third reply information based on the second loss function.

[0097] Optionally, a second loss function is used to calculate the difference between the preset thinking path and the third thinking path, that is, the third loss value. The difference between the preset reply information and the third reply information is calculated, that is, the fourth loss value. The second loss function can consider the logical coherence of the thinking path, the accuracy of the reply information, and the correctness of the keyword completion. Assuming that in the preliminary output of the model, the correctness of the keyword completion is low, and the logical structure of the thinking path is relatively reasonable, the second loss value is calculated comprehensively based on these indicators. For example, the third loss value between the preset thinking path and the third thinking path calculated using the second loss function is 0.15, and the fourth loss value between the preset reply information and the third reply information is 0.3.

[0098] The third step is to update the parameters of the initial text processing model based on the third loss value and the fourth loss value, and repeatedly perform the steps of inputting the preset prompt words and the processed preset text samples into the updated initial text processing model, calculating the loss value between the thinking path output by the updated initial text processing model and the preset thinking path, and the loss value between the reply information output by the updated initial text processing model and the preset reply information, until the calculated loss value meets the second preset condition, and the updated initial text processing model is used as the target text processing model.

[0099] Optionally, based on the calculated third loss value and fourth loss value, the back propagation algorithm can be used to update the parameters of the initial text processing model to reduce the third loss value and the fourth loss value. In this step, the model gradually adjusts its internal parameters through continuous iterative learning to more accurately generate outputs that are consistent with the preset thinking path and preset response information. The updated initial text processing model will again accept the same preset prompt words and processed preset text samples, output a new third thinking path and third response information, and calculate the third loss value and the fourth loss value again until the two loss values ​​meet the second preset condition, that is, reach a set threshold (for example, less than 0.1), indicating that the model output is highly consistent with the preset answer, at which point the updated initial text processing model is considered to be trained.

[0100] In summary, through the above training steps, the initial text processing model can learn how to perform effective reasoning and keyword completion in the absence of information. This enables the target text processing model to demonstrate more powerful performance when processing natural language processing tasks, more accurately understand and generate complex and diverse text information, and meet a wide range of intelligent application needs, especially in scenarios that require retrieval enhancement and logical reasoning capabilities.

[0101] Optional, Figure 3 A diagram of the training process of a text processing model is provided. Follow the steps below to train and obtain the target text processing model:

[0102] The first step is data preparation. Randomly extract paragraphs from unlabeled text sources. Identify keywords such as named entities and numbers in the paragraphs, and randomly mask 1-10 keywords in the paragraphs.

[0103] The second step is to prepare initial data. Design specific prompt words, call the initial text processing model, and generate a thinking path with search calls and the final answer. These initial data will be used to train a small model (i.e., sample generation model) for production data to reduce costs. At the same time, the inference effect of the sample generation model can be better than that of the untrained initial text processing model.

[0104] The third step is to train the sample generation model. The sample generation model is trained using the initial data generated in the second step. The sample generation model has the ability to call and infer search tools.

[0105] Step 4: Prepare training data with thinking paths. Use the sample generation model obtained in step 3 to generate a large amount of training data with thinking paths.

[0106] Step 5: Data filtering. The sample generation model in step 4 generates a large amount of training data with thinking paths, and then filters them to verify whether the response information in the training data is correct (i.e., whether it is consistent with the content before masking). Only the training data with correct response information is retained to form the final training data set to ensure the reliability and accuracy of the reasoning path.

[0107] Step 6: Train the target text processing model. Use the filtered training data obtained in step 5 to train the target text processing model.

[0108] In summary, the target text processing model is trained through the steps in the above example, so that the logical reasoning ability of the target text processing model is stronger, thereby improving the accuracy of the response information output by the target text processing model.

[0109] It should be noted that the steps shown in the flowcharts of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and that, although a logical order is shown in the flowcharts, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0110] Example 2

[0111] Under the above operating environment, this application provides Figure 4 The information processing method based on the target text processing model is shown. Figure 4 It is a flowchart of the information processing method based on the target text processing model according to Example 2 of the present application.

[0112] Step S401, receiving question information.

[0113] Optionally, the above question information can be text information input by the user or the system, in order to obtain a response or processing result. The question information can be any form of natural language expression, including questions, commands, statements or requests, etc. The question information is the initial input of the information processing process, which directly determines the nature and direction of subsequent processing and responses. For example, the question information is "Which astronomical events have had a measurable impact on the Earth's climate in the past decade?".

[0114] Step S402: input question information into the target text processing model.

[0115] Optionally, the above-mentioned target text processing model refers to a model trained with training samples containing thinking paths generated by a sample generation model, which has better logical reasoning and external tool calling capabilities. The above-mentioned thinking path can refer to the logical flow, decision-making process or a series of steps when solving a problem or completing a task. The target text processing model generated by such trained content can better perform logical reasoning and tool use when facing new problems, and generate content that is more in line with logical and practical needs.

[0116] Step S403: outputting reply information corresponding to the question information through the target text processing model.

[0117] Optionally, the above-mentioned reply information may include a complete thinking path and a final answer constructed by the target text processing model based on the content of the question information, internal knowledge, and external retrieval information.

[0118] In summary, the accuracy of the reply information output by the target text processing model in the above steps is higher. This is because the target text processing model is trained with training samples generated by the sample generation model. The target text processing model has stronger logical reasoning and information retrieval capabilities, thereby providing users with more accurate and rich information replies, meeting the growing needs of natural language processing.

[0119] In order to utilize the reasoning ability of the target text processing model to generate more accurate reply information, optionally, in the information processing method based on the target text processing model provided in the embodiment of the present application, the reply information corresponding to the question information output by the target text processing model includes:

[0120] The first step is to generate the target thinking link corresponding to the problem information through the target text processing model.

[0121] Optionally, the above target thinking chain represents a series of step-by-step reasoning steps of the target text processing model when solving the problem. These steps can clearly show how the model starts from known information and finally gets the answer through logical analysis, information retrieval, calculation and other means. The target thinking chain can not only improve the reasoning transparency of the target text processing model, but also help the target text processing model learn more complex logical structures and reasoning strategies, so that it can perform better when facing tasks that require multi-step reasoning. The target text processing model can first have a deep understanding of the problem information, construct an initial framework of the target thinking chain, and fill in this initial framework through retrieval to obtain the above target thinking chain.

[0122] For example, a user asks a question, "Which astronomical events have had a measurable impact on the Earth's climate in the last decade?" The target text processing model first builds an initial framework and identifies three key concepts: "last decade," "astronomical events," and "Earth climate impact." The target text processing model fills this initial framework by retrieving information, and the target thought chain obtained may involve the types of astronomical events (such as solar activity, comet impacts, and asteroid flybys), analysis of the direct and indirect impacts of events on the Earth's climate, and the scientific community's evaluation and measurement of these impacts.

[0123] In the second step, the target text processing model is used to obtain the requirement description information based on the target thinking chain and problem information, wherein the requirement description information is used to indicate the information that needs to be supplemented in the target thinking chain.

[0124] For example, after analyzing the target thinking chain and problem information in the above example, the information that needs to be supplemented in the target thinking chain (corresponding to the above-mentioned demand information) is a list of major astronomical events that have occurred in the past ten years, an analysis of the theoretical impact of each event on the earth's climate, and so on.

[0125] The third step is to call external information sources to obtain target requirement information corresponding to the requirement description information.

[0126] Optionally, the external information source called refers to an information database or service stored outside the target text processing model. The external information source provides information that is not included in the training data of the target text processing model itself or is not rich and specific enough. The type of external information source can be a network search engine, a database and knowledge base, an API interface service, etc.

[0127] For example, the target text processing model in the above example will call on astronomical databases, climatological research materials, data from global weather stations, scientific journals and news reports to collect detailed records of astronomical events such as solar flare activity, close flybys of asteroids, changes in the lunar orbit, as well as theoretical predictions and actual observational data on these events' impact on the Earth's climate patterns.

[0128] The fourth step is to fill the target demand information into the target thinking chain to obtain the target thinking path.

[0129] For example, after collecting target demand information, the information is integrated into the constructed thinking chain to form a complete target thinking path that includes event identification and climate change impact analysis. The target thinking path describes in detail how the solar activity cycle affects the temperature and radiation balance of the earth, how asteroid flybys affect the climate by changing the particle distribution in the earth's atmosphere, and the long-term impact of changes in the lunar orbit on tides and ocean temperatures.

[0130] Step 5: Generate response information based on the target thinking path.

[0131] For example, the target text processing model can generate a response message based on its complex reasoning process: "In the past decade, abnormalities in the solar activity cycle, including the outbreak of multiple strong solar flares, have affected the Earth's radiation levels and temperature patterns. Asteroid A flew close to the Earth in February 2013. Although it did not directly impact the Earth, its particles and dust entered the atmosphere, causing slight changes in local light and temperature in the short term. In addition, the long-term slight changes in the moon's orbit indirectly affect the Earth's tidal patterns and ocean circulation, and have subtle but long-term effects on the global climate system."

[0132] In summary, the target thinking path is generated through the above steps, and then the reply information is generated based on the target thinking path. This process fully demonstrates the advantages of the target text processing model in information mining and integration, as well as its excellent performance in logical reasoning and expression ability, thereby providing users with a scientific, comprehensive and accurate reply information.

[0133] It should be noted that the steps shown in the flowcharts of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and that, although a logical order is shown in the flowcharts, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0134] Example 3

[0135] The embodiment of the present application also provides an information processing device based on a target text processing model. It should be noted that the information processing device based on a target text processing model in the embodiment of the present application can be used to execute the information processing method based on a target text processing model provided in the embodiment of the present application. The information processing device based on a target text processing model provided in the embodiment of the present application is introduced below.

[0136] According to an embodiment of the present application, a device for implementing the above-mentioned information processing method based on the target text processing model is also provided, such as Figure 5 As shown, the device includes: a receiving unit 501, an input unit 502 and an output unit 503.

[0137] Specifically, the receiving unit 501 is used to receive question information;

[0138] An input unit 502 is used to input question information into a target text processing model;

[0139] The output unit 503 is used to output the reply information corresponding to the question information through the target text processing model.

[0140] In the information processing device based on the target text processing model in the embodiment of the present application, the receiving unit 501 receives the question information; the input unit 502 inputs the question information into the target text processing model; the output unit 503 outputs the reply information corresponding to the question information through the target text processing model, which solves the technical problem that the accuracy of the reply information output by the obtained model is low when the initial text processing model is trained with manually annotated training samples. The target text processing model generated by training content with training sample data containing thinking paths can better perform logical reasoning and tool use, and generate content that is more in line with logic and actual needs. Thereby improving the accuracy of the reply information output by the target text processing model.

[0141] Optionally, in the information processing device based on the target text processing model provided in the embodiment of the present application, the output unit 502 includes: a first generation module, used to generate a target thinking chain corresponding to the problem information through the target text processing model; a first determination module, used to obtain requirement description information based on the target thinking chain and the problem information through the target text processing model, wherein the requirement description information is used to indicate the information that needs to be supplemented in the target thinking chain; an acquisition module, used to call an external information source to obtain target requirement information corresponding to the requirement description information; a filling module, used to fill the target requirement information into the target thinking chain to obtain the target thinking path; a second generation module, used to generate reply information based on the target thinking path.

[0142] It should be noted that the receiving unit 501, the input unit 502 and the output unit 503 correspond to steps S401 to S403 in Embodiment 1, and the three units and the corresponding steps implement the same examples and application scenarios, but are not limited to the contents disclosed in Embodiment 1. It should be noted that the modules or units may be hardware components or software components stored in a memory (e.g., memory 104) and processed by one or more processors (e.g., processors 102a, 102b, ..., 102n), and the modules may also be part of a device and run in the computer terminal 10 provided in Embodiment 1.

[0143] Example 4

[0144] According to an embodiment of the present application, a device for implementing the above-mentioned text processing model training method is also provided. The text processing model training device of the embodiment of the present application can be used to execute the text processing model training method provided in the embodiment of the present application. Figure 6 As shown, the device includes: an acquisition unit 601, a generation unit 602 and a first training unit 603.

[0145] Specifically, the acquisition unit 601 is used to acquire a sample generation model, wherein the sample generation model is used to generate training sample data containing a thinking path;

[0146] A generating unit 602, configured to generate training sample data including a thinking path using a sample generating model;

[0147] The first training unit 603 is used to train the initial text processing model based on the training sample data to obtain a target text processing model, wherein the target text processing model is used to perform retrieval enhancement processing on the question information and output reply information based on the processing result.

[0148] The training device of the text processing model provided in the embodiment of the present application obtains the sample generation model through the acquisition unit 601, wherein the sample generation model is used to generate training sample data containing the thinking path; the generation unit 602 uses the sample generation model to generate training sample data containing the thinking path; the first training unit 603 trains the initial text processing model based on the training sample data to obtain the target text processing model, which solves the technical problem that the accuracy of the response information output by the obtained model is low when the initial text processing model is trained with manually annotated training samples. The target text processing model generated by the training content using the training sample data containing the thinking path can better perform logical reasoning and tool use, and generate content that is more in line with logic and actual needs. Thereby improving the accuracy of the response information output by the target text processing model.

[0149] Optionally, in the training device of the text processing model provided in the embodiment of the present application, the device also includes: an extraction unit, used to extract paragraph samples from a text sample set; an identification unit, used to identify keywords in the paragraph samples, and use a preset device to mask at least one keyword in the paragraph sample to obtain a target paragraph; a determination unit, used to determine an initial training sample based on the target paragraph; and a second training unit, used to train the sample model based on the initial training sample to obtain a sample generation model.

[0150] Optionally, in the training device of the text processing model provided in the embodiment of the present application, the determination unit includes: a second determination module, used to determine the prompt word corresponding to the target paragraph; an input module, used to input the target paragraph and the prompt word into the initial text processing model, and output a first thinking path for the target paragraph and a first reply information for the target paragraph; a third determination module, used to use the target paragraph, the prompt word corresponding to the target paragraph, the first thinking path and the first reply information as the initial training sample when the contents of the first reply information and the paragraph sample are the same.

[0151] Optionally, in the training device of the text processing model provided in the embodiment of the present application, the second training unit includes: a first output module, used to input the target paragraph and the prompt words corresponding to the target paragraph into the sample model, and output the second thinking path for the target paragraph and the second reply information for the target paragraph; a first calculation module, used to calculate the first loss value of the first thinking path and the second thinking path based on the first loss function, and calculate the second loss value of the first reply information and the second reply information based on the first loss function; a second calculation module, used to update the parameters of the sample model based on the first loss value and the second loss value, and repeatedly execute the steps of inputting the target paragraph and the prompt words corresponding to the target paragraph into the updated sample model, and calculating the loss value between the thinking path output by the updated sample model and the first thinking path, and the loss value between the reply information output by the updated sample model and the first reply information based on the first loss function, until the calculated loss value meets the first preset condition, and the updated sample model is used as the sample generation model.

[0152] Optionally, in the training device of the text processing model provided in the embodiment of the present application, the generation unit 602 includes: a third generation module, used to generate preset training samples using a sample generation model, wherein the preset training samples include preset prompt words, preset text samples, processed preset text samples, preset thinking paths and preset reply information; a fourth determination module, used to use the preset prompt words, processed preset text samples, preset thinking paths and preset reply information as training sample data when the preset text samples and the preset reply information are consistent.

[0153] Optionally, in the training device of the text processing model provided in the embodiment of the present application, the first training unit 603 includes: a second output module, used to input the preset prompt words and the processed preset text samples into the initial text processing model, and output a third thinking path for the processed preset text samples and a third reply information for the processed preset text samples; a third calculation module, used to calculate the third loss value of the preset thinking path and the third thinking path based on the second loss function, and calculate the fourth loss value of the preset reply information and the third reply information based on the second loss function; a fourth calculation module, used to update the parameters of the initial text processing model based on the third loss value and the fourth loss value, and repeatedly perform the steps of inputting the preset prompt words and the processed preset text samples into the updated initial text processing model, calculating the loss value between the thinking path output by the updated initial text processing model and the preset thinking path, and the loss value between the reply information output by the updated initial text processing model and the preset reply information, until the calculated loss value meets the second preset condition, and the updated initial text processing model is used as the target text processing model.

[0154] It should be noted that the acquisition unit 601, the generation unit 602 and the first training unit 603 correspond to steps S201 to S203 in Example 1, and the three units and the corresponding steps implement the same examples and application scenarios, but are not limited to the contents disclosed in the above-mentioned Example 2. It should be noted that the above-mentioned modules or units may be hardware components or software components stored in a memory (e.g., memory 104) and processed by one or more processors (e.g., processors 102a, 102b, ..., 102n), and the above-mentioned modules may also be part of the device and may be run in the computer terminal 10 provided in Example 1.

[0155] Example 5

[0156] The embodiment of the present application may provide a computer terminal, which may be any computer terminal device in a computer terminal group. Optionally, in this embodiment, the computer terminal may also be replaced by a terminal device such as a mobile terminal or an electronic device.

[0157] Optionally, in this embodiment, the computer terminal may be located in at least one network device among a plurality of network devices of the computer network.

[0158] In this embodiment, the above-mentioned computer terminal can execute the program code of the following steps in the training method of the text processing model: obtaining a sample generation model, wherein the sample generation model is used to generate training sample data containing a thinking path; using the sample generation model to generate training sample data containing a thinking path; training the initial text processing model based on the training sample data to obtain a target text processing model, wherein the target text processing model is used to perform retrieval enhancement processing on the question information, and output reply information based on the processing results.

[0159] Optionally, the above-mentioned computer terminal can execute the program code of the following steps in the training method of the text processing model: the sample generation model is obtained by the following steps: extracting paragraph samples from a text sample set; identifying keywords in the paragraph samples, and using a preset method to mask at least one keyword in the paragraph samples to obtain a target paragraph; determining an initial training sample based on the target paragraph; training the sample model based on the initial training sample to obtain a sample generation model.

[0160] Optionally, the above-mentioned computer terminal can execute the program code of the following steps in the training method of the text processing model: determining the initial training sample based on the target paragraph includes: determining the prompt word corresponding to the target paragraph; inputting the target paragraph and the prompt word into the initial text processing model, and outputting the first thinking path for the target paragraph and the first reply information for the target paragraph; when the content of the first reply information and the paragraph sample is the same, taking the target paragraph, the prompt word corresponding to the target paragraph, the first thinking path and the first reply information as the initial training sample.

[0161] Optionally, the computer terminal can execute the following steps in the training method of the text processing model: the initial training sample includes a target paragraph, a prompt word corresponding to the target paragraph, a first thinking path and a first reply message, and the sample model is trained based on the initial training sample to obtain a sample generation model including: inputting the target paragraph and the prompt word corresponding to the target paragraph into the sample model, outputting a second thinking path for the target paragraph and a second reply message for the target paragraph; calculating a first loss value of the first thinking path and the second thinking path based on a first loss function, and calculating a second loss value of the first reply message and the second reply message based on the first loss function; updating the parameters of the sample model based on the first loss value and the second loss value, and repeatedly inputting the target paragraph and the prompt word corresponding to the target paragraph into the updated sample model, and calculating the loss value between the thinking path output by the updated sample model and the first thinking path, and the loss value between the reply message output by the updated sample model and the first reply message based on the first loss function, until the calculated loss value meets the first preset condition, and the updated sample model is used as the sample generation model.

[0162] Optionally, the above-mentioned computer terminal can execute the program code of the following steps in the training method of the text processing model: using the sample generation model to generate training sample data containing a thinking path includes: using the sample generation model to generate a preset training sample, wherein the preset training sample includes a preset prompt word, a preset text sample, a processed preset text sample, a preset thinking path and a preset reply information; when the preset text sample and the preset reply information are consistent, the preset prompt word, the processed preset text sample, the preset thinking path and the preset reply information are used as the training sample data.

[0163] Optionally, the above-mentioned computer terminal can execute the program code of the following steps in the training method of the text processing model: the training sample data includes preset prompt words, processed preset text samples, preset thinking paths and preset reply information, and training the initial text processing model based on the training sample data includes: inputting the preset prompt words and the processed preset text samples into the initial text processing model, outputting a third thinking path for the processed preset text samples and a third reply information for the processed preset text samples; calculating a third loss value of the preset thinking path and the third thinking path based on the second loss function, and calculating a fourth loss value of the preset reply information and the third reply information based on the second loss function; based on the third loss value and the fourth loss value, updating the parameters of the initial text processing model, repeatedly executing the steps of inputting the preset prompt words and the processed preset text samples into the updated initial text processing model, calculating the loss value between the thinking path output by the updated initial text processing model and the preset thinking path, and the loss value between the reply information output by the updated initial text processing model and the preset reply information, until the calculated loss value meets the second preset condition, and taking the updated initial text processing model as the target text processing model.

[0164] In this embodiment, the above-mentioned computer terminal can also execute the program code of the following steps in the information processing method based on the target text processing model: receiving question information; inputting the question information into the above-mentioned target text processing model; and outputting reply information corresponding to the question information through the target text processing model.

[0165] Optionally, the above-mentioned computer terminal can execute the program code of the following steps in the information processing method based on the target text processing model: outputting reply information corresponding to the question information through the target text processing model includes: generating a target thinking chain corresponding to the question information through the target text processing model; obtaining requirement description information based on the target thinking chain and the question information through the target text processing model, wherein the requirement description information is used to indicate the information that needs to be supplemented in the target thinking chain; calling an external information source to obtain the target requirement information corresponding to the requirement description information; filling the target requirement information into the target thinking chain to obtain the target thinking path; generating reply information based on the target thinking path.

[0166] Optionally, Figure 7 is a structural block diagram of an electronic device according to an embodiment of the present application. Figure 7 As shown, the electronic device may include: one or more ( Figure 7 Only one is shown) processor 702, memory 704, storage controller, and peripheral interface, wherein the peripheral interface is connected to the radio frequency module, audio module and display.

[0167] Among them, the memory can be used to store software programs and modules, such as the program instructions / modules corresponding to the training method and device of the text processing model in the embodiment of the present application. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, that is, realizing the training method of the above-mentioned text processing model. The memory may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory may further include a memory remotely arranged relative to the processor, and these remote memories can be connected to the terminal via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0168] The processor can call the information and application programs stored in the memory through the transmission device to execute the above steps in the training method of the above text processing model.

[0169] The embodiment of the present application provides a method for training a text processing model. By obtaining a sample generation model, wherein the sample generation model is used to generate training sample data containing a thinking path; the sample generation model is used to generate training sample data containing a thinking path; the initial text processing model is trained based on the training sample data to obtain a target text processing model, thereby solving the technical problem that the accuracy of the response information output by the obtained model is low when the initial text processing model is trained using manually annotated training samples, thereby improving the accuracy of the response information output by the target text processing model.

[0170] It can be understood by those skilled in the art that Figure 7 The structure shown is for illustration only, and the electronic device may also be a smart phone (such as an Android phone, an iOS phone, etc.), a tablet computer, a PDA, a mobile Internet device (Mobile Internet Devices, MID), a PAD, or other terminal devices. Figure 7 The structure of the electronic device is not limited. Figure 7 More or fewer components (such as network interfaces, display devices, etc.) shown in, or with Figure 7 Different configurations are shown.

[0171] A person of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing the hardware related to the terminal device through a program, and the program can be stored in a computer-readable storage medium, and the storage medium may include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.

[0172] Example 6

[0173] The embodiment of the present application also provides a computer-readable storage medium. Optionally, in this embodiment, the storage medium can be used to store the program code executed by the information processing method based on the target text processing model or the training method of the text processing model provided in the first embodiment.

[0174] Optionally, in this embodiment, the above storage medium may be located in any computer terminal in a computer terminal group in a computer network, or in any mobile terminal in a mobile terminal group.

[0175] Optionally, the storage medium is configured to store program code for performing the following steps: obtaining a sample generation model, wherein the sample generation model is used to generate training sample data containing a thinking path; using the sample generation model to generate training sample data containing a thinking path; training an initial text processing model based on the training sample data to obtain a target text processing model, wherein the target text processing model is used to perform retrieval enhancement processing on question information and output reply information based on the processing results.

[0176] Optionally, the storage medium is also configured to store program code for executing the following steps: the sample generation model is obtained by the following steps: extracting paragraph samples from a text sample set; identifying keywords in the paragraph samples, and masking at least one keyword in the paragraph samples using a preset method to obtain a target paragraph; determining an initial training sample based on the target paragraph; and training the sample model based on the initial training sample to obtain a sample generation model.

[0177] Optionally, the storage medium is also configured to store program code for executing the following steps: determining an initial training sample based on a target paragraph includes: determining a prompt word corresponding to the target paragraph; inputting the target paragraph and the prompt word into an initial text processing model, and outputting a first thinking path for the target paragraph and a first reply information for the target paragraph; when the contents of the first reply information and the paragraph sample are the same, taking the target paragraph, the prompt word corresponding to the target paragraph, the first thinking path and the first reply information as the initial training sample.

[0178] Optionally, the storage medium is also configured to store program codes for executing the following steps: an initial training sample includes a target paragraph, a prompt word corresponding to the target paragraph, a first thinking path and a first reply message, and a sample model is trained based on the initial training sample to obtain a sample generation model including: inputting the target paragraph and the prompt word corresponding to the target paragraph into the sample model, and outputting a second thinking path for the target paragraph and a second reply message for the target paragraph; calculating a first loss value of the first thinking path and the second thinking path based on a first loss function, and calculating a second loss value of the first reply message and the second reply message based on the first loss function; updating the parameters of the sample model based on the first loss value and the second loss value, and repeatedly executing the steps of inputting the target paragraph and the prompt word corresponding to the target paragraph into the updated sample model, and calculating the loss value between the thinking path output by the updated sample model and the first thinking path, and the loss value between the reply message output by the updated sample model and the first reply message based on the first loss function, until the calculated loss value meets the first preset condition, and the updated sample model is used as the sample generation model.

[0179] Optionally, the storage medium is also configured to store program code for executing the following steps: using a sample generation model to generate training sample data containing a thinking path includes: using a sample generation model to generate preset training samples, wherein the preset training samples include preset prompt words, preset text samples, processed preset text samples, preset thinking paths and preset reply information; when the preset text samples and preset reply information are consistent, the preset prompt words, processed preset text samples, preset thinking paths and preset reply information are used as training sample data.

[0180] Optionally, the storage medium is also configured to store program codes for executing the following steps: the training sample data includes preset prompt words, processed preset text samples, preset thinking paths and preset reply information, and training the initial text processing model based on the training sample data includes: inputting the preset prompt words and the processed preset text samples into the initial text processing model, outputting a third thinking path for the processed preset text samples and a third reply information for the processed preset text samples; calculating a third loss value of the preset thinking path and the third thinking path based on a second loss function, and calculating a fourth loss value of the preset reply information and the third reply information based on the second loss function; based on the third loss value and the fourth loss value, updating the parameters of the initial text processing model, and repeatedly executing the steps of inputting the preset prompt words and the processed preset text samples into the updated initial text processing model, calculating the loss value between the thinking path output by the updated initial text processing model and the preset thinking path, and the loss value between the reply information output by the updated initial text processing model and the preset reply information, until the calculated loss value meets the second preset condition, and taking the updated initial text processing model as the target text processing model.

[0181] Optionally, in this embodiment, the storage medium is also configured to store program codes for executing the following steps: receiving question information; inputting the question information into the above-mentioned target text processing model; and outputting reply information corresponding to the question information through the target text processing model.

[0182] Optionally, the storage medium is also configured to store program code for executing the following steps: outputting reply information corresponding to the question information through the target text processing model includes: generating a target thinking chain corresponding to the question information through the target text processing model; obtaining requirement description information based on the target thinking chain and the question information through the target text processing model, wherein the requirement description information is used to indicate the information that needs to be supplemented in the target thinking chain; calling an external information source to obtain target requirement information corresponding to the requirement description information; filling the target requirement information into the target thinking chain to obtain the target thinking path to generate reply information.

[0183] The present application also provides a computer program product, which, when executed on a data processing device, is suitable for executing the program steps of an information processing method based on a target text processing model or a training method of a text processing model.

[0184] The serial numbers of the above-mentioned embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.

[0185] In the above embodiments of the present application, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, please refer to the relevant description of other embodiments.

[0186] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.

[0187] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0188] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.

[0189] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions for a computer device (which can be a personal computer, server or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), mobile hard disk, disk or optical disk, etc., which can store program code.

[0190] The above is only a preferred implementation of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.

Claims

1. A training method for a text processing model, characterized in that: include: Obtaining a sample generation model, wherein the sample generation model is used to generate training sample data containing a thinking path; Using the sample generation model to generate training sample data containing thinking paths; The initial text processing model is trained based on the training sample data to obtain a target text processing model, wherein the target text processing model is used to perform retrieval enhancement processing on the question information and output reply information based on the processing result.

2. The method according to claim 1, characterized in that The sample generation model is obtained by the following steps: Extracting paragraph samples from a text sample collection; Identify keywords in the paragraph sample, and use a preset method to mask at least one keyword in the paragraph sample to obtain a target paragraph; Determining an initial training sample based on the target paragraph; The sample model is trained based on the initial training sample to obtain a sample generation model.

3. The method according to claim 2, characterized in that Determining an initial training sample based on the target paragraph includes: Determine the prompt word corresponding to the target paragraph; Inputting the target paragraph and the prompt word into an initial text processing model, and outputting a first thinking path for the target paragraph and first response information for the target paragraph; In the case where the contents of the first reply information and the paragraph sample are identical, the target paragraph, the prompt word corresponding to the target paragraph, the first thinking path and the first reply information are used as the initial training sample.

4. The method according to claim 2, characterized in that: The initial training sample includes the target paragraph, the prompt word corresponding to the target paragraph, the first thinking path and the first reply information. The sample model is trained based on the initial training sample to obtain the sample generation model including: Inputting the target paragraph and the prompt words corresponding to the target paragraph into the sample model, and outputting a second thinking path for the target paragraph and second response information for the target paragraph; Calculating first loss values ​​of the first thinking path and the second thinking path based on a first loss function, and calculating second loss values ​​of the first reply information and the second reply information based on the first loss function; Based on the first loss value and the second loss value, the parameters of the sample model are updated, and the steps of inputting the target paragraph and the prompt word corresponding to the target paragraph into the updated sample model are repeated, and based on the first loss function, the loss value between the thinking path output by the updated sample model and the first thinking path, and the loss value between the reply information output by the updated sample model and the first reply information are calculated, until the calculated loss value meets the first preset condition, and the updated sample model is used as the sample generation model.

5. The method according to any one of claims 1 to 4, characterized in that: Using the sample generation model to generate training sample data containing thinking paths includes: The sample generation model is used to generate preset training samples, wherein the preset training samples include preset prompt words, preset text samples, processed preset text samples, preset thinking paths and preset response information; In the case where the preset text sample and the preset reply information are consistent, the preset prompt word, the processed preset text sample, the preset thinking path and the preset reply information are used as the training sample data.

6. The method according to claim 1, characterized in that The training sample data includes preset prompt words, processed preset text samples, preset thinking paths and preset response information. Training the initial text processing model based on the training sample data includes: Inputting the preset prompt word and the processed preset text sample into the initial text processing model, and outputting a third thinking path for the processed preset text sample and third reply information for the processed preset text sample; Calculating a third loss value of the preset thinking path and the third thinking path based on a second loss function, and calculating a fourth loss value of the preset reply information and the third reply information based on the second loss function; Based on the third loss value and the fourth loss value, the parameters of the initial text processing model are updated, and the steps of inputting the preset prompt word and the processed preset text sample into the updated initial text processing model, calculating the loss value between the thinking path output by the updated initial text processing model and the preset thinking path, and the loss value between the reply information output by the updated initial text processing model and the preset reply information are repeatedly performed, until the calculated loss value meets the second preset condition, and the updated initial text processing model is used as the target text processing model.

7. An information processing method based on a target text processing model, characterized in that: include: Receive problem information; Inputting the question information into the target text processing model described in any one of claims 1 to 6; The target text processing model outputs reply information corresponding to the question information.

8. The method according to claim 7, characterized in that Outputting the reply information corresponding to the question information through the target text processing model includes: Generate a target thinking link corresponding to the question information through the target text processing model; Obtaining requirement description information based on the target thinking link and the question information through the target text processing model, wherein the requirement description information is used to indicate information that needs to be supplemented in the target thinking link; Calling an external information source to obtain target demand information corresponding to the demand description information; Fill the target demand information into the target thinking link to obtain the target thinking path; The reply information is generated according to the target thinking path.

9. A training device for a text processing model, characterized in that: include: An acquisition unit, used for acquiring a sample generation model, wherein the sample generation model is used for generating training sample data including a thinking path; A generating unit, used for generating training sample data including a thinking path by using the sample generating model; The first training unit is used to train the initial text processing model based on the training sample data to obtain a target text processing model, wherein the target text processing model is used to perform retrieval enhancement processing on the question information and output reply information based on the processing result.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a stored executable program, wherein when the executable program is run, the device where the computer-readable storage medium is located is controlled to execute the text processing model training method described in any one of claims 1 to 6.

11. An electronic device, characterized in that: include: A memory storing an executable program; A processor is used to run the program, wherein the program, when running, executes the training method of the text processing model described in any one of claims 1 to 6.

12. A computer program product comprising computer instructions, characterized in that: When the computer instructions are executed by a processor, the steps of the text processing model training method described in any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Keyword-based question and answer method and device and medium

    CN112487165A

  • Word weight generation model training method and device and word weight generation method and device

    CN114417863A

  • Natural language reasoning method, device and equipment based on reasoning path

    CN116596073A

  • Text reply method and training sample generation method

    CN118013278A

  • Retrieval sample selection and multi-round error correction method for solving mathematical problem of large language model

    CN119537567A

Cited By

  • Text processing model training method, text processing method and dialogue processing method

    CN120780815A

  • Text processing model training methods, text processing methods, and dialogue processing methods

    CN120780815B

  • Task interaction method and device, equipment, medium and program product

    CN121168668A

  • Model training method, text processing method and related device

    CN121188470A