Text processing model training methods, text processing methods, and dialogue processing methods
By training a large model with sample text, questions, and answers, the problem of inaccurate understanding of user intent in dialogue data analysis by large models is solved, and the analytical capabilities and accuracy of text processing models are improved.
Patent Information
- Application Number
- CN202511215560.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-28
- Publication Date
- 2026-01-30
- Estimated Expiration
- 2045-08-28
AI Technical Summary
Existing large models struggle to accurately understand user intent when analyzing dialogue data with multiple topics or constantly shifting focus, leading to inaccurate dialogue analysis results.
By acquiring sample text, sample questions, and sample answers, and using a large model to generate sample questions and answers, the text processing model is trained until a fully trained text processing model is obtained, thereby improving the model's text analysis capabilities.
This improves the accuracy of the text processing model's understanding of dialogue data, ensuring the accuracy of the analysis results.
Smart Images

Figure CN120780815B_ABST
Abstract
Description
Technical Field
[0001] The embodiments in this specification relate to the fields of computer technology and artificial intelligence technology, and in particular to text processing model training methods, text processing methods, and dialogue processing methods. Background Technology
[0002] In practical applications, general-purpose large models have made significant progress in language understanding and generation, and can often be used for text analysis and response generation. However, some problems still exist in real-world scenarios. For example, when analyzing a dialogue dataset, because a dialogue may contain multiple topics or focuses, and the context of the dialogue is constantly changing, the model may struggle to accurately capture the user's true intent, thus failing to accurately understand the meaning of the dialogue data. This leads to inaccurate dialogue analysis results or responses from the model. Therefore, an effective technical solution is urgently needed to address these issues. Summary of the Invention
[0003] In view of this, embodiments of this specification provide two methods for training text processing models. One or more embodiments of this specification also relate to two text processing model training devices, two text processing methods, two text processing apparatuses, a dialogue processing method, a dialogue processing apparatus, a training data generation method, a computing device, a computer-readable storage medium, and a computer program product, to address the technical deficiencies existing in the prior art.
[0004] According to a first aspect of the embodiments of this specification, a text processing model training method is provided, comprising:
[0005] Obtain sample text, a corresponding sample question, and a sample answer. The sample question and the sample answer are generated based on the text analysis elements corresponding to the sample text. The sample question is used to perform text analysis on the sample text, and the sample answer is the text analysis result of the sample text. The text analysis elements are used to represent the text content of the sample text.
[0006] The text processing model is trained based on the sample text, the sample question, and the sample answer until a fully trained text processing model is obtained.
[0007] According to a second aspect of the embodiments of this specification, a text processing model training apparatus is provided, comprising:
[0008] The acquisition module is configured to acquire sample text, a sample question corresponding to the sample text, and a sample answer, wherein the sample question and the sample answer are generated based on the text analysis elements corresponding to the sample text, the sample question is used to perform text analysis on the sample text, the sample answer is the text analysis result of the sample text, and the text analysis elements are used to represent the text content of the sample text;
[0009] The training module is configured to train the text processing model based on the sample text, the sample question, and the sample answer until a fully trained text processing model is obtained.
[0010] According to a third aspect of the embodiments of this specification, a text processing method is provided, comprising:
[0011] Identify the text to be processed;
[0012] The text to be processed is input into the text processing model to obtain the text analysis result output by the text processing model, wherein the text processing model is trained according to the text processing model training method provided in the embodiments of this specification.
[0013] According to a fourth aspect of the embodiments of this specification, a text processing apparatus is provided, comprising:
[0014] The determination module is configured to determine the text to be processed.
[0015] The input module is configured to input the text to be processed into the text processing model and obtain the text analysis result output by the text processing model, wherein the text processing model is trained according to the text processing model training method provided in the embodiments of this specification.
[0016] According to a fifth aspect of the embodiments of this specification, another text processing model training method is provided, applied to a cloud-based device, comprising:
[0017] The receiving end device sends a model training request, and obtains the sample text, sample question, and sample answer corresponding to the model training request. The sample question and sample answer are generated based on the text analysis elements corresponding to the sample text. The sample question is used to perform text analysis on the sample text, and the sample answer is the text analysis result of the sample text. The text analysis elements are used to represent the text content of the sample text.
[0018] The text processing model is trained based on the sample text, the sample question, and the sample answer until a fully trained text processing model is obtained, and the model parameters of the text processing model are sent to the edge device.
[0019] According to a sixth aspect of the embodiments of this specification, another text processing model training apparatus is provided, applied to a cloud-based device, comprising:
[0020] The receiving module is configured to receive a model training request sent by the end-side device, and obtain the sample text corresponding to the model training request, the sample question corresponding to the sample text, and the sample answer. The sample question and the sample answer are generated based on the text analysis elements corresponding to the sample text. The sample question is used to perform text analysis on the sample text, and the sample answer is the text analysis result of the sample text. The text analysis elements are used to represent the text content of the sample text.
[0021] The training module is configured to train the text processing model based on the sample text, the sample question, and the sample answer until a trained text processing model is obtained, and then send the model parameters of the text processing model to the edge device.
[0022] According to a seventh aspect of the embodiments of this specification, a dialogue processing method is provided, comprising:
[0023] Identify pending dialogues;
[0024] The dialogue to be processed is input into the text processing model to obtain the dialogue analysis results output by the text processing model, wherein the text processing model is trained according to the text processing model training method provided in the embodiments of this specification.
[0025] According to an eighth aspect of the embodiments of this specification, a dialogue processing apparatus is provided, comprising:
[0026] The "Determine" module is configured to determine pending dialogues.
[0027] The input module is configured to input the dialogue to be processed into the text processing model and obtain the dialogue analysis results output by the text processing model, wherein the text processing model is trained according to the text processing model training method provided in the embodiments of this specification.
[0028] According to a ninth aspect of the embodiments of this specification, another text processing method is provided, applied to a cloud-side device, comprising:
[0029] A text processing task sent by a receiving end-side device, wherein the text processing task carries text to be processed;
[0030] The text to be processed is input into the text processing model to obtain the text analysis result output by the text processing model, and the text analysis result is sent to the terminal device. The text processing model is trained according to the text processing model training method provided in the embodiments of this specification.
[0031] According to a tenth aspect of the embodiments of this specification, another text processing apparatus is provided, applied to a cloud-side device, comprising:
[0032] The receiving module is configured to receive a text processing task sent by the end-side device, wherein the text processing task carries text to be processed;
[0033] The sending module is configured to input the text to be processed into the text processing model, obtain the text analysis result output by the text processing model, and send the text analysis result to the end device, wherein the text processing model is trained according to the text processing model training method provided in the embodiments of this specification.
[0034] According to an eleventh aspect of the embodiments of this specification, a training data generation method is provided, including:
[0035] Obtain sample text;
[0036] Based on the text analysis elements corresponding to the sample text, a sample question and a sample answer are generated for the sample text. The sample question is used to perform text analysis on the sample text, and the sample answer is the text analysis result of the sample text. The text analysis elements are used to represent the text content of the sample text.
[0037] Training data is constructed based on the sample text, the sample question, and the sample answer.
[0038] According to a twelfth aspect of the embodiments of this specification, a computing device is provided, comprising:
[0039] Memory and processor;
[0040] The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions, which, when executed by the processor, implement the steps of the above method.
[0041] According to a thirteenth aspect of the embodiments of this specification, a computer-readable storage medium is provided that stores a computer program / instructions that, when executed by a processor, implement the steps of the above-described method.
[0042] According to page fourteen of the embodiments of this specification, a computer program product is provided, including a computer program / instructions that, when executed by a processor, implement the steps of the above-described method.
[0043] One embodiment of this specification provides a text processing model training method, comprising: acquiring sample text, a sample question corresponding to the sample text, and a sample answer, wherein the sample question and the sample answer are generated based on text analysis elements corresponding to the sample text, the sample question is used to perform text analysis on the sample text, the sample answer is the text analysis result of the sample text, and the text analysis elements are used to represent the text content of the sample text; and training the text processing model based on the sample text, the sample question, and the sample answer until a trained text processing model is obtained.
[0044] In the above method, after obtaining the sample text, a sample question for text analysis can be generated based on the text analysis elements corresponding to the sample text, and a sample answer corresponding to the sample question can be generated as the text analysis result of the sample text. The text processing model is trained based on the sample text, sample question, and sample answer, so that the text processing model can learn the text analysis ability to perform element analysis on the sample text from the sample question and sample answer. This enables the trained text processing model to accurately understand the meaning of the sample text, further improve the text analysis ability of the text processing model, and ensure the accuracy of the text analysis results. Attached Figure Description
[0045] Figure 1 This is a schematic diagram illustrating an application scenario of a text processing model training method provided in one embodiment of this specification.
[0046] Figure 2 This is a flowchart illustrating a text processing model training method provided in one embodiment of this specification;
[0047] Figure 3 This is a schematic diagram of the synthesis of the first training data in a text processing model training method provided in one embodiment of this specification;
[0048] Figure 4 This is a schematic diagram illustrating the training of the first text processing model in a text processing model training method provided in one embodiment of this specification;
[0049] Figure 5 This is a schematic diagram illustrating the training of a second text processing model in a text processing model training method provided in one embodiment of this specification.
[0050] Figure 6This is a training diagram of a third text processing model in a text processing method provided in one embodiment of this specification;
[0051] Figure 7 This is a flowchart illustrating the processing procedure of a text processing model training method provided in one embodiment of this specification.
[0052] Figure 8 This is a schematic diagram of the structure of a text processing model training device provided in one embodiment of this specification;
[0053] Figure 9 This is a flowchart illustrating a text processing method provided in one embodiment of this specification;
[0054] Figure 10 This is a schematic diagram of the structure of a text processing device provided in one embodiment of this specification;
[0055] Figure 11 This is a flowchart of another text processing model training method provided in one embodiment of this specification;
[0056] Figure 12 This is a schematic diagram of the structure of another text processing model training device provided in one embodiment of this specification;
[0057] Figure 13 This is a flowchart of a dialogue processing method provided according to an embodiment of this specification;
[0058] Figure 14 This is a schematic diagram of the structure of a dialogue processing device provided in one embodiment of this specification;
[0059] Figure 15 This is a flowchart of another text processing method provided in one embodiment of this specification;
[0060] Figure 16 This is a schematic diagram of the structure of another text processing device provided in one embodiment of this specification;
[0061] Figure 17 This is a structural block diagram of a computing device provided in one embodiment of this specification. Detailed Implementation
[0062] Many specific details are set forth in the following description to provide a full understanding of this specification. However, this specification can be implemented in many other ways than those described herein, and those skilled in the art can make similar extensions without departing from the spirit of this specification. Therefore, this specification is not limited to the specific implementations disclosed below.
[0063] The terminology used in one or more embodiments of this specification is for the purpose of describing particular embodiments only and is not intended to be limiting of the one or more embodiments of this specification. The singular forms “a,” “described,” and “the” as used in one or more embodiments of this specification and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used in one or more embodiments of this specification refers to and includes any or all possible combinations of one or more associated listed items.
[0064] It should be understood that although the terms first, second, etc., may be used to describe various information in one or more embodiments of this specification, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, first may also be referred to as second without departing from the scope of one or more embodiments of this specification, and similarly, second may also be referred to as first. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to a determination."
[0065] Furthermore, it should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in one or more embodiments of this specification are all information and data authorized by the user or fully authorized by all parties. Moreover, the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation entry points are provided for users to choose to authorize or refuse.
[0066] In one or more embodiments of this specification, a large model refers to a deep learning model with a large number of model parameters, typically containing hundreds of millions, tens of billions, hundreds of billions, trillions, or even tens of trillions of model parameters. A large model can also be called a foundation model. It is pre-trained using large-scale unlabeled corpora to produce a pre-trained model with hundreds of millions of parameters. Such models can adapt to a wide range of downstream tasks and have good generalization ability. Examples include Large Language Models (LLMs) and multi-modal pre-training models.
[0067] In practical applications, large models only require a small number of samples to fine-tune the pre-trained model before they can be applied to different tasks. Large models can be widely used in fields such as Natural Language Processing (NLP) and Computer Vision. Specifically, they can be applied to computer vision tasks such as Visual Question Answering (VQA), Image Captioning (IC), and Image Generation, as well as NLP tasks such as text-based sentiment classification, text summarization, and machine translation. The main application scenarios for large models include digital assistants, intelligent robots, search, online education, office software, e-commerce, and intelligent design.
[0068] First, the terms and concepts used in one or more embodiments of this specification will be explained.
[0069] LLM: Large Language Model, refers to a deep learning model with a huge number of parameters (usually in the hundreds of millions to trillions).
[0070] SFT: Supervised Fine-Tuning, is a process of supervising the training of a pre-trained large model (LLM) using a high-quality dataset with manually labeled input-output pairs to adapt it to a specific task or domain.
[0071] IOPO: Input-Output Preference Optimization. A training method that optimizes the ranking of input-output pairs based on human or model preferences.
[0072] GRPO: Group Relative Policy Optimization estimates the baseline by using relative rewards within a group, thus avoiding the use of additional value function models. GRPO significantly reduces memory and computational resource consumption by calculating the average reward from multiple outputs of the same problem.
[0073] QPS: Queries Per Second, a system performance metric that indicates the number of requests a system can process per second.
[0074] AMPO: Adaptive Mode Policy Optimization, is a reinforcement learning optimization method based on adaptive reasoning optimization using multiple thought patterns.
[0075] RL: Reinforcement Learning, a machine learning paradigm in which an agent adjusts its strategy based on the rewards it receives by interacting with the environment in order to maximize long-term cumulative rewards.
[0076] This specification provides two text processing model training methods, and also relates to two text processing model training devices, two text processing methods, two text processing devices, a dialogue processing method, a dialogue processing device, a computing device, a computer-readable storage medium, and a computer program product, which are described in detail in the following embodiments.
[0077] See Figure 1 , Figure 1 The illustration shows an application scenario diagram of a text processing model training method according to an embodiment of this specification, which includes the following steps.
[0078] Obtain sample text, a corresponding sample question, and a sample answer. The sample question and the sample answer are generated based on the text analysis elements corresponding to the sample text. The sample question is used to perform text analysis on the sample text, and the sample answer is the text analysis result of the sample text. The text analysis elements are used to represent the text content of the sample text.
[0079] The text processing model is trained based on the sample text, the sample question, and the sample answer until a fully trained text processing model is obtained.
[0080] Figure 1 It includes end-side device 102 and cloud-side device 104.
[0081] In practice, the user sends a model training request to the cloud device 104 via the edge device 102. The cloud device 104 responds to the request by obtaining the sample text corresponding to the request, and generating sample questions and answers based on the text analysis elements of the sample text. The text processing model is then trained using the sample text, sample questions, and sample answers until a fully trained text processing model is obtained. The cloud device 104 can send the model parameters of the trained text processing model to the edge device 102, allowing the edge device 102 to deploy the text processing model locally. Alternatively, the cloud device can send the model call interface of the trained text processing model to the edge device 102, enabling the edge device 102 to remotely invoke the text processing model through this interface.
[0082] The edge device 102 may include a browser, an app (application), or a web application such as an H5 (Hypertext Markup Language 5) application, a lightweight application (also known as a mini-program), or a cloud application. The edge device can be developed based on a software development kit (SDK) provided by the server, such as a real-time communication (RTC) SDK. The edge device can be deployed in an electronic device and depends on the device's operation or certain apps within the device to run. The electronic device may have a display screen and support information browsing, such as a personal mobile terminal like a mobile phone, tablet, or personal computer. Various other types of applications can also be configured in the electronic device, such as human-computer interaction applications, model training applications, text processing applications, web browser applications, shopping applications, search applications, instant messaging tools, email clients, and social media platform software.
[0083] Cloud-side device 104 can be understood as a server providing various services, including physical servers and cloud servers. Examples include servers providing communication services to multiple clients, servers supporting backend training of models used on clients, and servers processing data sent by clients. It should be noted that cloud-side device 104 can be implemented as a distributed server cluster composed of multiple servers, or as a single server. Cloud-side device 104 can also be a server for a distributed system, or a server integrated with blockchain. Cloud-side device 104 can also be a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms, or an intelligent cloud computing server or intelligent cloud host with artificial intelligence technology.
[0084] It is worth noting that the image processing method provided in the embodiments of this specification can be executed by the cloud-side device 104. In other embodiments of this specification, the image processing model can be deployed in the edge device 102, so that the edge device 102 can also have similar functions to the cloud-side device 104, thereby executing the image processing method provided in the embodiments of this specification. In other embodiments, the image processing method provided in the embodiments of this specification can also be jointly executed by the edge device 102 and the cloud-side device 104.
[0085] See Figure 2 , Figure 2A flowchart of a text processing model training method according to an embodiment of this specification is shown, which specifically includes the following steps.
[0086] Step 202: Obtain sample text, sample question and sample answer corresponding to the sample text, wherein the sample question and sample answer are generated based on the text analysis elements corresponding to the sample text, the sample question is used to perform text analysis on the sample text, the sample answer is the text analysis result of the sample text, and the text analysis elements are used to represent the text content of the sample text.
[0087] Specifically, the text processing model training method provided in the embodiments of this specification can be applied to the field of text analysis. Text analysis can be understood as analyzing elements such as theme, meaning, emotion, and intent of a text. For example, for a lesson text, its theme, meaning, scene, and the emotions it expresses can be analyzed. In one embodiment of this specification, the text can specifically be a dialogue, which can include dialogue between people, between people and machines, between machines, and complex dialogues involving multiple parties. For example, it can be a dialogue between a user and intelligent customer service and / or human customer service in the service field. Dialogue analysis can be understood as identifying key information in various dialogues, deeply analyzing potential reasons, summarizing dialogue patterns, inferring potential intents, and proposing solutions. Based on text analysis, the goal of continuously improving related capabilities can be achieved, further improving user experience, reducing service complaint rates, and simplifying cumbersome manual processes.
[0088] The sample text can be understood as the text that needs to be analyzed and used for training the text processing model. The sample text may be, for example, a dialogue or a lesson text; however, this specification does not limit the scope of the examples. For ease of understanding, this specification uses dialogue analysis as an example for illustration.
[0089] Specifically, multiple sample texts can be obtained from data sources in multiple fields for subsequent text processing model training.
[0090] In practical applications, multiple domains can be included, but are not limited to, customer service, finance, and service platforms. Sample texts can include, but are not limited to, Chinese and English texts. Based on this, a large amount of Chinese and / or English dialogue data can be obtained from data sources across multiple domains as sample texts to ensure the diversity of domain tasks and further improve the applicability of subsequent text processing models in different domains. Specifically, a large amount of dialogue data can be obtained from shared corpora as sample texts, allowing public participation in determining behaviors that influence the text processing model.
[0091] In addition, the sample text can be the chat log text between the parties, or the text converted from the audio of the telephone communication between the parties. This specification does not limit this.
[0092] In practical applications, obtaining the sample text, the corresponding sample question, and the sample answer includes:
[0093] Obtain sample text;
[0094] Based on the text analysis elements corresponding to the sample text, generate the sample question and the sample answer for the sample text.
[0095] Specifically, after obtaining the sample text, for each sample text, a sample question and a sample answer corresponding to the sample question can be generated based on the text analysis elements corresponding to the sample text.
[0096] Among them, text analysis elements can be understood as the elements contained in the sample text, sample questions can be understood as questions posed to the text analysis elements corresponding to the sample text, sample questions can include prompts, which can be understood as rules for generating sample answers, such as "Please answer in no more than 100 words", and sample answers can be understood as answers to the sample questions based on the sample text. Optionally, when the sample text is a sample dialogue, the sample text can include dialogue data and dialogue turns between the parties.
[0097] For example, the sample text could be “[1] Customer Service: Hello, I am a staff member of XXX Service Center. I am happy to serve you. How can I help you? [2] User: Uh, I would like you to help me cancel those XX ones. [3] User: Cancel all the bills I borrowed before. [4] Customer Service: Please provide the account you want to consult. [5] User: xxx……
[14] User: Yes.
[15] Customer Service: Are you the account holder?……
[35] User: Did you call?
[36] Customer Service: Well, he just doesn't know how to operate it, so he called to ask how to do it……”, where “[ 1, 2, 3, 14… can be understood as dialogue rounds. Therefore, the sample question for this sample text could be, “In the third round of dialogue, the user mentions ‘cancel’ the bill. Please analyze the user’s specific intention in using the expression ‘cancel’ and explain their possible mental state or demands in the context (please answer in no more than 100 words).” The sample answer could be, “The user’s use of ‘cancel’ indicates that they want to completely terminate the service relationship. This may be to avoid future additional costs or inconvenience caused by account management. It shows distrust of the function or no further demand for it.”
[0098] In practical applications, when the sample text is a dialogue, the text analysis elements may include, but are not limited to, information about the dialogue participants, the dialogue context, the dialogue theme, and the dialogue goal. The dialogue participant information may further include the basic information of the dialogue participants (name, gender, age, occupation, education, etc.), interests, personality, language and expression style, social relationships, and viewpoints. The dialogue context may further include linguistic background, physical background, relationships between people, and identities. The dialogue theme may further include viewpoints, intentions, emotions, moods, strategies, and summaries. The dialogue goal may further include the specific content of the dialogue and the degree to which the goal is achieved.
[0099] It is understandable that different text analysis elements can correspond to different sample questions. Therefore, based on multiple text analysis elements corresponding to the sample text, sample questions corresponding to each text analysis element can be generated, and sample answers corresponding to the sample questions can be generated. Furthermore, each text analysis element can also correspond to multiple sample questions, which is not limited in this embodiment of the specification.
[0100] In specific implementation, before generating the sample question and sample answer for the sample text based on the text analysis elements corresponding to the sample text, the method further includes:
[0101] The sample text is input into the text generation model, and the text generation model is used to extract the text analysis elements corresponding to the sample text.
[0102] The step of generating the sample question and the sample answer for the sample text based on the text analysis elements corresponding to the sample text further includes:
[0103] Using the text generation model, a sample question and a sample answer are generated for the sample text based on the text analysis elements.
[0104] In this context, the text generation model can have a larger model size than the subsequent text processing model that needs to be trained; that is, the text generation model has more parameters than the subsequent text processing model. In practical applications, the text generation model can be a large model.
[0105] Specifically, sample text can be input into a large model, where features are extracted and sample questions and answers are generated based on the extracted text analysis features. In other words, by inputting sample text into a large model, the model outputs sample questions and answers specific to that sample text.
[0106] Furthermore, text analysis elements can be extracted from sample text based on rule-based and keyword matching. Specifically, keywords, regular expressions, or semantic rules can be predefined to match specific words or sentence structures in the sample text, thereby extracting the corresponding text analysis elements. For example, if the sample text is a dialogue, and keywords such as "refund," "dissatisfaction," or "complaint" appear, the corresponding text analysis element dialogue scenario can be determined to be "after-sales complaint." If keywords such as "appointment," "time," or "doctor" appear, the corresponding text analysis element dialogue scenario can be determined to be "medical consultation." Alternatively, machine learning models or deep learning models can be used to classify the sample text to identify its text analysis elements. This specification does not limit this approach.
[0107] In summary, by using a text generation model with a larger model size to generate sample questions and sample answers, the training data for subsequent text processing models can be generated.
[0108] Furthermore, the step of extracting the text analysis elements corresponding to the sample text using the text generation model includes:
[0109] Using the text generation model, based on a pre-built text element system, the text analysis elements corresponding to the sample text are extracted;
[0110] The text element system is constructed based on the target task and contains reference text analysis elements in multiple dimensions.
[0111] The target task can be understood as the text analysis task of the text processing model, and the reference text analysis elements can be understood as the text analysis elements used for reference. For example, if the text analysis task of the text processing model is a dialogue analysis task, then the text element system can be understood as the dialogue scene elements. The text element system includes text analysis elements. For example, if the pre-built text element system includes dialogue participant information, then the text analysis elements corresponding to the sample text can be the dialogue participants "User A" and "Customer Service B" extracted from the sample text.
[0112] Specifically, a pre-built text element system can be input into the text generation model, enabling the model to learn the text element system and extract text analysis elements corresponding to the sample text based on it. In other words, the text element system can be understood as prompt words in the input text generation model to prompt the model to extract elements from the sample text.
[0113] In practical applications, within the field of dialogue analysis, the text element system includes multiple dimensions of reference text analysis elements, such as those prior to the dialogue, those during the dialogue, and those after the dialogue. When constructing the text element system, the corresponding dialogue analysis capabilities (points, lines, and surfaces) can be identified based on the scenarios before, during, and after the dialogue, ensuring the completeness of the constructed text element system.
[0114] Specifically, dialogue scene elements can include multiple dimensions of text analysis elements, such as participants, context, theme, and goals. Participants can further include basic information, interests, personality, language and expression, social relationships, viewpoints, and a summary of their personal information. Basic information can include name, gender, age, occupation, and education. Name and gender can be obtained from a resource library, and age can be enumerated using a finite set. The large model can generate participants' occupation, education, interests, language and expression, social relationships, and viewpoints based on names, enumerable personality attributes, and guided questions. Language and expression can include participants' social skills and decision-making styles. Viewpoints can include participants' cultural background, values, and character. Participants' personalities can be obtained by enumerating from a constructed personality resource pool. The summary of participants' personal information can be generated based on the aforementioned basic information, interests, personality, language and expression, social relationships, and viewpoints.
[0115] The context can include linguistic background, physical background, and relationships and identities of the participants. Linguistic background can include the topic and atmosphere of the conversation, while physical background can include the time, occasion, and environment of the conversation. Occasions can include online chat logs, telephone communication, face-to-face communication, etc., and environments can include physical descriptions such as weather and temperature. Relationships and identities can include familiarity and identity relationships. The familiarity of the participants can be defined based on a discrete scale of 0-10, and identity relationships can include superior-subordinate relationships, friendship relationships, supply and demand relationships, etc. The large model can extract the linguistic and physical background from the conversation and generate the relationships and identities of the participants.
[0116] The main theme can include viewpoints, intentions, emotions, moods, strategies, summaries, and workflows. A workflow can be understood as the steps and processes for completing a task or a series of related tasks. It defines how the task is carried out sequentially from start to finish, including the person responsible for each step, the resources required, the inputs and outputs, and the logical relationships between each step. The large model can generate dialogue processes / after completion.
[0117] The objective can include the specific content of the dialogue and the degree to which the objective is achieved. The large model can directly extract the specific content of the dialogue and score the degree to which the objective is achieved based on the dialogue content. The score of 0-10 controls the degree to which the dialogue objective is achieved.
[0118] Among the aforementioned dialogue scenario elements, understanding and modeling based on individual elements of a single-channel dialogue determines the participants and context of the dialogue, achieving point-based dialogue analysis capabilities. Understanding and modeling based on the context and timing of single / multi-channel dialogue determines the viewpoints, intentions, emotions, and strategies of the dialogue, achieving line-based dialogue analysis capabilities. Understanding and modeling based on the overall summary and analysis of single / multi-channel dialogue determines the dialogue summary and workflow, achieving surface-based dialogue analysis capabilities.
[0119] In summary, by constructing a text element system, we can serve as the basis for the text generation model to automatically generate sample questions and sample answers, thus providing a data foundation for the training of the text processing model.
[0120] Step 204: Train the text processing model based on the sample text, the sample question, and the sample answer until a fully trained text processing model is obtained.
[0121] Specifically, after obtaining sample questions and sample answers for the sample text, the sample text and sample questions can be used as training samples, and the sample answers can be used as training labels to train the text processing model until a fully trained text processing model is obtained.
[0122] In practical applications, the text processing model includes a first text processing model and a second text processing model. The first text processing model is used to output the answer, and the second text processing model is used to output the answer and the thought process.
[0123] The step of training the text processing model based on the sample text, the sample question, and the sample answer until a fully trained text processing model is obtained includes:
[0124] Construct the first training data based on the sample text, the sample question, and the sample answer;
[0125] Based on the first training data, the first text processing model is trained until the first text processing model is obtained after training.
[0126] Based on the sample text, the sample question, and the sample answer, construct second training data containing the thought process;
[0127] The second text processing model is trained based on the second training data until a fully trained second text processing model is obtained.
[0128] The first text processing model outputs a predicted answer to the input sample text and sample question. The predicted answer is the text analysis result of the first text processing model on the sample text. The second text processing model outputs the thought process while outputting the predicted answer to the input sample text and sample question.
[0129] Specifically, based on sample text, sample questions, and sample answers, first and second training data can be constructed. Using the first training data, the first text processing model undergoes three stages of training: supervised fine-tuning, offline reinforcement learning, and / or online reinforcement learning, until a fully trained first text processing model is obtained. Using the second training data, the second text processing model is trained using an adaptive reinforcement learning algorithm, until a fully trained second text processing model is obtained.
[0130] In summary, by constructing the first training data and the second training data, the first text processing model and the second text processing model can be trained.
[0131] In one embodiment of this specification, generating sample questions and sample answers for the sample text based on the text analysis elements corresponding to the sample text includes:
[0132] Based on the text analysis elements corresponding to the sample text, generate multiple sample questions for the sample text, as well as sample answers for each sample question;
[0133] The step of constructing the first training data based on the sample text, the sample question, and the sample answer includes:
[0134] Based on the sample text, a quality assessment is performed on the sample question and the sample answer to obtain the quality assessment result;
[0135] Based on the quality assessment results, target sample questions and target sample answers are determined from the plurality of sample questions and sample answers;
[0136] The sample text, the target sample question, and the target sample answer are used as the first training data.
[0137] The quality assessment of sample questions and answers can be understood as evaluating their reasonableness, correctness, and difficulty. The assessment result can be a score, where a higher score indicates greater reasonableness, correctness, and difficulty. The difficulty of the sample questions and answers can be understood as the difficulty of analyzing the sample text. Higher difficulty indicates a more in-depth analysis. For example, if sample question 1 is "Please determine the information of the participants in this dialogue," and sample question 2 is "Please analyze the user's request in this dialogue and analyze the possible reasons why the user made this request," then answering sample question 1 is less difficult than answering sample question 2. Therefore, the sample answer generated for sample question 2 provides a more in-depth analysis of the dialogue in the sample text.
[0138] Specifically, based on the sample text, a quality assessment can be performed on each sample question and sample answer corresponding to the sample text to obtain a quality assessment result. Sample questions and sample answers whose quality assessment results meet preset quality conditions are determined as target sample questions and target sample answers, and the sample text, target sample questions, and target sample answers are used as the first training data. Meeting the preset quality conditions can mean that the quality assessment score reaches a preset score threshold, or it can mean selecting the sample question and sample answer with the highest quality assessment result from multiple sample questions and sample answers as the target sample questions and target sample answers. This embodiment of the specification does not limit this. It is understood that there can be multiple target sample answers.
[0139] Furthermore, during the training of the first text processing model based on the first training data, the sample text and the target sample question can be used as training samples, and the target sample answer can be used as training labels to train the first text processing model.
[0140] In practical applications, large language models can be used as a kind of "referee" or "judge" to conduct quality assessments on sample questions and sample answers and obtain quality assessment results.
[0141] In summary, by conducting quality assessments on the sample questions and answers, and determining the target sample questions and answers that meet the requirements based on the quality assessment results, the high quality of the constructed first training data is ensured, further improving the training effect of the first text processing model.
[0142] In another embodiment of this specification, constructing the first training data based on the sample text, the sample question, and the sample answer further includes:
[0143] At least two sample problems are selected from the plurality of sample problems and combined to obtain combined sample problems;
[0144] From the sample answers corresponding to each sample question, determine the combined sample answer corresponding to the combined sample question;
[0145] The sample text, the combined sample question, and the combined sample answer are used as the first training data.
[0146] Specifically, when constructing the first training data, at least two different sample questions can be selected from multiple sample questions to combine them to obtain combined sample questions. Then, from the sample answers corresponding to each sample question, the combined sample answer corresponding to the combined sample question is determined. The sample text, combined sample question, and combined sample answer are used as the first training data.
[0147] Furthermore, during the training of the first text processing model based on the first training data, the sample text and the combined sample question can be used as training samples, and the combined sample answer can be used as training labels to train the first text processing model.
[0148] In one embodiment of this specification, see [link to embodiment]. Figure 3 , Figure 3 This diagram illustrates the synthesis of first training data in a text processing model training method according to an embodiment of this specification. Figure 3 As shown, a text element system can be pre-constructed based on the target task, and dialogue data can be collected from multiple domains. Each dialogue data point is then sequentially input into the text generation model (LLM) as sample text. Based on this model, multiple sample questions and corresponding sample answers are generated from the sample text. The quality of the sample questions and answers is then evaluated using a large model. Specifically, role recognition, dialogue turn recognition, and sample question validity assessment can be performed on the sample text to obtain quality evaluation results. Based on these results, a task combination is determined from the multiple sample questions and answers. This task combination serves as the first training data. For example, a task combination may include dialogue content C (i.e., sample text), sample question Q1, sample answer R1 corresponding to sample question Q1, sample question Qn, and sample answer Rn corresponding to sample question Qn. It is understandable that a task combination can include multiple sample questions and their corresponding sample answers. This sample combination is then used as training corpus for subsequent training.
[0149] Optionally, when constructing the first training data, after determining the target sample question and target sample answer from multiple sample questions and sample answers based on the quality assessment results, at least two target sample questions can be selected from the target sample questions to combine them to obtain a combined sample question. The combined sample answer corresponding to the combined sample question can be determined from the target sample answers corresponding to each target sample question. The sample text, combined sample question, and combined sample answer are then used as the first training data.
[0150] In summary, by selecting different sample questions from multiple sample questions corresponding to the sample text and combining them to synthesize more complex combined sample questions, the difficulty of the training data is increased, thereby further improving the training effect of the first text processing model.
[0151] In specific implementation, the step of constructing second training data containing the thought process based on the sample text, the sample question, and the sample answer includes:
[0152] The sample answers were expanded to obtain sample answers that included the thought process;
[0153] The sample text, the sample question, and the sample answer containing the thought process are used as the second training data.
[0154] Specifically, when constructing the second training data, the sample answers can be expanded to obtain sample answers that include the thinking process. This means that the sample answers containing the thinking process contain the thinking process itself, i.e., the process of how the sample answers were generated. The sample text, sample questions, and sample answers containing the thinking process are then used as the second training data.
[0155] Furthermore, the expansion of the sample answer to obtain a sample answer that includes the thought process includes:
[0156] Based on the preset thinking pattern, construct the reasoning link corresponding to the preset thinking pattern;
[0157] The sample answer is expanded based on the reasoning link to obtain a sample answer that includes the reasoning link, which is then used as the sample answer that includes the thinking process.
[0158] The preset thinking mode can be understood as the thinking pattern for generating sample answers. Preset thinking modes can include multiple different thinking modes, each containing different thinking steps. For example, thinking mode 1 is a basic thinking mode that does not include a thinking process. Thinking mode 2 includes a repetition and an answer; under thinking mode 2, the second text processing model can repeat certain information and provide an answer based on thinking mode 2. Thinking mode 3 adds a step of re-collecting information to thinking mode 2 before providing an answer, indicating that relevant information needs to be reconfirmed or collected before answering the question. Thinking mode 4 adds a deductive reasoning process to thinking mode 3, meaning that after repetition and re-collection of information, logical reasoning is still required to arrive at an answer. Thinking mode 5 adds a decomposition step to thinking mode 4, indicating that the problem needs to be broken down into smaller parts for analysis before deductive reasoning. Thinking mode 6 adds verification and integration steps to thinking mode 5, meaning that after completing all the previous steps, the correctness of the results can be verified, and the various parts can be integrated to form the final answer. Understandably, thinking patterns 1 through 6 can be understood as a process of thinking from simple to complex. Each thinking pattern adds new thinking elements compared to the previous one to ensure a more comprehensive and accurate solution to problems.
[0159] Specifically, based on multiple different preset thinking modes, a reasoning link corresponding to each preset thinking mode can be constructed, and the sample answer can be expanded based on the reasoning link to obtain a sample answer containing the thinking process corresponding to the reasoning link as a sample answer containing the thinking process.
[0160] In practical applications, six preset thinking modes with different depths of thinking can be designed, and second training data corresponding to each preset thinking mode can be constructed. The second training data includes sample text, sample questions, and sample answers containing the thinking process. Based on a certain proportion of the constructed second training data, the second text processing model is cold-started and trained. The model parameters of the second text processing model are adjusted based on the adaptive reinforcement learning algorithm until the trained second text processing model is obtained.
[0161] In summary, by constructing preset thinking patterns with different depths of thought and fully modeling various thinking patterns, the second text processing model can cope with problems of different difficulty and complexity, enabling it to achieve adaptive and efficient reasoning after training.
[0162] In practical applications, training the first text processing model based on the first training data until a fully trained first text processing model is obtained includes:
[0163] Based on the first training data, perform online reinforcement learning and / or offline reinforcement learning on the first text processing model until a fully trained first text processing model is obtained.
[0164] For details, see Figure 4 , Figure 4 This diagram illustrates the training of a first text processing model in a text processing model training method according to an embodiment of this specification. Figure 4 As shown, a large-scale model, LLM, can be used to extract elements and generate questions from sample text (i.e., raw dialogue data) based on a text element system (i.e., a dialogue scene element classification system). This yields question-and-answer pairs (i.e., sample questions and sample answers) for different text analysis elements (i.e., different tasks). The first text processing model is then trained using the sample text and question-and-answer pairs through three stages: supervised fine-tuning (SFT), offline reinforcement learning (IOPO), and online reinforcement learning (GRPO), achieving reinforcement learning (RL) until a fully trained first text processing model is obtained. Furthermore, the quality, diversity, and difficulty of the question-and-answer pairs for different tasks can be controlled, as described in the process of constructing the first training data. By synthesizing extensive training data through data engineering and performing effective modeling during training, more QPS calls can be supported with limited training resources. This allows the smaller-sized first text processing model to achieve performance comparable to larger-sized models, thereby reducing model training costs.
[0165] In practical applications, during the supervised fine-tuning stage, sample text and sample questions can be input into the first text processing model to obtain predicted answers. The first model loss value is calculated based on the first model loss function, the predicted answer, and the sample answer, and the first text processing model is trained based on the first model loss value. During the offline reinforcement learning stage, sample text and sample questions can be input into the first text processing model to obtain predicted answers. The second model loss value is calculated based on the second model loss function, the predicted answer, and the sample answer, and the first text processing model is trained based on the second model loss value. Specifically, during the offline reinforcement learning stage, training is performed using existing training data (i.e., the first training data) without real-time interaction. During the online reinforcement learning stage, sample text and sample questions can be input into the first text processing model to obtain predicted answers. The third model loss value is calculated based on the third model loss function, the predicted answer, and the sample answer, and the first text processing model is trained based on the third model loss value. In one embodiment of this specification, the first model loss function, the second model loss function, and the third model loss function can be different loss functions.
[0166] The loss function for the second model is shown in the following formula.
[0167]
[0168] in, These can be model parameters of the first text processing model. It can be training corpus, and It can be the model input of the first text processing model. It could be sample text 1 and the corresponding sample question 1. It can be sample text 1 and the corresponding sample question 2. Sample question 2 can be obtained by transforming sample question 1 to make it... and There are differences. and This is the output of the first text processing model. The first text processing model is designed for The output, The first text processing model is designed for The output, It can be a reference model. It can be a hyperparameter. It can be an activation function.
[0169] The loss function for the third model is shown in the following formula.
[0170]
[0171] in, It could be due to sample advantage. It can represent the probability ratio between the old and new strategy models. This can represent the calculation of KL divergence. It is a pruning parameter used to limit the magnitude of policy updates.
[0172] Specifically, for each sample problem q, from the current policy model (That is, a set of outputs o1, o2...oG (i.e., the predicted answer) is sampled from the second text processing model before training, which can be based on a pre-trained language model or a model that has been fine-tuned under supervision. The size of the group, using a reward model for each output. Rate the scores and receive corresponding reward values. Calculate the average reward value within the group and the standard deviation of the reward value within the group. Normalize the reward value of each output to obtain the normalized reward. Based on the normalized reward, calculate the dominance function, that is, the dominance function of each output at each time step t. The normalized reward is set to update the policy model (i.e., the second text processing model) by maximizing the third loss function mentioned above.
[0173] Further, the step of training the second text processing model based on the second training data until a fully trained second text processing model is obtained includes:
[0174] Based on a preset amount of second training data, the second text processing model is cold-started to obtain the cold-started second text processing model.
[0175] The sample question is input into the second text processing model after cold start training to obtain the predicted answer corresponding to each preset thinking mode.
[0176] Based on reinforcement learning algorithms, the sample-level dominance value corresponding to each predicted answer and the pattern-level dominance value corresponding to each preset thinking pattern are calculated.
[0177] The second text processing model after cold start training is trained based on the sample-level dominance value and the pattern-level dominance value until a fully trained second text processing model is obtained.
[0178] Here, the preset number of second training data can be understood as a certain proportion of the constructed second training data. The reinforcement learning algorithm can be the AMPO algorithm.
[0179] See Figure 5 , Figure 5 This diagram illustrates the training of a second text processing model in a text processing model training method according to an embodiment of this specification. Long-chain reasoning is an effective approach to solving high-order complex problems, and problems of different difficulties require different levels of thinking depth. Based on this, preset thinking patterns with different levels of thinking depth can be designed, and the adaptive reinforcement learning algorithm AMPO can be used to more fully model various thinking patterns, achieving the goal of adaptive and efficient reasoning. In practical applications, the second text processing model can be trained based on the first text processing model.
[0180] like Figure 5As shown, a text generation model can be used to construct second training data corresponding to each preset thinking pattern. This second training data includes sample text, sample questions, and sample answers containing the thinking process. A certain proportion of this constructed second training data is used to perform cold start training on the second text processing model. The model parameters are then adjusted using an adaptive reinforcement learning algorithm until a fully trained second text processing model is obtained. Cold start training can be understood as addressing the problem of insufficient initial data or historical information in the new second text processing model, which makes accurate predictions or effective decisions difficult at the beginning. Therefore, second training data can be constructed based on multiple preset thinking patterns for cold start training. Specifically, sample text and sample questions can be input into the second text processing model to obtain predicted answers. Based on the fourth model loss function, the fourth model loss value is calculated using the sample answers containing the thinking process and the predicted answers. The second text processing model is then trained based on the fourth model loss value.
[0181] Specifically, based on the AMPO algorithm, the reward values of multiple thinking modes can be considered simultaneously. The second text processing model can be trained based on the reward values corresponding to these multiple thinking modes, enabling it to select the most suitable thinking mode to obtain the highest reward value. Optionally, such as... Figure 5 As shown, in the behavior pattern cloning stage, six preset thinking patterns, M1, M2, M3, M4, M5 and M6, can be used to construct the second training data. The second training data includes sample text, sample questions and sample answers corresponding to various preset thinking patterns. Based on the second training data, the second text processing model is cold-started to obtain the cold-started second text processing model.
[0182] In the adaptive mode strategy optimization phase, the cold-start trained second text processing model outputs multiple predicted answers o1, o2, o3...oG for sample question q based on various preset thinking modes (M1 to Mn). Different predicted answers correspond to different preset thinking modes. For different preset thinking modes, the mean and variance of each preset thinking mode are calculated. Based on the mean and variance corresponding to each preset thinking mode, a mode-level group calculation is performed to obtain the dominance value corresponding to each preset thinking mode, i.e., the mode-level dominance value. A reward model is used to calculate the corresponding reward value for each output predicted answer. Based on the reward value corresponding to each output predicted answer, a sample-level group calculation is performed to obtain the dominance value of each sample (i.e., the predicted answer), i.e., the sample-level dominance value. Finally, the mode-level and sample-level dominance values are fed back to the strategy model (i.e., the second text processing model) through KL divergence to optimize the parameters of the cold-start trained second text processing model. A reference model can also be used to provide a benchmark to help evaluate the performance of the current strategy model. Different strategy outputs are generated through multiple thinking modes, and the performance of these strategies is evaluated by combining the reward model and the reference model. Then, the policy model is optimized through pattern-level and sample-level advantage calculations, enabling it to better adapt to the environment and obtain higher rewards. This multi-pattern, multi-level optimization approach helps improve the robustness and generalization ability of the policy model.
[0183] The loss function for the fourth model is shown in the following formula.
[0184]
[0185] in, Indicates the advantages at the pattern level. It can represent the advantage at the sample level. and It's a hyperparameter. This represents the probability ratio between the old and new strategy models. This represents the total number of pre-defined thought patterns. Indicates the total number of samples collected. Indicates the first sample in the sample group Sample The reward value.
[0186] Furthermore, the text processing model also includes a third text processing model;
[0187] The method further includes:
[0188] Based on the first text processing model and / or the second text processing model, the third text processing model is distilled and trained until a fully trained third text processing model is obtained.
[0189] The model size of the third text processing model is smaller than that of the first text processing model and the second text processing model.
[0190] Specifically, the third text processing model can be used to respond to user questions during a dialogue, outputting corresponding answers and sending them to the user to complete the dialogue. Therefore, the first and / or second text processing models can be used as teacher models, and the third text processing model as student models, for online distillation training and / or offline distillation training until a fully trained third text processing model is obtained. The knowledge and capabilities of the first and / or second text processing models are then injected into the third text processing model.
[0191] In practical applications, offline distillation training fixes the parameters of the trained teacher model, preventing updates during distillation. The student model learns by reading the teacher model's predicted answers to sample texts as soft labels. In online distillation training, the teacher and student models are trained simultaneously. The teacher model's parameters can be updated during distillation, while the student model learns in real-time from the teacher model. The student model learns not only the true labels (sample answers) but also the teacher model's current predicted answers (soft labels).
[0192] See Figure 6 , Figure 6 This diagram illustrates the training of a third text processing model in a text processing method according to an embodiment of this specification, as shown below. Figure 6As shown, a method combining reinforcement learning (RL) and knowledge distillation is used to transfer the knowledge of strong models (the first and / or second text processing models) to a smaller policy model (the third text processing model). Sample text and sample questions are input into the strong models to obtain predicted answers. To improve diversity and robustness, multiple strong models can be used for sampling during the inference phase to obtain different output results (i.e., predicted answers). These predicted answers serve as references for subsequent fine-tuning and training. Furthermore, consistency verification of predicted answers ensures the consistency of different models or sampling results, improving the reliability of the final output. Based on the predicted answers, the strong models are fine-tuned on specific tasks to better adapt them to the project scenario. The fine-tuned model further enhances its performance on specific tasks. Using data generated by the fine-tuned strong model, a weak model with fewer parameters (i.e., the third text processing model) is trained. Through knowledge distillation, the weak model learns the knowledge and capabilities of the strong model, thereby reducing computational resource consumption while maintaining performance. The third text processing model is a small, lightweight model responsible for making decisions or generating content in real-world project scenarios. It receives input (such as user requests and contextual information) and outputs corresponding actions. The model is trained using a reward mechanism, which includes process rewards, outcome rewards, and format rewards. Process rewards are given based on intermediate states or behaviors (s1, s2, ..., sn) during the execution process, encouraging the model to take reasonable steps. Outcome rewards are given based on the quality of the final output (label), ensuring the model generates content that meets expectations. Format rewards evaluate the format and syntax of the output content, ensuring the standardization and readability of the generated results. GRPO is a reinforcement learning algorithm that incorporates a verifiable reward mechanism. By continuously iterating and updating the third text processing model, it achieves higher overall rewards across different dimensions, gradually improving its overall performance. Based on the above reward signals, the third text processing model is continuously updated and optimized. Through continuous training and adjustments, the third text processing model gradually approaches or even surpasses the performance of strong models while maintaining its lightweight characteristics.
[0193] In practical applications, the text processing model trained above can be used to perform text analysis or question response based on different tasks. For example, for a complete, concluded dialogue, the first text processing model can be used to directly obtain the analysis results. Alternatively, the second text processing model can be used to obtain the analysis results and the thought process of the second model during the analysis. For an ongoing dialogue, the third text processing model can be used to analyze the ongoing dialogue and output the analysis results, thus achieving in-process analysis of the current dialogue.
[0194] In summary, the above method, after obtaining the sample text, generates sample questions for text analysis based on the text analysis elements corresponding to the sample text, and generates sample answers to these sample questions as the text analysis results of the sample text. The text processing model is then trained based on the sample text, sample questions, and sample answers, enabling the model to learn text analysis capabilities for element analysis of the sample text from the sample questions and answers. This allows the trained text processing model to accurately understand the meaning of the sample text, further improving its text analysis capabilities and ensuring the accuracy of the text analysis results.
[0195] The following is in conjunction with the appendix Figure 7 Taking the text processing model training method provided in this specification as an example in model training, the text processing model training method will be further explained. Among them, Figure 7 The flowchart of a text processing model training method provided in one embodiment of this specification is shown, which specifically includes the following steps.
[0196] Step 702: Data Engineering.
[0197] Specifically, sample texts (i.e. dialogues) can be obtained from multiple domains, and sample questions and answers for the sample text can be generated based on a large model. After quality evaluation and combination of multiple sample questions and answers for the same sample question, the first training data is obtained.
[0198] Step 704: Enhance and fine-tune the first text processing model.
[0199] Specifically, based on the first training data, the first text processing model can be sequentially subjected to supervised fine-tuning, offline reinforcement learning, and online reinforcement learning to obtain the trained first text processing model.
[0200] Step 706: Design a pre-defined mindset.
[0201] Specifically, multiple preset thinking modes can be designed, each containing different thinking elements and reasoning steps.
[0202] Step 708: Perform adaptive long-chain inference reinforcement learning on the second text processing model.
[0203] Specifically, the second training data obtained by expanding the first training data based on multiple preset thinking modes can be used to perform adaptive long-chain inference reinforcement learning on the second text processing model to obtain the trained second text processing model.
[0204] Step 710: Distill the third text processing model.
[0205] Specifically, the third text processing model can be trained by distillation using the first and / or second text processing models as teacher models to obtain the trained third text processing model.
[0206] In summary, based on the text element system, an automated data construction chain was built, accumulating training data with high coverage and deep relevance to project problems. Through offline / online reinforcement learning, the first text processing model fully learns domain knowledge / specific capabilities, achieving comparable / better performance with a smaller model size. Based on the Socratic thinking system, multiple preset thinking modes were designed, and an adaptive long-chain reasoning method was proposed to expand the capability boundary of the second text processing model in handling difficult examples. Based on the strong model distillation system, a "large-to-small" approach was adopted, utilizing offline / online distillation to enable the smaller model to mimic the distribution of the strong model, allowing the third text processing model to meet latency requirements.
[0207] Corresponding to the above method embodiments, this specification also provides embodiments of a text processing model training device. Figure 8 A schematic diagram of a text processing model training device according to one embodiment of this specification is shown. Figure 8 As shown, the device includes:
[0208] The acquisition module 802 is configured to acquire sample text, a sample question corresponding to the sample text, and a sample answer, wherein the sample question and the sample answer are generated based on the text analysis elements corresponding to the sample text, the sample question is used to perform text analysis on the sample text, the sample answer is the text analysis result of the sample text, and the text analysis elements are used to represent the text content of the sample text;
[0209] The training module 804 is configured to train the text processing model based on the sample text, the sample question, and the sample answer until a fully trained text processing model is obtained.
[0210] In an optional embodiment, the apparatus further includes a generation module configured to:
[0211] Obtain sample text;
[0212] Based on the text analysis elements corresponding to the sample text, generate the sample question and the sample answer for the sample text.
[0213] In an optional embodiment, the generation module is further configured to:
[0214] The sample text is input into a text generation model, and the text generation model is used to extract the text analysis elements corresponding to the sample text.
[0215] Using the text generation model, a sample question and a sample answer are generated for the sample text based on the text analysis elements.
[0216] In an optional embodiment, the generation module is further configured to:
[0217] Using the text generation model, based on a pre-built text element system, the text analysis elements corresponding to the sample text are extracted;
[0218] The text element system is constructed based on the target task and contains reference text analysis elements in multiple dimensions.
[0219] In one optional embodiment, the text processing model includes a first text processing model and a second text processing model;
[0220] The training module 804 is further configured as follows:
[0221] Construct the first training data based on the sample text, the sample question, and the sample answer;
[0222] Based on the first training data, the first text processing model is trained until the first text processing model is obtained after training.
[0223] Based on the sample text, the sample question, and the sample answer, construct second training data containing the thought process;
[0224] The second text processing model is trained based on the second training data until a fully trained second text processing model is obtained.
[0225] In an optional embodiment, the generation module is further configured to:
[0226] Based on the text analysis elements corresponding to the sample text, generate multiple sample questions for the sample text, as well as sample answers for each sample question;
[0227] The training module 804 is further configured as follows:
[0228] Based on the sample text, a quality assessment is performed on the sample question and the sample answer to obtain the quality assessment result;
[0229] Based on the quality assessment results, target sample questions and target sample answers are determined from the plurality of sample questions and sample answers;
[0230] The sample text, the target sample question, and the target sample answer are used as the first training data.
[0231] In an optional embodiment, the training module 804 is further configured to:
[0232] At least two sample problems are selected from the plurality of sample problems and combined to obtain combined sample problems;
[0233] From the sample answers corresponding to each sample question, determine the combined sample answer corresponding to the combined sample question;
[0234] The sample text, the combined sample question, and the combined sample answer are used as the first training data.
[0235] In an optional embodiment, the training module 804 is further configured to:
[0236] Based on the first training data, perform online reinforcement learning and / or offline reinforcement learning on the first text processing model until a fully trained first text processing model is obtained.
[0237] In an optional embodiment, the training module 804 is further configured to:
[0238] The sample answers were expanded to obtain sample answers that included the thought process;
[0239] The sample text, the sample question, and the sample answer containing the thought process are used as the second training data.
[0240] In an optional embodiment, the training module 804 is further configured to:
[0241] Based on the preset thinking pattern, construct the reasoning link corresponding to the preset thinking pattern;
[0242] The sample answer is expanded based on the reasoning link to obtain a sample answer that includes the reasoning link, which is then used as the sample answer that includes the thinking process.
[0243] In an optional embodiment, the training module 804 is further configured to:
[0244] Based on a preset amount of second training data, the second text processing model is cold-started to obtain the cold-started second text processing model.
[0245] The sample question is input into the second text processing model after cold start training to obtain the predicted answer corresponding to each preset thinking mode.
[0246] Based on reinforcement learning algorithms, the sample-level dominance value corresponding to each predicted answer and the pattern-level dominance value corresponding to each preset thinking pattern are calculated.
[0247] The second text processing model after cold start training is trained based on the sample-level dominance value and the pattern-level dominance value until a fully trained second text processing model is obtained.
[0248] In an optional embodiment, the text processing model further includes a third text processing model;
[0249] The training module 804 is further configured as follows:
[0250] Based on the first text processing model and / or the second text processing model, the third text processing model is distilled and trained until a fully trained third text processing model is obtained.
[0251] The model size of the third text processing model is smaller than that of the first text processing model and the second text processing model.
[0252] In summary, in the above-described device, after acquiring the sample text, a sample question for text analysis of the sample text can be generated based on the text analysis elements corresponding to the sample text. A sample answer corresponding to the sample question is then generated as the text analysis result of the sample text. The text processing model is trained based on the sample text, sample question, and sample answer, enabling the model to learn the text analysis capabilities for element analysis of the sample text from the sample question and sample answer. This allows the trained text processing model to accurately understand the meaning of the sample text, further improving its text analysis capabilities and ensuring the accuracy of the text analysis results.
[0253] The above is an illustrative scheme of a text processing model training device according to this embodiment. It should be noted that the technical solution of this text processing model training device and the technical solution of the text processing model training method described above belong to the same concept. For details not described in detail in the technical solution of the text processing model training device, please refer to the description of the technical solution of the text processing model training method described above.
[0254] See Figure 9 , Figure 9 A flowchart of a text processing method according to an embodiment of this specification is shown, which specifically includes the following steps.
[0255] Step 902: Determine the text to be processed;
[0256] Step 904: Input the text to be processed into the text processing model to obtain the text analysis result output by the text processing model, wherein the text processing model is trained according to the text processing model training method provided in the embodiments of this specification.
[0257] The text processing model can include a first text processing model, a second text processing model, and a third text processing model.
[0258] Specifically, depending on the task, a first, second, or third text processing model can be selected for text analysis or question response. For example, for a complete, concluded dialogue, the first text processing model can be used to directly obtain the analysis results. Alternatively, the second text processing model can be used to obtain the analysis results and the thought process of the second model during the analysis. For an ongoing dialogue, the third text processing model can be used to analyze the ongoing dialogue and output the analysis results, thus achieving in-process analysis of the current dialogue.
[0259] The above is an illustrative scheme of a text processing method according to this embodiment. It should be noted that the technical solution of this text processing method belongs to the same concept as the technical solution of the text processing model training method described above. For details not described in detail in the technical solution of the text processing method, please refer to the description of the technical solution of the text processing model training method described above.
[0260] Corresponding to the above method embodiments, this specification also provides embodiments of a text processing device. Figure 10 A schematic diagram of the structure of a text processing apparatus according to one embodiment of this specification is shown. Figure 10 As shown, the device includes:
[0261] Module 1002 is configured to determine the text to be processed.
[0262] The input module 1004 is configured to input the text to be processed into the text processing model and obtain the text analysis result output by the text processing model, wherein the text processing model is trained according to the text processing model training method provided in the embodiments of this specification.
[0263] The above is an illustrative scheme of a text processing device according to this embodiment. It should be noted that the technical solution of this text processing device and the technical solution of the text processing model training method described above belong to the same concept. For details not described in detail in the technical solution of the text processing device, please refer to the description of the technical solution of the text processing model training method described above.
[0264] See Figure 11 , Figure 11 A flowchart of another text processing model training method according to an embodiment of this specification is shown, applied to a cloud-side device, and specifically includes the following steps.
[0265] Step 1102: Receive the model training request sent by the receiving end device, and obtain the sample text, sample question and sample answer corresponding to the model training request, wherein the sample question and sample answer are generated based on the text analysis elements corresponding to the sample text, the sample question is used to perform text analysis on the sample text, the sample answer is the text analysis result of the sample text, and the text analysis elements are used to represent the text content of the sample text;
[0266] Step 1104: Train the text processing model based on the sample text, the sample question, and the sample answer until a fully trained text processing model is obtained, and send the model parameters of the text processing model to the edge device.
[0267] The above is an illustrative scheme of a text processing model training method according to this embodiment. It should be noted that the technical solution of this text processing model training method belongs to the same concept as the technical solution of the text processing model training method described above. For details not described in detail in the technical solution of the text processing model training method, please refer to the description of the technical solution of the text processing model training method described above.
[0268] Corresponding to the above method embodiments, this specification also provides embodiments of a text processing model training device applied to cloud-based devices. Figure 12 A schematic diagram of another text processing model training apparatus provided in one embodiment of this specification is shown. Figure 12 As shown, the device includes:
[0269] The receiving module 1202 is configured to receive a model training request sent by a terminal device, and obtain the sample text corresponding to the model training request, the sample question corresponding to the sample text, and the sample answer. The sample question and the sample answer are generated based on the text analysis elements corresponding to the sample text. The sample question is used to perform text analysis on the sample text, and the sample answer is the text analysis result of the sample text. The text analysis elements are used to represent the text content of the sample text.
[0270] The training module 1204 is configured to train the text processing model based on the sample text, the sample question, and the sample answer until a trained text processing model is obtained, and to send the model parameters of the text processing model to the edge device.
[0271] The above is an illustrative scheme of a text processing model training device according to this embodiment. It should be noted that the technical solution of this text processing model training device and the technical solution of the text processing model training method described above belong to the same concept. For details not described in detail in the technical solution of the text processing model training device, please refer to the description of the technical solution of the text processing model training method described above.
[0272] See Figure 13 , Figure 13 A flowchart of a dialogue processing method according to an embodiment of this specification is shown, which specifically includes the following steps.
[0273] Step 1302: Identify the dialogue to be processed;
[0274] Step 1304: Input the dialogue to be processed into the text processing model to obtain the dialogue analysis result output by the text processing model, wherein the text processing model is trained according to the text processing model training method provided in the embodiments of this specification.
[0275] Specifically, the dialogue to be processed can be understood as the text to be processed, and the dialogue analysis results can be understood as the text analysis results.
[0276] The above is an illustrative scheme of a dialogue processing method according to this embodiment. It should be noted that the technical solution of this dialogue processing method and the technical solution of the text processing model training method described above belong to the same concept. For details not described in detail in the technical solution of the dialogue processing method, please refer to the description of the technical solution of the text processing model training method described above.
[0277] Corresponding to the above method embodiments, this specification also provides embodiments of a dialogue processing device. Figure 14 A schematic diagram of a dialogue processing apparatus according to one embodiment of this specification is shown. Figure 14As shown, the device includes:
[0278] Module 1402 is configured to identify dialogues to be processed.
[0279] The input module 1404 is configured to input the dialogue to be processed into the text processing model and obtain the dialogue analysis result output by the text processing model, wherein the text processing model is trained according to the text processing model training method provided in the embodiments of this specification.
[0280] The above is an illustrative scheme of a dialogue processing device according to this embodiment. It should be noted that the technical solution of this dialogue processing device and the technical solution of the above-described text processing model training method belong to the same concept. For details not described in detail in the technical solution of the dialogue processing device, please refer to the description of the technical solution of the above-described text processing model training method.
[0281] See Figure 15 , Figure 15 A flowchart of another text processing method according to an embodiment of this specification is shown, applied to a cloud-side device, and specifically includes the following steps.
[0282] Step 1502: Receive a text processing task sent by the receiving end device, wherein the text processing task carries text to be processed;
[0283] Step 1504: Input the text to be processed into the text processing model, obtain the text analysis result output by the text processing model, and send the text analysis result to the terminal device, wherein the text processing model is trained according to the text processing model training method provided in the embodiments of this specification.
[0284] The above is an illustrative scheme of a text processing method according to this embodiment. It should be noted that the technical solution of this text processing method belongs to the same concept as the technical solution of the text processing model training method described above. For details not described in detail in the technical solution of the text processing method, please refer to the description of the technical solution of the text processing model training method described above.
[0285] Corresponding to the above method embodiments, this specification also provides embodiments of a text processing device applied to cloud-side equipment. Figure 16 A schematic diagram of another text processing apparatus provided in one embodiment of this specification is shown. Figure 16 As shown, the device includes:
[0286] The receiving module 1602 is configured to receive a text processing task sent by the end-side device, wherein the text processing task carries text to be processed.
[0287] The sending module 1604 is configured to input the text to be processed into the text processing model, obtain the text analysis result output by the text processing model, and send the text analysis result to the end device, wherein the text processing model is trained according to the text processing model training method provided in the embodiments of this specification.
[0288] The above is an illustrative scheme of a text processing device according to this embodiment. It should be noted that the technical solution of this text processing device and the technical solution of the text processing model training method described above belong to the same concept. For details not described in detail in the technical solution of the text processing device, please refer to the description of the technical solution of the text processing model training method described above.
[0289] Corresponding to the above method embodiments, this specification also provides a training data generation method, including:
[0290] Obtain sample text;
[0291] Based on the text analysis elements corresponding to the sample text, a sample question and a sample answer are generated for the sample text. The sample question is used to perform text analysis on the sample text, and the sample answer is the text analysis result of the sample text. The text analysis elements are used to represent the text content of the sample text.
[0292] Training data is constructed based on the sample text, the sample question, and the sample answer.
[0293] Specifically, when constructing training data, first training data can be constructed based on sample text, sample questions, and sample answers, and second training data containing the thought process can be constructed based on sample text, sample questions, and sample answers. The specific construction process is similar to that described above, and will not be repeated in the embodiments of this specification. After constructing the training data, the above text processing model can be trained based on the constructed training data. The specific training process is similar to that described above, and will not be limited in the embodiments of this specification.
[0294] Figure 17 A structural block diagram of a computing device 1700 according to one embodiment of this specification is shown. The components of the computing device 1700 include, but are not limited to, a memory 1710 and a processor 1720. The processor 1720 is connected to the memory 1710 via a bus 1730, and a database 1750 is used to store data.
[0295] The computing device 1700 also includes an access device 1740, which enables the computing device 1700 to communicate via one or more networks 1760. Examples of these networks include a Public Switched Telephone Network (PSTN), a Local Area Network (LAN), a Wide Area Network (WAN), a Personal Area Network (PAN), or a combination of communication networks such as the Internet. The access device 1740 may include one or more of any type of wired or wireless network interface (e.g., a Network Interface Card (NIC)), such as an IEEE 802.11 Wireless Local Area Network (WLAN) interface, a Wi-MAX interface, an Ethernet interface, a Universal Serial Bus (USB) interface, a cellular network interface, a Bluetooth interface, a Near Field Communication (NFC) interface, and so on.
[0296] In one embodiment of this specification, the above-described components of the computing device 1700 and Figure 17 Other components, not shown, can also be connected to each other, for example, via a bus. It should be understood that... Figure 17 The block diagram of the computing device shown is for illustrative purposes only and is not intended to limit the scope of this specification. Those skilled in the art can add or replace other components as needed.
[0297] The computing device 1700 can be any type of stationary or mobile computing device, including mobile computers or mobile computing devices (e.g., tablet computers, personal digital assistants, laptop computers, notebook computers, netbooks, etc.), mobile phones (e.g., smartphones), wearable computing devices (e.g., smartwatches, smart glasses, etc.) or other types of mobile devices, or stationary computing devices such as desktop computers or PCs. The computing device 1700 can also be a mobile or stationary server.
[0298] The processor 1720 is used to execute the following computer program / instructions, which, when executed by the processor, implement the steps of the above method.
[0299] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the computing device embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions of the method embodiments.
[0300] An embodiment of this specification also provides a computer-readable storage medium storing a computer program / instructions that, when executed by a processor, implement the steps of the above-described method.
[0301] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the computer-readable storage medium embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions of the method embodiments.
[0302] An embodiment of this specification also provides a computer program product, including a computer program / instructions that, when executed by a processor, implement the steps of the above-described method.
[0303] The above is an illustrative scheme of a computer program product according to this embodiment. It should be noted that the technical solution of this computer program product and the technical solution of the above method belong to the same concept, and all details not described in detail in the technical solution of the computer program product can be referred to the description of the technical solution of the above method.
[0304] The foregoing has described specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired result. In some embodiments, multitasking and parallel processing are possible or may be advantageous.
[0305] The computer instructions include computer program code, which may be in the form of source code, object code, executable file, or certain intermediate forms. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording media, USB flash drive, portable hard drive, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content included in the computer-readable medium may be appropriately added or removed according to the requirements of patent practice. For example, in some regions, according to patent practice, computer-readable media may not include electrical carrier signals and telecommunication signals.
[0306] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that the embodiments in this specification are not limited to the described order of actions, because according to the embodiments in this specification, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in this specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the embodiments in this specification.
[0307] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0308] The preferred embodiments disclosed above are merely illustrative of this specification. The optional embodiments do not exhaustively describe all details, nor do they limit the invention to the specific implementations described. Clearly, many modifications and variations can be made based on the embodiments described herein. These embodiments are selected and specifically described in this specification to better explain the principles and practical applications of the embodiments, thereby enabling those skilled in the art to better understand and utilize this specification. This specification is limited only by the claims and their full scope and equivalents.
Claims
1. A text processing model training method, comprising: obtaining a sample text, a sample question corresponding to the sample text, and a sample answer, wherein the sample question and the sample answer are generated according to a text analysis element corresponding to the sample text, the sample question is used for text analysis of the sample text, the sample answer is a text analysis result of the sample text, and the text analysis element is used for representing text content of the sample text; training a text processing model according to the sample text, the sample question, and the sample answer until a trained text processing model is obtained, wherein the text processing model comprises a second text processing model, the second text processing model is used for outputting an answer and a thinking process, the second text processing model is trained according to a sample level advantage value and a mode level advantage value, the sample level advantage value corresponds to a predicted answer output by the second text processing model, and the mode level advantage value corresponds to a preset thinking mode, and the predicted answer is output by the second text processing model based on the preset thinking mode and the sample question.
2. The method of claim 1, wherein the obtaining a sample text, a sample question corresponding to the sample text, and a sample answer comprises: obtaining a sample text; generating the sample question and the sample answer for the sample text according to a text analysis element corresponding to the sample text.
3. The method of claim 2, wherein before the generating the sample question and the sample answer for the sample text according to the text analysis element corresponding to the sample text, the method further comprises: inputting the sample text into a text generation model, and extracting the text analysis element corresponding to the sample text by using the text generation model; the generating the sample question and the sample answer for the sample text according to the text analysis element corresponding to the sample text further comprises: generating the sample question and the sample answer for the sample text according to the text analysis element by using the text generation model.
4. The method of claim 3, wherein the extracting the text analysis element corresponding to the sample text by using the text generation model comprises: extracting the text analysis element corresponding to the sample text based on a pre-constructed text element system by using the text generation model; wherein the text element system is constructed according to a target task, and the text element system contains reference text analysis elements of multiple dimensions.
5. The method of claim 1, wherein the text processing model comprises a first text processing model and a second text processing model, the first text processing model is used for outputting an answer, and the second text processing model is used for outputting an answer and a thinking process; the training the text processing model according to the sample text, the sample question, and the sample answer until the trained text processing model is obtained comprises: constructing first training data according to the sample text, the sample question, and the sample answer. training, according to the first training data, a first text processing model until a trained first text processing model is obtained; constructing, according to the sample text, the sample question and the sample answer, second training data containing a thinking process; training, according to the second training data, a second text processing model until a trained second text processing model is obtained.
6. The method of claim 5, wherein generating, according to the text analysis element corresponding to the sample text, a sample question and a sample answer for the sample text comprises: generating, according to the text analysis element corresponding to the sample text, a plurality of sample questions for the sample text and a sample answer corresponding to each sample question; constructing, according to the sample text, the sample question and the sample answer, first training data comprises: performing quality evaluation on the sample question and the sample answer according to the sample text to obtain a quality evaluation result; determining a target sample question and a target sample answer from the plurality of sample questions and sample answers according to the quality evaluation result; taking the sample text, the target sample question and the target sample answer as the first training data.
7. The method of claim 6, wherein constructing, according to the sample text, the sample question and the sample answer, first training data further comprises: selecting at least two sample questions from the plurality of sample questions to combine to obtain a combined sample question; determining a combined sample answer corresponding to the combined sample question from the sample answers corresponding to each sample question; taking the sample text, the combined sample question and the combined sample answer as the first training data.
8. The method of claim 5, wherein training, according to the first training data, a first text processing model until a trained first text processing model is obtained comprises: performing online reinforcement learning and / or offline reinforcement learning on the first text processing model according to the first training data until a trained first text processing model is obtained.
9. The method of claim 5, wherein constructing, according to the sample text, the sample question and the sample answer, second training data containing a thinking process comprises: performing expansion on the sample answer to obtain a sample answer containing a thinking process; taking the sample text, the sample question and the sample answer containing a thinking process as the second training data.
10. The method of claim 9, wherein performing expansion on the sample answer to obtain a sample answer containing a thinking process comprises: constructing an inference link corresponding to a preset thinking mode according to the preset thinking mode; performing expansion on the sample answer based on the inference link to obtain a sample answer containing the inference link as the sample answer containing a thinking process.
11. The method of claim 10, wherein training, according to the second training data, a second text processing model until a trained second text processing model is obtained comprises: training the second text processing model according to the second preset number of training data, to obtain a cold start trained second text processing model; inputting the sample question into the cold start trained second text processing model, to obtain a predicted answer corresponding to each preset thinking mode; calculating a sample level advantage value corresponding to each predicted answer and a mode level advantage value corresponding to each preset thinking mode based on a reinforcement learning algorithm; training the cold start trained second text processing model according to the sample level advantage value and the mode level advantage value, until a trained second text processing model is obtained.
12. The method of claim 5, wherein the text processing model further comprises a third text processing model; the method further comprises: training the third text processing model according to the first text processing model and / or the second text processing model, until a trained third text processing model is obtained; wherein the model size of the third text processing model is smaller than the model size of the first text processing model and the second text processing model.
13. A text processing method, comprising: determining a text to be processed; inputting the text to be processed into the text processing model, to obtain a text analysis result output by the text processing model, wherein the text processing model is trained according to the method of any one of claims 1-12.
14. A dialogue processing method, comprising: determining a dialogue to be processed; inputting the dialogue to be processed into the text processing model, to obtain a dialogue analysis result output by the text processing model, wherein the text processing model is trained according to the method of any one of claims 1-12.
15. A text processing model training method applied to a cloud side device, comprising: receiving a model training request sent by an end side device, and obtaining a sample text corresponding to the model training request, a sample question corresponding to the sample text, and a sample answer, wherein the sample question and the sample answer are generated according to a text analysis element corresponding to the sample text, the sample question is used for text analysis of the sample text, the sample answer is a text analysis result of the sample text, and the text analysis element is used for representing text content of the sample text; training a text processing model according to the sample text, the sample question, and the sample answer, until a trained text processing model is obtained, and sending model parameters of the text processing model to the end side device, wherein the text processing model comprises a second text processing model, the second text processing model is used for outputting an answer and a thinking process, the second text processing model is trained according to a sample level advantage value and a mode level advantage value, the sample level advantage value corresponds to a predicted answer output by the second text processing model, and the mode level advantage value corresponds to a preset thinking mode, the predicted answer is output by the second text processing model based on the preset thinking mode and the sample question.
16. A text processing method applied to a cloud side device, comprising: Receiving a text processing task sent by an end-side device, wherein the text processing task carries a text to be processed; Inputting the text to be processed into the text processing model, obtaining a text analysis result output by the text processing model, and sending the text analysis result to the end-side device, wherein the text processing model is trained according to the method of any one of claims 1-12.
17. A training data generation method, comprising: Obtaining a sample text; generating a sample question and a sample answer for the sample text according to a text analysis element corresponding to the sample text, wherein, The sample question is used for text analysis on the sample text, the sample answer is a text analysis result of the sample text, and the text analysis element is used for representing the text content of the sample text; According to the sample text, the sample question and the sample answer, constructing training data, wherein the training data is used for training a text processing model, the text processing model comprises a second text processing model, the second text processing model is used for outputting an answer and a thinking process, the second text processing model is trained according to a sample level advantage value and a mode level advantage value, the sample level advantage value corresponds to a predicted answer output by the second text processing model, the mode level advantage value corresponds to a preset thinking mode, and the predicted answer is output by the second text processing model based on the preset thinking mode for predicting the sample question.
18. A computing device, comprising: A memory and a processor; The memory is used for storing computer programs / instructions, and the processor is used for executing the computer programs / instructions, which realize the steps of the method of any one of claims 1-17 when executed by the processor.
19. A computer readable storage medium storing computer programs / instructions, which realize the steps of the method of any one of claims 1-17 when executed by the processor.
20. A computer program product comprising computer programs / instructions, which realize the steps of the method of any one of claims 1-17 when executed by the processor.
Citation Information
Patent Citations
Intelligent decision-making method and device based on offline-online hybrid reinforcement learning
CN117648548A
Text processing model training method, text processing method and question and answer processing method and device
CN118627543A
Text processing model training method and device and storage medium
CN120012942A