Question processing method, model training method and related devices
By supervising and fine-tuning the pre-trained large language model and iteratively training the target rewriting model, the problem of large rewriting deviation in traditional rewriting methods is solved, and a more accurate rewriting effect is achieved.
Patent Information
- Application Number
- CN202410101181.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-24
- Publication Date
- 2025-07-29
AI Technical Summary
Traditional problem rewriting methods are easy to expand additional information, resulting in deviations in rewriting problems and unable to obtain better rewriting results.
By supervising fine-tuning the pre-trained large language model, iterative training obtains the target rewrite model, and the target rewrite model is used to determine the characters in the problem to be rewrite and update it, and the problem rewriting process is completed.
Effectively reduce rewriting deviation, improve the rewriting effect of problem, and ensure that the rewriting results are more accurate.
Smart Images

Figure CN120386835A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of artificial intelligence technology, and particularly to a method for problem processing, a method for model training, and related devices. Background Art
[0002] In the process of daily conversation, problems such as mutual reference and citation, or other information omissions often occur. The user can correctly understand the intention of the speaker, but devices such as robots cannot easily interpret the intention of the conversation. For example, in the early-stage consultation, sales, and later underwriting stages of insurance, robots are often required to automatically handle the consultation questions of insurance customers. At this time, correctly understanding the intention of the speaker is particularly important for robots. To enable devices such as robots to correctly understand the true intention of the speaker, it can be achieved by rewriting the question raised by the speaker.
[0003] However, in the traditional way of question rewriting, a non-large model is usually used to rewrite the question. However, this method is prone to expanding additional information, resulting in deviation in the obtained rewritten question and unable to obtain a better rewriting effect. Summary of the Invention
[0004] The embodiments of the present application provide a method for problem processing, a method for model training, and related devices, which can complete question rewriting without expanding additional information, effectively reduce the rewriting deviation, and improve the question rewriting effect.
[0005] In a first aspect, the embodiments of the present application provide a method for problem processing. The method includes: obtaining a question to be rewritten and the target conversation content corresponding to the question to be rewritten; inputting the question to be rewritten and the target conversation content into a target rewriting model to determine the first character in the question to be rewritten, where the first character is the character used for rewriting in the question to be rewritten, and the target rewriting model is a machine learning model obtained by iteratively training an initial rewriting model with sample questions, sample conversation content corresponding to the sample questions, and sample rewritten questions as training data, and the sample rewritten questions are obtained by the initial rewriting model performing question rewriting processing on the sample conversation content and the sample questions, and the initial rewriting model is a model obtained by performing supervised fine-tuning on a pre-trained large language model; inputting the characters in the target conversation content into the target rewriting model to update the first character at the corresponding position in the question to be rewritten, and obtaining the target question corresponding to the question to be rewritten.
[0006] Second aspect, an embodiment of the present application provides a method for model training. The method includes: obtaining training data, where the training data includes sample questions and sample conversation contents corresponding to the sample questions; inputting the sample conversation contents and the sample questions into an initial rewriting model to determine a second character in the sample question, where the second character is the character used for rewriting in the sample question, and the initial rewriting model is a model obtained by performing supervised fine-tuning on a pre-trained large language model; inputting the characters in the sample conversation contents into the initial rewriting model to update the second character at the corresponding position in the sample question, obtaining a sample rewritten question corresponding to the sample question; based on the sample conversation contents, the sample question, and the sample rewritten question, iteratively training the initial rewriting model to obtain a target rewriting model, where the target rewriting model is used to perform question rewriting processing on a question to be rewritten and a target conversation content corresponding to the content to be rewritten, obtaining a target question corresponding to the question to be rewritten.
[0007] Third aspect, an embodiment of the present application provides a problem processing device. The problem processing device includes an acquisition unit and a processing unit. Among them, the acquisition unit is used to obtain a question to be rewritten and a target conversation content corresponding to the question to be rewritten. The processing unit inputs the question to be rewritten and the target conversation content into the target rewriting model to determine a first character in the question to be rewritten, where the first character is the character used for rewriting in the question to be rewritten, and the target rewriting model is a machine learning model obtained by iteratively training the initial rewriting model using the sample question, the sample conversation content corresponding to the sample question, and the sample rewritten question as training data, and the sample rewritten question is obtained by the initial rewriting model performing question rewriting processing on the sample conversation content and the sample question, and the initial rewriting model is a model obtained by performing supervised fine-tuning on a pre-trained large language model. The processing unit is used to input the characters in the target conversation content into the target rewriting model to update the first character at the corresponding position in the question to be rewritten, obtaining a target question corresponding to the question to be rewritten.
[0008] In some optional implementation manners, the first character includes one or more of a character capable of anaphora resolution, an omitted character, a character for typo correction, and a replacement character corresponding to the character capable of anaphora resolution.
[0009] In some other alternative embodiments, the processing unit is configured to: calculate the first part-of-speech similarity between each character in the problem to be rewritten and the corresponding character in the target dialogue content based on the target rewriting model; when the first part-of-speech similarity is greater than a preset reference resolution judgment threshold, determine the corresponding character in the target dialogue content as the character that can be resolved by reference in the problem to be rewritten, or the replacement character corresponding to the character that can be resolved by reference; or, when the first part-of-speech similarity is greater than a preset ellipsis judgment threshold, determine the corresponding character in the target dialogue content as the omitted character in the problem to be rewritten; or, when the first part-of-speech similarity is greater than a preset typo correction threshold, determine the corresponding character in the target dialogue content as the typo correction character in the problem to be rewritten.
[0010] In some other alternative embodiments, the processing unit is further configured to, after inputting the characters in the target dialogue content into the target rewriting model to update the first characters at the corresponding positions in the problem to be rewritten to obtain the target problem corresponding to the problem to be rewritten, extract the characters to be replaced in the target problem, where the characters to be replaced include one or more of abbreviated words and colloquial words; match the characters to be replaced with each proper noun in the preset proper noun mapping table to obtain the target proper noun of the characters to be replaced; and update the characters to be replaced with the target proper noun to obtain the updated target problem.
[0011] In some other alternative embodiments, the processing unit is further configured to: after inputting the characters in the target dialogue content into the target rewriting model to update the first characters at the corresponding positions in the problem to be rewritten to obtain the target problem corresponding to the problem to be rewritten, perform retrieval and recall on the problem to be rewritten and the target problem respectively to obtain the retrieval answer of the problem to be rewritten and the retrieval answer of the target problem; perform evaluation processing on the retrieval answer of the problem to be rewritten and the retrieval answer of the target problem respectively based on a preset evaluation index to obtain a first evaluation result and a second evaluation result, where the first evaluation result is used to indicate the recall situation of the retrieval answer of the problem to be rewritten, and the second evaluation result is used to indicate the recall situation of the retrieval answer of the target problem; when the second evaluation result is greater than the first evaluation result, splice the retrieval answer of the target problem and the target problem, and fine-tune the target rewriting model based on the spliced data.
[0012] Fourthly, an embodiment of the present application provides a model training device. The model training device includes an acquisition module and a processing module. Among them, the acquisition module is used to acquire training data, and the training data includes sample questions and sample dialogue contents corresponding to the sample questions. The processing module is used to input the sample dialogue content and the sample question into the initial rewriting model to determine the second character in the sample question. The second character is the character used for rewriting in the sample question, and the initial rewriting model is a model obtained by performing supervised fine-tuning on a pre-trained large language model. The processing module is used to input the characters in the sample dialogue content into the initial rewriting model to update the second character at the corresponding position in the sample question, and obtain a sample rewritten question corresponding to the sample question. The processing module is used to iteratively train the initial rewriting model based on the sample dialogue content, the sample question, and the sample rewritten question to obtain a target rewriting model. The target rewriting model is used to perform question rewriting processing on the question to be rewritten and the target dialogue content corresponding to the question to be rewritten, and obtain the target question corresponding to the question to be rewritten.
[0013] In some alternative embodiments, the second character includes one or more of a character capable of anaphora resolution, an omitted character, a character for correcting typos, and a replacement character corresponding to the character capable of anaphora resolution.
[0014] In some other alternative embodiments, the processing module is configured to: calculate the second part-of-speech similarity between each character in the sample question and the corresponding character in the sample dialogue content based on the initial rewriting model; when the second part-of-speech similarity is greater than a preset anaphora resolution judgment threshold, determine the corresponding character in the sample dialogue content as the character capable of anaphora resolution in the sample question, or the replacement character corresponding to the character capable of anaphora resolution; or, when the second part-of-speech similarity is greater than a preset omission judgment threshold, determine the corresponding character in the sample dialogue content as the omitted character in the sample question; or, when the second part-of-speech similarity is greater than a preset typo correction threshold, determine the corresponding character in the sample dialogue content as the character for correcting typos in the sample question.
[0015] In some alternative embodiments, the processing module is configured to: perform retrieval and recall on the sample rewritten question to obtain a retrieval answer for the sample rewritten question; calculate the difference between the retrieval answer and the answer label corresponding to the sample question to obtain a loss value; update the model parameters of the initial rewriting model based on the loss value to obtain the target rewriting model.
[0016] A fifth aspect of the embodiments of the present application provides a problem processing device, including: a memory, an input / output interface, and a memory. The memory is used to store program instructions. The processor is used to execute the program instructions in the memory to execute the problem processing method corresponding to the embodiment of the first aspect above; or, execute the model training method corresponding to the embodiment of the second aspect above.
[0017] In a sixth aspect of the embodiments of the present application, a computer-readable storage medium is provided. Instructions are stored in the computer-readable storage medium. When it runs on a computer, it causes the computer to execute the method for problem handling corresponding to the implementation manner of the first aspect above; or, execute the method for model training corresponding to the implementation manner of the second aspect above.
[0018] In a seventh aspect of the embodiments of the present application, a computer program product containing instructions is provided. When it runs on a computer or a processor, it causes the computer or the processor to execute the method for problem handling corresponding to the implementation manner of the first aspect above; or, execute the method for model training corresponding to the implementation manner of the second aspect above.
[0019] As can be seen from the above technical solutions, the embodiments of the present application have the following advantages:
[0020] In the embodiments of the present application, by supervised fine-tuning of the pre-trained large language model, and then using the obtained initial rewriting model to complete the problem rewriting of the sample dialogue content in the training data and the sample questions proposed for the sample dialogue content, and then obtaining the corresponding sample rewritten questions. Then, through the sample dialogue content, sample questions, and sample rewritten questions, the iterative training of the initial rewriting model is completed, and thus the target rewriting model is trained. In this way, after obtaining the problem to be rewritten and the target dialogue content corresponding to the problem to be rewritten, the first character used for rewriting in the problem to be rewritten can be determined with the help of the target rewriting model to complete the problem rewriting process of the problem to be rewritten and the target dialogue content, and then the target rewriting model is used to update the first character at the corresponding position in the problem to be rewritten based on the characters in the target dialogue content, so as to obtain the target problem corresponding to the problem to be rewritten. That is to say, in the present application, through efficient supervised fine-tuning of the pre-trained large language model, the obtained initial rewriting model can be capable of the problem rewriting task, and then the iterative training of the initial rewriting model is completed on the training data, so that the target rewriting model can be effectively applied to the downstream problem rewriting task, with a clearer intention expression, without the need to expand additional information to complete the problem rewriting, but only using the characters in the target dialogue content itself to update the first character at the corresponding position in the problem to be rewritten, effectively reducing the rewriting deviation and improving the problem rewriting effect. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0022] Figure 1 Shows a schematic diagram of an application scenario provided by an embodiment of the present application;
[0023] Figure 2 Shows a schematic diagram of a system framework provided by an embodiment of the present application;
[0024] Figure 3 Shows a schematic flowchart of a method for model training provided by an embodiment of the present application;
[0025] Figure 4 Shows a schematic flowchart of a method for problem handling provided by an embodiment of the present application;
[0026] Figure 5 Shows a schematic diagram of functional modules of a problem handling device provided in an embodiment of the present application;
[0027] Figure 6 Shows a schematic diagram of functional modules of a model training device provided in an embodiment of the present application;
[0028] Figure 7 Shows a schematic diagram of the hardware structure of a problem handling device provided in an embodiment of the present application. Detailed implementation manners
[0029] An embodiment of the present application provides a method for problem handling, a method for model training, and related devices, which can complete problem rewriting without expanding additional information, effectively reduce rewriting deviation, and improve the problem rewriting effect.
[0030] It can be understood that in the specific implementation manners of the present application, data related to user information, etc. is involved. When the above embodiments of the present application are applied to specific products or technologies, user permission or consent needs to be obtained, and the collection, use, and processing of relevant data need to comply with relevant laws, regulations, and standards of relevant countries and regions.
[0031] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0032] In the description and claims of this application and the above-mentioned drawings, terms such as "first", "second", "third", "fourth", etc. (if any) are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of this application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that comprises a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0033] With the research and progress of artificial intelligence (AI) technology, AI technology has been studied and applied in multiple fields, such as common smart homes, smart wearable devices, virtual assistants, smart speakers, driverless, autonomous driving, drones, robots, smart healthcare, smart customer service, etc. It is believed that with the development of technology, AI technology will be applied in more fields and play an increasingly important role. In addition, with the development of computer technology, natural language processing (NLP) technology has also emerged. NLP technology has achieved what people have long pursued, namely "communicating with a computer in natural language". For example, it can achieve human-machine dialogue, intelligent language translation, and dialogue reading comprehension through a trained model, etc.
[0034] During the process of daily conversation, often due to the mutual reference of the previous and subsequent conversation content or the omission of information in the conversation, the robot cannot accurately understand the true intention of the speaker like a user. For example, Figure 1 FIG. shows a schematic diagram of an application scenario provided by an embodiment of this application.
[0035] As Figure 1 shown, the conversation between the user and the robot is as follows: User: "What's the weather like in City A today?"; Robot: "The weather in City A today is sunny turning to cloudy, 24 degrees to 32 degrees, north wind level 3"; User: "Is it suitable to wear a skirt?"
[0036] However, the robot cannot accurately understand the specific intention of the original question "Is it suitable to wear a skirt?", that is, whether an answer needs to be based on the original conversation content. Therefore, it is necessary to rewrite this original question. However, in traditional question rewriting methods, usually a non-large model is used to complete the rewriting of the question, which is likely to cause a large deviation in question rewriting due to the addition of extra information and cannot obtain a better rewriting effect.
[0037] It should be noted that the mentioned problem of paraphrasing hallucination can be understood as the content that seems reasonable after paraphrasing but is actually inaccurate or does not conform to the facts.
[0038] To solve the above-mentioned technical problems, the embodiments of the present application provide a method for model training. Correspondingly, the embodiments of the present application also provide a method for problem processing. Through the model training method provided by the present application, supervised fine-tuning of the pre-trained large language model is realized. After obtaining the initial paraphrasing model, iterative training of the initial paraphrasing model is completed based on the training data, and then the target paraphrasing model is trained. After training the target paraphrasing model, in the method for problem processing, the target paraphrasing model is used to complete the paraphrasing process of the problem to be paraphrased, without the need to expand additional information to assist in completion, effectively reducing the paraphrasing deviation and improving the paraphrasing effect of the problem.
[0039] Exemplarily, the model training method and the problem processing method provided by the embodiments of the present application are both implemented based on artificial intelligence. Artificial intelligence is a comprehensive technology in computer science. It attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines, enabling the machines to have the functions of perception, reasoning, and decision-making. Artificial intelligence technology is an interdisciplinary subject, involving a wide range of fields, including both hardware-level technologies and software-level technologies. Artificial intelligence basic technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, pre-trained model technology, operation / interaction systems, and mechatronics. Among them, the pre-trained model is also called the large model or the basic model, and can be widely applied to downstream tasks in various directions of artificial intelligence after fine-tuning. Artificial intelligence software technologies mainly include several major directions such as computer vision technology, speech technology, natural language processing technology, and machine learning (ML) / deep learning.
[0040] In the embodiments of the present application, the artificial intelligence technologies mainly involved include the above-mentioned machine learning, natural language processing technology, etc. For example, it may involve text processing, semantic understanding, robot question answering, etc. in natural language processing technology; it may also involve deep learning in machine learning, including artificial neural networks, etc., which are not specifically described in the embodiments of the present application.
[0041] The method for model training provided in this application can be applied to a problem processing device with data processing capabilities. Exemplarily, the method for problem processing provided in this application can also be applied to the problem processing device mentioned above. As a schematic description, the mentioned problem processing device includes, but is not limited to, terminal devices, servers, question-and-answer robots, etc. Among them, terminal devices can include, but are not limited to, smartphones, desktop computers, laptop computers, tablet computers, smart speakers, in-vehicle devices, smart watches, wearable smart devices, intelligent voice interaction devices, smart home appliances, aircraft, virtual reality (VR) devices, augmented reality (AR) devices, mixed reality (MR) devices, etc. A server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery network (CDN), and big data and artificial intelligence platforms. This application does not make specific limitations. In addition, the mentioned terminal devices and servers can be directly or indirectly connected through wired communication or wireless communication, etc. This application does not make specific limitations.
[0042] The above-mentioned problem processing device can be capable of implementing the above-mentioned natural language processing technology. The mentioned natural language processing technology is an important direction in the fields of computer science and artificial intelligence. It studies various theories and methods that can achieve effective communication between humans and computers in natural language. Natural language processing involves natural language, that is, the language people use in daily life, and is closely related to linguistics research; at the same time, it involves computer science and mathematics. The pre-training model, an important technology for model training in the field of artificial intelligence, is developed from the large language model (LLM) in the NLP field. After fine-tuning, the large language model can be widely applied to downstream tasks. Natural language processing technology usually includes technologies such as text processing, semantic understanding, machine translation, robot question answering, and knowledge graphs. In the embodiments of this application, the problem processing device can use this natural language processing technology to achieve semantic understanding of sample questions, sample conversation contents, etc., and to achieve semantic understanding and other processing of the problem to be rewritten and the target conversation content corresponding to the problem to be rewritten.
[0043] In addition, the problem processing device may also have machine learning capabilities. Machine learning is an interdisciplinary field that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers can simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as neural networks. In the model training method and the problem processing method provided in the embodiments of the present application, the use of artificial intelligence models mainly involves the application of neural networks. Through neural networks, training processing is performed on sample problems, sample dialogue contents, and sample rewritten problems to train a target rewritten model, and then the target rewritten model is used to perform problem rewriting processing on the problem to be rewritten and the target dialogue content, etc.
[0044] Exemplarily, Figure 2 FIG. shows a schematic diagram of the system framework provided by the embodiments of the present application.
[0045] As Figure 2 shown, the system framework at least includes a model training stage and a model usage stage. Among them, in the model training stage, training data needs to be obtained, and the training data includes sample problems and corresponding sample dialogue contents. As a schematic description, a generalization dataset can be generated based on, for example, a natural language processing model, or a public dataset can be obtained, and then the above-mentioned training data can be obtained from one or more of the generalization dataset and the public dataset. For example, the data in the mentioned generalization dataset and public dataset can be examples proposed for error correction types, or examples proposed for anaphora resolution types, or examples proposed for substitution types, etc., which are not specifically limited in the present application. The described natural language processing model can include, but is not limited to, the GPT-4 model, etc., which is not limited in the present application.
[0046] After obtaining the training data, it is also necessary to perform supervised fine-tuning (SFT) on the pre-trained large language model to obtain an initial rewritten model. For example, the model parameters of the pre-trained large language model can be fine-tuned based on a preset low-rank adaptation LoRA model and a preset DeepSpeed model in a supervised manner to fine-tune and obtain an initial rewritten model. In this way, according to the instructions of the sample dialogue content and using the initial rewritten model, the problem rewriting of the sample problem is completed, and thus the sample rewritten problem corresponding to the sample problem is obtained. Further, the initial rewritten model is iteratively trained using the sample problem, the sample rewritten problem, and the sample dialogue content to obtain a target rewritten model.
[0047] During the model usage phase, the problem to be rewritten and the corresponding target dialogue content can be obtained from, for example, a dialogue system. Then, using the target rewriting model trained in the above model training phase, the problem rewriting process of the problem to be rewritten is completed according to the instructions of the target dialogue content, thereby obtaining the target problem corresponding to the problem to be rewritten.
[0048] Optionally, after completing the problem rewriting process for the problem to be rewritten, a retrieval and recall process can be performed on the target problem, or after this retrieval and recall process, it can also be applied to a vertical domain to update the target rewriting model, and then a large model applicable to the vertical domain can be obtained, etc. Specifically, it is not limited in this application.
[0049] It should be noted that the mentioned target rewriting model is a machine learning model obtained by training an initial rewriting model using sample problems, sample dialogue content, and sample rewritten problems. Subsequently, the training process of the target rewriting model described can be specifically referred to Figure 3 for understanding, and it will not be described here first.
[0050] The above-mentioned target rewriting model, initial rewriting model, pre-trained large language model, etc. can be deployed in a server or in a terminal device. Specifically, it is not limited in this application.
[0051] Exemplarily, the model training method and problem processing method provided in this application can also be applied to scenarios such as cloud technology, artificial intelligence, intelligent transportation, assisted driving, big data, knowledge retrieval, vertical domain large models, dialogue robots, etc. Specifically, it is not limited in this application.
[0052] As a schematic description, since the execution of the problem processing method described above depends on the target rewriting model trained by the model training method in the early stage. Therefore, first, from the perspective of embodiments, the model training method provided in the embodiments of this application will be described in detail. Exemplarily, Figure 3 shows a schematic flowchart of the model training method provided in the embodiments of this application. As Figure 3 shown, the model training method at least includes the following steps:
[0053] 301. Obtain training data, where the training data includes sample problems and the sample dialogue content corresponding to the sample problems.
[0054] In this example, the sample problem can be understood as a problem raised for the sample dialogue content, usually the last interrogative sentence or the last sentence in the sample dialogue content. The described sample dialogue content can be understood as an interactive dialogue between the two parties.
[0055] In some examples, it can be from the foregoingFigure 2 In the publicly available datasets described therein and the generalized datasets generated by natural language processing models, training data is obtained. The described training data can be training data of the information completion type, or it can also be training data of the error correction type, or coreference resolution type training data, or substitution type training data, etc., and no specific limitation is made.
[0056] For example, the following example describes training data of the information completion type, that is, the sample dialogue content is: Question: "What's the weather like in Beijing today?"; Answer: "The weather in Beijing today is sunny turning to cloudy, 24 degrees to 32 degrees, and the northerly wind is level 3." Question: "Is it suitable to wear a skirt?" At this time, the last question in the sample dialogue content, that is, "Is it suitable to wear a skirt?", can be used as the sample question.
[0057] Or, the following example describes training data of the error correction type, that is, the sample dialogue content is: Question: "How long do you want to stay in Dalian?" Answer: "If I find a job, I want to stay there for about three years." At this time, the last sentence in the sample dialogue content, that is, "If I find a job, I want to stay there for about three years.", can be used as the sample question.
[0058] Optionally, the sample dialogue content can also be used as a prompt to generate more datasets with the help of the aforementioned natural language processing model. For example, sample dialogue content and sample questions can be generated according to the examples and requirements in the dataset.
[0059] For example, taking a single-round dialogue as an example, its example is shown in the following code, that is:
[0060] {{"input":{utterance},"output":{rewrite_utterance}}};
[0061] The corresponding requirements can include the following: 1. The scenario of the dialogue is insurance consultation; 2. In the example, "input" represents that the input is the sample dialogue content, and the delimiter between each dialogue statement in the sample dialogue content is "\n"; 3. In the example, "output" represents the sample rewritten question obtained by rewriting the last dialogue statement (i.e., the sample question) in the sample dialogue content; 4. The number of training data generated each time is 10, and the generated sample dialogue content and sample rewritten questions conform to the form of the example; 5. The data format of each generated data is: {{"input": sample dialogue content (including sample question), "output": sample rewritten question}}, and finally the saved and returned data.
[0062] Alternatively, taking multi-turn conversations as an example, the examples are as follows: This is a sample conversation content: {history}. Please rewrite the content in single quotes according to the historical conversation: {query}. The corresponding rewriting requirements are as follows: 1. Make full use of the context information in the sample conversation content, and rewrite it from the dimensions of anaphora resolution, removal of special characters, correction of typos, etc., without creating content by yourself; 2. Only rewrite the content in single quotes, do not answer questions, and do not rewrite the historical conversation; 3. Do not change the original tone. An affirmative sentence remains an affirmative sentence after rewriting, and an interrogative sentence remains an interrogative sentence after rewriting; 4. For affirmative statements such as "mm-hmm", only return the statement of the previous question; 5. The sample questions for rewriting are the questions or answers of the user, and the sample conversation content is the conversation between the user and the insurance customer service; 6. Return one piece of data, that is, the data format of the sample rewritten question meets: {output}.
[0063] As a schematic description, for the training data in the generalization dataset generated by the natural language processing model, it can also be manually cleaned and adjusted again to obtain training data with better performance. In this way, combined with the aforementioned public dataset, a sufficient amount of training data can be obtained. For example, taking insurance consultation as an example, the data format of its corresponding training data is uniformly adjusted to the following format, for example: {"input": "I want to consult some questions about ** insurance.\nHello, ** insurance is mainly divided into two categories. One is the protection type of ** insurance, which mainly protects against death risks; the other is the investment type of ** insurance, which has both protection functions and certain financial management functions.\nI want to ask what the premium of the protection type of ** insurance is?","output": "What is the premium of the protection type of ** insurance?","instruction": "Please rewrite the last sentence of the input from the dimensions of anaphora resolution, information completion, word order adjustment, etc."}.
[0064] It should be noted that the various examples mentioned above are only for describing the sample conversation content and sample questions. In actual applications, there can also be other sample conversation content and sample questions, which are not specifically limited here.
[0065] 302. Input the sample conversation content and the sample question into the initial rewriting model to determine the second character in the sample question. The second character is the character used for rewriting in the sample question. The initial rewriting model is a model obtained by performing supervised fine-tuning on a pre-trained large language model.
[0066] In this example, since the pre-trained large language model has more parameters and a deeper model depth compared to a conventional neural network model, before using the initial rewriting model to perform question rewriting on the sample conversation content and sample questions, the pre-trained large language model can also be fine-tuned with supervision to obtain the initial rewriting model. As a schematic description, for example, a pre-trained large language model can be fine-tuned with supervision using a preset low-rank adaptation (LoRA) model and a preset DeepSpeed model, so as to accelerate the fine-tuning to obtain this initial rewriting model. In this way, the initial rewriting model is then applied to the training data, and then based on the initial rewriting model, question rewriting is performed on the sample conversation content and sample questions, thereby obtaining the sample rewritten questions corresponding to the sample questions.
[0067] The described preset low-rank adaptation model can be understood as a framework for efficiently fine-tuning a pre-trained large language model through low-rank matrix multiplication. The described preset DeepSpeed model can be understood as a tool for distributed training of new large-scale models, which can complete parallel training and accelerate the training process.
[0068] Exemplarily, after obtaining the initial rewriting model, the sample conversation content and sample questions can be input into the initial rewriting model to determine the second character in the sample question.
[0069] It should be noted that the described second character includes, but is not limited to, one or more of the characters in the sample question that can be resolved for anaphora, the omitted characters, the characters for correcting typos, and the replacement characters corresponding to the characters that can be resolved for anaphora.
[0070] The anaphora mentioned can be understood as using a referent to refer back to a previously mentioned linguistic unit in the conversation content. Generally, the referent is called the anaphor, and the object or content being referred to is called the antecedent. Usually, the antecedent can be before or after the anaphor. For example, in the example shown in step 301 above, that is: "How long do you want to stay in Dalian?", Answer: "If I find a job, I want to stay there for about three years." "There" can be used as the anaphor (referent), and "Dalian" can be used as the antecedent. In addition, anaphora resolution is to determine the corresponding relationship between the anaphor and the antecedent (such as "there" and the "antecedent"), and the same anaphor can also refer to different antecedents. The process of determining the antecedent of the anaphor is the process of anaphora resolution.
[0071] As a schematic description, in the process of determining the second character in the sample question based on the initial rewriting model, it can be understood with reference to the following method, that is: calculate the second part-of-speech similarity between each character in the sample question and the corresponding character in the sample conversation content based on the initial rewriting model. For example, the calculation and processing of the second part-of-speech similarity can be completed through the part-of-speech analysis function in the initial rewriting model. For example, the similarity distance between the character in the sample question and the character in the sample conversation content can be calculated to obtain the second part-of-speech similarity. The described second part-of-speech similarity can include but is not limited to cosine similarity, similarity based on Euclidean distance, etc., and is not limited here. In this way, after calculating the sum of the second part-of-speech similarities, then compare the relationship between the second part-of-speech similarity and each judgment threshold, and use the comparison result to determine the second character.
[0072] For example, in the case where it is determined that the second part-of-speech similarity is greater than the preset reference resolution judgment threshold, the corresponding character in the sample conversation content is determined as the character in the sample question that can be resolved by reference, or as the replacement character corresponding to the character that can be resolved by reference. Or, in the case where it is determined that the second part-of-speech similarity is greater than the preset ellipsis judgment threshold, the corresponding character in the sample conversation content is determined as the omitted character in the sample question. Or, in the case where it is compared that the second part-of-speech similarity is greater than the preset typo correction threshold, the corresponding character in the sample conversation content is determined as the typo correction character in the sample question.
[0073] It should be noted that the above-mentioned preset reference resolution judgment threshold can be understood as the critical value for judging whether a character conforms to reference resolution. The mentioned preset ellipsis judgment threshold is understood as the critical value for judging whether a character conforms to character ellipsis. The mentioned preset typo correction threshold is understood as the critical value for judging whether a character conforms to typo correction. In addition, for the preset reference resolution judgment threshold, the preset ellipsis judgment threshold, and the preset typo correction threshold, their values can be determined according to business requirements, and the specific values are not limited in this application.
[0074] 303. Input the characters in the sample conversation content into the initial rewriting model to update the second character at the corresponding position in the sample question, and obtain the sample rewritten question corresponding to the sample question.
[0075] In this example, after determining the second character of the sample question, the characters in the sample conversation content can be input into the initial rewriting model, so that the initial rewriting model uses the characters in the sample conversation content to update the second character at the corresponding position in the sample question, thereby obtaining the sample rewritten question corresponding to the sample question. For example, the characters in the sample conversation content can be used to replace the characters that can be resolved by anaphora in the sample question, and the sample question obtained after replacement is determined as the corresponding sample rewritten question. Another example is to fill the characters in the sample conversation content into the positions where the characters omitted in the sample question are located, so as to complete the information filling operation of the sample question, thereby obtaining the corresponding sample rewritten question. Another example is to use the characters in the sample conversation content to correct the misspelled characters in the sample question, thereby obtaining the corresponding sample rewritten question.
[0076] It should be noted that in practical applications, the characters in the sample conversation content can also be used to replace the abbreviated characters or colloquial characters in the sample question, etc., so as to obtain the corresponding sample rewritten question. Specific details are not limited in this application.
[0077] 304. Based on the sample conversation content, the sample question, and the sample rewritten question, the initial rewriting model is iteratively trained to obtain a target rewriting model, which is used to perform question rewriting processing on the question to be rewritten and the target conversation content, and obtain the target question corresponding to the question to be rewritten.
[0078] In this example, after obtaining the sample rewritten question, the sample conversation content, the sample question, and the sample rewritten question can be used as the input of the initial rewriting model, so that the initial rewriting model processes the sample conversation content, the sample question, and the sample rewritten question, and thus the target rewriting model is trained.
[0079] As a schematic description, since it is desired that the output of the deep neural network is as close as possible to the value that is truly desired to be predicted, the weight vectors of each layer of the neural network can be updated based on the difference between the predicted value of the current network and the truly desired target value (of course, there is usually an initialization process before the first update, that is, parameters are pre-configured for each layer in the deep neural network). For example, if the predicted value of the network is too high, the weight vector is adjusted to make it predict lower, and continuous adjustment is made until the neural network can predict the truly desired target value. Therefore, it is necessary to pre-define "how to compare the difference between the predicted value and the target value", which is the loss function or the objective function. They are important equations for measuring the difference between the predicted value and the target value. Among them, taking the loss function as an example, the higher the output value (loss) of the loss function, the greater the difference. Then, the training of the deep neural network becomes a process of minimizing this loss as much as possible.
[0080] Therefore, during the specific training process, the loss function can be added synchronously to improve the learning ability of the rewriting model. In other words, the retrieval and recall process can be performed on the sample rewriting problem to obtain the retrieval answer to the sample rewriting problem. In this way, by calculating the difference between the retrieval answer and the answer label corresponding to the sample problem, the loss value is obtained, and then the model parameters of the initial rewriting model are updated using the loss value to obtain the target rewriting model.
[0081] In this way, after training the target rewriting model, the target rewriting model can be used to process the problem to be rewritten and the target dialogue content. Specifically, it can be understood with reference to Figure 4 the problem processing method provided in the embodiments of the present application shown. As Figure 4 shown, the problem processing method provided by the present application at least includes the following steps:
[0082] 401. Obtain the problem to be rewritten and the target dialogue content corresponding to the problem to be rewritten.
[0083] In this example, the problem to be rewritten can be understood as the last question or the last statement in the target dialogue content, and there is no specific limitation. As a schematic description, the target dialogue content and the corresponding problem to be rewritten can be collected from the data stored in the dialogue system; or, the interactive dialogue between the user and the device can also be collected in real time to obtain the target dialogue content and the corresponding problem to be rewritten. Specifically, the present application does not limit the manner of obtaining the problem to be rewritten and the corresponding target dialogue content.
[0084] 402. Input the problem to be rewritten and the target dialogue content into the target rewriting model to determine the first character in the problem to be rewritten, where the first character is the character used for rewriting in the problem to be rewritten.
[0085] In this example, the target rewriting model described here is a machine learning model obtained by iteratively training an initial rewriting model using sample problems, sample dialogue content corresponding to the sample problems, and sample rewritten problems as training data. The mentioned sample rewritten problems are obtained by the initial rewriting model through problem rewriting processing of the sample dialogue content and the sample problems. The initial rewriting model is a model obtained by performing supervised fine-tuning on a pre-trained large language model. Regarding how to train the target rewriting model, it can be understood by referring to the training process shown in the foregoing Figure 3 and will not be elaborated here. After training the target rewriting model, the problem to be rewritten and the target dialogue content can be used as the input of the target rewriting model to process the problem to be rewritten and the target dialogue content through the target rewriting model, so as to obtain the target problem corresponding to the problem to be rewritten.
[0086] Exemplarily, after training the target rewriting model, the problem to be rewritten and the target object content obtained in step 401 can be used as the input of the target rewriting model to determine the first character in the problem to be rewritten through the target rewriting model. It should be noted that the described first character includes, but is not limited to, one or more of the characters that can be resolved by anaphora, the omitted characters, the characters for correcting typos, and the replacement characters corresponding to the characters that can be resolved by anaphora in the problem to be rewritten.
[0087] As a schematic description, in the process of determining the first character of the problem to be rewritten based on the target rewriting model, it can be understood by referring to the following method, that is: calculate the first part-of-speech similarity between each character in the problem to be rewritten and the corresponding character in the target dialogue content based on the target rewriting model. For example, the first part-of-speech similarity can be calculated through the part-of-speech analysis function in the target rewriting model. For example, calculate the similarity distance between the characters in the problem to be rewritten and the characters in the target dialogue content based on the part-of-speech analysis function, including but not limited to the cosine similarity distance, the Euclidean distance, etc., to obtain the first part-of-speech similarity. In this way, then compare the first part-of-speech similarity with each judgment threshold, and thus determine the first character based on the comparison result.
[0088] For example, in the case where it is determined that the first part-of-speech similarity is greater than the preset reference resolution judgment threshold, the corresponding character in the target dialogue content is determined as the character that can be resolved by reference in the problem to be rewritten, or as the replacement character corresponding to the character that can be resolved by reference. Or, in the case where it is determined that the first part-of-speech similarity is greater than the preset ellipsis judgment threshold, the corresponding character in the target dialogue content is determined as the ellipsis character in the problem to be rewritten. Or, in the case where it is determined that the first part-of-speech similarity is greater than the preset typo correction threshold, the corresponding character in the target dialogue content is determined as the typo correction character in the problem to be rewritten.
[0089] It should be noted that the above-mentioned preset reference resolution judgment threshold, preset ellipsis judgment threshold, and preset typo correction threshold can be understood with reference to the content described in step 302 above, and will not be elaborated here. Figure 3 in step 302 above, and will not be elaborated here.
[0090] 403. Input the characters in the target dialogue content into the target rewriting model to update the first character at the corresponding position in the problem to be rewritten, and obtain the target problem corresponding to the problem to be rewritten.
[0091] In this example, after determining the first character of the problem to be rewritten, the characters in the target object content can also be used as the input to the target rewriting model, so that the target rewriting model uses the characters in the target dialogue content to update the first character at the corresponding position in the problem to be rewritten, thereby obtaining the sample rewritten problem corresponding to the problem to be rewritten. For example, the characters in the target dialogue content can be used to replace the characters that can be resolved by reference in the problem to be rewritten, and the rewritten problem obtained after replacement is determined as the corresponding sample rewritten problem. Another example is to use the characters in the target dialogue content to fill in the positions of the ellipsis characters in the problem to be rewritten, thereby completing the information filling operation of the problem to be rewritten and obtaining the corresponding sample rewritten problem. Another example is to use the characters in the target dialogue content to correct the typo correction characters in the problem to be rewritten, thereby obtaining the corresponding target problem.
[0092] It should be noted that in practical applications, the characters in the target dialogue content can also be used to replace the abbreviated characters or colloquial characters in the problem to be rewritten to obtain the corresponding target problem. Specific details are not limited in this application.
[0093] In some alternative examples, after determining the target problem corresponding to the problem to be rewritten in step 403, the target problem can also be updated. For example, first extract the characters to be replaced in the target problem, and the characters to be replaced include one or more of abbreviated words and colloquial words. Exemplarily, the characters to be replaced can also include synonyms, etc., which are not limited herein. The abbreviated words described refer to the words obtained by abbreviating a certain word or phrase without changing the original meaning. For example, the abbreviated words include, but are not limited to, the abbreviated words of professional terms in different fields, English abbreviated words, etc., which are not specifically limited in this application. In addition, the colloquial words described can be understood as non-professional or non-written words. In this way, after extracting the characters to be replaced, match each of the characters to be replaced with each proper noun in the preset proper noun mapping table, so as to match the target proper noun of the character to be replaced, and then update the character to be replaced in the target problem with the target proper noun to obtain the updated target problem.
[0094] In the above manner, using proper nouns to replace the characters to be replaced in the target problem makes the target problem more clearly express the true intention, laying a foundation for subsequent retrieval and recall processing, and greatly improving the effect of subsequent retrieval and recall.
[0095] In some other alternative examples, since the question rewriting task is an upstream task of retrieval-augmented generation (RAG), the retrieval recall effect after rewriting the question should be at least not worse than that before rewriting the question. Therefore, after determining the target question in step 403, it is also possible to perform retrieval recall on the question to be rewritten and the target question respectively, to obtain the retrieval answers of the question to be rewritten and the retrieval answers of the target question. It should be noted that the target question described here can be the target question determined in step 403, or the target question updated using proper nouns, which is not limited here. In this way, after obtaining the retrieval answers of the question to be rewritten and the retrieval answers of the target question, based on a preset evaluation metric, the retrieval answers of the question to be rewritten and the retrieval answers of the target question are respectively evaluated to obtain a first evaluation result and a second evaluation result. The described first evaluation result can indicate the recall situation of the retrieval answer of the question to be rewritten. The described second evaluation result can indicate the recall situation of the retrieval answer of the target question. The described preset evaluation metric can include, but is not limited to, recall rate, accuracy, etc., which is not limited here. In this way, by comparing the first evaluation result and the second evaluation result, and in the case where the second evaluation result is greater than the first evaluation result, the retrieval answer of the target question and the target question are concatenated, and the concatenated data is used as a new instruction to fine-tune the target rewriting model, so that the fine-tuned target rewriting model can be applied to scenarios in the vertical domain.
[0096] It should be noted that in the case where the second evaluation result is greater than the first evaluation result, it indicates that after rewriting the question to be rewritten, the recall effect in the process of retrieving answers using the target question will be better than that of retrieving using the question to be rewritten, improving the recall effect. Moreover, using the target question and the corresponding retrieval answer to form a new instruction to fine-tune the target rewriting model can reduce the semantic retrieval deviation problem and solve the hallucination problem in the rewriting process, etc.
[0097] In the embodiments of the present application, by efficiently performing supervised fine-tuning on the pre-trained large language model, the initially rewritten model obtained by fine-tuning can be capable of the question rewriting task. Then, iterative training of the initially rewritten model is completed on the training data, enabling the target rewriting model to be effectively applied to the downstream question rewriting task, with a clearer intention expression, without the need to expand additional information to complete the question rewriting, effectively reducing the rewriting deviation and improving the question rewriting effect.
[0098] The above mainly introduces the solution provided by the embodiments of the present application from the perspective of methods. It can be understood that in order to implement the above functions, it includes the corresponding hardware structure and / or software module for executing each function. Those skilled in the art should easily realize that, in combination with the modules and algorithm steps of each example described in the embodiments disclosed in the present application, the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a certain function is executed in the way of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described function for each specific application, but such implementation should not be considered to exceed the scope of the present application.
[0099] The embodiments of the present application can divide the device into functional modules according to the above method examples. For example, each functional module can be divided corresponding to each function, or two or more functions can be integrated into one processing module. The above integrated module can be implemented in the form of hardware or in the form of a software functional module. It should be noted that the division of modules in the embodiments of the present application is illustrative, only a logical function division, and there may be other division methods in actual implementation.
[0100] In the embodiments of the present application, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other related parts to achieve a predetermined goal, and can be fully or partially implemented by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, one processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of the overall module or unit that includes the function of the module or unit.
[0101] The problem processing device in the embodiments of the present application will be described in detail below. Figure 5 It is a schematic diagram of an embodiment of the problem processing device provided in the embodiments of the present application. As Figure 5 shown, the problem processing device may include an acquisition unit 501 and a processing unit 502.
[0102] Among them, the acquisition unit 501 is used to acquire the problem to be rewritten and the target dialogue content corresponding to the problem to be rewritten. Specifically, it can be understood by referring to the content described in step 401 above, and details are not described here. Figure 4 The content described in step 401 above is not repeated here.
[0103] The processing unit 502 is configured to input the problem to be rewritten and the target dialogue content into the target rewriting model to determine the first character in the problem to be rewritten. The first character is the character used for rewriting in the problem to be rewritten. The target rewriting model is a machine learning model obtained by iteratively training an initial rewriting model using sample problems, corresponding sample dialogue contents, and sample rewritten problems as training data. The sample rewritten problems are obtained by the initial rewriting model performing problem rewriting processing on the sample dialogue contents and sample problems. The initial rewriting model is a model obtained by performing supervised fine-tuning on a pre-trained large language model. Specifically, reference can be made to the content described in step 402 above Figure 4 for understanding, which will not be elaborated here
[0104] The processing unit 502 is configured to input the characters in the target dialogue content into the target rewriting model to update the first character at the corresponding position in the problem to be rewritten, so as to obtain the target problem corresponding to the problem to be rewritten. Specifically, reference can be made to the content described in step 403 above Figure 4 for understanding, which will not be elaborated here.
[0105] In some alternative embodiments, the first character includes one or more of a character capable of anaphora resolution, an omitted character, a character for correcting typos, and a replacement character corresponding to the character capable of anaphora resolution.
[0106] In some other alternative embodiments, the processing unit 502 is configured to: calculate the first part-of-speech similarity between each character in the problem to be rewritten and the corresponding character in the target dialogue content based on the target rewriting model; when the first part-of-speech similarity is greater than a preset anaphora resolution judgment threshold, determine the corresponding character in the target dialogue content as the character capable of anaphora resolution in the problem to be rewritten, or the replacement character corresponding to the character capable of anaphora resolution; or, when the first part-of-speech similarity is greater than a preset omission judgment threshold, determine the corresponding character in the target dialogue content as the omitted character in the problem to be rewritten; or, when the first part-of-speech similarity is greater than a preset typo correction threshold, determine the corresponding character in the target dialogue content as the character for correcting typos in the problem to be rewritten.
[0107] In some other alternative embodiments, after the processing unit 502 inputs the characters in the target dialogue content into the target rewriting model to update the first character at the corresponding position in the problem to be rewritten to obtain the target problem corresponding to the problem to be rewritten, it extracts the characters to be replaced in the target problem. The characters to be replaced include one or more of abbreviated words and colloquial words; performs matching processing on the characters to be replaced with each proper noun in a preset proper noun mapping table to obtain the target proper noun of the characters to be replaced; and updates the characters to be replaced with the target proper noun to obtain the updated target problem.
[0108] In some other alternative embodiments, the processing unit 502 is further configured to: after inputting the characters in the target conversation content into the target rewriting model to update the first character at the corresponding position in the question to be rewritten, obtaining the target question corresponding to the question to be rewritten, respectively perform retrieval and recall on the question to be rewritten and the target question, to obtain the retrieval answer of the question to be rewritten and the retrieval answer of the target question; based on a preset evaluation metric, respectively perform evaluation processing on the retrieval answer of the question to be rewritten and the retrieval answer of the target question, to obtain a first evaluation result and a second evaluation result, the first evaluation result is used to indicate the recall situation of the retrieval answer of the question to be rewritten, and the second evaluation result is used to indicate the recall situation of the retrieval answer of the target question; when the second evaluation result is greater than the first evaluation result, splice the retrieval answer of the target question and the target question, and fine-tune the target rewriting model based on the spliced data.
[0109] The above Figure 5 The problem processing device has been mainly described from the perspective of functional modules. Next, the model training device will be described from the perspective of functional modules. Figure 6 FIG. is a schematic diagram of an embodiment of the model training device provided in the embodiments of the present application. As Figure 6 shown, the model training device includes an acquisition module 601 and a processing module 602.
[0110] Among them, the acquisition module 601 is configured to acquire training data, and the training data includes sample questions and sample conversation contents corresponding to the sample questions. Specifically, reference may be made to the content described in step 301 above Figure 3 for understanding, and details are not described here.
[0111] The processing module 602 is configured to input the sample conversation content and the sample question into the initial rewriting model to determine the second character in the sample question, where the second character is the character used for rewriting in the sample question, and the initial rewriting model is a model obtained by performing supervised fine-tuning on a pre-trained large language model. Specifically, reference may be made to the content described in step 302 above Figure 3 for understanding, and details are not described here.
[0112] The processing module 602 is configured to input the characters in the sample conversation content into the initial rewriting model to update the second character at the corresponding position in the sample question, obtaining the sample rewritten question corresponding to the sample question. Specifically, reference may be made to the content described in step 303 above Figure 3 for understanding, and details are not described here.
[0113] A processing module 602 is configured to iteratively train an initial rewriting model based on sample conversation content, a sample question, and a sample rewritten question to obtain a target rewriting model, where the target rewriting model is used to perform question rewriting processing on a question to be rewritten and the target conversation content corresponding to the question to be rewritten to obtain a target question corresponding to the question to be rewritten. Specifically, reference may be made to the content described in step 304 above, which will not be elaborated here. Figure 3 for understanding the content described in step 304 above, which will not be elaborated here.
[0114] In some alternative embodiments, the second character includes one or more of a character capable of anaphora resolution, an omitted character, a character for typo correction, and a replacement character corresponding to the character capable of anaphora resolution.
[0115] In some other alternative embodiments, the processing module 602 is configured to: calculate a second part-of-speech similarity between each character in the sample question and the corresponding character in the sample conversation content based on the initial rewriting model; when the second part-of-speech similarity is greater than a preset anaphora resolution judgment threshold, determine the corresponding character in the sample conversation content as a character capable of anaphora resolution in the sample question, or a replacement character corresponding to the character capable of anaphora resolution; or, when the second part-of-speech similarity is greater than a preset omission judgment threshold, determine the corresponding character in the sample conversation content as an omitted character in the sample question; or, when the second part-of-speech similarity is greater than a preset typo correction threshold, determine the corresponding character in the sample conversation content as a character for typo correction in the sample question.
[0116] In some alternative embodiments, the processing module 602 is configured to: retrieve and recall the sample rewritten question to obtain a retrieval answer for the sample rewritten question; calculate the difference between the retrieval answer and the answer label corresponding to the sample question to obtain a loss value; and update the model parameters of the initial rewriting model based on the loss value to obtain the target rewriting model.
[0117] The problem processing device in the embodiments of the present application is described above from the perspective of modular functional entities. Next, the problem processing device in the embodiments of the present application will be described from the perspective of hardware processing. Figure 6 is a schematic structural diagram of the problem processing device provided by the embodiments of the present application. The problem processing device may vary greatly due to configuration or performance, including but not limited to the Figure 5 problem processing device shown above, the Figure 6 model processing device shown above.
[0118] Such as Figure 7As shown, the problem processing device 300 can vary significantly due to differences in configuration or performance, and may include one or more central processing units (CPUs) 322 (e.g., one or more processors) and a memory 332, and one or more storage media 330 (e.g., one or more mass storage devices) for storing application programs 342 or data 344. Among them, the memory 332 and the storage media 330 can be transient storage or persistent storage. The programs stored in the storage media 330 can include one or more modules (not shown in the figure), and each module can include a series of instruction operations for the problem processing device. Further, the central processing unit 322 can be configured to communicate with the storage media 330 and execute a series of instruction operations in the storage media 330 on the problem processing device 300. Exemplarily, the central processing unit 322 is used to execute the application program 342 stored in the storage media 330, thereby implementing the model training method or the problem processing method provided in the foregoing embodiments of the present application.
[0119] The problem processing device 300 may further include one or more power supplies 326, one or more wired or wireless network interfaces 350, one or more input / output interfaces 358, and / or one or more operating systems 341.
[0120] Exemplarily, the central processing unit 322 in can cause the problem processing device to execute as by calling the computer-executable instructions stored in the memory 332. the method in the corresponding method embodiment.
[0121] Specifically, the function / implementation process of the processing module 602 in and the processing unit 502 in can be implemented by the central processing unit 322 in calling the computer-executable instructions stored in the memory 332. the function / implementation process of the acquisition module 601 in and the acquisition unit 501 in can be implemented by the input / output interface 358 in.
[0122] The steps performed by the problem processing device in the foregoing embodiments can be based on the shown problem processing device structure.
[0123] An embodiment of the present application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the methods described in the foregoing embodiments are implemented.
[0124] An embodiment of the present application further provides a computer program product, including a computer program which, when executed by a processor, implements the steps of the methods described in the foregoing embodiments.
[0125] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0126] In several embodiments provided by the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, indirect couplings or communication connections of devices or units, and can be in electrical, mechanical, or other forms.
[0127] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place, or can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0128] In addition, the functional units in each embodiment of the present application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.
[0129] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods of various embodiments of this application. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs.
[0130] The computer program product includes one or more computer instructions. When the computer execution instructions are loaded and executed on a computer, the processes or functions according to the embodiments of this application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website, a computer, a server, or a data center to another website, a computer, a server, or a data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that a computer can store, or a data storage device such as a server or a data center that includes one or more integrated available media. The available medium can be a magnetic medium (for example, floppy disk, hard disk, magnetic tape), an optical medium (for example, DVD), or a semiconductor medium (for example, SSD), etc.
[0131] The above embodiments are only used to illustrate the technical solutions of this application, rather than to limit it; although this application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of various embodiments of this application.
Claims
1. A method for problem handling, characterized in that, Including: Obtain the problem to be rewritten and the target dialogue content corresponding to the problem to be rewritten; Input the problem to be rewritten and the target dialogue content into the target rewriting model to determine the first character in the problem to be rewritten. The first character is the character used for rewriting in the problem to be rewritten. The target rewriting model is a machine learning model obtained by iteratively training an initial rewriting model based on sample problems, sample dialogue content corresponding to the sample problems, and sample rewritten problems. The sample rewritten problems are obtained by the initial rewriting model performing problem rewriting processing on the sample dialogue content and the sample problems. The initial rewriting model is a model obtained by performing supervised fine-tuning on a pre-trained large language model; Input the characters in the target dialogue content into the target rewriting model to update the first character at the corresponding position in the problem to be rewritten, and obtain the target problem corresponding to the problem to be rewritten.
2. The method according to claim 1, wherein The first character includes one or more of a character that can perform anaphora resolution, an omitted character, a character for correcting typos, and a replacement character corresponding to the character that can perform anaphora resolution.
3. The method according to claim 2, wherein Inputting the problem to be rewritten and the target dialogue content into the target rewriting model to determine the first character in the problem to be rewritten includes: Calculating the first part-of-speech similarity between each character in the problem to be rewritten and the corresponding character in the target dialogue content based on the target rewriting model; When the first part-of-speech similarity is greater than a preset anaphora resolution judgment threshold, determine the corresponding character in the target dialogue content as the character that can perform anaphora resolution in the problem to be rewritten, or the replacement character corresponding to the character that can perform anaphora resolution; or, When the first part-of-speech similarity is greater than a preset omission judgment threshold, determine the corresponding character in the target dialogue content as the omitted character in the problem to be rewritten; or, When the first part-of-speech similarity is greater than a preset typo correction threshold, determine the corresponding character in the target dialogue content as the character for correcting typos in the problem to be rewritten.
4. The method according to any one of claims 1 to 3, characterized in that After inputting the characters in the target dialogue content into the target rewriting model to update the first character at the corresponding position in the problem to be rewritten and obtaining the target problem corresponding to the problem to be rewritten, the method further includes: Extract the characters to be replaced in the target problem. The characters to be replaced include one or more of abbreviated words and colloquial words; Match the characters to be replaced with each proper noun in a preset proper noun mapping table to obtain the target proper noun of the characters to be replaced; Update the characters to be replaced with the target proper noun to obtain an updated target problem.
5. The method according to any one of claims 1 to 3, characterized in that After inputting the characters in the target dialogue content into the target rewriting model to update the first character at the corresponding position in the problem to be rewritten and obtaining the target problem corresponding to the problem to be rewritten, the method further includes: Retrieve and recall the problem to be rewritten and the target problem respectively to obtain the retrieval answer for the problem to be rewritten and the retrieval answer for the target problem; Based on a preset evaluation metric, evaluate the retrieval answer for the problem to be rewritten and the retrieval answer for the target problem respectively to obtain a first evaluation result and a second evaluation result. The first evaluation result is used to indicate the recall situation of the retrieval answer for the problem to be rewritten, and the second evaluation result is used to indicate the recall situation of the retrieval answer for the target problem; When the second evaluation result is greater than the first evaluation result, concatenate the retrieval answer for the target problem and the target problem, and fine-tune the target rewriting model based on the concatenated data.
6. A method for model training, characterized in that, It includes: Obtain training data, where the training data includes sample problems and sample dialogue contents corresponding to the sample problems; Input the sample dialogue content and the sample problem into an initial rewriting model to determine the second character in the sample problem. The second character is the character used for rewriting in the sample problem. The initial rewriting model is a model obtained by performing supervised fine-tuning on a pre-trained large language model; Input the characters in the sample dialogue content into the initial rewriting model to update the second character at the corresponding position in the sample problem, obtaining the sample rewritten problem corresponding to the sample problem; Based on the sample dialogue content, the sample problem, and the sample rewritten problem, perform iterative training on the initial rewriting model to obtain a target rewriting model. The target rewriting model is used to perform problem rewriting processing on the problem to be rewritten and the target dialogue content corresponding to the content to be rewritten, obtaining the target problem corresponding to the problem to be rewritten.
7. The method according to claim 6, characterized in that The second character includes one or more of a character that can be resolved by anaphora, an omitted character, a character for correcting typos, and a replacement character corresponding to the character that can be resolved by anaphora.
8. The method according to claim 7, wherein Inputting the sample dialogue content and the sample problem into the initial rewriting model to determine the second character in the sample problem includes: Calculate the second part-of-speech similarity between each character in the sample problem and the corresponding character in the sample dialogue content based on the initial rewriting model; When the second part-of-speech similarity is greater than a preset anaphora resolution judgment threshold, determine the corresponding character in the sample dialogue content as the character that can be resolved by anaphora in the sample problem, or the replacement character corresponding to the character that can be resolved by anaphora; or, When the second part-of-speech similarity is greater than a preset omission judgment threshold, determine the corresponding character in the sample dialogue content as the omitted character in the sample problem; or, When the second part-of-speech similarity is greater than a preset typo correction threshold, determine the corresponding character in the sample dialogue content as the character for correcting typos in the sample problem.
9. The method according to any one of claims 6 to 8, characterized in that Performing iterative training on the initial rewriting model based on the sample dialogue content, the sample problem, and the sample rewritten problem to obtain a target rewriting model includes: Retrieve and recall for the sample rewriting problem to obtain the retrieval answer for the sample rewriting problem; Calculate the difference between the retrieval answer and the answer label corresponding to the sample problem to obtain a loss value; Update the model parameters of the initial rewriting model based on the loss value to obtain a target rewriting model.
10. A problem processing device, characterized in that, Comprising: An acquisition unit for acquiring a problem to be rewritten and the target dialogue content corresponding to the problem to be rewritten; A processing unit for inputting the problem to be rewritten and the target dialogue content into the target rewriting model to determine a first character in the problem to be rewritten, where the first character is the character used for rewriting in the problem to be rewritten. The target rewriting model is a machine learning model obtained by iteratively training the initial rewriting model using the sample problem, the sample dialogue content corresponding to the sample problem, and the sample rewriting problem as training data. The sample rewriting problem is obtained by the initial rewriting model performing problem rewriting processing on the sample dialogue content and the sample problem. The initial rewriting model is a model obtained by performing supervised fine-tuning on a pre-trained large language model; The processing unit for inputting the characters in the target dialogue content into the target rewriting model to update the first character at the corresponding position in the problem to be rewritten to obtain the target problem corresponding to the problem to be rewritten.
11. A model training device, characterized in that, Comprising: An acquisition module for acquiring training data, where the training data includes a sample problem and the sample dialogue content corresponding to the sample problem; A processing module for inputting the sample dialogue content and the sample problem into the initial rewriting model to determine a second character in the sample problem, where the second character is the character used for rewriting in the sample problem. The initial rewriting model is a model obtained by performing supervised fine-tuning on a pre-trained large language model; The processing module for inputting the characters in the sample dialogue content into the initial rewriting model to update the second character at the corresponding position in the sample problem to obtain the sample rewriting problem corresponding to the sample problem; The processing module for performing iterative training on the initial rewriting model based on the sample dialogue content, the sample problem, and the sample rewriting problem to obtain a target rewriting model, where the target rewriting model is used to perform problem rewriting processing on the problem to be rewritten and the target dialogue content corresponding to the problem to be rewritten to obtain the target problem corresponding to the problem to be rewritten.
12. A problem processing device, characterized in that, Comprising: An input / output (I / O) interface, a processor, and a memory, where program instructions are stored in the memory; The processor is used to execute the program instructions stored in the memory and execute the method according to any one of claims 1 to 9.
13. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes instructions, and when the instructions run on a computer device, the computer device is caused to execute the method according to any one of claims 1 to 9.
14. A computer program product, characterized in that, The computer program product includes instructions, and when the instructions run on a computer device, the computer device is caused to execute the method according to any one of claims 1 to 9.