Question processing method and apparatus, device, storage medium and program product

By combining natural language thinking chain and procedural thinking chain methods, key information of the problem is extracted and a procedural solution is generated and verified. This solves the problem of insufficient accuracy and robustness of machine learning models in mathematical reasoning problems and achieves more efficient and accurate solutions.

WO2026090810A1PCT designated stage Publication Date: 2026-05-07BEIJING YOUZHUJU NETWORK TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
BEIJING YOUZHUJU NETWORK TECH CO LTD
Filing Date
2024-10-28
Publication Date
2026-05-07

AI Technical Summary

Technical Problem

Existing machine learning models suffer from insufficient accuracy and robustness when dealing with complex problems, especially mathematical reasoning problems. Their single-thinking-chain approach has poor generalization ability and cannot effectively improve the accuracy of the solution.

Method used

This approach combines natural language thinking and procedural thinking. By using a machine learning model to extract key information from the problem, generating a procedural language solution and verifying its accuracy, and then converting it into a natural language solution, the complementarity of the two thinking chains is utilized to improve the accuracy and robustness of the solution.

Benefits of technology

It improves the accuracy and robustness of problem handling, especially in mathematical reasoning problems, enhancing the logic and precision of solutions and compensating for the shortcomings of a single thought process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024127913_07052026_PF_FP_ABST
    Figure CN2024127913_07052026_PF_FP_ABST
Patent Text Reader

Abstract

Provided in the embodiments of the present disclosure are a question processing method and apparatus, a device, a storage medium and a program product. The method comprises: using a machine learning model to extract key information of a target question, the machine learning model being based on a language model; on the basis of the key information and the target question, using the machine learning model to determine a programmatic language answer for the target question by means of a program of thoughts, the programmatic language answer comprising a description of a process of answering the target question in a programming language; and, on the basis of the key information, the target question and the programmatic language answer, using the machine learning model to determine a natural language answer for the target question by means of a natural language chain of thoughts, the natural language answer at least comprising a verification result of the programmatic language answer. Thus, answers can be determined by combining the program of thoughts and the natural language chain of thoughts, thus improving the accuracy of question processing.
Need to check novelty before this filing date? Find Prior Art

Description

Methods, apparatus, devices, storage media, and program products for problem handling. Technical Field

[0001] The exemplary embodiments disclosed herein generally relate to the field of computers, and particularly to methods, apparatuses, electronic devices, computer-readable storage media, and computer program products for problem-solving. Background Technology

[0002] With the development of information technology, various terminal devices can provide people with a variety of services in work and life. For example, terminal devices can deploy applications that provide problem-solving services. Applications with problem-solving capabilities can output corresponding solutions based on user-input questions. Terminal devices or applications can use trained machine learning models (e.g., language models) to determine the corresponding solutions to questions, aiming to improve the accuracy of the determined solutions.

[0003] Summary of the Invention

[0004] In a first aspect of this disclosure, a problem-solving method is provided. The method includes: extracting key information about a target problem using a machine learning model, the machine learning model being based on a language model; determining a procedural language solution to the target problem using the machine learning model through a procedural thought chain based on the key information and the target problem, the procedural language solution including a description of the solution process to the target problem in procedural language; and determining a natural language solution to the target problem using the machine learning model through a natural language thought chain based on the key information, the target problem, and the procedural language solution, the natural language solution including at least a verification result of the procedural language solution.

[0005] In a second aspect of this disclosure, an apparatus for problem processing is provided. The apparatus includes: a key information extraction module configured to extract key information about a target problem using a machine learning model based on a language model; a first solution determination module configured to determine a procedural language solution to the target problem based on the key information and the target problem using the machine learning model through a procedural thought chain, the procedural language solution including a description of the solution process for the target problem in a procedural language; and a second solution determination module configured to determine a natural language solution to the target problem based on the key information, the target problem, and the procedural language solution using a machine learning model through a natural language thought chain, the natural language solution including at least a verification result of the procedural language solution.

[0006] In a third aspect of this disclosure, an electronic device is provided. The electronic device includes: at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions causing the electronic device to perform the method according to a first aspect of this disclosure when executed by the at least one processing unit.

[0007] In a fourth aspect of this disclosure, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, causes the processor to perform the method according to a first aspect of this disclosure.

[0008] In a fifth aspect of this disclosure, a computer program product is provided. The computer program product is tangibly stored in a computer storage medium and includes computer-executable instructions that, when executed by a device, cause the device to perform the method of the first aspect.

[0009] It should be understood that the content described in this content section is not intended to limit the key or essential features of the embodiments of this disclosure, nor is it intended to restrict the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0010] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. In the drawings, the same or similar reference numerals denote the same or similar elements, wherein:

[0011] Figure 1 shows a schematic diagram of an example environment in which embodiments of the present disclosure may be implemented;

[0012] Figure 2 shows a flowchart of a method for problem handling according to some embodiments of the present disclosure;

[0013] Figure 3 illustrates an example process of model training according to some embodiments of the present disclosure;

[0014] Figure 4 shows a schematic structural block diagram of a problem-solving apparatus according to some embodiments of the present disclosure; and

[0015] Figure 5 shows a block diagram of an electronic device that can implement one or more embodiments of the present disclosure. Detailed Implementation

[0016] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0017] It should be noted that the headings of any section / subsection provided herein are not limiting. Various embodiments are described throughout this document, and embodiments of any type may be included under any section / subsection. Furthermore, embodiments described in any section / subsection may be combined in any way with any other embodiments described in the same section / subsection and / or different sections / subsections.

[0018] In the description of embodiments of this disclosure, the term "comprising" and similar terms should be understood as open-ended inclusion, i.e., "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". Other explicit and implicit definitions may also be included below. The terms "first", "second", etc., may refer to different or the same objects. Other explicit and implicit definitions may also be included below.

[0019] The embodiments of this disclosure may involve user data, data acquisition, and / or use. All of these aspects comply with applicable laws, regulations, and relevant provisions. In the embodiments of this disclosure, all data collection, acquisition, processing, manipulation, forwarding, and use are conducted with the user's knowledge and confirmation. Accordingly, in implementing the embodiments of this disclosure, the type, scope of use, and usage scenarios of any data or information that may be involved should be communicated to the user and their authorization obtained in accordance with relevant laws and regulations through appropriate means. The specific methods of notification and / or authorization may vary depending on the actual situation and application scenario, and the scope of this disclosure is not limited in this respect.

[0020] In this specification and the embodiments, any processing of personal information will be carried out only under the premise of legality (such as obtaining the consent of the personal information subject, or being necessary for the performance of a contract), and will only be carried out within the scope stipulated or agreed upon. A user's refusal to process personal information other than that necessary for basic functions will not affect the user's use of basic functions.

[0021] As used in this paper, the term "model" refers to a model that learns the relationship between inputs and outputs from training data, enabling it to generate corresponding outputs for a given input after training. Model generation can be based on machine learning techniques. Deep learning is a machine learning algorithm that processes inputs and provides corresponding outputs using multiple layers of processing units. A neural network model is an example of a deep learning-based model. In this paper, "model" may also be referred to as a "machine learning model," "learning model," "machine learning network," or "learning network," and these terms are used interchangeably.

[0022] A neural network is a machine learning network based on deep learning. A neural network processes input and provides a corresponding output, typically consisting of an input layer, an output layer, and one or more hidden layers between the input and output layers. Neural networks used in deep learning applications often include many hidden layers, thus increasing the network's depth. The layers of a neural network are connected sequentially, so that the output of the previous layer is provided as the input to the next layer. The input layer receives the input to the neural network, while the output layer's output serves as the final output. Each layer of a neural network includes one or more nodes (also called processing nodes or neurons), each node processing the input from the layer above.

[0023] Machine learning typically comprises three phases: training, testing, and application (also known as inference). In the training phase, a given model is trained using a large amount of training data, iteratively updating parameter values ​​until the model can consistently generate inferences that meet the expected goals from the training data. Through training, the model can be considered to have learned the relationship between inputs and outputs (also known as an input-output mapping) from the training data. The parameter values ​​of the trained model are determined. In the testing phase, test inputs are applied to the trained model to test whether it can provide the correct output, thus determining the model's performance. The testing phase can sometimes be integrated into the training phase. In the application or inference phase, the trained model can be used to process actual model inputs based on the trained parameter values ​​to determine the corresponding model output.

[0024] Figure 1 illustrates a schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented. In this example environment 100, an application 120 is installed on a terminal device 110. A user 150 can interact with the application 120 via the terminal device 110 and / or an attachment device of the terminal device 110. For example, the application 120 can capture the user 150's voice via a voice capture device (e.g., a microphone) of the terminal device 110, can receive text input by the user via the display screen of the terminal device 110, and so on.

[0025] In embodiments of this disclosure, application 120 can be any suitable application with problem-solving capabilities. For example, application 120 can be a social application, a chat application, a media application, an educational application, and so on. Application 120 may provide a digital assistant for human-computer dialogue. This digital assistant supports text-based dialogue services, voice-based dialogue services, and content dialogue in other modalities with user 150. In some embodiments, application 120 or its digital assistant may utilize machine learning model 130. For example, application 120 or its digital assistant may utilize machine learning model 130 to provide problem-solving services to user 150. The digital assistant's response to the user may be determined based on the model output of machine learning model 130.

[0026] Machine learning model 130 may include at least one of machine learning model 130-1 deployed locally on terminal device 110 and machine learning model 130-2 deployed on other devices (e.g., machine learning model 130-2 on server 140). In this document, machine learning model 130-1 and machine learning model 130-2 may be collectively referred to as machine learning model 130. Machine learning model 130-1 deployed locally on terminal device 110 may be referred to as an offline machine learning model. Machine learning model deployed on other devices, such as machine learning model 130-2, may be referred to as an online machine learning model.

[0027] Machine learning model 130 can be based on any suitable model architecture, including but not limited to Transformer models, convolutional neural networks (CNNs), recurrent neural networks (RNNs), deep neural networks (DNNs), and so on. In some embodiments, machine learning model 114 and / or machine learning model 130 can be based on a language model (LM). A language model, by learning from a large corpus, is capable of question-answering. In some embodiments, machine learning model 130 can be a content-generating model, capable of generating corresponding outputs based on model inputs. The response to a question can be determined based on the model output. It should be noted that machine learning model 130 can include one or more machine learning models. If multiple machine learning models are included, their functions, structures, uses, etc., can be the same or different.

[0028] In environment 100, if application 120 is active, terminal device 110 can display interface 160 of application 120. Interface 160 may include various pages that application 120 can provide, such as a dialogue page between the user and a digital assistant (which may display the current dialogue and historical dialogue, including text dialogue content), and so on. In some embodiments, terminal device 110 may receive questions and provide corresponding answers to user 150 via interface 160.

[0029] In some embodiments, terminal device 110 communicates with server 140 to provide services to application 120. For example, server 140 may invoke machine learning model 130-2 to support the problem-solving functions of application 120 based on the output of machine learning model 130-2.

[0030] Terminal device 110 can be any type of mobile terminal, fixed terminal, or portable terminal, including mobile phones, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, media computers, multimedia tablets, personal communication system (PCS) devices, personal navigation devices, personal digital assistants (PDAs), audio / video players, digital cameras / camcorders, positioning devices, television receivers, radio broadcast receivers, e-book devices, gaming devices, or any combination thereof, including accessories and peripherals of these devices or any combination thereof. In some embodiments, terminal device 110 may also support any type of user-facing interface (such as "wearable" circuitry).

[0031] Server 140 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks, and big data and artificial intelligence platforms. Server 140 may include, for example, computing systems / servers such as mainframes, edge computing nodes, computing devices in a cloud environment, etc. Server 140 may be implemented based on a cloud environment.

[0032] It should be understood that the structure and function of the various elements in environment 100 are described for illustrative purposes only and do not imply any limitation on the scope of this disclosure.

[0033] As mentioned above, terminal devices or applications can use trained machine learning models to determine the answers to questions. In question-answering based on language models (including large language models), prompts are typically input into the language model to guide it in reasoning about the input question and determining the answer. In some question-answering scenarios, the language model can be guided through Chain-Of-Thought (COT). COT technology aims to improve the performance of machine learning models on complex reasoning tasks, such as arithmetic reasoning, common sense reasoning, and symbolic reasoning. COT technology guides the machine learning model to deduce a series of intermediate steps or sub-goals before generating the final answer. These intermediate steps constitute a "chain of thought," ultimately guiding the model to the correct result. Machine learning models can solve problems through COT, based on a series of thinking, analysis, and reasoning steps. COT can include Natural Language Chain of Thought (N-COT) and Procedural Chain of Thought (P-COT, which can be simply referred to as a procedural chain). P-COT requires machine learning models to execute COT using a procedural language (also known as a formal language, such as C, C++, Python, etc.), while N-COT requires machine learning models to execute COT using natural language.

[0034] While N-COT can improve the robustness and accuracy of inference by proceduralizing the inference process, and P-COT can represent these inference steps using a procedural language and be executed with the help of an interpreter to improve computational accuracy, both still have some limitations. For example, machine learning models using a single approach have poor generalization ability and are usually only suitable for handling specific types of problems. Traditionally, to improve the accuracy of the final solution, one approach is to determine the solution to the problem based on N-COT and P-COT separately, and then compare the two solutions to select the more accurate one. This approach does not consider the mutual influence of the two paths and cannot effectively improve the model's inference ability. Another approach is to execute N-COT before P-COT to use the N-COT inference path as a prefix for the P-COT inference path, using N-COT to guide the generation of the P-COT path. This approach only considers the influence of N-COT on the generation of P-COT paths, and similarly does not consider the influence of P-COT on the generation of N-COT paths, thus offering limited improvement to the model's inference ability.

[0035] In view of this, embodiments of this disclosure propose an improved problem-solving approach. This approach includes: utilizing a machine learning model to extract key information about a target problem, the machine learning model being based on a language model; based on the key information and the target problem, using the machine learning model to determine a procedural language solution to the target problem through a procedural thought chain, the procedural language solution including a description of the solution process for the target problem in procedural language; and based on the key information, the target problem, and the procedural language solution, using the machine learning model to determine a natural language solution to the target problem through a natural language thought chain, the natural language solution including at least a verification result of the procedural language solution. In this manner, embodiments of this disclosure can utilize P-COT to achieve problem-solving, and require the model to verify the solution through natural language via N-COT, which can improve the robustness and accuracy of problem-solving, particularly the robustness and accuracy of mathematical reasoning.

[0036] The following section provides a detailed description of various example implementations of this scheme, with reference to the accompanying drawings.

[0037] Figure 2 shows a flowchart of a method 200 for problem handling according to some embodiments of the present disclosure. For ease of discussion, method 200 will be described with reference to the environment 100 of Figure 1. Method 200 may be implemented at terminal device 110 and / or server 140. For ease of description only, it will be described by way of example that method 200 is implemented at server 140.

[0038] It should be noted that, taking the implementation of method 200 at terminal device 110 as an example, some operations described with reference to terminal device 110 may require the assistance of server 140 to complete. The operations performed by terminal device 110 may specifically be performed by relevant applications installed on terminal device 110.

[0039] In box 210, server 140 uses machine learning model 130 to extract key information about the target problem, which is based on a language model.

[0040] In some embodiments, terminal device 110 can receive a target question input by a user in natural language. The target question can be any suitable type of question, such as text or speech. It is understood that a speech question can be any suitable speech of any duration, language, and tone, and a text question can be any suitable text of any number of words and language. Terminal device 110 can provide the target question to server 140 via a communication connection between terminal device 110 and server 140. In some embodiments, if the target question is a non-text type (e.g., speech type), terminal device 110 can convert the target question into a text type question and provide it to server 140. Alternatively or additionally, in some embodiments, terminal device 110 can also directly provide a non-text type target question to server 140, and server 140 can convert the non-text type target question into a text type itself.

[0041] The target problem can be any suitable problem. In some embodiments, the target problem may include a mathematical reasoning problem. Mathematical reasoning problems are typically problems that require the application of mathematical concepts, principles, rules, and logical reasoning abilities to solve. Mathematical reasoning problems can usually be broken down into a series of steps or stages, each step being a logical progression towards the final answer. The answers to mathematical reasoning problems are usually precise and require rigorous reasoning to support them. In the embodiments of this disclosure, key information is first obtained through problem analysis, and then the reasoning process of the procedural language is implemented using P-COT, and the reasoning verification is implemented using N-COT, which better matches the mathematical reasoning and verification process.

[0042] Server 140 may determine the prompt word input (which may be referred to as the first prompt word input) for machine learning model 130 based on the target question, for example. In some embodiments, server 140 may also obtain a prompt word template (which may be referred to as the first prompt word template) for machine learning model 130 and determine the first prompt word input for machine learning model 130 by filling the first prompt word template with the target question. The first prompt word input may, for example, guide machine learning model 130 to extract key information from the target question. Key information may include, for example, numbers, variables, etc., in the target question. Referring to Table 1, Table 1 shows examples of target questions and key information:

[0043] Table 1

[0044] As shown in Table 1, server 140 can, for example, determine the corresponding first prompt input based on the target problem: "In a town, there is a multi-story parking garage that can hold 425 cars. The garage has 5 floors, each the same size. If 23 cars are already parked on a floor, how many more cars can that floor hold?" The reasoning for the answer is: "To solve this problem, we first need to find all the numerical information in the problem." Server 140 can, for example, provide the first prompt input to machine learning model 130. Machine learning model 130 can extract key information from the target problem based on the first prompt input, such as "The parking garage can hold 425 cars, the parking garage has 5 floors, and 23 cars are already parked on the first floor," as shown in Table 1.

[0045] In box 220, server 140, based on key information and target question, uses machine learning model 130 to determine a procedural language solution to the target question through a procedural thought chain. The procedural language solution includes a description of the solution process to the target question in a procedural language.

[0046] By leveraging the model's information extraction capabilities, key information from the input question can be summarized in the first stage. This key information is then used to assist in generating the P-COT inference chain, utilizing the rigorous logic and precise computational capabilities of a procedural language. The procedural language solution can include solutions in any suitable procedural language such as C, C++, or Python.

[0047] In some embodiments, server 140 may determine the prompt word input (which may be referred to as the second prompt word input) for machine learning model 130 based, for example, on key information and the target question. In some embodiments, server 140 may also obtain a prompt word template (which may be referred to as the second prompt word template) for machine learning model 130 and determine the second prompt word input for machine learning model 130 by filling the second prompt word template with the target question and key information. The second prompt word input may, for example, guide machine learning model 130 to determine a procedural language solution to the target question through a procedural thought process. Table 2 shows examples of procedural language solutions to the example questions in Table 1:

[0048] Table 2

[0049] As shown in Tables 1 and 2, server 140, for example, can determine the second prompt input corresponding to the target problem shown in Table 1: "In a town, there is a multi-story parking garage that can park 425 cars. The garage has 5 floors, each of the same size. If 23 cars are already parked on a floor, how many more cars can that floor hold?" and the corresponding key information: "The parking garage can park 425 cars, the parking garage has 5 floors, and 23 cars are already parked on a floor." The prompt input is: "In a town, there is a multi-story parking garage that can park 425 cars. The parking garage has 5 floors, each of the same size. If 23 cars are already parked on a floor, how many more cars can that floor hold?" The reasoning is: "To solve this problem, we first find all the numerical information in the problem: 'The parking garage can park 425 cars, the parking garage has 5 floors, and 23 cars are already parked on a floor.' Please refer to this numerical information to complete a Python-style solution." Server 140 can then provide this second prompt input to machine learning model 130. Machine learning model 130 can determine the procedural language solution as shown in Table 2 based on this second cue word input.

[0050] In box 230, server 140 uses machine learning model 130 to determine a natural language solution for the target question based on key information, the target question, and the programmed language solution. The natural language solution includes at least the verification results of the programmed language solution.

[0051] The natural language solution can be in any suitable language. The language of the natural language solution can, for example, be the same as the language of the target question. In some embodiments, the natural language solution may also include a description of the procedural language solution in natural language. Server 140 may, for example, determine the prompt word input (which may be referred to as the third prompt word input) for machine learning model 130 based on the target question, key information, and the procedural language solution. In some embodiments, server 140 may also obtain a prompt word template (which may be referred to as the third prompt word template) for machine learning model 130 and determine the third prompt word input for machine learning model 130 by filling the third prompt word template with the target question, key information, and procedural language solution. The third prompt word input may, for example, guide machine learning model 130 to determine the natural language solution for the target question through a natural language thought process. Table 3 shows examples of natural language solutions for the procedural language solution examples in Table 2:

[0052] Table 3

[0053] As shown in Tables 1, 2, and 3, server 140 can, for example, determine the third prompt input for the target problem "In a town, there is a multi-story parking lot that can park 425 cars. The parking lot has 5 levels, each of the same size. If 23 cars are already parked on a level, how many more cars can that level hold?" based on the target problem shown in Table 1, the key information corresponding to the target problem "The parking lot can park 425 cars, the parking lot has 5 levels, and 23 cars are already parked on a level", and the procedural language solution shown in Table 2. The prompt is: "In a town, there is a multi-story parking lot that can park 425 cars. The parking lot has 5 levels, each of the same size. If 23 cars are already parked on a level, how many more cars can that level hold?". The reasoning for the answer is: To solve this problem, we first find all the numerical information in the problem: The parking lot can park 425 cars, the parking lot has 5 levels, and 23 cars are already parked on a level. Please refer to this numerical information to complete the Python-style solution: def solution(): 425 cars. The parking lot has 5 levels, each of the same size. How many more cars can fit on one level if there are already 23 parked cars on that level? \"\"\n cars_total=425\n levels=5\n cars_per_level=cars_total / levels\n cars_parked=23\n cars_left=cars_per_level-cars_parked\n result=cars_left\n return result\nPlease convert this Python code-style solution into a natural language-style solution for human understanding. Therefore, the natural language-style solution is:\n". Server 140 can, for example, provide this third prompt input to machine learning model 130. Machine learning model 130 can determine the natural language solution as shown in Table 3 based on this third prompt input. "The answer is: 62" in the natural language solution shown in Table 3 can be considered as a verification result of the procedural language solution.

[0054] Therefore, the information extraction capabilities of machine learning models can be leveraged to extract key information from the target problem. This key information can then be used to assist in generating procedural language solutions, enabling precise computation. Procedural language solutions can be translated into natural language solutions to verify their logic and accuracy. Besides verifying procedural language solutions using natural language, translation can also yield easily expressible natural language reasoning schemes, compensating for the difficulty in expressing procedural language solutions. Furthermore, the concise and precise reasoning steps of procedural language solutions can improve performance (natural language reasoning redundancy), achieving a synergistic improvement and enhancing the accuracy of problem-solving.

[0055] Regarding the training method of the machine learning model 130, in some embodiments, the machine learning model 130 can be trained on the terminal device 110, server 140, or any other suitable electronic device. For ease of description, the device / system for training the machine learning model 130 can be referred to as a model training system. Figure 3 illustrates an example process 300 of model training according to some embodiments of this disclosure. The example process 300 can be implemented at the model training system.

[0056] As shown in Figure 3, the training of the machine learning model 130 includes two stages. The model training system can train the machine learning model 130 through supervised learning 301 in the first stage, based on the sample problem and the labeled information for the sample problem. After the first stage of supervised learning 301, the machine learning model 130 is further trained through reinforcement learning 302 in the second stage. That is, the model training system can first perform the first stage of training on the machine learning model 130, and after the first stage of training is completed, perform the second stage of training on the machine learning model 130 trained in the first stage.

[0057] Regarding the specific method of supervised learning 301 in the first stage, in some embodiments, the model training system can determine three tasks: extracting key information, generating procedural language solutions, and generating natural language solutions. The model training system can acquire a set of sample questions and a corresponding set of annotation information for these three tasks. The set of sample questions includes one or more sample questions. It should be noted that the sample questions for different tasks can be different. For example, the sample questions for the task of extracting key information can be different from the sample questions for the task of generating procedural language solutions. Furthermore, the sample questions for each task can include multiple sample questions. For ease of description, the following example only illustrates that the sample questions for each task include one sample question.

[0058] In some embodiments, the annotation information corresponding to the sample question for the task of extracting key information may include sample key information. The model training system can obtain a first sample question for the task of extracting key information and annotation information 310 for the first sample question (e.g., first sample key information). The model training system can determine the first expected key information corresponding to the first sample question by providing the first sample question for the task of extracting key information to the machine learning model 130.

[0059] The model training system can then train the machine learning model 130 based on the difference (for example, referred to as the first difference) between the first expected key information and the first sample key information corresponding to the first sample problem. The model training system can, for example, train the machine learning model 130 by reducing the first difference. The model training system can determine that the machine learning model 130 has completed training for the task of extracting key information in response to the first difference being less than a corresponding first threshold. The first threshold can, for example, be any pre-determined appropriate threshold.

[0060] In some embodiments, the annotation information for a sample question for the task of generating a procedural language solution may include sample key information and sample procedural language solution. The model training system can obtain a second sample question for the task of generating a procedural language solution and annotation information 320 for the second sample question (e.g., second sample key information and second sample procedural language solution). The model training system can determine a second predicted procedural language solution corresponding to the second sample question by providing the second sample question and the second sample key information corresponding to the second sample question to the machine learning model 130.

[0061] The model training system can then train the machine learning model 130 based on the difference between the second predicted procedural language solution and the second sample procedural language solution corresponding to the second sample problem (which may be referred to as the second difference, for example). Similarly, the model training system can train the machine learning model 130, for example, by reducing the second difference. The model training system can determine that the machine learning model 130 has completed training for the task of generating procedural language solutions in response to the second difference being less than a corresponding second threshold. The second threshold may be any pre-determined appropriate threshold.

[0062] In some embodiments, for a sample question in the task of generating a natural language answer, the corresponding annotation information includes sample key information, sample procedural language answer, and sample natural language answer. The model training system can obtain a third sample question for the task of generating a natural language answer and annotation information 330 for the third sample question (e.g., third sample key information, third sample procedural language answer, and third sample natural language answer). The model training system can determine the third expected natural language answer corresponding to the third sample question by providing the third sample question, the third sample key information corresponding to the third sample question, and the third sample procedural language answer corresponding to the third sample question to the machine learning model 130.

[0063] The model training system can then train the machine learning model 130 based on the difference between the third predicted natural language answer and the third sample natural language answer corresponding to the third sample question (which may be referred to as the third difference, for example). Similarly, the model training system can, for example, train the machine learning model 130 by reducing the third difference. The model training system can determine that the machine learning model 130 has completed training for the task of generating natural language answers in response to the third difference being less than a corresponding third threshold. The third threshold may, for example, be any pre-determined appropriate threshold.

[0064] The model training system can determine that the machine learning model 130 has completed the training of the three tasks of extracting key information, generating procedural language solutions, and generating natural language solutions, and that the machine learning model 130 has completed the first stage of training.

[0065] Furthermore, regarding the specific method of the second-stage reinforcement learning 302, in some embodiments, the model training system can acquire a fourth sample question 340 and, by providing the fourth sample question 340 to the machine learning model 130, determine the fourth predicted key information 350, the fourth predicted procedural language solution 360, and the fourth predicted natural language solution 370 corresponding to the fourth sample question 340. The specific method for using the machine learning model 130 to determine the key information, the procedural language solution, and the natural language solution has been described in detail above and will not be repeated here.

[0066] The model training system can determine the first reward feedback corresponding to the fourth predicted procedural language solution 360 in the second stage of reinforcement learning, and the second reward feedback corresponding to the fourth predicted natural language solution 370 in the second stage of reinforcement learning. The model training system can determine the reward feedback (i.e., the first and second reward feedback) in any suitable manner. For example, the model training system can provide the fourth predicted procedural language solution 360 and the fourth predicted natural language solution 370 to the user to obtain reward feedback provided by the user. Alternatively, the model training system can use a trained reward model to determine the scores of the fourth predicted procedural language solution 360 and the fourth predicted natural language solution 370, and determine the first and second reward feedback based on their respective scores.

[0067] The model training system can train the machine learning model 130 based on first and second reward feedback. The reward feedback can guide the machine learning model 130 to learn and generate more accurate fourth predicted procedural language solutions 360 and fourth predicted natural language solutions 370. Specifically, since the accuracy of the fourth predicted procedural language solution 360 is affected by the fourth predicted key information 350 (e.g., more accurate fourth predicted key information 350 can help the model generate a more accurate fourth predicted procedural language solution 360), the model training system can adjust the extraction of key information by the machine learning model 130 based on the first reward feedback. That is, the model training system can extract more accurate fourth predicted key information 350 based on the first reward feedback for the fourth predicted procedural language solution 360.

[0068] Since the accuracy of the fourth predicted natural language solution 370 is affected by the fourth predicted procedural language solution 360 (e.g., a more accurate fourth predicted procedural language solution 360 can help the model generate a more accurate fourth predicted natural language solution 370), the model training system can adjust the machine learning model 130's generation of procedural language solutions based on the first reward feedback and the second reward feedback. That is, the model training system can generate a more accurate fourth predicted procedural language solution 360 based on the first reward feedback for the fourth predicted procedural language solution 360 and the second reward feedback for the fourth predicted natural language solution 370.

[0069] The model training system can also adjust the generation of natural language answers by the machine learning model 130 based on the second reward feedback. That is, the model training system can generate a more accurate fourth predicted natural language answer 370 based on the second reward feedback for the fourth predicted natural language answer 370.

[0070] Therefore, the model training system can acquire multiple sample questions for the second-stage reinforcement learning 302, and repeatedly use the reward feedback for the expected procedural language solution and the reward feedback for the expected natural language solution to train the machine learning model. This can help the machine learning model to better complete the translation task and improve the quality of its inference chain.

[0071] Embodiments of this disclosure also provide corresponding apparatus for implementing the methods or processes described above. Figure 4 shows a schematic structural block diagram of an apparatus 400 for problem handling according to some embodiments of this disclosure. The apparatus 400 may be implemented as or included in the terminal device 110 and / or the server 140. The various modules / components in the apparatus 400 may be implemented by hardware, software, firmware, or any combination thereof.

[0072] As shown in Figure 4, the device 400 includes: a key information extraction module 410, configured to extract key information of a target problem using a machine learning model, the machine learning model being based on a language model; a first solution determination module 420, configured to determine a procedural language solution to the target problem based on the key information and the target problem using a machine learning model through a procedural thought chain, the procedural language solution including a description of the solution process to the target problem in procedural language; and a second solution determination module 430, configured to determine a natural language solution to the target problem based on the key information, the target problem, and the procedural language solution using a machine learning model through a natural language thought chain, the natural language solution including at least a verification result of the procedural language solution.

[0073] In some embodiments, the natural language solution also includes a description of the procedural language solution in natural language.

[0074] In some embodiments, the target problem includes a mathematical reasoning problem.

[0075] In some embodiments, the machine learning model is trained by: training the machine learning model through a first-stage supervised learning based on a sample problem and labeled information for the sample problem; and retraining the machine learning model through a second-stage reinforcement learning after the first-stage supervised learning.

[0076] In some embodiments, the annotation information includes sample key information, and training the machine learning model through supervised learning in the first stage includes: providing a first sample question to the machine learning model, determining first expected key information corresponding to the first sample question; and training the machine learning model based on the difference between the first expected key information and the first sample key information corresponding to the first sample question.

[0077] In some embodiments, the annotation information includes sample key information and sample procedural language solutions, and training the machine learning model through supervised learning in the first stage includes: determining the second expected procedural language solution corresponding to the second sample question by providing the second sample key information corresponding to the second sample question to the machine learning model; and training the machine learning model based on the difference between the second expected procedural language solution and the second sample procedural language solution corresponding to the second sample question.

[0078] In some embodiments, the annotation information includes sample key information, sample procedural language solutions, and sample natural language solutions, and training the machine learning model through supervised learning in the first stage includes: providing the machine learning model with the third sample question, the third sample key information corresponding to the third sample question, and the third sample procedural language solution corresponding to the third sample question to determine the third expected natural language solution corresponding to the third sample question; and training the machine learning model based on the difference between the third expected natural language solution and the third sample natural language solution corresponding to the third sample question.

[0079] In some embodiments, retraining the machine learning model through a second-stage reinforcement learning includes: providing a fourth sample question to the machine learning model; determining a fourth predicted key information, a fourth predicted procedural language solution, and a fourth predicted natural language solution corresponding to the fourth sample question; determining a first reward feedback corresponding to the fourth predicted procedural language solution in the second-stage reinforcement learning, and a second reward feedback corresponding to the fourth predicted natural language solution in the second-stage reinforcement learning; and training the machine learning model based on the first reward feedback and the second reward feedback.

[0080] In some embodiments, training a machine learning model based on a first reward feedback and a second reward feedback includes: adjusting the machine learning model's extraction of key information based on the first reward feedback; adjusting the machine learning model's generation of procedural language answers based on the first reward feedback and the second reward feedback; and adjusting the machine learning model's generation of natural language answers based on the second reward feedback.

[0081] The units and / or modules included in device 400 can be implemented in various ways, including software, hardware, firmware, or any combination thereof. In some embodiments, one or more units and / or modules can be implemented using software and / or firmware, such as machine-executable instructions stored on a storage medium. In addition to or as an alternative to machine-executable instructions, some or all of the units and / or modules in device 400 can be implemented at least partially by one or more hardware logic components. By way of example and not limitation, exemplary types of hardware logic components that can be used include field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-chips (SoCs), complex programmable logic devices (CPLDs), and so on.

[0082] It should be understood that one or more steps in the above methods can be performed by appropriate electronic devices or combinations of electronic devices. Such electronic devices or combinations of electronic devices may, for example, include terminal device 110 and / or server 140 in FIG. 1.

[0083] Figure 5 shows a block diagram of an electronic device 500 in which one or more embodiments of the present disclosure may be implemented. It should be understood that the electronic device 500 shown in Figure 5 is merely exemplary and should not be construed as limiting the functionality and scope of the embodiments described herein. The electronic device 500 shown in Figure 5 can be used to implement the terminal device 110 of Figure 1, the server 140, or the device 400 of Figure 4.

[0084] As shown in Figure 5, the electronic device 500 is in the form of a general-purpose electronic device. Components of the electronic device 500 may include, but are not limited to, one or more processors or processing units 510, memory 520, storage devices 530, one or more communication units 540, one or more input devices 550, and one or more output devices 560. The processing unit 510 may be a physical or virtual processor and is capable of performing various processes according to programs stored in the memory 520. In a multiprocessor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing capability of the electronic device 500.

[0085] Electronic device 500 typically includes multiple computer storage media. Such media can be any accessible media that is accessible to electronic device 500, including but not limited to volatile and non-volatile media, removable and non-removable media. Memory 520 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. Storage device 530 can be removable or non-removable media and can include machine-readable media, such as flash drives, disks, or any other media that can be used to store information and / or data and can be accessed within electronic device 500.

[0086] Electronic device 500 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in FIG. 5, disk drives for reading from or writing to removable, non-volatile disks (e.g., "floppy disks") and optical disk drives for reading from or writing to removable, non-volatile optical disks may be provided. In these cases, each drive may be connected to a bus (not shown) via one or more data media interfaces. Memory 520 may include computer program product 525 having one or more program modules configured to perform various methods or actions of various embodiments of the present disclosure.

[0087] Communication unit 540 enables communication with other electronic devices via a communication medium. Additionally, the functionality of components of electronic device 500 can be implemented using a single computing cluster or multiple computing machines capable of communicating via communication connections. Therefore, electronic device 500 can operate in a networked environment using logical connections to one or more other servers, network personal computers (PCs), or another network node.

[0088] Input device 550 can be one or more input devices, such as a mouse, keyboard, trackball, etc. Output device 560 can be one or more output devices, such as a monitor, speaker, printer, etc. Electronic device 500 can also communicate with one or more external devices (not shown) via communication unit 540 as needed. These external devices include storage devices, display devices, etc., and can communicate with one or more devices that enable user interaction with electronic device 500, or with any device that enables electronic device 500 to communicate with one or more other electronic devices (e.g., network card, modem, etc.). Such communication can be performed via input / output (I / O) interface (not shown).

[0089] According to an exemplary implementation of this disclosure, a computer-readable storage medium is provided that stores computer-executable instructions thereon, wherein the computer-executable instructions are executed by a processor to implement the methods described above. According to an exemplary implementation of this disclosure, a computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions that are executed by a processor to implement the methods described above.

[0090] Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatuses, devices, and computer program products implemented according to this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0091] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processing unit of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.

[0092] Computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions that execute on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.

[0093] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some, as newer, implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0094] Various implementations of this disclosure have been described above. The foregoing description is exemplary and not exhaustive, nor is it limited to the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is chosen to best explain the principles, practical applications, or improvements to technology in the market, or to enable others skilled in the art to understand the various implementations disclosed herein.

Claims

1. A problem-solving method, comprising: The machine learning model is used to extract key information about the target problem, and the machine learning model is based on a language model. Based on the key information and the target problem, the machine learning model is used to determine a programmatic language solution to the target problem through a programmatic thinking chain. The programmatic language solution includes a description of the solution process to the target problem in a programmatic language. as well as Based on the key information, the target question, and the procedural language solution, the machine learning model is used to determine a natural language solution for the target question through a natural language thought chain. The natural language solution includes at least the verification result of the procedural language solution.

2. The method of claim 1, wherein the natural language solution further includes a description of the procedural language solution in natural language.

3. The method according to claim 1, wherein the target problem includes a mathematical reasoning problem.

4. The method of claim 1, wherein the machine learning model is trained via: Based on the sample problem and the annotation information for the sample problem, the machine learning model is trained through supervised learning in the first stage; and After the first stage of supervised learning, the machine learning model is retrained through the second stage of reinforcement learning.

5. The method according to claim 4, wherein the annotation information includes key sample information, and training the machine learning model through supervised learning in the first stage comprises: By providing a first sample question to the machine learning model, the first predicted key information corresponding to the first sample question is determined; as well as The machine learning model is trained based on the difference between the first predicted key information and the first sample key information corresponding to the first sample problem.

6. The method according to claim 4, wherein the annotation information includes sample key information and sample procedural language solutions, and training the machine learning model through supervised learning in the first stage comprises: By providing the machine learning model with the second sample question and the corresponding key information of the second sample question, a second predicted procedural language solution corresponding to the second sample question is determined; and The machine learning model is trained based on the difference between the second predicted procedural language solution and the second sample procedural language solution corresponding to the second sample problem.

7. The method according to claim 4, wherein the annotation information includes sample key information, sample procedural language solutions, and sample natural language solutions, and training the machine learning model through supervised learning in the first stage includes: By using the third sample question, the key information of the third sample corresponding to the third sample question, and the third... The third sample procedural language solution corresponding to the sample question is provided to the machine learning model to determine the third predicted natural language solution corresponding to the third sample question. as well as The machine learning model is trained based on the difference between the third predicted natural language solution and the third sample natural language solution corresponding to the third sample question.

8. The method of claim 4, wherein retraining the machine learning model through a second stage of reinforcement learning comprises: By providing the fourth sample question to the machine learning model, the fourth predicted key information, the fourth predicted procedural language solution, and the fourth predicted natural language solution corresponding to the fourth sample question are determined. Determine the first reward feedback corresponding to the fourth predicted procedural language solution in the second stage of reinforcement learning, and the second reward feedback corresponding to the fourth predicted natural language solution in the second stage of reinforcement learning; as well as The machine learning model is trained based on the first reward feedback and the second reward feedback.

9. The method of claim 8, wherein training the machine learning model based on the first reward feedback and the second reward feedback comprises: The machine learning model is adjusted to extract key information based on the first reward feedback. The machine learning model is adjusted to generate solutions in the programmed language based on the first reward feedback and the second reward feedback. as well as The machine learning model is adjusted to generate natural language answers based on the second reward feedback.

10. An apparatus for problem processing, comprising: The key information extraction module is configured to extract key information about the target problem using a machine learning model, wherein the machine learning model is based on a language model; The first solution determination module is configured to determine a programmatic language solution to the target problem based on the key information and the target problem, using the machine learning model through a program thinking chain. The programmatic language solution includes a description of the solution process to the target problem in a programmatic language. as well as The second solution determination module is configured to determine a natural language solution for the target question based on the key information, the target question, and the procedural language solution using the machine learning model through a natural language thought chain. The natural language solution includes at least the verification result of the procedural language solution.

11. An electronic device, comprising: At least one processing unit; as well as At least one memory, coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, which, when executed by the at least one processing unit, cause the electronic device to perform the method according to any one of claims 1 to 9.

12. A computer-readable storage medium having a computer program stored thereon, the computer program being executable by a processor to implement the method according to any one of claims 1 to 9.

13. A computer program product tangibly stored in a computer storage medium and comprising computer-executable instructions that, when executed by a device, cause the device to perform the method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Language processing question answering system and method based on AIGC large model

    CN118093834A

  • Data trend analysis method and device, equipment, storage medium and program product

    CN118227742A

  • Knowledge question-answering method based on natural language processing

    CN118673109A

  • Natural language query processing based on machine learning to perform a task

    US20240112074A1