Processing method and device of reasoning task, storage medium and electronic equipment
By generating a target reasoning model trained using adversarial reinforcement learning, the problem of low accuracy in reasoning models in the fintech field is solved, achieving higher reasoning accuracy and logical coherence.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- INDUSTRIAL AND COMMERCIAL BANK OF CHINA
- Filing Date
- 2026-02-09
- Publication Date
- 2026-05-29
AI Technical Summary
In the existing fintech field, inference models suffer from low inference accuracy when handling inference tasks.
A target reasoning model trained by generative adversarial reinforcement learning is adopted. By preprocessing the target reasoning task, a vector representation is generated, and reasoning is performed using an encoding layer, a reasoning layer, a decision fusion layer, and an output layer to generate logically coherent reasoning results.
It improves the accuracy and logical coherence of reasoning tasks and enhances the reliability of reasoning results.
Smart Images

Figure CN122114160A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of financial technology, and more specifically, to a method and apparatus for processing reasoning tasks, a storage medium, and an electronic device. Background Technology
[0002] To meet the increasingly complex analysis and decision-making needs of financial institutions, related technologies mainly employ statistical analysis and traditional machine learning models to handle reasoning tasks. These models are typically based on pattern recognition using fixed rules and historical data, over-relying on historical data and lacking complex reasoning capabilities, resulting in low reasoning accuracy when handling reasoning tasks.
[0003] There is currently no effective solution to the problem of low inference accuracy in inference models used in related technologies when handling inference tasks. Summary of the Invention
[0004] The main objective of this application is to provide a method, apparatus, storage medium, and electronic device for processing reasoning tasks, in order to solve the problem that reasoning models in related technologies have low reasoning accuracy when processing reasoning tasks.
[0005] To achieve the above objectives, according to one aspect of this application, a method for processing reasoning tasks is provided. The method includes: receiving a target reasoning task input from a target object; preprocessing the target reasoning task to obtain a vector representation corresponding to the target reasoning task; and performing reasoning processing based on the vector representation using a target reasoning model to obtain a reasoning result corresponding to the target reasoning task, wherein the target reasoning model is obtained by training an initial reasoning model and an initial discriminant model using generative adversarial reinforcement learning based on a sample reasoning task set.
[0006] Furthermore, the reasoning process based on the vector representation by the target reasoning model to obtain the reasoning result corresponding to the target reasoning task includes: performing semantic analysis on the vector representation through the encoding layer of the target reasoning model to obtain a semantic representation; performing step-by-step reasoning based on the semantic representation through the reasoning layer of the target reasoning model to obtain multiple reasoning steps; generating a target reasoning chain based on the multiple reasoning steps through the decision fusion layer of the target reasoning model, and determining the first reasoning result based on the target reasoning chain; and converting the format of the first reasoning result through the output layer of the target reasoning model to obtain the reasoning result.
[0007] Further, the target reasoning model is obtained through the following steps: obtaining a sample reasoning task set, wherein the sample reasoning task set includes multiple sample task questions and the true answer corresponding to each sample task question; for each sample task question, preprocessing the sample task question to obtain the vector representation corresponding to the sample task question; and performing generative adversarial reinforcement learning training on the initial reasoning model and the initial discriminant model based on the vector representation corresponding to the sample task question until the preset conditions are met, thereby obtaining the target reasoning model.
[0008] Furthermore, the initial inference model and the initial discriminant model are trained using generative adversarial reinforcement learning based on the vector representation corresponding to the sample task problem until preset conditions are met, resulting in the target inference model. This process includes: generating a problem-solving inference chain step by step using the initial inference model based on the vector representation corresponding to the sample task problem, wherein the problem-solving inference chain contains the initial answer corresponding to the sample task problem; splitting the problem-solving inference chain according to preset rules to obtain multiple inference slices, wherein each inference slice consists of a logically complete sequence of inference steps; evaluating each inference slice using the initial discriminant model to obtain the evaluation result corresponding to each inference slice; and training the initial inference model and the initial discriminant model using generative adversarial reinforcement learning based on the evaluation results corresponding to each inference slice until preset conditions are met, resulting in the target inference model.
[0009] Furthermore, based on the evaluation results corresponding to each reasoning slice, the initial reasoning model and the initial discriminant model are trained using generative adversarial reinforcement learning until preset conditions are met to obtain the target reasoning model. This includes: determining the reward score corresponding to each reasoning slice based on the evaluation results corresponding to each reasoning slice; determining the reward score corresponding to the initial answer and the reward score corresponding to the initial discriminant model based on the actual answer; determining the reward score corresponding to the initial reasoning model based on the reward score corresponding to each reasoning slice and the reward score corresponding to the initial answer; and iteratively training the initial reasoning model and the initial discriminant model based on the reward scores corresponding to the initial reasoning model and the initial discriminant model until preset conditions are met to obtain the target reasoning model.
[0010] Furthermore, based on the actual answer, determining the reward score corresponding to the initial answer and the reward score corresponding to the initial discrimination model includes: comparing the initial answer with the actual answer to obtain a first comparison result, and determining the reward score corresponding to the initial answer based on the first comparison result; comparing the evaluation result corresponding to each reasoning slice with the actual answer to obtain a second comparison result, and determining the reward score corresponding to the initial discrimination model based on the second comparison result.
[0011] Furthermore, determining the reward score for the initial reasoning model based on the reward score for each reasoning slice and the reward score for the initial answer involves: performing a weighted summation of the reward score for each reasoning slice and the reward score for the initial answer to obtain the reward score for the initial reasoning model.
[0012] To achieve the above objectives, according to another aspect of this application, a processing apparatus for a reasoning task is provided. The apparatus includes: a receiving unit for receiving a target reasoning task input by a target object; a first processing unit for preprocessing the target reasoning task to obtain a vector representation corresponding to the target reasoning task; and a second processing unit for performing reasoning processing based on the vector representation using a target reasoning model to obtain a reasoning result corresponding to the target reasoning task, wherein the target reasoning model is obtained by training an initial reasoning model and an initial discriminant model using generative adversarial reinforcement learning based on a sample reasoning task set.
[0013] Furthermore, the second processing unit includes: a first processing subunit, used to perform semantic analysis on the vector representation through the encoding layer of the target inference model to obtain a semantic representation; a second processing subunit, used to perform step-by-step inference based on the semantic representation through the inference layer of the target inference model to obtain multiple inference steps; a third processing subunit, used to generate a target inference chain based on the multiple inference steps through the decision fusion layer of the target inference model, and determine a first inference result based on the target inference chain; and a fourth processing subunit, used to perform format conversion on the first inference result through the output layer of the target inference model to obtain an inference result.
[0014] Furthermore, the device also includes the following units for obtaining the target inference model through the following steps: an acquisition unit for acquiring a sample inference task set, wherein the sample inference task set includes multiple sample task questions and the true answer corresponding to each sample task question; a third processing unit for preprocessing the sample task question for each sample task question to obtain the vector representation corresponding to the sample task question; and a fourth processing unit for performing generative adversarial reinforcement learning training on the initial inference model and the initial discriminant model based on the vector representation corresponding to the sample task question until a preset condition is met to obtain the target inference model.
[0015] Furthermore, the fourth processing unit includes: a fifth processing subunit, used to generate a problem-solving reasoning chain step by step based on the vector representation corresponding to the sample task problem using the initial reasoning model, wherein the problem-solving reasoning chain contains the initial answer corresponding to the sample task problem; a sixth processing subunit, used to split the problem-solving reasoning chain according to preset rules to obtain multiple reasoning slices, wherein each reasoning slice consists of a logically complete sequence of reasoning steps; a seventh processing subunit, used to evaluate each reasoning slice using the initial discriminant model to obtain the evaluation result corresponding to each reasoning slice; and an eighth processing subunit, used to perform generative adversarial reinforcement learning training on the initial reasoning model and the initial discriminant model based on the evaluation result corresponding to each reasoning slice, until the preset conditions are met to obtain the target reasoning model.
[0016] Furthermore, the eighth processing subunit includes: a first determining module, used to determine the reward score corresponding to each inference slice based on the evaluation result corresponding to each inference slice; a second determining module, used to determine the reward score corresponding to the initial answer and the reward score corresponding to the initial discriminant model based on the true answer; a third determining module, used to determine the reward score corresponding to the initial inference model based on the reward score corresponding to each inference slice and the reward score corresponding to the initial answer; and a first processing module, used to iteratively train the initial inference model and the initial discriminant model based on the reward scores corresponding to the initial inference model and the initial discriminant model until the preset conditions are met to obtain the target inference model.
[0017] Furthermore, the second determining module includes: a first determining submodule, used to compare the initial answer with the real answer to obtain a first comparison result, and determine the reward score corresponding to the initial answer based on the first comparison result; and a second determining submodule, used to compare the evaluation result corresponding to each reasoning slice with the real answer to obtain a second comparison result, and determine the reward score corresponding to the initial discrimination model based on the second comparison result.
[0018] Furthermore, the third determining module includes a third determining submodule, which is used to perform a weighted summation calculation on the reward score corresponding to each reasoning slice and the reward score corresponding to the initial answer to obtain the reward score corresponding to the initial reasoning model.
[0019] According to another aspect of the present invention, an electronic device is also provided, comprising: a memory storing an executable program; and a processor for running the program, wherein the program, when running, performs a processing method for a reasoning task as described above.
[0020] According to another aspect of the present invention, a computer-readable storage medium is also provided, wherein a processing method is provided for controlling the device where the storage medium is located to perform the reasoning task described above during program execution.
[0021] In this embodiment, the following steps are employed: receiving a target reasoning task input from a target object; preprocessing the target reasoning task to obtain a vector representation corresponding to the target reasoning task; and performing reasoning processing based on the vector representation using a target reasoning model to obtain the reasoning result corresponding to the target reasoning task. The target reasoning model is trained using generative adversarial reinforcement learning on an initial reasoning model and an initial discriminant model based on a sample reasoning task set. This solves the technical problem of low reasoning accuracy in related technologies when processing reasoning tasks. In this solution, by receiving target object input, preprocessing it into a vector representation, and using a target reasoning model trained with generative adversarial reinforcement learning for reasoning processing, the accuracy and logical coherence of reasoning are improved, and the reliability of the reasoning result is enhanced. Attached Figure Description
[0022] The accompanying drawings, which form part of this application, are used to provide a further understanding of this application. The illustrative embodiments and descriptions of this application are used to explain this application and do not constitute an undue limitation of this application. In the drawings:
[0023] Figure 1 A hardware block diagram of a computer terminal for implementing a processing method for reasoning tasks is shown.
[0024] Figure 2 This is a flowchart of a reasoning task processing method provided according to an embodiment of this application;
[0025] Figure 3 This is a schematic diagram of a processing apparatus for inference tasks provided according to an embodiment of this application;
[0026] Figure 4 This is a structural block diagram of an electronic device according to an embodiment of this application. Detailed Implementation
[0027] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.
[0028] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0029] It should be noted that the information collected in this application (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for display, data used for analysis, etc.) are information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, storage, use, processing, transmission, provision, disclosure, and application of this data all comply with relevant laws, regulations, and standards, necessary confidentiality measures have been taken, and they do not violate public order and good morals. Corresponding access points are provided for users to choose to authorize or refuse. For example, interfaces are set up between this system and relevant users or organizations, providing users with corresponding access points to choose to agree to or refuse automated decision-making results; if the user chooses to refuse, the process proceeds to the expert decision-making stage.
[0030] Example 1
[0031] According to an embodiment of this application, a method embodiment for processing inference tasks is also provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.
[0032] The method embodiment provided in Embodiment 1 of this application can be executed on a mobile terminal, computer terminal, or similar computing device. Figure 1 A hardware block diagram of a computer terminal (or mobile device) for implementing a processing method for reasoning tasks is shown. Figure 1As shown, the computer terminal 10 (or mobile device) may include one or more processors 102 (shown as 102a, 102b, ..., 102n in the figure) 102 (processor 102 may include, but is not limited to, a microprocessor MCU or a programmable logic device FPGA, etc.), a memory 104 for storing data, and a transmission device 106 for communication functions. In addition, it may also include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of a BUS bus), a network interface, a power supply, and / or a camera. Those skilled in the art will understand that... Figure 1 The structure shown is for illustrative purposes only and does not limit the structure of the aforementioned electronic device. For example, computer terminal 10 may also include... Figure 1 The more or fewer components shown, or having the same Figure 1 The different configurations shown.
[0033] It should be noted that the aforementioned one or more processors 102 and / or other data processing circuits are generally referred to herein as "data processing circuits". These data processing circuits may be embodied, in whole or in part, in software, hardware, firmware, or any other combination thereof. Furthermore, the data processing circuits may be a single, independent processing module, or may be integrated, in whole or in part, into any other element within the computer terminal 10 (or mobile device). As involved in the embodiments of this application, the data processing circuits serve as a processor control mechanism (e.g., selection of a variable resistor termination path connected to an interface).
[0034] The memory 104 can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the reasoning task processing method in the embodiments of this application. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, thereby realizing the above-mentioned reasoning task processing method. The memory 104 may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include memory remotely located relative to the processor 102, and these remote memories can be connected to the computer terminal 10 via a network. Examples of the above-mentioned networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.
[0035] The transmission device 106 is used to receive or send data via a network. Specific examples of the network described above may include a wireless network provided by the communication provider of the computer terminal 10. In one example, the transmission device 106 includes a Network Interface Controller (NIC), which can connect to other network devices via a base station to communicate with the Internet. In another example, the transmission device 106 may be a Radio Frequency (RF) module, used for wireless communication with the Internet.
[0036] The display may be a touchscreen liquid crystal display (LCD) that allows the user to interact with the user interface of the computer terminal 10 (or mobile device).
[0037] Under the aforementioned operating environment, this application provides the following: Figure 2 The method for handling reasoning tasks is shown. Figure 2 This is a flowchart of a reasoning task processing method according to Embodiment 1 of this application. The reasoning task processing method includes:
[0038] Step S201: Receive the target reasoning task input by the target object;
[0039] Step S202: Preprocess the target reasoning task to obtain the vector representation corresponding to the target reasoning task;
[0040] Step S203: The target reasoning model performs reasoning processing based on vector representation to obtain the reasoning result corresponding to the target reasoning task. The target reasoning model is obtained by generative adversarial reinforcement learning training on the initial reasoning model and the initial discriminant model based on the sample reasoning task set.
[0041] Optionally, the target inference task can be credit risk assessment, abnormal transaction detection, financial product recommendation, etc. Taking credit risk assessment as an example, the system receives the credit risk assessment task input by the user (such as an account manager), including the basic information of the entity to be assessed, credit history, debt situation, loan purpose, etc., to assess its loan default risk. Then, preprocessing is performed, for example, converting the basic information into structured data, the credit history into credit score, etc., and converting this data into vector representations for easy model processing. After receiving the vector representation, the target inference model (a large language model trained by generative adversarial reinforcement learning) begins to infer and assess the loan risk step by step, ultimately forming a coherent credit risk assessment report as the inference result.
[0042] In summary, this solution addresses the technical problem of low inference accuracy in inference models when handling inference tasks in related technologies. In this scheme, by receiving target object input, preprocessing it into a vector representation, and then utilizing a target inference model trained through generative adversarial reinforcement learning for inference processing, the accuracy and logical coherence of inference are improved, thus enhancing the reliability of the inference results.
[0043] Optionally, in the reasoning task processing method provided in this application embodiment, the reasoning process based on the vector representation by the target reasoning model to obtain the reasoning result corresponding to the target reasoning task includes: performing semantic analysis on the vector representation through the encoding layer of the target reasoning model to obtain a semantic representation; performing step-by-step reasoning based on the semantic representation through the reasoning layer of the target reasoning model to obtain multiple reasoning steps; generating a target reasoning chain based on the multiple reasoning steps through the decision fusion layer of the target reasoning model, and determining a first reasoning result based on the target reasoning chain; and converting the format of the first reasoning result through the output layer of the target reasoning model to obtain the reasoning result.
[0044] In an optional embodiment, semantic analysis is performed on the vector representation through the encoding layer of the target reasoning model to obtain a semantic representation. The encoding layer consists of one or more encoders, which are responsible for transforming the input vector representation into a deeper semantic representation. The encoders perform deep semantic analysis on the vector representation through components such as self-attention mechanisms and feedforward neural networks, capturing long-distance dependencies and features in the input, and outputting an initial hidden state vector (i.e., a semantic representation) rich in the problem context.
[0045] In an optional embodiment, the reasoning layer of the target reasoning model performs step-by-step reasoning based on semantic representation to obtain multiple reasoning steps. The reasoning layer consists of a series of identical or different generation layers. Each generation layer is responsible for generating a step in the reasoning chain. The reasoning process is started based on the initial hidden state vector output by the encoding layer. The process is repeated until the reasoning is completed. In each loop: the generation unit generates the description and conclusion of the current reasoning step, and the state update unit updates the internal state of the model based on the generated conclusion to provide input for the next reasoning step.
[0046] In an optional embodiment, the decision fusion layer of the target reasoning model generates a target reasoning chain based on multiple reasoning steps, and determines a first reasoning result based on the target reasoning chain. The decision fusion layer integrates all reasoning steps output by the reasoning layer to form a coherent reasoning chain, and extracts a preliminary reasoning result from it.
[0047] In an optional embodiment, the output layer of the target inference model performs format conversion on the first inference result to obtain the inference result. The output layer converts the inference result obtained from the decision fusion layer into natural language form, making it easier for users to understand and verify.
[0048] Through multi-layered processing within the model, complex reasoning tasks are decomposed and synthesized, ensuring the quality of the reasoning chain and the rationality and clarity of the final reasoning result.
[0049] Optionally, in the reasoning task processing method provided in this application embodiment, the target reasoning model is obtained through the following steps: obtaining a sample reasoning task set, wherein the sample reasoning task set includes multiple sample task questions and the true answer corresponding to each sample task question; for each sample task question, preprocessing the sample task question to obtain the vector representation corresponding to the sample task question; and performing generative adversarial reinforcement learning training on the initial reasoning model and the initial discriminant model based on the vector representation corresponding to the sample task question until the preset conditions are met to obtain the target reasoning model.
[0050] In an optional embodiment, a sample reasoning task set is first acquired. Then, for each sample task question, the sample task question is preprocessed to obtain a vector representation corresponding to the sample task question. Next, generative adversarial reinforcement learning is performed on the initial reasoning model and the initial discriminant model based on the vector representation corresponding to the sample task question until preset conditions are met, resulting in the target reasoning model. For example, the system acquires sample task questions as input, which can be mathematical problems, logical reasoning problems, or general question-and-answer questions. The input format can be natural language descriptions or structured data. The input questions are then preprocessed (e.g., format normalization, tokenization) to facilitate model processing. If the question contains background information or multimodal content (e.g., charts, formulas), it can be parsed and encoded. Optionally, for multimodal input, different encoding methods for different modalities can be specified. When the input is incomplete or the format does not meet expectations, an error can be returned and a request to resubmit the question can be made. If the question exceeds the system's capabilities, a prompt indicating that it cannot be processed can be given directly at this stage. Existing question parsing tools or formula parsing libraries for specific domains (e.g., mathematics) can be used to enhance the understanding and processing of the input.
[0051] Optionally, in the reasoning task processing method provided in this application embodiment, training the initial reasoning model and the initial discriminant model with generative adversarial reinforcement learning based on the vector representation corresponding to the sample task problem until a preset condition is met to obtain the target reasoning model includes: generating a problem-solving reasoning chain step by step based on the vector representation corresponding to the sample task problem using the initial reasoning model, wherein the problem-solving reasoning chain contains the initial answer corresponding to the sample task problem; splitting the problem-solving reasoning chain according to preset rules to obtain multiple reasoning slices, wherein each reasoning slice consists of a logically complete sequence of reasoning steps; evaluating each reasoning slice using the initial discriminant model to obtain an evaluation result corresponding to each reasoning slice; and training the initial reasoning model and the initial discriminant model with generative adversarial reinforcement learning based on the evaluation result corresponding to each reasoning slice until a preset condition is met to obtain the target reasoning model.
[0052] In an optional embodiment, a problem-solving reasoning chain is gradually generated based on the vector representation corresponding to the sample task problem using an initial inference model. For example, a large language model can be used as the initial inference model (i.e., the inferencer). Based on the vector representation corresponding to the sample task problem, the problem-solving reasoning chain is gradually generated. The initial inference model breaks down the problem into a series of intermediate inference steps, deriving an intermediate conclusion at each step, and then using that conclusion as the basis for the next step, until the final answer is derived. Optionally, the maximum number of inference steps or the maximum character length of each step can be limited to prevent the initial inference model from generating indefinitely. Parameters such as random temperature can also be set to diversify the inference process. If the initial inference model fails to arrive at an answer within a given number of steps, the inference chain can be truncated and marked as a non-convergent sample. If the model output is irrelevant to the problem, the inference process is re-initialized.
[0053] In an optional embodiment, the problem-solving reasoning chain is divided into several consecutive logical slices according to predefined rules (e.g., segmentation based on the number of steps, segmentation based on semantic boundaries, dynamic adjustment based on complexity, adaptive segmentation strategy, special labeling segmentation, etc.) for review by the initial discriminant model (i.e., the discriminator). Each slice contains a logically complete sequence of reasoning steps, and its length can be determined according to a predetermined strategy (e.g., each slice contains several steps, or segmentation is based on semantic boundaries). Slicing reduces evaluation overhead while ensuring that each segment is logically self-contained, facilitating the initial discriminant model's judgment of its correctness. Optionally, the slice length can be dynamically adjusted based on task complexity or adaptive segmentation can be used (e.g., based on breaks such as periods, reasoning stages, etc.). The review scheduling strategy may include selectively skipping obviously simple and correct steps to save computational resources. If the reasoning chain itself is very short (less than a threshold number of steps), no slicing is required or only one slice is generated. If the reasoning chain is too long and exceeds the processing capacity of the initial discriminant model, it can be truncated during slicing and the remaining part can be marked for subsequent special processing. For well-structured reasoning (e.g., numbered steps), slicing can be done by numbered segments. After slicing, you can append the position index of the segment in the overall inference chain to the beginning of each segment so that the initial discrimination model can refer to the context order.
[0054] In an optional embodiment, each reasoning slice is evaluated using an initial discriminant model to obtain an evaluation result for each slice. For each input reasoning fragment, the output of the initial discriminant model includes a binary judgment (i.e., whether the reasoning in the fragment is logically sound and error-free) and a structured reason (i.e., if a problem is found, the error is identified and briefly explained). For example, if the fragment contains mathematical calculations, the discriminator will verify the calculation process; if it involves common sense or theorem references, the discriminator will verify whether the references are appropriate. Along with the judgment, the model generates a concise explanation of the reason, which can be formatted for easy parsing. This process is repeated for each fragment in the slice set. Optionally, the discriminator can choose different levels of detail (brief or detailed) when generating the reason, and can also configure its inspection focus (e.g., emphasizing mathematical correctness or logical coherence). For successful fragments, the discriminator can also randomly provide positive feedback or repeat key steps to solidify the judgment. If the discriminator cannot determine the correctness of a fragment (e.g., beyond the scope of knowledge), it can be marked as uncertain. In this case, the system can choose to give the fragment a neutral reward or introduce manual review. When the discriminator's output does not meet the format requirements (e.g., the reason is too long or not given in the agreed format), the segment can be truncated or the evaluation can be retried. The initial discriminator model is a specially trained large language model, for example, pre-trained on a large amount of error-labeled reasoning chain data through reinforcement learning or supervised learning, making it adept at recognizing common reasoning errors. To improve accuracy, during the discrimination process, the discriminator can first generate its own internal chain of reasoning for the segment (but not feed it back to the inferencer) to more carefully check complex reasoning.
[0055] The structured reasoning output by the discriminator provides an evaluation basis for each reasoning step, which is equivalent to generating annotations and explanations of the process, improving the model's self-interpretability and decision transparency.
[0056] In an optional embodiment, based on the evaluation results corresponding to each inference slice, the initial inference model and the initial discriminant model are trained using generative adversarial reinforcement learning until preset conditions are met to obtain the target inference model. Applying the generator-discriminator adversarial collaborative training mechanism to the inference optimization of large language models can supervise the model's own inference; whenever an error occurs in the generation step, the internal discriminator points it out immediately. Through this self-adversarial game, the model develops the ability to automatically identify and correct its own inference errors, improving the reliability and consistency of the inference chain and reducing computational errors and unreasonable skips.
[0057] In an alternative embodiment, the discriminator is embedded as an independent module in the training framework and given a flexible and scalable configuration, enabling the framework to customize the discriminator evaluation criteria according to different application objectives. For example, expert problem-solving approaches can be incorporated into the discriminator to achieve teacher knowledge distillation, or the discriminator can be required to perform formal proof checks in the field of mathematics. This modular design enhances the system's versatility and customization capabilities.
[0058] Optionally, in the reasoning task processing method provided in this application embodiment, the initial reasoning model and the initial discriminant model are trained using generative adversarial reinforcement learning based on the evaluation results corresponding to each reasoning slice until a preset condition is met to obtain the target reasoning model. This includes: determining the reward score corresponding to each reasoning slice based on the evaluation results corresponding to each reasoning slice; determining the reward score corresponding to the initial answer and the reward score corresponding to the initial discriminant model based on the actual answer; determining the reward score corresponding to the initial reasoning model based on the reward score corresponding to each reasoning slice and the reward score corresponding to the initial answer; and iteratively training the initial reasoning model and the initial discriminant model based on the reward scores corresponding to the initial reasoning model and the initial discriminant model until a preset condition is met to obtain the target reasoning model.
[0059] Optionally, in the reasoning task processing method provided in this application embodiment, determining the reward score corresponding to the initial answer and the reward score corresponding to the initial discrimination model based on the real answer includes: comparing the initial answer with the real answer to obtain a first comparison result, and determining the reward score corresponding to the initial answer based on the first comparison result; comparing the evaluation result corresponding to each reasoning slice with the real answer to obtain a second comparison result, and determining the reward score corresponding to the initial discrimination model based on the second comparison result.
[0060] Optionally, in the reasoning task processing method provided in this application embodiment, determining the reward score corresponding to the initial reasoning model based on the reward score corresponding to each reasoning slice and the reward score corresponding to the initial answer includes: performing a weighted summation calculation on the reward score corresponding to each reasoning slice and the reward score corresponding to the initial answer to obtain the reward score corresponding to the initial reasoning model.
[0061] In an optional embodiment, a reward score is determined for each inference slice based on the evaluation result corresponding to each inference slice. For example, if the evaluation result is correct, the inferencer is given a positive reward (e.g., add 1 point), and if the evaluation result is incorrect, the inferencer is given a negative reward (e.g., deduct 1 point).
[0062] In an optional embodiment, based on the true answer, the reward score corresponding to the initial answer and the reward score corresponding to the initial discriminant model are determined separately. That is, the discriminator's evaluation result for each slice is combined with the correctness of the initial answer to form a reward signal for training. According to the output of the discriminator, a step-level reward is assigned to each slice: if the discriminator judges the slice correctly, a positive reward is given to the inferencer; if it judges incorrectly, a negative reward is given. The reward value can be adjusted according to the confidence level or error severity of the discriminator's reasoning. The initial answer in the problem-solving reasoning chain is compared with the preset correct answer (i.e., the true answer) of the sample task question to obtain a first comparison result (i.e., the answer is the same or the answer is different) to determine whether the initial answer generated by the inferencer is correct. Based on the first comparison result, the reward score corresponding to the initial answer is determined. For example, if the answers are the same, that is, the initial answer generated by the inferencer is correct, an additional global positive reward (sparse reward) is given; if the answers are different, that is, the initial answer generated by the inferencer is incorrect, a punitive negative reward is given.
[0063] In an optional embodiment, the evaluation result corresponding to each inference slice is compared with the true answer to obtain a second comparison result. Based on the second comparison result, the reward score corresponding to the initial discriminant model is determined. Specifically, for the discriminator, the discriminator's judgment result is compared with the preset correct answer to the sample task question to determine whether the discriminator's judgment is accurate and to give a corresponding reward. For example, if the discriminator successfully identifies an incorrect segment or correctly recognizes an error-free segment, it receives a positive reward; if the discriminator misjudges, it receives a negative reward, thus obtaining the discriminator's reward signal.
[0064] In an optional embodiment, the reward score for the initial inference model is determined based on the reward score for each inference slice and the reward score for the initial answer. This involves summing all step-level rewards and combining them with the initial answer reward to obtain the total reward score for the complete inference, which serves as the inferencer's reward signal. Optionally, the reward and penalty mechanisms can employ a weighted summation strategy to balance the importance of step rewards and global rewards. The reward score for the initial inference model is calculated by weighted summation of the reward score for each inference slice and the reward score for the initial answer. For example, 1 point is added for each correct step, 1 point is deducted for each incorrect step, and K points are added for a correct answer (K is set according to the task difficulty). Optionally, a discount factor can be introduced into the reward, with a smaller reward percentage for steps further from the answer, to encourage the inferencer to converge to the correct solution more quickly. For segments where the discriminator outputs uncertainties, no reward may be included or a zero score may be given to prevent noise from affecting training stability. Optionally, for extremely long inference chains, upper and lower limits can be set for the cumulative reward to prevent excessively large single training signals from affecting model stability. Optionally, reward variance can be reduced by employing a dominance function baseline method from reinforcement learning. For example, the difference between the inferencer's total reward and the rolling average baseline can be calculated as the actual dominance reward to improve training convergence stability. An independent reward smoothing mechanism can also be implemented for the discriminator to ensure that its evaluation criteria gradually converge to a consistent level.
[0065] In an optional embodiment, the initial inference model and the initial discriminator model are iteratively trained based on their reward scores until preset conditions are met, resulting in the target inference model. After obtaining the reward signal, reinforcement learning algorithms are used to simultaneously update the parameters of the inferencer and discriminator to improve their respective performance. For example, the policy gradient method is used to adjust the inferencer parameters, making them more likely to generate inference step sequences that yield high rewards; the reward signal obtained from the discriminator is used to optimize the discriminator parameters using policy gradient or a value network algorithm, making them more accurate in distinguishing between correct and incorrect inferences. Training is performed iteratively, with multiple rounds of alternating updates until convergence. Optionally, hyperparameters such as the learning rate and update batch size are set to balance convergence speed and stability. In adversarial training, a double buffering mechanism can be introduced, i.e., alternatingly freezing one model from updating to prevent oscillations caused by simultaneous and drastic updates of both models. If training becomes unstable (e.g., the rewards of both models do not improve for a long time or fluctuate drastically), the learning rate can be reduced or an entropy regularization term can be added to smooth the process. If an overly powerful discriminator leads to negative returns for the inference engine, the discriminator's reward weight can be reduced to maintain game balance. Optionally, every few rounds of adversarial training, a regular reinforcement learning update or human feedback based solely on the final answer can be introduced to calibrate the discriminator's criteria and prevent the accumulation of evaluation bias.
[0066] In an optional embodiment, after training, the inference model can be used independently for reasoning tasks, while the discriminator model can be used for online auxiliary interpretation and verification of reasoning results.
[0067] By introducing step-level feedback signals, the model can obtain learning information from multiple steps in each training iteration, reducing the number of trials and errors, improving training efficiency and sample utilization. Combined with the sparse reward of the final answer, the model is guided to ensure that each step of the reasoning process is reasonable while pursuing the correct answer, thus reducing the cost of continuous iterative optimization of the model.
[0068] The reasoning task processing method provided in this application includes the following steps: receiving a target reasoning task input from a target object; preprocessing the target reasoning task to obtain a vector representation corresponding to the target reasoning task; and performing reasoning processing based on the vector representation using a target reasoning model to obtain the reasoning result corresponding to the target reasoning task. The target reasoning model is obtained by training an initial reasoning model and an initial discriminant model using generative adversarial reinforcement learning on a sample reasoning task set. This solves the technical problem of low reasoning accuracy in related technologies when processing reasoning tasks. In this solution, by receiving target object input, preprocessing it into a vector representation, and using a target reasoning model trained with generative adversarial reinforcement learning for reasoning processing, the accuracy and logical coherence of reasoning are improved, and the reliability of the reasoning result is enhanced.
[0069] It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases the steps shown or described may be executed in a different order than that shown here.
[0070] Example 2
[0071] This application also provides a processing apparatus for reasoning tasks. It should be noted that the processing apparatus for reasoning tasks in this application can be used to execute the processing method for reasoning tasks provided in this application. The processing apparatus for reasoning tasks provided in this application will be described below.
[0072] According to an embodiment of this application, a processing apparatus for a reasoning task, such as a processing method for implementing the above-described reasoning task, is also provided. Figure 3 As shown, the device includes: a receiving unit 301, a first processing unit 302, and a second processing unit 303.
[0073] The receiving unit 301 is used to receive the target reasoning task input by the target object;
[0074] The first processing unit 302 is used to preprocess the target reasoning task to obtain the vector representation corresponding to the target reasoning task;
[0075] The second processing unit 303 is used to perform reasoning processing based on vector representation through the target reasoning model to obtain the reasoning result corresponding to the target reasoning task. The target reasoning model is obtained by performing generative adversarial reinforcement learning training on the initial reasoning model and the initial discriminant model based on the sample reasoning task set.
[0076] The reasoning task processing apparatus provided in this application embodiment receives a target reasoning task input by a target object through a receiving unit 301; a first processing unit 302 preprocesses the target reasoning task to obtain a vector representation corresponding to the target reasoning task; a second processing unit 303 performs reasoning processing based on the vector representation through a target reasoning model to obtain a reasoning result corresponding to the target reasoning task. The target reasoning model is obtained by performing generative adversarial reinforcement learning training on an initial reasoning model and an initial discriminant model based on a sample reasoning task set.
[0077] Optionally, in the processing apparatus for the reasoning task provided in the embodiments of this application, the second processing unit includes: a first processing subunit, used to perform semantic analysis on the vector representation through the encoding layer of the target reasoning model to obtain a semantic representation; a second processing subunit, used to perform step-by-step reasoning based on the semantic representation through the reasoning layer of the target reasoning model to obtain multiple reasoning steps; a third processing subunit, used to generate a target reasoning chain based on the multiple reasoning steps through the decision fusion layer of the target reasoning model, and determine a first reasoning result based on the target reasoning chain; and a fourth processing subunit, used to perform format conversion on the first reasoning result through the output layer of the target reasoning model to obtain a reasoning result.
[0078] Optionally, in the reasoning task processing apparatus provided in the embodiments of this application, the apparatus further includes the following units, used to obtain a target reasoning model through the following steps: an acquisition unit, used to acquire a sample reasoning task set, wherein the sample reasoning task set includes multiple sample task questions and the true answer corresponding to each sample task question; a third processing unit, used to preprocess the sample task question for each sample task question to obtain a vector representation corresponding to the sample task question; and a fourth processing unit, used to perform generative adversarial reinforcement learning training on the initial reasoning model and the initial discriminant model based on the vector representation corresponding to the sample task question, until a preset condition is met to obtain the target reasoning model.
[0079] Optionally, in the reasoning task processing apparatus provided in this application embodiment, the fourth processing unit includes: a fifth processing subunit, used to generate a problem-solving reasoning chain step by step based on the vector representation corresponding to the sample task problem using an initial reasoning model, wherein the problem-solving reasoning chain contains the initial answer corresponding to the sample task problem; a sixth processing subunit, used to split the problem-solving reasoning chain according to preset rules to obtain multiple reasoning slices, wherein each reasoning slice consists of a logically complete sequence of reasoning steps; a seventh processing subunit, used to evaluate each reasoning slice using an initial discriminant model to obtain an evaluation result corresponding to each reasoning slice; and an eighth processing subunit, used to perform generative adversarial reinforcement learning training on the initial reasoning model and the initial discriminant model based on the evaluation result corresponding to each reasoning slice, until a preset condition is met to obtain a target reasoning model.
[0080] Optionally, in the reasoning task processing apparatus provided in this application embodiment, the eighth processing subunit includes: a first determining module, used to determine the reward score corresponding to each reasoning slice based on the evaluation result corresponding to each reasoning slice; a second determining module, used to determine the reward score corresponding to the initial answer and the reward score corresponding to the initial discriminant model based on the true answer; a third determining module, used to determine the reward score corresponding to the initial reasoning model based on the reward score corresponding to each reasoning slice and the reward score corresponding to the initial answer; and a first processing module, used to iteratively train the initial reasoning model and the initial discriminant model based on the reward score corresponding to the initial reasoning model and the reward score corresponding to the initial discriminant model until a preset condition is met to obtain the target reasoning model.
[0081] Optionally, in the reasoning task processing apparatus provided in the embodiments of this application, the second determining module includes: a first determining submodule, used to compare the initial answer with the real answer to obtain a first comparison result, and determine the reward score corresponding to the initial answer based on the first comparison result; and a second determining submodule, used to compare the evaluation result corresponding to each reasoning slice with the real answer to obtain a second comparison result, and determine the reward score corresponding to the initial discrimination model based on the second comparison result.
[0082] Optionally, in the reasoning task processing apparatus provided in this application embodiment, the third determining module includes: a third determining submodule, used to perform a weighted summation calculation on the reward score corresponding to each reasoning slice and the reward score corresponding to the initial answer to obtain the reward score corresponding to the initial reasoning model.
[0083] It should be noted that the receiving unit 301, the first processing unit 302, and the second processing unit 303 mentioned above correspond to steps S201 to S203 in Embodiment 1. The three units and the corresponding steps implement the same instances and application scenarios, but are not limited to the content disclosed in Embodiment 1. It should be noted that the above modules or units can be hardware or software components stored in memory (e.g., memory 104) and processed by one or more processors (e.g., processors 102a, 102b, ..., 102n). The above units can also be part of a device and run in the computer terminal 10 provided in Embodiment 1.
[0084] Example 3
[0085] Embodiments of this application may provide an electronic device. Figure 4 This is a structural block diagram of an electronic device according to an embodiment of this application. Figure 4 As shown, the electronic device may include: one or more ( Figure 4 (Only one is shown) Processor 402, memory 404, memory controller, and peripheral interface, wherein the peripheral interface is connected to the radio frequency module, audio module and display.
[0086] The memory can be used to store software programs and modules, such as the program instructions / modules corresponding to the methods and apparatus in the embodiments of this application. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, thereby implementing the above-described methods. The memory may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory may further include memory remotely located relative to the processor, and these remote memories can be connected to the terminal via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.
[0087] The processor can call the information and application stored in the memory through the transmission device to perform the following steps: receiving the target inference task input by the target object; preprocessing the target inference task to obtain the vector representation corresponding to the target inference task; and performing inference processing based on the vector representation through the target inference model to obtain the inference result corresponding to the target inference task. The target inference model is obtained by generative adversarial reinforcement learning training on the initial inference model and the initial discriminant model based on the sample inference task set.
[0088] The processor can access information and applications stored in memory via a transmission device to execute the following steps: performing semantic analysis on the vector representation through the encoding layer of the target inference model to obtain a semantic representation; performing step-by-step inference based on the semantic representation through the inference layer of the target inference model to obtain multiple inference steps; generating a target inference chain based on the multiple inference steps through the decision fusion layer of the target inference model, and determining a first inference result based on the target inference chain; and converting the format of the first inference result through the output layer of the target inference model to obtain the inference result.
[0089] The processor can access the information and application programs stored in the memory via the transmission device to perform the following steps: acquiring a sample reasoning task set, wherein the sample reasoning task set includes multiple sample task questions and the true answer corresponding to each sample task question; for each sample task question, preprocessing the sample task question to obtain the vector representation corresponding to the sample task question; and training the initial reasoning model and the initial discriminant model using generative adversarial reinforcement learning based on the vector representation corresponding to the sample task question until the preset conditions are met to obtain the target reasoning model.
[0090] The processor can access information and applications stored in memory via a transmission device to execute the following steps: First, it generates a problem-solving reasoning chain step-by-step based on the vector representation corresponding to the sample task problem using an initial reasoning model. This reasoning chain contains the initial answer to the sample task problem. Second, it splits the problem-solving reasoning chain according to preset rules to obtain multiple reasoning slices. Each reasoning slice consists of a logically complete sequence of reasoning steps. Third, it evaluates each reasoning slice using an initial discriminant model to obtain an evaluation result for each slice. Fourth, it trains the initial reasoning model and the initial discriminant model using generative adversarial reinforcement learning based on the evaluation results for each reasoning slice, until preset conditions are met to obtain the target reasoning model.
[0091] The processor can access the information and application programs stored in the memory via the transmission device to execute the following steps: determining the reward score for each inference slice based on the evaluation results; determining the reward score for the initial answer and the reward score for the initial discriminant model based on the true answer; determining the reward score for the initial inference model based on the reward score for each inference slice and the reward score for the initial answer; and iteratively training the initial inference model and the initial discriminant model based on the reward scores for the initial inference model and the initial discriminant model until preset conditions are met to obtain the target inference model.
[0092] The processor can access the information and application programs stored in the memory via the transmission device to perform the following steps: compare the initial answer with the real answer to obtain a first comparison result, and determine the reward score corresponding to the initial answer based on the first comparison result; compare the evaluation result corresponding to each reasoning slice with the real answer to obtain a second comparison result, and determine the reward score corresponding to the initial discrimination model based on the second comparison result.
[0093] The processor can access the information and application stored in the memory via the transmission device to perform the following steps: calculate the weighted sum of the reward score corresponding to each reasoning slice and the reward score corresponding to the initial answer to obtain the reward score corresponding to the initial reasoning model.
[0094] Those skilled in the art will understand that Figure 4 The structure shown is for illustrative purposes only. Electronic devices can also be smartphones, tablets, handheld computers, mobile internet devices (MIDs), PADs, and other terminal devices. Figure 4 This does not limit the structure of the aforementioned electronic device. For example, electronic devices may also include components that are more... Figure 4 The more or fewer components shown (such as network interfaces, display devices, etc.), or having the same Figure 4 The different configurations shown.
[0095] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by a program instructing the hardware related to the terminal device. The program can be stored in a computer-readable storage medium, which may include: flash drive, read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.
[0096] Example 4
[0097] Embodiments of this application also provide a computer-readable storage medium. Optionally, in this embodiment, the storage medium can be used to store the program code executed by the reasoning task processing method provided in Embodiment 1.
[0098] Optionally, in this embodiment, the storage medium may be located in any computer terminal in a group of computer terminals in a computer network, or in any mobile terminal in a group of mobile terminals.
[0099] This application also provides a computer program product, which, when executed on a data processing device, is a program suitable for performing processing method steps for a reasoning task.
[0100] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.
[0101] In the above embodiments of this application, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0102] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual coupling, direct coupling, or communication connection may be through some interfaces; the indirect coupling or communication connection between units or modules may be electrical or other forms.
[0103] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0104] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0105] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard drive, magnetic disk, or optical disk.
[0106] The above description is only a preferred embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this application, and these improvements and modifications should also be considered within the scope of protection of this application.
Claims
1. A method for processing reasoning tasks, characterized in that, include: The target reasoning task is to receive input from the target object. The target reasoning task is preprocessed to obtain the vector representation corresponding to the target reasoning task; The inference result corresponding to the target inference task is obtained by performing inference processing based on the vector representation using the target inference model. The target inference model is obtained by performing generative adversarial reinforcement learning training on the initial inference model and the initial discriminant model based on the sample inference task set.
2. The method according to claim 1, characterized in that, The reasoning results corresponding to the target reasoning task are obtained by performing reasoning processing based on the vector representation using the target reasoning model, including: Semantic analysis is performed on the vector representation through the encoding layer of the target reasoning model to obtain a semantic representation; The reasoning layer of the target reasoning model performs step-by-step reasoning based on the semantic representation to obtain multiple reasoning steps; The decision fusion layer of the target reasoning model generates a target reasoning chain based on the multiple reasoning steps, and determines the first reasoning result based on the target reasoning chain. The first inference result is converted into a new format by the output layer of the target inference model to obtain the inference result.
3. The method according to claim 1, characterized in that, The target inference model is obtained through the following steps: Obtain the sample reasoning task set, wherein the sample reasoning task set includes multiple sample task questions and the true answer corresponding to each sample task question; For each sample task problem, the sample task problem is preprocessed to obtain the vector representation corresponding to the sample task problem; Generative adversarial reinforcement learning is performed on the initial inference model and the initial discriminant model based on the vector representation corresponding to the sample task problem until the preset conditions are met, thereby obtaining the target inference model.
4. The method according to claim 3, characterized in that, Based on the vector representation corresponding to the sample task problem, the initial inference model and the initial discriminant model are trained using generative adversarial reinforcement learning until preset conditions are met, resulting in the target inference model, which includes: The initial reasoning model generates a problem-solving reasoning chain step by step based on the vector representation corresponding to the sample task problem, wherein the problem-solving reasoning chain contains the initial answer corresponding to the sample task problem; The problem-solving reasoning chain is split according to preset rules to obtain multiple reasoning slices, wherein each reasoning slice consists of a logically complete sequence of reasoning steps; Each inference slice is evaluated using the initial discrimination model to obtain an evaluation result for each inference slice; Based on the evaluation results corresponding to each inference slice, the initial inference model and the initial discriminant model are trained using generative adversarial reinforcement learning until the preset conditions are met, thereby obtaining the target inference model.
5. The method according to claim 4, characterized in that, Based on the evaluation results corresponding to each inference slice, the initial inference model and the initial discriminant model are trained using generative adversarial reinforcement learning until preset conditions are met, resulting in the target inference model, which includes: Based on the evaluation results corresponding to each reasoning slice, the reward score corresponding to each reasoning slice is determined; Based on the actual answer, determine the reward score corresponding to the initial answer and the reward score corresponding to the initial discrimination model, respectively; Based on the reward score corresponding to each reasoning slice and the reward score corresponding to the initial answer, the reward score corresponding to the initial reasoning model is determined; Based on the reward scores corresponding to the initial inference model and the initial discrimination model, the initial inference model and the initial discrimination model are iteratively trained until the preset conditions are met, thereby obtaining the target inference model.
6. The method according to claim 5, characterized in that, Based on the actual answer, the reward score corresponding to the initial answer and the reward score corresponding to the initial discrimination model are determined respectively, including: The initial answer is compared with the actual answer to obtain a first comparison result, and the reward score corresponding to the initial answer is determined based on the first comparison result. The evaluation result corresponding to each reasoning slice is compared with the actual answer to obtain a second comparison result, and the reward score corresponding to the initial discrimination model is determined based on the second comparison result.
7. The method according to claim 5, characterized in that, Based on the reward score corresponding to each reasoning slice and the reward score corresponding to the initial answer, the reward score corresponding to the initial reasoning model is determined as follows: The reward score corresponding to each reasoning slice and the reward score corresponding to the initial answer are weighted and summed to obtain the reward score corresponding to the initial reasoning model.
8. A processing apparatus for a reasoning task, characterized in that, include: The receiving unit is used to receive the target inference task input from the target object. The first processing unit is used to preprocess the target reasoning task to obtain the vector representation corresponding to the target reasoning task; The second processing unit is used to perform reasoning processing based on the vector representation through the target reasoning model to obtain the reasoning result corresponding to the target reasoning task. The target reasoning model is obtained by performing generative adversarial reinforcement learning training on the initial reasoning model and the initial discriminant model based on the sample reasoning task set.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored executable program, wherein, when the executable program is executed, it controls the device containing the computer-readable storage medium to perform the processing method of the reasoning task according to any one of claims 1 to 7.
10. An electronic device, characterized in that, include: Memory, which stores executable programs; A processor for running the program, wherein the program, when running, performs the processing method for the reasoning task according to any one of claims 1 to 7.