Language model training method and device, equipment and storage medium
By dividing the language model into forward reasoning and backward verification modules, and conducting multi-round self-game training and logic distillation loops, the problem of low reasoning ability and accuracy of the language model is solved, and the stability and interpretability of the model in complex tasks are improved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING KNOWLEDGE ATLAS TECHNOLOGY CO LTD
- Filing Date
- 2025-12-03
- Publication Date
- 2026-04-24
AI Technical Summary
Existing language models have low reasoning ability and accuracy, and lack logical consistency and interpretability.
The language model is divided into a forward reasoning module and a backward verification module. Through multi-round self-game training and logical distillation loop, logical consistency is strengthened, including the verification of logical constraint equations and backward reasoning equations, self-correction mechanism and calculation of logical consistency score.
It improves the reasoning ability and accuracy of the language model, enhances the stability and interpretability of the model in complex tasks, and achieves adaptive optimization and continuous learning at the logical level.
Smart Images

Figure CN121920559A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of artificial intelligence technology, and in particular to a language model training method, apparatus, device, and storage medium. Background Technology
[0002] As the scale of large language models (LLMs) and multimodal language models (MLLMs) continues to expand, reasoning performance has become a key factor affecting the model's understanding ability, logical consistency, and intelligent interpretability.
[0003] Existing reinforcement reasoning learning and self-game optimization stages mostly focus on semantic layer rewriting and surface consistency optimization.
[0004] Currently, language models have low reasoning ability and low accuracy. Summary of the Invention
[0005] This disclosure provides a language model training method, apparatus, device, and storage medium to at least solve the problems of low reasoning ability and low accuracy of existing language models.
[0006] The technical solution disclosed herein is as follows: This disclosure provides a language model training method, including: The initial language model is divided into a forward reasoning module and a reverse verification module; The forward reasoning module generates a forward reasoning chain and a forward reasoning conclusion based on the input text; The reverse verification module, based on the forward reasoning conclusion and with the forward reasoning conclusion as the target, derives the reverse reasoning chain and the reverse output set in reverse. Based on the forward inference chain and the backward inference chain, the initial language model is subjected to multiple rounds of self-game training and logical distillation loops to obtain a trained target language model. The target language model is capable of performing language understanding on the input target language text to obtain the user's requirement text.
[0007] Optionally, the forward reasoning chain includes: a set of key variables and logical constraint equations. The forward reasoning module generates a forward reasoning chain and a forward reasoning conclusion based on the input text, including: The forward reasoning module generates the set of key variables, the logical constraint equations, and the forward reasoning conclusion based on the input text, the model's internal knowledge, and the parameter distribution.
[0008] Optionally, the reverse inference chain includes: a reverse inference equation; the reverse verification module, based on the forward inference conclusion and with the forward inference conclusion as the target, derives the reverse inference chain and the reverse output set in reverse, including: Based on the forward reasoning conclusion, the reverse verification module uses the forward reasoning conclusion as the target to deduce the reverse reasoning equation and the reverse output set.
[0009] Optionally, the forward inference chain includes: a logical constraint equation; the backward inference chain includes: a backward inference equation; the step of performing multi-round self-game training and logical distillation loops on the initial language model based on the forward inference chain and the backward inference chain to obtain the trained target language model includes: The logical consistency of the logical constraint equation and the reverse reasoning equation is verified to obtain a logical consistency score. Based on the logical consistency score and logical consistency threshold, the initial language model is subjected to multiple rounds of self-game training and logical distillation loops to obtain the trained target language model.
[0010] Optionally, the forward inference chain further includes: a set of key variables; the step of verifying the logical consistency of the logical constraint equation and the reverse inference equation to obtain a logical consistency score includes: Calculate the variable consistency score based on the set of key variables and the set of reverse outputs; Calculate the equation consistency score based on the logical constraint equation and the reverse reasoning equation; Calculate the logical consistency score based on the variable consistency score and the equation consistency score.
[0011] Optionally, based on the logical consistency score and logical consistency threshold, the initial language model undergoes multiple rounds of self-game training and logical distillation loops to obtain the trained target language model, including: If the logical consistency score is less than the logical consistency threshold, then a repair action is performed on the inconsistent reasoning path; The initial language model is subjected to multiple rounds of self-game training and logical distillation loops to obtain the trained target language model.
[0012] Optionally, the performing the repair action includes at least one of the following: Based on logical feedback, identify inconsistent variables or equation differences; input these inconsistent variables or equation differences into the subsequent reasoning process to guide the initial language model to reason again. The total loss function is determined based on the logical consistency loss and the task loss; the initial language model is then trained based on the total loss function. The initial language model is trained using logical closed-loop samples; wherein, the logical closed-loop samples are the samples corresponding to the logical consistency score being greater than or equal to the logical consistency threshold.
[0013] This disclosure also provides a language model training device, including: The partitioning module is used to divide the initial language model into a forward reasoning module and a reverse verification module; The generation module is used by the forward reasoning module to generate a forward reasoning chain and a forward reasoning conclusion based on the input text. The derivation module is used by the reverse verification module to derive the reverse reasoning chain and the reverse output set based on the forward reasoning conclusion and with the forward reasoning conclusion as the target. The training module is used to perform multiple rounds of self-game training and logical distillation loops on the initial language model according to the forward inference chain and the backward inference chain to obtain the trained target language model, wherein the target language model is capable of language understanding of the input target language text to obtain the user requirement text.
[0014] This disclosure also provides an electronic device, including: processor; Memory used to store processor-executable instructions; The processor is configured to execute instructions to implement the steps in the above method.
[0015] This disclosure also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the above-described method.
[0016] The technical solutions provided by the embodiments of this disclosure have at least the following beneficial effects: In some embodiments of this disclosure, the initial language model is divided into a forward reasoning module and a reverse verification module. The forward reasoning module generates a forward reasoning chain and a forward reasoning conclusion based on the input text. The reverse verification module, based on the forward reasoning conclusion and with the forward reasoning conclusion as the target, derives a reverse reasoning chain and a reverse output set. Based on the forward reasoning chain and the reverse reasoning chain, the initial language model undergoes multiple rounds of self-game training and logical distillation loops to obtain a trained target language model. The target language model is capable of language understanding of the input target language text to obtain the user's required text. This disclosure improves the reasoning ability of the language model and enhances its accuracy by rewriting and strengthening the logic layer from the existing semantic layer.
[0017] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description
[0018] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure, and are not intended to unduly limit this disclosure.
[0019] Figure 1 A flowchart illustrating a language model training method provided for an exemplary embodiment of this disclosure; Figure 2 A schematic diagram of the structure of a language model training device provided as an exemplary embodiment of this disclosure; Figure 3 A schematic diagram of the structure of an electronic device provided for an exemplary embodiment of this disclosure. Detailed Implementation
[0020] To enable those skilled in the art to better understand the technical solutions of this disclosure, the technical solutions in the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings.
[0021] It should be noted that the terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this disclosure are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this disclosure described herein can be implemented in orders other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this disclosure.
[0022] It should be noted that the user information involved in this disclosure includes, but is not limited to, user device information and user personal information; the collection, storage, use, processing, transmission, provision and disclosure of user information in this disclosure all comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0023] To address the aforementioned technical problems, in some embodiments of this disclosure, the initial language model is divided into a forward reasoning module and a reverse verification module. The forward reasoning module generates a forward reasoning chain and a forward reasoning conclusion based on the input text. The reverse verification module, based on the forward reasoning conclusion and using it as the target, derives a reverse reasoning chain and a reverse output set. Based on the forward and reverse reasoning chains, the initial language model undergoes multiple rounds of self-game training and logical distillation loops to obtain a trained target language model. The target language model is capable of performing language understanding on the input target language text to obtain the user's required text. This disclosure improves the language model's reasoning ability and accuracy by enhancing it from the existing semantic layer rewriting to the logical layer.
[0024] The technical solutions provided by the embodiments of this disclosure are described in detail below with reference to the accompanying drawings.
[0025] Figure 1 This is a flowchart illustrating a language model training method provided as an exemplary embodiment of the present disclosure. Figure 1 As shown, the method includes: S101: Divide the initial language model into a forward reasoning module and a reverse verification module; S102: The forward reasoning module generates a forward reasoning chain and a forward reasoning conclusion based on the input text; S103: The reverse verification module, based on the conclusion of the forward reasoning, uses the conclusion of the forward reasoning as the target to deduce the reverse reasoning chain and the reverse output set; S104: Based on the forward and backward reasoning chains, the initial language model is subjected to multiple rounds of self-game training and logical distillation loops to obtain the trained target language model. The target language model is capable of language understanding of the input target language text to obtain the user's requirement text.
[0026] In this embodiment, the subject executing the above method is a terminal device or a server.
[0027] The terminal device includes, but is not limited to, mobile stations (MS), mobile terminals, mobile phones, handsets, and portable equipment. This terminal device can communicate with one or more core networks via a radio access network (RAN). For example, the terminal device can be a mobile phone (or "cellular" phone), a computer with wireless communication capabilities, a computer with wireless transceiver capabilities, a virtual reality (VR) terminal device, an AR terminal device, a wireless terminal in industrial control, a wireless terminal in self-driving, a wireless terminal in remote medical care, a wireless terminal in a smart grid, a wireless terminal in transportation safety, a wireless terminal in a smart city, a wireless terminal in a smart home, etc. The operating systems installed on the terminal device include, but are not limited to, iOS, Android, Windows, Linux, and Mac OS. In different networks, terminals may be called by different names, such as: user equipment, mobile station, user unit, station, cellular phone, personal digital assistant, wireless modem, wireless communication device, handheld device, laptop, cordless phone, wireless local loop station, television, etc. For ease of description, this embodiment will simply refer to it as terminal device.
[0028] In this embodiment, the implementation form of the server is not limited. For example, the server can be a conventional server, a cloud server, a cloud host, a virtual center, or other server devices. The server mainly consists of a processor, hard disk, memory, system bus, and other common computer architecture types.
[0029] To provide a detailed and clear explanation of this disclosure, the following explanations are provided for the terms used in this disclosure.
[0030] Logical consistency reinforcement: This refers to the process by which the model corrects and reinforces the reasoning path by detecting the consistency of logical constraints (such as variable conservation and isomorphism of equation structure) between forward and backward reasoning during the reasoning process, thereby achieving adaptive optimization and stable convergence at the logical level.
[0031] Self-game learning refers to a model that internally divides itself into two roles: Forward Reasoner (FR) and Backward Verifier (BV). Through bidirectional logical verification and result mutual verification, a self-adversarial and self-correcting learning mechanism is formed to strengthen the model's reasoning closed-loop capability.
[0032] Consistency of key variable set: This refers to the system automatically extracting the core variable set (such as {u, a, t, v}) in forward and backward reasoning, and calculating the variable matching degree and tolerance difference. If they are completely consistent, the logical path structure is determined to be stable; otherwise, self-correction is triggered.
[0033] Constraint equation consistency verification: refers to the verification of the system of equations during the analytical forward and backward reasoning processes. , The symbolic normalization algorithm is used to determine whether the equations are algebraically isomorphic or equivalent, which is used to verify the structural consistency of the logical reasoning chain.
[0034] Logical Integrity Coefficient (LIC): This index comprehensively evaluates the logical consistency of the reasoning chain, and is a weighted fusion of variable consistency score and equation consistency score.
[0035] when When the system determines that the reasoning loop is successfully closed; when When this occurs, a logical self-correction mechanism is triggered.
[0036] Self-correction mechanism: When the system detects logical inconsistencies, it automatically executes gradient backpropagation, logical feedback injection, and re-inference processes to update model parameters to correct logical deviations, thus forming a self-evolving continuous learning capability.
[0037] Consistency distillation: refers to retraining or knowledge distillation using logically closed-loop samples (LIC≥τ) as high-confidence data to enhance the joint consistency of the model at the logical and semantic layers.
[0038] Logical perturbation refers to operations such as rearranging conditions, rewriting chains, and reverse reasoning in the reasoning process, which are used to test the stability and self-consistency of the model's reasoning under different logical perspectives.
[0039] Logical Loss Function: Defined as It is used to transform logical consistency signals into differentiable optimization targets, enabling end-to-end training and parameter updates of the logic layer.
[0040] This disclosure achieves a paradigm shift from semantic-level rewriting to logical-level self-evolutionary learning by constructing a closed-loop optimization framework of "forward reasoning - reverse verification - logical verification - consistency reinforcement" within the model. This improves the stability, interpretability, and generalization ability of the language model in complex reasoning, conditional back reasoning, and multi-step reasoning tasks.
[0041] In some embodiments of this disclosure, the initial language model is divided into a forward reasoning module and a reverse verification module. The present invention first initializes and configures the language model, dividing it into two collaborative sub-modules: a forward reasoning module, responsible for performing forward logical reasoning based on input conditions, generating a reasoning chain and a final conclusion; and a reverse verification module, responsible for performing reverse reasoning starting from the conclusion, verifying the reversibility and self-consistency of the logical chain. During the initialization phase, pre-trained model parameters are loaded, and a logical consistency threshold is set. and consistency weight parameters This provides a foundation for subsequent logical consistency calculations and reinforcement learning.
[0042] This disclosure introduces a logic-layer self-game mechanism, dividing the language model into two collaborative roles: a forward inferrer and a reverse verifier. These two roles alternately engage in a game and mutually verify each other within the inference space. In each round of the game, the system calculates a logical consistency score and automatically determines the self-consistency of the inference chain based on the score. When logical inconsistencies are detected, a self-correction mechanism is triggered, repairing logical deviations through re-inference and gradient backpropagation. This closed-loop design achieves a transformation from "unidirectional generation" to "bidirectional logical convergence," enabling the model to possess self-verification and self-correction capabilities, significantly improving inference stability and logical credibility in complex tasks.
[0043] In some embodiments of this disclosure, the forward inference module generates a forward inference chain and a forward inference conclusion based on the input text. One possible implementation is that the forward inference module generates a set of key variables, a set of logical constraint equations, and a forward inference conclusion based on the input text, internal model knowledge, and parameter distribution. The input text can be an input task or a problem sample. The forward inference chain includes the set of key variables and the logical constraint equations. Specifically, the forward inference agent (FR) receives the input task or problem sample. Based on the model's internal knowledge and parameter distribution, as well as the autoregressive generation mechanism, a forward inference chain is generated. Conclusions of forward reasoning The forward reasoning chain includes intermediate logical steps, conditional transformations, and equation relationships, used to characterize the reasoning path structure of the model. The forward reasoning module outputs: the forward reasoning conclusion. The set of key variables involved in the reasoning Logical constraint equations .
[0044] In some embodiments of this disclosure, the reverse inference chain includes a reverse inference equation. The reverse verification module, based on the forward inference conclusion and targeting the forward inference conclusion, derives the reverse inference chain and the reverse output set. One possible implementation is that the reverse verification module, based on the forward inference conclusion and targeting the forward inference conclusion, derives the reverse inference equation and the reverse output set. The reverse verifier (BV) receives the forward inference conclusion. Using this as the target, we can deduce the possible preconditions and intermediate processes in reverse, thus generating a reverse reasoning chain. With the reverse output set Extract the reverse reasoning equation in this step. This is then structured and aligned with the logical constraint equations to provide input for subsequent logical consistency checks. This process forms a reverse reasoning path of "conclusion → premise," which, together with forward reasoning, constitutes a self-game framework at the logical level.
[0045] In some embodiments of this disclosure, the initial language model is subjected to multiple rounds of self-game training and logical distillation cycles based on the forward and backward inference chains to obtain the trained target language model. One possible approach is to perform logical consistency verification on the logical constraint equations and backward inference equations to obtain a logical consistency score; based on the logical consistency score and logical consistency threshold, the initial language model is subjected to multiple rounds of self-game training and logical distillation cycles to obtain the trained target language model.
[0046] This disclosure strengthens the structural rationality of reasoning through dual constraints of key variable set consistency and equation consistency. Traditional semantic layer optimization methods often only focus on the correctness of the result or text similarity, neglecting the consistency of variable relationships and equation constraints during the reasoning process, leading to the problem of "semantic correct but logically incorrect". This disclosure uses a key variable set consistency detection mechanism and a constraint equation consistency verification mechanism. The former is used to detect the core variable set (such as...) in forward and backward reasoning. Whether they are consistent; the latter is used to determine whether the corresponding equations of the two are isomorphic in algebraic structure (e.g. and Through these two layers of constraints, this disclosure establishes feedback signals simultaneously on numerical consistency and logical structure consistency, ensuring that the model's reasoning path conforms to logical rules, and fundamentally improving the interpretability and verifiability of the reasoning.
[0047] In the above embodiments, logical consistency verification is performed on the logical constraint equation and the reverse inference equation to obtain a logical consistency score. One possible approach is to calculate the variable consistency score based on the key variable set and the reverse output set; calculate the equation consistency score based on the logical constraint equation and the reverse inference equation; and calculate the logical consistency score based on the variable consistency score and the equation consistency score.
[0048] Specifically, the system performs a two-layer logical consistency check on the results of both forward and reverse reasoning: (1) Consistency check of key variable set: The system compares the sets of key variables. and Calculate the variable consistency score based on whether the corresponding variable values are consistent within the tolerance range. .
[0049] (2) Consistency verification of constraint equations: The system... (Logical constraint equations) and The (reverse reasoning equation) is symbolized and normalized. If the two equations are algebraically isomorphic, the logical structure is considered consistent, and the equation consistency score is calculated. .
[0050] The formula for calculating the logical consistency score is: ; in, Consistency weight parameter of key variable set, Consistency weight parameters for constraint equations.
[0051] In the above embodiments, based on the logical consistency score and logical consistency threshold, the initial language model undergoes multiple rounds of self-game training and logical distillation loops to obtain the trained target language model. One possible approach is to perform a repair action on the inconsistent reasoning path if the logical consistency score is less than the logical consistency threshold; then, the initial language model undergoes multiple rounds of self-game training and logical distillation loops to obtain the trained target language model. If the logical consistency score is greater than or equal to the logical consistency threshold, the logical loop of the reasoning chain is considered successfully closed.
[0052] Specifically, when If the logical loop of the reasoning chain is successfully closed, then the chain is considered closed. When this occurs, the system triggers a self-correction mechanism to perform repair actions on inconsistent inference paths. Specifically, This is the logical consistency threshold.
[0053] It should be noted that the logical consistency threshold The τ value can be adaptively set according to model accuracy and task type. When the τ value is high (e.g., 0.98–1.0), the system achieves higher logical purity in the distilled samples; when the τ value is low (e.g., 0.90–0.95), the system can obtain a larger sample coverage. This adjustable mechanism ensures the quality of logically closed-loop samples while improving the system's scalability and training efficiency in multi-task environments.
[0054] In some embodiments of this disclosure, remedial actions are performed on inconsistent inference paths. These include, but are not limited to, the following remedial methods: Repair Method 1: Logical Feedback Injection. Based on logical feedback prompts, inconsistent variables or equation differences are identified; these are then input into subsequent reasoning processes, guiding the initial language model to re-reason. Specifically, the system generates logical feedback prompts, indicating inconsistent variables or equation differences, and injects them into the next round of reasoning input, guiding the model to re-reason. For example, in the reasoning process of solving a circuit problem, the user inputs: "A 12V power supply is connected in series with a 3Ω and a 6Ω resistor. Find the total current." The initial language model outputs: "Total resistance is 3+6=9Ω, current I=12 / 9≈1.33A," but incorrectly treats the two resistors as being in series. The system verifies through circuit topology rules that if the problem implicitly contains a "parallel" structure, the correct total resistance should be Req=(3×6) / (3+6)=2Ω, and the current should be 6A. Based on this, the system generates a logical feedback prompt: "You assumed the two resistors were connected in series, but according to circuit connection rules, they are actually connected in parallel. Please recalculate the equivalent resistance and total current." This feedback is injected into the next round of inference input, guiding the model to be corrected to: "Two resistors are connected in parallel, the equivalent resistance is 2Ω, and the total current I = 12V / 2Ω = 6A," thereby achieving accurate re-inference based on inconsistent variables (current values) and equation differences (misuse of series / parallel formulas).
[0055] Method 2: Gradient-based Correction. The total loss function is determined based on the logical consistency loss and the task loss; the initial language model is then trained using this total loss function. Specifically, the logical consistency loss... With mission loss Together, they form the total loss function: ; Specifically, the model parameters are automatically updated through gradient backpropagation, causing them to gradually converge toward logical consistency.
[0056] Repair Method 3: Logical Consistency Sample Distillation. The initial language model is trained using logically closed-loop samples; these samples are those whose logical consistency scores are greater than or equal to the logical consistency threshold. Specifically, all LICs... The samples identified as logically closed-loop samples are aggregated into a high-quality dataset for retraining or distillation learning. Through multiple rounds of self-game and distillation iterations, the model continuously strengthens its logical reasoning patterns and structural stability, achieving continuous evolution of the logical layer.
[0057] This disclosure achieves adaptive optimization of inference performance through a logical consistency reinforcement learning mechanism. Existing inference optimization mechanisms often lack learnable logical feedback signals and cannot directly incorporate logical consistency into the parameter update process. This disclosure constructs a reinforcement learning mechanism based on the logical consistency index (LIC), transforming the consistency score into a differentiable loss function. And jointly optimize with task loss:
[0058] When the LIC falls below the logical consistency threshold τ, the system automatically triggers a self-correction mechanism to regenerate inconsistent inference chains and update parameters, thereby achieving continuous optimization and adaptive convergence of the logic layer during training. This mechanism enables the model to not only focus on the output results during inference but also actively learn how to maintain the stability of the logical structure, thus achieving long-term performance evolution and continuous improvement.
[0059] Through multiple rounds of self-game training and logical distillation cycles, the model gradually develops logical self-correction capabilities. Ultimately, the system achieves simultaneous improvement in logical consistency and reasoning performance, enabling the language model to exhibit higher accuracy, stability, and interpretability in complex reasoning, reverse thinking, and multi-step logical tasks.
[0060] In summary, this disclosure enables language models to automatically detect logical errors, repair inference chains, and continuously strengthen logical structures without external annotations through a closed-loop mechanism of "logic layer self-game → consistency detection → self-correction feedback → data distillation optimization," thereby improving the inference performance and generalization ability of language models.
[0061] This disclosure enhances the generalization and robustness of the model through a multi-layered repair and self-distillation mechanism. Traditional models often rely solely on manual correction or external retraining after inference errors occur, lacking internal self-correction capabilities. This disclosure constructs a three-layered self-correction and distillation mechanism based on logical layer consistency feedback: Logical feedback layer: When inference is inconsistent, the system automatically generates logical prompts to guide the model to re-infer; Parameter repair layer: Updates parameters by backpropagating the logistic loss, enabling the model to converge toward logical consistency; Data distillation layer: This layer separates logically consistent samples (LIC) into a single layer. As high-quality data, it is redistilled to continuously optimize the model distribution.
[0062] This multi-layered repair structure enables the model to maintain stable performance when faced with different tasks or logical disturbances, improving its cross-task generalization ability and resistance to logical bias.
[0063] Figure 2 This is a schematic diagram of the structure of a language model training device 20 provided for an exemplary embodiment of this disclosure. (See diagram below.) Figure 2 As shown, the language model training device 20 includes: a partitioning module 21, a generation module 22, a derivation module 23, and a training module 24.
[0064] Among them, the partitioning module 21 is used to divide the initial language model into a forward reasoning module and a reverse verification module; The generation module 22 is used by the forward reasoning module to generate a forward reasoning chain and a forward reasoning conclusion based on the input text. Derivation module 23 is used by the reverse verification module to deduce the reverse reasoning chain and reverse output set based on the forward reasoning conclusion and with the forward reasoning conclusion as the target. Training module 24 is used to perform multiple rounds of self-game training and logical distillation loops on the initial language model according to the forward and backward inference chains to obtain the trained target language model. The target language model is able to perform language understanding on the input target language text to obtain the user's requirement text.
[0065] Optionally, the forward reasoning chain includes: a set of key variables and logical constraint equations. The generation module 22, when the forward reasoning module generates the forward reasoning chain and forward reasoning conclusion based on the input text, is used for: The forward reasoning module generates a set of key variables, logical constraint equations, and forward reasoning conclusions based on the input text, the model's internal knowledge, and parameter distribution.
[0066] Optionally, the reverse inference chain includes: a reverse inference equation; the derivation module 23 is used by the reverse verification module to derive the reverse inference chain and the reverse output set based on the forward inference conclusion and with the forward inference conclusion as the target: The reverse verification module, based on the conclusions of the forward reasoning, uses those conclusions as its target to derive the reverse reasoning equations and the reverse output set.
[0067] Optionally, the forward inference chain includes: a logical constraint equation; the backward inference chain includes: a backward inference equation; based on the forward and backward inference chains, the training module 24, when performing multi-round self-game training and logical distillation loops on the initial language model to obtain the trained target language model, is used for: Logical consistency is verified on the logical constraint equations and reverse reasoning equations to obtain a logical consistency score. Based on the logical consistency score and logical consistency threshold, the initial language model is subjected to multiple rounds of self-game training and logical distillation loops to obtain the trained target language model.
[0068] Optionally, the forward inference chain also includes: a set of key variables; the training module 24, when performing logical consistency verification on the logical constraint equation and the backward inference equation to obtain a logical consistency score, is used for: Calculate the variable consistency score based on the set of key variables and the reverse output set; Calculate the equation consistency score based on the logical constraint equation and the reverse reasoning equation. Calculate the logical consistency score based on the variable consistency score and the equation consistency score.
[0069] Optionally, when training module 24 performs multiple rounds of self-game training and logic distillation loops on the initial language model based on the logic consistency score and logic consistency threshold to obtain the trained target language model, it is used for: If the logical consistency score is less than the logical consistency threshold, then a repair action is performed on the inconsistent reasoning path; The initial language model is subjected to multiple rounds of self-game training and logical distillation loops to obtain the trained target language model.
[0070] Optionally, a repair action is performed, including at least one of the following: Based on logical feedback, identify inconsistent variables or equation differences; input the inconsistent variables or equation differences into the subsequent reasoning process to guide the initial language model to reason again. The total loss function is determined based on the logical consistency loss and the task loss; the initial language model is then trained based on the total loss function. The initial language model is trained using logical closed-loop samples; where logical closed-loop samples are those corresponding to logical consistency scores greater than or equal to the logical consistency threshold.
[0071] Regarding the apparatus in the above embodiments, the specific manner in which each module performs its operation has been described in detail in the embodiments related to the method, and will not be elaborated upon here.
[0072] Figure 3 This is a schematic diagram of the structure of an electronic device provided as an exemplary embodiment of the present disclosure. For example... Figure 3 As shown, the electronic device includes a memory 31 and a processor 32. Additionally, the electronic device also includes a power supply component 33 and a communication component 34.
[0073] Memory 31 is used to store computer programs and can be configured to store various other data to support operation on the electronic device. Examples of this data include instructions for any application or method used to operate on the electronic device.
[0074] The memory 31 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk or optical disk.
[0075] Communication component 34 is used for data transmission with other devices.
[0076] The processor 32 executes computer instructions stored in the memory 31 to: divide the initial language model into a forward reasoning module and a reverse verification module; the forward reasoning module generates a forward reasoning chain and a forward reasoning conclusion based on the input text; the reverse verification module, based on the forward reasoning conclusion and with the forward reasoning conclusion as the target, derives a reverse reasoning chain and a reverse output set; and performs multi-round self-game training and logical distillation loops on the initial language model based on the forward reasoning chain and the reverse reasoning chain to obtain a trained target language model, wherein the target language model is capable of language understanding of the input target language text to obtain the user's required text.
[0077] Accordingly, embodiments of this disclosure also provide a computer-readable storage medium storing a computer program. When the computer-readable storage medium stores a computer program, and the computer program is executed by one or more processors, it causes one or more processors to perform... Figure 1 Each step in the method embodiment.
[0078] Accordingly, embodiments of this disclosure also provide a computer program product, which includes a computer program / instructions that are executed by a processor. Figure 1 Each step in the method embodiment.
[0079] The above Figure 3 The communication component is configured to facilitate wired or wireless communication between the device containing the communication component and other devices. The device containing the communication component can access wireless networks based on communication standards, such as WiFi, 2G, 3G, 4G / LTE, 5G, or combinations thereof. In one exemplary embodiment, the communication component receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In one exemplary embodiment, the communication component also includes a Near Field Communication (NFC) module to facilitate short-range communication. For example, the NFC module may be implemented based on Radio Frequency Identification (RFID), Infrared Data Association (IrDA) technology, Ultra-Wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.
[0080] The above Figure 3The power supply component provides power to various components within the device in which it resides. The power supply component may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to the device in which it resides.
[0081] The aforementioned electronic devices also include a display screen and audio components.
[0082] The display includes a screen, which may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touchscreen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can sense not only the boundaries of touch or swipe actions, but also the duration and pressure associated with the touch or swipe operation.
[0083] An audio component can be configured to output or input audio signals. For example, an audio component includes a microphone (MIC) configured to receive external audio signals when the device containing the audio component is in an operating mode, such as call mode, recording mode, or voice recognition mode. The received audio signals can be further stored in memory or transmitted via a communication component. In some embodiments, the audio component also includes a speaker for outputting audio signals.
[0084] Those skilled in the art will understand that embodiments of this disclosure can be provided as methods, systems, or computer program products. Therefore, this disclosure can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this disclosure can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0085] This disclosure is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this disclosure. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create a machine for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0086] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0087] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0088] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0089] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.
[0090] Computer-readable media include both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.
[0091] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes the element.
[0092] The above are merely specific embodiments of this disclosure, enabling those skilled in the art to understand or implement this disclosure. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this disclosure. Therefore, this disclosure is not to be limited to these embodiments, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A language model training method, characterized in that, include: The initial language model is divided into a forward reasoning module and a reverse verification module; The forward reasoning module generates a forward reasoning chain and a forward reasoning conclusion based on the input text; The reverse verification module, based on the forward reasoning conclusion and with the forward reasoning conclusion as the target, derives the reverse reasoning chain and the reverse output set in reverse. Based on the forward inference chain and the backward inference chain, the initial language model is subjected to multiple rounds of self-game training and logical distillation loops to obtain a trained target language model. The target language model is capable of performing language understanding on the input target language text to obtain the user's requirement text.
2. The method according to claim 1, characterized in that, The forward reasoning chain includes: a set of key variables and logical constraint equations. The forward reasoning module generates a forward reasoning chain and a forward reasoning conclusion based on the input text, including: The forward reasoning module generates the set of key variables, the logical constraint equations, and the forward reasoning conclusion based on the input text, the model's internal knowledge, and the parameter distribution.
3. The method according to claim 1, characterized in that, The reverse inference chain includes: a reverse inference equation; the reverse verification module, based on the forward inference conclusion and with the forward inference conclusion as the target, derives the reverse inference chain and the reverse output set, including: Based on the forward reasoning conclusion, the reverse verification module uses the forward reasoning conclusion as the target to deduce the reverse reasoning equation and the reverse output set.
4. The method according to claim 1, characterized in that, The forward inference chain includes: a logical constraint equation; the reverse inference chain includes: a reverse inference equation; the step of performing multi-round self-game training and logical distillation loops on the initial language model based on the forward inference chain and the reverse inference chain to obtain the trained target language model includes: The logical consistency of the logical constraint equation and the reverse reasoning equation is verified to obtain a logical consistency score. Based on the logical consistency score and logical consistency threshold, the initial language model is subjected to multiple rounds of self-game training and logical distillation loops to obtain the trained target language model.
5. The method according to claim 4, characterized in that, The forward inference chain also includes: a set of key variables; the logical consistency verification of the logical constraint equation and the reverse inference equation to obtain a logical consistency score includes: Calculate the variable consistency score based on the set of key variables and the set of reverse outputs; Calculate the equation consistency score based on the logical constraint equation and the reverse reasoning equation; Calculate the logical consistency score based on the variable consistency score and the equation consistency score.
6. The method according to claim 4, characterized in that, Based on the logical consistency score and logical consistency threshold, the initial language model undergoes multiple rounds of self-game training and logical distillation loops to obtain the trained target language model, including: If the logical consistency score is less than the logical consistency threshold, then a repair action is performed on the inconsistent reasoning path; The initial language model is subjected to multiple rounds of self-game training and logical distillation loops to obtain the trained target language model.
7. The method according to claim 6, characterized in that, The performed repair action includes at least one of the following: Based on logical feedback, identify inconsistent variables or equation differences; input these inconsistent variables or equation differences into the subsequent reasoning process to guide the initial language model to reason again. The total loss function is determined based on the logical consistency loss and the task loss; the initial language model is then trained based on the total loss function. The initial language model is trained using logical closed-loop samples; wherein, the logical closed-loop samples are the samples corresponding to the logical consistency score being greater than or equal to the logical consistency threshold.
8. A language model training device, characterized in that, include: The partitioning module is used to divide the initial language model into a forward reasoning module and a reverse verification module; The generation module is used by the forward reasoning module to generate a forward reasoning chain and a forward reasoning conclusion based on the input text. The derivation module is used by the reverse verification module to derive the reverse reasoning chain and the reverse output set based on the forward reasoning conclusion and with the forward reasoning conclusion as the target. The training module is used to perform multiple rounds of self-game training and logical distillation loops on the initial language model according to the forward inference chain and the backward inference chain to obtain the trained target language model, wherein the target language model is capable of language understanding of the input target language text to obtain the user requirement text.
9. An electronic device, characterized in that, include: processor; Memory used to store processor-executable instructions; The processor is configured to execute instructions to implement the steps of the method as described in any one of claims 1-7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1-7.