A model inference method and apparatus

CN122549569APending Publication Date: 2026-08-11LENOVO (BEIJING) LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-03-31
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

[0003]1、推理效率低下:现有基于显性文本的思维链推理方法生成的推理步骤冗长冗余,需处理大量离散文本token,导致计算开销大、响应延迟高,难以满足自动驾驶、嵌入式智能设备、实时客服系统等实时性应用需求

Benefits of technology

[0006]本公开提供了一种模型推理方法、装置、电子设备及存储介质,以至少解决现有技术中存在的以上技术问题。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122549569A_ABST
    Figure CN122549569A_ABST
Patent Text Reader

Abstract

The present disclosure provides a model reasoning method and device, the method comprising: obtaining a to-be-processed question; preprocessing the to-be-processed question by using an input processing layer of a reasoning model to obtain a word segmentation sequence; identifying the word segmentation sequence by using a first representation processing branch of the reasoning model to generate a first latent representation; if a semantic information value of the first latent representation satisfies a condition, decoding the target answer according to the first latent representation; if the semantic information value of the first latent representation does not satisfy the condition, identifying the word segmentation sequence by using a second representation processing branch of the reasoning model to generate an explicit thinking chain, generating a second latent representation according to the explicit thinking chain, and decoding the target answer according to the second latent representation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of artificial intelligence technology, and in particular to a model reasoning method and apparatus. Background Technology

[0002] With the development of artificial intelligence technology, Large Language Models (LLMs) are widely used in natural language processing, complex reasoning, and other fields. Chain of Reasoning (CoT) technology, by guiding the model to generate a step-by-step reasoning process, significantly improves the model's performance in complex tasks such as mathematical computation and logical reasoning. However, existing technologies have the following key problems:

[0003] 1. Low reasoning efficiency: Existing reasoning methods based on explicit text generate lengthy and redundant reasoning steps, requiring the processing of a large number of discrete text tokens, resulting in high computational overhead and high response latency, making it difficult to meet the real-time application requirements of autonomous driving, embedded intelligent devices, real-time customer service systems, and other applications.

[0004] 2. Limited reasoning accuracy: Traditional thinking chains rely on discrete text tokens to transmit reasoning information. The redundancy of text expression can easily introduce cumulative errors, and the reasoning information encoding in potential reasoning schemes is insufficient, which affects the accuracy of the final reasoning result.

[0005] 3. Insufficient robustness and flexibility: Existing latent inference methods lack intermediate verification mechanisms, the validity of latent representations cannot be evaluated in real time, and deviations in intermediate steps are easily transmitted to the final result; moreover, the use of a uniform training and inference strategy for tasks of different difficulty leads to over-inference for simple tasks and under-inference for complex tasks, making it difficult to balance efficiency and accuracy. Summary of the Invention

[0006] This disclosure provides a model reasoning method, apparatus, electronic device, and storage medium to at least solve the above-mentioned technical problems existing in the prior art.

[0007] According to a first aspect of this disclosure, a model reasoning method is provided, the method comprising: Get the issues to be processed; The input processing layer of the inference model is used to preprocess the problem to be processed, and a word segmentation sequence is obtained; The first representation processing branch of the inference model is used to identify the word segmentation sequence and generate a first latent representation. If the semantic information value of the first latent representation meets the conditions, the target answer is obtained by decoding the first latent representation. If the semantic information value of the first latent representation does not meet the conditions, the second representation processing branch of the inference model is used to identify the word segmentation sequence and generate an explicit thought chain. A second latent representation is generated based on the explicit thought chain, and the target answer is obtained by decoding the second latent representation.

[0008] According to a second aspect of this disclosure, a model inference apparatus is provided, the apparatus comprising: The acquisition module is used to acquire issues to be processed. The processing module is used to preprocess the problem to be processed using the input processing layer of the inference model to obtain a word segmentation sequence; to identify the word segmentation sequence using the first representation processing branch of the inference model to generate a first latent representation; if the semantic information value of the first latent representation meets the conditions, then the target answer is obtained by decoding the first latent representation; if the semantic information value of the first latent representation does not meet the conditions, then the explicit thought chain is generated using the second representation processing branch of the inference model to identify the word segmentation sequence, a second latent representation is generated based on the explicit thought chain, and the target answer is obtained by decoding the second latent representation.

[0009] According to a third aspect of this disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the model inference method described in this disclosure.

[0010] According to a fourth aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions for causing the computer to perform the model inference method described in this disclosure.

[0011] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0012] The above and other objects, features, and advantages of this disclosure will become readily apparent from the following detailed description of exemplary embodiments, taken in conjunction with the accompanying drawings. Several embodiments of this disclosure are illustrated in the drawings by way of example and not limitation, in which: In the accompanying drawings, the same or corresponding reference numerals indicate the same or corresponding parts.

[0013] Figure 1 A flowchart illustrating a model reasoning method provided in an embodiment of this disclosure; Figure 2 This is a schematic diagram of the structure of a reasoning model provided in an embodiment of the present disclosure; Figure 3 A flowchart illustrating a training method for an inference model provided in an embodiment of this disclosure; Figure 4A flowchart illustrating a reasoning method for a reasoning model provided in this embodiment of the disclosure; Figure 5 This is a schematic diagram of the structure of a model inference device provided in an embodiment of the present disclosure; Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. Detailed Implementation

[0014] To make the objectives, features, and advantages of this disclosure more apparent and understandable, the technical solutions in the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this disclosure, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of this disclosure without creative effort are within the scope of protection of this disclosure.

[0015] In the following description, references are made to “some embodiments,” which describe a subset of all possible embodiments. However, it is understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.

[0016] In the following description, the terms "first" and "second" are used merely to distinguish similar objects and do not represent a specific ordering of objects. It is understood that "first" and "second" may be interchanged in a specific order or sequence where permitted, so that the embodiments of this disclosure described herein can be implemented in an order other than that illustrated or described herein.

[0017] Unless otherwise defined, all technical and scientific terms used in this disclosure have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs. The terminology used in this disclosure is for the purpose of describing embodiments of this disclosure only and is not intended to be limiting of this disclosure.

[0018] It should be understood that in the various embodiments of this disclosure, the sequence number of each implementation process does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this disclosure.

[0019] Figure 1 A flowchart illustrating a model inference method provided in an embodiment of this disclosure is shown; as follows: Figure 1 As shown, the method includes: Step 101: Obtain the issues to be processed; Step 102: Preprocess the problem to be processed using the input processing layer of the inference model to obtain the word segmentation sequence; Step 103: Use the first representation processing branch of the inference model to identify the word segmentation sequence and generate a first latent representation; if the semantic information value of the first latent representation meets the conditions, then decode the first latent representation to obtain the target answer; if the semantic information value of the first latent representation does not meet the conditions, use the second representation processing branch of the inference model to identify the word segmentation sequence and generate an explicit thought chain, generate a second latent representation based on the explicit thought chain, and decode the second latent representation to obtain the target answer.

[0020] Here, the problem to be processed is preprocessed to obtain a word segmentation sequence. The native tokenizer of the inference model can be used to segment the problem Q, converting it into a token sequence T=[t1, t2, ..., t] that the model can recognize. n ], each t i A token represents a word, which can be a word, subword, or character, depending on the design of the tokenizer. By segmenting and sorting the tokens, we can ensure the consistency of the input format with that of the pre-training stage, providing a unified data foundation for feature alignment.

[0021] The semantic information value characterizes whether the first latent representation deviates from the content to be expressed in the problem to be processed. In one example, the semantic information value can be represented by the variance of the first latent representation. Accordingly, the semantic information value of the first latent representation satisfies the following conditions: the variance of the first latent representation is less than the target threshold. The semantic information value of the first latent representation does not satisfy the following conditions: the variance of the first latent representation is greater than or equal to the target threshold.

[0022] The latent representation refers to the continuous vector representation output by the hidden layer of the model, which can encode key information without relying on explicit textual forms, and has the advantages of compactness and efficiency.

[0023] The Chain-of-Thought (CoT) refers to the step-by-step reasoning process generated by a large language model when solving complex problems. It improves the accuracy of the final answer through explicit intermediate reasoning steps. An explicit CoT is a continuous thought chain, meaning a reasoning chain that exists in a latent representation form. It transmits reasoning information through continuous vectors, replacing the traditional discrete text-based thought chain.

[0024] The inference model has two distinct processing branches (first and second representation processing branches), which ensures efficient extraction and generation of answers from the problem. The inclusion of a semantic information value judgment mechanism enhances the stability of the inference process, significantly improving robustness in complex scenarios compared to a unified strategy.

[0025] In some embodiments, the method further includes: Difficulty analysis is performed on the word segmentation sequence to obtain the difficulty coefficient; Based on the difficulty coefficient, query the corresponding potential representation dimension to determine the target potential representation dimension; Accordingly, generating a first latent representation based on the word segmentation sequence includes: The embedding features of the word segmentation sequence are mapped to 2k dimensions, where k is the dimension of the target latent representation, and k is greater than or equal to 1; The embedding features of the 2k-dimensional word segmentation sequence are compressed back to k dimensions to obtain the first latent representation.

[0026] Here, an evaluation function with a multi-dimensional difficulty coefficient (λ) can be constructed, for example: λ= f (len(Q),key_words(Q),task_type(Q)); Where len(Q) represents the length of the question text, how many characters or words the question consists of. Generally, the longer the question, the more complex it may be. key_words(Q) represents the keyword complexity (such as the number of mathematical formulas or logical connectors); task_type(Q) represents the task type, such as mathematical calculation, logical reasoning, etc. The problem is analyzed by evaluating the difficulty coefficient, and the difficulty coefficient (λ), λ∈[0,1], is output to provide a basis for dynamic adjustment of the potential representation dimension.

[0027] Here, a pre-designed correspondence between difficulty coefficients and potential representation dimensions can be established. For example, when λ ≥ 0.7, k = 256; when 0.3 < λ < 0.7, k = 128; and when λ ≤ 0.3, k = 64, where k represents the potential representation dimension. The target potential representation dimension is determined by querying the correspondence between difficulty coefficients and potential representation dimensions based on the difficulty coefficient.

[0028] In this way, by adaptively adjusting the latent representation dimension through the difficulty evaluation function of the input processing layer, it is possible to accurately match tasks of different difficulties, achieve the optimal allocation of inference resources under tasks of different difficulties, and avoid over- or under-inference caused by fixed strategies.

[0029] In some embodiments, generating a second latent representation based on the explicit thought chain includes: Map the hidden layer features of the explicit thought chain to 2k dimensions; By compressing the hidden layer features of the 2k-dimensional explicit thought chain back to k dimensions, a second latent representation is obtained.

[0030] Here, explicit thought chain refers to a clear reasoning path or logical chain formed after model processing, which includes intermediate results of problem analysis and reasoning.

[0031] The hidden layer features of an explicit thought chain refer to the feature representations after processing by a neural network. These features extract certain important information from the input data.

[0032] By mapping features to a higher dimension (2k), the expressive power of the model is increased, enabling it to capture more complex patterns or relationships. These features are then compressed back to k dimensions to reduce model complexity and improve computational efficiency, while still retaining sufficient information for subsequent reasoning or decision-making.

[0033] Thus, by compressing the explicit thought chain into a compact latent representation, the generation and processing of discrete tokens are significantly reduced, lowering computational overhead and response latency, thus meeting the requirements of real-time applications. Furthermore, by first mapping the hidden layer features to a higher dimension before compression, effective information can be extracted and retained.

[0034] In some embodiments, the method further includes: training an inference model, the training inference model comprising: Obtain a training dataset, wherein each training data point in the training dataset includes: a question, an explicit thought chain corresponding to the question, and a standard answer; The question for each training data is input into the first representation processing branch to generate a first training latent representation; the first training latent representation is input into the decoding module to decode the first training latent representation to obtain the first answer; The question for each training data is input into the second representation processing branch to generate an explicit training thought chain; the explicit training thought chain is input into the generation module to obtain a second training latent representation; the second training latent representation is input into the decoding module to decode the second training latent representation to obtain a second answer; The total loss function is calculated based on the first answer, the second answer, the trained explicit thought chain, and the explicit thought chain corresponding to the question. The parameters of the generation module and the decoding module of the inference model are adjusted according to the value of the total loss function.

[0035] Here, the inference model may include: an input processing layer, a first representation processing branch, a second representation processing branch, a generation module, and a decoding module.

[0036] Each training data set includes: a problem (such as a problem text or task text to be solved); explicit thought chains, which refer to the explicit thought paths related to the problem, i.e., the logical steps or reasoning process followed when solving the problem; and a standard answer, which refers to the correct answer to the problem.

[0037] The first representation processing branch generates a first training latent representation based on the problem, inputs the first training latent representation into the decoding module, and decodes the first answer from it.

[0038] The second representation processing branch generates a training explicit thought chain based on the question, which demonstrates the thought process of deriving the answer from the question; the generated training explicit thought chain is input into the generation module to obtain the second training latent representation; the second training latent representation is input into the decoding module to decode and obtain the second answer.

[0039] Based on the above process, a total loss function is calculated. This total loss function measures the gap between the model output (answer and thought chain) and the standard answer and explicit thought chain. The parameters of the generation and decoding modules in the inference model are adjusted according to the calculated value of the total loss function. This helps the model to utilize the inference chain and contextual information more effectively when solving complex problems.

[0040] In addition, during the training process, the results generated by the first representation processing branch, the second representation processing branch, the generation module, and the decoding module can be filtered according to the standard answer to obtain the filtering results, which are used to continue training the inference model.

[0041] In some embodiments, adjusting the parameters of the generation and decoding modules of the inference model based on the value of the total loss function includes: If the value of the total loss function does not meet the training requirements, adjust the parameters of the generation module and decoding module of the inference model, and continue training the inference model until the total loss function converges.

[0042] Here, model training requires multiple rounds of parameter updates. If the value of the total loss function does not meet the training requirements (e.g., it is greater than a threshold), the parameters of the generation and decoding modules of the inference model are adjusted, and the inference model continues to be trained until the total loss function converges. Convergence means that the value of the total loss function tends to stabilize and no longer decreases significantly.

[0043] Parameter tuning can be achieved through optimization algorithms (such as gradient descent), specifically by calculating the gradient of the loss function with respect to the model parameters; based on this gradient information, the parameters of the generation and decoding modules are updated to reduce the loss.

[0044] This ensures that the model can be continuously optimized, ultimately achieving better inference results.

[0045] In some embodiments, the total loss function is calculated based on the first answer, the second answer, the trained explicit thought chain, and the explicit thought chain corresponding to the question, including: Calculate the first loss based on the first answer and the second answer; The second loss is calculated based on the explicit thought chain corresponding to the problem and the second training latent representation; The third loss is calculated based on the variance and variance threshold of the second trained latent representation; The total loss function is calculated based on the first loss, the second loss, and the third loss.

[0046] Here, the total loss function adopts self-distillation total loss, which addresses the pain point of insufficient information encoding in traditional latent reasoning through a triple alignment mechanism (including feature alignment, result alignment, and gradient alignment), ensuring that the latent representation completely preserves the reasoning logic. Self-distillation refers to the model simultaneously acting as both a "teacher" and a "student," aligning its explicit and implicit reasoning feature representations to achieve internal knowledge transfer and optimization.

[0047] Feature alignment: refers to calculating the explicit thought chain (i.e., the explicit feature sequence, denoted as H). C The stepwise cosine similarity between the first training latent representation (Z) and the second training latent representation (Z) is used to construct the alignment loss L. align = 1 - cos(H C The loss (Z) is used to force the latent representation to align with the explicit reasoning features in the semantic space, ensuring that the core logic is not lost.

[0048] Result alignment: refers to the cross-entropy loss L obtained by calculating the answers from two processing branches, namely the first answer (i.e., the answer decoded by the first inference branch, denoted as A') and the second answer (i.e., the answer decoded by the second inference branch, denoted as A). ce = CrossEntropy(A, A'), this loss is used to constrain the consistency of two-branch inference results and avoid potential representation encoding bias.

[0049] Gradient alignment refers to the simultaneous optimization of the parameters of the latent representation generation and decoding modules through gradient backpropagation using a hybrid loss function. This achieves synergy in gradient updates across two branches, avoiding alignment imbalances caused by a single loss function. For example, the total loss function can be written as: L = α×L align + β×L ce , where α and β are weighting coefficients, which can both be set to 0.5.

[0050] The total loss function may also include: validation loss (L val That is, the total loss function includes: alignment loss L align Cross-entropy loss L ce , verification loss L val L valThe variance loss of the intermediate validation module is represented as the validation module is used at least for the validation of the variance and variance threshold of the second trained latent representation.

[0051] During model optimization, the total loss L can be minimized using the gradient descent algorithm, and the parameters of the generation and decoding modules can be updated to achieve knowledge self-distillation.

[0052] The training method of this disclosure compresses explicit thought chains into compact latent representations using a self-distillation framework, significantly reducing discrete token generation and processing, lowering computational overhead and response latency, and meeting the requirements of real-time applications. By aligning the features of explicit and implicit reasoning, the latent representation fully preserves key logic, and the intermediate verification mechanism promptly corrects deviations, avoiding accumulated errors. The reliability of the inference results far exceeds that of existing latent reasoning schemes.

[0053] In some embodiments, the training explicit thought chain is input into the generation module to obtain a second training latent representation, including: The hidden layer features of the explicit thought chain are mapped to 2k dimensions; k represents the training latent representation dimension. The hidden layer features of the 2k-dimensional explicit thought chain are subjected to low-rank adaptation, and the adapted 2k-dimensional hidden layer features of the explicit thought chain are compressed back to k dimensions to obtain the second training latent representation.

[0054] Here, the generation module includes: a first fully connected layer, a first LoRA adaptation layer, and a second fully connected layer; the first fully connected layer can map the input features (here referring to the hidden layer features of the explicit thought chain) to 2k dimensions, with ReLU as the activation function; the first LoRA adaptation layer adapts the 2k-dimensional hidden layer features of the explicit thought chain to a low-rank matrix; the second fully connected layer compresses the adapted 2k-dimensional hidden layer features of the explicit thought chain back to k dimensions to obtain the second training latent representation.

[0055] The first LoRA adaptation layer can be adapted using a low-rank matrix W = W0 + ΔW, where W0 is the pre-trained parameter, ΔW = A×B^T, A is a d×r matrix, B is an r×2k matrix, d is the dimension of the input feature or the dimension of the hidden layer, and r is the low-rank dimension, such as r=16.

[0056] In this way, the low-rank matrix enables flexible model adjustment, allowing for improved model performance on new tasks with relatively small parameter increases, while maintaining the integrity of pre-trained parameters. Furthermore, LoRA is compatible with mainstream large language models (LLaMA, Qwen, etc.), requiring no model structure reconstruction, resulting in low deployment costs and solving the problems of poor adaptability and high deployment difficulty.

[0057] In some embodiments, the method further includes: Difficulty coefficients are obtained by performing difficulty analysis on the word segmentation sequence corresponding to the question. Based on the difficulty coefficient, the corresponding potential representation dimension is queried to determine the value.

[0058] Here, we can construct an evaluation function with a multi-dimensional difficulty coefficient (λ), as follows: λ= f (len(Q),key_words(Q),task_type(Q)); Where len(Q) represents the length of the question text, how many characters or words the question consists of. Generally, the longer the question, the more complex it may be. key_words(Q) represents the keyword complexity (such as the number of mathematical formulas or logical connectors); task_type(Q) represents the task type, such as mathematical calculation, logical reasoning, etc. The problem is analyzed by evaluating the difficulty coefficient, and the difficulty coefficient (λ), λ∈[0,1], is output to provide a basis for dynamic adjustment of the potential representation dimension.

[0059] Here, a pre-designed correspondence between difficulty coefficients and latent representation dimensions can be established. For example, k=256 when λ≥0.7, k=128 when 0.3<λ<0.7, and k=64 when λ≤0.3, where k represents the latent representation dimension. The correspondence between difficulty coefficients and latent representation dimensions is queried based on the difficulty coefficients to determine the training latent representation dimension.

[0060] In this way, by adaptively adjusting the latent representation dimension through the difficulty evaluation function of the input processing layer, it is possible to accurately match tasks of different difficulties, achieve the optimal allocation of inference resources under tasks of different difficulties, and avoid over- or under-inference caused by fixed strategies.

[0061] In some embodiments, the second trained latent representation is input into the decoding module, and the second trained latent representation is decoded to obtain a second answer, including: The second training latent representation is mapped to the hidden layer dimension, and the second training latent representation in the hidden layer dimension is identified to obtain the second answer.

[0062] Here, the decoding module is used to decode the key reasoning logic from the latent representation and generate the final answer.

[0063] The decoding module can use a LoRA adaptation layer combined with a fully connected network, with the input being the latent representation and the output being the decoded inference features.

[0064] Specifically, the decoding module may include: a second LoRA adaptation layer and a third fully connected layer; the second LoRA adaptation layer maps the second training latent representation to the hidden layer dimension d (e.g., d=4096); and the third fully connected layer is then used to identify the second training latent representation in the hidden layer dimension to obtain the second answer.

[0065] The method provided in this disclosure is applicable to various tasks such as mathematical calculations and logical reasoning. Through dynamic adjustment and feature alignment techniques, cross-task reasoning performance is also significantly improved.

[0066] Figure 2 This is a schematic diagram of the structure of a reasoning model provided in an embodiment of the present disclosure; as shown below. Figure 2 As shown, for this reasoning module, self-distillation is used to enable the reasoning model to learn to compress explicit thought chains into compact latent representations. Then, a decoding module recovers the key reasoning logic from the latent representations. Simultaneously, intermediate verification and dynamic adjustment mechanisms are introduced to ensure a balance between reasoning efficiency and accuracy. The reasoning model includes: an input processing layer, a self-distillation training framework, and an output layer.

[0067] The self-distillation training framework includes: an implicit reasoning branch (equivalent to a first representation processing branch), an explicit reasoning branch (equivalent to a second representation processing branch), and a hierarchical adaptation module; the hierarchical adaptation module includes: a latent representation generation module (equivalent to the above generation module), a latent representation decoding module (equivalent to the above decoding module), and an intermediate verification module.

[0068] Specifically, the input processing layer is used to perform standardized preprocessing and difficulty quantification assessment of the input question, providing a foundation for subsequent alignment training and dynamic adaptation. Specifically, word segmentation technology can be used to segment the input question Q into token sequences T = [t1, t2, ..., t...]. n ]; and, the difficulty coefficient λ for each problem alignment is determined by the difficulty coefficient evaluation function.

[0069] The self-distillation training framework is used to achieve accurate knowledge transfer between explicit and implicit reasoning through a triple alignment mechanism, ensuring that the latent representation fully encodes the reasoning logic and addressing the pain point of insufficient information encoding in traditional latent reasoning. The triple alignment mechanism includes: 1. Feature Alignment: Calculate the dominant feature sequence H C The stepwise cosine similarity between the hidden layer features of the model corresponding to the explicit thought chain and the latent representation sequence Z is used to construct the alignment loss L. align = 1 - cos(H C (Z), which forces the latent representation and explicit reasoning features to align in the semantic space, ensuring that the core logic is not lost.

[0070] 2. Result Alignment: Calculate the cross-entropy loss L between the answer A′ from the implicit inference branch decoding and the answer A from the explicit inference branch. ce = CrossEntropy(A, A') constrains the consistency of two-branch inference results to avoid potential representation encoding bias.

[0071] 3. Gradient Alignment: Design a mixed total loss function L = α×L align + β×L ce (α and β are weight coefficients, both set to 0.5). The parameters of the latent representation generation module and the decoding module are optimized synchronously through gradient backpropagation to achieve the synergy of gradient updates in the dual-branch system and avoid alignment imbalance caused by a single loss.

[0072] Based on the self-distillation training framework described above, an alignment loss L is constructed. align and cross-entropy loss L ce A total loss function is established and used for model optimization. Model optimization can minimize the total loss L using gradient descent, updating the parameters of the latent representation generation and decoding modules to achieve knowledge self-distillation. This addresses the pain point of insufficient information encoding in latent inference, ensuring that the latent representation fully preserves the reasoning logic.

[0073] Both the latent representation generation and latent representation decoding modules in the hierarchical adaptation module are implemented using LoRA technology, training only the low-rank matrix parameters without changing the main structure of the pre-trained model. The latent representation generation module is used to convert the explicit thought chain or input token sequence into a compact latent representation. This module can be implemented using a two-layer fully connected network combined with a LoRA adaptation layer, with its input being the hidden layer features (H...) of the explicit thought chain. C ) or the embedding features of the input token sequence (H T The output is a latent representation (Z). Specifically, the first fully connected layer maps the input feature dimension to 2k dimensions, with ReLU as the activation function; the LoRA adaptation layer adapts the feature using a low-rank matrix W = W0 + ΔW; the second fully connected layer compresses the feature dimension to k dimensions, obtaining the latent representation Z. The latent representation generation module is also used to adjust the latent representation dimension k according to the difficulty coefficient λ of the input problem. For example, k = 256 when λ ≥ 0.7, k = 128 when 0.3 < λ < 0.7, and k = 64 when λ ≤ 0.3.

[0074] The latent representation decoding module is used to decode key reasoning logic from the latent representation and generate the final answer. It can employ a LoRA adaptation layer combined with a fully connected network, with the latent representation (Z) as input and the decoded reasoning features (H) as output. decSpecifically, the LoRA adaptation layer maps Z to the hidden layer dimension d of the model (d=4096 for example), and the fully connected network maps H... dec The input is processed and fed into the output layer of the pre-trained model to generate the final answer A.

[0075] The intermediate validation module is used to evaluate the validity of the latent representation in real time, correct inference biases, and improve robustness. It extracts the variance feature σ² = Var(Z) of the latent representation (Z), where Var represents the variance. If σ² < θ (θ is a threshold, set to 0.01), the latent representation is considered valid, and the process proceeds to the next decoding step. If σ² ≥ θ, inference bias is detected, triggering an explicit inference branch to supplement key information. One to two core inference steps are generated through the explicit inference branch, input into the latent representation generation module to update Z, and σ² is recalculated until σ² < θ is satisfied.

[0076] The output layer integrates the output of the decoding module with the intermediate verification results, outputting the final answer. The output layer can standardize the format of the answer generated by the decoding module. If the intermediate verification module has triggered a supplementary mechanism, a summary of key reasoning evidence is appended to the answer to ensure the interpretability of the result.

[0077] During the training phase of the inference model, the output layer has ample time to perform loss calculations. Based on the loss calculation results, the parameters of the inference model are adjusted until a suitable inference model is obtained.

[0078] Figure 3 A flowchart illustrating a training method for an inference model provided in this disclosure embodiment; as shown Figure 3 As shown, the method can be applied to Figure 2 The model architecture and methods shown include: Step 301: Obtain the training dataset; Each training data set includes: question Q, the corresponding explicit thought chain C, and the standard answer A. gt .

[0079] Step 302: The input processing layer performs word segmentation and difficulty assessment on question Q to obtain the token sequence T and the difficulty coefficient λ.

[0080] Step 303: Explicit thought chain C predicted by the explicit reasoning branch generation model. pred The implicit reasoning branch prediction yields answer A. pred .

[0081] Step 304: The latent representation generation module will generate C pred Transform into the latent representation Z.

[0082] Step 305: The latent representation decoding module decodes Z to obtain the answer A. dec .

[0083] Step 306: Calculate the total loss L from self-distillation; The total self-distillation loss L includes: alignment loss L align Cross-entropy loss L ce , verification loss L val L val This represents the variance loss of the intermediate validation module.

[0084] In one example, the alignment loss L is combined align Cross-entropy loss L ce L = α×L align +β×L ce (α and β are weighting coefficients, both set to 0.5); In another example, combining the above losses, L = α × L align + β×L ce +η×L val , where α, β, and η are weighting coefficients, and the sum of α, β, and η is 1.

[0085] Step 307: Update the model parameters until the training results are obtained; Here, the parameters of the hierarchical adaptation module can be updated using gradient descent if the loss converges (|L t - L t-1 If |<1e-5), then training ends; otherwise, return to step 402 to continue iterating.

[0086] Figure 4 A flowchart illustrating a reasoning method for a reasoning model provided in this disclosure embodiment; as shown Figure 4 As shown, the method includes: Step 401: Receive user input question Q. The input processing layer identifies the input question Q and generates token timing T and difficulty coefficient λ.

[0087] Step 402: The implicit reasoning branch of the latent representation generation module generates the initial latent representation Z0 based on T.

[0088] Step 403: The intermediate verification module calculates the variance of Z0 and compares the variance calculation result with the threshold. Wherein, the variance of Z0 is denoted as σ0², and the threshold is denoted as threshold. If σ0² < θ, proceed to step 404; otherwise, trigger explicit inference to supplement the core step, update Z0 to Z1, and perform intermediate verification again before proceeding to step 404.

[0089] Step 404: The latent representation decoding module decodes Z (Z0 or Z1) to obtain answer A.

[0090] Step 405: Output the standardized answer A from the output layer, and the reasoning ends.

[0091] Figure 5 This is a schematic diagram of the structure of a model inference device provided in an embodiment of the present disclosure; as shown below. Figure 5 As shown, the device includes: The acquisition module is used to acquire issues to be processed. The processing module is used to preprocess the problem to be processed using the input processing layer of the inference model to obtain a word segmentation sequence; to identify the word segmentation sequence using the first representation processing branch of the inference model to generate a first latent representation; if the semantic information value of the first latent representation meets the conditions, then the target answer is obtained by decoding the first latent representation; if the semantic information value of the first latent representation does not meet the conditions, then the explicit thought chain is generated using the second representation processing branch of the inference model to identify the word segmentation sequence, a second latent representation is generated based on the explicit thought chain, and the target answer is obtained by decoding the second latent representation.

[0092] In some embodiments, the processing module is further configured to perform difficulty analysis based on the word segmentation sequence to obtain a difficulty coefficient; Based on the difficulty coefficient, query the corresponding potential representation dimension to determine the target potential representation dimension; Correspondingly, the processing module is used to map the embedded features of the word segmentation sequence to 2k dimensions, where k is the target latent representation dimension and k is greater than or equal to 1; The embedding features of the 2k-dimensional word segmentation sequence are compressed back to k dimensions to obtain the first latent representation.

[0093] In some embodiments, the processing module is configured to map the hidden layer features of the explicit thought chain to 2k dimensions; By compressing the hidden layer features of the 2k-dimensional explicit thought chain back to k dimensions, a second latent representation is obtained.

[0094] In some embodiments, the apparatus further includes: a training module for training an inference model, the training inference model comprising: Obtain a training dataset, wherein each training data point in the training dataset includes: a question, an explicit thought chain corresponding to the question, and a standard answer; The question for each training data is input into the first representation processing branch to generate a first training latent representation; the first training latent representation is input into the decoding module to decode the first training latent representation to obtain the first answer; The question for each training data is input into the second representation processing branch to generate an explicit training thought chain; the explicit training thought chain is input into the generation module to obtain a second training latent representation; the second training latent representation is input into the decoding module to decode the second training latent representation to obtain a second answer; The total loss function is calculated based on the first answer, the second answer, the trained explicit thought chain, and the explicit thought chain corresponding to the question. The parameters of the generation module and the decoding module of the inference model are adjusted according to the value of the total loss function.

[0095] In some embodiments, the training module is configured to adjust the parameters of the generation module and the decoding module of the inference model if the value of the total loss function does not meet the training requirements, and continue training the inference model until the total loss function converges.

[0096] In some embodiments, the training module is configured to calculate a first loss based on the first answer and the second answer; The second loss is calculated based on the explicit thought chain corresponding to the problem and the second training latent representation; The third loss is calculated based on the variance and variance threshold of the second trained latent representation; The total loss function is calculated based on the first loss, the second loss, and the third loss.

[0097] In some embodiments, the training module is used to map the hidden layer features of the explicit thought chain to 2k dimensions; k represents the training latent representation dimension; The hidden layer features of the 2k-dimensional explicit thought chain are subjected to low-rank adaptation, and the adapted 2k-dimensional hidden layer features of the explicit thought chain are compressed back to k dimensions to obtain the second training latent representation.

[0098] In some embodiments, the training module is further configured to perform difficulty analysis based on the word segmentation sequence corresponding to the question to obtain a difficulty coefficient; The training latent representation dimension is determined by querying the corresponding latent representation dimension based on the difficulty coefficient.

[0099] In some embodiments, the training module is configured to map the second training latent representation to the hidden layer dimension, identify the second training latent representation in the hidden layer dimension, and obtain a second answer.

[0100] It is understood that, when implementing the corresponding model inference method, the model inference apparatus provided in the above embodiments can allocate the above processing to different program modules as needed to complete all or part of the processing described above. Furthermore, the server provided in the above embodiments and the embodiments of the corresponding methods belong to the same concept; its specific implementation process is detailed in the method embodiments and will not be repeated here.

[0101] This disclosure provides a computer-readable storage medium storing executable instructions, wherein the executable instructions are stored and, when executed by a processor, will trigger the processor to execute the model inference method provided in this disclosure.

[0102] In some embodiments, the computer-readable storage medium may be a ferroelectric random access memory (FRAM), a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), flash memory, magnetic surface memory, optical disc, or CD-ROM, etc.; or it may be a device that includes one or any combination of the above-mentioned memories.

[0103] In some embodiments, executable instructions may take the form of a program, software, software module, script, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as a standalone program or as a module, model, subroutine, or other unit suitable for use in a computing environment.

[0104] As an example, executable instructions can be deployed to execute on a single computing device, or on multiple computing devices located in one location, or on multiple computing devices distributed across multiple locations and interconnected via a communication network.

[0105] This disclosure provides a computer program product, which includes a computer program / instruction that, when executed by a processor, implements the model inference method described in this disclosure.

[0106] Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure; as shown below. Figure 6 As shown, the electronic device 60 includes: a processor 601 and a memory 602 for storing computer programs that can run on the processor; when the processor 601 runs the computer program, it executes the communication method provided in the embodiments of this disclosure.

[0107] In practical applications, the electronic device 60 may further include at least one network interface 603. The various components of the electronic device 60 are coupled together via a bus system 604. It is understood that the bus system 604 is used to implement communication between these components. In addition to a data bus, the bus system 604 also includes a power bus, a control bus, and a status signal bus. However, for clarity, in... Figure 6 Various buses are designated as bus system 604. The number of processors 601 can be at least one. Network interface 603 is used for wired or wireless communication between electronic device 60 and other devices.

[0108] The memory 602 in this embodiment is used to store various types of data to support the operation of the electronic device 60.

[0109] The methods disclosed in the above embodiments of this disclosure can be applied to processor 601, or implemented by processor 601. Processor 601 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the integrated logic circuit of the hardware in processor 601 or by instructions in the form of software. The processor 601 may be a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Processor 601 can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this disclosure. A general-purpose processor may be a microprocessor or any conventional processor, etc. The steps of the methods disclosed in the embodiments of this disclosure can be directly manifested as being executed by a hardware decoding processor, or being executed by a combination of hardware and software modules in the decoding processor. The software modules may be located in a storage medium, which is located in memory 602. Processor 601 reads the information in memory 602 and combines its hardware to complete the steps of the aforementioned method.

[0110] In some embodiments, the electronic device 60 may be implemented by one or more application-specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), general-purpose processors, controllers, microcontrollers (MCUs), microprocessors, or other electronic components to perform the aforementioned methods.

[0111] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.

[0112] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this disclosure, "a plurality of" means two or more, unless otherwise explicitly specified.

[0113] The above description is merely a specific embodiment of this disclosure, but the scope of protection of this disclosure is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this disclosure should be included within the scope of protection of this disclosure. Therefore, the scope of protection of this disclosure should be determined by the scope of the claims.

Claims

1. A model reasoning method, the method comprising: Get the issues to be processed; The input processing layer of the inference model is used to preprocess the problem to be processed, and a word segmentation sequence is obtained; The first representation processing branch of the inference model is used to identify the word segmentation sequence and generate a first latent representation. If the semantic information value of the first latent representation meets the conditions, the target answer is obtained by decoding the first latent representation. If the semantic information value of the first latent representation does not meet the conditions, the second representation processing branch of the inference model is used to identify the word segmentation sequence and generate an explicit thought chain. A second latent representation is generated based on the explicit thought chain, and the target answer is obtained by decoding the second latent representation.

2. The method according to claim 1, further comprising: Difficulty analysis is performed on the word segmentation sequence to obtain the difficulty coefficient; Based on the difficulty coefficient, query the corresponding potential representation dimension to determine the target potential representation dimension; Accordingly, generating a first latent representation based on the word segmentation sequence includes: The embedding features of the word segmentation sequence are mapped to 2k dimensions, where k is the dimension of the target latent representation, and k is greater than or equal to 1; The embedding features of the 2k-dimensional word segmentation sequence are compressed back to k dimensions to obtain the first latent representation.

3. The method according to claim 2, wherein generating a second latent representation based on the explicit thought chain comprises: Map the hidden layer features of the explicit thought chain to 2k dimensions; By compressing the hidden layer features of the 2k-dimensional explicit thought chain back to k dimensions, a second latent representation is obtained.

4. The method of claim 1, further comprising: Training an inference model, the training inference model comprising: Obtain a training dataset, wherein each training data point in the training dataset includes: a question, an explicit thought chain corresponding to the question, and a standard answer; The question for each training data is input into the first representation processing branch to generate a first training latent representation; the first training latent representation is input into the decoding module to decode the first training latent representation to obtain the first answer; The question for each training data is input into the second representation processing branch to generate an explicit training thought chain; the explicit training thought chain is input into the generation module to obtain a second training latent representation; the second training latent representation is input into the decoding module to decode the second training latent representation to obtain a second answer; The total loss function is calculated based on the first answer, the second answer, the trained explicit thought chain, and the explicit thought chain corresponding to the question. The parameters of the generation module and the decoding module of the inference model are adjusted according to the value of the total loss function.

5. The method according to claim 4, wherein adjusting the parameters of the generation module and the decoding module of the inference model based on the value of the total loss function comprises: If the value of the total loss function does not meet the training requirements, adjust the parameters of the generation module and decoding module of the inference model, and continue training the inference model until the total loss function converges.

6. The method according to claim 4, wherein calculating the total loss function based on the first answer, the second answer, the trained explicit thought chain, and the explicit thought chain corresponding to the question, comprises: Calculate the first loss based on the first answer and the second answer; The second loss is calculated based on the explicit thought chain corresponding to the problem and the second training latent representation; The third loss is calculated based on the variance and variance threshold of the second trained latent representation; The total loss function is calculated based on the first loss, the second loss, and the third loss.

7. The method according to claim 4, wherein the training explicit thought chain is input into the generation module to obtain a second training latent representation, comprising: The hidden layer features of the explicit thought chain are mapped to 2k dimensions; k represents the training latent representation dimension. The hidden layer features of the 2k-dimensional explicit thought chain are subjected to low-rank adaptation, and the adapted 2k-dimensional hidden layer features of the explicit thought chain are compressed back to k dimensions to obtain the second training latent representation.

8. The method according to claim 7, further comprising: Difficulty coefficients are obtained by performing difficulty analysis on the word segmentation sequence corresponding to the question. The training latent representation dimension is determined by querying the corresponding latent representation dimension based on the difficulty coefficient.

9. The method according to claim 4, wherein the second trained latent representation is input into the decoding module, and the second trained latent representation is decoded to obtain the second answer, comprising: The second training latent representation is mapped to the hidden layer dimension, and the second training latent representation in the hidden layer dimension is identified to obtain the second answer.

10. A model reasoning apparatus, the apparatus comprising: The acquisition module is used to acquire issues to be processed. The processing module is used to preprocess the problem to be processed using the input processing layer of the inference model to obtain the word segmentation sequence; The first representation processing branch of the inference model is used to identify the word segmentation sequence and generate a first latent representation. If the semantic information value of the first latent representation meets the conditions, the target answer is obtained by decoding the first latent representation. If the semantic information value of the first latent representation does not meet the conditions, the second representation processing branch of the inference model is used to identify the word segmentation sequence and generate an explicit thought chain. A second latent representation is generated based on the explicit thought chain, and the target answer is obtained by decoding the second latent representation.