Generation optimization method and device based on dynamic verification feedback, equipment and medium

By building a dynamic verification feedback mechanism, the problem of parameter update lag in generative model training is solved, the model's rapid response to multi-objective requirements and training convergence are achieved, and the quality and adaptability of the generated content are improved, especially in the fields of financial technology and medical health.

CN120745822APending Publication Date: 2025-10-03PING AN TECH (SHENZHEN) CO LTD
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510862161.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-25
Publication Date
2025-10-03

AI Technical Summary

Technical Problem

Existing generative model training methods lack a dynamic verification and feedback mechanism, resulting in delayed updates of training parameters. The model easily falls into local optimality and is difficult to adapt to changing business contexts and complex task requirements, especially in the fields of financial technology and healthcare.

Method used

A generative optimization method based on dynamic verification feedback is constructed. By pre-training the verification network, constructing a dynamic reward space with multiple orthogonal reward components, and injecting verification signals into multiple processing layers of the generative model, a real-time verification loop is formed. The exploration strategy is optimized and dual-loop feedback control is performed to dynamically adjust the training parameters.

Benefits of technology

The generative model's ability to quickly respond to multi-objective demands and improve training convergence speed have been achieved, the accuracy, logic and diversity of generated content have been improved, and the model's performance in complex tasks has been enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120745822A_ABST
    Figure CN120745822A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence, can be applied to business scenes such as financial science and technology and medical health, and discloses a generation optimization method and device based on dynamic verification feedback, equipment and a medium, and the method comprises the steps: pre-training a verification network used for analyzing and generating content quality; constructing a dynamic reward space containing orthogonal reward components; integrating the dynamic reward space into a generative model to form a real-time verification loop, and injecting verification signals into a plurality of processing layers; optimizing an exploration strategy of the generative model according to the verification signal; executing double-loop feedback control based on the real-time verification loop and the optimized exploration strategy, dynamically adjusting training parameters, and generating a target model; and outputting a reasoning result based on the target model. By constructing a real-time verification loop and an optimization strategy, a double-loop feedback control mechanism is realized, a verification signal is introduced into a training process, a training error and a strategy deviation are dynamically responded, and the training efficiency of the generative model is improved by combining rapid and slow adjustment paths.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a generation optimization method, device, equipment and storage medium based on dynamic verification feedback. Background Art

[0002] Current generative model training methods primarily rely on static training strategies and fixed verification mechanisms, making them difficult to adapt to the dynamic evolution of models during training. Existing technologies generally employ a single-directional training feedback path, lacking a real-time response mechanism for training errors and strategy exploration results. This results in delayed training parameter updates, making the model prone to falling into local optimality, which in turn affects overall performance and generalization.

[0003] In the fintech sector, generative models are used for tasks such as compliance text generation and intelligent review assistance, which place high demands on factual accuracy and logical consistency. However, due to the lack of dynamic adjustment mechanisms during training, generative models often suffer from rigid strategies and a monotonous text style when processing multiple rounds of approval rules and complex regulatory statements, making them difficult to adapt to changing business contexts and risk appetites.

[0004] In the healthcare sector, generative models are widely used in scenarios such as assisted medical record writing and clinical dialogue generation. Training often faces challenges with dense medical terminology and complex reasoning chains. Existing training feedback paths cannot dynamically adjust parameters based on real-time feedback from clinical context. This leads to unstable model processing of professional terminology and insufficient logical reasoning span, thus affecting the credibility and usability of diagnosis and treatment recommendations.

[0005] Furthermore, current technologies generally lack a dual-channel feedback mechanism in generative model training. There is a time lag and path inconsistency between training strategy updates and model performance evaluation, making it difficult to achieve coordinated optimization of training accuracy and strategy exploration efficiency. This deficiency is particularly pronounced when training large, complex models, severely restricting the efficiency and effectiveness of model training. In particular, in multi-layer structures, it is impossible to achieve hierarchical calibration and synchronous strategy updates. Existing methods also lack precise dynamic adjustment of training parameters, and are unable to jointly drive adaptive adjustments to the training process based on error feedback and strategy status, limiting the performance of generative models in complex tasks. Summary of the Invention

[0006] The main purpose of the present invention is to provide a generation optimization method, device, equipment and storage medium based on dynamic verification feedback, aiming to solve the technical problem that the existing technology lacks a dual-loop feedback mechanism based on dynamic verification signals to achieve adaptive adjustment of training parameters, resulting in the inability to respond to errors and strategy changes in real time during generative model training.

[0007] To achieve the above objectives, the present invention provides a generation optimization method based on dynamic verification feedback, comprising:

[0008] Pre-training a verification network for analyzing the quality of generated content;

[0009] Based on the verification network, construct a dynamic reward space including multiple orthogonal reward components;

[0010] Integrating the dynamic reward space into a generative model to be trained to form a real-time verification loop that injects verification signals into multiple processing layers of the generative model;

[0011] Optimizing the exploration strategy of the generative model according to the verification signal to obtain an optimized exploration strategy;

[0012] Based on the real-time verification loop and the optimized exploration strategy, performing dual-loop feedback control on the training process of the generative model to dynamically adjust training parameters and generate a target model;

[0013] Outputting an inference result based on the target model.

[0014] Furthermore, to achieve the above-mentioned purpose, the present invention provides a generation optimization device based on dynamic verification feedback, comprising:

[0015] Verification network building module, used to pre-train the verification network used to analyze the quality of generated content;

[0016] A reward space generation module, configured to construct a dynamic reward space comprising a plurality of orthogonal reward components based on the verification network;

[0017] A verification loop injection module, configured to integrate the dynamic reward space into the generative model to be trained to form a real-time verification loop, wherein the real-time verification loop injects verification signals into multiple processing layers of the generative model;

[0018] an exploration strategy optimization module, configured to optimize the exploration strategy of the generative model according to the verification signal to obtain an optimized exploration strategy;

[0019] A feedback control execution module is configured to perform dual-loop feedback control on the training process of the generative model based on the real-time verification loop and the optimized exploration strategy to dynamically adjust training parameters and generate a target model;

[0020] The inference result generation module is used to output the inference result based on the target model.

[0021] Furthermore, to achieve the above-mentioned purpose, the present invention also provides a computer device, which includes a memory, a processor, and a generation optimization program based on dynamic verification feedback stored in the memory and run on the processor. When the generation optimization program based on dynamic verification feedback is executed by the processor, the steps of the generation optimization method based on dynamic verification feedback as described above are implemented.

[0022] Furthermore, to achieve the above-mentioned purpose, the present invention also provides a computer-readable storage medium, on which a generation optimization program based on dynamic verification feedback is stored. When the generation optimization program based on dynamic verification feedback is executed by a processor, the steps of the generation optimization method based on dynamic verification feedback as described above are implemented.

[0023] Beneficial effects: The present invention relates to the field of artificial intelligence technology and can be applied to business scenarios such as financial technology and medical health. It discloses a generation optimization method, device, equipment and medium based on dynamic verification feedback, including: pre-training a verification network for analyzing the quality of generated content; constructing a dynamic reward space containing multiple orthogonal reward components based on the verification network; integrating the dynamic reward space into the generative model to be trained to form a real-time verification loop, and injecting verification signals into multiple processing layers of the generative model; optimizing the exploration strategy of the generative model according to the verification signal to obtain an optimized exploration strategy; performing dual-loop feedback control on the generative model based on the real-time verification loop and the optimized exploration strategy, dynamically adjusting the training parameters, and generating a target model; and outputting the inference results based on the target model. The present invention realizes a dual-loop feedback control mechanism in the training process by constructing a real-time verification loop and an optimized exploration strategy, injecting verification signals into the model training process, and can dynamically capture training errors and strategy deviations. It can also adjust the training parameters in real time through the synergy of fast response and slow optimization, thereby effectively improving the responsiveness of the generative model to multi-objective requirements and the training convergence speed. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] The present invention will be further described below with reference to the accompanying drawings and embodiments, in which:

[0025] Figure 1 A schematic diagram of an application environment of a generation optimization method based on dynamic verification feedback according to an embodiment of the present invention;

[0026] Figure 2 Schematic diagram of a flow chart of an embodiment of a generation optimization method based on dynamic verification feedback of the present invention;

[0027] Figure 3 Schematic diagram of functional modules of a preferred embodiment of a generation optimization device based on dynamic verification feedback of the present invention;

[0028] Figure 4A schematic diagram of the structure of a computer device according to an embodiment of the present invention;

[0029] Figure 5 FIG. 2 is another structural diagram of a computer device according to an embodiment of the present invention. DETAILED DESCRIPTION

[0030] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0031] The generation optimization method based on dynamic verification feedback provided by the embodiment of the present invention can be applied in the following situations: Figure 1 In an application environment, the user end communicates with the server end through a network. The server end can pre-train a verification network for analyzing the quality of generated content through the user end; construct a dynamic reward space containing multiple orthogonal reward components based on the verification network; integrate the dynamic reward space into the generative model to be trained to form a real-time verification loop, and inject verification signals into multiple processing layers of the generative model; optimize the exploration strategy of the generative model based on the verification signal to obtain an optimized exploration strategy; perform dual-loop feedback control on the generative model based on the real-time verification loop and the optimized exploration strategy, dynamically adjust training parameters, and generate a target model; and output inference results based on the target model. The present invention realizes a dual-loop feedback control mechanism during the training process by constructing a real-time verification loop and an optimized exploration strategy, injecting verification signals into the model training process, and can dynamically capture training errors and strategy deviations. Through the synergistic effect of fast response and slow optimization, the training parameters are adjusted in real time, thereby effectively improving the responsiveness of the generative model to multi-objective requirements and the training convergence speed. The user end can be, but is not limited to, various personal computers, laptops, smart phones, tablets, and portable wearable devices. The server end can be implemented as an independent server or a server cluster consisting of multiple servers. The present invention is described in detail below through specific examples.

[0032] See also Figure 2 , Figure 2 This is a flow chart of an embodiment of a generation optimization method based on dynamic verification feedback provided by the present invention. It should be noted that although a logical order is shown in the flow chart, in some cases, the steps shown or described may be performed in a different order than that shown here.

[0033] like Figure 2 As shown, the generation optimization method based on dynamic verification feedback proposed by the present invention includes the following steps:

[0034] S10, a pre-trained verification network for analyzing the quality of generated content;

[0035] In this embodiment, to effectively evaluate the quality of generated content, a discriminative quality assessment model is pre-built during the training phase. This model is capable of learning to identify potential flaws in generated text. First, a text sample dataset is generated based on a language modeling dataset frequently used in natural language processing tasks. This dataset can be extracted from general corpora or constructed from scenario-specific question-and-answer logs, summary data, or generated content sets, and should cover a variety of linguistic phenomena and text types. To improve the evaluation model's ability to identify different types of questions, structural flaws are introduced into the original text samples, including semantic traps, logical conflicts, and factual distortions. Semantic traps can be introduced through ambiguous references and ambiguous sentence construction. Logical conflicts can be achieved through reversed statements or conditional errors. Factual errors are introduced by introducing content that is inconsistent with real background knowledge, such as mismatches in time, quantity, or causal relationships. The constructed text sample dataset is annotated, and the annotations must accurately distinguish error types and mark the locations or segments of the errors.

[0036] After structural transformation, the above-mentioned labeled data forms a multi-dimensional verification indicator matrix. Each indicator dimension corresponds to a potential defect category and can be further subdivided into subcategories. For example, logical rationality can be divided into conditional consistency, causal chain integrity, and reasoning path correctness. This matrix is ​​used as a label reference in subsequent model training, and the confidence interval error control technology is used to screen the stability of each indicator dimension, eliminate extreme outliers or inconsistent annotations, and control training stability. The confidence interval error control technology can be used to dynamically generate an error control parameter table based on Bootstrap resampling and confidence upper bound calculation. The parameter table contains the deviation threshold and adjustment weights for each dimension. Finally, the controlled indicator matrix and error control parameter table are input into the verification network for training. The training method adopts a discriminative learning strategy, such as multi-label cross entropy loss combined with a weighted loss function mechanism, to ensure that the network has the ability to distinguish multiple types of defects and outputs a quantifiable score.

[0037] The aforementioned training process can be implemented using a multi-layer Transformer structure to implement a verification network. The input layer receives a token sequence and a corresponding error label matrix, and the output layer generates predicted probabilities for various quality indicators. Alternatively, a dual-channel encoding structure can be used, with one channel encoding the original text and the other encoding the text after introducing defects. Comparative learning can be used to obtain differential representations of text quality. Furthermore, indicator design can be adapted to specific tasks, such as strengthening the dimension of factual consistency in financial data generation and adding the dimension of domain terminology standardization in medical report generation. Confidence interval error control algorithms in different scenarios can adjust sample sampling strategies or confidence bound coefficients based on annotation quality to improve error estimation accuracy. Implementations can also incorporate pseudo-labeling mechanisms to expand the training set size for unlabeled data and improve the model's ability to identify boundary samples.

[0038] Example description: In the healthcare business scenario, a training dataset can be constructed based on the output of the electronic medical record automatic generation model, injecting relevant defects in the fields of medication dosage errors, diagnostic conclusion jumps, etc., and medical experts can accurately label the error locations and categories. The verification network is trained to identify risk expressions or unreasonable statements in medically generated content.

[0039] In the financial field, we can use natural language to generate content such as financial reports and market briefs, inject types of broken logical reasoning chains or data alignment errors, build a quality matrix through professional annotation, and train an evaluation network to identify financial logic inconsistencies or factual errors in the generated content, thereby improving the stability and credibility of the downstream generation system.

[0040] This embodiment builds a verification network for analyzing the quality of generated content, enabling continuous evaluation feedback from real text defects during training. By injecting different types of errors and performing structured annotation, the model is equipped to identify a variety of potential quality issues. An error control mechanism optimizes the stability of training labels, effectively reducing training bias in the evaluation model, thereby improving overall quality assessment accuracy and providing reliable input for subsequent dynamic reward generation.

[0041] S20, constructing a dynamic reward space including a plurality of orthogonal reward components based on the verification network;

[0042] In this embodiment, in order to build a training feedback system with discriminative and instructive features, it is necessary to further extract a multi-dimensional reward signal for training the generative model based on the output of the quality assessment model completed in the early stage of training. The data output by the verification network contains evaluation results in multiple dimensions, which correspond to the performance of the generated content in terms of factual accuracy, logical consistency, semantic fluency, and contextual alignment. In order to establish a more discriminative reward mechanism, it is necessary to first extract these dimensions as basic reward component basic data. These basic data can be obtained through vectorized scoring results or numerical conversion of discrete labels. They can be formally represented as several scoring vectors, each of which is mapped to a specific quality indicator.

[0043] To address the correlation issues between different basic reward components, such as the strong correlation between factual accuracy and logical consistency, direct training can easily lead to gradient redundancy or unclear reward guidance direction. Therefore, they need to be orthogonalized. Orthogonalization can be achieved through principal component stripping, eigenvector projection, or kernel function decoupling methods, so that each dimension of reward has an independent contribution to the feedback effect without interference. After orthogonalization, multiple independent reward components are formed, each corresponding to an independent quality feedback direction.

[0044] After obtaining multiple orthogonal reward components, it is necessary to further construct an adaptable reward combination system. To this end, a dynamic weight allocation model can be introduced to assign different weights to different reward components at different stages of training. This model can be adaptively updated based on performance feedback on the goal generation task. For example, in the early stages of the model, the weight of the reward for factual accuracy can be strengthened, while in later stages, the weight of the reward for innovation and coherence can be strengthened. The dynamic weight allocation model can be implemented using an attention mechanism, a hierarchical weight scheduler, or a reinforcement learning controller, dynamically assigning weights to each reward component to generate the final combined weight.

[0045] By combining orthogonal reward components and dynamic combination weights, a reward projection matrix is ​​constructed. This matrix serves as a mapping tool from the validation space to the generative model training feedback space. During the generation process, projection calculations are performed based on the current model output quality feedback, dynamically adjusting the reward value output. This creates a dynamic reward space that changes in real time with changes in the quality of the generated content. This space enables fine-grained discrimination of generated content and effectively guides the evolution of the generative model towards multi-dimensional quality optimization.

[0046] Orthogonalization can be performed using a principal component analysis method based on covariance matrix decomposition, or an asymmetric kernel correlation stripping method, so that the gradient direction of each reward component is independent of each other during training. In specific implementations, a pseudo-labeling mechanism can also be introduced to expand the reward components to more detailed quality attribute dimensions, such as introducing title consistency rewards in news generation tasks and answer coverage rewards in question and answer generation. The dynamic weight distribution model can be designed as a controller based on a reinforcement learning strategy, which inputs the performance indicators of the current model in different dimensions and outputs the weight distribution of each reward component in the current training cycle. The reward projection matrix can be implemented using a matrix weighted linear combination structure or a tensor mapping structure to adapt to the coupling relationship between high-order features.

[0047] Example description: In the structured medical record generation task in the healthcare field, reward components are constructed based on dimensions such as doctor terminology accuracy, timeline rationality, and consistency of treatment recommendations. These components correspond to factual accuracy, logical coherence, and scenario consistency, respectively. After orthogonalization, they can independently guide the model to generate optimized content in different dimensions.

[0048] In the task of generating intelligent investment analysis text in the financial field, the accuracy of economic data citation, consistency of causal reasoning, and policy sensitivity can be used as the basic dimensions of rewards. By dynamically adjusting the reward combination weights, the model can generate more professional and compliant analysis content when responding to market events.

[0049] This embodiment constructs a dynamic reward space containing multiple orthogonal reward components, enabling balanced optimization of the training process across multiple quality dimensions. Orthogonalization reduces redundancy between reward components and avoids gradient interference between multiple quality indicators. The dynamic weight model ensures that the optimization target can be automatically adjusted at different stages of training, giving the generated model greater adaptability and generalization capabilities. Ultimately, the dynamic reward space constructed through the reward projection matrix provides refined and real-time training feedback, significantly improving the overall performance of the model-generated content in terms of accuracy, logic, and diversity.

[0050] S30, integrating the dynamic reward space into the generative model to be trained to form a real-time verification loop, wherein the real-time verification loop injects verification signals into multiple processing layers of the generative model;

[0051] In this embodiment, the dynamic reward space is applied to the training process and needs to be embedded in the internal structure of the generative model to achieve real-time feedback and dynamic guidance during the training process. To this end, it is necessary to integrate the reward space into the generative model to be trained and design a mechanism that can provide real-time feedback on quality evaluation information to form a closed loop. This structure not only uses rewards as external signals to weight the training objective function, but also embeds the reward evaluation mechanism into multiple processing layers of the model in a modular manner, constructing an inner loop structure that can be sensed by gradients and dynamically adjusted.

[0052] The key to this structure is identifying the most suitable locations within the multiple processing layers of the generative model as injection points for the verification signal. To achieve this, modeling analysis is performed to select specific intermediate layers, which typically represent representative feature expressions or semantic generation nodes, such as the multi-head attention layer, feedforward sublayer, and decoder alignment position in the Transformer model. Based on this analysis, intermediate layer position information is generated and fixed as a mapping index structure for subsequent injection operations.

[0053] Each orthogonal reward component in the dynamic reward space must be processed by a gating mechanism and converted into a vector signal that matches the internal structure of the target model. The gating mechanism controls the flux and direction of the reward signal. It works by converting the external reward into an internal representation tensor through gating functions (such as sigmoid, softmax, or attention-weighted gating), which can both express the characteristics of the reward component and be compatible with the internal parameter dimensions of the model. The resulting validation signal can be injected as an additional input vector into the selected intermediate processing layer.

[0054] To achieve the injection effect, the verification signal needs to be combined with the original activation value of the intermediate layer. This can be done by methods such as vector concatenation, residual connection, feature fusion, or channel mixing. In order to ensure that the signal can be perceived by the backpropagation mechanism during model training, a bidirectional propagation path needs to be designed. That is, the gradient transmitted backward from the output loss layer can pass through the verification signal branch and fuse with the backbone model to affect its parameter update. This constitutes the first closed-loop path. In addition, feedback needs to be generated from the direction of the reward evaluation signal itself to control the gating mechanism and injection frequency. This constitutes the second closed-loop path, thus realizing a complete dual-path feedback loop. The two loops work together to form a real-time verification loop that integrates the reward mechanism and generation logic.

[0055] Specific intermediate layer positions can be selected based on a pre-training semantic importance scoring mechanism. For example, information bottleneck analysis can be used to determine the importance of representation at different layers in the model, and the top-K representation layers are recorded as injection locations. The gating mechanism can be implemented as a structured learnable weight function. All reward signals are uniformly enabled during initialization, and the optimal flux path is automatically learned based on gradient changes during training. The fusion operation of the verification signal and the intermediate layer activation vector can use dot-product attention to enhance information coupling, or low-rank tensor projection can be used to compress the high-dimensional reward signal into the backbone feature space. The bidirectional feedback mechanism can be explicitly implemented through a dual-channel parameter update path: one path updates the backbone model from the loss function, and the other path updates the gating weights from the evaluation module. This structure is compatible with existing mainstream Transformer architectures and large-scale generative models based on decoder architectures.

[0056] Example: In the healthcare field, generative models are used to generate electronic medical record summaries or medical history records. Rewards for factual accuracy and terminology standardization can be incorporated into the model's decoder layer. By injecting verification signals into the encoder-decoder interaction layer, the model's accuracy in disease descriptions and treatment recommendations can be improved, preventing errors in medication logic or mixed terminology in the generated content.

[0057] In the financial field, when generative models are used to generate compliance reports or financial summaries, verification signals regarding data consistency and reporting style stability can be injected into the key economic indicator description modules at each stage. Through real-time verification loops, the model's learning effects on time consistency, unit conversion, and expression style are continuously strengthened, thereby outputting analytical report content with more professional standards.

[0058] This embodiment integrates a dynamic reward space into the generative model and constructs a real-time verification loop, enabling the reward evaluation mechanism to directly participate in the model's intermediate representation adjustment process, changing the structural limitation of traditional RLHF training, which only provides feedback rewards at the output stage. By injecting verification signals into multiple processing layers, the model can perceive the current output quality at each step of generation, thereby dynamically adjusting its generation strategy and semantic construction direction. The dual-loop mechanism ensures a complete feedback path during training, making the optimization process more stable, while improving the coordination between training sample efficiency and multi-objective indicators.

[0059] S40, optimizing the exploration strategy of the generative model according to the verification signal to obtain an optimized exploration strategy;

[0060] In this embodiment, during generative model training, the exploration strategy determines how the model selects the next output in the action space. To ensure both diversity and quality constraints in generated behavior, the exploration strategy needs to be dynamically optimized, adapting to different contexts and feedback conditions. This optimization relies on a validation signal injected into the preceding sequence. This signal, derived from evaluation feedback in the dynamic reward space, contains information across multiple evaluation dimensions, such as factual accuracy, logical consistency, and linguistic fluency, and is used to quantify the quality of the current generated result.

[0061] First, the model's original high-dimensional action space, typically represented as a vocabulary or token-level selection set, is large in dimensionality and expensive to explore. To improve training efficiency and the effectiveness of the control policy distribution, this action space needs to be mapped to a low-dimensional manifold space. This mapping can be constructed based on word embedding space, sentence latent vector space, or semantic space reduced in dimensionality through principal component analysis, ensuring that the action set retains the important characteristic information of the generation path in the new space.

[0062] Verification signals often contain implicit information about the distribution of model behavior preferences, such as whether certain types of generated behaviors frequently trigger low or high-scoring rewards. To leverage this information to dynamically constrain policies, a policy entropy constraint bound is constructed within this low-dimensional manifold space. This bound is based on the entropy trends of historical generated behaviors and serves to limit the distribution of model-generated policies, preventing them from becoming too divergent or overly concentrated.

[0063] Within the scope of satisfying the entropy constraint, boundary conditions are constructed for the action space, and a subspace that meets the desired policy quality is selected to form a constrained action space. This subspace contains the set of candidate actions that are most likely to produce high-quality outputs under the current context and reward feedback, providing a limited range for subsequent optimization.

[0064] A probabilistic graphical model is built in this constrained action space to model the semantic dependencies between actions and the selection probability structure. This graphical model can be a directed acyclic graph (DAG), a Markov decision graph (MDP), or a conditional spanning graph (e.g., a Gibbs Sampling network), with edge weights reflecting the priority transition probabilities between different actions.

[0065] Finally, based on this graph model, the exploration paths generated during the original training process are pruned, removing low-probability and low-quality paths and retaining only the set of high-value paths. This results in an optimized exploration strategy. This strategy dynamically converges to a generation path that meets the evaluation criteria of multiple reward dimensions, thereby improving the quality, stability, and efficiency of the model's generation results.

[0066] The construction of a low-dimensional manifold space can be achieved by extracting the principal components of the language embedding space before training, such as screening the first 100 principal components in the BERT embedding space to form the embedding base. The policy entropy constraint can be defined as the maximum / minimum policy entropy difference in a rolling window. Exceeding the threshold triggers the reconstruction of the action space. The boundary construction of the constrained action space can be achieved by using k-means clustering and selecting the cluster center as the high-quality action node. The establishment of a probabilistic graph model can use an attention-weighted graph neural network to model the semantic similarity between actions, while injecting a verification signal as a node edge weight adjustment factor. The pruning operation uses a heuristic search algorithm (such as A*) combined with a reward signal to filter out low-scoring paths, retaining only the path with the maximum cumulative reward value as the current update policy.

[0067] Example: In the healthcare sector, when summarizing patient conversations, the generative model must accurately extract changes in the patient's condition and medical measures across multiple rounds of interaction. By verifying the quality of the current output in terms of medical terminology, logical order, and information completeness, constructing a compressed semantic space, and optimizing the output path based on policy entropy, the model prioritizes paths that accurately express medication regimens and disease descriptions, avoiding the generation of redundant or inappropriate content.

[0068] In the financial sector, for example, in models generating intelligent compliance Q&A content, optimized exploration strategies enable the model to accurately retain the core logic when handling regulatory inquiries and multi-layer rule interpretation scenarios, shielding against risky deviations from the answer path. Strategies formed after pruning verification signals can more quickly converge on answer paths that meet factuality, consistency, and compliance requirements during the generation process, thereby improving the model's applicability and approval rate in financial scenarios.

[0069] This embodiment introduces verification signals into the exploration strategy optimization process, eliminating the reliance of strategy updates solely on static loss feedback. Instead, it dynamically adjusts the generation path by integrating real-time quality assessment results, significantly improving the responsiveness and quality stability of the generation model. Low-dimensional mapping and policy entropy bounds effectively compress the action space, reducing the computational overhead of ineffective exploration. Combined with graph structure modeling and path pruning, generation accuracy is improved without sacrificing generation diversity, addressing the exploration-exploitation imbalance in traditional strategy optimization.

[0070] S50, based on the real-time verification loop and the optimized exploration strategy, performing dual-loop feedback control on the training process of the generative model to dynamically adjust training parameters and generate a target model;

[0071] In this embodiment, training optimization of the generative model requires not only minimizing static loss but also achieving real-time responsiveness to dynamic generation errors and policy changes. To this end, a dual-loop feedback control mechanism is introduced, enabling the training process to simultaneously achieve rapid response and structural adjustment capabilities. The real-time verification loop provides continuous feedback signals, while the optimized exploration strategy provides stable behavioral guidance. Together, they form a linked control system that dynamically adjusts training parameters and gradually converges to the target model.

[0072] The first feedback path establishes a fast parameter adjustment loop, relying on a validation signal injected into the real-time validation loop to evaluate the immediate training error of the current model output during each training round. This fast adjustment loop requires high responsiveness and is suitable for fine-grained parameter adjustments such as the learning rate and weight updates during gradient descent. To achieve high sensitivity and stability, a proportional-integral-derivative (PID) control mechanism is introduced to jointly calculate a fast control signal based on three components: the current error (P), the accumulated error (I), and the rate of change of the error (D).

[0073] The second feedback path constructs a slow architectural optimization loop, which periodically updates the model's structural parameters based on the optimized exploration strategy. These parameters include network depth, number of attention heads, token pruning threshold, batch size, and other training architecture parameters. This loop changes less frequently than the fast loop but can make global structural corrections when the strategy direction shifts or task characteristics change. The slow loop uses a Bayesian optimization algorithm to construct a response surface in the parameter space based on the reward feedback distribution of the exploration path. It then selects the hyperparameter configuration combination with the best convergence trend to output a slow control signal.

[0074] To avoid output signal conflicts between the two loops or oscillations caused by feedback superposition, a training stability guarantee function is designed as a constraint for training parameter updates. This stability guarantee function regularizes the amplitude, direction, and frequency variation range of the fast and slow control signals, constraining their weight adjustment range and gradient fluctuations, thereby forming an effective fusion control range.

[0075] Fast and slow control signals are jointly modeled with stability constraints to generate a fused control signal, which is used to uniformly adjust all training parameters of the current generative model. This fused control signal acts in real time on the various optimizers, schedulers, and gradient processing units in the training engine through a dynamic parameter adjustment mechanism, generating an updated set of training parameters. Based on this parameter set, the model is trained and updated in successive rounds, ultimately generating a target model with stable convergence and high consistency of verification signals.

[0076] The PID controller in the fast loop can automatically adjust the three parameters Kp, Ki, and Kd by setting a training error threshold. If the error fluctuates significantly, the proportional term can be weighted higher; when the error trend is stable, the integral term can be strengthened to accelerate convergence. The real-time training error can be composed of the loss function output and the validation signal matching score. The Bayesian optimization algorithm in the slow loop can use a Gaussian process as a surrogate model to evaluate the potential benefit of each architecture parameter combination in the policy feedback space and periodically sample and adjust the architecture (e.g., every 10 rounds). A stability constraint function sets the maximum gradient change rate and parameter fluctuation amplitude, and enforces boundary constraints through a gradient clipping mechanism. The final fused control signal can be generated using a gated aggregation mechanism, which weights and superimposes the fast and slow control signals and maps them to specific parameter update commands.

[0077] Example: In a healthcare scenario, when training a model for generating electronic medical record summaries, the fast loop can quickly correct generation deviations in key entity information such as drug names and dosages through real-time error feedback. Meanwhile, the slow loop periodically fine-tunes the attention structure based on reward feedback for semantic coherence and medical terminology coverage, ensuring that the model output better aligns with doctors' writing habits and factual description requirements.

[0078] In FinTech, when training a model for contract clause generation, the fast loop can quickly respond to issues such as logical inconsistencies or incorrect amounts, while the slow loop adjusts the encoder depth and maximum token length based on compliance verification signal analysis to optimize the completeness and compliance structure of the generated clauses. The synergistic effect of these two loops makes the model more accurate and stable when processing complex compliance expressions, effectively adapting to the demanding generation tasks in the financial sector.

[0079] This embodiment significantly improves the responsiveness and policy consistency of the model training process through a dual-loop feedback control mechanism. The fast loop can swiftly respond to immediate error fluctuations during training, improving the stability of gradient adjustment; the slow loop optimizes the overall policy path from a longer-term and structural perspective. The two types of loops work together to effectively address the problems of policy adjustment lag or local oscillation under a single control path, enhancing the convergence and target adaptability of the training process, and ensuring that the resulting model is stable, accurate, and generalizable across multiple quality dimensions.

[0080] S60: Outputting an inference result based on the target model.

[0081] In this embodiment, the output of the inference result is based on the trained target model, which is obtained by dynamically adjusting the training parameters under the aforementioned dual-loop feedback control mechanism, and is therefore highly adaptable, stable, and context-aligned. In actual execution, the inference process first needs to format the input data to be processed into a data vector that is consistent with the input structure of the target model, such as text tokens, structured indicators, or embedded vector representations. The input data should cover the scope of the model's contextual understanding and be segmented or batched according to a preset format to support model parallel computing.

[0082] When the target model performs inference, it propagates the input data forward through each layer of the model structure, including but not limited to the embedding encoding layer, the multi-head attention mechanism, the cross-layer normalization module, and the final output decoder. At each layer, the target model combines the parameter weights formed during training to perform nonlinear transformations and semantic aggregation on the input features until the final output is generated. The output format varies depending on the task type, such as text sequences in natural language generation tasks, structured fields in knowledge extraction tasks, and logical judgment values ​​or confidence scores in inference tasks.

[0083] To ensure the quality of inference output, additional processing logic is typically integrated at the end of the model for post-processing and confidence calibration. This includes temperature-adjusted softmax to control output distribution, confidence threshold clipping, content entity alignment verification mechanisms, and output legitimacy judgment functions. This process ensures that inference results remain consistent with training objectives in terms of expressive completeness, contextual coherence, and logical consistency.

[0084] In particular, some modules that are linked to the verification signal pathway are retained in the structure of the target model. For example, lightweight verification signal feedback can still be enabled during the inference phase to calibrate the risk of semantic drift or consistency collapse in the output process, thereby retaining the model's sensitivity to the original training feedback pathway without increasing the training burden.

[0085] In healthcare scenarios, input data can include patient electronic medical record summaries, time-series vectors of vital signs, or historical medical conversation records. The target model semantically encodes and analyzes this data, generating outputs such as auxiliary diagnostic recommendations, treatment plan summaries, or risk warnings. During the inference process, a specific medical vocabulary normalization module can be integrated to standardize terminology and spell-check the generated content, ensuring that the results meet medical professional requirements.

[0086] In FinTech business scenarios, input data includes investor profiles, responses to risk preference questionnaires, and historical transaction data. After inference, the target model generates risk assessment reports, contract generation recommendations, or automated compliance alerts. Post-processing can combine legal text matching algorithms and local policy database comparison modules to perform compliance review on the generated content, enhancing the regulatory consistency of the inference results.

[0087] The target model can also be deployed as a service-oriented module according to different tasks, which can be called by the business system in real time to form dynamic responses in multiple rounds of interaction, such as automatically generating reply content for customer service systems or forming preliminary conclusion summaries in the audit process.

[0088] Example: In a healthcare scenario, an auxiliary consultation system for primary care doctors receives input from patients regarding their chief complaint, vital signs, and drug allergy information. After receiving these inputs, the target model quickly generates a list of consultation suggestions. The inference output clearly lists the medical history information that needs to be further collected and the recommended preliminary examination items, reducing the doctor's operational burden and improving consultation efficiency.

[0089] In financial services, corporate clients upload financing contracts, historical transaction records, and qualification documents. The target model automatically verifies the contract terms and generates customized optimization suggestions. The inference results include not only a check of the contract's structured fields but also risk warnings compared to similar cases, thereby improving compliance security and the customer experience.

[0090] This embodiment constructs a dynamic feedback mechanism for the target model during the training phase, enabling the model to more fully utilize the structural information and verification signal mapping relationship learned during its training process during the reasoning phase, thereby achieving adaptive optimization of the reasoning output under complex input scenarios. Compared with the traditional one-time static reasoning model, the target model generated by this method has higher reasoning robustness and result consistency. The reasoning results are not only contextually consistent at the semantic level, but also have strong constraints at the structural and compliance levels, significantly improving the practical application value of the model in key tasks.

[0091] The present invention relates to the field of artificial intelligence technology and can be applied to business scenarios such as financial technology and medical health. A generation optimization method, device, equipment and medium based on dynamic verification feedback are disclosed, including: pre-training a verification network for analyzing the quality of generated content; constructing a dynamic reward space containing multiple orthogonal reward components based on the verification network; integrating the dynamic reward space into the generative model to be trained to form a real-time verification loop, and injecting verification signals into multiple processing layers of the generative model; optimizing the exploration strategy of the generative model according to the verification signal to obtain an optimized exploration strategy; performing dual-loop feedback control on the generative model based on the real-time verification loop and the optimized exploration strategy, dynamically adjusting training parameters, and generating a target model; and outputting inference results based on the target model. The present invention realizes a dual-loop feedback control mechanism during the training process by constructing a real-time verification loop and an optimized exploration strategy, injecting verification signals into the model training process, dynamically capturing training errors and strategy deviations, and adjusting training parameters in real time through the synergy of fast response and slow optimization, thereby effectively improving the generative model's responsiveness to multi-objective requirements and the training convergence speed.

[0092] In one embodiment, the above step S10 includes:

[0093] S101, generating an original text sample dataset;

[0094] S102, injecting semantic trap defects, logical conflict defects, and factual error defects into the original text sample dataset;

[0095] S103, marking the error type labels of the text sample dataset after the injection defect;

[0096] S104, constructing a multi-dimensional verification indicator matrix based on the text sample dataset annotated with error type labels;

[0097] S105, executing confidence interval error control of the multidimensional validation indicator matrix to generate an error control parameter table;

[0098] S106: Complete discriminant training of the pre-training verification network based on the error control parameter table.

[0099] In this embodiment, building a verification network for analyzing the quality of generated content requires high-quality data support. First, a dataset of raw text samples must be generated. This dataset can come from historical conversation records, a domain corpus, or initial text generated by a pre-trained model. Its purpose is to provide representative semantic expression, structural logic, and factual coverage, providing semantic space coverage for subsequent defect injection and annotation.

[0100] After constructing the original text sample dataset, semantic trap defects, logical conflict defects, and factual error defects need to be injected into it. Semantic trap defects refer to expressions that introduce semantic ambiguity and pragmatic ambiguity into the text, leading to semantic misunderstanding, such as the use of puns and out-of-context vocabulary. Logical conflict defects involve contradictions or logical inconsistencies between propositions in the reasoning structure, such as a fact being both affirmed and denied simultaneously. Factual error defects refer to information in the text that is inconsistent with the knowledge base, fact database, or objective facts, such as outputting "The Earth has two satellites." These defect injections can be achieved through template perturbation, knowledge inversion, and conditional trigger editing.

[0101] After injecting defects, text samples need to be annotated with error type labels. Each text piece is assigned a label based on the defect type, forming a supervised training sample for the target output of the classification task during subsequent training. This step relies on manual annotation, rule engines, or weakly supervised model-assisted annotation, and sample usage strategies can be set based on different annotation confidence levels.

[0102] A multidimensional validation indicator matrix is ​​constructed based on a dataset of labeled text samples. This matrix maps various error types to multiple quantifiable dimensions, such as semantic consistency scores, fact-checking confidence, and logical coherence probability. Each dimension in the matrix represents an evaluation dimension, forming a structured quality vector representation that provides a structural alignment mechanism for network input and supervision.

[0103] Confidence interval error control is performed on this multidimensional validation indicator matrix to assess the uncertainty and distribution stability of each dimension's indicator. By performing an upper bound analysis on the confidence intervals of the sample scores for each dimension, an error control parameter table with a controllable tolerance range is generated, thereby preventing performance instability of the discriminant model due to extreme sample perturbations during training.

[0104] Finally, discriminative training is performed based on the aforementioned error control parameter table. The training process employs a multi-objective cross-entropy loss function or label smoothing strategy to control the sensitivity and tolerance threshold for various types of errors. The training model must be able to perceive and independently judge errors at different levels within the text, forming a verification network with robust classification and generalization capabilities.

[0105] This implementation systematically injects multiple types of semantic, logical, and factual errors into raw text sample data to construct training samples with a realistic error distribution. Combined with annotation and a multidimensional validation indicator matrix, this method establishes a pre-trained validation network with fine-grained multi-label recognition capabilities. This significantly improves the model's ability to discern text generation quality during training. Leveraging a parameter optimization mechanism for confidence interval error control, the validation network achieves greater stability and adaptability in setting discriminant boundaries, avoiding false positives and missed detections caused by training sample bias. This provides accurate and high-resolution quality feedback signals for subsequent dynamic reward modeling and exploration strategy optimization.

[0106] In one embodiment, the above step S20 includes:

[0107] S201, obtaining basic data of reward components generated by the verification network;

[0108] S202, setting a factual accuracy reward component and a logical consistency reward component based on the reward component basic data;

[0109] S203, performing orthogonalization processing on the fact accuracy reward component and the logical consistency reward component to generate an orthogonalized reward component;

[0110] S204, establishing a dynamic weight distribution model based on the orthogonalized reward components;

[0111] S205 : Based on the orthogonalized reward components and the dynamic weight distribution model, a reward projection matrix is ​​constructed to form a dynamic reward space.

[0112] In this embodiment, when constructing a dynamic reward space containing multiple orthogonal reward components, the first step is to obtain the basic reward component data generated by the verification network. This basic data is derived from the verification network's analysis and scoring results of the generated text. It is expressed as a set of numerical data on quality assessment scores for different dimensions, such as semantic plausibility confidence, fact-checking probability, and logical consistency judgment results. This data is generally stored as a multidimensional vector, with each dimension mapped to a potential quality assessment factor.

[0113] On this basis, two core reward components are designed to capture the quality of content generation: factual accuracy and logical consistency. The factual accuracy component reflects whether the generated text is consistent with background knowledge, knowledge bases, or contextual facts. It can be calculated through mechanisms such as fact-checking tools, knowledge graph matching, or entity consistency detection. The logical consistency component is used to assess whether the generated text possesses structural integrity, reasoning closure, and contextual coherence. It is typically extracted through natural language inference models and context prediction accuracy. These two components represent two complementary dimensions of generative model output quality.

[0114] Orthogonalizing these two reward components is intended to enhance the independence of the reward signals, preventing a high score in one dimension from masking a deficit in another. Orthogonalization can be performed using principal component analysis (PCA), singular value decomposition (SVD), or normalization methods based on gradient independence to ensure that the two components are independently represented in feature space without redundancy, thereby improving the expressiveness and stability of the reward mechanism in multi-objective tasks.

[0115] Based on the obtained orthogonalized reward components, a dynamic weight distribution model is further established. This model is used to dynamically adjust the weight distribution of each reward component in the total reward based on the current training state, historical reward trends, and the evolution of the model's strategic behavior. For example, during the model exploration phase, logical consistency signals can be appropriately amplified, while during the model convergence phase, factual accuracy can be strengthened. The weight distribution model can be implemented using a temporal reinforcement learning strategy, an attention mechanism with memory, or a historical variance scheduling strategy, allowing for adaptive regulation of the reward space and tailoring it to the needs of the training phase.

[0116] Finally, by combining orthogonalized reward components with a dynamic weight distribution model, a reward projection matrix is ​​constructed to achieve high-dimensional mapping and unified output of reward signals during the training process. This projection matrix is ​​a tensor representation structure containing dynamic weights and orthogonal components. It is used to guide the loss function calculation or reward distribution sampling mechanism in the training strategy, thereby forming a complete dynamic reward space. This reward space is differentiable, real-time adjustable, and expression-independent, becoming a foundational element for subsequent exploration strategy optimization and the construction of training feedback loops.

[0117] This embodiment converts the quality scores output by the verification network into a multidimensional reward signal with an orthogonal structure and introduces a dynamic weight adjustment mechanism to construct a reward projection matrix. This allows for controllable decomposition and enhanced precision of quality feedback during generative model training, effectively addressing the limitations of traditional static reward mechanisms in terms of expressiveness, adaptability, and objective balance. The dynamic reward space not only improves signal resolution during training but also enhances the model's responsiveness to multi-objective constraints through orthogonal components and a dynamic scheduling mechanism. This provides an adjustable, directional, and directional evaluation foundation for exploring strategy optimization and training feedback control.

[0118] In one embodiment, the above step S30 includes:

[0119] S301, selecting a specific middle layer of the generative model to be trained as a signal injection position, and generating a specific middle layer position;

[0120] S302, mapping the dynamic reward space into a verification signal through a gating mechanism;

[0121] S303, injecting the verification signal at the specific intermediate layer position to generate an intermediate layer into which the verification signal is injected;

[0122] S304, designing a dual-path gradient propagation channel based on the intermediate layer injected with the verification signal;

[0123] S305, performing hardware-accelerated signal fusion processing based on the dual-path gradient propagation channel to generate an accelerated fusion processing result;

[0124] S306: Based on the accelerated fusion processing result, the construction of the real-time verification loop is completed.

[0125] In this embodiment, the process of integrating the dynamic reward space into the generative model is intended to achieve a high-frequency, fine-grained quality feedback mechanism. First, it is necessary to identify one or more intermediate layers from the generative model structure to be trained as the injection location of the feedback signal. This intermediate layer is usually located at the intermediate computing node of the Transformer, LSTM or other neural network structure, such as the attention layer, the output point of the feedforward network, or the tensor state before the residual connection. By analyzing the feature expression ability and gradient sensitivity of the layer, the position most suitable for feedback intervention is determined to form a specific intermediate layer position.

[0126] A gating mechanism is then used to map the dynamic reward space into an injectable verification signal. This gating mechanism determines the signal form and strength based on the current reward component and the model state. It is essentially a differentiable control function, which can be a gated recurrent unit (GRU), a multiply-add network with sigmoid-controlled weights, or a conditional transformation module. It outputs a set of tensor-based verification signals for multi-dimensional reward inputs. These signals not only encode the reward direction but also preserve the control properties of dynamic weight scheduling.

[0127] The generated verification signal is injected into a previously selected intermediate layer to obtain the intermediate layer state. This injection can be done through vector-level concatenation, weighted fusion, or contextual enhancement guided by an attention mechanism. This allows the original intermediate representation to incorporate quality feedback signals without disrupting the model's semantic construction process, thereby enhancing the representation layer's sensitivity to reward partial derivatives.

[0128] After injection, a dual gradient propagation channel is designed based on the intermediate layer state, transmitting gradient information from the main loss function and the verification signal branch respectively. The first channel maintains the forward learning path of the main task, while the second channel carries the quality feedback path. During training, the two channels are jointly propagated to form a collaborative update mechanism that combines semantic representation and quality guidance. This channel structure establishes a trade-off mechanism through structural separation and parameter sharing, avoiding signal interference while allowing gradient interaction.

[0129] To enhance training efficiency and feedback response speed, hardware-accelerated signal fusion processing is introduced. This fusion process is implemented through parallel matrix computation, tensor acceleration instructions, or customized graph processor architectures (such as Tensor Processing Units (TPUs) and CUDA cores). The gradient signals and intermediate representations generated by the two channels are uniformly processed on the accelerated hardware to generate an accelerated fusion result. This result reflects the quality response changes in the model's current representation state and serves as a direct input for dynamic training feedback.

[0130] Finally, a real-time verification loop is constructed based on the fusion processing results. This closed-loop structure allows the model to obtain dynamic quality feedback on the current output after each round of forward propagation and incorporates this feedback into the subsequent training path. This allows the model to continuously correct parameter trends during training, thereby enhancing exploration capabilities and output accuracy. This loop structure supports mechanisms such as periodic updates, adaptive learning rate adjustment, and online policy switching, achieving a coupled evolution of quality control and structural updates.

[0131] This embodiment builds a verification loop based on dynamic rewards and injects verification signals into multiple processing layers of the generative model. This not only enhances the model's real-time responsiveness to reward feedback, but also improves training effectiveness and computational efficiency through a dual-path gradient path and accelerated fusion mechanism. This structure continuously outputs directional feedback information during training, effectively avoiding the optimization path rigidity caused by a fixed loss function. Dynamic control enhances the diversity of expression and quality controllability during the exploration process, enabling the model to adaptively perceive and adjust to training errors and changes in reward strategies.

[0132] In one embodiment, the above step S303 includes:

[0133] S3031, based on the verification signal and the specific intermediate layer position, perform causal verification of the micro-generation unit to generate a micro-layer verification result;

[0134] S3032, based on the verification signal and the specific middle layer position, performing coherence verification of the meso-level semantic unit to generate a meso-level window verification result;

[0135] S3033, performing macro-session-level value alignment verification based on the verification signal and the specific intermediate layer position, and generating a macro-level verification result;

[0136] S3034, fusing the micro-level verification result, the meso-level window verification result, and the macro-level verification result to generate a layered fusion verification signal;

[0137] S3035: Inject the layered fusion verification signal into the specific middle layer position to generate an middle layer for injecting the verification signal.

[0138] In this embodiment, within the generative model's training structure, injecting verification signals into specific intermediate layers and shaping the intermediate layer state through a layered verification mechanism is a key step in achieving multi-granular dynamic quality supervision. This process uses specific intermediate layer locations as anchors and, combined with externally generated verification signals, performs semantic quality analysis at three different levels of abstraction, achieving comprehensive verification across micro, meso, and macro semantic dimensions.

[0139] Micro-level verification primarily addresses the correctness of local causal structures. At this stage, verification signals are combined with specific intermediate-layer tensors and fed into a lightweight causal discriminator module, typically employing an attention-based token sequence evaluator. This module analyzes the dependency structure between the currently generated token and the contextual history to determine whether any output behavior violates the linguistic causal structure, such as prematurely outputting conclusions before introducing previous concepts. Micro-level verification results are output by determining the consistency of conditional probabilities and the continuity of the causal logic chain.

[0140] Meso-window verification assesses the coherence of semantic units at the sentence or phrase level. In this process, a sliding window structure is constructed to extract semantic segments from the intermediate layer representation tensor at a fixed step size, and a verification signal is injected as a bias into the embedded representation within the window. The coherence verification module, which can be a local LSTM network or a bidirectional Transformer encoder, determines the naturalness of semantic transitions and the continuity of information based on representation consistency metrics (such as cosine similarity and KL divergence) across each window, generating a meso-window verification result.

[0141] Macro-level verification focuses on consistency of value orientation at the conversational level or across segments. This verification compares the current model generation path with historical high-value paths. Based on the target value direction indicated by the verification signal, it assesses whether the output aligns with the system's expectations or pragmatic goals. The comparison mechanism employs a contrastive learning architecture, mapping the current intermediate-layer representation to positive and negative samples in vector space. This outputs macro-level verification results by distancing negative samples and bringing them closer to positive samples.

[0142] After completing the three levels of verification, the system performs a hierarchical fusion operation. This fusion process does not rely on simple averaging, but rather a weighted combination of the verification signals at each level, based on their stability, information gain, and structural interference, forming a multi-component fusion strategy. This strategy can be implemented through dynamic weighting functions, attention aggregation networks, or learnable gating structures. The output is the hierarchical fusion verification signal.

[0143] Reinjecting the layered fusion verification signal into a specific intermediate layer is the final step in closing the semantic feedback loop. This injection process must be consistent with the initial injection path to ensure that the fused signal incorporates structural guidance without destroying the model's original representation. This fused signal acts as an additional condition in the subsequent propagation, influencing the gradient distribution and guiding the model to prioritize high-quality semantic structures during generation.

[0144] This embodiment constructs and injects fused multi-level verification signals into specific intermediate layers of the generative model. This allows the system to obtain dynamic feedback on local causality, sentence coherence, and overall value consistency during the training phase, significantly improving the model's adaptive perception of semantic quality. Joint verification at the micro, meso, and macro levels not only enables a detailed assessment of generation behavior at different levels of abstraction, but also transforms verification information into training control signals through a layered fusion mechanism, avoiding the problem of unbalanced feedback at a single scale, thereby improving the model's generalization, stability, and adaptability to task objectives.

[0145] In one embodiment, the above step S40 includes:

[0146] S401, mapping the original high-dimensional action space of the generative model to a low-dimensional manifold space to generate a low-dimensional manifold action space;

[0147] S402, establishing a policy entropy constraint boundary based on the historical entropy value in the verification signal;

[0148] S403, solving the action space boundary conditions that satisfy the policy entropy constraint boundary in the low-dimensional manifold action space to generate a constrained action space;

[0149] S404, constructing a probability graph model in the constrained action space;

[0150] S405 , performing a pruning operation on the exploration path generated by the generative model during the training process based on the probabilistic graphical model to generate an optimized exploration strategy.

[0151] In this embodiment, the generative model has a wide range of action selection space during the training process, while the high-dimensional action space often has too much redundancy and inefficient paths, which affects the exploration efficiency. Therefore, mapping the original high-dimensional action space of the generative model to a low-dimensional manifold space helps to retain the potential effective action structure while compressing meaningless dimensions. This process can be achieved through manifold learning algorithms, such as using embedding structures such as variational autoencoders (VAE) or t-SNE to convert high-dimensional policy vector representations into controllable, dense low-dimensional representations. The low-dimensional manifold action space reduces the complexity of sampling and learning while retaining the semantic structure.

[0152] To improve the balance of information entropy during action selection, a policy entropy constraint mechanism is introduced in the low-dimensional manifold action space. By analyzing the historical entropy values ​​in the verification signal, the entropy distribution trajectory of the model's corresponding action selections during each training round is extracted, and a policy entropy constraint boundary is established. This boundary is constructed based on the mean and variance of the entropy in the sliding window, defining the maximum and minimum allowable policy uncertainty ranges, thereby preventing the model from falling into extreme determinism or overexploration.

[0153] After setting the policy entropy constraint bounds, we need to find a valid action region that satisfies these bounds in the low-dimensional manifold action space. This process uses the intersection of the bounds function and the current policy mapping function to obtain a set of action boundaries, confining the entire action space to a controllable region, forming a constrained action space. This constrained action space reflects the acceptable range of behavioral variation in verification feedback and serves as the foundation for subsequent exploration path modeling.

[0154] Constructing a probabilistic graphical model in a constrained action space is intended to provide a structured basis for path selection in generative models during policy optimization. This construction can employ a Markov decision process (MDP) or Bayesian network, using each action state as a node and the transition probabilities driven by verification signals as edge weights to form a traversable policy graph. Each path in the graphical model represents a possible evolutionary trajectory of the generated sequence. By comparing the cumulative reward under multiple paths with the policy entropy, a pruning principle is provided.

[0155] Pruning exploration paths generated during training based on a probabilistic graphical model is a key step in achieving policy optimization. The pruning logic dynamically removes inefficient paths based on criteria such as cumulative reward falling below a threshold, policy entropy deviating from constraint boundaries, and low validation signal feedback. The remaining paths are retained for the next round of policy iteration. The resulting retained paths constitute the optimized exploration strategy, which effectively reduces inefficient searches during training while ensuring exploration diversity and consistent quality feedback, thereby improving training efficiency and generalization.

[0156] Example: Consider a language model used in a text generation task. During each generation cycle, the model must select the next output word from a vocabulary containing tens of thousands of terms. The model's action space dimension is the size of the vocabulary. If the vocabulary contains 50,000 tokens, then the action selection at each moment corresponds to a 50,000-dimensional probability distribution vector. This high-dimensional space not only contains a large number of semantically irrelevant or redundant tokens but also introduces significant complexity and sparsity that impacts training stability and policy gradient propagation. By introducing low-dimensional embeddings or semantic compression mechanisms, such as using a variational autoencoder (VAE) to reduce the dimensionality of the generated policy, the 50,000-dimensional action space can be compressed into a latent semantic vector space of 128 dimensions. In this 128-dimensional space, each action representation not only exhibits semantic clustering (e.g., "doctor," "nurse," and "patient" are compressed into adjacent positions), but also facilitates policy modeling and search, avoiding meaningless exploration in the sparse high-dimensional space. The model can select directions for semantic generation on this dense low-dimensional manifold, thereby improving exploration efficiency and sample utilization.

[0157] This embodiment embeds the high-dimensional policy space into a low-dimensional manifold space and, in conjunction with verification signal feedback, constructs a policy entropy boundary and probabilistic graphical model. This achieves dynamic pruning optimization of the generative model's exploration strategy, making the model training process more goal-oriented and feedback-responsive. This mechanism effectively balances the contradiction between exploration and exploitation, preventing the strategy from falling into local extremes during training, such as deterministic convergence or disordered expansion, while also improving resource utilization and the overall stability of generation quality.

[0158] In one embodiment, the above step S50 includes:

[0159] S501, establishing a fast parameter adjustment loop based on the real-time verification loop to respond to the real-time training error and generate a fast control signal;

[0160] S502, constructing a slow architecture optimization loop based on the optimized exploration strategy to adjust network hyperparameters and generate a slow control signal;

[0161] S503, setting a training stability guarantee function constraint parameter update process to generate stability constraints;

[0162] S504, based on the fast control signal, the slow control signal and the stability constraint, achieving collaborative fusion of the dual-loop control signals to generate a fused control signal;

[0163] S505, dynamically adjusting the training parameters of the generative model based on the fusion control signal to generate updated training parameters;

[0164] S506: Generate an optimized target model based on the updated training parameters.

[0165] In this embodiment, during the training process, a fast parameter adjustment loop is constructed to respond to real-time error changes in model training. This mechanism relies on the verification signal injected into the real-time verification loop to periodically collect the error signal between the model's predicted output and the target output during training iterations. Based on these error values, a feedback function is constructed to dynamically adjust the model's parameter update rate and amplitude through a proportional-integral-differential (PID) control mechanism. PID control has the characteristics of fast response and error stability, which enables the model to make corrections quickly when the error fluctuates, thereby improving the robustness and response sensitivity of the training process.

[0166] On the other hand, to achieve more structured policy optimization, a slow architectural optimization loop must be constructed based on the optimized exploration strategy. This loop is used to periodically evaluate and adjust model structure-level hyperparameter configurations during training, such as the number of attention heads, hidden layer width, dropout rate, or the optimizer's learning rate decay strategy. This slow loop uses Bayesian optimization as its primary scheduling mechanism. It establishes a nonlinear mapping function based on historical training trajectories and validation feedback, predicts the impact of current structural adjustments on overall performance, and then samples the optimal hyperparameter combination based on this function. Bayesian optimization provides a computationally efficient exploration method that avoids the high cost of traversal search while preserving the trend toward global optimality.

[0167] To ensure that parameter updates during training do not cause gradient explosion, model degradation, or unstable convergence, a training stability guarantee function is designed to constrain the overall training behavior. This function can be composed of multiple stability indicators, such as gradient norm constraints, parameter increment limits, overfitting risk monitoring, and policy entropy drift detection. These indicators constitute a set of soft constraints during training, imposing bounds on parameter updates and preventing them from deviating from a reasonable training trajectory.

[0168] The fast parameter adjustment loop, the slow architecture optimization loop, and the stability guarantee function represent the three information paths of immediate feedback, policy update, and structural homeostasis, respectively. Within the training system, these three paths must be fused together through a unified collaborative fusion mechanism to form a fused control signal. This fused control signal can be normalized and integrated through weighted averaging, an attention mechanism, or a meta-learning strategy to generate a unified instruction set that guides the next step of training parameter updates. This process must simultaneously consider the time delay of the control signal, weight adaptation, and target consistency to ensure that the overall optimization direction is consistent and oscillating.

[0169] Ultimately, the fused control signal serves as the driving input for parameter updates, dynamically adjusting the model's training parameters. This parameter adjustment process encompasses the backbone network parameters, optimizer parameters, and regularization parameters. Through iterative updates, the optimal parameter set for the training state is generated. From this foundation, one or more fine-tuning training cycles are performed to produce the final target model. This target model incorporates a fusion feedback mechanism and structural adaptability, resulting in improved generalization performance and output stability.

[0170] For example, in large language model training, the text generated after each iteration is compared with the reference answer. If phenomena such as a high factual error rate, increased repetitive content, or illogical language appear over a short period of time, the verification signal in the real-time verification loop will shift significantly, manifesting as fluctuations in the training error metric. At this point, the fast parameter adjustment loop immediately responds to these error signals, rapidly adjusting the learning rate for the current iteration through the PID control mechanism. For example, temporarily lowering the current learning rate from 1e-4 to 5e-5 can prevent violent fluctuations in model parameters caused by high gradients. This mechanism can also trigger dynamic adjustments to the short-term regularization coefficient, such as temporarily increasing the dropout ratio to suppress overfitting trends. This type of feedback adjustment is typically completed within one or several batches, with a response frequency of sub-second to second levels.

[0171] Similarly, during model training, if it is observed that the training accuracy tends to saturation but the performance of the validation set has not improved for a long time, it may indicate that the current model structure parameter configuration (such as the width of the hidden layer or the number of attention heads) is insufficient to support further generalization. At this time, the slow architecture optimization loop uses the historical training data and verification signal trends to build a Bayesian optimization process based on the optimized exploration strategy to make low-frequency adjustments to the model structure. For example, the structural parameter combination is re-evaluated every 5 epochs, and the current optimal combination is selected from multiple candidate hyperparameter combinations (such as {attention heads: 8, 12, 16}) for replacement, and several rounds of warm-up training are restarted. This process usually takes thousands of training steps to complete a complete evaluation and structural adjustment, and is a low-frequency optimization process at the minute to hour level.

[0172] Example: In a financial business scenario, to improve the compliance department's efficiency in generating regulatory filings, client notifications, and policy interpretation documents, a generative language model training system with a dynamic verification mechanism was deployed. The system aims to build a model for automatically generating financial compliance documents with factual accuracy, logical consistency, and linguistic standardization. The implementation process is as follows:

[0173] First, the system constructs a dataset of raw text samples by integrating historical financial policy texts, compliance audit documents, and actual reporting materials. The data engineering module then introduces three typical flaw types into each text: injecting ambiguous definitions, conflicting clauses, and inaccurate clauses, which constitute semantic traps, logical conflicts, and factual errors, respectively. Based on this, the annotation module assigns a corresponding error type label to each defective segment. Through manual review and verification with the audit rules engine, a high-confidence annotated sample set is constructed.

[0174] Based on this annotated sample set, a financial text verification network was trained. This network takes as input a multidimensional verification indicator matrix, including scoring dimensions such as factual verification, clause consistency, and logical chain closure. Using confidence interval error control techniques, the network adjusts the precision of parameters in each dimension, ultimately generating an error control parameter table capable of executing discriminant training. After training, the verification network possesses the ability to perform fine-grained defect analysis on generated content and outputs multiple types of reward component data.

[0175] Next, the system calls the verification network to perform scoring operations on the initially generated candidate sequences of compliant texts, obtaining basic scoring data for each text across different dimensions. Based on this data, the strategy modeling module constructs two reward components: factual accuracy and logical consistency. These components are then orthogonalized using principal component analysis and orthogonal transformation methods to generate two spatially independent orthogonalized reward component axes. Furthermore, the system incorporates dynamic weight adjustment models from the financial field, such as adaptively increasing the weight of the factual accuracy dimension based on quarterly audit priorities, thereby constructing a dynamic weight allocation model. Finally, using the orthogonalized reward components and weight model as input, a reward projection matrix is ​​constructed to generate a dynamic reward space.

[0176] This dynamic reward space is integrated into the generative language model being trained through a specific mechanism. The system analyzes multiple intermediate layer nodes in the Transformer architecture and selects specific intermediate locations, such as the output of the multi-head attention layer or the activation points of the feedforward network layer, as the points for injecting verification signals. At these locations, a gating mechanism is used to map the dynamic reward space into a formalized verification signal, which is then injected into the network.

[0177] After the injection is complete, the system performs three levels of verification tasks at each intermediate layer. The micro-level verifies the causal relationship between sentence-level clauses, such as whether the responsibility and exemption logic of the clauses are self-consistent; the meso-level analyzes the coherence of the entire policy explanation to prevent jumps in meaning or unclosed clause structures; the macro-level analyzes the value alignment of the entire text between regulatory objectives and factual support, such as whether the text accurately expresses the core viewpoints of the regulatory guidance. After the three-layer verification signal is fused, it is used to guide the current model to fine-tune the forward propagation output of this layer, while providing a gradient path for backpropagation to complete the signal fusion processing.

[0178] During model training, the system dynamically optimizes the generative model's exploration strategy by extracting feedback information from verification signals. The original high-dimensional action space encompasses dozens of potential operations, including legal citation methods, policy interpretation logic, and wording style. This is mapped to a low-dimensional action manifold space using the t-SNE manifold learning algorithm, and the policy entropy constraint boundary is set based on the entropy distribution trend in historical training rounds, such as controlling the degree of ambiguity and the range of semantic divergence in policy interpretation generation. On this basis, an action boundary function is constructed to generate a constrained action space, and a probabilistic graphical model based on a Bayesian network is constructed to model possible generation paths. If a path receives reward feedback below the entropy threshold in three consecutive rounds of training and deviates from the policy boundary, the path is pruned to improve the efficiency and quality of the generated path.

[0179] Finally, a dual-loop feedback mechanism is introduced during training to achieve dynamic parameter optimization. The fast feedback loop adjusts parameters such as the current learning rate and gradient clipping coefficient in real time based on the training errors fed back in the real-time verification loop, such as the surge in unreasonable reference terms during model generation. The slow loop periodically optimizes network structure hyperparameters, such as the depth of the Transformer layer or the configuration of the attention heads, based on the optimized exploration strategy. The dual loops form consistent parameter update instructions through a fusion control mechanism, ensuring that the training process strikes a balance between rapid convergence and long-term generalization. Combined with training stability constraint functions, such as the maximum perturbation amplitude limit for parameter updates, the convergence stability of the training phase is further improved.

[0180] Once trained, the generative model is able to automatically generate compliant documents that adhere to financial regulatory policies, given regulatory requirements, customer background, and data input. For example, given a policy summary regarding loan interest rate adjustments, the model can automatically generate a compliant customer notification letter, encompassing a summary of the policy text, the scope of impact, recommended actions, and customer rights information. The model then uses a verification module to generate a factual consistency score report for reviewers' reference, significantly improving the efficiency and accuracy of compliant document compilation.

[0181] In healthcare scenarios, to reduce the paperwork burden on doctors, improve medical record quality, and enhance clinical decision support, a generative language model with a dynamic validation loop was constructed to generate structured medical record summaries and patient health advice. This model must meet clinical standards across multiple dimensions, including output accuracy, logical consistency, and patient-friendly understanding.

[0182] First, the system constructs a dataset of raw text samples based on historical electronic medical record data, structured examination reports, and medical order records from the hospital information system (HIS). To train the verification module with defect recognition capabilities, the system introduces three types of simulated errors: intentional obfuscation of diagnostic results, introduction of contradictory medication information, and statements with inconsistent etiology-treatment logic. These correspond to semantic trap defects, logical conflict defects, and factual errors, respectively. The medical annotation system performs multi-label annotation according to ICD coding standards and guidelines, including error category, severity level, and potential impact on clinical intervention.

[0183] Based on this annotated dataset, the system trains a medical validation network. This network constructs a multidimensional validation indicator matrix, encompassing dimensions such as cause-symptom consistency, rationality of treatment pathways, and the technicality and accessibility of information. To ensure the statistical stability of the model's judgments, the system applies confidence interval error control to the evaluation outputs of each dimension, sets error tolerances, and generates an error control parameter table, ensuring that the model's judgments remain reliable even in small sample sizes and in edge cases.

[0184] During training, the system uses the verification network output as the incentive signal source to construct a dynamic reward space. Based on the scores of the output text across two key dimensions: cause alignment and treatment plan rationality, the system designs reward components for factual accuracy and logical consistency, respectively. To prevent interference between incentives across multiple dimensions, the system employs principal component axis rotation and orthogonal projection mechanisms to perform orthogonalization. The system also dynamically adjusts weights within the hospital's current medication regulations or guideline update cycle to construct a weight distribution model. The resulting reward projection matrix is ​​used to construct the dynamic reward space.

[0185] This reward space is integrated into the generative medical text model being trained in real time. The system analyzes key intermediate layers in the BERT or Transformer architecture, such as the encoder terminal output and the cross-attention layer, as signal injection locations. Based on a gating function, the system maps tensors in the reward space into verification signals consistent with the intermediate layer structure and injects them synchronously into multiple intermediate layers as needed.

[0186] To improve feedback quality, the injection process simultaneously performs three layers of verification: micro-level causal verification of drug dosage descriptions and unit compatibility; meso-level verification of coherence across multiple paragraphs to ensure the correct transmission of diagnostic evidence; and macro-level verification of global alignment of the entire generated text with the given ICD diagnostic criteria. These three types of verification results are integrated into a unified hierarchical verification signal via a fusion network, driving a dual-path gradient feedback channel for local model correction and policy adjustment.

[0187] The model training strategy is further optimized on this basis. The original model faces high-dimensional action spaces containing different symptom combinations, past medical history, multiple medication constraints, etc., and has problems with sparse path distribution and low learning efficiency. The system uses VAE (variational autoencoder) to learn the manifold structure and compress high-dimensional actions into a low-dimensional manifold action space. In this space, the historical policy entropy sequence of verification signal feedback is extracted, and the control upper and lower limits are set to construct the policy entropy constraint boundary. Then, the constrained action space under the boundary is solved, and a policy graph based on Markov transition probability is constructed to simulate the coverage of patient scenarios by different generated paths. Pruning operations based on policy entropy offset and reward component sparsity are performed on the paths in the graph to eliminate inefficient paths and generate new exploration strategies.

[0188] Combining an optimized exploration strategy with a real-time verification loop, the system introduces a dual-loop feedback control mechanism for the entire training process. The fast loop verifies the network's output based on instantaneous error feedback, such as errors in drug names or mismatches in disease duration, and adjusts model training parameters such as the optimizer step size and regularization coefficient in real time. The slow loop constructs a trend model based on the results of the entire training cycle and regularly and slowly adjusts the network structure configuration (such as the embedding dimension and the number of attention heads). During this process, the system introduces a stability guarantee function to set constraints on the parameter perturbation range and the frequency of structural changes to ensure that the model will not become unstable due to frequent changes.

[0189] After training, the generative model is able to output personalized, accurate, and logically coherent medical record summaries, diagnostic recommendations, and health guidance documents based on the patient's structured input information (symptoms, test results, and medical history). For example, if the system inputs the information of a 65-year-old patient hospitalized for hypertension and abnormal renal function, it can generate personalized recommendations for blood pressure management and kidney disease complication control. The system also generates a multidimensional quality score report through a verification network, including the standardization of medical terminology, coverage of medical history, and the completeness of the explanation of the medication plan, providing decision support for the doctor's final approval.

[0190] This embodiment achieves simultaneous management of real-time error changes and structural hyperparameters during model training by constructing a fast parameter adjustment loop and a slow architecture optimization loop. A training stability constraint function is introduced to ensure the controllability of the parameter update process, ultimately dynamically driving parameter adjustment and structural optimization through fusion control signals. This mechanism forms a dual-loop closed-loop control system based on verification feedback and strategy evolution, effectively improving the model's convergence speed, stability, and generalization ability during training, and significantly reducing the risk of performance degradation caused by fixed hyperparameter strategies or error oscillations.

[0191] In one embodiment, a generation optimization device based on dynamic verification feedback is provided, and the generation optimization device based on dynamic verification feedback corresponds one-to-one with the generation optimization method based on dynamic verification feedback in the above embodiment. Figure 3 , Figure 3 This is a functional module diagram of a preferred embodiment of a generation optimization device based on dynamic verification feedback. It includes a verification network construction module 10, a reward space generation module 20, a verification loop injection module 30, an exploration strategy optimization module 40, a feedback control execution module 50, and an inference result generation module 60. Each functional module is described in detail below:

[0192] A verification network building module 10 is used to pre-train a verification network for analyzing the quality of generated content;

[0193] A reward space generation module 20 is configured to construct a dynamic reward space comprising a plurality of orthogonal reward components based on the verification network;

[0194] A verification loop injection module 30 is used to integrate the dynamic reward space into the generative model to be trained to form a real-time verification loop, wherein the real-time verification loop injects verification signals into multiple processing layers of the generative model;

[0195] An exploration strategy optimization module 40 is configured to optimize the exploration strategy of the generative model according to the verification signal to obtain an optimized exploration strategy;

[0196] A feedback control execution module 50 is configured to perform dual-loop feedback control on the training process of the generative model based on the real-time verification loop and the optimized exploration strategy to dynamically adjust training parameters and generate a target model;

[0197] The inference result generating module 60 is configured to output an inference result based on the target model.

[0198] In one embodiment, the verification network construction module 10 is specifically configured to:

[0199] Generate a dataset of raw text samples;

[0200] Injecting semantic trap defects, logical conflict defects, and factual error defects into the original text sample dataset;

[0201] Annotate the error type labels of the text sample dataset after the injection defect;

[0202] Construct a multi-dimensional validation indicator matrix based on the text sample dataset annotated with error type labels;

[0203] Executing confidence interval error control of the multidimensional validation indicator matrix to generate an error control parameter table;

[0204] The discriminant training of the pre-training verification network is completed based on the error control parameter table.

[0205] In one embodiment, the reward space generation module 20 is specifically configured to:

[0206] Get the basic data of reward components generated by the verification network;

[0207] Based on the reward component basic data, setting a factual accuracy reward component and a logical consistency reward component;

[0208] performing an orthogonalization process on the factual accuracy reward component and the logical consistency reward component to generate an orthogonalized reward component;

[0209] Establishing a dynamic weight distribution model based on the orthogonalized reward components;

[0210] Based on the orthogonalized reward components and the dynamic weight distribution model, a reward projection matrix is ​​constructed to form a dynamic reward space.

[0211] In one embodiment, the verification loop injection module 30 is specifically configured to:

[0212] Select a specific middle layer of the generative model to be trained as the signal injection position to generate a specific middle layer position;

[0213] Mapping the dynamic reward space into a verification signal through a gating mechanism;

[0214] Injecting the verification signal at the specific intermediate layer position to generate an intermediate layer into which the verification signal is injected;

[0215] Based on the intermediate layer for injecting verification signals, a dual-path gradient propagation channel is designed;

[0216] Based on the dual-path gradient propagation channel, performing hardware-accelerated signal fusion processing to generate an accelerated fusion processing result;

[0217] Based on the accelerated fusion processing result, the construction of the real-time verification loop is completed.

[0218] In one embodiment, the verification loop injection module 30 is specifically configured to:

[0219] Based on the verification signal and the specific intermediate layer position, performing causal verification of the micro-generation unit to generate a micro-level verification result;

[0220] Based on the verification signal and the specific middle layer position, performing coherence verification of the meso-level semantic unit and generating a meso-level window verification result;

[0221] Based on the verification signal and the specific middle layer position, perform value alignment verification at the macro-session level to generate a macro-level verification result;

[0222] Fusion of the micro-level verification results, the meso-level window verification results, and the macro-level verification results to generate a layered fusion verification signal;

[0223] The layered fusion verification signal is injected into the specific middle layer position to generate the middle layer into which the verification signal is injected.

[0224] In one embodiment, the exploration strategy optimization module 40 is specifically configured to:

[0225] Mapping the original high-dimensional action space of the generative model to a low-dimensional manifold space to generate a low-dimensional manifold action space;

[0226] Establishing a policy entropy constraint boundary based on the historical entropy value in the verification signal;

[0227] Solving the action space boundary conditions that satisfy the policy entropy constraint boundary in the low-dimensional manifold action space to generate a constrained action space;

[0228] constructing a probabilistic graphical model in the constrained action space;

[0229] Based on the probabilistic graphical model, a pruning operation is performed on the exploration path generated by the generative model during the training process to generate an optimized exploration strategy.

[0230] In one embodiment, the feedback control execution module 50 is specifically configured to:

[0231] Establishing a fast parameter adjustment loop based on the real-time verification loop to respond to the real-time training error and generate a fast control signal;

[0232] Based on the optimized exploration strategy, a slow architecture optimization loop is constructed to adjust network hyperparameters and generate a slow control signal;

[0233] Set the training stability guarantee function to constrain the parameter update process and generate stability constraints;

[0234] Based on the fast control signal, the slow control signal and the stability constraint, a synergistic fusion of the dual-loop control signals is achieved to generate a fused control signal;

[0235] Dynamically adjusting the training parameters of the generative model based on the fusion control signal to generate updated training parameters;

[0236] An optimized target model is generated based on the updated training parameters.

[0237] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as follows: Figure 4 As shown. The computer device includes a processor, a memory, a network interface and a database connected via a system bus. The processor of the computer device is used to provide determination and control capabilities. The memory of the computer device includes a non-volatile and / or volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external user terminal via a network connection. When the computer program is executed by the processor, it realizes the functions or steps on the server side of a generation optimization method based on dynamic verification feedback.

[0238] In one embodiment, a computer device is provided. The computer device may be a user terminal, and its internal structure diagram may be as follows: Figure 5 As shown. The computer device includes a processor, memory, network interface, display screen and input device connected via a system bus. The processor of the computer device is used to provide determination and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external server via a network connection. When the computer program is executed by the processor, it realizes the functions or steps on the user side of a generation optimization method based on dynamic verification feedback.

[0239] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the following steps are performed:

[0240] Pre-training a verification network for analyzing the quality of generated content;

[0241] Based on the verification network, construct a dynamic reward space including multiple orthogonal reward components;

[0242] Integrating the dynamic reward space into a generative model to be trained to form a real-time verification loop that injects verification signals into multiple processing layers of the generative model;

[0243] Optimizing the exploration strategy of the generative model according to the verification signal to obtain an optimized exploration strategy;

[0244] Based on the real-time verification loop and the optimized exploration strategy, performing dual-loop feedback control on the training process of the generative model to dynamically adjust training parameters and generate a target model;

[0245] Outputting an inference result based on the target model.

[0246] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:

[0247] Pre-training a verification network for analyzing the quality of generated content;

[0248] Based on the verification network, construct a dynamic reward space including multiple orthogonal reward components;

[0249] Integrating the dynamic reward space into a generative model to be trained to form a real-time verification loop that injects verification signals into multiple processing layers of the generative model;

[0250] Optimizing the exploration strategy of the generative model according to the verification signal to obtain an optimized exploration strategy;

[0251] Based on the real-time verification loop and the optimized exploration strategy, performing dual-loop feedback control on the training process of the generative model to dynamically adjust training parameters and generate a target model;

[0252] Outputting an inference result based on the target model.

[0253] It should be noted that the above functions or steps that can be implemented by the computer-readable storage medium or computer device can be found in the relevant descriptions of the server side and the user side in the aforementioned method embodiment. To avoid repetition, they will not be described one by one here.

[0254] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).

[0255] Those skilled in the art will clearly understand that for the sake of convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0256] It should be noted that if any software tools or components other than those of the Company appear in the embodiments of this application, they are merely for illustration and do not represent actual use. The above embodiments are intended only to illustrate the technical solutions of the present invention, not to limit them. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or replace some of the technical features therein with equivalents. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included in the scope of protection of the present invention.

Claims

1. A generation optimization method based on dynamic verification feedback, characterized in that: The following steps are involved: Pre-training a verification network for analyzing the quality of generated content; Based on the verification network, construct a dynamic reward space including multiple orthogonal reward components; Integrating the dynamic reward space into a generative model to be trained to form a real-time verification loop that injects verification signals into multiple processing layers of the generative model; Optimizing the exploration strategy of the generative model according to the verification signal to obtain an optimized exploration strategy; Based on the real-time verification loop and the optimized exploration strategy, performing dual-loop feedback control on the training process of the generative model to dynamically adjust training parameters and generate a target model; Outputting an inference result based on the target model.

2. The generation optimization method based on dynamic verification feedback according to claim 1, characterized in that: Pre-trained verification network for analyzing the quality of generated content, including: Generate a dataset of raw text samples; Injecting semantic trap defects, logical conflict defects, and factual error defects into the original text sample dataset; Annotate the error type labels of the text sample dataset after the injection defect; Construct a multi-dimensional validation indicator matrix based on the text sample dataset annotated with error type labels; Executing confidence interval error control of the multidimensional validation indicator matrix to generate an error control parameter table; The discriminant training of the pre-training verification network is completed based on the error control parameter table.

3. The generation optimization method based on dynamic verification feedback according to claim 1, characterized in that: Based on the verification network, a dynamic reward space containing multiple orthogonal reward components is constructed, including: Get the basic data of reward components generated by the verification network; Based on the reward component basic data, setting a factual accuracy reward component and a logical consistency reward component; performing an orthogonalization process on the factual accuracy reward component and the logical consistency reward component to generate an orthogonalized reward component; Establishing a dynamic weight distribution model based on the orthogonalized reward components; Based on the orthogonalized reward components and the dynamic weight distribution model, a reward projection matrix is ​​constructed to form a dynamic reward space.

4. The generation optimization method based on dynamic verification feedback according to claim 1, characterized in that: Integrating the dynamic reward space into the generative model to be trained to form a real-time verification loop, wherein the real-time verification loop injects verification signals into multiple processing layers of the generative model, including: Select a specific middle layer of the generative model to be trained as the signal injection position to generate a specific middle layer position; Mapping the dynamic reward space into a verification signal through a gating mechanism; Injecting the verification signal at the specific intermediate layer position to generate an intermediate layer into which the verification signal is injected; Based on the intermediate layer for injecting verification signals, a dual-path gradient propagation channel is designed; Based on the dual-path gradient propagation channel, performing hardware-accelerated signal fusion processing to generate an accelerated fusion processing result; Based on the accelerated fusion processing result, the construction of the real-time verification loop is completed.

5. The generation optimization method based on dynamic verification feedback according to claim 4, characterized in that: Injecting the verification signal at the specific intermediate layer position to generate an intermediate layer for injecting the verification signal includes: Based on the verification signal and the specific intermediate layer position, performing causal verification of the micro-generation unit to generate a micro-level verification result; Based on the verification signal and the specific middle layer position, performing coherence verification of the meso-level semantic unit and generating a meso-level window verification result; Based on the verification signal and the specific middle layer position, perform value alignment verification at the macro-session level to generate a macro-level verification result; Fusion of the micro-level verification results, the meso-level window verification results, and the macro-level verification results to generate a layered fusion verification signal; The layered fusion verification signal is injected into the specific middle layer position to generate the middle layer into which the verification signal is injected.

6. The generation optimization method based on dynamic verification feedback according to claim 1, characterized in that: Optimizing the exploration strategy of the generative model according to the verification signal to obtain an optimized exploration strategy, including: Mapping the original high-dimensional action space of the generative model to a low-dimensional manifold space to generate a low-dimensional manifold action space; Establishing a policy entropy constraint boundary based on the historical entropy value in the verification signal; Solving the action space boundary conditions that satisfy the policy entropy constraint boundary in the low-dimensional manifold action space to generate a constrained action space; constructing a probabilistic graphical model in the constrained action space; Based on the probabilistic graphical model, a pruning operation is performed on the exploration path generated by the generative model during the training process to generate an optimized exploration strategy.

7. The generation optimization method based on dynamic verification feedback according to claim 1, characterized in that: Based on the real-time verification loop and the optimized exploration strategy, dual-loop feedback control is performed on the training process of the generative model to dynamically adjust training parameters to generate a target model, including: Establishing a fast parameter adjustment loop based on the real-time verification loop to respond to the real-time training error and generate a fast control signal; Based on the optimized exploration strategy, a slow architecture optimization loop is constructed to adjust network hyperparameters and generate a slow control signal; Set the training stability guarantee function to constrain the parameter update process and generate stability constraints; Based on the fast control signal, the slow control signal and the stability constraint, a synergistic fusion of the dual-loop control signals is achieved to generate a fused control signal; Dynamically adjusting the training parameters of the generative model based on the fusion control signal to generate updated training parameters; An optimized target model is generated based on the updated training parameters.

8. A generation optimization device based on dynamic verification feedback, characterized in that: The generation optimization device based on dynamic verification feedback includes: Verification network building module, used to pre-train the verification network used to analyze the quality of generated content; A reward space generation module, configured to construct a dynamic reward space comprising a plurality of orthogonal reward components based on the verification network; A verification loop injection module, configured to integrate the dynamic reward space into the generative model to be trained to form a real-time verification loop, wherein the real-time verification loop injects verification signals into multiple processing layers of the generative model; an exploration strategy optimization module, configured to optimize the exploration strategy of the generative model according to the verification signal to obtain an optimized exploration strategy; A feedback control execution module is configured to perform dual-loop feedback control on the training process of the generative model based on the real-time verification loop and the optimized exploration strategy to dynamically adjust training parameters and generate a target model; The inference result generation module is used to output the inference result based on the target model.

9. A computer device, characterized in that: The computer device includes a memory, a processor, and a generation optimization program based on dynamic verification feedback stored in the memory and capable of running on the processor. When the generation optimization program based on dynamic verification feedback is executed by the processor, the steps of the generation optimization method based on dynamic verification feedback as described in any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium, characterized in that The storage medium stores a generation optimization program based on dynamic verification feedback, which, when executed by a processor, implements the steps of the generation optimization method based on dynamic verification feedback as described in any one of claims 1 to 7.

Citation Information

Cited By

  • Intelligent agent model adaptive optimization method and system based on error feedback information

    CN121434371A

  • Laboratory equipment group power supply and heat dissipation resource allocation method based on reinforcement learning

    CN121436571A

  • High-temperature rotating part dynamic reliability prediction method and device based on digital-analog driving

    CN121936076A

  • Combined diagnosis method for self-regulation learning and emotion incentives

    CN122196643A