A Method and System for Illusion Correction Based on a Large Language Model of Self-Reflection Reinforcement Learning

By employing a self-reflective reinforcement learning method to correct illusions in a large language model, and by optimizing the model using a self-reflective dataset and gradient regularization strategy, the problems of data fidelity deviation and logical consistency breakage in exchange rate forecasting by large language models are solved, achieving high-precision and highly interpretable exchange rate forecasting.

CN122492356APending Publication Date: 2026-07-31GUANGDONG UNIV OF FINANCE
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
GUANGDONG UNIV OF FINANCE
Filing Date
2026-04-30
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

In exchange rate forecasting based on large language models, the models are prone to deviations in data fidelity, breaks in logical consistency, deviations from facts in unstructured text parsing, incomplete attribution logic, and inversion of primary and secondary attention allocation, which leads to a decline in forecast accuracy and the scientific nature of decision-making.

Method used

We employ a self-reflective reinforcement learning approach, using distillation to process unstructured news data into key fact summaries. We then construct an ST fusion model for nonlinear weighted integration, comparing and contrasting prediction signals with market trends in real time. This generates a self-reflective dataset, and we optimize the model using gradient regularization. We also introduce format normalization rewards and composite reward functions for fine-tuning.

Benefits of technology

It improves the accuracy and scientific nature of exchange rate forecasting, achieves a balance between high precision and high interpretability, effectively suppresses the illusion phenomenon in attribution analysis, and enhances the forecasting stability in complex market environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122492356A_ABST
    Figure CN122492356A_ABST
Patent Text Reader

Abstract

This invention relates to the field of large language model technology, providing a method for correcting illusions in large language models based on self-reflective reinforcement learning. It captures long- and short-term dependencies and multi-period features of structured data by constructing a fusion model. Simultaneously, it utilizes a large model to distill unstructured news into core evidence, employing a self-reflective mechanism that requires no manual annotation. By comparing predicted signals with actual market trends, it automatically generates high-quality training data containing explanations and reflections. A gradient regularization strategy is used to fine-tune the large language model, and a composite reward function is combined to achieve a balance between high accuracy and high interpretability. This invention also provides a system for executing the above method. Through a dual-path parallel processing mechanism, it achieves collaborative mining of temporal multi-period features and semantic incremental information, improving prediction stability in complex market environments. Furthermore, through reinforcement learning fine-tuning with enforced logical constraints, it suppresses illusion phenomena in attribution analysis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of large language model technology, and in particular to a method and system for correcting illusions in large language models based on self-reflective reinforcement learning. Background Technology

[0002] With the rapid development of the digital economy, financial data is exhibiting significant characteristics of multi-source and heterogeneity. Traditional forecasting models that rely solely on historical transaction data are no longer adequate for the complex and volatile exchange rate market environment. Against this backdrop, multi-source information fusion forecasting methods that integrate structured transaction data with unstructured financial news data are gradually becoming the mainstream trend in the field of exchange rate forecasting. Unstructured texts such as financial news and institutional research reports contain rich information on macroeconomic policies, geopolitical dynamics, and market sentiment. This information can effectively compensate for the limitations of traditional transaction data, which can only reflect historical fluctuation patterns.

[0003] CN121563670A discloses a method for calculating multi-currency quotations, including the following steps: obtaining the original dataset related to the quotation, which contains the transaction currency; obtaining the spot exchange rate and historical exchange rate data of the transaction currency against the local currency, and forming a market dataset; performing cost analysis based on the original dataset and the market dataset, calculating the cost currency valuation of each cost component, and summing them to obtain the total direct cost amount; obtaining survey data, and obtaining the risk level by analyzing the survey data, and providing corresponding hedging solutions based on the risk level; combining user cash flow, historical exchange rate data, and hedging solutions to calculate the cash flow valuation under the influence of exchange rates, obtaining a sample set of contract quotation amounts by pre-setting gross profit margin, confidence level, and multiple sets of transaction currency exchange rates, and calculating the final contract quotation amount based on the lowest contract quotation amount in the sample set of contract quotation amounts and the pre-set premium parameter.

[0004] When using large language models for attribution analysis of exchange rate forecasts, three types of illusions are prone to occur: deviation from data fidelity, break in logical consistency, and weakening of information relevance. These phenomena lead to deviations from the facts in the analysis of unstructured text, incomplete attribution logic, and inversion of primary and secondary attention allocation, further affecting the accuracy of exchange rate forecasts and the scientific nature of decision-making. Summary of the Invention

[0005] Long-term practice has shown that in attribution analysis of exchange rate forecasting based on large language models, model illusion leads to serious deviations in data fidelity and breaks in logical consistency. This results in technical problems such as unstructured text parsing that deviates from the facts, incomplete attribution logic, and inverted attention allocation, which directly weaken the accuracy of forecast results and the scientific nature of decision-making.

[0006] In view of this, the present invention provides a method for correcting illusions in large language models based on self-reflective reinforcement learning, the method comprising: Step S1: Distill the unstructured news data to form a standardized summary of key facts; Step S2: Construct an ST fusion model based on historical transaction data and macroeconomic indicators. The ST fusion model performs nonlinear weighted integration of the extracted feature outputs through a gradient boosting meta-model; and outputs a binary rise and fall prediction signal for the next trading day. Step S3: The binary rise and fall prediction signal is compared with the actual market trend data in real time. If the binary rise and fall prediction signal is consistent with the actual market trend data, the ST fusion model extracts the key macroeconomic indicators or policy events that support the trend based on the key fact summary, and constructs the causal attribution chain data from the key macroeconomic indicators or policy events to the correct prediction conclusion. If the binary rise and fall prediction signal is inconsistent with the actual market trend data, the ST fusion model performs failure analysis on the decision-making process, identifies the dominant factors that are ignored or the contradictory signals that are misjudged in the summary of key facts, and generates reflective text data containing erroneous attributions and logical corrections. Construct a self-reflective dataset using the causal attribution chain and the reflective text data; Step S4: Using the self-reflective dataset, the model is fine-tuned using a gradient regularization optimization algorithm to output the trained ST fusion model.

[0007] Preferably, step S4 further includes, Step S41: Force the model to be deployed within the first tag by using preset XML tags to complete causal inference based on news evidence; Step S42: Output the predicted signal within the second tag of the XML; Step S43: The completeness of the generated content's tags is scored using a format compliance reward.

[0008] Preferably, step S4 further includes, Step S44: Apply a reward function to each item in the first and second tags to provide feedback.

[0009] in, , , Assign weights to each reward signal; Let be the total reward value for the i-th generated content. The i-th element in a set of content generated by the model based on the current strategy; RXML is a format conformity reward, measuring whether the model follows a predefined structured label; Rcorrectness is a reward for prediction accuracy, which evaluates the alignment between the generated content and the facts. Rcount is a logical completeness reward, used to evaluate the logicality of an interpreted statement.

[0010] Preferably, step S4 further includes, Step S45, the comprehensive reward value for the i-th generated content. Calculate the first group using within-group statistical characteristics. Advantage function for generating content The combined reward value of a set of N generated content by the current strategy. The average value is used as the baseline, and the standard deviation is normalized.

[0011] in, Let x be the advantage function value of the i-th generated content; x is the input sample. The i-th element in a set of content generated by the model based on the current strategy; The average reward value of N generated content items in the group; Let be the standard deviation of the reward values ​​of the N generated content in the group.

[0012] Preferably, in step S4, Fine-tuning the model involves adjusting its parameters through gradient updates.

[0013] Where b is the expected value of the baseline, In strategy Given x as the input sample, generate content. The logarithmic probability; Gradient operator, calculating the gradient of logarithmic probability.

[0014] Preferably, the model's current policy is optimized by maximizing the objective function. ,

[0015] in, This is the old strategy; This refers to the strategy after initial pre-training. To truncate the threshold, limit the update step size; The weighting coefficients for the KL divergence penalty term; This is the KL divergence penalty term; clip is the truncation operation, which forcibly restricts the value to [ Within the specified interval. Preferably, in the ST fusion model, a gradient boosting meta-model is used to non-linearly weight and integrate the feature outputs extracted from Space-Time and Times-Net; Fusion features for,

[0016] in, For the m-th regression tree, The corresponding weights; Output features for Space-Time; The output features of Times-Net are M, where M is the total number of fusions and M is a positive integer greater than 2.

[0017] This invention also discloses a system for hallucination correction based on a large language model using self-reflective reinforcement learning, as described above. The system includes... Standardized units are used to distill unstructured news data to form standardized summaries of key facts. The construction unit is used to build an ST fusion model based on historical transaction data and macroeconomic indicators. The ST fusion model performs nonlinear weighted integration of the extracted feature outputs through a gradient boosting meta-model and outputs a binary rise and fall prediction signal for the next trading day. The comparison unit is used to compare the binary rise and fall prediction signal with the actual market trend data in real time. If the binary rise and fall prediction signal is consistent with the actual market trend data, the ST fusion model extracts the key macroeconomic indicators or policy events that support the trend based on the key fact summary, and constructs the causal attribution chain data from the key macroeconomic indicators or policy events to the correct prediction conclusion. If the binary rise and fall prediction signal is inconsistent with the actual market trend data, the ST fusion model performs failure analysis on the decision-making process, identifies the dominant factors that are ignored or the contradictory signals that are misjudged in the summary of key facts, and generates reflective text data containing erroneous attributions and logical corrections. Construct a self-reflective dataset using the causal attribution chain and the reflective text data; The output unit is used to fine-tune the model using the gradient regularization strategy optimization algorithm on the self-reflective dataset, thereby outputting the trained ST fusion model.

[0018] The present invention also discloses an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that the processor, when executing the computer program, implements the above-described method for correcting illusions based on a large language model of self-reflective reinforcement learning.

[0019] The present invention also discloses a machine-readable storage medium storing instructions for causing a machine to execute the large language model illusion correction method based on self-reflective reinforcement learning as described in any of the preceding claims.

[0020] This invention provides a method for correcting illusions in large language models based on self-reflective reinforcement learning. First, in step S1, unstructured news data is distilled to obtain a standardized key fact summary. Step S2 constructs an ST fusion model based on historical trading data and macroeconomic indicators, and uses a gradient boosting meta-model to perform nonlinear weighted integration of feature outputs, outputting a binary rise / fall prediction signal for the next trading day. Step S3 compares the predicted signal with the actual market trend in real time. When the predictions are consistent, macroeconomic indicators or policy events supporting the trend are extracted from the key fact summary to establish a causal attribution chain pointing to the prediction conclusion. When the predictions are inconsistent, decision failure analysis is performed to identify overlooked dominant factors or misjudged contradictory signals, generating reflective text containing erroneous attributions and logical corrections. The causal attribution chain and reflective text are then integrated to construct a self-reflective dataset. Step S4 uses this self-reflective dataset to fine-tune the model using a gradient regularization optimization algorithm, finally obtaining the trained ST fusion model. This invention also discloses a system for implementing the aforementioned illusion correction method for large language models based on self-reflective reinforcement learning. The core innovation of this method and system lies in upgrading from traditional scalar-level subtraction to operator-level projection. By constructing an ST fusion model, it combines Space-Time and Times-Net to capture the long- and short-term dependencies and multi-period features of structured data. Simultaneously, it utilizes a large model to distill unstructured news into core evidence, providing high-quality initial signals and factual basis for attribution analysis. A self-reflective mechanism without manual annotation is employed to automatically generate high-quality training data containing explanations and reflections by comparing predicted signals with actual market trends, effectively addressing the problem of scarce labeled data in the financial field. A gradient regularization strategy is used to fine-tune the large language model, combined with a composite reward function, guiding the model to generate logically rigorous and evidence-based predictive explanations, achieving a balance between high accuracy and high interpretability. A dual-path parallel processing mechanism enables the collaborative mining of multi-period features of exchange rate time series and incremental information of news semantics, effectively improving prediction stability in complex market environments. Furthermore, through reinforcement learning fine-tuning with mandatory logical constraints, the illusion phenomenon in attribution analysis is suppressed in principle. Attached Figure Description

[0021] The accompanying drawings, which form part of this invention, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute an undue limitation of the invention. In the drawings: Figure 1 This is a schematic diagram of a large language model illusion correction method based on self-reflective reinforcement learning, according to one embodiment of the present invention. Detailed Implementation

[0022] The specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are for illustration and explanation only and are not intended to limit the present invention.

[0023] In exchange rate forecasting attribution analysis scenarios based on large language models, the combined effects of insufficient training data and model illusion can easily lead to problems such as deviations in data fidelity, breaks in logical consistency, unstructured text parsing deviating from actual facts, incomplete attribution logic, and inverted attention allocation. These issues severely reduce the accuracy of exchange rate forecasts and weaken the scientific nature of market decisions. This invention proposes a method for correcting large language model illusions based on self-reflective reinforcement learning, such as... Figure 1 As shown, the large language model illusion correction method based on self-reflective reinforcement learning includes: Step S1: Distill the unstructured news data to form a standardized summary of key facts; Step S2: Construct an ST fusion model based on historical transaction data and macroeconomic indicators. The ST fusion model performs nonlinear weighted integration of the extracted feature outputs through a gradient boosting meta-model; and outputs a binary rise and fall prediction signal for the next trading day. Step S3: The binary rise and fall prediction signal is compared with the actual market trend data in real time. If the binary rise and fall prediction signal is consistent with the actual market trend data, the ST fusion model extracts the key macroeconomic indicators or policy events that support the trend based on the key fact summary, and constructs the causal attribution chain data from the key macroeconomic indicators or policy events to the correct prediction conclusion. If the binary rise and fall prediction signal is inconsistent with the actual market trend data, the ST fusion model performs failure analysis on the decision-making process, identifies the dominant factors that are ignored or the contradictory signals that are misjudged in the summary of key facts, and generates reflective text data containing erroneous attributions and logical corrections. Construct a self-reflective dataset using the causal attribution chain and the reflective text data; Step S4: Using the self-reflective dataset, the model is fine-tuned using a gradient regularization optimization algorithm to output the trained ST fusion model.

[0024] This invention provides a method for correcting illusions in large language models based on self-reflective reinforcement learning. First, in step S1, unstructured news data is distilled to obtain a standardized key fact summary. Step S2 constructs an ST fusion model based on historical trading data and macroeconomic indicators, and uses a gradient boosting meta-model to perform nonlinear weighted integration of feature outputs, outputting a binary rise / fall prediction signal for the next trading day. Step S3 compares the predicted signal with the actual market trend in real time. When the predictions are consistent, macroeconomic indicators or policy events supporting the trend are extracted from the key fact summary to establish a causal attribution chain pointing to the prediction conclusion. When the predictions are inconsistent, decision failure analysis is performed to identify overlooked dominant factors or misjudged contradictory signals, generating reflective text containing erroneous attributions and logical corrections. The causal attribution chain and reflective text are then integrated to construct a self-reflective dataset. Step S4 uses this self-reflective dataset to fine-tune the model using a gradient regularization optimization algorithm, finally obtaining the trained ST fusion model. The core innovation of this method lies in upgrading from traditional scalar-level subtraction to operator-level projection. By constructing an ST fusion model, it combines Space-Time and Times-Net to capture the long-term and short-term dependencies and multi-period features of structured data. Simultaneously, it utilizes a large model to distill unstructured news into core evidence, providing high-quality initial signals and factual basis for attribution analysis. A self-reflective mechanism without manual annotation is employed to automatically generate high-quality training data containing explanations and reflections by comparing predicted signals with actual market trends, effectively addressing the problem of scarce labeled data in the financial field. A gradient regularization strategy is used to fine-tune the large language model, combined with a composite reward function, guiding the model to generate logically rigorous and evidence-based predictive explanations, achieving a balance between high accuracy and high interpretability. A dual-path parallel processing mechanism enables the collaborative mining of multi-period features of exchange rate time series and incremental information from news semantics, effectively improving prediction stability in complex market environments. Furthermore, reinforcement learning fine-tuning with mandatory logical constraints fundamentally suppresses the illusion phenomenon in attribution analysis.

[0025] To guide the model in generating explanatory text that combines high prediction accuracy with deep causal logic, in a more preferred embodiment of the present invention, step S4 further includes: Step S41: Force the model to be deployed within the first tag by using preset XML tags to complete causal inference based on news evidence; Step S42: Output the predicted signal within the second tag of the XML; Step S43: The completeness of the generated content's tags is scored using format conformity rewards. For example, the system employs a format constraint mechanism, a core technical constraint method that forces the model to generate a specific logical structure through preset XML tags. This mechanism forces the model to first... <reasoning>Only after deploying and completing causal inference based on news evidence within the tag can it be implemented. <answer>The predicted signal is output within the label. Simultaneously, the completeness of the generated content's labels is scored using RXML, a format-standardization reward, thus achieving a technical closed loop from mandatory structure output to standardization reward feedback. This prevents the model from skipping inference and directly guessing results at the underlying logic level. Furthermore... <answer>The predicted signal is output within the label, thus avoiding the model skipping inference and directly guessing the result.

[0026] To go beyond simply evaluating the correctness of the results, this invention provides a comprehensive quantitative score for the output based on three dimensions: format completeness, prediction accuracy, and logical consistency. In a more preferred embodiment, step S4 further includes... Step S44: Apply a reward function to each item in the first and second tags to provide feedback.

[0027] in, , , Assign weights to each reward signal; Let be the total reward value for the i-th generated content. The i-th element in a set of content generated by the model based on the current strategy; RXML is a format conformity reward, measuring whether the model follows a predefined structured label; Rcorrectness is a reward for prediction accuracy, which evaluates the alignment between the generated content and the facts. Rcount is a logical completeness reward, used to evaluate the logical integrity of an explained statement. For example, the weighting ratio is set as follows: , , .

[0028] The Group Relative Policy Optimization (GRPO) strategy innovatively abandons the independent value network (Critic) used to estimate state values ​​in the traditional Proximal Policy Optimization (PPO) algorithm, thereby significantly reducing memory usage and computational overhead during training. Instead, this strategy employs a group sampling mechanism to dynamically estimate the advantage value. In a more preferred embodiment of the invention, step S4 further includes... Step S45, the comprehensive reward value for the i-th generated content. Calculate the first group using within-group statistical characteristics. Advantage function for generating content The combined reward value of a set of N generated content by the current strategy. The average value is used as the baseline, and the standard deviation is normalized.

[0029] in, Let x be the advantage function value of the i-th generated content; x is the input sample. The i-th element in a set of content generated by the model based on the current strategy; The average reward value of N generated content items in the group; Let be the standard deviation of the reward values ​​of the N generated content in the group.

[0030] For each input sample x, which is the combination of the predicted signal and the news summary, the model follows the current strategy. Sample to generate a set of One candidate answer The system calculates the overall reward value for each piece of content generated within the group. And use within-group statistical characteristics to calculate the first Advantage function for generating content The calculation process dynamically sets the baseline to the average reward of the current batch and normalizes it using the standard deviation. This intra-group relative ranking mechanism not only effectively eliminates the variance in the reward signal but also allows the model to keenly identify high-quality attribution logic that outperforms the average under the current policy, thus providing a more robust learning signal.

[0031] During the parameter update phase, the system adjusts the model parameters through gradient updates, specifically GRPO reinforcement learning gradient updates. To prevent the model from overfitting the reward function during reinforcement learning, which could lead to a decline in language expression ability or incorrect reward responses, in a more preferred embodiment of the invention, in step S4... Fine-tuning the model involves adjusting its parameters through gradient updates.

[0032] Where b is the expected value of the baseline, In strategy Given x as the input sample, generate content. The logarithmic probability; Gradient operator, calculating the gradient of logarithmic probability.

[0033] The model's policy πθ is optimized by maximizing the objective function, making its generated explanations more likely to receive higher rewards. To prevent the model from overfitting the reward function during reinforcement learning, which could lead to a decline in language expression ability or incorrect reward responses, in a more preferred embodiment of this invention, the current policy of the model is optimized by maximizing the objective function. ,

[0034] in, This is the old strategy; This refers to the strategy after initial pre-training. To truncate the threshold, limit the update step size; The weighting coefficients for the KL divergence penalty term; This is the KL divergence penalty term; clip is the truncation operation, which forcibly restricts the value to [ Within the specified interval. A KL divergence (Kullback-Leibler Divergence) is introduced as a regularization term into the objective function. This term is used to constrain the current policy. Compared with the initial reference model The distribution differences between them ensure that the model does not lose general language capabilities while learning domain logic. Furthermore, the optimization process employs a clipping mechanism to limit the probability ratio of the new and old strategies within a certain range, for example... This prevents training from crashing due to excessively large single gradient update steps.

[0035] By leveraging the Space-Time architecture, this invention overcomes the memory bottleneck of traditional models for long sequences, accurately capturing the long-term dependencies in exchange rate data. Simultaneously, it combines Times-Net to transform one-dimensional time series into two-dimensional spatial tensors, effectively extracting multi-period fluctuation features. In a more preferred embodiment of this invention, the ST fusion model uses a gradient boosting meta-model to non-linearly weight and integrate the feature outputs extracted from Space-Time and Times-Net. Fusion features for,

[0036] in, The m-th regression tree is the weak learner. The corresponding weights. Output features for Space-Time; M represents the output features of Times-Net, and M is the total number of fusions, where M is a positive integer greater than 2. The ST fusion model is a neural network model that integrates the Space-Time model and the Times-Net model. It takes input samples into the Space-Time model and the Times-Net model respectively, and then fuses the output features of the two models.

[0037] After the nonlinear weighted fusion process of the aforementioned meta-model, the final output is a binary price prediction signal for the next trading day. The prediction and judgment process is as follows:

[0038] in, This is the closing price of the exchange rate between the first currency and the second currency for the next trading day.

[0039] This is the closing exchange rate of the first currency against the second currency for the current trading day. When it is determined to be an upward signal, This is considered a downtrend signal.

[0040] This invention also discloses a system for hallucination correction based on a large language model using self-reflective reinforcement learning, as described above. The system includes... Standardized units are used to distill unstructured news data to form standardized summaries of key facts. The construction unit is used to build an ST fusion model based on historical transaction data and macroeconomic indicators. The ST fusion model performs nonlinear weighted integration of the extracted feature outputs through a gradient boosting meta-model and outputs a binary rise and fall prediction signal for the next trading day. The comparison unit is used to compare the binary rise and fall prediction signal with the actual market trend data in real time. If the binary rise and fall prediction signal is consistent with the actual market trend data, the ST fusion model extracts the key macroeconomic indicators or policy events that support the trend based on the key fact summary, and constructs the causal attribution chain data from the key macroeconomic indicators or policy events to the correct prediction conclusion. If the binary rise and fall prediction signal is inconsistent with the actual market trend data, the ST fusion model performs failure analysis on the decision-making process, identifies the dominant factors that are ignored or the contradictory signals that are misjudged in the summary of key facts, and generates reflective text data containing erroneous attributions and logical corrections. Construct a self-reflective dataset using the causal attribution chain and the reflective text data; The output unit is used to fine-tune the model using the gradient regularization strategy optimization algorithm on the self-reflective dataset, thereby outputting the trained ST fusion model.

[0041] This system enables the model to generate explanatory and reflective data by comparing it with actual market trends. This effectively solves the problem of scarce fine-tuning datasets in the financial vertical field of reinforcement learning. While significantly reducing reliance on expensive human expert annotation, it also significantly reduces the time and manpower costs of model training, achieving low-cost iteration of domain knowledge. Through the aforementioned dual-path parallel processing mechanism, it achieves collaborative mining of multi-period features of exchange rate time series and incremental information of news semantics, effectively improving the predictive stability in complex market environments. At the same time, by introducing a composite reward function with format constraint mechanism and multi-dimensional logical evaluation to fine-tune the large language model, it significantly enhances the representational ability of generated text in terms of logical consistency and data fidelity from the underlying principle. The ST fusion model effectively captures the multi-period features of time series data, and combines it with the GRPO reinforcement learning strategy to fine-tune the large model, guiding it to follow the logical paradigm of reasoning first and then giving conclusions. This not only ensures high accuracy of exchange rate prediction, but also endows the model with the ability to generate logically rigorous and evidence-based attribution analysis capabilities, solving the pain point of unexplainable decision-making in traditional artificial intelligence models and improving the scientific nature of financial decision support.

[0042] The present invention also discloses an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that the processor, when executing the computer program, implements the above-described method for correcting illusions based on a large language model of self-reflective reinforcement learning.

[0043] The present invention also discloses a machine-readable storage medium storing instructions for causing a machine to execute the large language model illusion correction method based on self-reflective reinforcement learning described above.

[0044] In the above embodiments, the descriptions of each embodiment have different focuses. Parts not described in detail in a certain embodiment can be referred to in the relevant descriptions of other embodiments. The above descriptions are merely preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention can have various modifications and variations. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.< / answer> < / answer> < / reasoning>

Claims

1. A method for correcting hallucinations using a large language model based on self-reflective reinforcement learning, characterized in that, The large language model illusion correction method based on self-reflective reinforcement learning includes: Step S1: Distill the unstructured news data to form a standardized summary of key facts; Step S2: Construct an ST fusion model based on historical transaction data and macroeconomic indicators. The ST fusion model performs nonlinear weighted integration of the extracted feature outputs through a gradient boosting meta-model; and outputs a binary rise and fall prediction signal for the next trading day. Step S3: The binary rise and fall prediction signal is compared with the actual market trend data in real time. If the binary rise and fall prediction signal is consistent with the actual market trend data, the ST fusion model extracts the key macroeconomic indicators or policy events that support the trend based on the key fact summary, and constructs the causal attribution chain data from the key macroeconomic indicators or policy events to the correct prediction conclusion. If the binary rise and fall prediction signal is inconsistent with the actual market trend data, the ST fusion model performs failure analysis on the decision-making process, identifies the dominant factors that are ignored or the contradictory signals that are misjudged in the summary of key facts, and generates reflective text data containing erroneous attributions and logical corrections. Construct a self-reflective dataset using the causal attribution chain and the reflective text data; Step S4: Using the self-reflective dataset, the model is fine-tuned using a gradient regularization optimization algorithm to output the trained ST fusion model.

2. The hallucination correction method based on self-reflective reinforcement learning for large language models according to claim 1, characterized in that, Step S4 also includes, Step S41: Force the model to be deployed within the first tag by using preset XML tags to complete causal inference based on news evidence; Step S42: Output the predicted signal within the second tag of the XML; Step S43: The completeness of the generated content's tags is scored using a format compliance reward.

3. The hallucination correction method based on self-reflective reinforcement learning for large language models according to claim 2, characterized in that, Step S4 also includes, Step S44: Apply a reward function to each item in the first and second tags to provide feedback. in, , , Assign weights to each reward signal; Let be the total reward value for the i-th generated content. The i-th element in a set of content generated by the model based on the current strategy; RXML is a format conformity reward, measuring whether the model follows a predefined structured label; Rcorrectness is a reward for prediction accuracy, which evaluates the alignment between the generated content and the facts. Rcount is a logical completeness reward, used to evaluate the logicality of an interpreted statement.

4. The hallucination correction method based on self-reflective reinforcement learning for large language models according to claim 3, characterized in that, Step S4 also includes, Step S45, the comprehensive reward value for the i-th generated content. Calculate the first group using within-group statistical characteristics. Advantage function for generating content The combined reward value of a set of N generated content by the current strategy. The average value is used as the baseline, and the standard deviation is normalized. in, Let x be the advantage function value of the i-th generated content; x is the input sample. The i-th element in a set of content generated by the model based on the current strategy; The average reward value of N generated content items in the group; Let be the standard deviation of the reward values ​​of the N generated content in the group.

5. The hallucination correction method based on self-reflective reinforcement learning for large language models according to claim 4, characterized in that, In step S4, Fine-tuning the model involves adjusting its parameters through gradient updates. Where b is the expected value of the baseline, In strategy Given x as the input sample, generate content. The logarithmic probability; Gradient operator, calculating the gradient of logarithmic probability.

6. The method for correcting hallucinations based on large language models using self-reflective reinforcement learning according to claim 5, characterized in that, The current strategy of optimizing the model by maximizing the objective function. , in, This is the old strategy; This refers to the strategy after initial pre-training. To truncate the threshold, limit the update step size; The weighting coefficients for the KL divergence penalty term; This is the KL divergence penalty term; clip is the truncation operation, which forcibly restricts the value to [ Within the range.

7. The method for correcting hallucinations based on large language models using self-reflective reinforcement learning according to any one of claims 1-6, characterized in that, In the ST fusion model, the feature outputs extracted from Space-Time and Times-Net are non-linearly weighted and integrated through a gradient boosting meta-model. Fusion features for, in, For the m-th regression tree, The corresponding weights; Output features for Space-Time; The output features of Times-Net are M, where M is the total number of fusions and M is a positive integer greater than 2.

8. A system for a large language model illusion correction method based on self-reflective reinforcement learning as described in any one of claims 1-7, characterized in that, The system includes, Standardized units are used to distill unstructured news data to form standardized summaries of key facts. The construction unit is used to build an ST fusion model based on historical transaction data and macroeconomic indicators. The ST fusion model performs nonlinear weighted integration of the extracted feature outputs through a gradient boosting meta-model and outputs a binary rise and fall prediction signal for the next trading day. The comparison unit is used to compare the binary rise and fall prediction signal with the actual market trend data in real time. If the binary rise and fall prediction signal is consistent with the actual market trend data, the ST fusion model extracts the key macroeconomic indicators or policy events that support the trend based on the key fact summary, and constructs the causal attribution chain data from the key macroeconomic indicators or policy events to the correct prediction conclusion. If the binary rise and fall prediction signal is inconsistent with the actual market trend data, the ST fusion model performs failure analysis on the decision-making process, identifies the dominant factors that are ignored or the contradictory signals that are misjudged in the summary of key facts, and generates reflective text data containing erroneous attributions and logical corrections. Construct a self-reflective dataset using the causal attribution chain and the reflective text data; The output unit is used to fine-tune the model using the gradient regularization strategy optimization algorithm on the self-reflective dataset, thereby outputting the trained ST fusion model.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the large language model illusion correction method based on self-reflective reinforcement learning as described in any one of claims 1-7.

10. A machine-readable storage medium, characterized in that, The machine-readable storage medium stores instructions for causing the machine to execute the large language model illusion correction method based on self-reflective reinforcement learning as described in any one of claims 1-7.