Prompt word optimization method and device and electronic equipment
By constructing a pre-defined template library and an illusion detection model within a large language model, embedding constraints and terminology from the financial field, and combining model parameter and task complexity detection, the problem of generating illusions in the financial field using a large language model is solved. This achieves efficient illusion optimization and professional adaptation, reducing financial business risks.
Patent Information
- Application Number
- CN202511674370.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-14
- Publication Date
- 2026-03-13
AI Technical Summary
The concealment and misleading nature of large language models in generating illusory content in the financial field leads to increased credit risk and violations of regulatory requirements, which existing technologies have failed to effectively address.
By constructing a pre-defined template library to match the constraint rules of financial instructions, enhanced thought chain prompts are generated. These prompts are then detected and optimized using an illusion detection model. Domain-specific terminology and industry rules are embedded, and model parameters and task complexity are combined to eliminate unsuitable scenarios. Multi-path consistency verification and iterative optimization are then performed.
It reduces the probability of illusion in the financial field of large language models, improves professional adaptability, reduces unreliable output, reduces the risk of financial business and the risk of logical breakage, and improves efficiency.
Smart Images

Figure CN121660070A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, and more specifically to a prompt word optimization method, a prompt word optimization device, and an electronic device. Background Technology
[0002] With the widespread application of Large Language Models (LLMs) across various industries, the problem of "illusions" (i.e., outputting incorrect or unsubstantiated information) in their generated content has become increasingly prominent. From simple factual errors in the early days to complex logical contradictions today, the concealment and misleading nature of illusory content are constantly increasing, not only affecting the accuracy of information transmission but also posing potential risks to fields that rely on model-based decision-making.
[0003] In the financial industry, large-scale models are frequently used in critical operations such as loan approval, credit rating, and compliance verification. The potential harm of illusory content is particularly prominent—for example, miscalculating the monthly payment / income ratio, fabricating credit records, and misjudging the scope of regulatory provisions. This could lead to increased credit risk, violations of regulatory requirements, and even financial losses and reputational damage. Therefore, how to optimize large-scale language models for illusion in the financial field has become an urgent problem to be solved. Summary of the Invention
[0004] The purpose of this invention is to provide a prompt word optimization method, apparatus, and electronic device for achieving illusion optimization of large language models in the financial field.
[0005] To achieve the above objectives, embodiments of the present invention provide a prompt word optimization method, including: Obtain an input task, wherein the input task includes at least a financial instruction; The example and constraint rules corresponding to the financial instruction are determined based on a preset template library, and the reasoning path corresponding to the financial instruction is generated; the preset template library includes constraint rules associated with different financial keywords; The financial instructions, the examples, the constraint rules, and the reasoning path are used to construct enhanced thought chain prompts; The enhanced thought chain cue words are subjected to hallucination detection based on a hallucination detection model, and an optimization strategy for the enhanced thought chain cue words is determined based on the hallucination detection results.
[0006] Optionally, the input task further includes a target model; after obtaining the input task, it further includes: The task complexity of the financial instruction is detected using a task complexity detection model to obtain a complexity score; The target model is subjected to model parameter detection to obtain the model parameter values of the target model; If the complexity score is greater than or equal to a set score threshold and / or the model parameter value is greater than or equal to a set parameter threshold, the financial instruction is determined to meet the requirements.
[0007] Optionally, the inference path includes multiple paths, and the inference path for generating the financial instruction includes: Multiple initial inference paths corresponding to the financial instruction are randomly generated; The path perplexity of each initial inference path is calculated using a path perplexity model; From the multiple initial inference paths, at least two initial inference paths with a path perplexity less than or equal to a set perplexity threshold are selected as the multiple inference paths corresponding to the financial instruction.
[0008] Optionally, the process of performing hallucination detection on the enhanced thought chain cue words based on the hallucination detection model, and determining the optimization strategy for the enhanced thought chain cue words based on the hallucination detection results, includes: Based on the hallucination detection model, hallucination detection is performed on the enhanced thought chain cue words to obtain the proportion of hallucination sentences and the number of erroneous reasoning steps of the enhanced thought chain cue words; The hallucination score is calculated based on the proportion of hallucinatory sentences and the number of incorrect reasoning steps. If the hallucination score is less than or equal to a set score threshold, the hallucination detection result is determined to meet the requirements. If the hallucination detection results meet the requirements, the enhanced thought chain prompt is output.
[0009] Optionally, calculating the hallucination score based on the proportion of hallucinatory sentences and the number of erroneous reasoning steps includes: The hallucination score is calculated based on the weighted sum of the proportion of hallucinatory sentences and the number of erroneous reasoning steps.
[0010] Optionally, the method further includes: Determine the conclusions corresponding to all the aforementioned reasoning paths; If, among all the conclusions corresponding to the inference paths, there are fewer than a set threshold of identical conclusions, calculate the path perplexity of each inference path. The inference path with the minimum path perplexity among all the inference paths is identified as the target inference path. Delete all reasoning paths from the enhanced thought chain prompts that exclude the target reasoning path.
[0011] Optionally, the method further includes: Repeat the following steps until the set termination condition is met: If the hallucination score is greater than a set score threshold, determine the hallucination type of the current enhanced mind chain cue word; The hallucination cause corresponding to the hallucination type is determined based on a preset mapping relationship; the preset mapping relationship includes multiple hallucination types and the hallucination cause associated with each hallucination type; Based on the causes of the hallucinations, a cue word adjustment strategy was determined; The current enhanced mind chain prompts are adjusted based on the prompt adjustment strategy; Calculate the illusion score of the adjusted enhanced mind chain cue; wherein the adjusted enhanced mind chain cue is used as the current enhanced mind chain cue in the next iteration.
[0012] Optionally, the constraint rules may include domain terminology and / or industry rules.
[0013] On the other hand, embodiments of the present invention also provide a prompt word optimization device, comprising: An acquisition module is used to acquire input tasks, wherein the input tasks include at least financial instructions; The first determining module is used to determine the example and constraint rules corresponding to the financial instruction based on a preset template library, and to generate the reasoning path corresponding to the financial instruction; the preset template library includes constraint rules associated with different financial keywords; A construction module is used to construct enhanced thought chain prompts using the financial instructions, the examples, the constraint rules, and the reasoning path; The second determining module is used to perform hallucination detection on the enhanced thought chain cue words based on the hallucination detection model, and to determine the optimization strategy for the enhanced thought chain cue words based on the hallucination detection results.
[0014] On the other hand, the present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the above-mentioned prompt word optimization method when executing the program.
[0015] On the other hand, the present invention also provides a machine-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the above-described prompt word optimization method.
[0016] On the other hand, the present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the above-mentioned prompt word optimization method.
[0017] Through the above technical solution, this invention matches the constraint rules corresponding to financial instructions using a preset template library, and constructs enhanced thought chain prompts based on the financial instructions, the example, the constraint rules, and the reasoning path. This allows the invention to embed financial-specific constraint rules into the enhanced thought chain prompts, improving their professional adaptability to the financial field. Furthermore, an illusion detection model is used to detect illusions in the enhanced thought chain prompts and determine optimization strategies. This invention reduces the probability of illusions in prompts by optimizing their professional adaptability and detecting illusions, thereby achieving illusion optimization of large language models in the financial field.
[0018] Other features and advantages of the embodiments of the present invention will be described in detail in the following detailed description section. Attached Figure Description
[0019] The accompanying drawings are provided to further illustrate embodiments of the present invention and form part of the specification. They are used together with the following detailed description to explain the embodiments of the present invention, but do not constitute a limitation thereof. In the drawings: Figure 1 This is one of the flowcharts illustrating the prompt word optimization method provided by the present invention; Figure 2 This is the second flowchart illustrating the prompt word optimization method provided by the present invention; Figure 3 This is a schematic diagram of the prompt word optimization device provided by the present invention; Figure 4 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation
[0020] The specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are for illustration and explanation only and are not intended to limit the scope of the present invention.
[0021] Method Implementation Examples Please refer to Figure 1 This invention provides a method for optimizing prompt words, including: Step 100: Obtain input tasks, wherein the input tasks include at least financial instructions.
[0022] Electronic devices acquire user input tasks. These input tasks include at least financial instructions. In other embodiments, the input task may also include a target model of the user's needs (a large language model). The financial instructions are the financial-related needs that the user submits to the large language model. For example, in one embodiment, the financial instruction could be "Calculate the approvable loan amount and risk level for Ms. Zhang, who has a monthly income of 20,000 yuan and no debt, applying for a 3-year, 300,000 yuan credit loan." The large language model can be any commercially available large language model, such as Gemini, ChatGPT, or Wenxin Yiyan.
[0023] Step 200: Determine the example and constraint rules corresponding to the financial instruction based on the preset template library, and generate the reasoning path corresponding to the financial instruction; the preset template library includes constraint rules associated with different financial keywords.
[0024] Step 300: Construct enhanced thought chain prompts using the financial instructions, the examples, the constraint rules, and the reasoning path.
[0025] This invention generates high-precision enhanced thought chain prompts adapted to the financial field (such as personal loan calculation and corporate credit rating) based on a preset template library. The preset template library can be stored in a MySQL database. In one embodiment, the preset template library can be a table structure containing template ID, scenario name, constraint rule library, and example library. The scenario name can be a personal loan calculation scenario, a corporate credit rating example, etc. The examples in the example library represent the reasoning path and answer for the sample task. For example, each example in the example library can be set as a "question-reasoning-answer" triple. For example, in one embodiment, a financial scenario is illustrated: "Question: What is the monthly payment for a 200,000 yuan loan over 2 years (LPR 3.45%)? → Reasoning: Monthly payment = 200,000 × (3.45% / 12) × (1 + 3.45% / 12)^24 / [(1 + 3.45% / 12)^24 - 1] ≈ 8,740 yuan, Monthly payment / income ratio = 8,740 / 15,000 ≈ 58% (high risk) → Answer: Monthly payment 8,740 yuan, high risk." The constraint rule base includes a domain terminology base and / or an industry rule base. That is, in one embodiment, the constraint rule base includes a domain terminology base. In another embodiment, the constraint rule base includes an industry rule base. In yet another embodiment, the constraint rule base includes both a domain terminology base and an industry rule base. That is, the constraint rules include domain terms and / or industry rules. The domain terminology base includes various professional terms in the financial field. For example, the domain terminology library includes the LPR interest rate, monthly payment / income ratio, and current effective values (such as the LPR interest rate of 3.45%). The industry rule library includes various regulatory requirements and calculation standards in the financial field. For example, the industry rule library includes requirements to calculate the monthly payment / income ratio, marking "high risk" if it is >50%, and includes the formula for equal principal and interest payments. The reasoning path represents the reasoning and analysis path for financial instructions. For example, for the question "What is the monthly payment for a 200,000 yuan loan with a 2-year term (LPR 3.45%)?", the reasoning and analysis path could be Let's think step by step, first calculating the monthly payment, and then analyzing the risk level. The preset template library of this invention embodiment needs to be reviewed by financial risk control experts with more than 5 years of experience (score ≥80 points) and updated quarterly according to policy. It should be noted that when facing instructions from other fields (such as instructions in the medical field), the preset template library is the scenario name, constraint rule library, and example library of the financial field.
[0026] Electronic devices can access a pre-defined template library and determine the corresponding examples and constraint rules for financial instructions through keyword matching. For example, for a financial instruction: calculate the monthly payment and risk level of user Zhang San's 3-year personal consumption loan of 300,000 yuan, which requires step-by-step reasoning, the electronic device, through keyword matching of "calculate," "monthly payment," and "risk," obtains the following example: "Question: Monthly payment for a 2-year loan of 200,000 yuan (LPR 3.45%)? → Reasoning: Monthly payment = 200,000 × (3.45% / 12) × (1 + 3.45% / 12)^24 / [(1 + 3.45% / 12)^24 - 1] ≈ 8,740 yuan, Monthly payment / income ratio = 8,740 / 15,000 ≈ 58% (high risk) → Answer: Monthly payment 8,740 yuan, high risk." The constraint rule is also obtained: the current LPR interest rate (3.45%) must be used, and a monthly payment / income ratio > 50% is marked as "high risk." It should be noted that multiple (3) adaptation examples can be randomly selected from the example library. The path sampling of the electronic device uses the transformers.pipeline interface to generate the inference path, and the temperature is set to 0.8. Finally, the electronic device integrates the prompts into a fixed structure, including [Instruction], [Example], [Constraint Rule], and [Inference Path], which clearly defines the task requirements, examples, rules, and step guidance. That is, the electronic device integrates in the order of "Instruction → Example → Constraint Rule → Inference Path" to generate the final enhanced thinking chain prompts. In one embodiment, the example of the enhanced thinking chain prompts for the financial scenario is as follows: "Instruction: Calculate the monthly payment and risk level of user Zhang San's 300,000 yuan 3-year personal consumption loan, which requires step-by-step reasoning; Example: "Question: What is the monthly payment of a 200,000 yuan 2-year loan (LPR 3.45%)?" → Reasoning: Monthly payment = 200000 × (3.45% / 12) × (1 + 3.45% / 12)^24 / [(1 + 3.45% / 12)^24 - 1] ≈ 8740 yuan, Monthly payment / income ratio = 8740 / 15000 ≈ 58% (high risk) → Answer: Monthly payment 8740 yuan, high risk; Constraint: The current LPR interest rate (3.45%) must be used, and a monthly payment / income ratio > 50% should be marked as "high risk"; Reasoning guidance: Let's think step by step, first calculate the monthly payment, then analyze the risk level.
[0027] In another embodiment, an example of an enhanced thought chain prompt for a financial scenario is as follows:
Instruction
Example
[0028] In another embodiment, an example of an enhanced thought chain prompt for a financial scenario is as follows:
Instruction
Example
Constraints
Reasoning Guidance
[0029] Existing thought chains often use general examples and logical bases, lacking a "domain-specific terminology library and / or constraint rule library," thus failing to dynamically match the knowledge system of specialized tasks and causing model outputs to deviate from domain norms. This invention embeds domain-specific terminology and industry rules into existing thought chains, generating technically integrated financial terms (such as LPR interest rates and debt-to-equity ratios) and industry rules, thereby overcoming the professional adaptation deficiencies of existing thought chains.
[0030] Step 400: Perform hallucination detection on the enhanced thought chain cue words based on the hallucination detection model, and determine the optimization strategy for the enhanced thought chain cue words based on the hallucination detection results.
[0031] The electronic device performs hallucination detection on the enhanced thought chain cues using a hallucination detection model, and determines an optimization strategy for the enhanced thought chain cues based on the hallucination detection results. Specifically, the hallucination detection can employ a RoBERTa-large hallucination detector pre-trained on the FEVER fact-checking dataset and fine-tuned on a domain dataset. The hallucination detector is based on the RoBERTa-large model and pre-trained on a financial fact-checking dataset containing 200,000 samples (learning rate 1e). -5 The model is fine-tuned on a "Financial Reasoning Illusion Dataset" containing 8000 samples (batch 8, round 3, weighted cross-entropy loss). The illusion detection model detects illusions based on enhanced thought chain cues, outputs an illusion detection index, and the electronic device calculates an illusion score based on the illusion detection index. The illusion detection result is determined based on the illusion score, and the optimization strategy for the enhanced thought chain cues is determined based on the illusion detection result.
[0032] In one embodiment, the process of performing hallucination detection on the enhanced thought chain prompt based on a hallucination detection model, and determining an optimization strategy for the enhanced thought chain prompt based on the hallucination detection results, includes: performing hallucination detection on the enhanced thought chain prompt based on the hallucination detection model to obtain the proportion of hallucination sentences and the number of incorrect reasoning steps for the enhanced thought chain prompt; calculating a hallucination score based on the proportion of hallucination sentences and the number of incorrect reasoning steps; determining that the hallucination detection result meets the requirements if the hallucination score is less than or equal to a set score threshold; and outputting the enhanced thought chain prompt if the hallucination detection result meets the requirements. The electronic device uses a hallucination detection model to perform hallucination detection on the enhanced thought chain prompt, obtaining the S1 proportion of hallucination sentences and the S2 number of incorrect reasoning steps for the enhanced thought chain prompt. The electronic device calculates the hallucination score based on the sum of the proportion of hallucination sentences and the number of incorrect reasoning steps.
[0033] In other aspects of this invention, calculating the hallucination score based on the proportion of hallucinatory sentences and the number of erroneous reasoning steps includes: calculating the hallucination score based on a weighted sum of the proportion of hallucinatory sentences and the number of erroneous reasoning steps. The electronic device can calculate the hallucination score based on the weighted sum of the proportion of hallucinatory sentences and the number of erroneous reasoning steps. For example, the weight of the proportion of hallucinatory sentences is 0.5, and the weight of the number of erroneous reasoning steps is 0.5. That is, both are normalized to 0-1. That is, the hallucination score = 0.5 × S1 + 0.5 × S2.
[0034] In this embodiment of the invention, a pre-set scoring threshold is implemented. If the hallucination score is less than or equal to the pre-set threshold (e.g., 0.2), the hallucination detection result is determined to meet the requirements, and the enhanced thought chain prompt is output. If the hallucination score is greater than the pre-set threshold, the hallucination detection result is determined to not meet the requirements; if the hallucination detection result does not meet the requirements, the enhanced thought chain prompt is iteratively optimized.
[0035] This invention, in its embodiments, matches constraint rules corresponding to financial instructions using a preset template library. Based on the financial instructions, the example, the constraint rules, and the reasoning path, it constructs enhanced thought chain prompts. This allows the invention to embed financial-specific constraint rules into the enhanced thought chain prompts, improving their professional adaptability to the financial field. Furthermore, it uses an illusion detection model to detect illusions in the enhanced thought chain prompts and determine optimization strategies for them. This invention reduces the probability of illusions in prompts by optimizing their professional adaptability and detecting illusions, thereby achieving illusion optimization of large language models in the financial field.
[0036] In other aspects of this invention, the input task further includes a target model; after obtaining the input task in step 100, the method further includes: using a task complexity detection model to perform task complexity detection on the financial instruction to obtain a complexity score; performing model parameter detection on the target model to obtain the model parameter values of the target model; and determining that the financial instruction meets the requirements if the complexity score is greater than or equal to a set score threshold and / or the model parameter values are greater than or equal to a set parameter threshold.
[0037] To mitigate the risk of illusion at the data source, this invention establishes domain and / or model fit screening to exclude incompatible scenarios and ensure the effectiveness of subsequent enhanced thought chain prompts. For example, in one embodiment, the electronic device performs model parameter detection on the target model of the input task. Model parameter detection can utilize the AutoConfig utility class from the Hugging FaceTransformers library, obtaining the model parameter count through the model.config.num_parameters interface and converting it to the order of "B". By calling the configuration interface of a large model (such as model.config from the Hugging Face Transformers library), the model parameter values of the target model for the input task are extracted. Only target models with model parameter values greater than or equal to a set parameter threshold (e.g., model parameter value ≥ 20B) are allowed to proceed to the next stage. If the model parameter value is less than the set threshold (e.g., model parameter value < 20B), a prompt "Model does not fit thought chain (CoT), basic instruction prompts recommended" is output. Because small models lack sufficient pre-trained knowledge, they cannot understand the "atomic knowledge" and multi-step reasoning logic required by the thought chain; blindly applying the thought chain may lead to a break in reasoning logic, inducing illusion. For example, when existing thought chains are applied to small models with fewer than 20 parameters, the illusion rate increases significantly (e.g., a model with 7 parameters has an illusion rate of 35% in a medical reasoning task, while a model with 34 parameters has an illusion rate of only 8%). This invention addresses the problem of financial logic breakdowns caused by blindly applying thought chains by eliminating models with small parameters through model parameter detection.
[0038] In another embodiment, the electronic device performs task complexity detection on the financial instructions input into the task, obtaining a complexity score. The task complexity score can be achieved using a task complexity detection model pre-trained on BERT-based and fine-tuned on an "inference task complexity dataset." The task complexity detection model is built on a BERT-based pre-trained model, with the training dataset being a "financial task complexity annotation set" containing 100,000 samples. The training parameters are set to batch size 16, learning rate 2e-5, training epochs 3, the optimizer is AdamW, and the loss function is mean squared error. The financial instructions are input into the task complexity detection model, which outputs the number of high-weight keywords (e.g., "calculate," "derive," "analyze," "measure," "evaluate") and the number of low-weight keywords (e.g., "select," "summary," "query," "enter"). High-weight keywords add 0.3 points, low-weight keywords subtract 0.2 points, and keywords containing ≥2 high-weight keywords add an extra 0.1 points. The electronic device obtains the complexity score through keyword weighting and sentence structure analysis. For example, if the task complexity detection model detects 3 high-weight keywords and 1 low-weight keyword, then the complexity score = 0.3*3 - 0.2*1 + 0.1 = 0.8. A complexity score greater than or equal to a set threshold (complexity score ≥ 0.6) is considered a complex task (enabling enhanced thought chain prompts). A score less than the set threshold (complexity score < 0.6) is considered a simple task, and the message "Model is not compatible with thought chain (CoT), basic instruction prompts are recommended" is output. In existing simple tasks (such as single-choice questions and text summarization), thought chains cannot improve performance but instead increase redundant output. This invention eliminates simple tasks through task complexity detection, solving the problem of financial logic breaks caused by blindly applying thought chains (CoT).
[0039] Please refer to Figure 2 In another embodiment, the electronic device performs domain-model adaptation screening, eliminating incompatible scenarios through dual screening of the target model and the task. Specifically, after inputting financial instructions and the target model, the adaptation screening confirms that only models with parameter values greater than or equal to a set parameter threshold and task complexity greater than or equal to a set scoring threshold are eligible to initiate the generation process of enhanced thought chain prompts. This embodiment of the invention further addresses the financial logic break caused by blindly applying thought chains by excluding models with small parameter values (<20B) and simple tasks. By eliminating incompatible scenarios through dual screening of the target model and the task, the risk of illusion is further reduced from the source.
[0040] In other aspects of the embodiments of the present invention, the inference path includes multiple paths, and generating the inference path corresponding to the financial instruction includes: randomly generating multiple initial inference paths corresponding to the financial instruction; calculating the path perplexity of each initial inference path using a path perplexity model; and selecting at least two initial inference paths from the multiple initial inference paths whose path perplexity is less than or equal to a set perplexity threshold as multiple inference paths corresponding to the financial instruction.
[0041] The electronic device can randomly sample and generate 5-8 initial inference paths corresponding to financial instructions using the `transformers.pipeline` interface with a temperature coefficient of 0.7-0.9. Path perplexity is then calculated using a path perplexity model (e.g., the RoBERTa-large pre-trained model). The RoBERTa-large model can be pre-trained on the WikiText-103 financial subset dataset. From these initial inference paths, at least two initial inference paths with a perplexity less than or equal to a set perplexity threshold are selected as the multiple inference paths corresponding to the financial instruction. For example, the electronic device removes invalid inference paths with a perplexity > 10, retaining 2-5 high-confidence inference paths. The electronic device can generate multiple inference paths, facilitating subsequent multi-path voting for consistency and dynamic pruning of path consistency, balancing accuracy and efficiency.
[0042] In other aspects of this invention, after performing hallucination detection on the enhanced thought chain prompts based on the hallucination detection model and determining the optimization strategy for the enhanced thought chain prompts based on the hallucination detection results, the method further includes: determining the conclusions corresponding to all the reasoning paths; calculating the path perplexity of each reasoning path when there are fewer than a set threshold of identical conclusions among the conclusions corresponding to all the reasoning paths; determining the reasoning path with the smallest path perplexity among all the reasoning paths as the target reasoning path; and deleting all reasoning paths in the enhanced thought chain prompts that exclude the target reasoning path.
[0043] This invention verifies the consistency of multiple inference paths within an enhanced thought chain prompt. Specifically, the electronic device determines the conclusions corresponding to all inference paths and performs a "majority vote" on the conclusions of multiple inference paths (e.g., 2-5 inference paths). If the conclusions of multiple inference paths are consistent (e.g., ≥60%) (e.g., more than or equal to 3 inference paths out of 5 have consistent conclusions), then that conclusion is adopted, and the multiple inference paths are retained. If they are inconsistent, the perplexity of each inference path is recalculated. The electronic device calculates the path perplexity using a path perplexity model (e.g., a RoBERTa-large pre-trained model). The electronic device determines the inference path with the lowest path perplexity among all the inference paths as the target inference path, i.e., selects the path with the lowest path perplexity as the baseline conclusion (target inference path), and deletes all inference paths in the enhanced thought chain prompt that exclude the target inference path. Thus, this invention integrates self-consistent multi-path voting with dynamic pruning of path consistency, balancing accuracy and efficiency. Furthermore, illusion detection and path verification identify minor data contradictions (e.g., fabricated revenue) or logical errors (e.g., ignoring the impact of loan terms).
[0044] In other aspects of embodiments of the present invention, the method further includes: Repeat the following steps until the set termination condition is met: if the hallucination score is greater than the set score threshold, determine the hallucination type of the current enhanced mind chain prompt; determine the hallucination cause corresponding to the hallucination type based on the preset mapping relationship; determine the prompt adjustment strategy based on the hallucination cause; adjust the current enhanced mind chain prompt based on the prompt adjustment strategy; calculate the hallucination score of the adjusted enhanced mind chain prompt; wherein, the adjusted enhanced mind chain prompt is used as the current enhanced mind chain prompt for the next iteration.
[0045] When the hallucination score exceeds a set threshold, the electronic device iteratively optimizes the enhanced thought chain prompts. If the hallucination score exceeds the threshold, the hallucination detection model outputs the hallucination type (e.g., no reference to domain rules, fabricated data, incorrect interest rate calculation, missing industry rules, etc.), thus determining the hallucination type of the current enhanced thought chain prompt. The electronic device then determines the hallucination cause corresponding to the hallucination type based on a preset mapping relationship (i.e., a hallucination type-hallmark mapping table). This preset mapping relationship includes multiple hallucination types and the hallucination cause associated with each type. For example, in one embodiment, "no reference to domain rules" corresponds to "the prompt constraint rule library lacks this clause," "fabricated data" corresponds to "the domain terminology library has not been updated with the latest data," "incorrect interest rate calculation" corresponds to "no reference to the latest LPR," and "missing industry rules" corresponds to "missing constraint rule library." Once the hallucination cause is determined, a prompt adjustment strategy is then determined based on that cause. In this embodiment, the prompt adjustment strategy is implemented with targeted adjustments based on the hallucination cause. Similarly, a hallucination cause-adjustment strategy mapping table can also be established. The illusion cause-adjustment strategy mapping table records the adjustment strategies corresponding to different illusion causes. For example, if constraints are missing, industry rules are added ("LPR update time needs to be marked"); if there are insufficient examples, 1-2 examples containing missing steps are added; if the terminology is outdated, the domain terminology library is updated. The electronic device adjusts the current enhanced mind chain prompts based on the prompt adjustment strategy. The electronic device then calculates the illusion score of the adjusted enhanced mind chain prompts; wherein, the adjusted enhanced mind chain prompts are used as the current enhanced mind chain prompts for the next iteration. The illusion score can be calculated using the method in step 400 above, which will not be repeated here. The electronic device then repeats the "generate → detect" process for the adjusted enhanced mind chain prompts. That is, the electronic device repeatedly executes the above steps for iterative optimization. When the iteration termination condition (set end condition) is that the illusion score is less than or equal to the set score threshold (e.g., ≤0.2) for two consecutive times, or the number of iterations is greater than or equal to the set number threshold (e.g., ≥3), the iteration ends. In addition, if the number of iterations is greater than or equal to the set number threshold but not met, the electronic device prompts "Manual update of the preset template library is required".
[0046] Therefore, this application forms an optimization closed loop based on the hallucination detection results through "cause localization → rule mapping → prompt word adjustment," that is, a closed loop based on the hallucination detection results, including hallucination cause localization, adjustment strategies, and iteration termination conditions. Existing thought chains do not establish a mapping relationship between "hallucination indicators" and "prompt word adjustment rules," nor do they establish a closed-loop mechanism of "detection-analysis-optimization," thus failing to transform output problems into prompt word improvement actions. Existing thought chains only cover a unidirectional process of "prompt word generation → model output," failing to associate hallucination detection results with prompt word adjustment, leading to the repeated occurrence of the same type of hallucination. This invention reduces recurring hallucinations caused by template defects by enhancing the iterative optimization closed loop of thought chain prompt words. Furthermore, this method can reduce manual review workload and omissions, improving the efficiency of financial operations.
[0047] In other aspects of the embodiments of the present invention, model adaptation detection can be extended in the step of calculating model parameter values. For example, a "model training data matching degree" detection can be added (such as using cosine similarity to determine the overlap between the model's training data and the task domain to assess task complexity), further improving adaptation accuracy. Additionally, the embodiments of the present invention can also replace dynamic pruning algorithms. For example, "reasoning step redundancy analysis" can be used instead of perplexity detection (such as calculating the semantic repetition rate between steps and eliminating reasoning paths with a repetition rate higher than a set threshold, for example, eliminating reasoning paths with a repetition rate > 60%). Furthermore, the embodiments of the present invention can also expand hallucination detection indicators. For example, "self-contradiction rate" (such as the proportion of conflicting conclusions between inferences) and "factual basis missing rate" (the proportion of sentences without labeled data sources) can be added, and the "self-contradiction rate," "factual basis missing rate," and their weights can be added to the calculation of hallucination scores, refining the dimensions of hallucination evaluation.
[0048] In summary, this invention establishes a dual screening mechanism of "model parameters - task complexity" for the first time, avoiding blindly applying thought processes to small models and simple tasks, thus reducing the risk of illusions from the outset. Furthermore, this invention integrates self-consistent multi-path voting with dynamic pruning of path consistency, balancing accuracy and efficiency, while embedding domain terminology and constraints to improve adaptability to specialized scenarios. This invention constructs an iterative "detection-cause-optimization" mechanism, transforming output illusions into improved actions for prompting words, overcoming the limitations of the "one-way process" in existing technologies.
[0049] In other words, this invention proposes a large-model illusion suppression prompt word optimization method for financial scenarios. It aims to accurately and efficiently reduce errors or unfounded content generated by large models, preventing unreliable outputs from entering financial processes such as loan approval and risk assessment, thereby reducing fraud and decision-making risks. This is achieved through: ① "Domain-Model Fit Screening" to exclude low-parameter models (<20 bytes) and simple tasks, resolving financial logic breaks caused by blindly applying CoT (Cooperation of Thought) mechanisms; ② Dynamically enhancing CoT prompt words by integrating financial terminology (such as LPR interest rate and debt-to-equity ratio) with industry rules (or regulatory rules), compensating for the professional adaptation deficiencies of general prompt words; ③ Illusion detection and path verification modules to identify minor data contradictions (such as fabricated revenue) or logical errors (such as ignoring the impact of loan terms); and ④ Iterative prompt word optimization loops to reduce recurring illusions caused by template defects. This method can reduce manual review workload and omissions, improving the efficiency of financial operations.
[0050] In one embodiment, the user inputs the task "Calculate the approvable loan amount and risk level for Ms. Zhang, who has a monthly income of 20,000 yuan and no debt, applying for a 3-year credit loan of 300,000 yuan," and passes in a large model with 35B parameters. The system is adapted and screened: model parameter values 35B ≥ 20B, task complexity score 0.65 ≥ 0.6, and a dynamic enhanced thought chain prompt generation process is enabled. Enhanced thought chain prompts containing the constraint rule "debt-to-income ratio ≤ 50%", the loan amount calculation formula (average monthly income × 12 × 3 × 30%), and similar examples are generated. After model output, if the hallucination detection score is 0.1 (no hallucination sentences, 0 error steps), the result is directly output: an approvable loan amount of 216,000 yuan, low risk level. If the score exceeds 0.2, the "monthly payment needs to be calculated using LPR 3.45%" constraint is added to optimize the prompt, and the test is retested until the target is met.
[0051] In another embodiment, the technical effects are as follows: After inputting financial instructions and the target model, and confirming through adaptation and screening that the model parameter values are ≥20B and the task complexity is ≥0.6, a dynamic enhanced thought chain prompt generation process is initiated. Based on a preset template library, enhanced thought chain prompts containing terminology, industry rule constraints (such as "LPR 3.45%" and "asset-liability ratio ≤70%)) and examples are generated. After the model outputs, a score is calculated using an illusion detector. If the score is ≤0.2, it is directly output; otherwise, iterative optimization is performed (such as supplementing credit impact constraints). In practical applications, the illusion rate of complex financial reasoning is reduced by 40%-65% (e.g., the illusion rate of calculating mortgage monthly payments decreases from 22.5% to 3.8%), reasoning latency is shortened by more than 30%, and manual review is reduced to 15%-20%, significantly improving reliability and efficiency.
[0052] Device Examples Please refer to Figure 3 On the other hand, embodiments of the present invention also provide a prompt word optimization device, including: Acquisition module 301 is used to acquire input tasks, wherein the input tasks include at least financial instructions; The first determining module 302 is used to determine the example and constraint rules corresponding to the financial instruction based on a preset template library, and to generate the reasoning path corresponding to the financial instruction; the preset template library includes constraint rules associated with different financial keywords; Module 303 is used to construct enhanced thought chain prompts using the financial instructions, the examples, the constraint rules, and the reasoning path; The second determining module 304 is used to perform hallucination detection on the enhanced thought chain cue words based on the hallucination detection model, and to determine the optimization strategy of the enhanced thought chain cue words based on the hallucination detection results.
[0053] Optionally, the input task further includes a target model; the apparatus further includes: The first detection module is used to perform task complexity detection on the financial instruction using a task complexity detection model to obtain a complexity score. The second detection module is used to perform model parameter detection on the target model and obtain the model parameter values of the target model; The third determining module is used to determine that the financial instruction meets the requirements if the complexity score is greater than or equal to a set score threshold and / or the model parameter value is greater than or equal to a set parameter threshold.
[0054] Optionally, the inference path includes multiple paths, and the inference path for generating the financial instruction includes: Multiple initial inference paths corresponding to the financial instruction are randomly generated; The path perplexity of each initial inference path is calculated using a path perplexity model; From the multiple initial inference paths, at least two initial inference paths with a path perplexity less than or equal to a set perplexity threshold are selected as the multiple inference paths corresponding to the financial instruction.
[0055] Optionally, the process of performing hallucination detection on the enhanced thought chain cue words based on the hallucination detection model, and determining the optimization strategy for the enhanced thought chain cue words based on the hallucination detection results, includes: Based on the hallucination detection model, hallucination detection is performed on the enhanced thought chain cue words to obtain the proportion of hallucination sentences and the number of erroneous reasoning steps of the enhanced thought chain cue words; The hallucination score is calculated based on the proportion of hallucinatory sentences and the number of incorrect reasoning steps. If the hallucination score is less than or equal to a set score threshold, the hallucination detection result is determined to meet the requirements. If the hallucination detection results meet the requirements, the enhanced thought chain prompt is output.
[0056] Optionally, calculating the hallucination score based on the proportion of hallucinatory sentences and the number of erroneous reasoning steps includes: The hallucination score is calculated based on the weighted sum of the proportion of hallucinatory sentences and the number of erroneous reasoning steps.
[0057] Optionally, the device further includes: The fourth determining module is used to determine the conclusions corresponding to all the reasoning paths; The calculation module is used to calculate the path perplexity of each inference path when there are fewer than a set threshold of identical conclusions among all the conclusions corresponding to the inference paths. The fifth determining module is used to determine the inference path with the minimum path perplexity among all the inference paths as the target inference path; The deletion module is used to delete all reasoning paths in the enhanced thought chain prompts that exclude the target reasoning path.
[0058] Optionally, the device further includes: An iterative module is used to repeatedly execute the following steps until a set termination condition is reached: If the hallucination score is greater than a set score threshold, determine the hallucination type of the current enhanced mind chain prompt; determine the hallucination cause corresponding to the hallucination type based on a preset mapping relationship; the preset mapping relationship includes multiple hallucination types and the hallucination cause associated with each hallucination type; determine a prompt adjustment strategy based on the hallucination cause; adjust the current enhanced mind chain prompt based on the prompt adjustment strategy; calculate the hallucination score of the adjusted enhanced mind chain prompt; wherein the adjusted enhanced mind chain prompt is used as the current enhanced mind chain prompt for the next iteration.
[0059] Optionally, the constraint rules may include domain terminology and / or industry rules.
[0060] The prompt word optimization device includes a processor and a memory. The aforementioned acquisition module 301, first determination module 302, construction module 303, and second determination module 304 are all stored in the memory as program units. The processor executes the aforementioned program units stored in the memory to achieve the corresponding functions.
[0061] A processor contains a kernel, which retrieves the corresponding program units from memory. One or more kernels can be configured.
[0062] The memory may include non-permanent memory in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM, and the memory includes at least one memory chip.
[0063] Figure 4 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 4 As shown, the electronic device may include a processor 410, a communications interface 420, a memory 430, and a communication bus 440, wherein the processor 410, communications interface 420, and memory 430 communicate with each other via the communication bus 440. The processor 410 can call logical instructions in the memory 430 to execute a prompt word optimization method. This method includes: acquiring an input task, the input task including at least a financial instruction; determining an example, constraint rules, and inference path corresponding to the financial instruction based on a preset template library; the preset template library including constraint rules associated with different financial keywords; constructing an enhanced thought chain prompt word using the financial instruction, the example, the constraint rules, and the inference path; performing hallucination detection on the enhanced thought chain prompt word based on a hallucination detection model; and determining an optimization strategy for the enhanced thought chain prompt word based on the hallucination detection results.
[0064] Furthermore, the logical instructions in the aforementioned memory 430 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0065] On the other hand, the present invention also provides a computer program product, the computer program product comprising a computer program that can be stored on a machine-readable storage medium, wherein when the computer program is executed by a processor, the computer is capable of executing a prompt word optimization method, the method comprising: acquiring an input task, the input task including at least a financial instruction; determining an example, constraint rules corresponding to the financial instruction based on a preset template library, and generating a reasoning path corresponding to the financial instruction; the preset template library including constraint rules associated with different financial keywords; constructing an enhanced thought chain prompt word using the financial instruction, the example, the constraint rules, and the reasoning path; performing hallucination detection on the enhanced thought chain prompt word based on a hallucination detection model, and determining an optimization strategy for the enhanced thought chain prompt word based on the hallucination detection result.
[0066] In another aspect, the present invention also provides a machine-readable storage medium storing a computer program thereon, which, when executed by a processor, implements a prompt word optimization method. The method includes: acquiring an input task, the input task including at least a financial instruction; determining an example, constraint rules, and a reasoning path corresponding to the financial instruction based on a preset template library; the preset template library including constraint rules associated with different financial keywords; constructing enhanced thought chain prompt words using the financial instruction, the example, the constraint rules, and the reasoning path; performing hallucination detection on the enhanced thought chain prompt words based on a hallucination detection model; and determining an optimization strategy for the enhanced thought chain prompt words based on the hallucination detection results.
[0067] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0068] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0069] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for optimizing prompt words, characterized in that, include: Obtain an input task, wherein the input task includes at least a financial instruction; The example and constraint rules corresponding to the financial instruction are determined based on a preset template library, and the reasoning path corresponding to the financial instruction is generated; the preset template library includes constraint rules associated with different financial keywords; The financial instructions, the examples, the constraint rules, and the reasoning path are used to construct enhanced thought chain prompts; The enhanced thought chain cue words are subjected to hallucination detection based on a hallucination detection model, and an optimization strategy for the enhanced thought chain cue words is determined based on the hallucination detection results.
2. The prompt word optimization method according to claim 1, characterized in that, The input task also includes the target model; after obtaining the input task, it also includes: The task complexity of the financial instruction is detected using a task complexity detection model to obtain a complexity score; The target model is subjected to model parameter detection to obtain the model parameter values of the target model; If the complexity score is greater than or equal to a set score threshold and / or the model parameter value is greater than or equal to a set parameter threshold, the financial instruction is determined to meet the requirements.
3. The prompt word optimization method according to claim 1, characterized in that, The inference path includes multiple paths, and the inference path for generating the financial instruction includes: Multiple initial inference paths corresponding to the financial instruction are randomly generated; The path perplexity of each initial inference path is calculated using a path perplexity model; From the multiple initial inference paths, at least two initial inference paths with a path perplexity less than or equal to a set perplexity threshold are selected as the multiple inference paths corresponding to the financial instruction.
4. The prompt word optimization method according to claim 1, characterized in that, The process of performing hallucination detection on the enhanced thought chain cue words based on the hallucination detection model, and determining the optimization strategy for the enhanced thought chain cue words based on the hallucination detection results, includes: Based on the hallucination detection model, hallucination detection is performed on the enhanced thought chain cue words to obtain the proportion of hallucination sentences and the number of erroneous reasoning steps of the enhanced thought chain cue words; The hallucination score is calculated based on the proportion of hallucinatory sentences and the number of incorrect reasoning steps. If the hallucination score is less than or equal to a set score threshold, the hallucination detection result is determined to meet the requirements. If the hallucination detection results meet the requirements, the enhanced thought chain prompt is output.
5. The prompt word optimization method according to claim 4, characterized in that, The calculation of the hallucination score based on the proportion of hallucinatory sentences and the number of erroneous reasoning steps includes: The hallucination score is calculated based on the weighted sum of the proportion of hallucinatory sentences and the number of erroneous reasoning steps.
6. The prompt word optimization method according to claim 3, characterized in that, The method further includes: Determine the conclusions corresponding to all the aforementioned reasoning paths; If, among all the conclusions corresponding to the inference paths, there are fewer than a set threshold of identical conclusions, calculate the path perplexity of each inference path. The inference path with the minimum path perplexity among all the inference paths is identified as the target inference path. Delete all reasoning paths from the enhanced thought chain prompts that exclude the target reasoning path.
7. The prompt word optimization method according to claim 4, characterized in that, The method further includes: Repeat the following steps until the set termination condition is met: If the hallucination score is greater than a set score threshold, determine the hallucination type of the current enhanced mind chain cue word; The hallucination cause corresponding to the hallucination type is determined based on a preset mapping relationship; the preset mapping relationship includes multiple hallucination types and the hallucination cause associated with each hallucination type; Based on the causes of the hallucinations, a cue word adjustment strategy was determined; The current enhanced mind chain prompts are adjusted based on the prompt adjustment strategy; Calculate the illusion score of the adjusted enhanced mind chain cue; wherein the adjusted enhanced mind chain cue is used as the current enhanced mind chain cue in the next iteration.
8. The prompt word optimization method according to claim 1, characterized in that, The constraints include domain terminology and / or industry rules.
9. A prompt word optimization device, characterized in that, include: An acquisition module is used to acquire input tasks, wherein the input tasks include at least financial instructions; The first determining module is used to determine the example and constraint rules corresponding to the financial instruction based on a preset template library, and to generate the reasoning path corresponding to the financial instruction; the preset template library includes constraint rules associated with different financial keywords; A construction module is used to construct enhanced thought chain prompts using the financial instructions, the examples, the constraint rules, and the reasoning path; The second determining module is used to perform hallucination detection on the enhanced thought chain cue words based on the hallucination detection model, and to determine the optimization strategy for the enhanced thought chain cue words based on the hallucination detection results.
10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the prompt word optimization method according to any one of claims 1 to 8.
11. A machine-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the prompt word optimization method according to any one of claims 1 to 8.
12. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the prompt word optimization method according to any one of claims 1 to 8.