Interaction optimization method and device based on expectation confirmation, equipment and medium

By vectorizing and modeling user historical interaction data and performing satisfaction evolution modeling, the system identifies and optimizes dialogue turns that do not meet expectations. This solves the problem that large language models cannot dynamically capture user satisfaction in the financial and healthcare fields, and achieves more efficient user expectation perception and preference optimization.

CN121901467APending Publication Date: 2026-04-21PING AN TECH (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
PING AN TECH (SHENZHEN) CO LTD
Filing Date
2026-01-09
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

Existing large language models cannot dynamically capture user satisfaction trends in the financial and healthcare fields, making it difficult for the system to perceive and utilize relevant information on user expectations in a timely and accurate manner, and making it impossible to effectively perform multi-round preference optimization. The dynamic changes in user satisfaction are difficult to capture and utilize.

Method used

By acquiring historical user interaction data for vectorized modeling, constructing an expected representation object, performing user satisfaction evolution modeling, identifying dialogue turns that do not meet preset expected conditions, making preference corrections, inputting the preset user simulation model for expected confirmation feedback, and optimizing the interaction model parameters.

Benefits of technology

It enables large language models to dynamically capture user satisfaction trends, improves model output efficiency, enhances the perception and utilization of user expectations, optimizes preferences in multi-turn interaction processes, and improves the ability to capture dynamic changes in user satisfaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121901467A_ABST
    Figure CN121901467A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data processing, and discloses an interaction optimization method and device based on expectation confirmation, equipment and a medium, which can be applied to the financial and medical fields, and the method comprises the following steps: carrying out vectorization modeling processing and user satisfaction evolution modeling processing on user historical interaction data to obtain a satisfaction evolution object; performing recognition correction processing on the satisfaction evolution object to obtain a preference correction object; and performing expectation confirmation feedback processing and preference optimization processing on the preference correction object to update interaction model parameters, and generating an interaction result object. In the invention, aiming at the problem that the existing large language model cannot dynamically capture the trend of the user satisfaction, preference optimization processing can be carried out on the calculated expected confirmation feedback object to update the interaction model parameters, so that the corresponding interaction result object is generated, and therefore, the user satisfaction degree can be dynamically captured. Therefore, the big language model can dynamically capture the trend of user satisfaction, and the model output efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing technology, and in particular to an interaction optimization method, apparatus, device, and medium based on expected confirmation. Background Technology

[0002] In recent years, large language models have been applied in scenarios such as intelligent customer service, claims Q&A, and product recommendation in the financial and healthcare fields. However, existing technologies are mostly based on static recommendation strategies or simple human-machine alignment mechanisms. They usually only rely on immediate feedback signals such as clicks and ratings to make local adjustments to the model. They lack systematic vectorized modeling based on users' historical interaction data and multi-round satisfaction evolution modeling mechanisms. They cannot construct user expectation representations from multi-round dialogues and quantify the deviation between user expectations and actual experience in a unified framework. As a result, the system is unable to perceive and utilize relevant information on user expectation confirmation in a timely and accurate manner to complete multi-round preference optimization. The dynamic trend of user satisfaction is difficult to capture and utilize effectively. Summary of the Invention

[0003] This invention provides an interaction optimization method, apparatus, device, and medium based on expectation confirmation to solve the technical problem that existing large language models cannot dynamically capture trends in user satisfaction.

[0004] Firstly, an interaction optimization method based on expected confirmation is provided, including: Obtain user historical interaction data and perform vector modeling to obtain the desired representation object; The user satisfaction evolution model is performed using the aforementioned expectation representation object to obtain a satisfaction evolution object; The dialogue turns that do not meet the preset expectation conditions in the satisfaction evolution object are identified and corrected to obtain the preference correction object; The preference correction object is input into a preset user simulation model for expected confirmation feedback processing to obtain the expected confirmation feedback object; The desired confirmation feedback object is subjected to preference optimization processing to update the interaction model parameters and generate the corresponding interaction result object.

[0005] Secondly, an interaction optimization device based on expected confirmation is provided, comprising: The modeling and processing module is used to acquire user historical interaction data and perform vectorized modeling to obtain the desired representation object; The data evolution module is used to perform user satisfaction evolution modeling processing using the expected representation object to obtain a satisfaction evolution object; The identification and correction module is used to identify and correct dialogue turns in the satisfaction evolution object that do not meet the preset expectation conditions, so as to obtain the preference correction object. The expectation confirmation module is used to input the preference correction object into a preset user simulation model for expectation confirmation feedback processing to obtain the expectation confirmation feedback object. The preference optimization module is used to perform preference optimization processing on the expected confirmation feedback object to update the interaction model parameters and generate the corresponding interaction result object.

[0006] Thirdly, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the computer program, implements the steps of the above-described interactive optimization method based on expected confirmation.

[0007] Fourthly, a computer-readable storage medium is provided, which stores a computer program that, when executed by a processor, implements the steps of the above-described interactive optimization method based on expected confirmation.

[0008] In the above-described scheme implemented by the interaction optimization method, apparatus, device, and medium based on expectation confirmation, historical user interaction data can be acquired and vectorized for modeling to obtain an expectation representation object; user satisfaction evolution modeling can be performed using the expectation representation object to obtain a satisfaction evolution object; dialogue turns in the satisfaction evolution object that do not meet preset expectation conditions can be identified and corrected to obtain a preference correction object; the preference correction object can be input into a preset user simulation model for expectation confirmation feedback processing to obtain an expectation confirmation feedback object; preference optimization processing can be performed on the expectation confirmation feedback object to update the interaction model parameters and generate a corresponding interaction result object. In this invention, addressing the problem that existing large language models cannot dynamically capture trends in user satisfaction, preference optimization processing can be performed on the calculated expectation confirmation feedback object to update the interaction model parameters, thereby generating a corresponding interaction result object. This allows the large language model to dynamically capture trends in user satisfaction, improving model output efficiency. Attached Figure Description

[0009] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments of the present invention will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0010] Figure 1 This is a flowchart illustrating an interaction optimization method based on expected confirmation in one embodiment of the present invention; Figure 2 yes Figure 1 A detailed implementation flow diagram of step S10 Figure 1 ; Figure 3 yes Figure 1 A detailed implementation flow diagram of step S20 Figure 2 ; Figure 4 yes Figure 1 A detailed implementation flow diagram of step S30 Figure 3 ; Figure 5 yes Figure 1 A detailed implementation flow diagram of step S40 Figure 4 ; Figure 6 yes Figure 1 A detailed implementation flow diagram of step S50 Figure 5 ; Figure 7 yes Figure 1 A flowchart illustrating a specific implementation method following step S50. Figure 6 ; Figure 8 This is a schematic diagram of an interaction optimization device based on expected confirmation in one embodiment of the present invention; Figure 9 This is a schematic diagram of the structure of a computer device according to an embodiment of the present invention; Figure 10 This is another structural schematic diagram of a computer device according to one embodiment of the present invention. Detailed Implementation

[0011] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0012] Please see Figure 1 As shown, Figure 1 A flowchart illustrating an interaction optimization method based on expected confirmation, provided in an embodiment of the present invention, includes the following steps: S10: Obtain user historical interaction data and perform vectorized modeling to obtain the desired representation object.

[0013] In the context of intelligent financial and healthcare services, historical user interaction data is acquired. This historical interaction data includes at least user input text, historical query records, and records of interaction behaviors related to insurance products. The historical interaction data is then processed into a unified vector modeling model using a pre-defined semantic modeling network. This process outputs the explicit needs and implicit concerns of users in different rounds as expectation representation objects used to characterize multidimensional expected states.

[0014] Combination Figure 2 As shown, step S10 specifically includes the following steps: S101: Collect and preprocess user input text, historical query records, and interactive behaviors to obtain raw user interaction data objects.

[0015] For various business scenarios such as intelligent customer service in finance and healthcare, claims Q&A, insurance product recommendations, and medical consultation and health management, the system uniformly collects and preprocesses user input text, historical query records, and interactive behaviors related to insurance products or medical services in multi-turn conversations. The preprocessing includes at least sentence segmentation, word segmentation, noise reduction, and annotation of timestamps and business tags on the raw text, and desensitization and structured encoding of sensitive fields such as policy numbers, ID numbers, doctor's numbers, and examination report numbers, thereby constructing raw user interaction data objects for subsequent modeling.

[0016] S102: Input the original user interaction data object into a preset semantic understanding model for embedding operation to obtain a user interaction semantic embedding object.

[0017] The raw user interaction data objects are organized into a sequence input according to the dialogue rounds and fed into a preset semantic understanding model. The input text of each round is embedded through a domain semantic encoding network based on the Transformer structure to obtain the user interaction semantic embedding object. In the financial scenario, a pre-trained model in the financial domain (such as FinBERT, LLaMA-Fin, etc.) can be used to enhance the understanding of professional terms such as claims terms, coverage responsibilities, and rate structures. In the medical scenario, a semantic model optimized for electronic medical records and medical guidelines can be used to improve the ability to parse symptom descriptions, examination items, and treatment plan expressions.

[0018] S103: Extract the instantaneous intent vector and the historical expectation trend vector from the user interaction semantic embedding object and add them together to obtain the target expectation vector object.

[0019] Based on the user interaction semantic embedding object, the immediate intent vector corresponding to the current round of input and the historical expectation trend vector formed by the aggregation of past multi-round dialogue sequences are extracted through the immediate intent extraction sub-network and the historical context encoding sub-network, respectively. The two are then fused at the vector level to form a target expectation vector object that represents the comprehensive expectation state of the user in the current round. Its formula can be expressed as: in, This indicates the text input in the current round. Indicates the historical context of the dialogue. Indicates immediate intent to extract from the network. Represents the context encoding network. Represents the desired projection matrix. Used to characterize the weight of user attention across different dimensions of expectation. This represents the target expected vector object.

[0020] S104: Based on the target expectation vector object, construct the expectation dimension vector for multiple preset expectation targets to obtain the expectation dimension set object.

[0021] Based on the target expectation vector object, and combined with the expected target set predefined for different business scenarios, a multi-dimensional expectation space and its vectorized representation are constructed: In the financial insurance scenario, the expected targets may include at least the speed of claims processing and payment, the completeness of coverage, and the matching degree of premium discounts and value-added services; in the medical scenario, the expected targets may include at least the accuracy of diagnostic conclusions, the safety and conservatism of treatment plans, the cost of examinations and medications, and the waiting time for medical treatment; a corresponding expected dimension vector is generated for each type of expected target, and these are organized into an expected dimension set object to cover the core concerns of users in both financial and medical scenarios.

[0022] S105: Use the expected dimension set object to perform weighted processing and normalization on the target expected vector to obtain the expected representation object.

[0023] The target expected vector object is weighted and normalized using the aforementioned expected dimension set object: On the one hand, according to the above... The calculated attention weights are used to decompose and recombine the target expectation vector across various expectation dimensions, resulting in a weighted expectation vector that reflects the comprehensive preferences of multidimensional financial and medical needs. On the other hand, normalization operations constrain the vector norm and numerical range, ensuring that the representations obtained by different users and in different business scenarios are comparable and stable. Ultimately, an expectation representation object is generated for subsequent satisfaction evolution modeling and preference optimization calculations, providing a unified and computable user expectation representation foundation for the entire interaction optimization process based on expectation confirmation.

[0024] S20: Use the expected representation object to perform user satisfaction evolution modeling to obtain a satisfaction evolution object.

[0025] Using the aforementioned expectation representation object, combined with the evaluation results of each round of model responses in terms of semantic relevance, information completeness, and compliance, and input into a preset satisfaction evolution model, the user's expected value and actual experience value in the multi-round interaction process are dynamically correlated and modeled to obtain a satisfaction evolution object used to characterize the satisfaction level in each round and its changing trend with the dialogue process.

[0026] Combination Figure 3As shown, step S20 specifically includes the following steps: S201: Based on the expected representation object, the model responses of each round and multiple quality indicators are aggregated to obtain the multi-round dialogue evaluation input object.

[0027] Based on the obtained expectation representation, the quality of the model's responses in each round of the multi-turn dialogue process is evaluated. For financial insurance scenarios, quality indicators may include whether the answer covers the claim conditions and exclusions, whether the explanation of rates and coverage is complete, and whether the comparison recommendations are clear and compliant. For medical consultation scenarios, quality indicators may include whether the symptom analysis is accurate, whether the examination and medication recommendations conform to guidelines, and whether the risk warnings are sufficient. The expectation representations of each round of dialogue, the model output text, and the above-mentioned multiple quality indicators are aggregated in chronological order to obtain the multi-turn dialogue evaluation input object used to characterize the overall state of the multi-turn interaction.

[0028] S202: Calculate the actual experience values ​​for each round based on the multi-round dialogue evaluation input object, and pair the actual experience values ​​with the preset expected values ​​to obtain the satisfaction modeling input object.

[0029] Based on the multi-round dialogue evaluation input, the actual experience value and the preset expected value are calculated for each round of dialogue. The actual experience value quantifies the model's objective performance in that round and can be obtained by weighting indicators such as semantic matching degree, information completeness, compliance score, emotional friendliness, and service timeliness. The preset expected value is obtained by mapping the weights of the expected representation object on each expected dimension. For example, in a claims consultation scenario, users may value the speed of payment and the explanation of claims conditions more, while in a medical scenario, they may value the accuracy of diagnosis and the sufficiency of risk explanation more. The actual experience value for each round is... With the corresponding preset expected value Paired according to rounds, they form the input objects for satisfaction modeling, which are then used in the satisfaction evolution calculation stage.

[0030] S203: Based on the satisfaction modeling input object, add the first dynamic weight, the second dynamic weight, and the satisfaction value from the previous round to obtain a weighted satisfaction calculation object.

[0031] Based on the input objects for the satisfaction modeling, a first dynamic weight, a second dynamic weight, and the satisfaction value from the previous round are integrated: wherein, the first dynamic weight... The second dynamic weight is used to control the sensitivity of the difference between the actual experience and expectations in the current round to the satisfaction level in this round. Used to characterize the inertial influence of historical satisfaction levels on current subjective feelings; previous round satisfaction scores. This stems from the evolutionary results of the previous round. These three factors are then compared with the actual experience values ​​of the current round. and expected value The common combinations are used as inputs to construct a weighted satisfaction calculation object, whose core calculation relationship can be represented as follows: in, Indicates the first t Satisfaction rating of turn-based conversations; Indicates the first t The actual experience value of the round is used to comprehensively reflect the performance of the model response in terms of semantic quality, compliance, information completeness, and service friendliness. Indicates the first t The expected value of the round is obtained by mapping the weights of the expected representation object on different business dimensions (such as claims processing time, premium discounts, medical safety, and the acceptability of examination fees). The first dynamic weight is used to adjust for instantaneous differences. The strength of the impact on satisfaction can be adaptively adjusted based on user type, business scenario (finance or healthcare), and current session stage; The second dynamic weight is used to control for historical satisfaction levels. The inertial contribution to the current subjective feeling, to reflect the user's cumulative experience in multiple rounds of interaction; The activation function, preferably the sigmoid function, is used to compress the linearly weighted result to a preset range, so that... It falls between [0,1], which facilitates normalization comparison and subsequent judgment.

[0032] S204: Input the weighted satisfaction calculation object into the activation function for processing, and output a satisfaction sequence object.

[0033] The weighted satisfaction calculation object is input into the activation function processing unit, and the satisfaction sequence object that changes with the dialogue round is calculated in turn. The satisfaction sequence object records the satisfaction value and its time order in each round, which is used to reflect the dynamic evolution trend of the user's subjective feelings in scenarios such as financial consultation, claims processing, insurance plan recommendation, medical consultation, and examination interpretation.

[0034] S205: Based on the satisfaction sequence object, when the satisfaction value of the current round is lower than the preset expected value, mark this round as an unsatisfactory round and generate a trajectory to obtain a satisfaction evolution object.

[0035] Based on the satisfaction sequence object, a threshold detection is performed on the satisfaction value for each round: current round satisfaction... Lower than the expected value Or a reference standard obtained by pre-setting a desired threshold mapping, especially in When the expectation is not confirmed, the round is marked as a potential unsatisfactory round. The corresponding round index, the difference between expectation and actual experience, the business scenario (finance or medical) and the relevant dialogue context are recorded in the satisfaction evolution object to obtain a continuous satisfaction evolution trajectory, which provides a traceable quantitative basis for subsequent cause identification and preference correction.

[0036] S30: Identify and correct dialogue turns in the satisfaction evolution object that do not meet the preset expected conditions to obtain preference correction objects.

[0037] Based on the satisfaction evolution object, dialogue rounds that do not meet the preset expectation conditions are selected. The historical dialogue content, generation strategy features and context information corresponding to these rounds are collected and analyzed to obtain preference correction objects that reflect the deviation of the current strategy from the user's preference direction.

[0038] Combination Figure 4 As shown, step S30 specifically includes the following steps: S301: Based on the satisfaction evolution object, filter the rounds where the actual experience is lower than the preset expected threshold, and extract the corresponding historical dialogue fragments to obtain the historical dialogue window data object.

[0039] Based on the aforementioned satisfaction evolution object, the actual experience values ​​of each round in the multi-round dialogue are compared with the preset expected threshold round by round. When it is detected that the actual experience of a certain round is lower than the preset expected threshold, especially when the satisfaction trend is declining for several consecutive rounds, the round is marked as a potentially unsatisfactory round. At the same time, historical dialogue fragments covering the context of the round and several rounds before and after are extracted from the dialogue log. The content of the user's insurance consultation, clause interpretation, and claim progress inquiry in the financial insurance scenario, or symptom inquiry, examination result interpretation, and medication follow-up in the medical scenario, are organized in chronological order to construct a historical dialogue window data object, providing a complete context for subsequent cause analysis.

[0040] S302: Perform information sentiment logic annotation on the historical dialogue window data object to obtain an annotation description object.

[0041] The historical dialogue window data objects are input into a preset large language model, and a reflective analysis of the system responses within the window is performed in conjunction with predefined instruction templates. The model automatically extracts potential problems from three dimensions: information, emotion, and logic. In terms of information, it marks whether the response lacks key elements or is insufficiently explained, such as failing to clearly inform about insurance exclusions, failing to fully explain the materials required for claims, or failing to adequately explain the risks and necessity of a certain examination. In terms of emotion, it marks whether the response is harsh in tone and lacks reassurance and empathy, such as only providing cold, impersonal clauses in claims disputes or critical illness diagnosis scenarios without offering appropriate emotional support to the customer or patient. In terms of logic, it marks whether the response contradicts previous statements, such as inconsistent statements about the maximum coverage amount in two rounds, or conflicting interpretations of the same examination results. The results of this multidimensional analysis are structured into labeled descriptive objects, where each label carries information such as the round it belongs to, the type of problem, and its severity.

[0042] S303: Map the labeled description object to a numerical vector to obtain a vector object of reasons for dissatisfaction.

[0043] The system, based on preset encoding rules, aggregates information dimension-related annotations into information-type numerical components, emotion dimension-related annotations into emotion-type numerical components, and logic dimension-related annotations into logic-type numerical components, and then concatenates them in a fixed order to obtain the first... t The vector of reasons for dissatisfaction corresponding to the wheel The information component measures information gaps in responses regarding terms and conditions, risk warnings, and solution details; the emotional component measures deviations in tone of friendliness, reassurance, and empathetic expression; and the logical component measures consistency deviations between responses and historical statements regarding monetary calculations, scope of liability, and treatment pathways. Through this mapping, the contribution of different problem dimensions to feelings of dissatisfaction can be quantified simultaneously in both financial insurance and healthcare scenarios.

[0044] S304: Based on the dissatisfaction cause vector, calculate the preference gradient through the correction function to obtain the policy correction gradient object.

[0045] The dissatisfaction reason vector object and the current round satisfaction value are input into the policy correction network. The preference gradient is calculated through the correction function to obtain the policy correction gradient object. The core modeling relationship of the policy correction network can be represented as: in, Indicates the first t The dissatisfaction vector is composed of information components. Emotional weight and logical components composition; Used to indicate the degree of inadequacy of information, such as failure to specify the claim threshold or failure to explain the necessity of the inspection; Used to indicate the degree of emotional deviation, such as a cold tone, lack of comfort and encouragement, etc. Used to indicate the degree of logical conflict, such as inconsistencies in the interpretation of the same insurance liability or the same test result; Indicates the current input (Including the user's current problem and its historical context in financial or medical scenarios) the original policy of the interaction model for candidate outputs The probability distribution; This represents the updated policy probability distribution obtained after introducing corrections for reasons of dissatisfaction; Indicates the first t The satisfaction score is calculated by the aforementioned satisfaction evolution model, which integrates indicators such as semantic quality, compliance, information completeness, and emotional friendliness. This represents a correction function used to map the vector of reasons for dissatisfaction and its corresponding level of satisfaction to the direction and magnitude of the preference gradient. The larger the negative gradient is when the information is missing, the emotional bias is obvious, and the logical conflict is prominent. The learning rate parameter controls the step size of each policy update, preventing excessive fluctuations in model behavior under strict financial regulations or high medical safety requirements. Through this modeling relationship, the system can transform round-level dissatisfaction reasons into continuously learnable preference gradient signals without relying on manual scoring.

[0046] S305: The interaction policy output probability in the policy correction gradient object is synthesized with the preset learning rate weighted gradient to update the preference correction object.

[0047] The aforementioned strategy correction gradient object is weighted and synthesized with the output probability of the current interaction strategy to fine-tune the generation probability of different behavioral patterns and dialogue templates, thereby updating the preference correction object. In financial insurance scenarios, the preference correction object can increase the weight of answer patterns that employ step-by-step explanations of terms, provide case examples, and offer timeline breakdowns in complex claims issues, while reducing the weight of simple, generalized answers or vague statements. In medical scenarios, the preference correction object can enhance the probability of providing sufficient risk warnings, recommending follow-up visits or referrals, and suggesting supplementary examinations in serious or high-risk treatment scenarios. The preference correction object records the increase or decrease in strategy for different scenarios and question types in a structured form, and will serve as an important input for subsequent expectation confirmation feedback and preference optimization execution stages, achieving continuous preference alignment based on round-level feedback signals without relying on manual intervention for each item.

[0048] S40: Input the preference correction object into a preset user simulation model for expected confirmation feedback processing to obtain the expected confirmation feedback object.

[0049] The preference correction object is input into a preset user simulation model. Under the conditions of combining the target user profile and business scenario constraints, simulated interactive feedback is generated. The expectation satisfaction under different strategy adjustments is evaluated, and an expectation confirmation feedback object is output to comprehensively reflect the expectation confirmation result.

[0050] Combination Figure 5 As shown, step S40 specifically includes the following steps: S401: The preference correction object is encoded and aggregated to obtain a user simulated input object.

[0051] First, the obtained preference correction objects are encoded and aggregated. The historical dissatisfaction round markers, three types of dissatisfaction reason vectors (information, emotion, and logic), satisfaction values ​​for the corresponding rounds, and user profile features are packaged together to obtain the user simulation input object. The user profile features may include risk preference, insurance experience, family structure, income range, and current insurance policy combination in the financial insurance scenario, and may include age, gender, past medical history, chronic disease type, frequency of visits, and whether the patient is pre- or post-operative in the medical scenario. This provides sufficient business context and description of population differences for subsequent simulation feedback.

[0052] S402: Input the user simulation input object into the preset user simulation model to generate a simulation feedback object.

[0053] The system response content and user profile from the simulated user input are used as conditional inputs and fed into a pre-set user simulator model based on a large language model. The model learns the subjective feedback that different groups of people might give when receiving a certain type of answer during the offline training phase, and outputs simulated feedback objects. These simulated feedback objects include various structured and unstructured information such as satisfaction / dissatisfaction labels, textual descriptions of reasons, and expected trends. The conditional generation relationship can be represented as follows: in, This represents the feedback variables output by the user simulator, used to comprehensively represent results such as satisfaction / dissatisfaction, cause description, and increase / decrease in expectations; This indicates the content of the response given by the system in a certain round, which may correspond to insurance product recommendation scripts, explanations of claims progress, interpretations of medical examination results, or explanations of medication plans, etc. U This represents the feature vector of a user profile, such as the risk preferences and insurance experience of insurance customers, and the age and disease stage of medical patients; This represents a user simulator based on a large language model, used to output the feedback distribution given system responses and user profiles. This indicates that feedback will be received based on the current answer and profile. F The conditional probability is used to characterize the likelihood of different feedback types occurring; Indicates the first t The expected values ​​of users during the round of interaction can be obtained by comprehensively considering dimensions such as claims processing timeliness, coverage scope, premium pressure, medical safety, and the acceptability of examination costs. This represents the expected value for the next round after adjustments based on feedback from this round. Indicates the first t The actual experience value of the round is obtained by weighting the model's response based on indicators such as semantic quality, compliance, information completeness, sufficiency of risk warning, and emotional friendliness; This represents the expected update coefficient, used to control the impact of the current round's experience deviation on the expected update magnitude.

[0054] S403: Based on the simulated feedback object, calculate and normalize the conditional probabilities corresponding to various simulated feedbacks according to the preset generation mechanism to obtain the expected confirmation probability object.

[0055] The simulated feedback objects are analyzed, and the frequency of occurrence of different feedback categories (such as complete satisfaction, partial satisfaction but insufficient information, dissatisfaction due to emotional bias, dissatisfaction due to logical contradiction, etc.) is normalized into a conditional probability distribution. Based on this, an expected confirmation probability object is constructed to quantify the confidence level of the system in meeting expectations under a given answer and user profile.

[0056] S404: Calculate the expected update object by combining the expected confirmation probability with the difference between the expected and actual experience.

[0057] Use the expected value to confirm the probability object, and compare it with the expected value of the current round. and actual experience value The differences between them all affect the expected update formula mentioned above. When the confirmation deviation is small and the expected update amount is small, the expected update amount will be calculated accordingly. When the value is below the preset threshold, the strategy is considered to basically meet the expectations for this type of user group; conversely, when the confirmation deviation is large or the expectation drops significantly, the corresponding sample is marked as a key round that needs to enter the preference optimization, and the updated expected state is encapsulated as an expected update object.

[0058] S405: Based on the expected update object, trigger optimization and combine it with the user profile scenario to generate an expected confirmation feedback object.

[0059] The subsequent strategy optimization process is triggered based on the expected update object, and expected confirmation feedback objects are generated in combination with different user profile scenarios. For example, for insurance customers who prefer stable returns and are sensitive to claims services, the expected confirmation feedback object will highlight preference signals such as a more detailed explanation of the claims process and displaying historical claim cases; for chronic disease users with long-term follow-up, the expected confirmation feedback object will highlight preference signals such as medication safety, long-term risk management, and follow-up schedule. The expected confirmation feedback object is output to the preference optimization module in a structured form, serving as an important basis for the construction of subsequent rounds of secondary rewards and the updating of strategy parameters. This enables continuous adaptive alignment of the psychological expectations of different user groups in both financial and medical scenarios without the need for large-scale manual annotation.

[0060] S50: Perform preference optimization processing on the expected confirmation feedback object to update the interaction model parameters and generate the corresponding interaction result object.

[0061] Based on the expected confirmation feedback object, the strategy parameters and generation configuration of the interaction model are optimized and updated in a preference-oriented manner, so that the model gradually converges in the direction of meeting the user's long-term expectations in subsequent interactions. Based on the updated interaction model, an interaction result object corresponding to the current business request is generated, thereby completing a multi-round interaction optimization process based on expected confirmation.

[0062] Combination Figure 6 As shown, step S50 specifically includes the following steps: S501: The expected confirmation feedback object is aggregated with preset simulated interaction data to calculate the round reward and obtain the round reward object.

[0063] The aforementioned expected feedback objects are aggregated with a pre-built simulated interaction dataset. In the financial insurance scenario, the simulated interaction data can cover typical dialogues such as auto insurance claims consultation, critical illness insurance product configuration, policy cancellation and reduced premium payment. In the medical scenario, it can cover typical dialogues such as initial consultation inquiry, interpretation of test results, medication follow-up, and rehabilitation guidance. For each round of dialogue, the satisfaction value, expected update amount, and feedback level from the LLM user simulator are associated, mapping satisfaction / dissatisfaction and the magnitude of expected changes into round-level reward signals, thereby calculating the round-level reward object reflecting the effectiveness of the strategy in that round.

[0064] S502: Use a preset preference gradient to train an interaction model to update the reward object for the round, and obtain an offline optimization strategy parameter object.

[0065] Based on the reward objects in each round, the interaction model is optimized offline through a pre-defined preference gradient training mechanism. On one hand, proximal policy optimization or direct preference gradient methods are used, treating the reward of each round as a round-level preference signal. This strengthens the answer patterns and information organization methods adopted in high-reward rounds and weakens information gaps, emotional biases, or logical conflicts in low-reward rounds, resulting in a set of optimized policy parameters trained offline, which is then encapsulated as an offline optimized policy parameter object. On the other hand, during training, the reward function is explicitly constructed using changes in satisfaction and expectation, and the learning rate is dynamically scheduled. Its unified form can be expressed as: in, Indicates the first t The reward value of a round is used to comprehensively reflect the quality of that round in meeting user expectations, and can be derived from the satisfaction score of that round. Compared with expected change The functions constitute; Indicates the first t The satisfaction level is calculated by the satisfaction evolution model based on indicators such as semantic quality, compliance, information completeness, risk warning adequacy, and emotional friendliness. Indicates the relationship with the first t The expected change in the wheel association can be, for example, determined by the aforementioned expected update relationship. or The derivation yields a method used to characterize the direction and intensity of the impact of the current round's actual experience on the future expected baseline; Indicates the first t The dynamic learning rate used when updating round parameters; This represents the initial learning rate, used to limit the maximum update step size; This represents the learning rate adjustment coefficient, used to control the sensitivity of satisfaction fluctuations to learning rate decay. This represents the absolute value of the difference between the current round and the previous round of satisfaction. When satisfaction fluctuates significantly, It automatically decreases, thus ensuring smoother parameter convergence during periods when the strategy's effectiveness is unstable.

[0066] S503: Adjust the interaction model parameters of each round of dialogue strategy using the offline optimization strategy parameter object to obtain an online strategy adjustment object.

[0067] The offline optimization strategy parameter object is loaded into the online interaction model, and the parameters of each round of dialogue strategy are adjusted to obtain the online strategy adjustment object: In the financial insurance scenario, the strategy weight of adopting step-by-step clause decomposition, exemplified claims process explanation, and scenario-based solution comparison is increased in situations such as complex claims disputes and explanations of long-term insurance liabilities; In the medical scenario, the strategy weight of providing more detailed risk explanations, follow-up examination suggestions, and multidisciplinary consultation prompts is enhanced when there is uncertainty in the test results or high medication risks, so as to be closer to the preferences of the target population in the actual service process.

[0068] S504: Calculate the dynamic learning rate of the object according to the change in satisfaction based on the online strategy, and obtain the learning rate update object.

[0069] Based on the real-time changes in satisfaction collected during the online phase, and combined with the online strategy adjustment targets, the learning rate is dynamically calculated and a learning rate update target is formed: when a large fluctuation in the satisfaction curve of a certain user group (such as high-net-worth insurance clients or critical illness follow-up users) is detected, the above-mentioned... The formula reduces the global learning rate, but can set relatively high local update weights for subnetworks or policy heads related to the group, thereby achieving an adaptive learning mechanism that is stable overall and accelerates key areas. Conversely, when long-term satisfaction remains stable and expected deviation is small, the learning rate is gradually decayed, causing the model policy to converge.

[0070] S505: Update the hierarchical parameters of the learning rate update object and generate the corresponding interaction result object.

[0071] The interaction model is updated hierarchically based on the learning rate update object: the parameters of the bottom general language understanding and reasoning remain relatively stable, while the middle layer financial and medical domain adaptation module and the top layer dialogue strategy and speech control module are differentiated and iterated according to the learning rate of different levels. This achieves multi-level preference optimization that combines in-round optimization (local correction of unsatisfactory content in the current round), inter-round optimization (adjusting the overall dialogue strategy according to the satisfaction trend), and user group optimization (forming preference templates for people with different risk preferences and different disease types based on clustering results). Finally, a new interaction result object is generated under the updated parameter configuration, so that the system can continuously approach the long-term expectations of different user groups in both financial and medical scenarios.

[0072] Combination Figure 7 As shown, the interaction optimization method based on expected confirmation further includes the following steps: S1: Perform semantic parsing on the interaction result object to extract and generate text data, and construct the text object to be detected.

[0073] Semantic parsing is performed on the aforementioned interaction result objects to extract the text content generated by the interaction model in this round and construct the text object to be detected. Specifically, the natural language responses to be displayed to insurance customers or medical patients are first separated from the interaction result objects. These responses are then segmented into sentences and paragraphs and structured by combining intent tags and business scenario tags (such as life insurance claims consultation, car insurance recommendation, examination result interpretation, long-term medication follow-up, etc.). This results in a text object to be detected containing original text fragments, location indexes, and business context, providing an input basis for subsequent risk word identification and compliance judgment.

[0074] S2: Perform risk word identification on the text object to be detected to obtain the text annotation object.

[0075] A database of risk words and high-risk expressions is pre-built for both the insurance and healthcare sectors. In the insurance context, this includes prohibited or sensitive expressions such as guaranteed returns, absolute safety, sure profits, and 100% reimbursement. In the healthcare context, it includes expressions that easily violate medical advertising and medication safety regulations, such as radical cure, 100% cure rate, and no risk whatsoever. Then, based on a combination of dictionary matching and contextual semantic recognition, each segment of the text to be tested is scanned and labeled, recording the location, category, and intensity level of each risk word or high-risk expression, thereby generating text-labeled objects with risk tags and location markers.

[0076] S3: Input the text annotation object into the preset compliance model to calculate the security probability and obtain the compliance confidence object.

[0077] The text-annotated object and its corresponding original text are input into a preset compliance model to quantitatively evaluate the compliance of the current response and calculate the compliance confidence level. The compliance model can be a contrastive classification model. Using regulatory rules and industry standards as conditions, a probabilistic determination is made as to whether the generated text is safe. The core calculation relationship can be expressed as follows: in, Indicates compliance confidence level, used to quantify the current level. t The probability that a loop will be deemed a safe expression under regulatory and industry standards; This represents the probability of the security category output by the compliance model; Indicates the first t The generated text data extracted from the interaction result object includes explanatory text about insurance product benefits, coverage responsibilities, and claims process, or explanatory text about examination results, treatment plans, and medication risks. This represents a set of coded compliance rules, including at least regulatory provisions and sales script guidelines in the insurance sector, and advertising guidelines, treatment recommendations, and medication advice standards in the medical field. The model outputs its results after comprehensively considering risk term annotation, contextual semantics, and business scenario tags. This is used to construct compliance confidence objects for subsequent threshold comparison and control.

[0078] S4: Compare the compliance confidence object with a preset confidence threshold. When the compliance confidence object is lower than the preset confidence threshold, generate a regeneration control instruction to obtain a compliance control object.

[0079] when When the generated text is greater than or equal to a preset reliability threshold, it is considered acceptable from a compliance perspective, requiring no mandatory rewriting and only minor polishing when necessary; when... When the text falls below the threshold, or when certain absolutely prohibited high-risk expressions appear in the text annotation object, the regeneration control logic is triggered to generate a compliance control object. This object clearly records the round identifier to be processed, the triggering reason (such as the presence of prohibited words, excessive promises in the benefit description, excessive affirmation of medical effects, etc.), and the recommended rewriting strategy pattern (such as deleting promise-making words, adding risk warnings, weakening the efficacy description, etc.) so as to drive the subsequent generation process into the compliance correction or regeneration mode.

[0080] S5: Introduce a tone constraint vector based on the compliance control object, and use the tone constraint vector to adjust the interaction result object to obtain an optimized interaction result object.

[0081] To prevent the model from exhibiting emotional responses or overly extreme tones during multiple rounds of correction and regeneration, multiple tone prototypes are preset in the emotion generation layer, and corresponding target tone vectors are set for different business scenarios. For example, a professional + reassuring tone vector is used for insurance claim disputes, and a cautious + empathetic tone vector is used for interpreting serious illnesses. A tone constraint vector matching the current user profile and business scenario is selected and used as an additional conditional input to the decoder during text regeneration or rewriting to constrain word choice, sentence structure, and emotional intensity. This achieves the goal of maintaining stable, polite, and service-compliant language while deleting or weakening non-compliant content. The text formed after these adjustments is repackaged back into the interaction result object as the final optimized interaction result object output to insurance customers or medical users, achieving integrated processing of semantic compliance and service tone control.

[0082] It is evident that the interaction optimization method based on expectation confirmation provided in this application has significant advantages in technical effectiveness compared to existing financial insurance recommendation and question-answering systems: By introducing expectation confirmation theory and satisfaction evolution modeling into the multi-round dialogue process, it no longer relies solely on static recommendation strategies or single-round feedback signals, but combines users' historical interaction behavior, current round actual experience, and expectation deviation to dynamically track and optimize long-term interaction goals, enabling the system to continuously converge towards users' long-term preferences; by structurally modeling expectation representation, expectation differences, reasons for dissatisfaction, and strategy correction processes, a recordable satisfaction evolution trajectory and preference correction objects are formed, making the adjustment process of interaction strategies clearly interpretable and auditable, facilitating traceability analysis by financial insurance institutions in compliance review, service quality assessment, and model iteration; by constructing a basic... The user simulator based on a large language model automatically generates round-by-round feedback signals during the offline phase, replacing large-scale manually labeled preference data. This significantly reduces the manpower and time costs of preference optimization. Combined with semantic compliance detection and risk word recognition mechanisms, it performs security probability assessment and regeneration control on the generated text, thereby effectively reducing the risk of illegal expressions in key aspects such as information disclosure, benefit description, and risk warning, and improving the security and credibility of the system's output content. On this basis, in business scenarios such as intelligent recommendation of insurance products, claims consultation, and policy change guidance, this application can more accurately match the needs and psychological expectations of different customer groups, improve the overall user satisfaction in terms of consultation efficiency, service experience, and trust, and help improve user retention and interaction conversion rates, thereby enhancing the comprehensive competitiveness of the intelligent service system of financial and insurance institutions.

[0083] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0084] In one embodiment, an interaction optimization device based on expectation confirmation is provided, which corresponds one-to-one with the interaction optimization method based on expectation confirmation in the above embodiments. For example... Figure 8 As shown, the interaction optimization device based on expectation confirmation includes: a modeling processing module 100, a data evolution module 200, an identification correction module 300, an expectation confirmation module 400, and a preference optimization module 500.

[0085] Detailed descriptions of each functional module are as follows: The modeling and processing module 100 is used to acquire user historical interaction data and perform vectorized modeling processing to obtain the desired representation object; The data evolution module 200 is used to perform user satisfaction evolution modeling processing using the expected representation object to obtain a satisfaction evolution object; The identification and correction module 300 is used to identify and correct dialogue turns in the satisfaction evolution object that do not meet the preset expectation conditions, so as to obtain the preference correction object. The expectation confirmation module 400 is used to input the preference correction object into a preset user simulation model for expectation confirmation feedback processing to obtain the expectation confirmation feedback object. The preference optimization module 500 is used to perform preference optimization processing on the expected confirmation feedback object to update the interaction model parameters and generate the corresponding interaction result object.

[0086] In one embodiment, the interaction optimization device based on expected confirmation is further configured to: Semantic parsing is performed on the interaction result object to extract and generate text data, and a text object to be detected is constructed; Risk word identification is performed on the text object to be detected to obtain the text annotation object; The text annotation object is input into a preset compliance model to calculate the security probability and obtain a compliance confidence object; The compliance confidence object is compared with a preset confidence threshold. When the compliance confidence object is lower than the preset confidence threshold, a regeneration control instruction is generated to obtain a compliance control object. Based on the compliance control object, a tone constraint vector is introduced, and the interaction result object is adjusted using the tone constraint vector to obtain an optimized interaction result object.

[0087] In one embodiment, the modeling processing module 100 is specifically used for: The user input text, historical query records and interactive behaviors are collected and preprocessed to obtain the raw user interaction data object; The original user interaction data object is input into a preset semantic understanding model for embedding operation to obtain a user interaction semantic embedding object. The instantaneous intent vector and the historical expectation trend vector are extracted from the user interaction semantic embedding object and then added together to obtain the target expectation vector object; Based on the target expectation vector object, multiple preset expectation targets are constructed with expectation dimension vectors to obtain an expectation dimension set object; The expected vector is weighted and normalized using the expected dimension set object to obtain the expected representation object.

[0088] In one embodiment, the data evolution module 200 is specifically used for: Based on the expected representation object, the model responses of each round and multiple quality indicators are aggregated to obtain the multi-round dialogue evaluation input object; Based on the multi-turn dialogue evaluation input object, calculate the actual experience value for each round, and pair the actual experience value with the preset expected value to obtain the satisfaction modeling input object; Based on the satisfaction modeling input object, the first dynamic weight, the second dynamic weight, and the previous round satisfaction value are added to the integrated object to obtain the satisfaction weighted calculation object. The weighted satisfaction calculation object is input into the activation function for processing, and the output is a satisfaction sequence object; When the satisfaction value in the current round is lower than the preset expected value, based on the satisfaction sequence object, this round is marked as an unsatisfactory round and a trajectory is generated to obtain the satisfaction evolution object.

[0089] In one embodiment, the identification correction module 300 is specifically used for: Based on the satisfaction evolution object, the rounds in which the actual experience is lower than the preset expected threshold are selected, and the corresponding historical dialogue fragments are extracted to obtain the historical dialogue window data object; The historical dialogue window data object is annotated with information sentiment logic to obtain an annotated description object; The labeled description object is mapped to a numerical vector to obtain a vector object of reasons for dissatisfaction; Based on the aforementioned dissatisfaction cause vector, the preference gradient is calculated using a correction function to obtain the policy correction gradient object; The interaction policy output probability in the policy correction gradient object is synthesized with the preset learning rate weighted gradient to update the preference correction object.

[0090] In one embodiment, the expected confirmation module 400 is specifically used for: The preference correction object is encoded and aggregated to obtain a user-simulated input object; The user simulation input object is input into the preset user simulation model to generate a simulation feedback object; Based on the simulated feedback object, the conditional probabilities corresponding to various simulated feedbacks are calculated and normalized according to the preset generation mechanism to obtain the expected confirmation probability object; The expected update object is calculated by combining the expected confirmation probability with the difference between the expected and the actual experience. Based on the expected update object, optimization is triggered and combined with the user profile scenario to generate the expected confirmation feedback object.

[0091] In one embodiment, the preference optimization module 500 is specifically used for: The expected confirmation feedback object is aggregated with preset simulated interaction data to calculate the round reward and obtain the round reward object; The reward objects for each round are updated using a preset preference gradient training interaction model to obtain an offline optimization strategy parameter object; The interaction model parameters of each round of dialogue strategy are adjusted using the offline optimization strategy parameter object to obtain the online strategy adjustment object; Based on the online strategy, the dynamic learning rate of the adjusted object is calculated according to the change in satisfaction level, and the learning rate update object is obtained. The learning rate update object is updated with hierarchical parameters, and a corresponding interactive result object is generated.

[0092] Specific limitations regarding the interaction optimization device based on expectation confirmation can be found in the limitations of the interaction optimization method based on expectation confirmation above, and will not be repeated here. Each module in the aforementioned interaction optimization device based on expectation confirmation can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device in hardware form, or stored in the memory of a computer device in software form, so that the processor can call and execute the operations corresponding to each module.

[0093] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 9 As shown, the computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile and / or volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The network interface is used to communicate with external clients via a network connection. When the computer program is executed by the processor, it implements a server-side function or step based on an interactive optimization method with expected confirmation.

[0094] In one embodiment, a computer device is provided, which may be a client, and its internal structure diagram may be as follows: Figure 10 As shown, the computer device includes a processor, memory, network interface, display screen, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The network interface is used to communicate with an external server via a network connection. When the computer program is executed by the processor, it implements client-side functions or steps based on an interactive optimization method with expected confirmation.

[0095] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed, can perform the steps provided in the above embodiments.

[0096] It should be noted that the functions or steps that can be implemented by the computer-readable storage medium or computer device described above can be referred to the relevant descriptions on the server side and client side in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.

[0097] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0098] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0099] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.

Claims

1. An interaction optimization method based on expected confirmation, characterized in that, include: Obtain user historical interaction data and perform vector modeling to obtain the desired representation object; The user satisfaction evolution model is performed using the aforementioned expectation representation object to obtain a satisfaction evolution object; The dialogue turns that do not meet the preset expectation conditions in the satisfaction evolution object are identified and corrected to obtain the preference correction object; The preference correction object is input into a preset user simulation model for expected confirmation feedback processing to obtain the expected confirmation feedback object; The desired confirmation feedback object is subjected to preference optimization processing to update the interaction model parameters and generate the corresponding interaction result object.

2. The interaction optimization method based on expected confirmation according to claim 1, characterized in that, Also includes: Semantic parsing is performed on the interaction result object to extract and generate text data, and a text object to be detected is constructed; Risk word identification is performed on the text object to be detected to obtain the text annotation object; The text annotation object is input into a preset compliance model to calculate the security probability and obtain a compliance confidence object; The compliance confidence object is compared with a preset confidence threshold. When the compliance confidence object is lower than the preset confidence threshold, a regeneration control instruction is generated to obtain a compliance control object. Based on the compliance control object, a tone constraint vector is introduced, and the interaction result object is adjusted using the tone constraint vector to obtain an optimized interaction result object.

3. The interaction optimization method based on expected confirmation according to claim 1, characterized in that, The process of acquiring user historical interaction data and performing vectorized modeling to obtain the desired representation object includes: The user input text, historical query records and interactive behaviors are collected and preprocessed to obtain the raw user interaction data object; The original user interaction data object is input into a preset semantic understanding model for embedding operation to obtain a user interaction semantic embedding object. The instantaneous intent vector and the historical expectation trend vector are extracted from the user interaction semantic embedding object and then added together to obtain the target expectation vector object; Based on the target expectation vector object, multiple preset expectation targets are constructed with expectation dimension vectors to obtain an expectation dimension set object; The expected vector is weighted and normalized using the expected dimension set object to obtain the expected representation object.

4. The interaction optimization method based on expected confirmation according to claim 1, characterized in that, The process of using the expected representation object to perform user satisfaction evolution modeling to obtain a satisfaction evolution object includes: Based on the expected representation object, the model responses of each round and multiple quality indicators are aggregated to obtain the multi-round dialogue evaluation input object; Based on the multi-turn dialogue evaluation input object, calculate the actual experience value for each round, and pair the actual experience value with the preset expected value to obtain the satisfaction modeling input object; Based on the satisfaction modeling input object, the first dynamic weight, the second dynamic weight, and the previous round satisfaction value are added to the integrated object to obtain the satisfaction weighted calculation object. The weighted satisfaction calculation object is input into the activation function for processing, and the output is a satisfaction sequence object; When the satisfaction value in the current round is lower than the preset expected value, based on the satisfaction sequence object, this round is marked as an unsatisfactory round and a trajectory is generated to obtain the satisfaction evolution object.

5. The interaction optimization method based on expected confirmation according to claim 1, characterized in that, The process of identifying and correcting dialogue turns that do not meet the preset expectation conditions in the satisfaction evolution object to obtain preference correction objects includes: Based on the satisfaction evolution object, the rounds in which the actual experience is lower than the preset expected threshold are selected, and the corresponding historical dialogue fragments are extracted to obtain the historical dialogue window data object; The historical dialogue window data object is annotated with information sentiment logic to obtain an annotated description object; The labeled description object is mapped to a numerical vector to obtain a vector object of reasons for dissatisfaction; Based on the aforementioned dissatisfaction cause vector, the preference gradient is calculated using a correction function to obtain the policy correction gradient object; The interaction policy output probability in the policy correction gradient object is synthesized with the preset learning rate weighted gradient to update the preference correction object.

6. The interaction optimization method based on expected confirmation according to claim 1, characterized in that, The step of inputting the preference correction object into a preset user simulation model for expected confirmation feedback processing to obtain the expected confirmation feedback object includes: The preference correction object is encoded and aggregated to obtain a user-simulated input object; The user simulation input object is input into the preset user simulation model to generate a simulation feedback object; Based on the simulated feedback object, the conditional probabilities corresponding to various simulated feedbacks are calculated and normalized according to the preset generation mechanism to obtain the expected confirmation probability object; The expected update object is calculated by combining the expected confirmation probability with the difference between the expected and the actual experience. Based on the expected update object, optimization is triggered and combined with the user profile scenario to generate the expected confirmation feedback object.

7. The interaction optimization method based on expected confirmation according to claim 1, characterized in that, The step of performing preference optimization processing on the expected confirmation feedback object to update the interaction model parameters and generate the corresponding interaction result object includes: The expected confirmation feedback object is aggregated with preset simulated interaction data to calculate the round reward and obtain the round reward object; The reward objects for each round are updated using a preset preference gradient training interaction model to obtain an offline optimization strategy parameter object; The interaction model parameters of each round of dialogue strategy are adjusted using the offline optimization strategy parameter object to obtain the online strategy adjustment object; Based on the online strategy, the dynamic learning rate of the adjusted object is calculated according to the change in satisfaction level, and the learning rate update object is obtained. The learning rate update object is updated with hierarchical parameters, and a corresponding interactive result object is generated.

8. An interaction optimization device based on expected confirmation, characterized in that, include: The modeling and processing module is used to acquire user historical interaction data and perform vectorized modeling to obtain the desired representation object; The data evolution module is used to perform user satisfaction evolution modeling processing using the expected representation object to obtain a satisfaction evolution object; The identification and correction module is used to identify and correct dialogue turns in the satisfaction evolution object that do not meet the preset expectation conditions, so as to obtain the preference correction object. The expectation confirmation module is used to input the preference correction object into a preset user simulation model for expectation confirmation feedback processing to obtain the expectation confirmation feedback object. The preference optimization module is used to perform preference optimization processing on the expected confirmation feedback object to update the interaction model parameters and generate the corresponding interaction result object.

9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the interaction optimization method based on expected confirmation as described in any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the interaction optimization method based on expected confirmation as described in any one of claims 1 to 7.