Intelligent session policy reply method and system

By performing preliminary matching of user input statements and state judgment of the large model, and updating historical state data with timestamps, the problem of multiple states coexisting and dynamic control in existing technologies is solved, realizing structured management and lifecycle control of dialogue, and improving user experience and system security.

CN120975102APending Publication Date: 2025-11-18HI-THINK YONDERVISION (BEIJING) TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511090024.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-05
Publication Date
2025-11-18

AI Technical Summary

Technical Problem

Existing technologies cannot effectively achieve multi-state coexistence, state combination judgment, and dynamic control, and lack a lifecycle control mechanism. This leads to problems such as state ambiguity, logical incoherence, and semantic drift in user dialogues with multiple topics and state transitions across rounds, especially posing uncontrollable risks in sensitive or security-constrained scenarios.

Method used

By performing preliminary matching on the current user input to obtain request prompts, using a large model to obtain the status and response, and storing them in a historical status database, updating the data with timestamps, determining in real time whether the session has ended and destroying the historical status data, supporting status combination judgment, and updating historical status data with timestamps, the system achieves structured management and lifecycle control of the dialogue.

Benefits of technology

It enables the time-series recording and state combination update of user states, forming a structured semantic trajectory, ensuring data security, and is suitable for business scenarios with semantic jumps, complex intents, and behavioral sensitivity requirements. It improves the user experience, enhances the user experience, strengthens the structured understanding of user intent, supports multi-task recognition capabilities, avoids semantic drift, and provides dialogue tracking and context evolution management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120975102A_ABST
    Figure CN120975102A_ABST
Patent Text Reader

Abstract

The invention provides an intelligent session policy reply method and system, and the method comprises the steps: storing a first state in a historical state database corresponding to a current session to form historical state data, and transmitting a subsequent user input statement of the current session and the latest historical state data to a large model, obtaining subsequent states and subsequent responses corresponding to the statements input by the subsequent user, and storing the subsequent states corresponding to the statements in a historical state database based on the timestamps to update the historical state data; and moreover, whether the current session is ended or not is judged in real time, and if the current session is ended, historical state data are destroyed, so that time sequence recording, state combination updating, formation of a structured semantic track, setting of the action range of state history as the current session and automatic destroying along with session termination or deletion are performed on the user state, and the data security is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of artificial intelligence and natural language processing technology, and designs a strategy reply method based on a large language model, more specifically, relates to an intelligent conversation strategy reply method and system. BACKGROUND

[0002] The large language model has powerful general language processing capability, but its original form is often a "input-output" closed loop structure, which only responds to the current input, and lacks structured management of long historical states. This mechanism will produce problems such as state ambiguity, incoherent logic, and semantic drift when facing real user conversations with cross-round multi-topic and multi-state jumps, especially in sensitive or security constraint scenarios, and its uncontrollability will amplify potential risks.

[0003] Traditional dialogue systems mostly use rule-based state machines or NLU+state machine solutions (such as BERT+Rasa), which can achieve limited intent recognition and flow control, but do not support dynamic multi-state combination judgment, lack fine-grained context semantic tracking and flexible control capabilities, and cannot meet the diversified and personalized public accumulation fund business and other financial needs. Some systems try to solve this problem by simply concatenating contexts, or introduce rule-based dialogue trees (such as Rasa, DialogFlow, etc.) to realize dialogue state tracking, but these systems still cannot truly realize multi-state coexistence, state combination judgment, dynamic control, and do not have strong reasoning capabilities, and lack a life cycle control mechanism to effectively isolate and clean up user data.

[0004] Therefore, an intelligent conversation strategy reply method and system that can be applied to business scenarios with semantic jumps, intent combinations, and behavior sensitivity requirements, integrating state classification, history tracking, semantic control, and life cycle management is urgently needed. SUMMARY

[0005] In view of the above problems, the purpose of the present application is to provide an intelligent conversation strategy reply method and system to solve the technical problems that the prior art cannot truly realize multi-state coexistence, state combination judgment, dynamic control, and does not have strong reasoning capabilities, and lacks a life cycle control mechanism to effectively isolate and clean up user data.

[0006] The intelligent conversation strategy reply method provided by the present application comprises:

[0007] performing preliminary matching on the initial user input sentence of the current conversation to obtain a demand prompt word;

[0008] sending the demand prompt word and the user input sentence to a preset large model to obtain a first state and a first response corresponding to the first state;

[0009] receiving the first state and the first response, storing the first state in a history state database corresponding to the current session to form history state data;

[0010] sending a subsequent user input sentence of the current session and the latest history state data to the large model to obtain a subsequent state and a subsequent response corresponding to each sentence of the subsequent user input sentence, and storing each subsequent state corresponding to each sentence in the history state database based on a time stamp to update the history state data; and determining in real time whether the current session is ended, and if the current session is ended, destroying the history state data.

[0011] Optionally, the large model obtains the first state, wherein the first state comprises:

[0012] performing classification and division according to the demand prompt word through a preset multi-label classifier to determine a main intent corresponding to the demand prompt word;

[0013] performing keyword extraction according to the user input sentence to obtain a first keyword, and matching a first operation type corresponding to the user input sentence according to the first keyword;

[0014] determining the first state according to the main intent and the first operation type.

[0015] Optionally, the large model obtains a subsequent state corresponding to each sentence of the subsequent user input sentence, wherein the subsequent state comprises:

[0016] performing keyword extraction on a single sentence of the subsequent user input sentence to obtain a subsequent keyword, and matching a subsequent operation type and an auxiliary intent corresponding to the user input sentence according to the subsequent keyword;

[0017] determining the subsequent state according to the subsequent operation type, the auxiliary intent, and the main intent.

[0018] Optionally, the first state or the subsequent state comprises a question, an indication, a viewpoint, a fact, a sensitivity, a risk, and a repetition.

[0019] Optionally, the first response or the subsequent response is obtained, comprising:

[0020] determining the first response according to the first state;

[0021] determining the subsequent response according to the subsequent state;

[0022] wherein,

[0023] If the first state or the subsequent state is a question, the first response or the subsequent response is: according to the first operation type or the subsequent operation type, calling a target document in a preset database to send to the current session, or based on the target document, extracting key information to form a policy reply, and sending the policy reply to the current session;

[0024] If the first state or the subsequent state is an indication, the first response or the subsequent response is: according to the first operation type or the subsequent operation type, changing the business state corresponding to the indication;

[0025] If the first state or the subsequent state is a viewpoint or a fact, the first response or the subsequent response is: selecting a reply sentence corresponding to the main intention or the auxiliary intention in the database, and sending the reply sentence to the current session;

[0026] If the first state or the subsequent state is sensitive, the first response or the subsequent response is: refusing to answer the question, sending a sensitive prompt to the current session, and waiting for the next subsequent user input sentence;

[0027] If the first state or the subsequent state is a risk, the first response or the subsequent response is: sending a risk prompt to the current session.

[0028] Optionally, the subsequent state corresponding to each sentence is stored in the historical state database based on a timestamp to update the historical state data, including:

[0029] According to the chronological order of the timestamp, the state data of the subsequent state corresponding to each sentence is arranged in sequence to form a structured state data list, and the newly added subsequent state data is spliced in the form of a new table at the tail of the structured state data list to update the historical state data.

[0030] The application also provides a kind of, intelligent conversation policy reply system, including:

[0031] The prologue recognition module is used to preliminarily match the initial user input sentence of the current session to obtain a demand prompt word;

[0032] The large model is used to obtain a first state according to the demand prompt word, and a first response corresponding to the first state; it is also used to obtain a subsequent state and a subsequent response corresponding to each sentence of the subsequent user input sentence according to the subsequent user input sentence of the current session and the latest historical state data;

[0033] The historical state module is configured to receive the first state and the first response, store the first state in a historical state database corresponding to a current session to form historical state data, store a subsequent state corresponding to each sentence based on a time stamp in the historical state database to update the historical state data, and determine whether the current session is ended in real time, and if the current session is ended, destroy the historical state data.

[0034] Optionally, the large model obtains the first state, including:

[0035] The multi-label classifier is preset to classify and divide the demand prompt word to determine a main intent corresponding to the demand prompt word.

[0036] The keyword extraction is performed on the user input sentence to obtain a first keyword, and a first operation type corresponding to the user input sentence is matched according to the first keyword.

[0037] The first state is determined according to the main intent and the first operation type.

[0038] Optionally, the large model obtains a subsequent state corresponding to each sentence of the subsequent user input sentence, including:

[0039] The keyword extraction is performed on a single sentence of the subsequent user input sentence to obtain a subsequent keyword, and a subsequent operation type and an auxiliary intent corresponding to the user input sentence are matched according to the subsequent keyword.

[0040] The subsequent state is determined according to the subsequent operation type, the auxiliary intent, and the main intent.

[0041] Optionally, the historical state module sequentially arranges state data of the subsequent state corresponding to each sentence according to a time stamp to form a structured state data list, and splices the newly added subsequent state data in the form of a new table at the tail of the structured state data list to update the historical state data.

[0042] From the above technical solutions, the intelligent conversation strategy reply method and system provided by the application first performs preliminary matching on the initial user input sentence of the current conversation to obtain a demand prompt word, then sends the demand prompt word and the user input sentence to a preset large model to obtain a first state and a first response, then stores the first state in a historical state database corresponding to the current conversation to form historical state data, and then sends the subsequent user input sentence of the current conversation and the latest historical state data to the large model to obtain subsequent states and subsequent responses corresponding to each sentence of the subsequent user input sentence, and stores each sentence corresponding subsequent state in the historical state database based on a timestamp to update the historical state data; and, the current conversation is judged in real time whether it is ended, if the current conversation is ended, the historical state data is destroyed, so that the prompt word is used to let the large model understand the state classification rule, and each conversation is classified and attributed, the user state is recorded in time sequence and updated in state combination, the structured semantic track is formed, the action range of the state history is set as the current conversation, and the state history is automatically destroyed with the termination or deletion of the conversation, the data safety is ensured, and an intelligent question and answer scheme integrated with state classification, historical tracking, semantic control and life cycle management is constructed. BRIEF DESCRIPTION OF DRAWINGS

[0043] Other objects and results of the present application will become more apparent and easily understood with reference to the following description of the application taken in conjunction with the accompanying drawings, in which:

[0044] Figure 1 A flow chart of the intelligent conversation strategy reply method according to the embodiment of the application;

[0045] Figure 2 A core control path diagram of the intelligent conversation strategy reply method according to the embodiment of the application;

[0046] Figure 3 A system block diagram of the intelligent conversation strategy reply system according to the embodiment of the application. DETAILED DESCRIPTION

[0047] The closest prior art is a dialogue system combined with "NLU (natural language understanding model) + state machine", for example, using a BERT classifier to cooperate with Rasa state logic to manage intent. However, the state determination of such a system usually depends on single sentence input, and the state transition is based on a preset path, and does not have the ability of real-time historical integration and combined state judgment.

[0048] In view of the above problems, the present application provides an intelligent conversation strategy reply method and system, which will be described in detail below in combination with the accompanying drawings.

[0049] To illustrate the intelligent conversation strategy response method and system provided by this invention, Figures 1-3 The embodiments of the present invention are illustrated by way of example.

[0050] The following description of exemplary embodiments is merely illustrative and is in no way intended to limit the invention or its application or use. Techniques and equipment known to those skilled in the art may not be discussed in detail, but where appropriate, such techniques and equipment should be considered part of the specification.

[0051] like Figure 1 , Figure 2 As shown in the embodiments of the present invention, the intelligent conversation strategy response method mainly includes:

[0052] S1: Perform preliminary matching on the initial user input statement of the current session to obtain the required prompt words;

[0053] S2: Send the demand prompt and the user input statement to a preset large model to obtain a first state and a first response corresponding to the first state;

[0054] S3: Receive the first state and the first response, and store the first state in the historical state database corresponding to the current session to form historical state data;

[0055] S4: Send the subsequent user input statements of the current session and the latest historical state data to the large model to obtain the subsequent states and responses corresponding to each statement of the subsequent user input statements, and store the subsequent states corresponding to each statement in the historical state database based on the timestamp to update the historical state data; and determine in real time whether the current session has ended. If the current session has ended, destroy the historical state data.

[0056] In this embodiment, step S1 is the process of performing preliminary matching on the initial user input statement of the current session to obtain the demand prompt words. The execution subject of this process is the pre-recognition module, that is, the pre-set pre-recognition module performs preliminary matching on the initial user input statement of the current session to obtain the demand prompt words.

[0057] In the embodiment, the pre-sequencing recognition module, also referred to as an LLM calling guide module, is the core of constructing the demand prompt PromptX, which is used to construct system prompts, inject field classification rules, tone, behavior restrictions and other information, guide the model to work in a conventional manner, and greatly improve the efficiency and accuracy of the model in answering professional problems in the field. In a specific embodiment, the demand prompt PromptX adopts a modular construction method, that is, through system instructions (System) + state context + initial user input sentence acquisition. The PromptX content used in the training of the LLM calling guide module in the early stage is jointly written by engineering and business personnel to improve the deployment professionalism, accuracy and specialization. The LLM calling guide module can adapt to multiple models: DeepSeekV3, DeepSeekR1, Qwen, ChatGLM, etc. are all adaptable, and have strong adaptability.

[0058] In the embodiment, step S2 is the process of sending the demand prompt and the user input sentence to a preset large model to obtain a first state and a first response corresponding to the first state. In this process, the first state is obtained by the large model, which includes:

[0059] SA211: classifying and dividing according to the demand prompt through a preset multi-label classifier to determine the main intent corresponding to the demand prompt;

[0060] SA212: extracting the first keyword according to the user input sentence to obtain the first keyword, and matching the first operation type corresponding to the user input sentence according to the first keyword;

[0061] SA213: determining the first state according to the main intent and the first operation type.

[0062] Step S3 is the process of receiving the first state and the first response, storing the first state in the historical state database corresponding to the current session to form the historical state data, and the execution subject is a preset historical state module.

[0063] Step S4 is the process of repeatedly responding to the user's input in the current session, that is, sending the subsequent user input sentence of the current session and the latest historical state data to the large model to obtain the subsequent state and the subsequent response corresponding to each sentence of the subsequent user input sentence, and storing each sentence corresponding to the subsequent state in the historical state database based on the timestamp to update the historical state data; and, judging in real time whether the current session is ended, if the current session is ended, destroying the historical state data.

[0064] How to judge whether the current session is ended is not limited here. The current session can be ended if no new statement is sent in a preset time interval after the last statement in the subsequent user input statement, or the keywords such as "received" and "completed" are captured in the subsequent user input statement.

[0065] In the embodiment, the steps in S2 and part of the steps in S4 are specifically executed by the large model, and the steps in S3 and part of the steps in S4 are executed by the historical state module.

[0066] In the embodiment, the large model obtains the subsequent state corresponding to each statement of the subsequent user input statement, which includes:

[0067] SA411: keyword extraction is performed on a single statement of the subsequent user input statement to obtain a subsequent keyword, and a subsequent operation type and an auxiliary intent corresponding to the user input statement are matched according to the subsequent keyword;

[0068] SA412: the subsequent state is determined according to the subsequent operation type, the auxiliary intent, and the main intent.

[0069] The first state or the subsequent state includes questioning, indication, viewpoint, fact, sensitivity, risk, repetition, and manual operation.

[0070] The first response or the subsequent response is obtained, including:

[0071] SA214: the first response is determined according to the first state;

[0072] SA413: the subsequent response is determined according to the subsequent state;

[0073] Wherein,

[0074] If the first state or the subsequent state is questioning, the first response or the subsequent response is that a target document is called in a preset database according to the first operation type or the subsequent operation type and sent to the current session, or a strategy reply is formed by key information extraction based on the target document and sent to the current session.

[0075] If the first state or the subsequent state is indication, the first response or the subsequent response is that a business state change corresponding to the indication is performed according to the first operation type or the subsequent operation type.

[0076] If the first state or the subsequent state is a viewpoint or a fact, the first response or the subsequent response is: selecting a reply sentence corresponding to the main intent or the auxiliary intent in the database, and sending the reply sentence to the current session;

[0077] If the first state or the subsequent state is sensitive, the first response or the subsequent response is: refusing to answer the question, issuing a sensitive prompt to the current session, and waiting for the next subsequent user input sentence;

[0078] If the first state or the subsequent state is risk, the first response or the subsequent response is: issuing a risk prompt to the current session.

[0079] Storing the subsequent state corresponding to each sentence in the history state database based on the timestamp to update the history state data, comprising:

[0080] SA421: arranging the state data of the subsequent state corresponding to each sentence in order according to the timestamps, forming a structured state data list, and splicing the newly added subsequent state data in the form of a new table at the tail of the structured state data list to update the history state data.

[0081] In one specific embodiment, the large model can be referred to as an LLM state recognition module, which is the core of state label determination. It injects a state label system into the LLM based on the demand prompt PromptX to realize multi-label classification judgment of user input sentences, supports state combination (Multi-Label) and cross state (Overlapping Intents). The state judgment of this large model uses a Softmax multi-label classifier (used for BERT / RoBERTa replacement model) or an LLM generated classification method, etc.

[0082] In this embodiment, the first state or the subsequent state is divided into a main state and an auxiliary state when it is obtained, which is used to represent the double-layer structure of "main intent + operation type" or "background demand + task request" to improve the recognition accuracy.

[0083] In this embodiment, the UUID and the timestamp are retained when the first state or the subsequent state is obtained, and each state record is bound with a timestamp and a UUID to ensure semantic backtracking consistency.

[0084] And as mentioned above, in this embodiment, only the first state with the main intent needs to be obtained by obtaining the demand prompt for the initial user input sentence of the current session. The first-confirmed main intent can be directly used as a reference to determine the subsequent state when the subsequent state is obtained. In this way, the overall conversation reply time can be saved, and the response speed can be improved.

[0085] In another embodiment, the subsequent user input sentence of the current session can also be preliminarily matched to obtain a subsequent demand prompt word, and the subsequent demand prompt word, the subsequent user input sentence of the current session and the latest historical state data are sent to the large model to obtain a subsequent state and a subsequent response corresponding to each sentence of the subsequent user input sentence. In this way, the subsequent state is also classified and divided by the preset multi-label classifier to determine a subsequent main intent corresponding to the subsequent demand prompt word, the subsequent keywords are extracted according to the user subsequent input sentence to obtain subsequent keywords, and the subsequent operation type corresponding to the user input sentence is matched according to the subsequent keywords.

[0086] The subsequent state is determined according to the main intent and the subsequent operation type.

[0087] In the embodiment, how to determine the first state according to the main intent and the first operation type, how to determine the subsequent state according to the main intent and the subsequent operation type, and how to determine the subsequent state according to the subsequent operation type, the auxiliary intent and the main intent are not limited herein. The score evaluation method can be used, for example, the subsequent operation type / first operation type is an operation type, the category can be divided into calling documents / artificial reply / rule interpretation, etc., and the score can account for 40%. If the sensitive word or risk word is designed, the score is 0. The main intent / auxiliary intent is an intent, the category can be divided into approval / repayment / loan / recharge / withdrawal, etc., and can account for 60%. If the main intent and the auxiliary intent exist at the same time, the main intent can account for 30%, the auxiliary intent can account for 30%, etc. If the sensitive word or risk word is involved, the score is 0, so that the comprehensive score is obtained to match the state (first state or subsequent state) corresponding to the score.

[0088] In the embodiment, the historical state module can be referred to as a state history record module, which is a semantic trajectory record core. The function thereof is used for structurally recording all historical state judgment results and supporting calling by the LLM in each round of interaction to constitute a "time axis" of user intent evolution. The storage structure of the historical state module is a combination of a chain structure (such as a JSON array) and a graph structure (such as a DAG) to facilitate subsequent auditing and visualization. The historical state module is not changeable but can be appended to ensure integrity. The historical state module can be encrypted and stored in a memory type cache (such as a PostgreSQL) or a Session DB, has a time limit automatic destruction mechanism, ensures that record information is not propagated across sessions, and can integrate a graph neural network to optimize multi-round task path prediction and processing to optimize user multi-round task path matching efficiency.

[0089] In addition, in the present embodiment, an artificial response mechanism is also included. When the first state or the subsequent state is "manual", or the manual background check finds that the user needs manual intervention, the artificial response is performed through a preset regulation system module based on the current state and historical state data and business rules, and the large model is controlled, which can include operations such as refusing to reply, interrupting, transferring, adjusting tone, etc.

[0090] The regulation system module adopts a special small model: such as bge-large-zh-v1.5, which assists in understanding and matching the intent of user input; the control rules are written as: the control rules support DSL (Domain-Specific Language) expression (such as expression: if state in [7,9] and conversation duration > 3min then refuse to answer); support for manual insertion of strategy points, including:

[0091] Risk state filtering: when the state judged by the model contains a preset risk state (such as state 9->inhibit output), the regulation system performs corresponding operations;

[0092] State conflict detection: such as "state 2 = express opinion" and "state 5 = request fact" appearing at the same time, the regulation system can determine the priority according to the context or guide the user to split;

[0093] Behavior forwarding: when the state is specific (such as state 6 = call database information), the regulation system automatically calls external modules. External plug-ins can be inserted, such as business process engines, API calling modules (retirement cancellation calling interfaces, etc.). State overflow (state surge monitoring, frequent loop warning, etc.): when the number of states exceeds a certain number and the combination is too complex to evolve, the regulation system controls the termination of the answer and prompts the user to simplify the problem;

[0094] When updating the business rules, the public accumulation fund business rule table can be loaded to realize dynamic rule updating.

[0095] And the regulation system module can:

[0096] Historical state review: summarize the state evolution path in the current conversation;

[0097] Thought organization: classify the user's multiple questions into several clues and guide them to logically comb;

[0098] Risk state prompt: if the state evolution contains repetition, contradiction, sensitive state, etc., the system issues a risk prompt and suggests termination or rewriting;

[0099] State chain identification: such as identifying the risk state chain (such as state 9->7->3) which can be interrupted or reported;

[0100] Semantic search assistant: allows users to search for keywords in the state history, refocus the issue focus.

[0101] To better illustrate this embodiment, the specific user usage scenario and process are as follows:

[0102] Step 1, use the demand prompt sample data for training, model deployment stage, domain experts and engineering personnel construct the state classification prompt PromptX that adapts to the field knowledge and the language of the large model, and clearly inform the language model of the number and brief definition of all identifiable states (and update it continuously). The prompt will be injected into the language model in the form of a system prompt, and will exist throughout the conversation, serving as a basis for judgment.

[0103] Step 2, the user issues the first natural language input (sentence 1), and the control system rewrites the input, performs the prompt engineering enabled by the professional field (such as public accumulation fund business) database, and converts it into a user demand language more suitable for the large language model. The rewritten user demand (demand 1) is combined with PromptX and delivered to the large language model.

[0104] Step 3, the large model makes a state judgment on demand 1 based on the classification criteria provided by PromptX, selects a state number (such as state 1), and generates a relevant response sentence under the premise of this state category and returns it to the user. This step effectively reduces the judgment range of the large model due to PromptX. It also reduces the reaction time of the large model and the user waiting time, enhancing the user interaction experience.

[0105] Step 4, record the current time point and the judged state number to form a structured record in the following format: "T1: state 1"

[0106] This record is bound to the current conversation and is valid throughout the conversation lifecycle. After the user actively ends the conversation or the system times out, the record is automatically deleted. At this stage, the control system can determine whether to allow the state to appear according to the artificially formulated strategy, such as setting "state 7 as a sensitive state, appearing immediately rejected", then the system immediately blocks the output and returns a rejection prompt.

[0107] Here, the control system acts as a buffer zone, sensitively and controllably identifying special states and taking action to control the use of the large model and the returned information, avoiding the emergence of sensitive or inappropriate remarks.

[0108] Step 5, the state history record appends a new structured record:

[0109] "T1: state 1; T2: state 1, 3"

[0110] Each record cannot be modified once generated and is used for subsequent control logic and review analysis.

[0111] Step 6, the user issues a new input statement 2 again, the system rewrites statement 2 as requirement 2, and transmits it into the language model together with the state history record.

[0112] For example: current input: requirement 2 + "T1: state 1".

[0113] Step 7, the large model integrates the content of requirement 2 and the historical state (T1: state 1) to rejudge the current user state. For example, the current state judgment result is "state 1, 3", which means that the user's intention contains the continuation of state 1 and the addition of new state 3.

[0114] Step 8, the regulation system module judges whether to allow answering, how to adjust the tone, and whether to interrupt the topic according to the current state and state combination, and refers to the control rules (such as state combination conflict, state repetition, risk state triggering, etc.).

[0115] Here, the regulation system continues to play a buffering control role.

[0116] In addition, if the regulation system detects that the classification judgment returned by the large model is ambiguous (such as the returned state label is not in the classification prompt PromptX), it will recombine the input PromptX.

[0117] Step 9, clear user dialogue and historical state data at the end of the conversation to ensure user information and privacy security.

[0118] More specifically, taking the simple public accumulation fund business process as an example:

[0119] Scenario one: cross-regional transfer and account freezing process (user resigns and plans to work in another city)

[0120] 1.1. The user asks: "I have resigned and plan to work in another city. What should I do with my public accumulation fund?"

[0121] State recognition: state 1 = "identity change consultation", state 2 = "cross-regional transfer intention", state 3 = "freezing judgment";

[0122] Control strategy: allow to continue, guide to understand "cross-regional transfer process" and explain the freezing-transfer processing method;

[0123] Response: inform that the freezing-transfer processing method will be adopted, and explain the cross-regional transfer process.

[0124] 1.2. The user asks: "What should I do after that to open a public accumulation fund?"

[0125] State evolution: add state 4 = "cross-regional transfer business inquiry";

[0126] Control strategy: Combine the "cross-provincial transfer" specification in the public accumulation fund business database to explain the conditions and scope of the time (regional policy support).

[0127] Response: Inform the company's dedicated staff who will be responsible for the operation, explain the conditions and scope of the time, call the prompt template and prompt the required list of materials.

[0128] Scenario two: Multi-intention complex question processing (user asks for extraction and loan at the same time)

[0129] 2.1. User input: "I want to extract public accumulation fund for house decoration, can I also loan to buy a car?"

[0130] State judgment: State 8 = "extraction purpose recognition (decoration)", state 9 = "loan purpose judgment (non-housing)";

[0131] Control strategy: The module recognizes it as a complex intention and divides the task into two policy consultations for extraction and loan;

[0132] Query the database to confirm the compliance of the decoration purpose (regional policy support);

[0133] At the same time, prompt the loan purpose as non-supporting range, trigger the "policy restriction notification" template;

[0134] Response: Explain the decoration extraction policy basis and quota conditions in sections, and explain that public accumulation fund loans cannot be used for car purchases under the current policy, and guide the user to consult the supporting range of housing loans.

[0135] Scenario three: Pre-retirement extraction + account closure consultation process (two-stage policy understanding)

[0136] 3.1. User input: "I'm about to retire, can I take out my public accumulation fund first?"

[0137] State recognition: State 5 = "retirement related", state 8 = "extraction request recognition";

[0138] Control strategy: Trigger the pre-retirement extraction process;

[0139] Response: Query the "pre-retirement extraction" condition field in the public accumulation fund business database (such as retirement age, contribution years) to inform the user of the necessary preconditions, feasibility, output relevant policy interpretation, and extraction quota calculation explanation.

[0140] 3.2. User follow-up question: "Do I still need to manage the account after that?"

[0141] State evolution: Add state 7 = "account closure consultation";

[0142] Control strategy: Recognize it as a time sequence type complex state chain, trigger the pre-retirement extraction → account closure policy guidance closed loop process;

[0143] Response: Query and output the cancellation condition policy.

[0144] Scenario four: loan progress + extraction limit + repayment method three consecutive questions (complex multi-intention)

[0145] 4.1. User input: "Did I get the loan approved? How much money can I take for decoration? How to repay?"

[0146] State combination: state 9 = "loan application progress query", state 8 = "decoration extraction limit consultation", state 10 = "repayment method policy explanation";

[0147] Control strategy: identified as three parallel states, need to answer in stages;

[0148] Return the loan status explanation;

[0149] Use the historical state path to determine the extraction intention chain, and explain the limit calculation method (such as monthly deposit proportion). If it cannot be determined, ask back;

[0150] Retrieve the policy field of the public accumulation fund business database, output the standard repayment method classification and repayment channel;

[0151] Response: Answer the user's question in segments, and ask the user after each segment to avoid generating a full answer that takes too long and affects user experience.

[0152] In summary, the intelligent conversation strategy reply method provided by the present application injects a state classification system into the model with demand prompt words, enabling the model to have multi-label state understanding capabilities, improving the structured understanding of user intentions, enhancing multi-task recognition capabilities, and avoiding semantic drift; Time sequence records state trajectory, which cannot be modified and can be traced back, providing a basis for subsequent semantic judgment and policy execution, enabling dialogue tracking and context evolution management; All state record prompts are limited within a single conversation, and are destroyed with the end of the conversation, protecting user privacy, preventing semantic pollution, and controlling system resource usage range; Support for answering interruption, rewriting, guiding or refusing based on state judgment results; Support for inserting human rules or dynamically inserting review modules to improve system security and controllability, adapt to multiple scene requirements (such as sensitive questions, business specifications); Introduce a traceable intention summary and risk review to help users clarify their thoughts based on state evolution trajectory, enhance system explainability and user experience, and provide visual support for task management and customer service assistance.

[0153] Thus, the following advantages are achieved:

[0154] Intelligence improvement: large language models combined with state classification prompts can more accurately identify user intention categories and their evolution process;

[0155] Controllability enhancement: The system has multi-dimensional response control capabilities such as state awareness, rule execution, and manual insertion.

[0156] Clear data structure: State records have a unified format, clear life cycle, cannot be tampered with, and good structure.

[0157] Good system scalability: Rules and state modules can be flexibly extended to adapt to different industry tasks and output control requirements.

[0158] Security and governance advantages:

[0159] User data isolation: State records are not passed across sessions to protect user conversation data from misuse.

[0160] Sensitive topic warning: A specific state can trigger a prohibited answer mechanism to prevent the model from responding beyond its authority or violating rules.

[0161] Task audit support: State monitoring tools can be used to review user intent flow trajectories for business analysis.

[0162] Commercial and practical advantages:

[0163] Low deployment cost: Based on classification and rule control, no need to retrain the model, suitable for rapid deployment of enterprises.

[0164] Improved user experience: The model better understands the user's intent trajectory, and the answer is more relevant, avoiding off-topic and misunderstanding.

[0165] Strong compliance: Especially suitable for sensitive scenarios such as finance, healthcare, and education that require strong behavior.

[0166] In addition, the present application also provides an intelligent conversation strategy reply system. Figure 3 The system block diagram of the intelligent conversation strategy reply system according to the embodiment of the present application is shown. As shown in the figure, Figure 3 The intelligent conversation strategy reply system 100 of the present embodiment comprises:

[0167] The preliminary recognition module 110 is used to preliminarily match the initial user input sentence of the current conversation to obtain a demand prompt word.

[0168] The large model 120 is used to obtain a first state according to the demand prompt word, and a first response corresponding to the first state; and is also used to obtain a subsequent state and a subsequent response corresponding to each sentence of a subsequent user input sentence according to the subsequent user input sentence and the latest historical state data of the current conversation.

[0169] The history state module 130 is configured to receive the first state and the first response, store the first state in a history state database corresponding to the current session to form history state data, store a subsequent state corresponding to each sentence based on a time stamp in the history state database to update the history state data, and determine whether the current session is ended in real time, and if the current session is ended, the history state data is destroyed.

[0170] The large model 120 obtains the first state, including:

[0171] The demand prompt word is classified and divided by a preset multi-label classifier to determine a main intent corresponding to the demand prompt word.

[0172] The first keyword is extracted from the user input sentence, and a first operation type corresponding to the user input sentence is matched based on the first keyword.

[0173] The first state is determined based on the main intent and the first operation type.

[0174] The large model 120 obtains a subsequent state corresponding to each sentence of the subsequent user input sentence, including:

[0175] The subsequent keyword is extracted from the single sentence of the subsequent user input sentence, and a subsequent operation type and an auxiliary intent corresponding to the user input sentence are matched based on the subsequent keyword.

[0176] The subsequent state is determined based on the subsequent operation type, the auxiliary intent, and the main intent.

[0177] The history state module 130 sequentially arranges state data of the subsequent state corresponding to each sentence based on the time stamp to form a structured state data list, and updates the history state data by splicing the newly added subsequent state data in the form of a new table at the tail of the structured state data list.

[0178] For more specific embodiments of the intelligent conversation strategy reply system described above, refer to the specific embodiments of the intelligent conversation strategy reply method described above, and no further description is made here.

[0179] The intelligent conversation strategy reply method and system according to the present application are described above with reference to the accompanying drawings by way of example. However, those skilled in the art should understand that various improvements can be made to the intelligent conversation strategy reply method and system according to the present application without departing from the content of the present application. Therefore, the protection scope of the present application should be determined by the content of the appended claims.

Claims

1. A smart conversational strategy response method, characterized in that, include: Perform preliminary matching on the initial user input statement in the current session to obtain request prompts; The request prompt and the user input statement are sent to a preset large model to obtain a first state and a first response corresponding to the first state; Receive the first state and the first response, and store the first state in the historical state database corresponding to the current session to form historical state data; The subsequent user input statements of the current session and the latest historical state data are sent to the large model to obtain the subsequent states and subsequent responses corresponding to each statement of the subsequent user input statements, and the subsequent states corresponding to each statement are stored in the historical state database based on the timestamp to update the historical state data. Furthermore, it determines in real time whether the current session has ended, and if the current session has ended, it destroys the historical state data.

2. The intelligent conversation strategy response method as described in claim 1, characterized in that, The large model obtains a first state, which includes: A preset multi-label classifier is used to obtain the idea image corresponding to the requirement prompt words; Based on the user input statement, keywords are extracted to obtain a first keyword, and based on the first keyword, a first operation type corresponding to the user input statement is matched. The first state is determined based on the concept diagram and the first operation type.

3. The intelligent conversation strategy response method as described in claim 2, characterized in that, The large model acquires subsequent states corresponding to each subsequent user input statement, including: Keyword extraction is performed on each individual statement of the subsequent user input statement to obtain subsequent keywords, and the subsequent operation type and auxiliary intent corresponding to the user input statement are matched based on the subsequent keywords; The subsequent state is determined based on the type of subsequent operation, the auxiliary intent, and the idea graph.

4. The intelligent conversation strategy response method as described in claim 3, characterized in that, The first state or the subsequent states include question, instruction, opinion, fact, sensitivity, risk, and repetition.

5. The intelligent conversation strategy response method as described in claim 4, characterized in that, Obtaining a first or subsequent response, including: Determine the first response based on the first state; Determine the subsequent response based on the subsequent state; in, If the first state or the subsequent state is a question, then the first response or the subsequent response is: according to the first operation type or the subsequent operation type, call the target document in the preset database and send it to the current session, or extract key information based on the target document to form a strategy response and send the strategy response to the current session; If the first state or the subsequent state is an indication, then the first response or the subsequent response is: to perform a business state change corresponding to the indication based on the first operation type or the subsequent operation type; If the first state or the subsequent state is an opinion or a fact, then the first response or the subsequent response is: selecting a response statement corresponding to the idea or the auxiliary intention from the database, and sending the response statement to the current session; If the first state or the subsequent state is sensitive, then the first response or the subsequent response is: refuse to answer the question, issue a sensitive prompt to the current session, and wait for the next subsequent user input statement; If the first state or the subsequent state is a risk, then the first response or the subsequent response is: to issue a risk warning to the current session.

6. The intelligent conversation strategy response method as described in claim 4, characterized in that, Based on timestamps, the subsequent states corresponding to each statement are stored in the historical state database to update the historical state data, including: The status data of each statement's subsequent status are arranged sequentially according to the timestamps to form a structured status data list. Newly added subsequent status data is appended to the end of the structured status data list in the form of a new table to update the historical status data.

7. An intelligent conversational strategy response system, characterized in that, include: The pre-entry recognition module is used to perform preliminary matching on the initial user input statement of the current session to obtain the required prompt words; The large model is used to obtain a first state and a first response corresponding to the first state based on the demand prompt words; it is also used to obtain subsequent states and subsequent responses corresponding to each statement of the subsequent user input statement based on the subsequent user input statements of the current session and the latest historical state data. The historical state module is used to receive the first state and the first response, and store the first state in the historical state database corresponding to the current session to form historical state data. It is also used to store the subsequent state corresponding to each statement in the historical state database based on the timestamp to update the historical state data; and to determine in real time whether the current session has ended, and if the current session has ended, to destroy the historical state data.

8. The intelligent conversational strategy response system as described in claim 7, characterized in that, The large model obtains its first state, including: The concept image corresponding to the requirement prompt is determined by classifying the requirement prompt based on the requirement prompt using a preset multi-label classifier. Based on the user input statement, keywords are extracted to obtain a first keyword, and based on the first keyword, a first operation type corresponding to the user input statement is matched. The first state is determined based on the concept diagram and the first operation type.

9. The intelligent conversational strategy response system as described in claim 8, characterized in that, The large model acquires subsequent states corresponding to each subsequent user input statement, including: Keyword extraction is performed on each individual statement of the subsequent user input statement to obtain subsequent keywords, and the subsequent operation type and auxiliary intent corresponding to the user input statement are matched based on the subsequent keywords; The subsequent state is determined based on the type of subsequent operation, the auxiliary intent, and the idea graph.

10. The intelligent conversational strategy response system as described in claim 9, characterized in that, The historical state module arranges the status data of the subsequent states corresponding to each statement in chronological order according to the timestamps, forming a structured status data list. Newly added subsequent status data is appended to the end of the structured status data list in the form of a new table to update the historical status data.

Citation Information

Cited By

  • Household user dialogue method and system based on life cycle design

    CN121706968A