Context and intention prediction-based verbal skill recommendation method and device, equipment and medium

By acquiring user input and dialogue history data, and combining intent recognition and reasoning graphs, targeted responses and guiding messages are generated, solving the problem of the lack of specificity and logical coherence in the messages in traditional customer service systems, and improving the interaction quality and efficiency of intelligent customer service systems.

CN121996765APending Publication Date: 2026-05-08PING AN HEALTH CLOUD CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
PING AN HEALTH CLOUD CO LTD
Filing Date
2026-01-23
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

Traditional intelligent customer service systems struggle to capture the dynamic evolution of user needs, resulting in a lack of targeted and logical coherence in their communication, which reduces the accuracy of their responses.

Method used

By acquiring the user's current input information and multi-turn dialogue history data, combined with intent recognition and reasoning graphs, potential subsequent intents are predicted, and targeted response and guiding dialogues are generated to improve the relevance and accuracy of the dialogues.

Benefits of technology

It improves the accuracy of customer service responses, optimizes user experience and message reply efficiency, and provides more accurate responses in complex dialogue scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121996765A_ABST
    Figure CN121996765A_ABST
Patent Text Reader

Abstract

The invention provides a verbal skill recommendation method based on context and intention prediction, and the method comprises the steps: introducing intention recognition and a potential subsequent intention prediction mechanism based on a affair graph, and fusing the current input information, a business knowledge context and a multi-round dialogue context corresponding to multi-round dialogue historical data; and the current intention of the user is responded, and a more accurate target reply verbal skill is generated. Then, the intention evolution of the potential follow-up intention of the current user is further predicted based on the affair graph, so that the potential demand of the user is pre-judged, and the target guide verbal skill is generated aiming at the potential demand, so that the verbal skill pertinence of the generated verbal skill in a complex dialogue scene is improved through a recommendation mode of combining the reply verbal skill and the guide verbal skill; therefore, the verbal skill generation accuracy is improved, and the user experience and the message reply efficiency are improved. The method can be applied to a seat verbal skill recommendation function in the financial field or the medical field, so that the seat verbal skill generation accuracy is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of intelligent decision-making, and in particular to a method, apparatus, computer device, and computer-readable storage medium for recommending speech based on context and intent prediction. Background Technology

[0002] Traditional intelligent customer service systems use keyword matching and regular expressions to process user input. However, this method struggles to capture the coherence between conversations, leading to gaps in understanding user needs and an inability to accurately identify the dynamic evolution of user intent. This results in generated responses lacking focus and logical coherence, thus reducing accuracy. For example, in the medical field, a patient might start by asking, "Can people with high blood pressure eat bananas?", then move on to "How many bananas should they eat each day?", and further to "What other fruits are suitable for people with high blood pressure?". Similarly, in the financial sector, a customer might begin by asking, "How do I apply for credit card installment plans?", followed by "How are installment fees calculated?", and then "What are the consequences of early repayment?". Therefore, improving the accuracy of customer service responses has become a pressing technical challenge. Summary of the Invention

[0003] The main purpose of this application is to provide a method, apparatus, computer device, and computer-readable storage medium for recommending customer service responses based on context and intent prediction, with the aim of improving the accuracy of customer service response scripts.

[0004] To achieve the above objectives, this application provides a dialogue recommendation method based on context and intent prediction, the dialogue recommendation method comprising the following steps: Obtain the current input information of the current user and the multi-turn dialogue history data between the current user and the agent, wherein the multi-turn dialogue history data includes historical input information and its corresponding agent response information; The current input information is subjected to intent recognition to obtain the current intent recognition result, and a target response script is generated based on the script generation model, which includes a preset business knowledge context, the current input information, and the multi-turn dialogue history data. Based on the preset reasoning graph and the current intent recognition result, the potential subsequent intent of the current user is predicted, and based on the speech generation model, the multi-turn dialogue history data, the business knowledge context, and the target guiding speech corresponding to the potential subsequent intent are generated. Based on the target response script and / or the target guidance script, the recommended script for the current user is obtained.

[0005] Furthermore, to achieve the above objectives, this application also provides a dialogue recommendation device based on context and intent prediction, the dialogue recommendation device comprising: The context acquisition module is used to acquire the current input information of the current user and the multi-turn dialogue history data between the current user and the agent. The multi-turn dialogue history data includes historical input information and its corresponding agent response information. The response script generation module is used to perform intent recognition on the current input information, obtain the current intent recognition result, and generate a target response script corresponding to the preset business knowledge context, the current input information, and the multi-turn dialogue history data based on the script generation model. The guidance dialogue generation module is used to predict the potential subsequent intent of the current user based on a preset reasoning graph and the current intent recognition result, and to generate the target guidance dialogue corresponding to the multi-turn dialogue history data, the business knowledge context, and the potential subsequent intent based on the dialogue generation model. The customer service script recommendation module is used to obtain recommended scripts for the current user based on the target response script and / or the target guidance script.

[0006] In addition, to achieve the above objectives, this application also provides a computer device, the computer device including a processor, a memory, and a context- and intent-based speech recommendation program stored in the memory and executable by the processor, wherein when the context- and intent-based speech recommendation program is executed by the processor, it implements the steps of the context- and intent-based speech recommendation method as described above.

[0007] In addition, to achieve the above objectives, this application also provides a computer-readable storage medium storing a context- and intent-based speech recommendation program, wherein when the context- and intent-based speech recommendation program is executed by a processor, it implements the steps of the context- and intent-based speech recommendation method described above.

[0008] This application provides a dialogue recommendation method based on context and intent prediction. The method acquires the current user's current input information and multi-turn dialogue history data between the current user and the agent, including historical input information and corresponding agent responses. It then performs intent recognition on the current input information to obtain a current intent recognition result, and generates a target response dialogue based on a dialogue generation model, using a preset business knowledge context, the current input information, and the multi-turn dialogue history data. Based on a preset event graph and the current intent recognition result, it predicts the current user's potential subsequent intent, and generates a target guiding dialogue based on the multi-turn dialogue history data, the business knowledge context, and the potential subsequent intent. Finally, it obtains a recommended dialogue for the current user based on the target response dialogue and / or the target guiding dialogue. Through this approach, this application introduces intent recognition and a potential subsequent intent prediction mechanism based on an event graph. Furthermore, based on intent recognition, it integrates the current input information, business knowledge context, and the multi-turn dialogue history data corresponding to the multi-turn dialogue context to respond to the user's current intent and generate more accurate target response dialogues. Then, based on the event graph, the potential evolution of the user's subsequent intentions is predicted, thereby anticipating the user's potential needs and generating target guidance messages for these needs. By combining the recommendation of response messages and guidance messages, the relevance of the generated messages in complex dialogue scenarios is improved, thus enhancing the accuracy of the generated messages, improving user experience, and increasing message reply efficiency. Attached Figure Description

[0009] Figure 1 A flowchart illustrating a context- and intent-based speech recommendation method provided in this application; Figure 2 A flowchart illustrating another context- and intent-based speech recommendation method provided in this application; Figure 3 A flowchart illustrating another context- and intent-based speech recommendation method provided in this application; Figure 4 A schematic diagram of the functional modules of a context- and intent-based speech recommendation device provided in this application; Figure 5 A schematic block diagram of the structure of a computer device provided in this application.

[0010] The realization of the purpose, functional features and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0011] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0012] The flowchart shown in the attached diagram is for illustrative purposes only and does not necessarily include all content and operations / steps, nor does it necessarily have to be performed in the described order. For example, some operations / steps can be broken down, combined, or partially merged, so the actual execution order may change depending on the actual situation.

[0013] It should be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the scope of the application. As used in this specification and the appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms unless the context clearly indicates otherwise.

[0014] It should also be understood that the term “and / or” as used in this application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.

[0015] The context- and intent-based speech recommendation method described in this application is mainly applied to computer devices, such as PCs, laptops, and mobile terminals, which have display and processing capabilities.

[0016] The following detailed description of some embodiments of this application is provided in conjunction with the accompanying drawings. Unless otherwise specified, the following embodiments and features can be combined with each other.

[0017] Reference Figure 1 , Figure 1 This is a flowchart illustrating a context- and intent-based speech recommendation method provided in this application.

[0018] like Figure 1 As shown in the figure, this application provides a dialogue recommendation method based on context and intent prediction. By analyzing user input and dialogue history and combining intent prediction, it provides users with appropriate replies or guiding suggestions, thereby improving the interaction quality and efficiency of the intelligent customer service system. The dialogue recommendation method based on context and intent prediction includes steps S101 to S104.

[0019] In this embodiment, the context- and intent-based speech recommendation method includes the following steps: Step S101: Obtain the current input information of the current user and the multi-turn dialogue history data between the current user and the agent. The multi-turn dialogue history data includes historical input information and its corresponding agent response information. In this embodiment, the current input information refers to the text or voice information submitted by the user to the intelligent customer service system at the current moment. This information forms the basis for the system's subsequent processing and response. Multi-turn dialogue history data refers to records of multiple interactions between the user and the intelligent customer service system or agent. This data typically includes the user's input information in different rounds and the corresponding responses from the system or agent, providing context for the dialogue.

[0020] Specifically, the system acquires the current user's current input information and the user's multi-turn dialogue history with the agent. This multi-turn dialogue history includes historical input information and its corresponding agent responses. For example, the current input information could be text typed by the user in the chat window, such as "I want to check my credit card statement." The multi-turn dialogue history could be all previous dialogue records between the user and customer service stored in the system, such as the user asking "How do I apply for a credit card," and the customer service representative replying with the application process. In one implementation, this information can be acquired by directly receiving text input from the user interface and retrieving dialogue records associated with the user ID from the backend database. In another implementation, this information can be obtained by converting the user's speech into text using a speech recognition module and extracting historical dialogue data from a distributed storage system.

[0021] Step S102: Perform intent recognition on the current input information to obtain the current intent recognition result, and generate the target response script corresponding to the preset business knowledge context, the current input information, and the multi-round dialogue history data based on the script generation model; In this embodiment, business knowledge context refers to a set of structured or unstructured knowledge related to a specific business domain. This knowledge may include product information, service processes, frequently asked questions, industry standards, etc., to provide professional background and accuracy for script generation. Target response script refers to the response text directly generated by the script generation model based on the user's current input information and identified intent, aimed at resolving the user's current problem.

[0022] Specifically, the current input information is subjected to intent recognition to obtain the current intent recognition result. Based on the dialogue generation model, a target response dialogue corresponding to the preset business knowledge context, the current input information, and the historical data of the multi-turn dialogue is generated. Intent recognition aims to understand the deeper purpose of the user's current input. For example, for the input "I want to check my credit card statement," the intent recognition result might be "Check statement." The dialogue generation model then uses this information to construct a response that directly answers the user's question. In one implementation, intent recognition can be performed using rule-based matching or keyword matching, comparing the user input with a preset intent template. The dialogue generation model can be a template-based system that selects and fills relevant information from a predefined dialogue library based on the identified intent and business knowledge context. For example, if the intent is "Check statement," then "Which month's statement do you want to check?" is generated.

[0023] For example, step S102 includes: The key text of the current input information, the business knowledge context, and the multi-turn dialogue history data (multi-turn dialogue context) are concatenated to form a comprehensive input sequence; The comprehensive input sequence is then fed into the speech generation model. The speech generation model encodes the comprehensive input sequence to generate a context vector containing sequence semantics; Based on the context vector, the first candidate recommendation phrase text sequence is generated by the decoder in an autoregressive manner; A cluster search strategy is adopted to retain the N candidate sequence paths with the highest probability during the decoding process, where N is an integer greater than 1, thereby generating N candidate recommendation phrases that differ in text expression; The at least one candidate recommended dialogue is scored in multiple dimensions, including consistency with the business knowledge context, coherence with the multi-turn dialogue history data (multi-turn dialogue context), and matching degree with the sentiment polarity label in the first inspection result. The multi-dimensional scores are weighted and fused based on preset weights to generate a confidence score for each candidate recommendation phrase; Based on the confidence score, the final recommended response is selected from the at least one candidate response as the target response.

[0024] In another embodiment, step S102 includes: The key text of the current input information, the business knowledge context, and the multi-turn dialogue history data (multi-turn dialogue context) are integrated to generate a comprehensive contextual representation that includes the current question, background knowledge, and historical dialogue state. Based on the comprehensive context representation, the core semantics of the current dialogue and the task objectives to be completed are understood, and semantic planning of the speech content is generated. Based on the semantic planning, candidate recommended texts containing at least one of the following elements are generated: question answering, operation guidance, or emotional resonance.

[0025] Assess the degree of consistency between the candidate recommended statements and the factual knowledge contained in the business knowledge context; Evaluate the logical coherence and topic relevance between the candidate recommended dialogues and the multi-turn dialogue history data (multi-turn dialogue context); Evaluate the degree of match between the candidate recommended statements and the sentiment polarity tags contained in the first inspection results; The confidence score is generated by integrating the evaluation results of consistency, logical coherence, and matching degree.

[0026] Based on the confidence score, a final recommended response is determined from the one or more candidate response responses, which serves as the target response response.

[0027] Step S103: Based on the preset reasoning graph and the current intent recognition result, predict the potential subsequent intent of the current user, and based on the speech generation model, generate the multi-turn dialogue history data, the business knowledge context, and the target guiding speech corresponding to the potential subsequent intent. In this embodiment, the event graph refers to a knowledge graph that describes the causal, sequential, or conditional relationships between events. This graph uses events as nodes and the relationships between events as edges, and is used to reason about and predict the development path of events or the user's potential needs. Potential follow-up intentions refer to the intentions the user might express in subsequent dialogues, predicted by the system based on the current user's intentions and the event graph. This prediction helps the system prepare in advance and provides proactive guidance. Targeted guidance dialogue refers to text generated by a dialogue generation model based on the predicted potential follow-up intentions, designed to guide the user to continue the dialogue in a specific direction or explore relevant information.

[0028] Specifically, based on a preset event graph and the current intent recognition result, the potential subsequent intents of the current user are predicted. Then, based on the dialogue generation model, the multi-turn dialogue history data, the business knowledge context, and the target guiding dialogue corresponding to the potential subsequent intents are generated. The event graph is used to capture the logical relationships between events, thereby predicting other intents that the user may have after the current intent. For example, if the current intent is "checking the bill," the event graph may predict potential subsequent intents such as "understanding installment payments" or "applying for bill adjustments." The target guiding dialogue then proactively provides relevant information or guides the user to the next step based on these predictions. In one implementation, the event graph can be a manually constructed knowledge base containing fixed transition paths between intents in various business scenarios. The prediction of potential subsequent intents can be accomplished by simply finding the direct successor intent of the current intent in the event graph. The dialogue generation model can select and generate corresponding guiding statements from preset guiding dialogue templates based on these predicted subsequent intents, such as "Do you need to understand credit card installment services?".

[0029] For example, step S103 includes: Based on the current intent recognition result, locate the corresponding current intent node in the event graph; Starting from the current intent node, retrieve all directly connected subsequent intent nodes in the event graph and obtain their corresponding historical transition probabilities; Based on customer profiles and historical behavior data, the historical transfer probabilities are weighted and corrected to generate personalized intent evolution predictions that reflect specific customer behavior patterns. Based on one or more of the most probable subsequent intentions in the personalized intention evolution prediction, corresponding pre-answer dialogues, related knowledge or operation suggestions are generated as proactively recommended content, i.e., target-guided dialogues.

[0030] In another embodiment, step S103 includes: Using the current intent recognition result as the query keyword, a matching query is performed in the intent node index of the event graph to locate the corresponding current intent node; Using a graph database query language, all outgoing edges are traversed from the current intent node, and the subsequent intent nodes pointed to by each outgoing edge and the transition probability weights attached to the edges are obtained. The feature vectors in the customer profile are concatenated with the intent sequence features in the historical behavior data and input into a probability adjustment model, which outputs the probability adjustment coefficient for each subsequent intent node. The probability adjustment coefficient is multiplied by the corresponding historical transfer probability weight to obtain the personalized transfer probability, and the top N subsequent intent nodes with the highest probabilities are selected, where N is a preset positive integer. Based on the selected subsequent intent node, related script templates and knowledge content are retrieved from the preset script template library and knowledge card library, and after variable filling, the proactive recommendation content is generated.

[0031] For example, step S103 includes: The multi-turn dialogue history data is concatenated and encoded with the business knowledge context to generate a unified dialogue state representation vector; The potential subsequent intent is encoded into an intent embedding vector and fused with the dialogue state representation vector to generate a target-guided contextual representation; The target-guided contextual representation is input into the speech generation model, and a candidate speech text sequence is generated in an autoregressive manner through a decoder; The generated candidate dialogue text sequence is subjected to fluency verification and context consistency check, and the final target guiding dialogue is output.

[0032] In another embodiment, step S103 further includes: By integrating the multi-turn dialogue history data with the business knowledge context, a dialogue context vector that comprehensively represents the current dialogue state and available knowledge is generated. The potential follow-up intention is used as the generation target and combined with the dialogue context vector to generate a guiding context representation, so that the generated utterance serves to guide the dialogue toward the potential follow-up intention; Based on the aforementioned guiding context representation, one or more candidate guiding statements are generated that respond to historical dialogues in terms of content, integrate business knowledge in terms of information, and point to the potential subsequent intentions in terms of objectives. The candidate guiding statements are evaluated and optimized to output target guiding statements that meet preset quality standards.

[0033] Step S104: Based on the target response script and / or the target guidance script, obtain the recommended script for the current user.

[0034] In this embodiment, the recommended dialogue refers to the dialogue that is ultimately presented to the user. It can be a target response dialogue, a target guidance dialogue, or a combination of both, in order to meet the user's needs and optimize the dialogue process.

[0035] Specifically, based on the target response and / or the target guidance, a recommended response is generated for the current user. The system presents the generated target response and / or guidance to the user as a recommendation. This allows the system not only to directly answer the user's question but also to proactively guide the user to explore relevant information, enhancing the depth and breadth of the conversation. In one implementation, the system can choose to display only the target response or only the target guidance based on preset priority rules. For example, if the target response can completely resolve the user's current problem, only the response response is displayed. If the target response is unclear or has multiple potential follow-up intentions, both the response and guidance can be displayed simultaneously, or only the guidance can be displayed.

[0036] For example, the context- and intent-based speech recommendation method further includes incremental learning and model optimization steps: Collect records of agents' adoption of recommended scripts and customer feedback data; Based on the collected data, incremental learning training samples are constructed to update the processing model; Update the multimodal knowledge graph and the event graph based on new business documents and customer questions.

[0037] For example, the context- and intent-based speech recommendation method also includes an intent conflict resolution mechanism: Perform intent recognition on each message in a multi-turn dialogue to generate an intent sequence; Analyze the intent evolution patterns in the intent sequence and detect intent conflicts; Based on preset rules, the detected intent conflicts are resolved to identify the end customer's intent; Based on the end customer's intent, generate new business knowledge context information, response scripts, and guidance scripts.

[0038] In one embodiment, before generating the preset business knowledge context, the current input information, the current intent recognition result, and the target response script corresponding to the multi-turn dialogue history data based on the script generation model, the method further includes: The current input information is subjected to a semantic validity check to obtain a first check result. Based on the first check result, the sentiment intensity and question type of the current input information are determined. The first check result includes a completeness score, sentiment polarity label, and question word detection result. Based on the first inspection result, the current intent recognition result, and the current input information, a message feature vector of the current user is generated; Obtain the current customer service system load, and based on the routing decision model, generate the selection probability distributions of the general large model, the domain fine-tuning model, and the small parameter model under the current customer service system load and the message feature vector; When the current customer system load is lower than a preset load threshold, the model with the highest probability of the selection probability distribution is used as the script generation model; When the current customer system load is not lower than a preset load threshold, the small parameter model is used as the script generation model.

[0039] In this embodiment, the multi-model routing mechanism includes: The intent complexity and intent category are parsed from the current intent recognition result; Extract the emotional intensity and problem type from the first examination results; Feature extraction is performed on the key text to obtain message length features and the number of key entities; The feature vector of the current input information is generated by combining the features of intent complexity, intent category, sentiment intensity, question type, message length, and number of key entities.

[0040] The feature vector, along with the current system load and time information, is used as the state input and fed into the routing decision model trained by reinforcement learning. Based on the state, the routing decision model calculates and outputs a selection probability distribution that includes a general large model, a domain fine-tuning model, and a small parameter model.

[0041] Monitor real-time system load metrics; Compare the system load index with a preset threshold; When the system load index is lower than the preset threshold, the processing model with the highest probability is selected as the model selection result according to the selection probability distribution. When the system load index is higher than or equal to the preset threshold, the small parameter model is selected as the model selection result.

[0042] Specifically, the semantic validity check includes: performing integrity analysis on the text messages in the first input data to generate an integrity score; performing sentiment analysis on the first input data to generate a multimodal sentiment vector; detecting interrogative semantics in the first input data and identifying the question type; calculating a comprehensive semantic validity score based on the integrity score, the multimodal sentiment vector, and the question type, and generating the first check result.

[0043] Semantic validity checks aim to assess the quality and usability of user input, ensuring that subsequent processing is based on meaningful input. This can be achieved through grammar and spell checking, using natural language processing tools to detect grammatical and spelling errors in the input text; or through information density analysis, calculating the ratio of the number of entities, keywords, or information units in the text to the total text length to assess the richness of information. The first check result is the output of the semantic validity check, used to quantify the characteristics of the input information. The completeness score reflects the grammatical, semantic, or informational completeness of the input information, for example, it can be calculated based on the completeness of the syntax tree or the coverage of key information elements. Sentiment polarity labels identify the sentiment tendency expressed by the input information, such as positive, negative, or neutral, which can be achieved through sentiment analysis models or dictionary matching methods. Interrogative word detection results identify whether the input information contains interrogative words to determine whether the user has asked a clear question, which can be achieved through pre-defined interrogative word list matching or methods based on part-of-speech tagging and syntactic analysis. Sentiment intensity refers to the strength of the emotional expression in the input information, which can be quantified based on the confidence score of the sentiment polarity label or the density of sentiment words. The question type refers to the specific category of the question posed by the input information, which can be identified based on interrogative word detection results, keyword matching, or a pre-trained question classification model. The message feature vector is a vector that numerically represents the user's current input information and its related attributes. Its function is to encode multi-dimensional, heterogeneous information into a format that can be processed by machine learning models. This can be achieved through feature concatenation, combining the first check result, the current intent recognition result, and the current input information; or through neural network encoding, using neural network structures such as multilayer perceptrons or Transformer encoders, using each input as a feature, and learning to obtain a compact and semantically rich vector representation. The current customer service system load refers to the workload or resource consumption of the customer service system at a certain moment. It can be obtained by monitoring systems to collect metrics such as CPU utilization, memory usage, and concurrent requests in real time; or by calling internal API interfaces to obtain current system load data from load balancers or resource schedulers. A routing decision model is a machine learning model used to dynamically select the optimal processing path or model based on input features. Its implementation can include training a multi-classifier, taking the customer service system load and message feature vectors as input, and outputting the selection probability for each dialogue generation model; or using a reinforcement learning model, treating model selection as a decision-making process and training a strategy through a reinforcement learning algorithm. A general-purpose large model refers to a language model with a large number of parameters, pre-trained on massive general corpora, possessing strong generalization capabilities. A domain-fine-tuned model refers to a model further trained on data from a specific business domain, based on the general-purpose large model. A small-parameter model refers to a dialogue generation model with relatively few parameters, low computational resource requirements, and fast inference speed.The selection probability distribution represents the likelihood of the routing decision model choosing each speech generation model. The preset load threshold is a pre-defined system load limit used to distinguish whether the system is in normal operation or under high load.

[0044] By meticulously checking the semantic validity of the current user's input, analyzing its completeness score, sentiment polarity label, and interrogative word detection results, the system not only assesses the quality of the input but also further determines the user's sentiment intensity and question type. This rich semantic information, along with the current intent recognition results and the original input information, is integrated and encoded into a comprehensive message feature vector. This vector accurately characterizes the complexity and characteristics of the user's current needs. Simultaneously, the system acquires real-time information about the current customer service system load. Subsequently, a pre-trained routing decision model combines this message feature vector with the current system load to dynamically evaluate the applicability of general-purpose models, domain-fine-tuned models, and small-parameter models, outputting a selection probability distribution. This distribution reflects the probability of each model being selected in the current specific context. When the system load is below a preset threshold, the system prioritizes the quality and accuracy of the generated dialogue, thus selecting the model with the highest predicted probability from the routing decision model. This typically means that when resources are sufficient, the system tends to use models with better performance and more comprehensive knowledge. However, when the system load is not lower than a preset threshold, in order to ensure system stability and response speed, the system strategically selects a low-parameter model with lower resource consumption and faster inference speed as the script generation model. Through this dynamic model selection strategy, the script recommendation method of this application can intelligently achieve a balance between script generation quality and system response efficiency based on the characteristics of user input information and the real-time operating status of the customer service system. This not only optimizes the utilization of system resources and avoids unnecessary resource waste, but also ensures that users can still obtain timely script recommendation services in high-concurrency scenarios, thereby significantly improving the overall user experience and system operating efficiency.

[0045] Through the above technical solution, the script recommendation method of this application can dynamically select the most suitable script generation model based on the semantic characteristics of user input information and the real-time operating status of the customer service system. This intelligent routing decision mechanism avoids the limitations of a single model or a fixed model in different scenarios, effectively solving the problem of balancing system resource consumption and response efficiency while ensuring the quality of script generation. Specifically, when the system load is low, it can fully utilize the high-performance model to provide high-quality and high-accuracy script recommendations; while when the system load is high, it can quickly switch to a lightweight model to ensure that the system can maintain a fast response under high concurrency pressure, avoiding service delays or system crashes caused by excessively long model inference time, thereby significantly improving the user experience and the overall robustness and efficiency of the system.

[0046] In one embodiment, generating the message feature vector of the current user based on the current intent recognition result and the current input information includes: Extract message length features and key entity counts from the current input information, and parse the current user's intent complexity and intent category from the current intent recognition result; The message feature vector is generated based on the emotional intensity, question type, message length features, number of key entities, intent complexity, and intent category.

[0047] In this embodiment, the message length feature is used to quantify the redundancy or conciseness of the current input information. Its implementation can include: one method is to count the number of characters in the current input information, for example, by calculating the string length; another method is to segment the current input information and then count the number of words contained within it to reflect the amount of information. The key entity count aims to identify and count the number of entity words in the current input information that have specific meaning or business relevance; these entities are often the core carriers of user intent. Specifically, named entity recognition (NER) technology can be used to identify and count the number of entities in predefined categories such as names of people, places, organizations, product names, and dates; or, a preset business dictionary or keyword extraction algorithm can be used to match and count the number of keywords strongly related to a specific business domain. Intent complexity is used to assess the structural complexity of the current user intent, such as whether it is a single, explicit intent or contains multiple sub-intents or implicit intents. One approach is to determine whether the intent belongs to a deep or complex intent within a multi-level structure based on the intent classification results output by the intent recognition model, and assign a corresponding complexity score. Another approach is to analyze the keyword density and semantic correlation between keywords in the intent recognition results; high density and strong correlation may indicate a more complex intent. Intent categories specify the predefined category to which the intent expressed by the current user input belongs, such as inquiry, complaint, query, purchase, appointment, etc. This can be achieved by directly outputting intent labels from the intent recognition model as intent categories; or, if the intent system is hierarchical, extracting first-level intents or more granular second-level intents as intent categories. Generating the message feature vector aims to integrate the various heterogeneous features extracted above, including sentiment intensity, question type, message length features, number of key entities, intent complexity, and intent category, into a unified numerical representation. One approach is to concatenate all numerical features and normalize them, while simultaneously encoding or embedding categorical features one-hot and concatenating them with the numerical features. Another approach is to assign different weights to different features based on their importance to the model selection, and then perform weighted summation or weighted concatenation to form the final feature vector for subsequent routing decision models to process.

[0048] To address the issue of insufficient refinement in message feature vector construction, this application proposes a more comprehensive message feature vector generation mechanism. This mechanism first performs in-depth analysis of the current user's input information, meticulously extracting message length features and the number of key entities to quantify the surface-level information and core concerns of the user input. Simultaneously, combined with the obtained current intent recognition results, it further analyzes the complexity and category of the current user's intent, thereby revealing the deep structure and specific direction of the user's intent. These features extracted from different dimensions, along with the sentiment intensity and question type obtained from semantic validity checks, are systematically integrated. By unifying and fusing these heterogeneous features with numerical representation, a comprehensive and refined message feature vector is finally generated. This vector accurately captures the deep semantics, sentiment tendency, complexity, and the business intent implied by the user message. This message feature vector is then used in a routing decision model to more accurately evaluate the applicability of the general large model, the domain fine-tuning model, and the small parameter model under the current customer service system load, and to generate their probability selection distributions. This refined feature engineering allows model selection to move beyond rough judgments and instead be based on a deep understanding of user needs, thereby significantly improving the accuracy and relevance of the recommended messages.

[0049] By meticulously extracting message length features and the number of key entities from the current input information, and parsing intent complexity and intent category from the current intent recognition results, this application can construct a more comprehensive and refined message feature vector. This vector not only contains the basic attributes of user input but also deeply reflects the deep structure and complexity of its intent. Therefore, in the subsequent dialogue generation model selection process, the routing decision model can more accurately evaluate the applicability of different models based on this high-dimensional, information-rich message feature vector, thereby selecting the dialogue generation model that best matches the current user needs and system state. This significantly improves the accuracy of dialogue recommendation and user satisfaction, effectively solves the problem of model selection bias caused by insufficient feature extraction, and thus improves the overall intelligence level and response efficiency of the dialogue recommendation system.

[0050] Reference Figure 2 , Figure 2 A flowchart illustrating another context- and intent-based speech recommendation method provided in this application.

[0051] To further address the above issues, such as Figure 2 As shown, this application embodiment provides a speech recommendation method based on context and intent prediction, wherein step S102 specifically includes: Step S1021: The current input information is segmented into sentences to obtain each sentence. Semantic role recognition is performed on each sentence to obtain the semantic role recognition result of each word in each sentence. Based on the semantic role recognition result, the semantic correlation degree between each sentence and the interrogative word detection result is calculated. Step S1022: Based on the attention mechanism, the multi-turn dialogue history data, and the semantic relevance of each sentence, calculate the attention weight of each sentence, and generate the key text of the current input information based on the sentences whose attention weight exceeds the preset weight threshold. Step S1023: Perform intent classification on the key text to obtain multiple candidate intents and their corresponding probability distributions, and generate the current intent recognition result based on the sentiment polarity label, the question type, the multiple candidate intents and their corresponding probability distributions.

[0052] In this embodiment, intent recognition is performed on the current input information to obtain the current intent recognition result, including: The text in the first input data is segmented into sentences to obtain multiple sentences; Dependency parsing is performed on each clause to identify its subject-verb-object structure and core predicate; Based on the results of the dependency parsing, semantic roles are labeled for the words in each clause.

[0053] Based on the semantic role recognition results, the semantic relevance between each clause and the interrogative word detection results in the first inspection result is calculated; By combining multi-turn dialogue history data (multi-turn dialogue context) and using an attention mechanism, the attention weight of each sentence is calculated. Select one or more clauses with the highest attention weight, remove redundant and irrelevant information, and then concatenate them to generate the key text.

[0054] Extract the detected interrogative words or phrases from the first inspection results; For each clause, extract its predicate core word and object entity determined by semantic role recognition results; Calculate the semantic similarity between the predicate core word and object entity and the interrogative word or phrase, and use the semantic similarity as the semantic relevance of the clause.

[0055] Use the semantic relevance of all clauses in the current round as the initial weight; Encode the key texts from previous rounds in the multi-turn dialogue history data (multi-turn dialogue history data (multi-turn dialogue context)) and each sentence in the current round to generate context-aware sentence representation vectors; Based on the self-attention mechanism, the correlation between the representation vectors of each sentence in the current round is calculated, and the initial weights are weighted and corrected to obtain the final attention weight of each sentence.

[0056] All sentences are sorted in descending order according to the attention weights mentioned above; Select the top K clauses, where K is a preset positive integer; Redundant information is filtered from the selected clauses, including the removal of semantically repetitive modifiers and interjections; The filtered sentences are concatenated in the order they appear in the original input to generate the key text.

[0057] The key text is input into a pre-trained intent classification model; The intent classification model outputs multiple candidate intents and their probability distributions corresponding to a preset intent category set.

[0058] Obtain the emotional polarity label and question type from the first examination results; The probability distribution of the candidate intentions is adjusted based on preset rules: when the emotional polarity is negative, the weight of intentions related to "complaint" or "fault reporting" is increased; when the problem type is "solution-oriented problem", the weight of intentions related to "seeking solutions" is increased. The candidate intent with the highest probability after weight adjustment is determined as the current intent recognition result.

[0059] Specifically, the current input information is first segmented into sentences. Sentence segmentation aims to break down potentially long user input texts containing multiple independent semantic units into smaller, easier-to-analyze independent sentence units. This can be achieved through text segmenters based on punctuation (such as periods, question marks, and exclamation marks) and grammatical rules, or by using deep learning-based sequence labeling models, such as Transformer-based sentence boundary detection models, which learn from a large amount of labeled data to identify sentence boundaries. Then, semantic role identification is performed on each sentence to obtain the semantic role identification results for each word in each sentence. The purpose of semantic role identification is to gain a deep understanding of the predicate-argument structure of each sentence, that is, to identify semantic elements such as the performer, receiver, time, place, and manner of the action. This can be achieved by using rule-based semantic role labeling tools based on predefined grammatical patterns and dictionaries; or by using neural network-based semantic role identification models, such as fine-tuned models based on pre-trained language models like BERT or RoBERTa, which learn from a large amount of labeled data to identify semantic roles in sentences. Based on the semantic role recognition results, the semantic relevance between each clause and the interrogative word detection results is calculated. This step aims to quantify the closeness of each clause to the user's core question or intent. One approach is to calculate the semantic relevance by analyzing the frequency and position of arguments (such as "goal" and "reason") related to interrogative words (such as "what," "how," and "why") in the semantic role recognition results, combined with the interrogative word detection results. For example, if a clause's predicate or core argument is highly correlated with the interrogative word detection results, its relevance is considered high. Another approach is to construct a semantic similarity model that compares the semantic role recognition results (e.g., vector representations of predicate-argument pairs) with the interrogative word detection results (e.g., embedding vectors of interrogative words), calculating the cosine similarity or other distance metrics between them as the semantic relevance.

[0060] Building upon this foundation, attention weights are calculated for each sentence based on the attention mechanism, multi-turn dialogue history data, and the semantic relevance of each sentence. The attention mechanism dynamically assigns importance scores to each sentence, comprehensively considering its inherent semantic relevance and its relevance to the overall dialogue context. This can be achieved using the self-attention mechanism in the Transformer architecture, taking the embedded representation of each sentence, the encoded representation of the multi-turn dialogue history data, and the semantic relevance as input, and calculating the attention weight for each sentence through a multi-head attention layer. Alternatively, an attention network based on a recurrent neural network (RNN) or a long short-term memory network (LSTM) can be designed, taking the sentence sequence, dialogue history context vector, and semantic relevance as input, and learning and outputting the weight of each sentence through an attention layer.

[0061] Then, based on sentences whose attention weight exceeds a preset weight threshold, key text for the current input information is generated. This step aims to extract the most core and relevant parts from the user's original input, forming a concise text that represents the user's main intent. One implementation is to set a fixed weight threshold and concatenate all sentences with attention weights higher than that threshold to form the key text. Another implementation is to use a dynamic threshold strategy, for example, selecting the top N sentences in terms of attention weight, or selecting sentences whose cumulative attention weight reaches a certain proportion, as the key text.

[0062] Subsequently, intent classification is performed on the key text to obtain multiple candidate intents and their corresponding probability distributions. Intent classification aims to identify the potential goals or intents expressed by users in the key text. This can be achieved by using deep learning-based text classification models, such as TextCNN, Bi-LSTM, and BERT, to train and predict key texts, outputting probability distributions for predefined intent categories. Alternatively, a hybrid classification approach can be used, combining rule-based and keyword matching classifiers with machine learning models. First, rule matching is used to initially filter intents, followed by refined classification and probability prediction using machine learning models.

[0063] Finally, based on the sentiment polarity label, question type, multiple candidate intentions, and their corresponding probability distributions, the current intention recognition result is generated. This step aims to integrate multiple aspects of information to form a comprehensive and robust intention recognition result. One implementation is to input the sentiment polarity label (e.g., positive, negative, neutral), question type (e.g., interrogative, declarative), and candidate intentions and their probability distributions as features into a multimodal fusion model (e.g., a multilayer perceptron or a small neural network). This model learns how to integrate this information and output the final current intention recognition result. Another implementation is to use post-processing rules. For example, if the sentiment polarity is negative and the question type is interrogative, then intentions related to negative or questioning issues such as complaints or inquiries are prioritized from the candidate intentions, and their probability distributions are adjusted to ultimately determine the current intention recognition result.

[0064] Through the aforementioned technical solution, this application can effectively handle the complex semantic structure and multiple intents in user input information. By segmenting the sentence, identifying semantic roles, calculating semantic relevance, and employing attention mechanisms, it accurately extracts key text from user input, thereby making intent classification more focused and accurate. Furthermore, by combining sentiment polarity tags and question types, the intent recognition results are further refined, ensuring a comprehensive and detailed understanding of user intent. This refined intent recognition process significantly improves the accuracy and robustness of the current intent recognition results, providing high-quality input for subsequent script generation models. This, in turn, enhances the accuracy and relevance of target response scripts and target guidance scripts, thereby optimizing the user experience of the entire script recommendation system.

[0065] In one embodiment, step S104 further includes: Step S1024: In the preset multimodal knowledge graph, retrieve the key text of the current input information and the multimodal business knowledge related to the current intent recognition result as the business knowledge context; Step S1025: Based on the key text, the business knowledge context, and the multi-turn dialogue history data, generate a target message. The target message includes the current input information, information background knowledge, and the comprehensive state of the historical dialogue state. Step S1026: Based on the target message, determine the task target corresponding to the current input information, and generate a script containing the target question answer or operation guidance information based on the task target, as the target response script.

[0066] In this embodiment, the preset multimodal knowledge graph is a knowledge representation structure that integrates multiple modalities such as text, images, audio, and video. It describes entities, concepts, and their relationships through nodes and edges. This knowledge graph can be a triplet knowledge base containing entities, relationships, and attributes, where entities and relationships can associate data from different modalities; or it can be a knowledge network built on a graph database (e.g., Neo4j, JanusGraph), where nodes represent concepts or entities, edges represent relationships between them, and it supports storing pointers to multimodal data. Its function is to provide rich, structured business knowledge, supporting deeper semantic understanding and reasoning. Retrieving the key text of the current input information and the multimodal business knowledge related to the current intent recognition result serves as the business knowledge context, aiming to filter the knowledge most relevant to the current user's needs from a vast knowledge base, providing accurate background information for speech generation. This retrieval process can utilize semantic matching-based retrieval algorithms to vectorize key text and intent recognition results, then perform similarity matching in a multimodal knowledge graph to recall relevant nodes and edges, and extract their associated multimodal information; alternatively, it can employ graph traversal algorithms, using key text and intent recognition results as starting points or query conditions, to perform multi-hop queries in the knowledge graph to obtain knowledge fragments that are semantically or logically closely related to this information, including text descriptions, images, video links, etc. A target message is generated, which includes the current input information, background information, and a comprehensive state of historical dialogue. The purpose is to integrate scattered input information, retrieved knowledge, and historical dialogue context into a unified representation, providing comprehensive context for subsequent task target determination and dialogue generation. The target message can be generated by sequentially concatenating the current input information, background knowledge retrieved from a multimodal knowledge graph (e.g., relevant product parameters, solution steps, frequently asked questions), and historical data from multiple rounds of dialogue (after encoding or summarizing) to form a long text sequence; or by employing multimodal fusion techniques, such as using attention mechanisms or Transformer encoders, to extract and fuse features from information from different sources (current input, background knowledge, and historical dialogue) to generate a high-dimensional vector representation, which represents the comprehensive state of the target message. Determining the task objective corresponding to the current input information aims to clarify the ultimate purpose of the user's current dialogue or the core problem that needs to be solved, guiding the dialogue generation model to generate targeted responses.The determination of the task objective can be based on a pre-trained intent classification model, combined with semantic information in the target message, to map the user intent to a series of predefined task objectives (e.g., "check order status," "process a return," "seek technical support," "understand product features," etc.); or it can utilize a rule-based matching system to trigger corresponding task objective identification rules based on keywords, entities, and intent types identified in the target message, thereby determining the specific task corresponding to the user's current dialogue. Based on the task objective, a response script containing answers to the target question or operational guidance information is generated as the target response script. Its function is to generate a response script that directly solves the user's problem or provides specific operational steps based on the clearly defined task objective, thereby improving the effectiveness and practicality of the response. This script generation process can use the task objective as one of the inputs to the script generation model. Combined with the target message, it can perform conditional generation through a sequence-to-sequence (Seq2Seq) model or a large language model (LLM). The model will adjust the generation strategy according to the task objective, prioritizing the generation of answers or guidance information that are strongly related to the task objective. Alternatively, a template-based script generation system can be built, which can preset multiple script templates for different task objectives. After the task objective is determined, relevant entities and information are extracted from the multimodal knowledge graph and filled into the template to generate structured scripts.

[0067] This application's solution, when generating target response scripts, no longer relies solely on preset business knowledge context, but instead introduces a preset multimodal knowledge graph. This graph can store and associate business knowledge in various forms, such as text, images, and videos. When a user inputs information, the system first retrieves the most relevant multimodal business knowledge from this multimodal knowledge graph, based on the key text of the current input and the current intent recognition result, thereby constructing a more accurate and richer business knowledge context. Subsequently, the system deeply integrates the current input information, the retrieved background information, and multi-turn dialogue history data to generate a comprehensive target message containing all relevant information. Based on this comprehensive target message, the system can more accurately determine the task objective corresponding to the current input information; for example, whether the user wants to query the parameters of a product or needs guidance on a certain operation. Finally, the script generation model generates scripts containing answers to the target questions or operational guidance information, serving as the final target response script. This mechanism ensures that the generated response scripts are not only contextually coherent but also directly address the user's core needs, significantly improving the relevance and practicality of the response.

[0068] Therefore, this embodiment can gain a deeper understanding of user intent and obtain richer and more accurate background knowledge from the multimodal knowledge graph, thereby generating highly relevant and instructive target response scripts. This makes the response content no longer general, but can directly solve the specific problems raised by users or provide clear operational guidance, significantly improving user experience and customer service efficiency. Especially when dealing with complex consultation scenarios or those requiring multi-faceted information support, it can effectively reduce the need for users to ask repeated questions and for agents to provide repetitive explanations.

[0069] Reference Figure 3 , Figure 3 This is a flowchart illustrating another context- and intent-based speech recommendation method provided in this application.

[0070] To further address the above issues, such as Figure 3 As shown, this application embodiment provides a speech recommendation method based on context and intent prediction, wherein step S103 includes: Step S1031: Based on the current intent recognition result, perform a matching query in the intent node index of the event graph to obtain the current intent node; Step S1032: Query the graph database, traverse all outgoing edges from the current intent node, and obtain the subsequent intent node pointed to by each outgoing edge and the transition probability weight attached to the edge. Step S1033: Based on the current user's user profile and historical behavior data, select the node with the highest transition probability from the subsequent intent nodes as the potential subsequent intent.

[0071] In this embodiment, the current intent recognition result is the user intent determined by the system after analyzing the user's current input information. It can be an intent tag, intent ID, or a semantic vector representation of the intent. The event graph is a structured knowledge base containing a large number of intent nodes and their relationships. These intent nodes represent various intents that the user may generate. The intent node index is constructed for efficient retrieval of intent nodes in the event graph, and can be, for example, a hash table, an inverted index, or a dedicated graph index structure. Through matching queries, the system can accurately locate the identified user's current intent to a specific node in the event graph, i.e., the current intent node, thus laying the foundation for subsequent intent prediction.

[0072] By querying a graph database, all outgoing edges from the current intent node are traversed, and the subsequent intent nodes pointed to by each outgoing edge and their associated transition probability weights are obtained. A graph database is a database specifically designed for storing and managing graph-structured data, capable of efficiently handling queries about relationships between nodes and edges. Traversing all outgoing edges from the current intent node means exploring all possible subsequent intents that can evolve from the current intent. Each outgoing edge connects to a subsequent intent node and carries a transition probability weight, which represents the likelihood of transitioning from the current intent to that subsequent intent. These transition probability weights can be obtained based on historical dialogue data statistics, expert experience, or trained through machine learning models; they reflect the general correlation patterns between intents.

[0073] Finally, based on the current user's profile and historical behavior data, the system selects the nodes with the highest transition probability from among the subsequent intent nodes as potential subsequent intents. The user profile is a comprehensive description of a user's characteristics, which may include basic information, preferences, interests, and consumption habits. Historical behavior data records the user's past interactions, such as query history, purchase history, click behavior, and problem-solving status. By combining this personalized information, the system can personalize the prediction results based on the subsequent intent nodes and their transition probability weights obtained in the previous step. For example, if the user profile shows that the user has a strong interest in a specific area, or if historical behavior data indicates that the user has mentioned a related issue multiple times, even if the general transition probability weight of that subsequent intent is not the highest, the system can still prioritize it based on the user's personalized information, thereby selecting the potential subsequent intents that best meet the current user's actual needs.

[0074] This embodiment achieves accurate prediction of potential subsequent intentions by combining the user's current intent with a structured reasoning graph and incorporating personalized user information. Specifically, the system first maps the current intent recognition result to the current intent node in the reasoning graph, placing the user's current intent within a rich knowledge context. Next, leveraging the powerful query capabilities of the graph database, the system explores all possible subsequent intent nodes and their corresponding transition probability weights, starting from the current intent node, thus obtaining a set of intent transition possibilities based on general knowledge. Building upon this, this solution goes beyond general rules, further incorporating the user's profile and historical behavioral data. This personalized information is used to fine-tune or filter the transition probabilities of subsequent intent nodes, ensuring that the final selected potential subsequent intents not only conform to reasoning logic but also align with the user's specific context and personalized needs. This prediction mechanism, combining general knowledge and personalized data, enables the system to more intelligently predict the user's next intention, providing a solid foundation for generating more forward-looking and guiding target-oriented dialogue, significantly improving the intelligence level of dialogue recommendation.

[0075] Therefore, this embodiment overcomes the limitations of relying solely on preset event graphs and current intent recognition results for subsequent intent prediction. By introducing user profiles and historical behavior data to personalize the filtering of subsequent intent nodes, the system can more accurately capture users' unique behavioral patterns and potential needs. This makes the predicted potential subsequent intents not only logically related to the current intent but also highly matched to the user at a personalized level. Therefore, the target guidance dialogue generated based on this potential subsequent intent will be more targeted and forward-looking, effectively guiding users to take the next step or express deeper needs, thereby significantly improving the accuracy of dialogue recommendations, user satisfaction, and overall dialogue efficiency, and reducing the need for users to repeatedly express their intent.

[0076] In one embodiment, step S103 further includes: Step S1034: The multi-turn dialogue history data and the business knowledge context are fused and concatenated to generate a dialogue state representation vector; Step S1035: Encode the potential subsequent intent into an intent embedding vector, and perform potential intent fusion decoding on the intent embedding vector and the dialogue state representation vector to generate a contextual representation for target guidance; Step S1036: Based on the speech generation model, the decoder generates a candidate speech text sequence corresponding to the context representation of the target guidance in an autoregressive manner. Step S1037: Perform fluency verification and context consistency check on the candidate speech text sequence, and determine the target guiding speech in the candidate speech text sequence based on the verification results and consistency check results.

[0077] In this embodiment, multi-turn dialogue history data and business knowledge context are fused and concatenated to generate a dialogue state representation vector. The aim is to integrate information from different sources into a unified numerical representation to comprehensively capture the background information and relevant business knowledge of the current dialogue. This process can be implemented in various ways. For example, pre-trained language models (such as BERT, GPT, etc.) can be used to encode the multi-turn dialogue history data and business knowledge context separately, obtaining their respective vector representations, and then fused through vector concatenation or weighted summation. Alternatively, an attention-based approach can be used, allowing the model to learn how to dynamically allocate weights between dialogue history and business knowledge context based on the importance of the current task, thereby generating a more representative dialogue state representation vector.

[0078] Encoding potential follow-up intentions into intent embedding vectors involves converting the predicted potential follow-up intentions into a numerical vector form that can be processed by machine learning models for subsequent integration with the dialogue context. Specifically, if the potential follow-up intentions are discrete category labels, one-hot encoding can be used, followed by an embedding layer to convert them into dense vectors. Alternatively, a specially trained intent embedding model can be utilized, which maps semantically similar intentions to nearby positions in the vector space, thereby better capturing the semantic relationships between intentions.

[0079] The intent embedding vector and the dialogue state representation vector are fused and decoded to generate a goal-guided contextual representation. This aims to deeply integrate specific potential subsequent intent information with the current dialogue state and business knowledge, forming a comprehensive contextual representation that encompasses both the current context and explicitly points to future intents. This fusion process can employ various techniques. For example, the intent embedding vector and the dialogue state representation vector can be concatenated and then nonlinearly transformed using a feed-forward neural network or a Transformer layer to learn the complex interactions between them. Alternatively, a gated fusion mechanism can be used, such as the gating unit used in Long Short-Term Memory (LSTM) networks or Gated Recurrent Units (GRUs), to control how intent information is selectively incorporated into the dialogue state, thereby generating a more instructive contextual representation.

[0080] Based on a speech generation model, candidate speech text sequences corresponding to the contextual representation of the target prompt are generated through an autoregressive mechanism using a decoder. This involves using a pre-trained or fine-tuned generative model to progressively generate multiple possible prompt texts based on the fused contextual representation. This process is typically handled by the model's decoder. For example, a Transformer-based decoder (such as the decoder in the GPT series models) can be used, which predicts the next word based on the preceding and currently generated words and generates a complete text sequence word by word in an autoregressive manner. Alternatively, a Recurrent Neural Network (RNN)-based decoder, such as LSTM or GRU, can be employed. These decoders process sequence information by maintaining internal states and progressively generate text. During the generation process, strategies such as Beam Search or Top-k / Top-p sampling can be combined to generate diverse and high-quality candidate sequences.

[0081] Fluency and contextual consistency checks are performed on candidate speech text sequences to assess the quality of the generated speech, ensuring it conforms to linguistic norms and remains logically connected to the dialogue context. Fluency checks can utilize language models to calculate the text's perplexity to evaluate its grammatical correctness and naturalness, or integrate grammar checking tools to detect syntactic and lexical errors. Contextual consistency checks can be performed by calculating the semantic similarity between candidate speech and multi-turn dialogue history data and business knowledge context (e.g., using models like Sentence-BERT to calculate vector similarity), or by using Natural Language Inference (NLI) models to determine whether candidate speech contradicts or implies any relationship with the context.

[0082] Determining the target guiding dialogue from candidate dialogue text sequences based on the verification and consistency check results means selecting the highest quality and most suitable dialogue from multiple candidate sequences as the final target guiding dialogue recommended to the user, according to the aforementioned evaluation results. Specifically, a scoring function can be designed to comprehensively consider fluency score, contextual consistency score, and other relevant indicators (such as intent matching degree) to calculate a comprehensive score for each candidate sequence, and then select the sequence with the highest score. Alternatively, a ranking model can be trained, which takes various evaluation features of candidate dialogues as input and directly outputs the ranking of candidate sequences, thereby selecting the optimal guiding dialogue.

[0083] Therefore, this embodiment effectively integrates multi-turn dialogue history data, business knowledge context, and predicted potential subsequent intentions, and performs rigorous quality control on the generated script text sequence. This ensures that the generated target guidance script is not only highly relevant to the current dialogue context and business knowledge semantically, but also fluent, natural, and grammatically correct. Compared to directly generating guidance scripts based solely on the script generation model, this solution significantly improves the accuracy, usability, and user experience of the guidance scripts through refined data fusion, intent embedding, context representation generation, and subsequent verification and checking mechanisms. This enables the system to more accurately predict user needs and proactively guide the dialogue, thereby improving user satisfaction and dialogue efficiency, and effectively solving the technical challenge of generating high-quality, highly relevant target guidance scripts in complex dialogue scenarios.

[0084] Please see Figure 4 , Figure 4 A schematic diagram of the functional modules of a context- and intent-based speech recommendation device provided in this application.

[0085] like Figure 4 As shown, the context- and intent-based speech recommendation device 400 includes: The context acquisition module 401 is used to acquire the current input information of the current user and the multi-turn dialogue history data between the current user and the agent, wherein the multi-turn dialogue history data includes historical input information and its corresponding agent response information; The response script generation module 402 is used to perform intent recognition on the current input information, obtain the current intent recognition result, and generate a target response script corresponding to the preset business knowledge context, the current input information, and the multi-round dialogue history data based on the script generation model. The guidance dialogue generation module 403 is used to predict the potential subsequent intent of the current user based on the preset reasoning graph and the current intent recognition result, and generate the target guidance dialogue corresponding to the multi-turn dialogue history data, the business knowledge context and the potential subsequent intent based on the dialogue generation model. The customer service script recommendation module 404 is used to obtain recommended scripts for the current user based on the target reply script and / or the target guidance script.

[0086] It should be noted that those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the above-described apparatus and modules can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0087] The aforementioned apparatus can be implemented as a computer program, which can be used in, for example... Figure 5 It runs on the computer device shown.

[0088] Please see Figure 5 , Figure 5 This application provides a schematic block diagram of the structure of a computer device. The computer device may be a server.

[0089] See Figure 5 The computer device includes a processor, memory, and network interface connected via a system bus, wherein the memory may include non-volatile storage media and internal memory.

[0090] Non-volatile storage media can store operating systems and computer programs. These computer programs include program instructions that, when executed, cause the processor to perform any context- and intent-based speech recommendation method.

[0091] The processor provides computing and control capabilities, supporting the operation of the entire computer device.

[0092] Internal memory provides an environment for the execution of computer programs in non-volatile storage media. When executed by a processor, the computer program enables the processor to perform any discourse recommendation method based on context and intent prediction.

[0093] This network interface is used for network communication, such as sending assigned tasks. Those skilled in the art will understand that... Figure 5 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0094] It should be understood that the processor can be a Central Processing Unit (CPU), or it can be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor. The processor executes the program instructions to implement any of the context- and intent-prediction-based speech recommendation methods provided in the embodiments of this application.

[0095] The embodiments of this application also provide a computer-readable storage medium storing a computer program, the computer program including program instructions, and the processor executing the program instructions to implement any of the context- and intent-based speech recommendation methods provided in the embodiments of this application.

[0096] The computer-readable storage medium may be an internal storage unit of the computer device described in the foregoing embodiments, such as the hard disk or memory of the computer device. The computer-readable storage medium may also be an external storage device of the computer device, such as a plug-in hard disk, SmartMedia Card (SMC), Secure Digital (SD) card, or Flash Card equipped on the computer device.

[0097] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this application, and these modifications or substitutions should all be covered within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A speech recommendation method based on context and intent prediction, characterized in that, The script recommendation method includes the following steps: Obtain the current input information of the current user and the multi-turn dialogue history data between the current user and the agent, wherein the multi-turn dialogue history data includes historical input information and its corresponding agent response information; The current input information is subjected to intent recognition to obtain the current intent recognition result, and a target response script is generated based on the script generation model, which includes a preset business knowledge context, the current input information, and the multi-turn dialogue history data. Based on the preset reasoning graph and the current intent recognition result, the potential subsequent intent of the current user is predicted, and based on the speech generation model, the multi-turn dialogue history data, the business knowledge context, and the target guiding speech corresponding to the potential subsequent intent are generated. Based on the target response script and / or the target guidance script, the recommended script for the current user is obtained.

2. The script recommendation method as described in claim 1, characterized in that, Before generating the preset business knowledge context, the current input information, the current intent recognition result, and the target response dialogue corresponding to the multi-turn dialogue history data based on the dialogue generation model, the method further includes: The current input information is subjected to a semantic validity check to obtain a first check result. Based on the first check result, the sentiment intensity and question type of the current input information are determined. The first check result includes a completeness score, sentiment polarity label, and question word detection result. Based on the first inspection result, the current intent recognition result, and the current input information, a message feature vector of the current user is generated; Obtain the current customer service system load, and based on the routing decision model, generate the selection probability distributions of the general large model, the domain fine-tuning model, and the small parameter model under the current customer service system load and the message feature vector; When the current customer system load is lower than a preset load threshold, the model with the highest probability of the selection probability distribution is used as the script generation model; When the current customer system load is not lower than a preset load threshold, the small parameter model is used as the script generation model.

3. The script recommendation method as described in claim 2, characterized in that, The step of generating the message feature vector of the current user based on the current intent recognition result and the current input information includes: Extract message length features and key entity counts from the current input information, and parse the current user's intent complexity and intent category from the current intent recognition result; The message feature vector is generated based on the emotional intensity, question type, message length features, number of key entities, intent complexity, and intent category.

4. The script recommendation method as described in claim 2, characterized in that, The step of performing intent recognition on the current input information to obtain the current intent recognition result includes: The current input information is segmented into sentences to obtain each sentence. Semantic role recognition is performed on each sentence to obtain the semantic role recognition result of each word in each sentence. Based on the semantic role recognition result, the semantic relevance between each sentence and the interrogative word detection result is calculated. Based on the attention mechanism, the multi-turn dialogue history data, and the semantic relevance of each sentence, the attention weight of each sentence is calculated, and the key text of the current input information is generated based on the sentences whose attention weight exceeds the preset weight threshold. The key text is classified into intents to obtain multiple candidate intents and their corresponding probability distributions. Based on the sentiment polarity label, the question type, the multiple candidate intents and their corresponding probability distributions, the current intent recognition result is generated.

5. The script recommendation method as described in claim 1, characterized in that, The prediction of the potential subsequent intentions of the current user based on the preset event graph and the current intention recognition result includes: Based on the current intent recognition result, a matching query is performed in the intent node index of the event graph to obtain the current intent node; By querying the graph database, all outgoing edges from the current intent node are traversed, and the subsequent intent nodes pointed to by each outgoing edge and the transition probability weights attached to the edges are obtained. Based on the current user's profile and historical behavior data, the node with the highest probability of transition is selected from the subsequent intent nodes as the potential subsequent intent.

6. The script recommendation method as described in claim 1, characterized in that, The step of generating the multi-turn dialogue history data, the business knowledge context, and the target guiding dialogue corresponding to the potential subsequent intent based on the dialogue generation model includes: The multi-turn dialogue history data is fused and concatenated with the business knowledge context to generate a dialogue state representation vector. The potential subsequent intent is encoded into an intent embedding vector, and the intent embedding vector is fused and decoded with the dialogue state representation vector to generate a contextual representation for target guidance; Based on the aforementioned dialogue generation model, a decoder generates a sequence of candidate dialogue texts corresponding to the contextual representation guided by the target in an autoregressive manner. The candidate speech text sequence is subjected to fluency verification and context consistency check, and the target guiding speech is determined from the candidate speech text sequence based on the verification results and consistency check results.

7. The script recommendation method as described in any one of claims 1-6, characterized in that, The step of generating a target response script based on a script generation model, which includes a preset business knowledge context, the current input information, and the multi-turn dialogue history data, includes: In a preset multimodal knowledge graph, key text of the current input information and multimodal business knowledge related to the current intent recognition result are retrieved as the business knowledge context; Based on the key text, the business knowledge context, and the multi-turn dialogue history data, a target message is generated. The target message includes the current input information, information background knowledge, and the comprehensive state of the historical dialogue. Based on the target message, the task target corresponding to the current input information is determined, and a script containing the answer to the target question or operation guidance information is generated based on the task target as the target response script.

8. A speech recommendation device based on context and intent prediction, characterized in that, The script recommendation device includes: The context acquisition module is used to acquire the current input information of the current user and the multi-turn dialogue history data between the current user and the agent. The multi-turn dialogue history data includes historical input information and its corresponding agent response information. The response script generation module is used to perform intent recognition on the current input information, obtain the current intent recognition result, and generate a target response script corresponding to the preset business knowledge context, the current input information, and the multi-turn dialogue history data based on the script generation model. The guidance dialogue generation module is used to predict the potential subsequent intent of the current user based on a preset reasoning graph and the current intent recognition result, and to generate the target guidance dialogue corresponding to the multi-turn dialogue history data, the business knowledge context, and the potential subsequent intent based on the dialogue generation model. The customer service script recommendation module is used to obtain recommended scripts for the current user based on the target response script and / or the target guidance script.

9. A computer device, characterized in that, The computer device includes a processor, a memory, and a context- and intent-based speech recommendation program stored in the memory and executable by the processor, wherein when the context- and intent-based speech recommendation program is executed by the processor, it implements the steps of the context- and intent-based speech recommendation method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a context- and intent-based speech recommendation program, wherein when the context- and intent-based speech recommendation program is executed by a processor, it implements the steps of the context- and intent-based speech recommendation method as described in any one of claims 1 to 7.

Citation Information

Cited By

  • A user intention mining and active recommendation method for a financial knowledge middle platform

    CN122196144A