Context intention recognition-based instruction chain generation method

Through contextual intent recognition and real-time emotion perception models, combined with explicit and implicit intent networks, a command chain is generated, which solves the problem of inaccurate intent recognition in existing intelligent conversational systems and achieves efficient and accurate understanding of user needs and high-risk event handling.

CN120724262AActive Publication Date: 2025-09-30JIANGSU ELECTRIC POWER INFORMATION TECH

Patent Information

Application Number
CN202511197091.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-26
Publication Date
2025-09-30
Estimated Expiration
2045-08-26

AI Technical Summary

Technical Problem

Existing intelligent conversational systems find it difficult to accurately understand users' potential needs and emotional changes, and lack full utilization of multi-round interactions, resulting in inaccurate intent recognition, inaccurate responses, and difficulty in ensuring the efficiency of high-risk event handling and service quality.

Method used

Through context-based intent recognition methods, combined with explicit and implicit intent recognition networks, we generate command chains, real-time emotion perception models, and event risk grading mechanisms, dynamically capture user input information, match appropriate customer service personnel, and ensure that high-risk events are handled manually in a timely manner.

Benefits of technology

It improves the accuracy and processing efficiency of intent recognition, enhances the understanding and response effects of multi-round conversations, ensures timely handling of high-risk events, reduces manual intervention, and improves user satisfaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120724262A_ABST
    Figure CN120724262A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence session processing, and discloses a context intention recognition-based instruction chain generation method, which comprises the following steps of: extracting an event type from a user session request, determining an event risk level, establishing an intelligent session window for a low-risk or medium-risk event, and generating a context intention recognition instruction chain; dynamically capturing user input and context information to form basic session features, and inputting the basic session features into a dominant intention recognition network to obtain first intention information and confidence; generating and displaying a reply based on the first intention, and performing real-time scoring through an emotion perception model; when the emotion or intention confidence of the user is lower than a threshold value, constructing an enhanced session feature and inputting the enhanced session feature into the implicit intention recognition network to obtain second intention information; and finally, generating an ordered instruction chain in combination with the first intention and the second intention. Therefore, accurate intention recognition, intelligent reply and automatic business operation in multiple rounds of sessions are realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of artificial intelligence conversation processing, and in particular to a method for generating an instruction chain based on contextual intent recognition. Background Art

[0002] With the rapid development of artificial intelligence and natural language processing technologies, conversation-based intelligent interaction systems have been widely used in scenarios such as State Grid's customer service, intelligent question-and-answer (Q&A), and task execution. In existing technologies, intelligent conversation systems typically identify the intent of user input text to match corresponding business operations or generate responses. However, existing intelligent conversation systems often suffer from the following technical issues: First, existing intelligent conversational systems struggle to accurately understand users' underlying needs and emotional shifts. They often fail to fully grasp user intent during multiple rounds of interaction, leading to inaccurate responses and inefficient processing. Second, existing intelligent conversations typically rely solely on single-turn user input for intent recognition, lacking full utilization of multi-turn context. This also prevents the effective allocation of appropriate human agents to high-risk incident handling. This results in inaccurate intent recognition, irrational instruction generation, and difficulty ensuring high-risk incident handling efficiency and service quality. Third, existing technologies evaluate user emotions through single-modal information, resulting in inaccurate emotion judgments, which in turn affects timely responses and accurate judgments of user status during intelligent conversations. Summary of the Invention

[0003] This summary is intended to briefly introduce concepts that will be described in detail in the detailed description below. This summary is not intended to identify key features or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.

[0004] The present invention proposes a method for generating an instruction chain based on contextual intent recognition to solve one or more of the technical problems mentioned in the above background technology section.

[0005] The present invention provides a method for generating an instruction chain based on contextual intent recognition, comprising: upon receiving a user session request, extracting an event type from the user session request, and determining an event risk level according to the event type; If the event risk level is low or medium, an intelligent conversation window is established and the context data corresponding to the user's current input information is dynamically captured to generate basic conversation features. The basic conversation features are input into the explicit intent recognition network to obtain the first intent information and the first confidence level. Based on the first intent information, generate reply information for the user's question and display it in the intelligent conversation window; During the conversation, the user's emotions are scored in real time through the emotion perception model; When the user's emotion score is lower than the preset score threshold; or the first confidence is less than the preset threshold and the current conversation round in the intelligent conversation window reaches the first preset round threshold, an enhanced conversation feature is constructed and input into the implicit intent recognition network to obtain the second intent information; Based on the first intention information and the second intention information, a command chain is generated.

[0006] Optionally, construct enhanced session features and input them into the implicit intent recognition network, including: The user emotion score and user attribute information are extracted to obtain user features; the user features are spliced ​​with the current basic session features to obtain enhanced session features and input them into the implicit intent recognition network.

[0007] Optionally, the current basic session features are extracted through the following steps: Extracting session data within a preset time window within the intelligent session window and performing a redundancy removal operation on the session data to obtain session data after redundancy removal; If the number of complete conversation rounds included in the conversation data after redundancy removal is less than a second preset turn threshold, updating the time window based on the preset time window until the number of non-redundant complete conversation rounds included in the updated time window is greater than or equal to the second preset turn threshold; The conversation data within the update time window is word-embedded to obtain the current basic conversation features.

[0008] Optionally, generating a command chain based on the first intent information and the second intent information includes: The first intention information and the second intention information are matched in a pre-set set of operation items to obtain a plurality of matching operation items; wherein each operation item corresponds to a set of instructions; According to the dependency relationship and time sequence between the instructions, multiple groups of instructions corresponding to multiple operation items are sorted to form an instruction chain.

[0009] Optionally, the method for generating a command chain based on contextual intent recognition of the present invention further includes: If the event risk level is high, a manual conversation window is created and the corresponding customer service representative is matched to the user conversation request. The manual conversation window is used to display the conversation between the customer service representative and the user. Among them, customer service personnel are matched through the following steps: Match the corresponding customer service personnel for the user session request based on user attribute information or historical session records.

[0010] Optionally, match the user's session request with a corresponding customer service representative based on user attribute information or historical session records, including: If the query finds multiple historical conversation records for the user, sort them in descending order based on the user's feedback score for each historical conversation record, and select the customer service representative with the highest feedback score as the candidate customer service representative; if the candidate customer service representative's current status is available, then the candidate customer service representative is determined to be a matching customer service representative; If the query finds that the user has no historical conversation record, the matching customer service personnel will be selected from the currently available customer service personnel based on the user's attribute information; If the query finds that the user has historical conversation records, but the user has not rated the historical conversation records, the customer service staff with the highest comprehensive customer service score is screened from multiple customer service staff corresponding to the historical conversation records and determined as the matching customer service staff.

[0011] Optionally, user attribute information includes user geographic location and account type, as well as Filter matching agents from currently available agents based on user attribute information, including: Determine the corresponding customer service information based on the customer service personnel whose current status is idle, where the customer service information includes the customer service ID and the skill tag corresponding to the customer service ID; According to the user's geographic location, regional matching is performed from customer service personnel who are currently available to obtain at least one first customer service personnel to be selected; If there are multiple first candidate customer service personnel, match the first candidate customer service personnel with the services according to the account type and select at least one second candidate customer service personnel; If there are multiple second customer service personnel to be selected, the skill tags of the customer service personnel are matched with the event type selected by the user to obtain at least one target customer service personnel; and the target customer service personnel is determined as the matching customer service personnel.

[0012] Optionally, the overall customer service score is determined by the following steps: Collect historical service data for each of the multiple customer service representatives, including historical response time, problem resolution rate, customer satisfaction score, and complaint rate. Perform weighted calculation on each indicator in the historical service data according to the preset weights to obtain the weighted total score of each customer service staff; the weighted total score is used as the comprehensive customer service score.

[0013] The present invention has the following beneficial effects: 1. Improve intent recognition accuracy and processing efficiency. Specifically, through explicit and implicit intent recognition networks, it not only understands explicit user needs but also captures latent or implicit needs, improving multi-turn conversation comprehension. It also provides real-time emotion perception and feedback mechanisms, enabling policy adjustments when users express negative emotions to improve satisfaction. Command chain generation logic automatically sorts and schedules action items, enabling fast and accurate business processing and reducing manual intervention. An event risk grading mechanism ensures that high-risk events are promptly diverted to human customer service, while low- and medium-risk events are routed intelligently, optimizing resources.

[0014] 2. Improve the understanding and response of multi-round conversations, and ensure that high-risk events can be handled promptly and reasonably. Specifically, through redundant deletion, time window update and word embedding, multi-round conversation data is converted into streamlined and sufficient basic conversation features to ensure the accuracy of explicit and implicit intent recognition. Match multi-round intention information with a set of preset operation items, and generate instruction chains based on dependencies and time sequence to achieve automated and accurate execution of multi-step business operations. Through user attribute information and historical conversation records, combined with feedback scores or comprehensive scores, intelligent matching of idle customer service personnel is carried out, so that high-risk events can be handled promptly and accurately, while reducing the blindness of manual intervention and improving service efficiency and user satisfaction.

[0015] 3. Improve the accuracy and real-time performance of emotion recognition. Specifically, text, facial, and voice features are extracted using a semantic emotion recognition model, a facial expression recognition model, and a voice emotion analysis model, respectively. Pre-trained models are used to ensure accurate feature extraction. Emotional information from the three modalities is quantified into comparable scores (facial emotion score, voice emotion score, and semantic emotion score), providing a foundation for fusion. The three emotion scores are fused together using a weighted summation with preset weights or feature concatenation input into the fusion model to obtain the final user emotion score. This fusion method can adaptively adjust the importance of different modalities, improving the accuracy and robustness of emotion assessment. The resulting user emotion score can be used in the intelligent conversation window to dynamically adjust response strategies or trigger enhanced intent recognition, thereby improving interaction quality and user satisfaction. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] The above and other features, advantages, and aspects of the various embodiments of the present invention will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and that the elements are not necessarily drawn to scale.

[0017] Figure 1 It is a flow chart of a method for generating a command chain based on contextual intention recognition according to the present invention; Figure 2It is a flowchart of creating an artificial conversation window in a method for generating a command chain based on contextual intention recognition according to the present invention; Figure 3 This is a matching flow chart of customer service personnel in a command chain generation method based on contextual intent recognition of the present invention. DETAILED DESCRIPTION

[0018] The present invention will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present invention are shown in the accompanying drawings, it should be understood that the present invention can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the present invention. It should be understood that the drawings and embodiments of the present invention are for illustrative purposes only and are not intended to limit the scope of protection of the present invention.

[0019] It should also be noted that, for ease of description, only the parts related to the invention are shown in the drawings. In the absence of conflict, the embodiments and features of the embodiments of the present invention may be combined with each other.

[0020] It should be noted that the concepts of "first" and "second" mentioned in the present invention are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.

[0021] It should be noted that the modifications of "one" and "multiple" mentioned in the present invention are illustrative rather than restrictive. Those skilled in the art should understand that unless otherwise clearly indicated in the context, it should be understood as "one or more".

[0022] The names of the messages or information exchanged between multiple devices of the present invention are only used for illustrative purposes and are not used to limit the scope of these messages or information.

[0023] The present invention will be described in detail below with reference to the accompanying drawings and in conjunction with embodiments.

[0024] like Figure 1 FIG. 1 shows a flow chart of a method for generating an instruction chain based on contextual intent recognition according to the present invention, which specifically includes the following steps: Step 101: upon receiving a user session request, extracting an event type from the user session request and determining an event risk level based on the event type; In some embodiments, the execution entity of a method for generating a command chain based on contextual intent recognition according to the present invention is a backend server. Upon receiving a user session request, the backend server receives the event type selected by the user in the session interface. Based on the event type, the backend server queries a risk grading table pre-stored on the backend server for the corresponding risk level. The risk grading table includes event types and corresponding event risk levels. For example, if the event type is "repair report," the corresponding event risk level is medium. If the event type is "query," the corresponding event risk level is low. A user session request refers to an interaction request initiated by a user in the session interface. The session interface provides several preset event type buttons, and the user actively selects the corresponding event type during the session. The event type refers to the business processing category label selected by the user, which is used to quickly clarify the business context and processing priority of the session. Event types include, but are not limited to, inquiries, such as electricity bill inquiries and electricity usage inquiries; complaints, such as service attitude complaints and billing disputes; repair reports, such as meter damage and line outages; and emergency reports, such as transformer sparks and distribution room fires. Event risk levels refer to the preset risk levels for different event types and are used to determine handling strategies and diversion paths.

[0025] Step 102: If the event risk level is low or medium, establish an intelligent conversation window and dynamically capture context data corresponding to the user's current input information to generate basic conversation features. Input the basic conversation features into the explicit intent recognition network to obtain first intent information and a first confidence level. In some embodiments, when the event risk level is low or medium, the backend server creates an intelligent conversation window on the user terminal to display the content of multiple rounds of conversation interactions with the user and support real-time response to user input. During this process, the backend server dynamically captures the user's current input and its related contextual data. Specifically, during the conversation, the backend server automatically and in real time obtains the conversation history, user attribute information, and service status information related to the current user input. This dynamically captured contextual data is then cleaned and encoded along with the user's current input to generate basic conversation features. The backend server then inputs these basic conversation features into an explicit intent recognition network for inference. The explicit intent recognition network utilizes a relatively simple and computationally efficient natural language understanding model, making it suitable for quickly identifying the user's direct intent in real-time conversation scenarios. For example, a TextCNN (a text classification model based on a convolutional neural network) can be used. After word segmentation and vectorization of the input text, this model uses a lightweight network structure to extract and classify features. A softmax classification layer is used at the output layer to convert the results into a probability distribution for each intent category, thereby generating first intent information and a corresponding first confidence score. For example, the model outputs a first intent (e.g., "electricity bill inquiry") and a corresponding first confidence score (e.g., 0.92). The intelligent conversation window refers to the interactive interface displayed on the user terminal (e.g., mobile app, website). This interface supports real-time information exchange with the backend server and dynamically displays responses based on user input. The intelligent conversation window can load contextual data, display automatically generated responses, and provide multi-turn dialogue. Contextual data refers to historical conversation content, user attributes, and service status information related to the user's current input. This primarily includes previous messages and replies sent by the user in the current conversation, as well as necessary historical service records (e.g., payment history, repair history, etc.). Basic conversation features refer to the vectorized representation of the user's current input and its contextual data, obtained through feature extraction. These features serve as input to the explicit intent recognition network to determine the user's direct intent. The explicit intent recognition network is a deep learning-based natural language understanding model designed to identify direct and explicit intent in user input. Explicit intent can typically be directly determined through textual expressions, such as "I want to check my electricity bill" or "I want to file a complaint." The first intent information refers to the classification result of the user's current input intent, output by the explicit intent recognition network. Examples include "electricity bill inquiry," "equipment repair," and "account top-up." The first confidence level refers to the explicit intent recognition network's confidence in the first intent information. It typically ranges from 0 to 1, with higher values ​​indicating greater confidence in the correct intent classification. The Softmax classification layer is a commonly used output layer in deep learning and is used for multi-classification tasks. Its core function is to convert the output of the neural network into a probability distribution for each category.TextCNN is a text classification model commonly used in natural language processing (NLP) and is based on the convolutional neural network (CNN) architecture.

[0026] Step 103: Based on the first intent information, generate a reply message to the user's question and display it in the smart conversation window; Step 104: During the conversation, the user's emotions are scored in real time using the emotion perception model; In some embodiments, the execution entity matches a response to the first intent message in a predefined response template library. For example, if the first intent message is "Electricity bill inquiry," the corresponding response in the template library is "Your July electricity bill is 325 yuan, and the payment deadline is August 31st." The generated response message is sent to the user through the intelligent conversation window, completing the information display. As an example, if the first intent message is "Report an electricity meter repair," the backend server matches the template or generates a dynamic response: "The electricity meter failure in your area has been registered. We will arrange for personnel to visit and address it as soon as possible. Please keep your phone open." During the conversation, the user's latest input text is captured, tokenized, and vectorized. This is then fed into a trained emotion perception model to output a semantic emotion score. The emotion perception model can be an LSTM (Long Short-Term Memory) emotion classification model. Based on this, the user's latest input text is fed into the LSTM emotion classification model. The output hidden state passes through a fully connected layer and a softmax classification layer to obtain the probability distribution of each emotion category (e.g., positive, negative, neutral). The user's emotion score can be the probability value of the corresponding category or combined with a weighted method to generate a continuous score. A reply message is a targeted text or multimedia message generated based on user intent, answering a user's question or guiding the next action. It can be a predefined template. An emotion perception model is an AI-based analysis model used to determine and quantify a user's emotional state during a conversation. It can identify emotions based on multimodal information such as text sentiment analysis, voice intonation, and facial expression recognition. The user emotion score is a quantitative indicator output by the emotion perception model that indicates the degree of positivity, negativity, or intensity of the user's current emotional state. It typically ranges from 0 to 1, with lower values ​​indicating more negative emotions or dissatisfaction. LSTM is a commonly used recurrent neural network variant that excels at processing sequential data and capturing long-range dependencies. In text sentiment analysis, LSTM can understand the relationship between words in a sentence to determine the overall emotional tendency. For example, a user inputs: "My electric meter is broken, why haven't you fixed it yet?!" The text sentiment analysis model determines that the text has a negative sentiment and outputs a user emotion score of 0.3 (below the preset threshold of 0.5), indicating that the user's emotions are relatively intense.

[0027] Step 105: When the user emotion score is lower than a preset score threshold; or the first confidence level is lower than a preset threshold and the current conversation round in the intelligent conversation window reaches the first preset round threshold, an enhanced conversation feature is constructed and input into the implicit intent recognition network to obtain second intent information; Step 106: Generate a command chain based on the first intention information and the second intention information.

[0028] In some embodiments, when a user's emotion score falls below a preset threshold, or when a first confidence level falls below a preset threshold and the current session turn in the intelligent session window reaches the first preset turn threshold, feature extraction is performed on the user's emotion score and user attribute information to generate a unified user feature vector. The user features are then concatenated with the current basic session features to generate enhanced session features, which are then fed into the implicit intent recognition network. Concatenation is a vector-level operation that directly concatenates the user feature vector and the current basic session feature vector along a single dimension to form a new vector, the enhanced session feature vector. User attribute information refers to basic user information and service-related information, such as account type, historical repair records, user location, and meter number. Feature extraction involves converting the raw information (user emotion score and user attribute information) into a numerical vector or high-dimensional feature representation for neural network processing. For example, user attributes can be encoded as numerically normalized values, with the emotion score used directly as a numerical feature. Enhanced session features concatenate user features with basic session features to form a new feature vector that incorporates both the user's semantic context and emotional and personalized information for implicit intent recognition. The implicit intent recognition network is a natural language understanding model based on deep learning. For example, it builds on the BERT pre-trained model with a multi-layer attention mechanism or a deep Transformer stack, and adds a Softmax classification layer to the output layer. This allows for inference within a longer context and multi-dimensional features. Compared to the explicit intent recognition network, the implicit network has a more complex structure and slower inference speed, but it can capture latent needs or unexpressed intent beyond the explicit expression. It takes in enhanced conversational features and outputs secondary intent information. Secondary intent information refers to the intent category output by the implicit intent recognition network and is used to assist in generating more accurate or comprehensive instruction chains. The preset score threshold is a pre-set numerical limit for determining whether user emotions are negative or dissatisfied. If the user emotion score falls below this threshold, it indicates that the user's emotions may be intense or dissatisfied, and further action is required (such as triggering implicit intent recognition or manual customer service intervention). The preset threshold is a confidence limit set to determine the reliability of the first intent prediction output by the explicit intent recognition network. When the first confidence score falls below this threshold, it indicates that the explicit intent recognition may be uncertain, and further confirmation of user intent is required through enhanced conversational features or implicit intent recognition. The current conversation turn refers to the number of interaction rounds between the user and the backend server in the intelligent conversation window (each round includes user input and response). It is used to determine whether the conversation depth or rounds have reached the conditions that require further processing. The first preset turn threshold is the pre-set conversation turn limit to determine whether enhanced intent recognition is still needed after multiple rounds of dialogue.BERT (Bidirectional Encoder Representations from Transformers) is a well-known pre-trained language model in natural language processing. Based on the Transformer encoder architecture, it employs a bidirectional attention mechanism to simultaneously consider the dependencies between previous and subsequent words in context, generating word or sentence vectors that contain rich semantic information. In specific tasks, BERT can be fine-tuned by adding a task-specific classifier (such as a softmax classification layer) to its output layer for applications such as intent recognition, sentiment analysis, and question answering. The Transformer is a neural network architecture based on a self-attention mechanism, consisting of an encoder and a decoder. The softmax classification layer is a common output layer in deep learning models, typically used for multi-classification tasks. Its function is to map the model's output real-valued vector into a probability distribution, with each class corresponding to a probability value. In intent recognition tasks, the softmax classification layer converts the input feature vector (such as the BERT encoding result) into predicted probabilities for each intent class, thereby obtaining the most likely intent class and its corresponding confidence score.

[0029] Based on this, the first and second intents are matched against the action item library to obtain the corresponding action items and instruction sets. These instructions are sorted based on their dependencies and priorities to form an instruction chain. An instruction chain is a sequence of action instructions generated based on user intent, used for backend business processing or automated execution. Each instruction corresponds to a specific action item (e.g., querying bills, generating work orders, assigning staff).

[0030] These embodiments improve intent recognition accuracy and processing efficiency. Specifically, through explicit and implicit intent recognition networks, not only can explicit user needs be understood, but also potential or implicit needs can be captured, improving the ability to understand multiple rounds of conversations. A real-time emotion perception and feedback mechanism is also provided, enabling policy adjustments when users express negative emotions, thereby improving satisfaction. Furthermore, command chain generation logic can automatically sort and schedule action items, enabling fast and accurate business processing and reducing manual intervention. An event risk grading mechanism ensures that high-risk events are promptly diverted to human customer service, while low- and medium-risk events are routed intelligently, achieving resource optimization.

[0031] In some embodiments, to further address the second technical problem described in the background technology section, namely, "Existing intelligent conversations typically rely solely on a single round of user input for intent recognition, lacking full utilization of multi-round context, and failing to effectively allocate appropriate human customer service personnel to high-risk event handling, resulting in inaccurate intent recognition, irrational instruction generation, and difficulty ensuring high-risk event handling efficiency and service quality." In some embodiments of the present invention, current basic conversation features are extracted through the following steps: Step 1: extracting session data within a preset time window in the intelligent session window and performing a redundancy removal operation on the session data to obtain redundancy-removed session data; In some embodiments, duplicate or invalid information is removed from conversation data to avoid noise interference, improve the quality of input data, and ensure more accurate and efficient feature extraction results. When the number of conversation turns is insufficient, context is supplemented by expanding the time window or incorporating historical messages to ensure sufficient contextual information in the generated conversation features, thereby enhancing semantic integrity. Pre-trained models or deep learning methods are used to convert text into high-dimensional semantic vector representations, which accurately represent the semantics of user input and context, providing more effective input for intent recognition. Specifically, the current basic conversation feature is a numerical vector extracted from the intelligent conversation window. It represents the user's current input and multi-turn conversation context information and serves as the input basis for explicit and implicit intent recognition. The preset time window means that in a multi-turn conversation, in order to extract context, the entire conversation is not retained indefinitely. Instead, conversation data is extracted within a limited time range. The time window can be a fixed duration (e.g., conversations within the past 5 minutes), a fixed number of turns (e.g., the last 5 conversations), or a fixed number of messages (e.g., the last 10 messages). Based on this, a window size is set, such as "last 5 conversations." Each time a user inputs, the backend server extracts only the most recent five rounds of conversation data as context input. It then performs redundancy removal (such as removing repeated questions, irrelevant short conversations, or merging similar statements) to obtain a more refined conversation context. Session data refers to the interaction information occurring within the intelligent conversation window, including user input and the backend server's responses. Redundancy removal involves removing repetitive, useless, or irrelevant content from the conversation to avoid information redundancy during model processing, which can affect feature quality. Examples include deleting repeated user input (such as "Are you there?", "Are you there?"), removing stop words (such as "ah" and "hmm"), and filtering out irrelevant short conversations (such as "The weather is nice"). This processed, more concise and valuable session data is used for subsequent feature extraction, preserving core semantic information relevant to the task.

[0032] Step 2: If the number of complete conversation turns included in the conversation data after redundancy removal is less than a second preset turn threshold, the time window is updated based on the preset time window until the number of non-redundant complete conversation turns included in the updated time window is greater than or equal to the second preset turn threshold; Step 3: Perform word embedding on the conversation data within the update time window to obtain the current basic conversation features.

[0033] In some embodiments, if the number of complete conversation turns after redundancy removal falls below a second preset turn threshold (e.g., the most recent three turns), the backend server can expand the search scope of historical conversations in the time dimension (e.g., if only two valid conversation turns were identified within the last five minutes, the search can be expanded to 10 minutes, thereby adding three complete conversation turns and meeting the second preset turn threshold). Alternatively, the number of retrieved historical messages can be increased in the message dimension (e.g., if the user's most recent five conversation turns contain only two valid business conversation turns, the search can be expanded to the most recent ten turns, thereby adding three complete conversation turns). This ensures that the final number of non-redundant complete conversation turns obtained is no less than the second preset turn threshold. A complete conversation turn refers to a user input sentence and the backend server reply sentence, constituting a complete interaction. For example, if the user says, "I want to check my electricity bill," and the backend server responds, "Which month's electricity bill do you want to inquire about?", this constitutes one complete turn. The second preset turn threshold is a lower limit set to ensure sufficient context before entering intent recognition. For example, if the second preset turn threshold is set to 3, at least three complete conversation turns are required. Time window updating means that if there are too few valid conversation turns within the initial time window, the backend server will expand to older historical conversations until the threshold requirement is met. Time window updates can be implemented by expanding the time range (for example, from the last 5 minutes to the last 10 minutes) or increasing the number of conversation turns (for example, from the last 5 turns to the last 10 turns). Non-redundant complete conversation turns refer to valid conversation turns after redundancy removal to ensure they are relevant to the business and free of duplication. Subsequently, the multi-turn conversation text within the updated time window is encoded into vectors using a pre-trained word embedding model (such as BERT) and concatenated or average-pooled to form the current basic conversation features. For example, a user can send consecutive messages such as "My meter isn't working" and "Did anyone come to my repair call last night?" and reply. This process generates a basic feature vector for intent recognition, providing reliable input for intent judgment and command chain generation in multi-turn conversations. Word embedding maps words or sentences in text into continuous vector representations, ensuring that semantically similar words are close in distance in the vector space. The goal of word embedding is to convert discrete language symbols into numerical representations that can be processed by neural networks. The current basic session features refer to the vectorized session features extracted and processed by the background server within the "update time window".

[0034] The command chain is generated based on the first intention information and the second intention information, including: Step 1: Match the first intent information and the second intent information in a pre-set set of operation items to obtain a plurality of matching operation items; wherein each operation item corresponds to a set of instructions; Step 2: Sort the multiple groups of instructions corresponding to the multiple operation items according to the dependency relationship and time sequence between the instructions to form an instruction chain.

[0035] In some embodiments, the action item set refers to a business operation mapping table pre-established by the backend server. Each action item corresponds to a set of specific executable instructions for completing a specific business processing task. The first and second intent information are matched with the tags or semantic representations in the action item set to filter out action items relevant to the user's intent. For each matched action item, a corresponding set of instructions is selected to perform the actual business operation. The matching method can be implemented based on keyword matching or vector similarity calculation, such as calculating the cosine similarity between the intent vector and the action item vector, and selecting action items with similarity above a threshold. As an example, a user explicitly asks "Request repair for an electric meter" in the intelligent conversation window, while implicit intent recognition infers that the user may want to "Query historical repair status." The backend server matches two action items in the action item set: ① "Generate repair ticket" and ② "Query historical work order status." Each action item corresponds to a set of instructions. For example, the instructions for the "Generate repair ticket" action item include creating a work order record, assigning maintenance personnel, and notifying relevant departments; the instructions for the "Query historical work order status" action item include retrieving historical repair records and returning processing status information. Through this matching process, the backend server can generate a complete, multi-dimensional set of business instructions for the user, laying the foundation for the subsequent formation of an orderly instruction chain. The instructions corresponding to the operation items are the specific steps or command sequences to implement the operation items, such as generating a work order, querying records, and assigning staff.

[0036] In some embodiments, dependencies between instructions refer to the order or logical constraints that exist between different instructions during execution. For example, some instructions must execute only after other instructions have completed. Temporal order refers to arranging the execution order of multiple instructions according to the actual business process or processing sequence to ensure reasonable operation and business continuity. The multiple sets of instructions corresponding to the matched operation items are analyzed, and the dependency fields of each instruction are extracted, such as the ID (identifier) ​​of the operation item that must be executed first and the priority field (e.g., a larger value indicates a higher priority). An instruction dependency graph is then constructed, treating each instruction as a node. If instruction B depends on the completion of instruction A, a directed edge A→B is added to the graph. The dependency graph is then topologically sorted to generate an instruction execution sequence, ensuring that each instruction executes after its dependent instructions have completed. Based on the topological sort, a secondary sort can be performed based on instruction priority or timestamps to achieve the optimal execution order for business logic. The backend server packages the sorted instruction sequence into an instruction chain object, which contains the instruction ID, execution order, and dependency information. An instruction chain is a sequence of sorted instructions used to automatically execute a complete business process, ensuring that each instruction is executed according to dependency and chronological order. An instruction dependency graph, a directed acyclic graph (DAG), is a structure used to represent the execution order and dependencies between instructions in a generated instruction chain. For example, a user reports an electric meter repair and queries its historical status. The backend server matches the actions "Generate Repair Ticket" and "Query Historical Records." The instruction dependencies are resolved: creating a work order must occur first, assigning a repair person and notifying the department can occur in parallel, and querying historical records is performed after user confirmation. The DAG is constructed with nodes representing Create Work Order (A), Assign Personnel (B), Notify Department (C), and Query Historical Records (D); the edges are A→B and A→C. A topological sort is used to generate the instruction chain: A→(B, C in parallel)→D. Finally, execution is performed sequentially to ensure that repair order generation, personnel assignment, and historical record queries are completed smoothly according to business logic.

[0037] The present invention provides a method for generating a command chain based on contextual intent recognition, further comprising: If the event risk level is high, a manual conversation window is created and the corresponding customer service representative is matched to the user conversation request. The manual conversation window is used to display the conversation between the customer service representative and the user. In some embodiments, as Figure 2The flowchart for creating a manual conversation window is shown below. When the event risk level is high, the backend server creates a manual conversation window to display the interaction between customer service personnel and users, ensuring that high-risk events are handled promptly. To match appropriate customer service personnel, the backend server selects the most appropriate personnel based on user attributes or historical user conversation records, thereby improving the efficiency and quality of manual intervention. The manual conversation window is an interactive interface displayed on user terminals (such as mobile apps or webpages) that presents real-time conversations between users and human customer service representatives. It supports text, voice, and other interaction methods to ensure that high-risk events are handled promptly. Therefore, when the event risk level is high, a manual conversation window is created to transfer user conversation requests to human customer service representatives, ensuring that emergency or critical situations can be handled manually. This effectively reduces the risks caused by misjudgments or delayed processing by intelligent customer service representatives, ensuring the safety of the power supply system and users.

[0038] Among them, customer service personnel are matched through the following steps: Matching customer service representatives to user session requests based on user attribute information or historical session records, including: Step 1: If the query finds multiple historical conversation records for the user, sort them in descending order based on the user's feedback score for each historical conversation record, and select the customer service representative with the highest feedback score as the candidate customer service representative; if the candidate customer service representative is currently available, then the candidate customer service representative is determined to be a matching customer service representative; In some embodiments, as Figure 3The customer service agent matching flowchart shown in the figure shows that historical conversation records refer to past interactions between users and customer service agents, including conversation content, processing results, and user feedback. Feedback scores are users' subjective evaluations of customer service quality after each conversation, typically expressed as a numerical value (e.g., 1-5), reflecting the efficiency and satisfaction of the customer service agent's handling of their issue. Based on this, the backend server first queries the database to retrieve all of the user's historical conversation records. If multiple historical conversation records are found for the user, the server extracts the customer service agent ID and corresponding user feedback score for each record. These records are sorted from high to low by score. The customer service agent corresponding to the record with the highest score is selected as a candidate customer service agent. The server checks the status of the candidate customer service agent: If the candidate customer service agent is currently available (not serving other users), the candidate customer service agent is directly assigned to establish a manual conversation window with the user. If the candidate customer service agent is busy, the server proceeds to the next matching step (e.g., searching for the next high-scoring customer service agent or switching to attribute matching). The candidate customer service agent refers to the set of customer service agents with whom the user interacted in the historical conversation records, sorted by feedback score, to select the most prioritized customer service agent. The idle state refers to a state where a customer service representative is not currently handling other user sessions and is immediately available to handle new user requests. For example, user Mr. Zhang has previously reported an electric meter problem through intelligent customer service: the first time, it was handled by Customer Service A, with a feedback score of 3; the second time, it was handled by Customer Service B, with a feedback score of 5 (very satisfied); and the third time, it was handled by Customer Service C, with a feedback score of 4. When Mr. Zhang initiates another high-risk repair request (such as "electricity meter smoking"), the backend server queries his historical session records and ranks the feedback scores from highest to lowest: Customer Service B (5 points), Customer Service C (4 points), and Customer Service A (3 points). The backend server identifies Customer Service B as a candidate. If Customer Service B is currently idle, Mr. Zhang is immediately matched with Customer Service B and a manual conversation window is generated on the user terminal, allowing Mr. Zhang to communicate directly with the customer service representative who has provided the best service in the past, improving processing efficiency and user experience.

[0039] Step 2: If the query finds that there is no historical session record for the user, then a matching customer service representative is selected from the currently available customer service representatives based on the user's attribute information; In some embodiments, if the query finds no historical conversation records for the user, all currently available customer service representatives are screened. Then, matching customer service representatives are selected from these currently available representatives based on the user's attribute information. First, the candidate pool is narrowed down based on the user's geographic location. Then, within the candidate pool, agents are screened based on the account type to identify those skilled in handling that type of service. Based on this, the matching is finally determined by matching the customer service representative's skill tags stored in the backend server with the user's attributes.

[0040] Step 3: If the query finds that the user has historical conversation records, but the user has not rated the historical conversation records, the customer service staff with the highest comprehensive customer service score is selected from the multiple customer service staff corresponding to the historical conversation records to determine as the matching customer service staff.

[0041] In some embodiments, if a user has a historical conversation record but has not provided a rating, the server searches a pre-stored comprehensive rating table for the corresponding customer service representative's comprehensive rating based on the customer service representative's ID. The representative with the highest comprehensive rating is selected as the matching customer service representative. If the highest-scoring representative is busy, the next representative is selected. The comprehensive rating table includes the customer service representative's ID and the corresponding comprehensive rating. The customer service representative's comprehensive rating is a service quality score calculated based on the representative's multi-dimensional performance. The comprehensive rating is typically a numerical indicator (e.g., 0 to 100) that measures the overall service level of the representative. For example, user B has contacted the intelligent customer service representative several times regarding electricity usage issues but has not previously rated the customer service representatives. The backend server searches historical conversations and finds the following customer service representatives involved: Zhang, Li, and Wang. Because the user did not provide a rating, the backend server retrieves the comprehensive ratings of the three representatives: Zhang: 97 (high resolution rate, low complaint rate), Li: 95, and Wang: 92. The backend server determines that Zhang has the highest comprehensive rating and is currently available, so it matches user B with the representative for this service. If agent Zhang is busy, the system prioritizes agent Li. In summary, if a user has multiple past conversations, the system prioritizes the agent who provided the best service experience based on feedback ratings. This enhances user trust and service continuity, prevents users from repeatedly explaining the same issue, and improves problem resolution efficiency. If a user has no past conversations, available agents are precisely selected based on attributes such as location, account type, and skill tags. This ensures a good match between people and positions, ensuring agents have relevant business experience and improving first-response resolution rates. If a user has past conversations but no ratings, a comprehensive rating table maintained in the backend is used to match the user. This rating is based on multiple metrics (such as resolution rate, complaint rate, and response speed). This ensures the assigned agent has a high overall service level and enhances professionalism and stability. This system not only considers ratings and matching, but also checks agent availability in real time to enable dynamic scheduling. This eliminates waiting in queues and ensures that high-risk incidents are addressed promptly.

[0042] Among them, user attribute information includes user geographic location and account type, as well as Filter matching agents from currently available agents based on user attribute information, including: Determine the corresponding customer service information based on the customer service personnel whose current status is idle, where the customer service information includes the customer service ID and the skill tag corresponding to the customer service ID; According to the user's geographic location, regional matching is performed from customer service personnel who are currently available to obtain at least one first customer service personnel to be selected; If there are multiple first candidate customer service personnel, match the first candidate customer service personnel with the services according to the account type and select at least one second candidate customer service personnel; If there are multiple second customer service personnel to be selected, the skill tags of the customer service personnel are matched with the event type selected by the user to obtain at least one target customer service personnel; and the target customer service personnel is determined as the matching customer service personnel.

[0043] In some embodiments, user attribute information includes the user's geographic location (e.g., region or power supply bureau jurisdiction) and account type (e.g., residential, industrial, or commercial). The user's geographic location refers to the user's actual location or service jurisdiction, and is used to determine the power supply bureau, city, or region to which the user belongs. For example, it can be represented by a specific administrative division, latitude and longitude coordinates, or power supply bureau code. Geographic location is primarily used for regional matching, ensuring that user requests are assigned to a customer service representative with service coverage in the same region. Account type refers to the user's electricity usage category or account identity within the power system and is used for service matching. For example, common account types include: residential (for daily household use); commercial (for business premises); and industrial (for production plants and businesses). Based on this, the ID and corresponding skill tags of each customer service representative are retrieved from a list of currently available customer service representatives. Based on the geographic location information in the user's attributes, customer service representatives in the same region or service jurisdiction are selected to obtain a first candidate customer service representative. If multiple customer service representatives are included in the first candidate, the customer service representative is selected based on the user's account type, specifically those who are skilled in handling services of that account type, to obtain a second candidate customer service representative. If multiple agents are still available in the second candidate pool, the agent's skill tags are matched with the user's current event type, and the agent who best meets the event handling requirements is selected as the target agent. The target agent is assigned to the user's session request to ensure that the user receives appropriate human service. Agent information refers to agent-related information recorded in the backend server, including the agent's unique ID and skill tags (the agent's professional skills or business expertise, such as repair, complaint, and inquiry). The first candidate pool refers to a set of available agents selected by regional matching based on the user's geographic location. These agents cover the user's area and can handle user requests in that area, but may still require further business or skill matching. The second candidate pool refers to a set of agents selected by business matching based on account type within the first candidate pool. These agents not only cover the user's area but also specialize in the user's business type (e.g., residential electricity, industrial electricity, etc.). The target agent is the final agent selected from the second candidate pool by matching their skill tags with the event type selected by the user. The target customer service personnel meet the regional, business, and skill requirements and are selected as the customer service personnel to establish a manual conversation window with the user. As an example, user Mr. Zhang initiates a repair request on the app. The event type is "repair report", the account type is "residential electricity", and the location is District A, City A. First, all customer service IDs and skill tags are obtained from the currently idle customer service personnel. Based on Mr. Zhang's geographical location, regional matching is performed to obtain the first candidate customer service personnel A, B, and C. Because Mr. Zhang's account type is residential electricity, further screening is performed on customer service personnel who specialize in residential business, resulting in the second candidate customer service personnel A and B.The second candidate's skill tag was matched to the event type "Repair Report." It was found that Customer Service B specialized in meter repair reports, while Customer Service A primarily handled inquiries. Customer Service B was ultimately selected as the matching customer service, and a manual conversation window was created on the user's terminal, allowing Mr. Zhang to communicate directly with Customer Service B, achieving efficient repair report processing.

[0044] The comprehensive customer service score is determined through the following steps: Collect historical service data for each of the multiple customer service representatives, including historical response time, problem resolution rate, customer satisfaction score, and complaint rate. Perform weighted calculation on each indicator in the historical service data according to the preset weights to obtain the weighted total score of each customer service staff; the weighted total score is used as the comprehensive customer service score.

[0045] In some embodiments, the historical service data of each customer service representative is extracted from the database of the backend server. A predefined weight is set for each indicator, for example, the user satisfaction score has the largest weight, the response time has the second largest weight, the complaint rate is a negative indicator, and the problem resolution rate is also weighted. The various indicators of each customer service representative are multiplied by the corresponding weights and summed up to obtain a weighted total score output, and the weighted total score is directly used as the comprehensive score of the customer service representative for subsequent matching decisions. Historical service data: refers to the relevant data generated by the customer service representative when handling user requests in the past, including response time (time required to process the request), problem resolution rate (the proportion of successfully resolved problems), user satisfaction score (user's subjective evaluation of the service, which can be 1 to 5 points), complaint rate (the frequency of user complaints), etc. As an example, for example, user Mr. Zhang initiates a repair request for "electricity meter abnormality". The database collects historical service data for three potential customer service representatives: Customer Service A, with a response time of 12 minutes, a problem resolution rate of 95%, a customer satisfaction rating of 4.8, and a complaint rate of 2%. Customer Service B, with a response time of 8 minutes, a problem resolution rate of 90%, a customer satisfaction rating of 4.5, and a complaint rate of 1%. Customer Service C, with a response time of 15 minutes, a problem resolution rate of 98%, a customer satisfaction rating of 4.7, and a complaint rate of 0%. Weights are assigned to response time (10%), problem resolution rate (30%), customer satisfaction rating (40%), and complaint rate (20%). A weighted total score is calculated based on the weights. The calculation shows that Customer Service A has the highest overall score and is currently available. Therefore, it is matched with Mr. Zhang as the customer service representative for handling his repair request, ensuring an efficient and satisfactory service experience.

[0046] In these embodiments, the understanding and response effects of multi-round conversations are improved, and high-risk events are ensured to be handled promptly and reasonably. Specifically, through redundancy deletion, time window update and word embedding, multi-round conversation data is converted into concise and sufficient basic conversation features to ensure the accuracy of explicit and implicit intent recognition. Multi-round intention information is matched with a set of preset operation items, and an instruction chain is generated based on dependencies and time sequence to achieve automated and accurate execution of multi-step business operations. Through user attribute information and historical conversation records, combined with feedback scores or comprehensive scores, idle customer service personnel are intelligently matched, so that high-risk events are handled promptly and accurately, while reducing the blindness of manual intervention and improving service efficiency and user satisfaction.

[0047] In some embodiments, in order to further solve the third technical problem described in the background technology section, namely, "the existing technology evaluates user emotions through single modal information, resulting in inaccurate emotion judgment, thereby affecting the timely response and accurate judgment of user status during intelligent conversation." In some embodiments of the present invention, the emotion perception model includes a semantic emotion recognition model, a facial expression recognition model, and a voice emotion analysis model, as well as Real-time user sentiment scoring through sentiment perception models, including: During the conversation, a semantic emotion score is generated through the semantic emotion recognition model; Collect user facial expression information and user voice information; analyze facial expression features through facial expression recognition model to generate facial emotion score; and generate voice emotion score through voice emotion analysis model; The facial emotion score, voice emotion score and semantic emotion score are integrated to generate the user emotion score.

[0048] In some embodiments, an emotion perception model, a model used to assess a user's emotional state in real time, comprises three components: a semantic emotion recognition model that analyzes potential emotional information in user input text (e.g., text or chat content); a facial expression recognition model that captures the user's facial expressions via a camera and extracts emotion-related features; and a speech emotion analysis model that analyzes features such as intonation, volume, and speaking rate in the user's speech to identify emotional state. The semantic emotion recognition model can use a pre-trained BERT model to extract text features, then outputs an emotion category (e.g., positive, negative, neutral) or an emotion score through a classification layer. Facial expression recognition models, such as the OpenFace model, can map emotional states (such as happiness, anger, sadness, etc.) through facial landmark detection and action unit recognition. A speech emotion analysis model, such as OpenSMILE, an open-source audio feature extraction tool, can extract features such as intonation, volume, and speaking rate. An LSTM based on the extracted features is used for emotion classification or emotion score prediction. Facial emotion scoring involves analyzing the emotional state reflected by the user's facial expressions and quantifying it into a numerical emotion score. Voice emotion scoring quantifies a user's emotional state into a numerical score by analyzing voice characteristics such as intonation, volume, and speaking rate. Semantic emotion scoring quantifies a user's emotional state by analyzing the semantic information of the text input (such as text content and the emotional tendency of the wording). Based on this, the facial emotion score, voice emotion score, and semantic emotion score are weighted and fused to produce a user emotion score. This weighted fusion can be achieved through methods such as weighted summation with preset weights or by concatenating features into a fusion model to generate a more accurate user emotion assessment.

[0049] In these embodiments, the accuracy and real-time performance of emotion recognition are improved. Specifically, text, facial, and voice features are extracted through a semantic emotion recognition model, a facial expression recognition model, and a voice emotion analysis model, respectively. A pre-trained model is used to ensure accurate feature extraction. The emotional information of the three modalities is quantified into comparable scores (facial emotion score, voice emotion score, and semantic emotion score), providing a basis for fusion. The three types of emotion scores are fused through a weighted summation with preset weights or feature splicing input into a fusion model to obtain a final user emotion score. The fusion method can adaptively adjust the importance of different modalities and improve the accuracy and robustness of emotion assessment. The obtained user emotion score can be used to dynamically adjust the reply strategy or trigger enhanced intent recognition in the intelligent conversation window, thereby improving the quality of interaction and user satisfaction.

[0050] The above descriptions are merely some preferred embodiments of the present invention and illustrate the underlying technical principles. Those skilled in the art should understand that the scope of the present invention is not limited to technical solutions formed by specific combinations of the aforementioned technical features. It also encompasses other technical solutions formed by any combination of the aforementioned technical features or their equivalents, without departing from the aforementioned inventive concept. For example, a technical solution formed by replacing the aforementioned features with (but not limited to) technical features with similar functions disclosed in this invention.

Claims

1. A method for generating a command chain based on contextual intent recognition, characterized in that: include: Upon receiving a user session request, extracting an event type from the user session request, and determining an event risk level based on the event type; If the event risk level is low or medium, establish an intelligent conversation window and dynamically capture context data corresponding to the user's current input information, and generate basic conversation features; Inputting the basic conversation features into an explicit intent recognition network to obtain first intent information and a first confidence level; Based on the first intention information, generating reply information for the user's question and displaying it in the smart conversation window; During the conversation, the user's emotions are scored in real time through the emotion perception model; When the user's sentiment score is lower than the preset score threshold; Or if the first confidence is less than a preset threshold and the current conversation round in the intelligent conversation window reaches the first preset round threshold, then an enhanced conversation feature is constructed and input into an implicit intent recognition network to obtain second intent information; Based on the first intention information and the second intention information, an instruction chain is generated.

2. The method for generating a command chain based on contextual intention recognition according to claim 1, characterized in that: The process of constructing enhanced conversation features and inputting the features into an implicit intent recognition network includes: The user emotion score and user attribute information are subjected to feature extraction to obtain user features; the user features are spliced ​​with the current basic session features to obtain enhanced session features and input them into the implicit intent recognition network.

3. The method for generating a command chain based on contextual intention recognition according to claim 2, characterized in that: The current basic session features are extracted through the following steps: Extracting session data within a preset time window within the intelligent session window, and performing a redundancy removal operation on the session data to obtain redundancy-removed session data; If the number of complete conversation turns included in the conversation data after redundancy removal is less than a second preset turn threshold, updating the time window based on the preset time window until the number of non-redundant complete conversation turns included in the updated time window is greater than or equal to the second preset turn threshold; The conversation data within the update time window is word-embedded to obtain the current basic conversation features.

4. The method for generating a command chain based on contextual intention recognition according to claim 3, characterized in that: The generating a command chain based on the first intent information and the second intent information includes: The first intention information and the second intention information are matched in a preset set of operation items to obtain a plurality of matching operation items; wherein each operation item corresponds to a set of instructions; According to the dependency relationship and time sequence between the instructions, multiple groups of instructions corresponding to multiple operation items are sorted to form an instruction chain.

5. The method for generating a command chain based on contextual intention recognition according to claim 4, characterized in that: Also includes: If the event risk level is high, a manual conversation window is created and a corresponding customer service representative is matched for the user conversation request. The manual conversation window is used to display the conversation between the customer service representative and the user. The customer service personnel are matched through the following steps: The user's session request is matched with a corresponding customer service representative through user attribute information or user's historical session records.

6. The method for generating a command chain based on contextual intention recognition according to claim 5, characterized in that: The matching of the user session request with a corresponding customer service representative based on the user attribute information or the user's historical session records includes: If the query finds that the user has multiple historical conversation records, sort them in descending order based on the user's feedback score for each historical conversation record, and select the customer service representative with the historical conversation record with the highest feedback score as the candidate customer service representative; if the candidate customer service representative is currently available, then it is determined as a matching customer service representative; If the query finds that the user has no historical conversation record, the matching customer service personnel will be selected from the currently available customer service personnel based on the user's attribute information; If the query finds that the user has historical conversation records, but the user has not rated the historical conversation records, the customer service staff with the highest comprehensive customer service score is screened from multiple customer service staff corresponding to the historical conversation records and determined as the matching customer service staff.

7. The method for generating a command chain based on contextual intention recognition according to claim 6, characterized in that: The user attribute information includes user geographic location and account type, and The step of screening matching customer service personnel from customer service personnel currently in an idle state according to user attribute information includes: Determine the corresponding customer service information based on the customer service personnel whose current status is idle, where the customer service information includes the customer service ID and the skill tag corresponding to the customer service ID; According to the user's geographic location, regional matching is performed from customer service personnel who are currently available to obtain at least one first customer service personnel to be selected; If there are multiple first customer service personnel to be selected, perform business matching from the first customer service personnel to be selected according to the account type, and screen at least one second customer service personnel to be selected; If there are multiple second customer service personnel to be selected, the skill tags of the customer service personnel are matched with the event type selected by the user to obtain at least one target customer service personnel; and the target customer service personnel is determined as the matching customer service personnel.

8. The method for generating a command chain based on contextual intention recognition according to claim 7, characterized in that: The customer service composite score is determined by the following steps: Collect historical service data for each of the multiple customer service representatives, including historical response time, problem resolution rate, customer satisfaction score, and complaint rate. A weighted calculation is performed on each indicator in the historical service data according to preset weights to obtain a weighted total score for each customer service staff member; the weighted total score is used as the comprehensive score of the customer service staff member.

Citation Information

Patent Citations

  • Text recognition method and device, electronic equipment and storage medium

    CN118095272A

  • Full-amount speech analysis method, device and equipment based on large model, and storage medium

    CN119155391A

  • Large-model complaint intention recognition method based on sentiment analysis

    CN120146056A

  • Intention recognition model training method, intention recognition method, intention recognition device and intention recognition equipment

    CN120235164A

  • Dialogue processing method and device based on historical session recognition, equipment and medium

    CN120450048A

Cited By

  • Instruction processing method and device based on social intention recognition, storage medium and equipment

    CN121960800A