Session handling method and program product

CN122838541APending Publication Date: 2026-09-29TIANJIU SHARING NETWORK TECH GRP CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610963361.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-30
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

然而,在实际应用场景中,用户意图具有高度的动态性和复杂性,会话状态会随交互轮次发生剧烈变化,例如从简单的信息查询迅速演变为高难度的业务谈判或突发的情绪化投诉

Benefits of technology

本公开实施例通过实时检测会话数据流并动态生成会话状态特征,实现了从静态固定服务向动态自适应服务的技术转变,能够精准识别用户意图的突变或复杂度的升级,在满足预设切换条件时,依据状态特征自动从执行单元池中检索并调度具备差异化专长、处理逻辑或服务层级的第二执行单元进行无缝接管;此外,本公开不仅显著提升了复杂业务场景下的意图识别准确率与问题解决效率,还通过基于反馈信号的确认机制确保了会话主导权切换的可靠性与平滑性,从而在保障低延迟响应的同时,实现了全链路会话处理质量的持续优化与用户体验的显著提升。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122838541A_ABST
    Figure CN122838541A_ABST
Patent Text Reader

Abstract

This disclosure provides a session processing method and program product, relating to the field of artificial intelligence technology. The method includes: acquiring a real-time session data stream between a first execution unit and a user; continuously monitoring the session data stream and generating session state features; if the session state features meet preset switching conditions, retrieving and determining a second execution unit from a preset execution unit pool based on the session state features, wherein the second execution unit and the first execution unit have different areas of expertise, processing logic configurations, or service levels; sending a takeover request to the second execution unit; and, upon receiving a feedback signal confirming the takeover from the second execution unit, setting the second execution unit as the dominant response node for the current session to continue processing the session data stream with the user. According to embodiments of this disclosure, the response quality of session processing can be improved, enhancing the user's interactive experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of artificial intelligence technology, and in particular to a session processing method and program product. Background Technology

[0002] In existing conversation processing architectures, a single execution unit (such as a fixed-configuration chatbot or a domain-specific agent) is typically used to respond to the user's real-time conversation data stream throughout the entire process. However, in real-world applications, user intent is highly dynamic and complex, and the conversation state can change drastically with each round of interaction, for example, rapidly evolving from a simple information query to a complex business negotiation or a sudden emotional complaint. Faced with these scenarios of drastic changes in conversation state, existing conversation processing solutions show a declining trend in response quality and struggle to maintain a continuous and stable interactive experience.

[0003] It should be noted that the information disclosed in the background section above is only used to enhance the understanding of the background of this disclosure, and therefore may include information that does not constitute prior art known to those skilled in the art. Summary of the Invention

[0004] This disclosure provides a session processing method and program product that can improve the response quality of session processing and enhance the user's interactive experience in scenarios where the session state changes drastically.

[0005] Other features and advantages of this disclosure will become apparent from the following detailed description, or may be learned in part by practice of this disclosure.

[0006] According to one aspect of this disclosure, a session processing method is provided, comprising: acquiring a real-time session data stream between a first execution unit and a user; continuously monitoring the session data stream and generating session state features; if the session state features meet preset switching conditions, retrieving and determining a second execution unit from a preset execution unit pool based on the session state features, wherein the second execution unit and the first execution unit have different areas of expertise, processing logic configurations, or service levels; sending a takeover request to the second execution unit; and, upon receiving a feedback signal confirming the takeover from the second execution unit, setting the second execution unit as the dominant response node of the current session to continue processing the session data stream with the user.

[0007] In one embodiment of this disclosure, retrieving and determining a second execution unit from a preset execution unit pool based on session state features includes: calculating a similarity score between the session state features and the capability description vectors of each candidate execution unit; wherein the capability description vectors characterize the expertise, processing logic configuration, or service level of each candidate execution unit; and selecting the candidate execution unit with the highest similarity score that meets a preset matching threshold as the second execution unit.

[0008] In one embodiment of this disclosure, the preset execution unit pool includes multiple agent instances and / or multiple human interaction nodes; retrieving and determining a second execution unit from the preset execution unit pool based on session state characteristics includes: determining the second execution unit from the candidate execution units based on the session state characteristics and the real-time running status of each candidate execution unit in the execution unit pool; wherein the real-time running status includes at least one of the following data: the concurrent processing load of the agent instance, the response latency index of the agent instance, or the current idle time of the human interaction node.

[0009] In one embodiment of this disclosure, continuous detection of the session data stream and generation of session state features include: performing real-time matching detection on the session data stream using a pre-set keyword rule base, identifying keyword sequences containing specific business intentions, and mapping the matching results to session state features; or, extracting deep semantic vectors from the session data stream, whereby the deep semantic vectors are used to characterize the user's potential emotional state or task complexity, and determining the deep semantic vectors as session state features; a second-dimensional feature vector; or, performing real-time matching detection on the session data stream using a pre-set keyword rule base, identifying keyword sequences containing specific business intentions, extracting deep semantic vectors from the session data stream, and generating session state features based on the keyword sequences and deep semantic vectors.

[0010] In one embodiment of this disclosure, the preset switching conditions include at least one of the following: the session state feature indicates that the user's emotional state index is lower than a preset emotional threshold, the emotional state index being calculated based on the user's emotional tendency represented in the deep semantic vector; the session state feature indicates that the task attribute of the current session exceeds the preset expertise domain of the first execution unit; the session state feature indicates that the user sends a preset manual intervention instruction, or detects that the user explicitly expresses the intention to seek manual assistance; a preset high-priority business keyword sequence is matched in the session data stream; the session state feature indicates that the task complexity of the current session exceeds the preset processing threshold of the first execution unit, the task complexity being determined comprehensively based on the number of dialogue turns, the depth of context dependency, or the confidence level of entity recognition.

[0011] In one embodiment of this disclosure, after retrieving and determining the second execution unit from a preset execution unit pool based on the session state characteristics, and before receiving a feedback signal confirming takeover from the second execution unit, the method further includes: when the session state characteristics indicate that the user's emotional state index is lower than a preset emotional threshold, controlling the first execution unit to call a preset emotional soothing strategy library to generate response text containing empathetic expressions, apology statements, or soothing guidance information; and sending the response text to the user.

[0012] In one embodiment of this disclosure, after setting the second execution unit as the dominant response node of the current session, the method further includes: generating session summary information or key feature snapshots based on the historical session context of the first execution unit; and pushing the session summary information or key feature snapshots to the second execution unit so that the second execution unit can inherit the previous session state and continue to process the session data stream.

[0013] In one embodiment of this disclosure, after receiving a feedback signal confirming takeover from the second execution unit, the method further includes: sending a session transfer notification message to the user, the notification message being used to inform the user that the current session is being taken over by a new execution unit; or, without interrupting the session data stream, downgrading the control authority of the first execution unit, so that the second execution unit can smoothly take over the conversation leadership.

[0014] In one embodiment of this disclosure, the method further includes: updating the last active timestamp corresponding to the current session each time a user's session data stream is received or a response text is generated; executing a background scheduling task with a preset period to traverse all pending session lists and obtain the last active timestamp of each session; for each traversed session, calculating the time difference between the current time and the last active timestamp; if the time difference exceeds a preset silence timeout threshold, controlling the first execution unit or a designated wake-up agent instance to retrieve and generate a corresponding wake-up prompt message from a preset session activation content library; and sending the wake-up prompt message to the user.

[0015] In one embodiment of this disclosure, when the dominant response node of the current session is a human interaction node, the method further includes: calculating the time difference between the current time and the last operation timestamp of the human interaction node on the current session data stream; if the time difference exceeds a preset human response timeout threshold, then calling the configuration rules in the preset waiting appeasement strategy library to generate an appeasement message, the appeasement message including at least one of queuing progress information, apology and explanation information or interactive guidance information; and sending the generated appeasement message to the user.

[0016] In one embodiment of this disclosure, the method further includes: analyzing the intent confidence and sentiment score in the session state features; if the intent confidence is higher than a preset high-priority threshold and a keyword sequence representing transaction-oriented behavior is matched in the deep semantic vector, then a first type of interaction intent label is generated; if the sentiment score is lower than a preset negative sentiment threshold and a keyword sequence representing objection / conflict is matched in the deep semantic vector, then a second type of interaction risk label is generated; the session metadata with the corresponding labels is written into a preset classification index library to construct a high-priority interaction queue and an abnormal risk queue; and an aggregated view of the high-priority interaction queue and the abnormal risk queue is pushed to the management terminal.

[0017] According to another aspect of this disclosure, a computer program product is provided, the computer program product including program instructions for implementing the session processing method described above.

[0018] The technical solutions provided in this disclosure can include the following beneficial effects: This disclosure achieves a technological shift from static, fixed services to dynamic, adaptive services by real-time detection of session data streams and dynamic generation of session state features. It can accurately identify sudden changes in user intent or escalation of complexity. When preset switching conditions are met, it automatically retrieves and schedules a second execution unit with differentiated expertise, processing logic, or service level from the execution unit pool to seamlessly take over based on the state features. In addition, this disclosure not only significantly improves the accuracy of intent recognition and problem-solving efficiency in complex business scenarios, but also ensures the reliability and smoothness of session leadership switching through a confirmation mechanism based on feedback signals. Thus, while ensuring low-latency response, it achieves continuous optimization of end-to-end session processing quality and a significant improvement in user experience.

[0019] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description

[0020] The accompanying drawings, which are incorporated in and form a part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure.

[0021] Obviously, the accompanying drawings described below are merely some embodiments of this disclosure. Those skilled in the art can obtain other drawings based on these drawings without any creative effort.

[0022] Figure 1 This diagram illustrates a flowchart of a session processing method according to an embodiment of the present disclosure; Figure 2 This diagram illustrates another session processing method according to an embodiment of the present disclosure. Figure 3 This diagram illustrates the process of handing over session history content in an embodiment of this disclosure. Figure 4 This diagram illustrates a user activation process in an embodiment of the present disclosure. Figure 5 This diagram illustrates a flowchart of the manual response delay handling process in an embodiment of this disclosure. Figure 6 This diagram illustrates the intended classification flowchart in an embodiment of the present disclosure. Figure 7 This diagram illustrates a session processing apparatus according to an embodiment of the present disclosure. Figure 8A structural block diagram of an electronic device according to an embodiment of the present disclosure is shown. Detailed Implementation

[0023] To facilitate understanding of the technical solutions of this disclosure, the disclosure will be further described below with reference to the accompanying drawings.

[0024] The terms "first" and "second," etc., used in this disclosure are only used to distinguish different objects and not to describe a specific order. The terms "comprising" and "having," and any variations thereof, in the embodiments of this disclosure are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the listed steps or units, but may also include steps or units not listed.

[0025] Throughout this disclosure, the references to "embodiments" do not necessarily refer to the same embodiments, nor are they independent or alternative embodiments mutually exclusive with other embodiments. Those skilled in the art will understand, explicitly and implicitly, that the embodiments described in this disclosure can be combined with other embodiments.

[0026] In the embodiments of this disclosure, "at least one" refers to one or more, "more" refers to two or more, "at least two" refers to two or three or more, and "and / or" is used to describe the relationship between associated objects. The character " / " generally indicates that the associated objects before and after are in an "or" relationship.

[0027] In real-world application scenarios of session processing, user intent is highly dynamic and complex, and session states can change drastically with each round of interaction. For example, a simple information query can quickly evolve into a complex business negotiation or a sudden emotional complaint. Faced with these scenarios of drastic changes in session states, existing session processing solutions show a declining trend in response quality and struggle to maintain a continuous and stable interactive experience.

[0028] The practical application scenarios of the aforementioned session processing can include enterprise-level instant messaging sales management scenarios. In these scenarios, taking a typical sales and customer interaction process as an example, the system typically deploys a single automated robot or a fixed-configuration intelligent agent as the first execution unit, responding to the user's real-time session data stream around the clock. In practical applications, user intent is highly dynamic and complex, and the session state can drastically evolve within a very short number of interaction rounds. For example, the initial stage of a conversation may only involve a simple product information query, belonging to a low-complexity, low-risk scenario; however, it may quickly evolve into a high-difficulty business negotiation involving price negotiations, discussions of customized solutions, or sudden emotional complaints due to delayed service response, belonging to a high-complexity, high-risk scenario.

[0029] Faced with the dramatic evolution of the aforementioned conversation state, existing technical solutions reveal significant shortcomings. First, rigid intent recognition leads to a decline in response quality. Existing technologies largely rely on static keyword matching or simple rule engines, lacking the ability to continuously perceive deep semantics and contextual features. When a user's intent abruptly shifts from information query to in-depth negotiation or emotional venting, a single execution unit cannot dynamically adjust its processing strategy, resulting in mechanical and illogical responses and a sharp drop in response quality. Second, the lack of human-computer collaboration mechanisms leads to a fragmented user experience.

[0030] Furthermore, in the existing architecture, the switch between artificial intelligence and human intervention heavily relies on manual intervention after monitoring chat logs. This passive switch has a significant time delay, causing key signals from high-intent customers to be ignored and the best follow-up opportunity to be missed.

[0031] The inventors discovered that, due to the inability to distinguish between low-intent casual conversation and high-intent closing signals in real time, valuable human sales resources are often wasted on repetitive questioning of low-value customers, while high-intent customers who truly require human intervention do not receive timely responses. This misallocation of resources directly leads to a precipitous drop in the interactive experience, not only causing the loss of a large number of potential customers but also severely restricting the company's service capabilities and conversion efficiency in the high-end market.

[0032] The deficiencies of the above solutions and the proposed solutions are the result of the inventor's practice and careful research. Therefore, the discovery process of the above problems and the solutions proposed in this disclosure below should be considered as the inventor's contribution to this disclosure.

[0033] It is understood that the data involved in this disclosure, including but not limited to the data itself, its acquisition, and its use, shall comply with the requirements of relevant laws, regulations, and provisions. Before using the technical solutions disclosed in the embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure in an appropriate manner in accordance with relevant laws and regulations, and their authorization should be obtained.

[0034] The following detailed description of this exemplary implementation method is provided in conjunction with the accompanying drawings and embodiments.

[0035] This disclosure provides a session processing method that can be executed by any electronic device with computing capabilities. The execution subject of the session processing method can be at least one of the terminal devices that can be configured to execute the session processing method provided in this disclosure, such as smartphones, smart tablets, wearable devices, desktop computers, laptops, and smart speakers. Alternatively, the execution subject of the session processing method can also be the client itself capable of executing the method.

[0036] Figure 1 A flowchart of a session processing method according to an embodiment of this disclosure is shown, such as Figure 1 As shown, the session processing method provided in this embodiment includes S101-S105.

[0037] In S101, the real-time session data stream between the first execution unit and the user is acquired.

[0038] The first execution unit refers to the initial session processing entity that is currently responding to the user's request.

[0039] In some embodiments, the execution unit (including the first execution unit and the second execution unit hereinafter) may be an AI agent instance or a human interaction node. An AI agent instance may be a chatbot configured with a basic knowledge base and standard response logic; a human interaction node may be a human agent.

[0040] Real-time session data streams refer to continuous, dynamic data sequences generated during the interaction between the user and the first execution unit, including text messages, speech-to-text transcription, multimedia instructions, and accompanying timestamps and metadata.

[0041] In some embodiments, the aforementioned session can be a session with a single user or a session between chat groups.

[0042] In S102, the session data stream is continuously monitored, and session state features are generated.

[0043] Continuous detection can be achieved by performing a continuous sliding window analysis on the incoming session data.

[0044] Session state features can be a set of multi-dimensional quantitative indicators abstracted and extracted from the raw data stream. Session state features may include, but are not limited to, user emotional tendency values ​​(such as anger, anxiety), intent complexity levels (such as information query vs. business negotiation), task complexity scores, and key business intent tags.

[0045] In some embodiments, pre-trained semantic analysis models, sentiment computing engines, or rule engines can be used to process real-time data streams in parallel, extract key semantic vectors, and combine them with preset statistical rules (such as the frequency of negative sentiment words, the proportion of long and difficult sentences, and the occurrence rate of professional terms) to calculate a comprehensive state index. For example, when it is detected that a user continuously uses rhetorical questions and the sentiment score is below a threshold, or mentions high-risk words such as "complaint" or "compensation," a state feature label of "high conflict risk" or "deep business negotiation" is generated.

[0046] The above steps transform unstructured natural language interactions into quantifiable structured state indicators, enabling a refined perception of the conversational situation. This allows for better capture of subtle signs of user intent shifting from simple queries to complex business or emotional scenarios. Regardless of whether the service is currently being provided by a machine or a human, it can objectively determine whether processing resources need to be upgraded, avoiding missed judgments caused by subjective judgments or static rules.

[0047] In S103, if the session state characteristics meet the preset switching conditions, the second execution unit is retrieved and determined from the preset execution unit pool based on the session state characteristics.

[0048] The second execution unit has different areas of expertise, processing logic configurations, or service levels than the first execution unit.

[0049] The preset switching conditions can be based on thresholds or combinations of Boolean rules defined by business logic, used to determine whether the current session exceeds the capability boundary of the first execution unit.

[0050] An execution unit pool can be a registry center containing multiple heterogeneous execution units, each labeled with specific metadata tags. These tags include areas of expertise (such as legal, financial, and technical support), processing logic configurations (such as multi-turn dialogue strategies and reassurance scripts), and service levels (such as junior customer service, senior experts, and advanced human agents). The pool contains both agents with different capabilities and human agents with different skill levels.

[0051] The above steps enable precise resource scheduling based on scenario characteristics, breaking the one-size-fits-all service model of a single execution unit. Through a dynamic matching mechanism, it ensures that only units with corresponding professional capabilities and processing levels (whether it is a more advanced intelligent agent or a more experienced human) can intervene in conversations of specific difficulty, thereby guaranteeing the professionalism and adaptability of subsequent responses from the source and significantly improving the success rate of problem solving and service quality in complex scenarios.

[0052] In S104, a takeover request is sent to the second execution unit.

[0053] A takeover request is a control signaling message that may include a context summary of the current session, generated state characteristics, fragments of historical interaction records, user profile information, and pending next instructions. This takeover request aims to notify the target unit that it is ready to intervene and inherit session control.

[0054] In S105, after receiving the feedback signal confirming the takeover from the second execution unit, the second execution unit is set as the dominant response node of the current session to continue processing the session data stream with the user.

[0055] The confirmation of takeover is a response returned by the second execution unit after processing the takeover request, indicating that it has received the context, assessed its own capacity, and agreed to take over the current session.

[0056] The dominant response node is the processing entity with the highest priority routing rights during the current session's lifecycle. All subsequent messages from the user will be directly routed to this node, which is responsible for generating responses and maintaining the session state.

[0057] Upon receiving confirmation, the main control module immediately updates the session routing table, pointing the current session ID to the second execution unit. Thereafter, new session data streams are no longer sent to the first execution unit but are directly forwarded to the second execution unit for processing. For human agents, the system automatically pops up a session window with the message "Please take over"; for intelligent agents, the system activates their corresponding expert policy. Simultaneously, the system retains the historical context of the first execution unit for the second unit's reference and, when necessary, allows the second unit to send back results to the unified output layer to complete the response to the user.

[0058] This disclosed embodiment dynamically identifies intent evolution and risks by detecting session state characteristics in real time. Once the switching condition is triggered, a second execution unit (AI agent or human) with a higher level of expertise or specialization is automatically retrieved and scheduled to take over the session. This achieves millisecond-level seamless switching between AI and human, ensuring that complex scenarios are handled accurately by experts, eliminating response delays and experience gaps, optimizing resource allocation efficiency, and significantly improving the conversion rate and service satisfaction of high-value customers.

[0059] In some embodiments, retrieving and determining a second execution unit from a preset execution unit pool based on session state characteristics may involve calculating the similarity score between the session state characteristics and the capability description vectors of each candidate execution unit; selecting the candidate execution unit with the highest similarity score that meets a preset matching threshold as the second execution unit. The capability description vector represents the area of ​​expertise, processing logic configuration, or service level of each candidate execution unit.

[0060] In some embodiments, the aforementioned session state features may include a set of vectors, also known as requirement vectors, describing the capability dimensions required by the current session. This requirement vector comprehensively characterizes the specific needs of the current session in terms of expertise, processing logic complexity, and service level. Furthermore, calculating the similarity score between the session state features and the capability description vectors of each candidate execution unit can be equivalent to calculating the similarity score between the requirement vectors and the capability description vectors of each candidate execution unit.

[0061] This disclosed embodiment transforms session requirements and unit capabilities into vectors, achieving precise matching through similarity calculation. This scheme replaces rigid rules with mathematical quantification, locking in the optimal execution unit within milliseconds; it supports fine-grained dynamic adaptation to complex scenario changes; and it utilizes a threshold mechanism to mitigate resource mismatch risks. Ultimately, it ensures efficient on-demand collaboration of heterogeneous resources (intelligent agents and human intervention), significantly improving service response quality and user experience.

[0062] In some embodiments, the preset execution unit pool includes multiple agent instances and / or multiple human interaction nodes.

[0063] In some embodiments, retrieving and determining a second execution unit from a preset execution unit pool based on session state characteristics may involve determining the second execution unit from among the candidate execution units based on session state characteristics and the real-time running status of each candidate execution unit in the execution unit pool. The real-time running status includes at least one of the following data: the concurrent processing load of the agent instance, the response latency metric of the agent instance, or the current idle time of the human interaction node.

[0064] The concurrent processing load of an agent instance refers to the number of sessions that the agent is currently processing simultaneously. The response latency metric of an agent instance refers to the actual time elapsed from receiving a request to outputting a result. The current idle time of a human interaction node refers to the time elapsed since the human agent last ended a session.

[0065] This disclosure introduces real-time runtime status as a supplement to the scheduling dimension, achieving dual optimization of capacity matching and load balancing. By dynamically sensing the concurrent pressure, latency performance, and idle status of human agents, it can accurately avoid congested nodes and seamlessly transfer complex sessions to both professional and idle execution units. This not only effectively prevents experience collapse caused by single-point overload but also maximizes the utilization of all resources, ensuring millisecond-level high-efficiency response and stable service even during peak periods.

[0066] In some embodiments, continuous monitoring of the session data stream and generation of session state features can be performed in any of the following ways: The system uses a pre-built keyword rule base to perform real-time matching and detection on the session data stream, identifies keyword sequences containing specific business intentions, and maps the matching results to session state features. Alternatively, deep semantic vectors can be extracted from the conversation data stream. These deep semantic vectors are used to characterize the user's potential emotional state or task complexity, and are thus identified as conversation state features; the second dimension of the feature vector. Alternatively, a pre-built keyword rule base can be used to perform real-time matching and detection on the session data stream, identify keyword sequences containing specific business intentions, extract deep semantic vectors from the session data stream, and generate session state features based on keyword sequences and deep semantic vectors.

[0067] The pre-built keyword rule base contains a multi-level business intent vocabulary (e.g., "refund" and "complaint" correspond to high-risk intent, while "price" and "cooperation" correspond to high-intent transaction intent). In some embodiments, regular expressions or the Aho-Corasick multi-pattern matching algorithm can be used to perform streaming matching on real-time session data streams to identify keyword sequences containing specific business intents. The matching results are converted into discrete state features through a pre-defined mapping function, such as generating "business intent type" and "keyword density" features.

[0068] In some embodiments, deep learning models (such as BERT) can be used to transform text into high-dimensional vectors, namely deep semantic vectors, which can capture the implicit information behind the literal meaning and accurately represent the user's potential emotions (such as anger and anxiety) and task complexity.

[0069] In some embodiments, a pre-trained language model based on the Transformer architecture (e.g., BERT, RoBERTa, or a lightweight variant thereof) can be employed. The model structure includes: an input embedding layer (converting segmented tokens into 768-dimensional vectors), a 12-layer Transformer encoder (each layer containing a multi-head self-attention mechanism and a feedforward neural network), and a pooling layer (taking the output of the [CLS] tokens as a sentence-level representation). The model output is a high-dimensional semantic vector (e.g., 768-dimensional), which is used to represent the user's potential emotional state and task complexity.

[0070] The specific training process for the model is as follows: Training data: Sentiment (three categories: positive, neutral, negative) and task complexity (four levels: L1 simple query, L2 multi-turn confirmation, L3 business negotiation, L4 complaint / negotiation) were manually labeled from historical conversation records. A total of no less than 100,000 dialogue turns were labeled.

[0071] Loss function: Cross-entropy loss is used for the sentiment classification task, and mean squared error (MSE) is used for the task complexity regression task, and they are trained together.

[0072] Optimizer: AdamW, learning rate 2e-5, batch size 32, training for 3 epochs.

[0073] Validation: The sentiment classification accuracy reached 89% on the retained test set, and the Pearson correlation coefficient for task complexity prediction reached 0.76.

[0074] During the real-time inference phase, the input session data stream is first preprocessed: special characters are removed, words are segmented, and the data is truncated or padded to a fixed length (e.g., 512 tokens). The preprocessed text sequence is then input into the pre-trained BERT model, which outputs a deep semantic vector. This vector is intrinsically related to the user's emotional state and task complexity: the projection of the vector onto the sentiment dimension can be mapped to a sentiment tendency score (-1 to +1, with negative values ​​representing negative emotions) through a linear classifier; the projection of the vector onto the complexity dimension is normalized using the Sigmoid function and multiplied by 4 to obtain a task complexity score (0~4). Simultaneously, by utilizing the weight distribution of the multi-head attention mechanism, key sentences or keywords that influence sentiment or complexity can be identified.

[0075] Feature Fusion Generation: The keyword rule matching results (discrete features) are fused with deep semantic vectors (continuous features). Specifically, the discrete features are One-Hot encoded and then mapped to a low-dimensional dense vector (e.g., 64-dimensional) through a fully connected layer. This vector is then concatenated with the 768-dimensional deep semantic vector to obtain an 832-dimensional comprehensive state feature vector. This comprehensive vector serves as the input for subsequent switching condition judgments and retrieval matching.

[0076] Using the methods described above, this embodiment transforms unstructured natural language interactions into quantifiable, high-dimensional structured state indicators, enabling refined perception of the conversation's dynamics. The generated state features include not only explicit business intent labels but also implicit sentiment tendencies and task complexity scores, providing a reliable data foundation for subsequent precise switching of execution units.

[0077] This disclosure ensures millisecond-level capture of key intents through keyword rules and accurately understands user emotions and task difficulty using deep semantic vectors. The fusion of these two approaches avoids the blind spots of a single mode, not only improving the accuracy of state features but also dynamically adapting to various scenarios ranging from simple queries to complex disputes, laying a solid foundation for subsequent intelligent scheduling.

[0078] In some embodiments, the preset switching conditions include at least one of the following: The conversation state feature indicates that the user's emotional state index (the user's emotional tendency score calculated based on deep semantic vectors, such as the level of anger and anxiety) is lower than the preset emotional threshold. The emotional state index is calculated based on the user's emotional tendency represented in the deep semantic vectors. The session state characteristics indicate that the task attributes of the current session (the business domain to which the user's intent belongs, such as "legal consultation" or "logistics inquiry") exceed the preset expertise domain of the first execution unit; The session state characteristics indicate when the user sends a preset manual intervention command, or detect when the user explicitly expresses an intention to seek manual assistance; A pre-defined high-priority business keyword sequence was matched in the session data stream; The session state feature indicates that the task complexity of the current session exceeds the preset processing threshold of the first execution unit. The task complexity is determined based on a combination of the number of dialogue rounds, the depth of context dependency, or the confidence level of entity recognition.

[0079] In some embodiments, if an emotional deterioration or human intervention command is triggered, the system can prioritize switching to a highly skilled human agent node to provide empathetic reassurance. If the task is mismatched or its complexity exceeds the limit, the system can switch to a higher-level agent instance with domain expertise or higher computing power. For high-priority keyword scenarios, the system can directly route to a VIP dedicated service channel (usually a senior human agent or an expert-level agent). This dynamic switching mechanism ensures that simple problems are handled quickly by AI, while complex / urgent / emotional problems are seamlessly taken over by humans or experts, achieving optimal resource allocation and a safety net for service experience.

[0080] Figure 2 A flowchart of a session processing method according to an embodiment of this disclosure is shown, such as Figure 2 As shown, the session processing method includes S201-S207, wherein S201-S204 and S207 are the same as S101-S105 above, and will not be described again here. In this embodiment, after retrieving and determining the second execution unit from the preset execution unit pool according to the session state characteristics, and before receiving the feedback signal confirming the takeover from the second execution unit, S205-S206 are also included.

[0081] In S205, when the session state characteristics indicate that the user's emotional state index is lower than the preset emotional threshold, the first execution unit controls the call of the preset emotional soothing strategy library to generate response text.

[0082] The response text contains at least one of the following: expressions of empathy, apologies, or reassuring guidance.

[0083] In S206, the response text is sent to the user.

[0084] Once the session state characteristics determine that the user's emotions have deteriorated (below a preset threshold), the current first execution unit (such as an agent) can proactively invoke the emotion soothing strategy library to generate response text containing empathetic expressions, apologies, or guiding information and send it to the user. This process occurs during the "window period" before the new execution unit (second execution unit) confirms its takeover, aiming to alleviate the user's negative emotions through rapid emotional feedback.

[0085] This disclosed embodiment fills the emotional response vacuum during the conversation transition, achieving an upgrade from a mechanical to a warm transition. By utilizing a pre-built strategy library to instantly output empathetic and reassuring content during the wait for human or higher-level intelligent agents to take over, it effectively reduces user anxiety and dissatisfaction, preventing experience collapse caused by queuing. This not only enhances the warmth and humanization of the service but also significantly improves user retention during the waiting period, creating a more stable psychological foundation for a seamless handover later.

[0086] In some embodiments, after setting the second execution unit as the dominant response node of the current session, it may further include Figure 3 S301-S302 are shown.

[0087] In S301, based on the historical session context of the first execution unit, session summary information or key feature snapshots are generated; In S302, session summary information or key feature snapshots are pushed to the second execution unit so that the second execution unit can inherit the previous session state and continue to process the session data stream.

[0088] Historical session context refers to all interaction records from the start of the session to the moment of switchover triggering, including the sequence of messages sent by the user, the response text generated by the first execution unit, timestamps, round identifiers, and intermediate state data (such as entity extraction results and intent classification labels). To achieve seamless inheritance when switching to the second execution unit, the most essential information needs to be extracted from the above context and compressed into a session summary or key feature snapshot.

[0089] In some embodiments, there may be multiple methods for generating session summary information.

[0090] As an example, session summary information can be generated by obtaining the complete historical session context. The system maintains a session cache, storing the interaction records of each session in chronological order. Each record includes: role identifier (user / first execution unit), message content, message timestamp, and dialogue round number. When the session state characteristics meet the switching conditions, the system reads the entire historical message sequence of the current session from the cache.

[0091] As an example, conversation summaries can be generated based on generative language models. These generative language models can employ pre-trained generative summarization models (such as BART, T5, or GPT series). Taking BART-base as an example, the model structure includes a 12-layer encoder and a 12-layer decoder. The encoder uses a bidirectional self-attention mechanism, and the decoder uses a unidirectional self-attention mechanism, connected by cross-attention links. The total number of model parameters is approximately 140 million.

[0092] Training data: Summarized manually from historical customer service conversation records. Each conversation corresponds to a reference summary, with a summary length controlled between 80 and 200 characters. Annotation rules: The summary must include the user's core request, confirmed key information (such as order number, product name), current resolution progress (such as answered question A, unanswered question B), and user emotional state tags. At least 50,000 conversation-summary pairs were annotated.

[0093] Input format: Concatenate historical sessions into a sequence "User: ...\nCustomer Service: ...\nUser: ...", adding a [CLS] flag at the beginning and a [SEP] flag at the end. Maximum input length: 1024 tokens.

[0094] Output format: The decoder generates a digest text for each token, ending with [EOS].

[0095] Loss function: Cross-entropy loss, based on standard autoregressive methods to predict the next token.

[0096] Optimizer: AdamW, learning rate 3e-5, linear learning rate decay, batch size 16, training for 3 epochs.

[0097] Evaluation metrics: ROUGE-1 / 2 / L reached 0.52 / 0.38 / 0.49 respectively, and the retention rate of manually evaluated content (the proportion of key information not lost) was 94%.

[0098] Inference phase: Input the historical session context into the model in the same format, and generate summary text using beam search (beam size=4). The generated summary is pushed to the second execution unit as session summary information.

[0099] A key feature snapshot is a structured data object used to supplement information that is difficult to quantify in a summary. Snapshots are generated without relying on complex models, but are extracted or computed directly from the session context. A snapshot contains the following fields: User profile tags: Read from the user information database, such as VIP level, historical purchase records, and common question preferences.

[0100] The list of identified entities, such as order number ORD-123456, product model X1000, amount 5000 yuan, etc., is derived from the results extracted by the first execution unit during the dialogue process using a Named Entity Recognition (NER) model (such as BiLSTM-CRF).

[0101] Unresolved issues list: Based on issues that have not been answered or explicitly denied in the last few rounds of user messages, these issues are automatically extracted using rules (such as detecting questions that are not directly answered in subsequent replies).

[0102] Emotional state trajectory: Records the sequence of changes in emotional tendency values ​​during the conversation (e.g., initial 0.2, mid-term -0.5, current -0.3), which comes from the time series of emotional scores mapped from the semantic vector in S102 above.

[0103] List of completed operations / provided answers: Extract confirmation information (such as "Found for you...", "Sent materials...") from the response of the first execution unit.

[0104] The system organizes the above fields into snapshot data in JSON format. The snapshot size is much smaller than the original conversation record, facilitating rapid transmission and parsing.

[0105] After generating session summary information and key feature snapshots, in S302, both are pushed to the second execution unit. Upon receiving them, the second execution unit inherits the previous session state in the following manner: First, the key feature snapshots are parsed to restore the structured state of the session (to-do items, entities, emotions, etc.) and directly loaded into memory as the initial state.

[0106] Then, read the conversation summary information to quickly understand the semantic content and development of the conversation, avoiding excessive length.

[0107] When the second execution unit needs to backtrack to the details of a specific round of dialogue, it can optionally request a complete record of the original context from the cache via the session ID (as a supplement, it does not affect real-time processing).

[0108] Through the above design, the second execution unit can reconstruct a session state that is almost identical to that of the first execution unit within milliseconds, without requiring the user to repeat the input, thus achieving seamless inheritance across execution units.

[0109] This disclosure achieves seamless state inheritance during execution unit switching by extracting historical context to generate summaries or snapshots. The second execution unit can quickly grasp the background and key features of the preceding session without requiring repeated user statements, eliminating service breakpoints. This not only significantly reduces the cost of repetitive communication for users and improves interaction efficiency, but also ensures the continuity and accuracy of complex task processing, significantly optimizing the overall service experience of multi-stage, cross-node sessions.

[0110] In some embodiments, after receiving a feedback signal confirming takeover from the second execution unit, the above session processing method may further include sending a session transfer notification message to the user, the notification message being used to inform the user that the current session is being taken over by a new execution unit; or, without interrupting the session data stream, downgrading the control authority of the first execution unit, so that the second execution unit can smoothly take over the conversation leadership.

[0111] In some embodiments, the above session processing method may further include Figure 4 S401-S405 are shown.

[0112] In S401, each time a user's session data stream is received or a response text is generated, the last active timestamp corresponding to the current session is updated; In S402, a background scheduling task with a preset cycle is executed, traversing all pending session lists and obtaining the last active timestamp of each session. In S403, for each session that is traversed, the time difference between the current time and the last active timestamp is calculated; In S404, if the time difference exceeds the preset silence timeout threshold, the first execution unit or the designated wake-up agent instance is controlled to retrieve and generate the corresponding wake-up prompt message from the preset session activation content library. In S405, a wake-up notification message is sent to the user.

[0113] This embodiment of the disclosure dynamically monitors session activity and proactively triggers a wake-up strategy using a silent timeout mechanism. This effectively prevents the natural loss of sessions caused by prolonged silence, transforming passive waiting into proactive care, and re-stimulating users' willingness to interact. This not only significantly improves session retention and conversion rates but also optimizes resource scheduling efficiency, ensuring that intelligent services remain online when potential user needs arise, and maximizing service reach value.

[0114] In some embodiments, when the dominant response node of the current session is a human interaction node, the above session processing method may further include: Figure 5 S501-S503 are shown.

[0115] In S501, the time difference between the current time and the last operation timestamp of the human interaction node on the current session data stream is calculated; In S502, if the time difference exceeds the preset manual response timeout threshold, the configuration rules in the preset waiting soothing strategy library are invoked to generate a soothing prompt message. The soothing prompt message includes at least one of the following: queuing progress information, apology and explanation information, or interactive guidance information. In S503, the generated reassuring message is sent to the user.

[0116] This disclosure presents a dynamic reassurance mechanism for scenarios with delayed responses from human agents. By monitoring operation timestamps in real time, a reassurance message containing queuing progress, apology, or guidance is automatically pushed when a timeout threshold is triggered. This effectively alleviates user anxiety and dissatisfaction while waiting for human assistance, eliminates the negative experience of waiting in a black box, and reduces the probability of users hanging up or complaining due to impatience. It significantly improves user satisfaction with human assistance and session retention, achieving a smooth transition in the service process.

[0117] In some embodiments, the above session processing method may further include Figure 6 S601-S605 are shown.

[0118] In S601, the intent confidence and sentiment score in the conversation state features are analyzed; In S602, if the confidence level of the intent is higher than the preset high priority threshold, and a keyword sequence representing transaction-oriented behavior is matched in the deep semantic vector, then the first type of interaction intent label is generated. In S603, if the sentiment tendency score is lower than the preset negative sentiment threshold and a keyword sequence representing disagreement and conflict is matched in the deep semantic vector, then a second type of interaction risk label is generated. In S604, session metadata with corresponding tags is written into a preset classification index library to build a high-priority interaction queue and an abnormal risk queue. In S605, an aggregated view of the high-priority interaction queue and the abnormal risk queue is pushed to the management terminal.

[0119] This disclosed embodiment achieves intelligent grading and risk warning of conversations through multi-dimensional semantic analysis (intent confidence, sentiment tendency), accurately identifies high-value transaction opportunities, prioritizes the service experience of high-priority users, improves conversion rate, and at the same time captures potential conflicts and complaint risks in real time, builds an abnormal queue and pushes it to the management end, realizes differentiated scheduling of resources, effectively reduces customer complaint rate, and optimizes overall operational efficiency and service quality.

[0120] The following specific example will illustrate in detail the implementation process of this disclosed session handling method.

[0121] First, you can configure the AI ​​smart managed parameters. As an example, the managed parameters can include those in Table 1 below.

[0122] Table 1

[0123] After configuring the hosting parameters, when the system receives a message from a customer, it can first check the hosting status. If the hosting status is 1 (hosting in progress), it forwards the message to the AI ​​agent. Then, the agent generates a reply based on the preset role and sends it to the customer. Furthermore, sensitive word detection can also be performed during the above process.

[0124] The dialogue process can also be analyzed in real time to determine customer intent, such as high intent and negative emotions.

[0125] As an example, the agent analyzes customer conversations in real time, identifying high-intent keywords in Table 2 and negative sentiment keywords in Table 3.

[0126] Table 2

[0127] Table 3

[0128] When high-intent keywords are identified, the system interface can be called back to query the configuration to determine whether to perform a switching operation.

[0129] After performing the switch operation, you can update the database hosting=0, delete the Redis hosted key (effective in milliseconds), add it to the prospective customer table, and notify the sales staff.

[0130] In some embodiments, the aforementioned session can be a session with a single user or a session between chat groups.

[0131] In some embodiments, the system can also automatically classify based on the recognition results: High-intent customer list: Records high-intent customers who inquire about prices, show interest in cooperation, or have a strong desire to purchase; Negative Emotion Customer List: Records customers who express negative emotions such as complaints, rejections, and doubts.

[0132] Sales staff can view the lists of high-intent customers and negative-emotion customers through the system and take targeted actions.

[0133] In some embodiments, the system also records the last response time of each customer, and a scheduled task scans the managed customer list. If the silence time threshold (also known as the silent timeout threshold) is exceeded and the previously configured parameter silence activation = 1, then silence content is automatically sent to reactivate the conversation, and then the timer is reset to continue monitoring.

[0134] This disclosed embodiment can accurately identify high-intent signals, improving the identification accuracy to over 85% and effectively avoiding human subjective bias. Leveraging Redis's millisecond-level state switching mechanism, the system response speed is optimized from several minutes to less than 100ms, enabling 24 / 7 automated follow-up and personalized scenario adaptation. Furthermore, the automatic activation during periods of inactivity and real-time sensitive word detection not only significantly improve customer activity and compliance but also drive an overall conversion rate increase of 40%, achieving a qualitative leap from passive response to proactive intelligent marketing.

[0135] Based on the same inventive concept, this disclosure also provides a session processing apparatus, such as... Figure 7 As shown, the session processing device includes a session data acquisition module 701, a state feature analysis module 702, an execution unit retrieval module 703, a takeover request sending module 704, and a node switching control module 705.

[0136] The session data acquisition module 701 is configured to acquire the real-time session data stream between the first execution unit and the user. The state feature analysis module 702 is configured to continuously detect the real-time session data stream and generate session state features based on the detection results. The execution unit retrieval module 703 is configured to retrieve and determine a second execution unit from a preset execution unit pool based on the session state characteristics if the session state characteristics meet the preset switching conditions. The second execution unit has different expertise, processing logic configuration or service level from the first execution unit. The takeover request sending module 704 is configured to send a takeover request to the determined second execution unit; The node switching control module 705 is configured to set the second execution unit as the dominant response node of the current session after receiving a feedback signal confirming takeover from the second execution unit, so as to continue processing the session data stream with the user.

[0137] Regarding the session processing apparatus in the above embodiments, the specific methods by which each module performs operations have been described in detail in the embodiments related to the session processing method, and will not be elaborated here.

[0138] It should be noted that although several modules or units for action execution have been mentioned in the detailed description above, this division is not mandatory. In fact, according to embodiments of this disclosure, the features and functions of two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.

[0139] Furthermore, some of the block diagrams shown in the attached figures are functional entities and do not necessarily correspond to physically or logically independent entities.

[0140] Based on the same inventive concept, embodiments of this disclosure also provide an electronic device. For example... Figure 8 As shown, the electronic device provided in this embodiment includes a processor 801 and a memory 802; the memory 802 is used to store program instructions; the processor 801 is used to call the program instructions stored in the memory 802 to implement the session processing method described in the above method embodiment.

[0141] In some embodiments, processor 801 is a device with data processing capabilities, including but not limited to a central processing unit (CPU), a field-programmable gate array (FPGA), and other programmable logic devices. Processor 801 may include one or more processing cores.

[0142] In some embodiments, the memory 802 is a device with data storage capability, including but not limited to random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), and flash memory.

[0143] In some embodiments, the processor 801 and memory 802 can be configured separately or integrated together. As an example, the processor 801 and memory 802 can be integrated on a single board or a system on a chip (SOC).

[0144] like Figure 8 As shown, the electronic device provided in this embodiment may further include an interface 803. The interface 803 is connected between the processor 801 and the memory 802, enabling information exchange between the processor 801 and the memory 802. In some embodiments, the processor 801, the memory 802, and the interface 803 are interconnected via a bus 804.

[0145] In some embodiments, there may be one or more processors 801.

[0146] In some embodiments, there may be one or more memories 802.

[0147] In some embodiments, there may be one or more interfaces 803.

[0148] In some embodiments, the memory 802 stores one or more program instructions. When the one or more program instructions are executed by the one or more processors 801, the one or more processors 801 implement the session processing method described in the above method embodiments and achieve the corresponding technical effects.

[0149] In some embodiments, the electronic device may also communicate with one or more external devices, such as a keyboard, pointing device, Bluetooth device, etc. In some embodiments, the electronic device may also communicate with one or more devices that enable a user to interact with the electronic device. In some embodiments, the electronic device may also communicate with devices that enable the electronic device to communicate with one or more other computing devices. In some embodiments, the above communication may be performed through interface 803.

[0150] Understandable, Figure 8 The illustrated structure does not constitute a specific limitation on the electronic device. In some embodiments, the electronic device may include... Figure 8 The number of components shown may be more or fewer. In some embodiments, the electronic device has... Figure 8 Based on the diagram, some components can be combined, some components can be separated, or different component arrangements can be made.

[0151] In some embodiments, Figure 8 The components shown can be implemented in hardware, software, or a combination of both.

[0152] Based on the same inventive concept, this disclosure also provides a computer-readable storage medium storing program instructions. When the program instructions are executed by a computer, they implement the session processing method described in the above method embodiments, thereby achieving the corresponding technical effects.

[0153] In some embodiments, the computer-readable storage medium can be any available medium capable of storing program instructions or a data storage device such as a data center containing one or more available media. In some embodiments, the available medium can be a magnetic medium, an optical medium, or a semiconductor medium, etc. As an example, a magnetic medium can be a floppy disk, a hard disk, a magnetic tape, etc. As an example, an optical medium can be a high-density digital video optical disc. As an example, a semiconductor medium can be a solid-state drive.

[0154] Based on the same inventive concept, this disclosure also provides a computer program product, which includes program instructions for implementing the session processing method described in the above method embodiments. When executed by a computer, the program instructions cause the computer to implement the session processing method described in the above method embodiments, achieving the corresponding technical effects. In specific implementations, the program instructions can be written using any combination of one or more programming languages. The program instructions can be executed entirely on the user device, partially on the user device, as a standalone software package, partially on the user device and partially on a remote device, or entirely on a remote device or server.

[0155] Those skilled in the art will understand that all or part of the steps of the above embodiments can be specifically implemented in the following forms: a completely hardware implementation, a completely software implementation, or an implementation combining hardware and software aspects.

[0156] The above description is merely a specific embodiment of this disclosure, but the scope of protection of this disclosure is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this disclosure should be included within the scope of protection of this disclosure.

Claims

1. A session processing method, characterized in that, include: Acquire the real-time session data stream between the first execution unit and the user; The session data stream is continuously monitored, and session state features are generated. If the session state characteristics meet the preset switching conditions, then a second execution unit is retrieved and determined from the preset execution unit pool according to the session state characteristics, wherein the second execution unit has different expertise, processing logic configuration or service level from the first execution unit; Send a takeover request to the second execution unit; Upon receiving a feedback signal confirming takeover from the second execution unit, the second execution unit is set as the dominant response node for the current session to continue processing the session data stream with the user.

2. The session processing method according to claim 1, characterized in that, The step of retrieving and determining the second execution unit from a preset execution unit pool based on the session state characteristics includes: Calculate the similarity score between the session state features and the capability description vectors of each candidate execution unit; wherein the capability description vectors represent the expertise, processing logic configuration, or service level of each candidate execution unit. The candidate execution unit with the highest similarity score and that meets the preset matching threshold is selected as the second execution unit.

3. The session processing method according to claim 2, characterized in that, The preset execution unit pool includes multiple agent instances and / or multiple human interaction nodes; The step of retrieving and determining the second execution unit from the preset execution unit pool based on the session state characteristics includes: determining the second execution unit from the candidate execution units based on the session state characteristics and the real-time running status of each candidate execution unit in the execution unit pool; The real-time operating status includes at least one of the following data: The concurrent processing load of the agent instance, the response latency metric of the agent instance, or the current idle time of the human interaction node.

4. The session processing method according to claim 1, characterized in that, The continuous monitoring of the session data stream and the generation of session state features include: The session data stream is matched and detected in real time using a pre-built keyword rule base to identify keyword sequences containing specific business intentions, and the matching results are mapped to the session state features. Alternatively, a deep semantic vector can be extracted from the session data stream. This deep semantic vector is used to characterize the user's potential emotional state or task complexity, and the deep semantic vector is determined as the session state feature; the second dimension feature vector. Alternatively, a pre-built keyword rule base can be used to perform real-time matching and detection on the session data stream to identify keyword sequences containing specific business intentions, extract deep semantic vectors from the session data stream, and generate the session state features based on the keyword sequences and the deep semantic vectors.

5. The session processing method according to claim 4, characterized in that, The preset switching conditions include at least one of the following: The session state feature indicates that the user's emotional state index is lower than a preset emotional threshold. The emotional state index is calculated based on the user's emotional tendency represented in the deep semantic vector. The session state characteristics indicate that the task attributes of the current session exceed the preset expertise area of ​​the first execution unit; The session state characteristics indicate that the user sends a preset manual intervention command, or detect that the user clearly expresses the intention to seek manual assistance; The session data stream was matched with a preset high-priority business keyword sequence; The session state feature indicates that the task complexity of the current session exceeds the preset processing threshold of the first execution unit. The task complexity is determined based on a combination of dialogue rounds, context dependency depth, or entity recognition confidence.

6. The session processing method according to claim 5, characterized in that, After retrieving and determining the second execution unit from a preset execution unit pool based on the session state characteristics, and before receiving a feedback signal confirming takeover from the second execution unit, the method further includes: When the session state characteristics indicate that the user's emotional state index is lower than the preset emotional threshold, the first execution unit is controlled to call the preset emotional soothing strategy library to generate response text containing empathetic expressions, apology statements or soothing guidance information. The response text is sent to the user.

7. The session processing method according to claim 1, characterized in that, After setting the second execution unit as the dominant response node for the current session, the method further includes: Based on the historical session context of the first execution unit, generate session summary information or key feature snapshots; The session summary information or key feature snapshot is pushed to the second execution unit so that the second execution unit can inherit the previous session state and continue to process the session data stream.

8. The method according to claim 1, characterized in that, After receiving the feedback signal confirming takeover from the second execution unit, the method further includes: Send a session transfer notification message to the user, the notification message being used to inform the user that the current session is being taken over by a new execution unit; Alternatively, without interrupting the session data stream, the control authority of the first execution unit can be downgraded, allowing the second execution unit to smoothly take over the control of the conversation.

9. The session processing method according to claim 1, characterized in that, The method further includes: Update the last active timestamp of the current session each time a user's session data stream is received or a response text is generated; Execute background scheduling tasks with a preset cycle, traverse all pending session lists, and obtain the last active timestamp of each session; For each session iterated through, calculate the time difference between the current time and the last active timestamp; If the time difference exceeds the preset silence timeout threshold, the first execution unit or the designated wake-up agent instance is controlled to retrieve and generate the corresponding wake-up prompt message from the preset session activation content library. The wake-up notification message is sent to the user.

10. The session processing method according to claim 1, characterized in that, When the dominant response node in the current session is a human-interacting node, the method further includes: Calculate the time difference between the current time and the timestamp of the last operation of the human interaction node on the current session data stream; If the time difference exceeds the preset manual response timeout threshold, the configuration rules in the preset waiting and soothing strategy library are invoked to generate a soothing prompt message. The soothing prompt message includes at least one of the following: queuing progress information, apology and explanation information, or interactive guidance information. The generated reassuring message is sent to the user.

11. The session processing method according to claim 1, characterized in that, The method further includes: Analyze the intent confidence and sentiment scores in the aforementioned conversation state features; If the confidence level of the intent is higher than the preset high priority threshold, and a keyword sequence representing transaction-oriented behavior is matched in the deep semantic vector, then a first type of interactive intent label is generated. If the sentiment score is lower than the preset negative sentiment threshold, and a keyword sequence representing disagreement and conflict is matched in the deep semantic vector, then a second type of interaction risk label is generated. Write session metadata with corresponding tags into a pre-defined classification index library to build a high-priority interaction queue and an anomaly risk queue; An aggregated view of the high-priority interaction queue and the abnormal risk queue is pushed to the management terminal.

12. A computer program product, characterized in that, The computer program product includes program instructions for implementing the session processing method as described in any one of claims 1 to 11.