A bimodal customer service reply generation method and system across visual angle intelligent reasoning
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- 安徽三七极光网络科技有限公司
- Filing Date
- 2026-05-19
- Publication Date
- 2026-08-07
AI Technical Summary
1、单一C端视角答复的局限性:传统客服系统生成的直接面向用户(C端)的答复,往往仅提供简洁明了的答案
[0011]与现有技术相比,本发明具有以下技术效果的至少之一:
Smart Images

Figure CN122529092A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, and in particular to a bimodal customer service response generation method and system for cross-perspective intelligent reasoning. Background Technology
[0002] In traditional customer service systems, AI-generated answers are usually limited to a single perspective. This situation exposes many drawbacks when facing complex and ever-changing user inquiry scenarios, seriously affecting the quality and efficiency of customer service. The specific manifestations are as follows: 1. Limitations of a Single-Perspective (C-end) Customer Service Response: Traditional customer service systems generate responses directly to users (C-end), often providing only concise and clear answers. This approach has advantages when handling simple issues, quickly meeting users' basic information needs and reducing their time spent obtaining answers. However, its limitations become apparent when dealing with complex issues. Due to a lack of detailed background information and explanations, users may struggle to fully understand the response. For example, in inquiries involving complex areas such as financial product rules and technical principles, a concise response cannot cover all key points and potential influencing factors. Users may have more questions due to insufficient information, or even doubt the authenticity and reliability of the response, leading to incomplete problem resolution, repeated communication, increased time and communication costs for users, and a decreased user experience.
[0003] 2. Limitations of a Single B2B Perspective in Responses: Traditional responses to customer service representatives (B2B) provide detailed knowledge points and background information. From the perspective of providing knowledge support to customer service representatives, this helps them gain a deep understanding of all aspects of the problem, laying the foundation for accurate answers to user questions. However, in actual customer service work, representatives need to spend a significant amount of extra time translating this complex knowledge into language suitable for users. When faced with a large number of user inquiries, this extra time consumption significantly reduces the efficiency of customer service representatives. For example, during e-commerce promotional events, the number of user inquiries increases dramatically. If customer service representatives need to translate every complex response, it will inevitably lead to excessively long waiting times for users, affecting their shopping experience and potentially causing customer churn. Moreover, under time pressure, customer service representatives may rush to answer due to insufficient knowledge translation, resulting in inconsistent response quality, failure to accurately and clearly convey key information, further impacting user satisfaction and damaging the company's image.
[0004] 3. Lack of intelligent reasoning and dynamic adjustment capabilities: Traditional customer service systems lack the ability to comprehensively analyze and intelligently reason based on multi-dimensional user information when generating responses. The generated responses lack specificity and fail to adequately meet the actual needs of users. Furthermore, when determining the content of the responses, their adaptability and flexibility across different scenarios are insufficient, making it impossible to provide personalized and high-quality service.
[0005] 4. Lack of effective response optimization and feedback mechanisms: Traditional customer service systems lack further optimization of responses after they are generated, making it difficult for customer service personnel to quickly grasp the key points of the response, affecting the accuracy and efficiency of the response. Furthermore, the absence of a robust feedback mechanism leads to slow system performance improvements, making it difficult to adapt to constantly changing user needs and market environments. Summary of the Invention
[0006] The purpose of this invention is to provide a bimodal customer service response generation method and system based on cross-perspective intelligent reasoning, which realizes cross-perspective intelligent reasoning and bimodal customer service response generation. It can provide responses adapted to different scenario needs for users and customer service by integrating multi-dimensional information, thereby improving the quality and efficiency of customer service and solving at least one of the aforementioned problems in the prior art.
[0007] In a first aspect, the present invention provides a dual-modal customer service response generation method for cross-perspective intelligent reasoning, the method specifically comprising: Real-time acquisition and analysis of users' voice tone and visual expression features when submitting questions, combined with question text to identify users' emotional state and question intent; Based on the user's question intent, relevant B-end knowledge nodes are retrieved from the unified knowledge base. A pre-trained graph neural network model is used to perform cross-perspective association reasoning on the B-end knowledge nodes, mapping out the associated C-end verbal nodes, forming a set of B-end node pairs with association weights. Based on user emotional state, historical user profiles, and question complexity, a dynamic weight adjustment algorithm is used to calculate the display weight ratio of B-end content and C-end content in the current scenario. Based on the display weight ratio, content is extracted from the BC node pair set to generate a dual-perspective response draft. The draft response from both perspectives is input into the intelligent summary model. Through difference extraction and comparative analysis, a natural language summary and structured comparison view that highlight the core information differences between the B-end and C-end are generated and presented to customer service personnel.
[0008] Secondly, this invention provides a dual-modal customer service response generation system with cross-perspective intelligent reasoning, the system specifically comprising: The data acquisition module is used to acquire and analyze the voice tone and visual expression features of users when they submit questions in real time, and combine the question text to identify the user's emotional state and the user's question intent. The intent analysis module is used to retrieve relevant B-end knowledge nodes from a unified knowledge base based on the user's question intent. It uses a pre-trained graph neural network model to perform cross-perspective association reasoning on the B-end knowledge nodes, mapping out the associated C-end verbal nodes and forming a set of B-end node pairs with association weights. The draft generation module is used to calculate the display weight ratio of B-end content and C-end content in the current scenario based on the user's emotional state, historical user profile and question complexity, and extract content from the BC node pair set according to the display weight ratio to generate a dual-perspective response draft. The response output module is used to input the dual-perspective response draft into the intelligent summary model. Through difference extraction and comparative analysis, it generates a natural language summary and structured comparison view that highlights the core information differences between the B-end and C-end, and presents them to customer service personnel.
[0009] Thirdly, the present invention provides a computer device, comprising: a memory and a processor, and a computer program stored in the memory, wherein when the computer program is executed on the processor, it implements a bimodal customer service response method for cross-perspective intelligent reasoning as described in any of the above methods.
[0010] Fourthly, the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements a bimodal customer service response method for cross-perspective intelligent reasoning as described in any of the above methods.
[0011] Compared with the prior art, the present invention has at least one of the following technical effects: 1. This invention realizes cross-perspective intelligent reasoning and dual-modal customer service response generation, which can provide users and customer service with responses that adapt to different scenario needs by integrating multi-dimensional information, thereby improving the quality and efficiency of customer service.
[0012] 2. This invention accurately identifies the user's emotional state and question intent by collecting multiple signals and extracting multiple feature vectors for joint reasoning, providing a foundation for generating targeted responses.
[0013] 3. This invention utilizes a cross-modal attention mechanism to generate joint feature vectors and simultaneously output classification labels, thereby enhancing the accuracy and comprehensiveness of user emotion and intent recognition.
[0014] 4. This invention generates structured query conditions through deep semantic parsing and integrates multiple retrieval methods to accurately retrieve a set of B-end knowledge nodes that meet the requirements.
[0015] 5. This invention constructs a heterogeneous graph knowledge base and uses a graph neural network model for reasoning to form a set of B-C node pairs with associated weights, thereby achieving an effective mapping from B-end knowledge to C-end rhetoric.
[0016] 6. This invention integrates multi-factor quantitative encoding and calculates the content display weight ratio between B-end and C-end to suit the current scenario through neural networks and business rule engines.
[0017] 7. This invention extracts and assembles content according to the display weight ratio to generate a dual-perspective response draft, ensuring that the response content meets the needs of the scenario and is logically clear.
[0018] 8. This invention generates natural language summaries and structured comparison views that highlight the differences in core information through parsing, difference analysis, and view template selection, making it easier for customer service personnel to quickly grasp the key points.
[0019] 9. This invention generates samples by capturing operation and feedback data to perform incremental learning on the model and algorithm, thereby achieving continuous optimization of the system and improving the quality of responses. Attached Figure Description
[0020] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0021] Figure 1 This is a flowchart illustrating a bimodal customer service response method based on cross-perspective intelligent reasoning, provided in an embodiment of the present invention. Figure 2 This is a schematic diagram of the structure of a dual-modal customer service response system with cross-perspective intelligent reasoning provided in an embodiment of the present invention; Figure 3 This is a schematic diagram of the structure of a computer device provided in an embodiment of the present invention. Detailed Implementation
[0022] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.
[0023] It should be understood that, when used in this application specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or a collection thereof.
[0024] It should also be understood that the term “and / or” as used in this application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.
[0025] As used in this application specification and the appended claims, the term "if" may be interpreted, depending on the context, as "when," "once," "in response to determination," or "in response to detection." Similarly, the phrase "if determined" or "if detected [the described condition or event]" may be interpreted, depending on the context, as meaning "once determined," "in response to determination," "once detected [the described condition or event]," or "in response to detection [the described condition or event]."
[0026] Furthermore, in the description of this application and the appended claims, the terms "first," "second," "third," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.
[0027] References to "one embodiment" or "some embodiments" as described in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.
[0028] In this application embodiment, the entity executing the process includes a terminal device. This terminal device includes, but is not limited to, devices capable of executing the methods disclosed in this application, such as servers, computers, smartphones, and tablets. Figure 1 A flowchart illustrating a bimodal customer service response method based on cross-perspective intelligent reasoning, according to an embodiment of the present invention, is shown below in detail: S101 acquires and analyzes the voice tone and visual expression features of users when submitting questions in real time, and combines the question text to identify the user's emotional state and the user's question intent. S102, based on the user's question intent, retrieve relevant B-end knowledge nodes from the unified knowledge base, use a pre-trained graph neural network model to perform cross-perspective association reasoning on the B-end knowledge nodes, map out the associated C-end verbal nodes, and form a set of BC node pairs with association weights. S103: Based on the user's emotional state, historical user profile, and question complexity, the display weight ratio of B-end content and C-end content in the current scenario is calculated using a dynamic weight adjustment algorithm. Based on the display weight ratio, content is extracted from the BC node pair set to generate a dual-perspective response draft. S104. Input the dual-perspective response draft into the intelligent summary model. Through difference extraction and comparative analysis, generate a natural language summary and structured comparison view that highlights the core information differences between the B-end and C-end, and present it to customer service personnel.
[0029] In this embodiment, in a customer service interaction scenario, when a user submits a question, the system simultaneously activates a multimodal information acquisition module. This module utilizes speech recognition technology to acquire the user's voice data in real time and analyzes its tone and intonation characteristics. For example, by analyzing parameters such as pitch, volume, and speech rate, it determines whether the user is currently in a calm, excited, angry, or confused emotional state. Simultaneously, using computer vision technology, it captures the user's visual facial expressions through a camera, such as subtle changes in facial muscles and eye movements, further assisting in judging the user's emotions. While acquiring voice and intonation characteristics and visual facial expressions, the system performs natural language processing on the user's submitted question text, using semantic analysis, keyword extraction, and other technologies to accurately identify the user's question intent. For example, if a user asks, "How well does this phone take photos when the battery is low?", the system can identify that the question intent is about the phone's performance in a specific usage scenario. By comprehensively analyzing voice tone, visual expressions, and question text, the system can comprehensively and accurately grasp the user's current emotional state and question intent, providing crucial evidence for generating subsequent responses.
[0030] Based on the identified user intent, the system retrieves relevant knowledge from a unified knowledge base. This base stores a wealth of B2B knowledge nodes, covering detailed information such as product technical specifications, business processes, and policies and regulations. For example, regarding the aforementioned mobile phone photography question, the system will retrieve B2B knowledge nodes related to the phone's camera hardware parameters and the principles of its photography algorithms. After retrieving these B2B knowledge nodes, the system uses a pre-trained graph neural network model to perform cross-perspective association reasoning on these nodes. The graph neural network model can analyze the inherent connections between B2B knowledge nodes and map these connections to associated C2C verbal nodes. C2C verbal nodes are user-friendly and easily understood language expressions. For example, regarding mobile phone camera hardware parameters, the corresponding C2C verbal node might be "This phone uses an advanced [specific model] camera, which can bring clearer photo results." During the mapping process, the system assigns a correlation weight to each B2B node pair, reflecting the degree of correlation between the B2B knowledge node and the C2C verbal node. In this way, the system forms a set of B2B node pairs with correlation weights, providing rich knowledge material for generating responses that meet user needs.
[0031] Based on user emotional state, historical user profiles, and question complexity, a dynamic weight adjustment algorithm calculates the display weight ratio of B2B and B2C content in the current scenario. User emotional state is a crucial factor. If the user is agitated, the system will appropriately increase the display weight of B2C content to soothe the user and answer questions in a more accessible and friendly manner. If the user is calm and shows interest in professional knowledge, the system will appropriately increase the display weight of B2B content to provide more detailed and accurate professional information. Historical user profiles include basic user information, past consultation records, and purchasing preferences. The system uses this information to determine the user's acceptance level of B2B knowledge and B2C language, thereby adjusting the display weight ratio. Question complexity is also a key factor. For simple questions, the system can increase the weight of B2C content to quickly provide a concise and clear answer; for complex questions, it increases the weight of B2B content to ensure the comprehensiveness and accuracy of the answer. Based on the calculated display weight ratio, the system extracts relevant content from the B2B node pair set, organically integrating B2B knowledge and B2C language to generate a dual-perspective response draft. For example, regarding the issue of mobile phone photography, if the displayed weighting is biased towards consumers (C-end), the draft response might be, "This phone has excellent photo quality. It uses an advanced [specific model] camera that can take clear photos even when the battery is low." If the weighting is biased towards businesses (B-end), the draft response might be, "The phone's camera hardware parameters are [specific parameters], and its photography algorithm is based on [specific principles]. In low-power mode, through optimized power management, it can still ensure a certain level of photography performance, specifically as shown in [detailed performance indicators]."
[0032] The generated dual-perspective response draft is input into the intelligent summarization model. The model first extracts differences between the B-end and C-end content, analyzing variations in presentation and information focus. For example, B-end content might emphasize technical parameters and principles, while C-end content focuses more on user experience and practical effects. Then, through comparative analysis, a natural language summary highlighting the core differences between the B-end and C-end content is generated. This summary concisely summarizes the key points of both content, allowing customer service personnel to quickly understand the core information. Simultaneously, the system generates a structured comparison view, displaying the differences between B-end and C-end content in intuitive charts or tables. For instance, technical parameters and user experience descriptions are listed in separate columns, enabling customer service personnel to easily compare and analyze the differences. Finally, the system presents the natural language summary and structured comparison view to customer service personnel, allowing them to quickly grasp the key points of the response, accurately understand the characteristics of both B-end and C-end content, and thus better communicate with users, improving the quality and efficiency of customer service.
[0033] In some embodiments, step S101 above, which involves acquiring and analyzing the user's voice tone and visual expression features when submitting a question in real time, and combining the question text to identify the user's emotional state and question intent, specifically includes: Collect audio and video signals when users submit questions and obtain the corresponding question text; perform noise reduction and alignment preprocessing on the audio and video signals. Based on the preprocessed audio and video signals and the question text, acoustic feature vectors encoding speech intonation are extracted through an acoustic model, visual feature vectors encoding facial expressions are extracted through a visual model, and textual semantic feature vectors encoding the meaning of the question are extracted through a semantic coding model. Acoustic feature vectors, visual feature vectors, and textual semantic feature vectors are concatenated and fused, and then input into a multi-task classification model with a cross-modal attention mechanism for joint inference. The output is a classification label corresponding to the user's emotional state and a classification label corresponding to the user's question intent.
[0034] In this embodiment, when a user begins submitting a question, audio signals are captured in real time using audio acquisition devices such as microphones. These audio signals fully record information such as the user's tone of voice. Simultaneously, video signals are captured using video acquisition devices such as cameras, which can capture subtle changes in the user's facial expressions. While acquiring audio and video signals, the system uses speech recognition technology to convert the user's speech into corresponding question text.
[0035] After acquiring audio, video, and question text signals, the system performs preprocessing operations. For audio signals, common audio noise reduction algorithms are used to remove potential environmental noise, equipment interference, and other interference factors, making the speech clearer and more intelligible. For video signals, video alignment technology is employed to ensure that each frame in the video accurately matches the corresponding time point in the audio signal, avoiding audio-visual desynchronization. These preprocessing operations lay the foundation for accurate feature vector extraction in the subsequent stages.
[0036] After preprocessing, the system uses different models to extract various feature vectors. For the preprocessed audio signal, the system calls the acoustic model. Trained on a large amount of speech data, the acoustic model can accurately analyze various acoustic features in the audio signal, such as pitch, volume, speech rate, and pitch variations, and encode these features into acoustic feature vectors. These acoustic feature vectors can comprehensively and accurately reflect the characteristics of the user's speech intonation.
[0037] For the pre-processed video signal, the system uses a visual model. Based on deep learning technology, the visual model analyzes facial expressions in each frame of the video, recognizing subtle facial muscle movements, changes in eye gaze, and other subtle facial features, encoding these features into visual feature vectors. These visual feature vectors effectively express the emotional information conveyed by the user through facial expressions.
[0038] Simultaneously, the system utilizes a semantic encoding model to process the question text. Trained on massive amounts of text data, this model deeply understands the semantic connotations of the question text and extracts a textual semantic feature vector that accurately encodes the meaning of the question. This textual semantic feature vector clearly reflects the core content and intent of the user's question.
[0039] The system concatenates and fuses the extracted acoustic feature vectors, visual feature vectors, and textual semantic feature vectors. The concatenation and fusion method involves linking these three different types of feature vectors together in a specific order to form a comprehensive feature vector. This comprehensive feature vector integrates information from the user's voice tone, facial expressions, and the question text, providing a more comprehensive and accurate reflection of the user's current state and question intent.
[0040] After the comprehensive feature vector is input into a multi-task classification model embedded with a cross-modal attention mechanism, the model begins joint inference. The cross-modal attention mechanism automatically learns the importance and correlation between features from different modalities, allowing the model to focus on features more critical to identifying the user's emotional state and question intent when processing the comprehensive feature vector. The multi-task classification model simultaneously undertakes two classification tasks: classifying the user's emotional state (e.g., whether the user is happy, angry, calm, or anxious); and classifying the user's question intent (e.g., whether the user is inquiring about product information, seeking technical support, or submitting a complaint). Through joint inference, the model ultimately outputs classification labels corresponding to the user's emotional state and question intent, providing accurate information for subsequent customer service response generation.
[0041] Furthermore, the process of concatenating and fusing acoustic feature vectors, visual feature vectors, and textual semantic feature vectors, and inputting them into a multi-task classification model embedded with a cross-modal attention mechanism for joint inference, outputting classification labels corresponding to the user's emotional state and classification labels corresponding to the user's question intent, specifically includes: The acoustic feature vector, visual feature vector, and text semantic feature vector are projected into the same dimensional space to obtain uniformly encoded text features, acoustic features, and visual features. Using uniformly encoded text features as query vectors, and simultaneously using uniformly encoded acoustic and visual features as key and value vectors respectively, the attention weight distribution of the query vector on the key and value vectors is calculated. By utilizing the attention weight distribution, the acoustic value vector and the visual value vector are weighted and summed respectively to generate the acoustic context features and visual context features of text perception. The unified encoded text features, the text-aware acoustic context features, and the text-aware visual context features are concatenated to form a joint feature vector. The joint feature vector is input into the multi-task classification model, and the features are abstracted through a shared neural network layer. Through two parallel task-specific classification heads, the classification labels of the user's emotional state and the user's question intent are output synchronously.
[0042] In this embodiment, after obtaining the acoustic feature vector, visual feature vector, and text semantic feature vector, these vectors may come from different feature extraction models, and their dimensions often differ. To enable effective fusion and computation later, the system performs a dimensionality projection operation on each of these three feature vectors. Specifically, a specific dimensionality transformation algorithm is used to map the acoustic feature vector, visual feature vector, and text semantic feature vector to the same pre-defined dimensional space. After this processing, the three feature vectors, which originally had inconsistent dimensions, are transformed into uniformly encoded text features, acoustic features, and visual features, maintaining dimensionality consistency and laying the foundation for subsequent cross-modal attention computation.
[0043] After dimensional projection, the system uses uniformly encoded text features as the query vector. This query vector is used to find other related feature information. Simultaneously, uniformly encoded acoustic and visual features are used as the key and value vectors, respectively. Next, the attention weight distribution of the query vector on the key and value vectors is calculated separately. Based on the features of the query vector, the importance of each part of the key vector is evaluated, thus determining the weight that should be assigned to each value vector part in the subsequent weighted summation. Through this calculation, the correlation and importance differences between text features and acoustic and visual features can be captured.
[0044] Based on the calculated attention weight distribution, a weighted summation operation is performed on the acoustic and visual value vectors respectively. For the acoustic value vector, the values of each part are weighted and summed according to their corresponding weights in the attention weight distribution to generate text-aware acoustic context features. This feature not only includes the acoustic information itself but also incorporates correlation information related to text features, better reflecting the characteristics of acoustic features in the text context. Similarly, the same operation is performed on the visual value vector to generate text-aware visual context features. These two context features are obtained by considering the correlation between text features and features from other modalities, and can more comprehensively express the integrated information of the user's question under different modalities.
[0045] After obtaining the uniformly encoded text features, text-aware acoustic context features, and text-aware visual context features, the system concatenates these three feature vectors. The concatenation method involves linking them together in a specific order to form a new, more comprehensive joint feature vector. This joint feature vector integrates information from the text, acoustic, and visual modalities and considers the correlations between them, enabling it to more accurately represent the comprehensive characteristics of the user's question and providing rich information support for subsequent classification tasks.
[0046] The joint feature vector is input into a multi-task classification model. This model has shared neural network layers that further abstract and extract features from the joint feature vector, uncovering deeper levels of information. After processing by the shared layers, the model performs classification operations using two parallel task-specific classification heads. One head outputs a label indicating the user's emotional state, such as whether they are happy, angry, anxious, or calm. The other head outputs a label indicating the user's intent, such as whether they are inquiring about product information, seeking technical support, or filing a complaint. This multi-task classification approach allows the system to simultaneously output labels for both the user's emotional state and intent, providing accurate information for generating more appropriate customer service responses.
[0047] In some embodiments, step S102 above, which involves retrieving relevant B-end knowledge nodes from the unified knowledge base based on the user's question intent, specifically includes: Perform semantic deep analysis on the category tags of user question intent, extract core query terms and supplement contextual related terms, and generate structured query conditions that include a set of required keywords, a set of preferred keywords, a set of excluded keywords and type constraints; Based on structured query conditions, the system performs keyword-based precise retrieval using an inverted index engine, semantic vector-based similarity retrieval using a vector index, and relationship-based extended retrieval using knowledge graph relationships. The results returned by the inverted index engine, vector index, and knowledge graph relationship are normalized, and the fusion weights are dynamically assigned according to the type of the current user's question intent to calculate the comprehensive score of each knowledge node. The knowledge nodes are sorted and deduplicated based on the comprehensive score. The sorted and deduplicated knowledge nodes are then filtered for timeliness, permissions, and redundancy, and a set of B-end knowledge nodes that meets the preset quantity requirements is output.
[0048] In this embodiment, after obtaining the category tags of the user's question intent, a deep semantic analysis is performed on these category tags to thoroughly analyze the semantic information contained in the tags and accurately extract the core query terms. Core query terms are words that directly reflect the key content of the user's question. For example, if a user asks "the battery life of a certain mobile phone," then "mobile phone battery life" is the core query term. Simultaneously, the system also supplements contextual keywords, which further clarify the scope and background of the question. For example, in the above example, it might supplement with keywords such as "the model of this mobile phone" and "normal usage." Through this operation, the system generates structured query conditions containing a set of required keywords, a set of preferred keywords, a set of excluded keywords, and type constraints. The words in the required keyword set are essential for the query; their absence may lead to inaccurate results. The words in the preferred keyword set improve the matching accuracy of the query results. The words in the excluded keyword set are used to exclude results that do not meet the requirements. Type constraints clarify the type range of knowledge nodes; for example, only querying knowledge nodes related to product technical parameters.
[0049] Based on the generated structured query conditions, three retrieval methods are employed to find relevant knowledge nodes from a unified knowledge base. The first is a keyword-based precise retrieval using an inverted index engine. The inverted index engine is similar to a mapping table between words and knowledge nodes, quickly locating knowledge nodes containing specific keywords. This method accurately finds knowledge nodes that directly match the keywords. The second is a semantic vector-based similarity retrieval using vector indexing. The system converts both query conditions and knowledge nodes into semantic vectors, calculating the similarity between vectors to find knowledge nodes semantically similar to the query conditions. This method captures semantic similarity, finding relevant knowledge nodes even if the keywords are not exactly the same. The third is a relationship-based extended retrieval based on knowledge graph relationships. The knowledge graph stores various relationships between knowledge nodes, such as causal relationships and inclusion relationships. Through this retrieval method, the system can expand the search scope and find more potential related knowledge nodes based on known knowledge nodes and their relationships.
[0050] The inverted index engine, vector index, and knowledge graph relationships each return their respective search results. Since the scoring criteria differ across search methods, the system normalizes the returned results to unify the scores from different methods into a single scoring range for comprehensive comparison. Then, it dynamically assigns fusion weights based on the type of the user's question intent. Different types of user question intents have varying degrees of dependence on different search methods. For example, for technical questions, semantic vector-based similarity retrieval may be more important; while for rule-based questions, keyword-based precise retrieval may be more crucial. The system assigns different weights to the three search methods based on this difference, and finally calculates a comprehensive score for each knowledge node. This comprehensive score takes into account the results of different search methods, more accurately reflecting the relevance of the knowledge node to the user's question intent.
[0051] Based on the calculated comprehensive score, the system sorts the knowledge nodes, placing those with higher scores at the top and removing duplicates to avoid redundant information. Next, the sorted and deduplicated knowledge nodes undergo timeliness, access control, and redundancy filtering. Timeliness filtering ensures that the retrieved knowledge nodes are the latest and most valid information; for example, outdated product parameters or rules should not appear in the results. Access control filtering returns only knowledge nodes that the user has permission to access, based on the user's permission settings. Redundancy filtering removes nodes with high content duplication with existing knowledge nodes, improving the simplicity of the search results. Finally, the system outputs a set of B-end knowledge nodes that meets a preset requirement. These knowledge nodes are highly relevant to the user's question intent and have undergone filtering, providing accurate and effective knowledge support for subsequent customer service response generation.
[0052] In some embodiments, in step S102 above, the step of using a pre-trained graph neural network model to perform cross-perspective association reasoning on B-end knowledge nodes, mapping out the associated C-end rhetoric nodes, and forming a set of B / C node pairs with association weights specifically includes: Construct a heterogeneous graph knowledge base that includes B-end knowledge nodes, C-end verbal nodes, and same-perspective relationship edges and cross-perspective relationship edges; The set of B-end knowledge nodes to be processed is located in a heterogeneous graph knowledge base. A pre-trained graph neural network encoder performs multiple rounds of message passing based on an attention mechanism to generate a deep encoding vector that integrates multi-hop graph context information for each node. Based on the deep encoding vector of the B-end node, the matching score between it and the encoding vector of each C-end node in the heterogeneous graph knowledge base is calculated by the cross-view association prediction head. The matching score is then normalized to obtain the association weight probability. Based on the association weight probability, select the C-end nodes with the highest weight for each B-end node to form an initial set of weighted B-end node pairs; Perform global consistency checks and verbal redundancy processing on the initial BC node pair set, and output the target BC node pair set.
[0053] In this embodiment, to achieve cross-perspective reasoning between B-end knowledge nodes and C-end dialogue nodes, a heterogeneous graph knowledge base must first be constructed. This knowledge base is not a simple collection of nodes of a single type, but rather contains two different types of nodes: B-end knowledge nodes and C-end dialogue nodes. B-end knowledge nodes mainly store detailed knowledge points and background information for customer service personnel (B-end), such as product technical parameters and business process rules; C-end dialogue nodes contain expressions suitable for direct interaction with users (C-end), which are more accessible, concise, and easy to understand. Simultaneously, the knowledge base also includes two types of relationship edges: same-perspective relationship edges, which connect nodes within the same perspective, such as connecting different B-end knowledge nodes, reflecting their logical relationships, such as causal relationships and inclusion relationships; and cross-perspective relationship edges, which connect B-end knowledge nodes and C-end dialogue nodes, establishing a connection between them and providing a foundation for subsequent cross-perspective reasoning. By constructing such a heterogeneous graph knowledge base containing multiple types of nodes and relationship edges, a rich data framework is built for the entire reasoning process.
[0054] Once the set of B-end knowledge nodes to be processed is obtained, these nodes need to be located within the pre-constructed heterogeneous graph knowledge base. Then, a pre-trained graph neural network encoder processes these nodes. The graph neural network encoder employs a multi-round attention-based message passing mechanism. In each round, the node updates its state based on its own information and the information of other connected nodes. Through multiple rounds of this passing, each node can fully integrate multi-hop graph context information. For example, a B-end knowledge node will consider not only the information of its directly connected nodes but also the information indirectly related through other nodes. Finally, a deep encoding vector is generated for each B-end knowledge node. This vector contains rich semantic and structural information about the node in the heterogeneous graph knowledge base, enabling a more accurate representation of the node's features.
[0055] After obtaining the deep encoding vectors of the B-end nodes, a cross-view association prediction head is used to calculate their matching scores with the encoding vectors of each C-end node in the heterogeneous graph knowledge base. The cross-view association prediction head is a pre-trained module that determines the degree of association between the B-end and C-end nodes based on their encoding vectors and provides corresponding matching scores. These matching scores reflect the tightness of the association between B-end and C-end nodes; higher scores indicate a tighter association. However, the numerical ranges of different matching scores may differ. To facilitate comparison and analysis, these matching scores need to be normalized. The normalized values are the association weight probabilities, representing the likelihood of each C-end node being associated with its corresponding B-end node, with values ranging from 0 to 1.
[0056] Based on the calculated association weight probabilities, several C-end nodes with the highest weights are selected for each B-end node. The purpose of this step is to identify the C-end nodes most closely associated with each B-end node, in order to generate more accurate responses later. For example, for a B-end knowledge node about product technical parameters, several C-end verbal nodes that can explain the parameter in simple language might be selected. Combining each B-end node with its selected C-end nodes forms an initial set of weighted B / C node pairs. In this set, each node pair contains the B-end knowledge node, the C-end verbal node, and the association weight probabilities between them, providing rich information for subsequent response generation.
[0057] The initial set of B2C node pairs may have some issues, such as conflicts in global consistency or redundant dialogue. Therefore, it is necessary to perform global consistency checks and dialogue redundancy processing on the initial set. Global consistency checks primarily ensure that the node pairs in the entire set are logically and semantically consistent, avoiding contradictions. Dialogue redundancy processing removes C-end dialogue nodes that are similar or repetitive, making the set more concise and efficient. After these processing steps, the target set of B2C node pairs is output. This set of node pairs more accurately and reasonably reflects the relationship between B-end knowledge nodes and C-end dialogue nodes, providing strong support for generating high-quality customer service responses.
[0058] In some embodiments, step S103 above, which involves calculating the display weight ratio of B-end content to C-end content in the current scenario based on the user's emotional state, historical user profile, and problem complexity using a dynamic weight adjustment algorithm, specifically includes: The user's emotional state, historical user profile, and problem complexity are quantified and encoded to generate numerical emotional feature vectors, user profile feature vectors, and complexity scores. The emotion feature vector, user profile feature vector and complexity score are concatenated and fused. The non-linear relationship between multi-dimensional features is learned through the feature interaction neural network to generate a fused high-order scene feature vector. The high-order scene feature vector is input into the pre-trained weighted decision network. After passing through multiple nonlinear transformations of the weighted decision network, the basic weight values of the B-end content are output through the Sigmoid function. Input the basic weight value of B-end content and the current scenario characteristics into the business rule engine. By matching the predefined weight adjustment rules, the basic weight value of B-end content is fine-tuned and truncated to obtain the display weight ratio of B-end content. Based on the content display weight ratio of B-end, a smoothing factor is calculated according to the requirements of session continuity to smooth the preset basic weight value of C-end content, and a mandatory boundary adjustment is performed through business security rules to obtain the content display weight ratio of C-end.
[0059] In this embodiment, to accurately measure the impact of user emotional state, historical user profile, and question complexity on the weighting of response content, these three factors must first be quantified. For user emotional state, a specific emotion analysis model is used to transform the emotions reflected in the user's tone of voice and visual expressions when submitting a question into a numerical emotion feature vector. This vector precisely represents the type and intensity of the user's emotion in numerical form; for example, different emotions such as anger, happiness, and calmness each have corresponding numerical ranges. For historical user profile, which includes information such as past consumption behavior, consultation records, and preferences, this information is used to generate a user profile feature vector through feature extraction and encoding techniques. This vector comprehensively reflects the user's characteristics and needs. As for question complexity, a complexity score is given based on factors such as the question type, the knowledge domain involved, and the required steps to solve it; a higher score indicates a more complex question. Through this quantification and encoding process, the originally abstract factors are transformed into calculable and processable numerical representations, providing the foundational data for subsequent calculations.
[0060] After obtaining the emotion feature vector, user profile feature vector, and complexity score, they need to be concatenated and fused. The concatenation operation combines these three different dimensions of features to form a more comprehensive feature set. Then, a feature interaction neural network is used to further process these concatenated features. The feature interaction neural network can learn the non-linear relationships between multi-dimensional features, uncovering potential correlations and interactions between different features through its internal neuronal structure and connection methods. For example, a user's emotional state may influence their acceptance of complex questions, while preference information from historical user profiles may, along with question complexity, influence the choice of response content. After processing by the feature interaction neural network, a fused high-order scene feature vector is generated. This vector more accurately reflects the comprehensive characteristics of the current user consultation scenario, providing a more precise basis for subsequent weight ratio calculations.
[0061] The generated high-order scene feature vector is input into a pre-trained weighted decision network. This weighted decision network is a neural network model trained on a large amount of data. It has a multi-layered nonlinear transformation structure, capable of performing complex calculations and processing on the input high-order scene feature vector. Internally, each layer of neurons performs weighted summation and nonlinear activation operations on the input features, progressively extracting and transforming feature information. After multiple layers of such nonlinear transformations, the network finally outputs a value between 0 and 1 through the Sigmoid function. This value is the basic weight value of the B-side content. The Sigmoid function maps the network output to a specific range, giving it a clear probabilistic meaning, representing the basic proportion of B-side content in the response under the current scenario.
[0062] After obtaining the basic weight value of the B-end content, further adjustments are needed based on the characteristics of the current scenario. The basic weight value of the B-end content and the characteristics of the current scenario are input into the business rule engine. The business rule engine predefines a series of weight adjustment rules, which are formulated based on actual business needs and experience. For example, for certain types of users or specific problem scenarios, it may be necessary to make specific adjustments to the basic weight value of the B-end content. The business rule engine will fine-tune and truncate the basic weight value of the B-end content according to these predefined rules. Fine-tuning involves making small adjustments to the weight value based on subtle differences in the characteristics of the current scenario to make it more consistent with the actual situation; truncation ensures that the weight value is within a reasonable range, avoiding unreasonable situations where it is too high or too low. After this processing, the final B-end content display weight ratio is obtained, which more accurately reflects the proportion that B-end content should occupy in the response under the current scenario.
[0063] After obtaining the display weight ratio of B-end content, the display weight ratio of C-end content needs to be calculated. First, based on the B-end content display weight ratio, a smoothing factor is calculated according to the conversation continuity requirement. The conversation continuity requirement is to ensure the continuity and fluency of the response content, avoiding unnatural responses caused by abrupt changes in the B-end and C-end content ratios. The smoothing factor smooths the preset base weight value of C-end content, making the changes in C-end content weight values smoother and more reasonable. Then, considering the requirements of business security rules, the smoothed C-end content weight values are subject to mandatory boundary adjustments. Business security rules are to ensure that the response content complies with the company's business specifications and security requirements, avoiding inappropriate content ratios. Through this calculation and adjustment process, the final C-end content display weight ratio is obtained. This ratio, in conjunction with the B-end content display weight ratio, jointly determines the proportion of B-end and C-end content in the dual-perspective response, thereby generating responses that better meet user needs and scenario characteristics.
[0064] In some embodiments, step S103 above, which involves extracting content from the BC node pair set according to the display weight ratio and generating a dual-perspective response draft, specifically includes: The BC node pair set is sorted and redundancy is removed according to the displayed weight ratio to form an ordered list of candidate node pairs. Calculate the target information quota for B-end knowledge nodes and C-end knowledge nodes based on the display weight ratio; Based on the target information volume quota, iteratively extract the content of B-end knowledge nodes and C-end knowledge nodes from the candidate node pair list until their respective target information volume quotas are met. The B-end knowledge node content and C-end knowledge node content are assembled within their respective perspectives. During the assembly process, the logical alignment relationship between the two perspectives is established and recorded, and a draft response is output from both perspectives.
[0065] In this embodiment, the node pairs in the BC node pair set are sorted according to their display weight ratio. The display weight ratio reflects the relative importance of B-end and C-end content in the response in the current scenario. Sorting the node pairs according to this ratio ensures that more important node pairs are prioritized, facilitating their selection in subsequent processing. Simultaneously, the set undergoes redundancy removal to check for similar or duplicated node pairs. Since the unified knowledge base may have overlapping information, different B-end knowledge nodes may map to similar C-end dialogue nodes. Redundancy removal eliminates these duplicates, reducing the complexity of subsequent processing and ultimately forming an ordered list of candidate node pairs, providing an ordered and concise data foundation for accurate content extraction.
[0066] Based on the established display weight ratio, the target information volume quotas for each of the B-end and C-end knowledge nodes are further calculated. The display weight ratio determines the proportion of B-end and C-end content in the response. Combined with the overall information volume required for the response, the information volume that should be included in the B-end and C-end can be reasonably allocated. For example, if the display weight ratio indicates that B-end content should account for 60% and C-end content should account for 40%, and the overall response needs to convey a certain amount of key information, then the target information volume that the B-end and C-end knowledge nodes need to carry can be calculated according to this ratio. This ensures that the generated response's content allocation meets the needs of the current scenario, neither overly emphasizing one side nor omitting important information.
[0067] Based on the calculated target information volume quota, B-end and C-end knowledge node content is iteratively extracted from an ordered list of candidate node pairs. During the iteration, each node pair is selected from the head of the list, and the amount of B-end and C-end content extracted from that pair is determined by its association weight and the remaining target information volume quota. If the target information volume quota is still not met after extraction, the next node pair is selected for extraction, continuing until the B-end and C-end knowledge node content respectively meet their respective target information volume quotas. This iterative extraction method fully utilizes the information in the candidate node pair list, gradually constructing the response content according to importance and information volume requirements, ensuring the completeness and rationality of the response content.
[0068] After extracting the knowledge nodes from both the B2B and B2C perspectives, they are assembled separately. For B2B knowledge nodes, they are sorted and integrated according to the logical relationships and importance of the knowledge points to form a coherent and organized knowledge system, facilitating understanding and use by customer service personnel. For B2C knowledge nodes, they are organized in a way that is easy for users to understand and using, ensuring the content is concise, clear, and easy to comprehend. During the assembly process, the logical alignment between the two perspectives is established and recorded. For example, a detailed explanation of a technical principle in the B2B perspective should correspond to a simple description in the B2C perspective. Recording this correspondence allows customer service personnel to clearly understand the connection between the B2B and B2C content, enabling them to provide more accurate and comprehensive answers to users, ultimately resulting in a draft response from both perspectives.
[0069] In some embodiments, step S104 above, which involves inputting the dual-perspective draft response into the intelligent summarization model and generating a natural language summary and structured comparison view that highlights the core information differences between the B-end and C-end through difference extraction and comparative analysis, specifically includes: The draft texts of the B-end and C-end in the dual-perspective response draft are parsed into semantic unit sequences, and an alignment matrix between cross-perspective semantic units is established based on pre-defined mapping and semantic similarity calculation. Multi-level difference analysis is performed based on the alignment relation matrix. Information content difference analysis is performed to identify information condensation, semantic expression difference analysis is performed to classify terminology and perspective shifts, and structural difference analysis is performed to identify changes in the overall narrative logic, thus obtaining the difference analysis results. Based on the difference analysis results and type distribution, the view template is dynamically selected, the semantic unit sequence is filled into the view template, and the core difference points are highlighted using visual elements to generate a natural language summary and a structured comparison view to describe the differences in core information.
[0070] In this embodiment, after obtaining the draft responses from both perspectives, natural language processing techniques are used to parse the B-end and C-end draft texts into sequences of semantic units. A semantic unit is the smallest linguistic unit capable of expressing a certain meaning. Through text segmentation and semantic understanding, complex text can be transformed into a series of meaningful semantic units. For example, text describing the rules governing the returns of financial products can be parsed into semantic units involving return calculation methods, time periods, etc.
[0071] After parsing, an alignment matrix between semantic units from different perspectives is established based on pre-defined mappings and semantic similarity calculations. Pre-defined mappings are pre-defined correspondences between semantic units from different perspectives; for example, there might be a mapping between detailed technical terminology in the B2B context and its corresponding colloquial explanation in the B2C context. Semantic similarity calculations determine the degree of similarity between semantic units by analyzing their semantic features. Using pre-defined mappings and semantic similarity calculations, the correspondences between semantic units in the B2B and B2C contexts can be determined, thus constructing an alignment matrix. This matrix clearly shows the associations between semantic units from different perspectives, providing a foundation for subsequent discrepancy analysis.
[0072] Based on the established alignment matrix, multi-level difference analysis is conducted. First, an information content difference analysis is performed, identifying condensed information by comparing the information content of semantic units in the B-end and C-end. For example, the B-end may contain a large amount of detailed technical parameters and background knowledge, resulting in a larger information content, while the C-end refines and simplifies this information, resulting in a relatively smaller information content. Through analysis, it can be determined which information in the B-end is condensed and expressed in the C-end.
[0073] Next, a semantic expression difference analysis was performed to categorize the semantic units of the B2B and B2C sides, identifying instances of terminology and perspective shifts. B2B typically uses technical jargon and detailed logical explanations, while B2C tends to use plain language and concise expressions. By analyzing the differences in semantic units, the technical terms in B2B can be categorized with their corresponding colloquial expressions in B2C, while also identifying potential semantic changes during perspective shifts.
[0074] Finally, structural difference analysis is performed to analyze the overall narrative logic of the B2B and B2C texts, identifying structural changes between them. B2B texts may be organized according to the logical order of the knowledge system, while B2C texts may be reorganized based on user comprehension habits and the focus of the questions. Through structural difference analysis, we can understand the differences in their narrative logic, thus gaining a more comprehensive grasp of their differences. After these three levels of difference analysis, detailed difference analysis results are finally obtained.
[0075] Based on the difference analysis results and the distribution of difference types within the whole, an appropriate view template is dynamically selected. Different difference types and distributions may require different view formats to clearly display them. For example, if the difference in information content is large, a view template that highlights the condensed information may be selected; if the difference in semantic expression is significant, a template that facilitates comparison of terms and perspective shifts may be selected.
[0076] After selecting a view template, populate it with the previously parsed semantic unit sequence. During the population process, ensure that the position and arrangement of the semantic units conform to the design requirements of the view template and accurately reflect the differences between the B-end and C-end. Simultaneously, use visual elements to highlight key differences, such as using different colors, font sizes, or special symbols to mark crucial differences, enabling customer service personnel to quickly grasp important information.
[0077] Through the above operations, a natural language summary and a structured comparison view are finally generated to describe the differences in core information. The natural language summary concisely summarizes the core differences between the B2B and B2C sides, making it easy for customer service personnel to quickly understand; the structured comparison view presents the details of the differences in intuitive graphical or tabular form, facilitating in-depth analysis and understanding by customer service personnel, thereby better serving users.
[0078] In some embodiments, in steps S101 to S105 above, the method further includes: Capture customer service personnel's adjustment operation data on natural language summaries and structured comparison views, collect user feedback data on the customer service personnel's final response, and package the operation data, feedback data, and complete contextual metadata of this session to generate a feedback data package; Based on the feedback data packet, positive, negative and reinforcement sample pairs are constructed according to the customer service personnel's adoption, editing or rejection of specific scripts; Incremental learning is performed on the graph neural network model using positive, negative, and reinforced samples associated with nodes, and the predicted association weight probabilities of the graph neural network model are fine-tuned. Based on the feedback data package, the ideal weight value in the current scenario is derived by combining the adjustment tendency of customer service personnel and user feedback in order to construct parameter calibration samples; The dynamic weight adjustment algorithm is incrementally learned by using parameter calibration samples, and the weighted decision network parameters of the dynamic weight adjustment algorithm are periodically fine-tuned.
[0079] In this embodiment, during the process of customer service personnel using a bimodal customer service response generation method based on cross-perspective intelligent reasoning to handle user inquiries, the system captures in real-time the adjustment operation data of the customer service personnel on the natural language summary and structured comparison view. These adjustment operations may include the customer service personnel modifying certain information in the summary, adding, deleting, or rearranging content in the structured comparison view, etc. Simultaneously, the system also collects user feedback data on the customer service personnel's final response, such as user satisfaction ratings and whether the problem was resolved. Afterwards, the system integrates and packages the operation data, feedback data, and complete contextual metadata of this session. The complete contextual metadata covers various information from when the user asked the question, such as the user's basic information, historical consultation records, and the specific content of the question. This information together constitutes the feedback data package, providing a comprehensive data foundation for subsequent model optimization.
[0080] Based on the generated feedback data packets, the system conducts in-depth analysis of customer service personnel's behavior regarding specific scripts. When a customer service representative adopts a script node, the system constructs a positive sample pair between the corresponding B-end knowledge node and C-end script node. This indicates that the association between these two nodes is reasonable and effective, contributing to improved response quality. If a customer service representative rejects a script node, the system constructs a negative sample pair between its corresponding nodes, suggesting that the association may not be applicable in the current scenario. When a customer service representative edits a script node, the system constructs a reinforced sample pair. The edited content reflects the customer service representative's optimization of the node association based on practical experience and user feedback. This reinforced sample pair further strengthens the model's learning of reasonable associations. In this way, the system constructs positive, negative, and reinforced sample pairs of node associations, providing rich sample data for the incremental learning of the graph neural network model.
[0081] By utilizing pre-constructed positive, negative, and reinforcement sample pairs of node associations, the system performs incremental learning on a pre-trained graph neural network model. Incremental learning is a method of updating the model with new data without retraining the entire model. During incremental learning, the model fine-tunes the association weight probabilities between B-end knowledge nodes and C-end discourse nodes based on the information in the sample pairs. For example, for positive sample pairs, the model increases the association weight probability between the corresponding nodes, making them easier to map in subsequent reasoning; for negative sample pairs, the model decreases the association weight probability to reduce unreasonable associations; for reinforcement sample pairs, the model makes targeted adjustments based on the edited content to make the node associations more consistent with actual needs. Through this incremental learning approach, the graph neural network model can continuously optimize the predicted association weight probabilities, improving the accuracy of cross-perspective association reasoning.
[0082] The system then comprehensively analyzes the adjustment tendencies of customer service personnel and user feedback based on the feedback data packets. The adjustment tendencies of customer service personnel reflect their experience and preferences in handling user inquiries; for example, they may prefer to increase the display ratio of B-end content in certain scenarios to provide more detailed information. User feedback directly reflects their satisfaction with the responses and their actual needs. Combining these two aspects, the system derives the ideal weight value for the current scenario. This ideal weight value balances the display ratio of B-end and C-end content in the responses, making it more in line with user needs and expectations. Then, the system uses this ideal weight value to construct parameter calibration samples, providing a basis for optimizing the dynamic weight adjustment algorithm.
[0083] Using pre-constructed parameter calibration samples, the system incrementally learns the dynamic weight adjustment algorithm. This algorithm calculates the display weight ratio between B-end and C-end content based on factors such as user emotional state, historical user profiles, and question complexity. During incremental learning, the algorithm periodically fine-tunes its weighted decision network parameters based on the ideal weight values in the parameter calibration samples. The weighted decision network is the core of the dynamic weight adjustment algorithm; it calculates the display weight ratio by weighting various input factors. By periodically fine-tuning the weighted decision network parameters, the algorithm continuously adapts to the needs of different scenarios, improving the accuracy and rationality of the calculated display weight ratio, thereby generating a dual-perspective response draft that better meets user needs.
[0084] Reference Figure 2 An embodiment of the present invention provides a dual-modal customer service response generation system 2 with cross-perspective intelligent reasoning, wherein the system 2 specifically includes: The data acquisition module 201 is used to acquire and analyze the voice tone and visual expression features of users when submitting questions in real time, and combine the question text to identify the user's emotional state and the user's question intent. The intent analysis module 202 is used to retrieve relevant B-end knowledge nodes from the unified knowledge base based on the user's question intent, and to perform cross-perspective association reasoning on the B-end knowledge nodes using a pre-trained graph neural network model, mapping out the associated C-end verbal nodes to form a set of B-end node pairs with association weights. The draft generation module 203 is used to calculate the display weight ratio of B-end content and C-end content in the current scenario based on the user's emotional state, historical user profile and question complexity, and extract content from the BC node pair set according to the display weight ratio to generate a dual-perspective response draft. The response output module 204 is used to input the dual-perspective response draft into the intelligent summary model. Through difference extraction and comparative analysis, it generates a natural language summary and structured comparison view that highlights the core information differences between the B-end and C-end, and presents them to customer service personnel.
[0085] It is understandable that, such as Figure 1 The content of the cross-perspective intelligent reasoning bimodal customer service response generation method embodiment shown is applicable to the cross-perspective intelligent reasoning bimodal customer service response generation system embodiment. The specific functions implemented by the cross-perspective intelligent reasoning bimodal customer service response generation system embodiment are as follows: Figure 1 The cross-perspective intelligent reasoning bimodal customer service response generation method shown in the embodiment is the same, and the beneficial effects achieved are the same as those described above. Figure 1 The beneficial effects achieved by the cross-perspective intelligent reasoning bimodal customer service response generation method embodiment shown are also the same.
[0086] It should be noted that the information interaction and execution process between the above systems are based on the same concept as the method embodiments of the present invention. For details on their specific functions and technical effects, please refer to the method embodiments section, which will not be repeated here.
[0087] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the system can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this application. The specific working process of the units and modules in the above system can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0088] Reference Figure 3 The present invention also provides a computer device 3, including: a memory 302 and a processor 301, and a computer program 303 stored in the memory 302. When the computer program 303 is executed on the processor 301, it implements the bimodal customer service response generation method of cross-perspective intelligent reasoning as described in any of the above methods.
[0089] The computer device 3 may be a desktop computer, laptop, handheld computer, or cloud server, etc. The computer device 3 may include, but is not limited to, a processor 301 and a memory 302. Those skilled in the art will understand that... Figure 3 The computer device 3 is merely an example and does not constitute a limitation on the computer device 3. It may include more or fewer components than shown in the figure, or combine certain components, or different components, such as input / output devices, network access devices, etc.
[0090] The processor 301 may be a Central Processing Unit (CPU), or it may be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor.
[0091] In some embodiments, the memory 302 may be an internal storage unit of the computer device 3, such as a hard disk or memory of the computer device 3. In other embodiments, the memory 302 may be an external storage device of the computer device 3, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc., equipped on the computer device 3. Furthermore, the memory 302 may include both internal and external storage units of the computer device 3. The memory 302 is used to store the operating system, applications, boot loader, data, and other programs, such as the program code of the computer program. The memory 302 can also be used to temporarily store data that has been output or will be output.
[0092] This invention also provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements a bimodal customer service response generation method for cross-perspective intelligent reasoning as described in any of the above methods.
[0093] In this embodiment, if the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include at least: any entity or device capable of carrying computer program code to a photographing device / terminal device, a recording medium, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium. Examples include USB flash drives, portable hard drives, magnetic disks, or optical disks. In some jurisdictions, according to legislation and patent practice, computer-readable media cannot be electrical carrier signals or telecommunication signals.
[0094] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0095] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0096] In the embodiments disclosed in this application, it should be understood that the disclosed devices / terminal equipment and methods can be implemented in other ways. For example, the device / terminal equipment embodiments described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual coupling or direct coupling or communication connection may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.
[0097] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
Claims
1. A dual-modal customer service response generation method based on cross-perspective intelligent reasoning, characterized in that, The method specifically includes: Real-time acquisition and analysis of users' voice tone and visual expression features when submitting questions, combined with question text to identify users' emotional state and question intent; Based on the user's question intent, relevant B-end knowledge nodes are retrieved from the unified knowledge base. A pre-trained graph neural network model is used to perform cross-perspective association reasoning on the B-end knowledge nodes, mapping out the associated C-end verbal nodes, forming a set of B-end node pairs with association weights. Based on user emotional state, historical user profiles, and question complexity, a dynamic weight adjustment algorithm is used to calculate the display weight ratio of B-end content and C-end content in the current scenario. Based on the display weight ratio, content is extracted from the BC node pair set to generate a dual-perspective response draft. The draft response from both perspectives is input into the intelligent summary model. Through difference extraction and comparative analysis, a natural language summary and structured comparison view that highlight the core information differences between the B-end and C-end are generated and presented to customer service personnel.
2. The method according to claim 1, characterized in that, The real-time acquisition and analysis of the user's voice tone and visual expression features when submitting a question, combined with the question text to identify the user's emotional state and question intent, specifically includes: Collect audio and video signals when users submit questions and obtain the corresponding question text; perform noise reduction and alignment preprocessing on the audio and video signals. Based on the preprocessed audio and video signals and the question text, acoustic feature vectors encoding speech intonation are extracted through an acoustic model, visual feature vectors encoding facial expressions are extracted through a visual model, and textual semantic feature vectors encoding the meaning of the question are extracted through a semantic coding model. Acoustic feature vectors, visual feature vectors, and textual semantic feature vectors are concatenated and fused, and then input into a multi-task classification model with a cross-modal attention mechanism for joint inference. The output is a classification label corresponding to the user's emotional state and a classification label corresponding to the user's question intent.
3. The method according to claim 2, characterized in that, The process involves concatenating and fusing acoustic feature vectors, visual feature vectors, and textual semantic feature vectors, then inputting this data into a multi-task classification model embedded with a cross-modal attention mechanism for joint inference. The resulting output includes classification labels corresponding to the user's emotional state and the user's question intent. Specifically, this includes: The acoustic feature vector, visual feature vector, and text semantic feature vector are projected into the same dimensional space to obtain uniformly encoded text features, acoustic features, and visual features. Using uniformly encoded text features as query vectors, and simultaneously using uniformly encoded acoustic and visual features as key and value vectors respectively, the attention weight distribution of the query vector on the key and value vectors is calculated. By utilizing the attention weight distribution, the acoustic value vector and the visual value vector are weighted and summed respectively to generate the acoustic context features and visual context features of text perception. The unified encoded text features, the text-aware acoustic context features, and the text-aware visual context features are concatenated to form a joint feature vector. The joint feature vector is input into the multi-task classification model, and the features are abstracted through a shared neural network layer. Through two parallel task-specific classification heads, the classification labels of the user's emotional state and the user's question intent are output synchronously.
4. The method according to claim 2, characterized in that, The process of retrieving relevant B-end knowledge nodes from a unified knowledge base based on the user's question intent specifically includes: Perform semantic deep analysis on the category tags of user question intent, extract core query terms and supplement contextual related terms, and generate structured query conditions that include a set of required keywords, a set of preferred keywords, a set of excluded keywords and type constraints; Based on structured query conditions, the system performs keyword-based precise retrieval using an inverted index engine, semantic vector-based similarity retrieval using a vector index, and relationship-based extended retrieval using knowledge graph relationships. The results returned by the inverted index engine, vector index, and knowledge graph relationship are normalized, and the fusion weights are dynamically assigned according to the type of the current user's question intent to calculate the comprehensive score of each knowledge node. The knowledge nodes are sorted and deduplicated based on the comprehensive score. The sorted and deduplicated knowledge nodes are then filtered for timeliness, permissions, and redundancy, and a set of B-end knowledge nodes that meets the preset quantity requirements is output.
5. The method according to claim 1, characterized in that, The method employs a pre-trained graph neural network model to perform cross-perspective association reasoning on B-end knowledge nodes, mapping them to associated C-end verbal nodes, forming a set of B / C node pairs with association weights, specifically including: Construct a heterogeneous graph knowledge base that includes B-end knowledge nodes, C-end verbal nodes, and same-perspective relationship edges and cross-perspective relationship edges; The set of B-end knowledge nodes to be processed is located in a heterogeneous graph knowledge base. A pre-trained graph neural network encoder performs multiple rounds of message passing based on an attention mechanism to generate a deep encoding vector that integrates multi-hop graph context information for each node. Based on the deep encoding vector of the B-end node, the matching score between it and the encoding vector of each C-end node in the heterogeneous graph knowledge base is calculated by the cross-view association prediction head. The matching score is then normalized to obtain the association weight probability. Based on the association weight probability, select the C-end nodes with the highest weight for each B-end node to form an initial set of weighted B-end node pairs; Perform global consistency checks and verbal redundancy processing on the initial BC node pair set, and output the target BC node pair set.
6. The method according to claim 1, characterized in that, The process involves calculating the display weight ratio between B-end content and C-end content in the current scenario based on user emotional state, historical user profiles, and problem complexity using a dynamic weight adjustment algorithm. Specifically, this includes: The user's emotional state, historical user profile, and problem complexity are quantified and encoded to generate numerical emotional feature vectors, user profile feature vectors, and complexity scores. The emotion feature vector, user profile feature vector and complexity score are concatenated and fused. The non-linear relationship between multi-dimensional features is learned through the feature interaction neural network to generate a fused high-order scene feature vector. The high-order scene feature vector is input into the pre-trained weighted decision network. After passing through multiple nonlinear transformations of the weighted decision network, the basic weight values of the B-end content are output through the Sigmoid function. Input the basic weight value of B-end content and the current scenario characteristics into the business rule engine. By matching the predefined weight adjustment rules, the basic weight value of B-end content is fine-tuned and truncated to obtain the display weight ratio of B-end content. Based on the content display weight ratio of B-end, a smoothing factor is calculated according to the requirements of session continuity to smooth the preset basic weight value of C-end content, and a mandatory boundary adjustment is performed through business security rules to obtain the content display weight ratio of C-end.
7. The method according to claim 1, characterized in that, The step of extracting content from the BC node pair set according to the display weight ratio and generating a dual-perspective response draft specifically includes: The BC node pair set is sorted and redundancy is removed according to the displayed weight ratio to form an ordered list of candidate node pairs. Calculate the target information quota for B-end knowledge nodes and C-end knowledge nodes based on the display weight ratio; Based on the target information volume quota, iteratively extract the content of B-end knowledge nodes and C-end knowledge nodes from the candidate node pair list until their respective target information volume quotas are met. The B-end knowledge node content and C-end knowledge node content are assembled within their respective perspectives. During the assembly process, the logical alignment relationship between the two perspectives is established and recorded, and a draft response is output from both perspectives.
8. The method according to claim 1, characterized in that, The process of inputting the draft response from both perspectives into the intelligent summarization model, and generating a natural language summary and structured comparison view that highlights the core information differences between the B-end and C-end through difference extraction and comparative analysis, specifically includes: The draft texts of the B-end and C-end in the dual-perspective response draft are parsed into semantic unit sequences, and an alignment matrix between cross-perspective semantic units is established based on pre-defined mapping and semantic similarity calculation. Multi-level difference analysis is performed based on the alignment relation matrix. Information content difference analysis is performed to identify information condensation, semantic expression difference analysis is performed to classify terminology and perspective shifts, and structural difference analysis is performed to identify changes in the overall narrative logic, thus obtaining the difference analysis results. Based on the difference analysis results and type distribution, the view template is dynamically selected, the semantic unit sequence is filled into the view template, and the core difference points are highlighted using visual elements to generate a natural language summary and a structured comparison view to describe the differences in core information.
9. The method according to any one of claims 1 to 8, characterized in that, The method further includes: Capture customer service personnel's adjustment operation data on natural language summaries and structured comparison views, collect user feedback data on the customer service personnel's final response, and package the operation data, feedback data, and complete contextual metadata of this session to generate a feedback data package; Based on the feedback data packet, positive, negative and reinforcement sample pairs are constructed according to the customer service personnel's adoption, editing or rejection of specific scripts; Incremental learning is performed on the graph neural network model using positive, negative, and reinforced samples associated with nodes, and the predicted association weight probabilities of the graph neural network model are fine-tuned. Based on the feedback data package, the ideal weight value in the current scenario is derived by combining the adjustment tendency of customer service personnel and user feedback in order to construct parameter calibration samples; The dynamic weight adjustment algorithm is incrementally learned by using parameter calibration samples, and the weighted decision network parameters of the dynamic weight adjustment algorithm are periodically fine-tuned.
10. A dual-modal customer service response generation system with cross-perspective intelligent reasoning, characterized in that, The system specifically includes: The data acquisition module is used to acquire and analyze the voice tone and visual expression features of users when they submit questions in real time, and combine the question text to identify the user's emotional state and the user's question intent. The intent analysis module is used to retrieve relevant B-end knowledge nodes from a unified knowledge base based on the user's question intent. It uses a pre-trained graph neural network model to perform cross-perspective association reasoning on the B-end knowledge nodes, mapping out the associated C-end verbal nodes and forming a set of B-end node pairs with association weights. The draft generation module is used to calculate the display weight ratio of B-end content and C-end content in the current scenario based on the user's emotional state, historical user profile and question complexity, and extract content from the BC node pair set according to the display weight ratio to generate a dual-perspective response draft. The response output module is used to input the dual-perspective response draft into the intelligent summary model. Through difference extraction and comparative analysis, it generates a natural language summary and structured comparison view that highlights the core information differences between the B-end and C-end, and presents them to customer service personnel.