Power customer service method and device and related product
By acquiring multimodal interaction data and combining it with knowledge graphs in the field of power marketing, the problem of insufficient accuracy in single-modal data recognition in existing technologies has been solved, achieving accurate recognition of emotions and intentions, and improving the intelligence of power customer service and user experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- GUANGXI POWER GRID CORP
- Filing Date
- 2026-01-15
- Publication Date
- 2026-05-29
AI Technical Summary
Existing power customer service systems rely on single-modal data (such as text or voice) for analysis and recognition, making it difficult to fully capture the complex information in customer interactions. This results in insufficient accuracy in emotion and intent recognition and a tendency to misjudge customers' core needs.
Acquire multimodal interaction data (voice call records, text interaction records, and customer behavior data), generate voice emotion feature vectors, text intent feature vectors, and behavioral context feature vectors through a multimodal feature extraction model, combine them with a knowledge graph in the power marketing field to generate cross-modal association representation vectors, and use a joint classifier to accurately identify emotions and intents.
It achieves accurate joint recognition of emotions and intentions, reduces misjudgment of core demands, and improves the intelligence level, response efficiency, and user experience of power customer service.
Smart Images

Figure CN122114931A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and more specifically, to a method, apparatus, and related products for providing electricity customer service. Background Technology
[0002] With the digital transformation of electricity marketing services, the core task of the electricity customer service system is to accurately and efficiently handle various customer requests, covering diverse scenarios such as electricity bill consultation, fault reporting, business application, and complaint feedback. In this process, accurately grasping customer emotions and intentions has become the key to improving service quality, and the demand for intelligent and precise emotion and intention recognition in customer service scenarios is becoming increasingly prominent.
[0003] Existing power customer service systems primarily rely on single-modal data (such as text or voice) for analysis and recognition when processing customer requests. However, this approach has significant drawbacks. By focusing on only one data type, it fails to comprehensively capture the complex information contained within customer interactions, leading to insufficient accuracy in emotion and intent recognition and a tendency to misinterpret core customer needs. Therefore, it is urgent to address this technical problem. Summary of the Invention
[0004] In view of the above situation, this application provides a method, apparatus and related products for electricity customer service, which are intended to solve the above problems or at least partially solve the above problems.
[0005] In a first aspect, embodiments of this application provide a method for providing electricity customer service, the method comprising: Acquire multimodal interaction data of the target electricity customer, the multimodal interaction data including: voice call records, text interaction records and customer behavior data; Based on a pre-built multimodal feature extraction model, the multimodal interaction data is processed to generate speech emotion feature vectors, text intent feature vectors, and behavioral context feature vectors. Based on a pre-built knowledge graph in the field of electricity marketing, cross-modal association representation vectors are generated using the multimodal interaction data, the voice emotion feature vector, the text intent feature vector, and the behavioral context feature vector. A comprehensive feature vector is generated based on the speech emotion feature vector, text intent feature vector, behavioral context feature vector, and cross-modal association representation vector. The comprehensive feature vector is processed using a pre-built joint classifier to generate sentiment category prediction results and intent type prediction results, and the target electricity customer is responded to based on the sentiment category prediction results and intent type prediction results.
[0006] Secondly, embodiments of this application also provide an electricity customer service device, the device comprising: The acquisition module is used to acquire multimodal interaction data of the target electricity customer, including: voice call records, text interaction records, and customer behavior data; The first feature extraction module is used to process the multimodal interaction data based on a pre-built multimodal feature extraction model to generate voice emotion feature vectors, text intent feature vectors, and behavioral context feature vectors. The second feature extraction module is used to generate a cross-modal association representation vector based on a pre-built knowledge graph in the field of electricity marketing, using the multimodal interaction data, the voice emotion feature vector, the text intent feature vector, and the behavioral context feature vector. The generation module is used to generate a comprehensive feature vector based on the speech emotion feature vector, text intent feature vector, behavioral context feature vector, and cross-modal association representation vector; The prediction module is used to process the comprehensive feature vector using a pre-built joint classifier to generate sentiment category prediction results and intent type prediction results, and to respond to the target electricity customer based on the sentiment category prediction results and intent type prediction results.
[0007] By utilizing the above technical solutions, the power customer service method, device, and related products provided in this application first acquire multimodal interaction data composed of voice call records, text interaction records, and customer behavior data, breaking through the limitations of existing technologies that rely on single-modal data and comprehensively covering various key information in customer interactions; then, based on a multimodal feature extraction model, fine-grained voice emotion feature vectors, text intent feature vectors, and behavioral context feature vectors are generated to accurately capture customers' emotional expressions, core demands, and behavioral tendencies, making up for the coarseness of single feature extraction; finally, a cross-modal association representation vector is generated by combining a knowledge graph in the power marketing field. This approach deeply integrates multimodal features with industry-specific business logic and entity relationships, avoiding recognition biases in general models within power marketing scenarios. By fusing four types of feature vectors to generate a comprehensive feature vector, it retains the core information of each modality while strengthening the contribution of key information through dynamic weight allocation, ensuring the comprehensiveness and relevance of feature representation. By utilizing a joint classifier to output emotion category and intent type prediction results, it achieves accurate joint recognition of emotion and intent, effectively reducing misjudgments of core demands, assisting service personnel in quickly grasping customers' true needs and providing accurate responses, ultimately significantly improving the intelligence level, response efficiency, and user experience of power customer service.
[0008] The above description is only an overview of the technical solution of this application. In order to better understand the technical means of this application and to implement it in accordance with the contents of the specification, and to make the above and other objects, features and advantages of this application more obvious and understandable, the following are specific embodiments of this application. Attached Figure Description
[0009] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings: Figure 1 A flowchart illustrating the electricity customer service method provided in an embodiment of this application is shown; Figure 2 A schematic diagram of the structure of the power customer service device provided in an embodiment of this application is shown; Figure 3 A schematic diagram of the structure of an electronic device provided in an embodiment of this application is shown. Detailed Implementation
[0010] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below in conjunction with specific embodiments and corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0011] It should be noted that similar labels and letters in the following figures indicate similar items. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures.
[0012] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such use can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the term "comprising" and its variations should be interpreted as open-ended terms meaning "including but not limited to."
[0013] As mentioned earlier, existing power customer service systems primarily rely on single-modal data (such as text or voice) for analysis and recognition when processing customer requests. However, this method, which focuses on only one type of data, has significant drawbacks. Because it focuses on only one data type, it struggles to comprehensively capture the complex information contained in customer interactions, leading to insufficient accuracy in emotion and intent recognition and a tendency to misjudge core customer needs. Therefore, this invention proposes a power customer service method, device, and related products. The following detailed description of specific embodiments further illustrates this application.
[0014] To facilitate understanding of this embodiment, a detailed description of the electricity customer service method disclosed in this application embodiment will be provided first. The execution entity of the electricity customer service method provided in this application embodiment is generally a computer device with a certain computing capability. This computer device may include, for example, a terminal device, a server, or other processing devices. The terminal device may be user equipment (UE), a mobile device, a user terminal, a terminal, a personal digital assistant (PDA), a handheld device, a computing device, etc. In some possible implementations, this electricity customer service method can be implemented by a processor calling computer-readable instructions stored in memory.
[0015] Figure 1 This application provides a schematic flowchart of a method for providing electricity customer service, which is illustrated in an embodiment of this application. Figure 1 It can be seen that the embodiments of this application include at least steps S101-S105: S101: Obtain multimodal interaction data of the target electricity customer, wherein the multimodal interaction data includes: voice call records, text interaction records and customer behavior data; S102: Based on a pre-built multimodal feature extraction model, the multimodal interaction data is processed to generate speech emotion feature vectors, text intent feature vectors, and behavioral context feature vectors; S103: Based on the pre-built knowledge graph of the power marketing domain, a cross-modal association representation vector is generated using the multimodal interaction data, the voice emotion feature vector, the text intent feature vector, and the behavioral context feature vector; S104: Generate a comprehensive feature vector based on the speech emotion feature vector, text intent feature vector, behavioral context feature vector, and cross-modal association representation vector; S105: Using a pre-built joint classifier, the comprehensive feature vector is processed to generate sentiment category prediction results and intent type prediction results, so as to reply to the target electricity customer based on the sentiment category prediction results and intent type prediction results.
[0016] As can be seen, this application first acquires multimodal interaction data composed of voice call records, text interaction records, and customer behavior data, breaking the limitations of existing technologies that rely on single-modal data and comprehensively covering various key information in customer interactions. Then, based on the multimodal feature extraction model, fine-grained voice emotion feature vectors, text intent feature vectors, and behavioral context feature vectors are generated to accurately capture customers' emotional expressions, core demands, and behavioral tendencies, making up for the coarseness of single feature extraction. Next, a cross-modal association representation vector is generated by combining the knowledge graph of the power marketing field, deeply binding multimodal features with industry-specific business logic and entity associations, avoiding the recognition bias of general models in power marketing scenarios. By fusing the four types of feature vectors to generate a comprehensive feature vector, the core information of each modality is retained, and the contribution of key information is strengthened through dynamic weight allocation, ensuring the comprehensiveness and relevance of feature representation. By using a joint classifier to output the prediction results of emotion category and intent type, accurate joint recognition of emotion and intent is achieved, effectively reducing misjudgment of core demands, assisting service personnel to quickly grasp the real needs of customers and give accurate responses, and ultimately significantly improving the intelligence level, response efficiency, and user experience of power customer service.
[0017] The following provides a detailed explanation of steps S101-S105.
[0018] Regarding step S101 above: This step, through real-time multi-channel data collection, provides timely and comprehensive input data for subsequent feature extraction and early warning. In this step, voice call logs contain the conversation between the customer and service personnel. During implementation, voice call logs can be obtained from the call center system in WAV or MP3 format.
[0019] Text interaction logs include text messages sent by customers via SMS, email, or online chat windows (such as "There is an error in the electricity bill"), which can be stored as structured text files in JSON format or plain text. In practice, text interaction logs can be retrieved from the online customer service platform.
[0020] Customer behavior data includes, for example, the frequency with which customers check their electricity bills, their click patterns on service options, or records of submitted complaints, such as real-time electricity consumption checks or repair requests. During implementation, customer behavior data can be collected from smart terminals or customer relationship management systems (CRM systems).
[0021] During implementation, a streaming data processing framework can be used to ensure low-latency input and to use timestamps to align speech, text, and behavioral data, ensuring data synchronization.
[0022] Regarding step S102 above: In this step, the voice emotion feature vector represents the customer's emotional state; the text intent feature vector represents the customer's core needs and intentions; and the behavioral context feature vector represents the customer's behavioral habits and business relevance tendencies.
[0023] Multimodal feature extraction models include speech encoders, text encoders, and behavior encoders. Specifically, for example, the speech encoder employs a convolutional neural network-based architecture, such as processing the Mel spectrogram of the audio signal, using multiple convolutional and pooling layers to capture acoustic features, and ultimately outputting a speech emotion feature vector to represent emotional state; the text encoder relies on pre-trained Transformer models such as BERT, encoding the input text through word embedding layers and self-attention mechanisms, extracting semantic information, and generating text intent feature vectors to reflect core needs and intentions; the behavior encoder is designed as a sequence model based on a long short-term memory network, processing users' historical interaction behavior data such as click sequences or dwell time, using recurrent layers to capture time dependencies, thereby outputting a behavioral context feature vector to characterize behavioral habits and business relevance tendencies.
[0024] In some embodiments, the multimodal interaction data is processed based on a pre-built multimodal feature extraction model to generate a speech emotion feature vector, including: The voice call record is processed using a preset acoustic feature extraction algorithm to generate a pitch change feature vector, a speech rate fluctuation feature vector, and a volume feature vector. Using a pre-trained sentiment decomposition model, the tone change feature vector, speech rate fluctuation feature vector, and volume feature vector are processed to generate basic sentiment feature vectors and composite sentiment feature vectors. Using a pre-trained emotion change analysis model, the tone change feature vector, speech rate fluctuation feature vector, and volume feature vector are processed to generate an emotion change feature vector. Using a pre-trained emotion intensity quantification model, the tone change feature vector, speech rate fluctuation feature vector, and volume feature vector are processed to generate an emotion intensity feature vector; Based on the pitch change feature vector, speech rate fluctuation feature vector, volume feature vector, basic emotion feature vector, composite emotion feature vector, emotion change feature vector, and emotion intensity feature vector, the speech emotion feature vector is generated.
[0025] In this embodiment, specifically, the voice call record is first processed using a preset acoustic feature extraction algorithm: the MFCC algorithm combined with the YIN pitch detection algorithm is used to extract the frame-by-frame pitch frequency sequence from the frequency domain correlation dimension of the voice call record, calculate the frequency difference and fluctuation amplitude between adjacent frames, and generate a 64-dimensional pitch change feature vector after processing by a 1D-CNN (2 convolutional layers + 1 max pooling layer); the energy threshold method (based on the energy dimension data of the initial vector) is used to distinguish between speech frames and silence frames, the higher-order coefficients of MFCC are used to identify syllable boundaries, the number of syllables per unit time and the speech rate difference between adjacent speech segments are counted, and the same structure 1D-CNN is input to generate a 64-dimensional speech rate fluctuation feature vector; the energy correlation dimension data in the voice call record is extracted, the peak value, average value, variance of the energy of each frame and the energy change rate between adjacent frames are calculated, and a 64-dimensional volume feature vector is generated after processing by a 1D-CNN with a batch normalization layer.
[0026] Next, these three acoustic feature vectors are input into a pre-trained sentiment decomposition model (structure: input layer → LSTM layer (128 hidden units) → Dropout layer (probability 0.2) → dual classification head). After the model is fine-tuned with a data set labeled with power customer service scenarios (including voice and sentiment labels for complaints, inquiries, etc.), one classification head outputs a one-hot encoded vector (e.g., 5-dimensional) of the basic sentiment category (e.g., anger, calmness, etc.), which is the basic sentiment feature vector. The other classification head identifies two or more superimposed emotions (e.g., anger + anxiety), and outputs a one-hot encoded vector (e.g., 8-dimensional) after matching predefined composite sentiment labels, which is the composite sentiment feature vector.
[0027] Simultaneously, the three acoustic feature vectors are input into a pre-trained emotion change analysis model (using a temporal convolutional network TCN, containing three causal convolutional layers (with kernel sizes of 3 and numbers of 64, 128, and 64 respectively), residual connections, and layer normalization). The model captures the dynamic changes of the three acoustic feature vectors in the time dimension through one-dimensional convolution (such as the temporal trajectory of pitch from steady to sudden rise and speech rate from slow to fast), extracts long-term dependencies, and after fine-tuning with emotional change annotation data in the power scene, outputs a 64-dimensional emotion change feature vector, representing the evolution trend of emotion over time.
[0028] The three acoustic feature vectors are then input into a pre-trained emotion intensity quantization model (structure: input layer → 2 fully connected layers (with 128 and 64 hidden units respectively) → Sigmoid activation layer). The model combines key indicators from the three vectors (such as volume peak, speech rate variation, and pause frequency) and is trained on high-intensity emotion samples in the power industry (such as excited voices during complaints). The model outputs 0-1 standardized emotion intensity values, which are then converted into 1-dimensional emotion intensity feature vectors.
[0029] Finally, the seven feature vectors of pitch change, speech rate fluctuation, volume, basic emotion, complex emotion, emotion change, and emotion intensity are concatenated and then fused and dimension-unified through a fully connected layer (output dimension 256) to generate a speech emotion feature vector containing multi-dimensional emotional information.
[0030] This embodiment accurately extracts three core acoustic subvectors—pitch variation, speech rate fluctuation, and volume—through a pre-set acoustic feature extraction algorithm, comprehensively capturing the physical attributes directly related to emotional expression in voice call records, providing accurate underlying data support for emotion analysis. Utilizing a pre-trained emotion decomposition model, it effectively distinguishes between basic and complex emotions, breaking the limitations of single emotion recognition and achieving fine-grained capture of complex emotional states. The emotion change analysis model captures the dynamic evolution of emotions over time, allowing emotion features to move beyond static categories and better reflect the actual scenarios of natural emotional changes during calls. Finally, the emotion intensity quantification model is used to quantify emotional expression. Transforming emotions into intuitive numerical features makes the intensity of emotions quantifiable and comparable, providing clear numerical basis for subsequent risk assessment. Finally, by integrating seven types of feature vectors to generate a complete voice emotion feature vector, a multi-dimensional, full-chain feature coverage from underlying acoustic attributes to high-level emotional expression is achieved. This ensures both the comprehensiveness and fine granularity of the features, as well as their structure and operability. It provides strong support for the accurate identification of emotions and intentions in power customer service scenarios, effectively improving the accuracy of service risk assessment and real-time response efficiency. It also helps service personnel quickly grasp the true emotional state and core needs of customers, thereby improving customer service quality and user experience.
[0031] In some embodiments, the multimodal interaction data is processed based on a pre-built multimodal feature extraction model to generate a text intent feature vector, including: The text interaction record is semantically segmented at the sentence and word levels to obtain segmentation results; and natural language processing technology is used to process the segmentation results to generate text semantic embedding vectors. Based on a pre-built hierarchical intent classifier, the segmentation results are processed to generate multi-level intent categories and confidence scores, wherein the multi-level intent categories include intent graphs and sub-intents; The text intent feature vector is generated based on the text semantic embedding vector, multi-level intent categories, and confidence level.
[0032] In this embodiment, specifically, when performing sentence-level and word-level semantic segmentation on text interaction records, customer messages in JSON or plain text format (such as "The electricity bill is incorrect, please check") can be obtained from the online customer service system first. A Chinese word segmentation tool (such as Jieba) combined with a sentence boundary recognition algorithm (based on punctuation marks such as ., !, ;) is used to complete sentence-level segmentation, splitting the above message into two independent sentences: "The electricity bill is incorrect" and "Please check." Then, keywords such as "electricity bill," "bill," and "error" are extracted through word-level segmentation. At the same time, dependency parsing is used to identify the subject-verb-object structure and modification relationships to obtain the segmentation result.
[0033] When generating text semantic embedding vectors, the segmentation results can be input into a pre-trained BERT model finely tuned with a labeled dataset in the power marketing field. Through the model's word embedding layer, multi-head attention mechanism, and feedforward neural network layer, the text is transformed into a high-dimensional text semantic embedding vector with fixed dimensions (such as 768 dimensions), preserving the semantic association between keywords and context.
[0034] The intent-level classifier can adopt a structure of "pre-trained BERT encoding layer → dropout layer (probability 0.2) → fully connected concept graph classification layer (64 hidden units, softmax activation function, outputting the concept graph probability distribution) → sub-intention branch selector based on the concept graph → fully connected sub-intention classification layer (32 hidden units, softmax activation function, outputting the sub-intention probability distribution)". For example, the segmentation result is first semantically encoded by the BERT encoding layer, and after the dropout layer prevents overfitting, the concept graph classification layer outputs the probability distribution of "consultation", "complaint", and "service request", and then selects the corresponding sub-intention branch selector based on the predicted main intent. Sub-intent branches output detailed probability distributions; examples of main graphs are "consultation", "complaint", and "service request", and examples of corresponding sub-intents are "consulting about electricity bills", "consulting about electricity consumption", "complaining about abnormal electricity bills", "complaining about metering failures", "requesting electricity connection", and "requesting fault repair"; when generating text intent feature vectors, the main graph and sub-intents are first converted into one-hot encoded vectors (e.g., the main graph "complaint" corresponds to [0,1,0], and the sub-intent "complaining about abnormal electricity bills" corresponds to [1,0,0,0]), and the confidence scores of the main graph and sub-intents are extracted (e.g., the confidence score of the main graph "complaint" is 0.85, and the confidence score of the sub-intent "complaining about abnormal electricity bills" is 0.90).
[0035] Finally, the text intent feature vector is concatenated in the following order: “text semantic embedding vector → main graph one-hot encoding vector → sub-intent one-hot encoding vector → main graph confidence score → sub-intent confidence score”. The feature is then fused and dimension-unified through a fully connected layer (output dimension 256), and after normalization, the text intent feature vector is obtained.
[0036] This embodiment can accurately capture the core semantics and industry terms of text in power marketing scenarios through refined semantic segmentation and pre-trained model adaptation. It uses a hierarchical classifier to achieve fine-grained recognition of main and sub-intents. The feature vector constructed by combining semantic embedding, intent category and confidence has both semantic richness and category recognition, which can provide reliable input for subsequent joint analysis of sentiment and intent. It helps service personnel to quickly locate the core needs of customers and improve service response efficiency and user experience.
[0037] Regarding step S103: Research has revealed that general multimodal models in the power marketing field often suffer from low recognition accuracy and unstable generation results due to a lack of industry knowledge, making it difficult to meet the demands for high-quality services. Therefore, this step utilizes a pre-constructed knowledge graph in the power marketing field, along with the multimodal interaction data, the speech emotion feature vector, the text intent feature vector, and the behavioral context feature vector, to generate a cross-modal association representation vector.
[0038] In some embodiments, the generation of cross-modal association representation vectors based on the pre-built knowledge graph of the electricity marketing domain, utilizing the multimodal interaction data, the voice emotion feature vector, the text intent feature vector, and the behavioral context feature vector, includes: Based on the multimodal interaction data, multiple related entities are extracted from the pre-built knowledge graph of the power marketing domain; Using the speech emotion feature vector, the text intent feature vector, and the behavior context feature vector, a pre-trained graph neural network is used to calculate the association strength between the multimodal interaction data and each of the associated entities; The correlation strengths of each modality are encoded to generate a cross-modal correlation representation vector.
[0039] In this embodiment, the pre-built knowledge graph for the power marketing domain includes nodes covering customer-related entities (including customer ID, customer type, electricity account, etc., with attributes such as address), business service entities (including electricity bills, fault reporting processes, electricity installation services, metering devices, etc., with attributes such as amount and cycle), and abstract concept entities (including emotion categories such as anger and calmness, intent categories such as complaints and inquiries, with attributes such as intensity values). The edges represent the semantic and business logic relationships between various types of nodes (including the "initiation / query" association between customers and services, the "pointing" association between intents and business entities, and the "association" relationship between emotions and intents). A knowledge system adapted to the power marketing scenario is constructed through node attributes and structured edge relationships.
[0040] When extracting related entities based on multimodal interaction data, key information can be extracted from voice call records (keywords such as "abnormal electricity bill" and "repair request" extracted through speech-to-text conversion), text interaction records (core expressions such as "electricity bill" and "metering failure" extracted through semantic segmentation), and customer behavior data (behavioral tags such as "frequent bill inquiries" and "submitting complaint applications"). Then, by combining keyword matching, semantic similarity calculation, and the node attributes and edge relationships of the knowledge graph, highly related entities can be selected (such as customer text "electricity bill is incorrect" + behavior "inquired about the bill 5 times in 3 days", corresponding to extract entities such as "electricity bill", "electricity bill dispute", and "bill verification process").
[0041] In this embodiment, the pre-trained graph neural network can be selected from graph convolutional networks (GCN), graph attention networks (GAT), graph SAGE, etc., and this embodiment does not limit it.
[0042] Regarding the step of calculating the association strength between the multimodal interaction data and each associated entity using the pre-trained graph neural network by utilizing the voice emotion feature vector, the text intent feature vector, and the behavior context feature vector, the three feature vectors can be concatenated into a multimodal comprehensive feature. This comprehensive feature, along with the embedding vectors of each associated entity in the knowledge graph (including semantic information of entity attributes and associated edges), is then input into a graph attention network (GAT). The network assigns weights of 0.4, 0.3, and 0.3 to the "anger emotion - complaint intent" association, the "customer - electricity bill" query relationship, and the "complaint intent - electricity bill" pointing relationship, respectively, through an attention mechanism. The semantic similarity and attribute matching degree between the comprehensive feature and the embedding vectors of each entity are calculated, ultimately yielding an association strength of 0.92 for "electricity bill", 0.88 for "electricity bill dispute", and 0.35 for "fault reporting process".
[0043] Finally, the association strength is encoded to generate a cross-modal association representation vector. Specifically, the association strength of each associated entity can be arranged into an initial one-dimensional vector according to the business priority order of the entities in the knowledge graph. Then, the initial vector is mapped to a high-dimensional space through a fully connected embedding layer, preserving the entity attributes and association details. Finally, principal component analysis (PCA) is used to perform dimensionality reduction, compressing the high-dimensional vector into a compact vector with a fixed dimension (such as 64 dimensions), which is the cross-modal association representation vector.
[0044] This embodiment accurately extracts associated entities related to multimodal interaction data from a pre-built knowledge graph in the power marketing domain, achieving deep binding between multimodal data and industry-specific business entities, rules, and historical cases. This effectively compensates for the lack of domain knowledge in general models. By leveraging a pre-trained graph neural network, it integrates voice emotion feature vectors, text intent feature vectors, and behavioral context feature vectors to calculate association strength, quantifying the association relationships between dispersed multimodal features and knowledge graph entities. This breaks down the barriers between different modal information and domain business, making the association logic between emotion, intent, behavior, and business entities computable and interpretable. By encoding the association strength to generate cross-modal association representation vectors, the domain-specific "emotion-intent-business" association relationship is transformed into a structured high-dimensional vector. This not only retains the core information of multimodal features but also incorporates the exclusive semantics and business logic of the power marketing scenario. This provides accurate input rich in domain knowledge for subsequent joint optimization models, significantly improving the industry adaptability and accuracy of emotion and intent recognition, avoiding the recognition bias of general models, and providing reliable association basis for emotion and intent interaction relationship mining and priority scoring calculation. This helps service personnel quickly locate the core needs and potential risks of customers, improving service response efficiency and user experience.
[0045] For steps S104-S105: Regarding step S104, the specific implementation steps for generating the comprehensive feature vector can be as follows: First, input the speech emotion feature vector, text intent feature vector, behavioral context feature vector, and cross-modal association representation vector into the dynamic fusion network in high-dimensional vector form. This network can be based on an attention mechanism architecture and can adaptively allocate weights according to the contribution of each feature to emotion and intent recognition. For example, when a customer complains about abnormal electricity bills, higher weights will be assigned to speech emotion features such as angry tone. At the same time, behavioral context features such as frequent bill inquiries will help enhance intent understanding. Then, a multilayer perceptron will be used to perform nonlinear mapping on the weighted four types of features to fully preserve the semantic information and interaction relationships of each modality feature. Furthermore, the interaction data of the electricity marketing scenario can be used to fine-tune and optimize the fusion effect of the network, ultimately generating a comprehensive feature vector that fully reflects the complex information of customer interactions.
[0046] For example, the specific structure of the dynamic fusion network is as follows: The input layer receives four high-dimensional feature vectors, including a 64-dimensional speech emotion feature vector (containing sub-vectors such as pitch, speech rate, and emotion intensity), a 256-dimensional text intent feature vector (containing semantic embedding, main / sub-intent encoding, and confidence score), a 64-dimensional behavioral context feature vector (containing customer query frequency, business operation trajectory, etc.), and a 64-dimensional cross-modal association representation vector (containing the association strength between multimodal data and knowledge graph entities). During input, layer normalization is used to standardize the values of each vector to the [-1,1] interval to eliminate dimensional differences. The attention weight calculation layer adopts a structure combining "shared fully connected + independent fully connected": First, a shared fully connected layer (input dimension = 64 + 256 + 64 + 64 = 456, output dimension = 128, activation function is ReLU) is used to perform preliminary semantic alignment on the four features; then, an independent fully connected layer (input dimension = 128, output dimension = 1) is configured for each feature to output the initial weight value of each feature; finally, the four initial weight values are normalized by the Softmax function to obtain dynamically allocated attention weights (e.g., in the scenario of abnormal electricity bills for customer complaints, the weight of voice emotion features is 0.4, the weight of text intent features is 0.3, the weight of behavioral context features is 0.15, and the weight of cross-modal association representation is 0.15), realizing the adaptive allocation of "high contribution features with high weights". The feature weighted fusion layer multiplies the normalized feature vector of each path with the corresponding attention weight to obtain a 4-path weighted feature vector. Then, it generates a 1-dimensional fusion feature vector (dimensionality = 456) by element-wise summation. This process preserves the core semantics and interaction relationships of each modality feature, preventing individual features from being marginalized. The nonlinear mapping layer adopts a 2-layer multilayer perceptron (MLP) structure: the first fully connected layer has an input dimension of 456, 256 hidden units, and uses GELU as the activation function (adapted to high-dimensional feature gradient propagation), with a Dropout layer of probability 0.2 added to prevent overfitting; the second fully connected layer has an input dimension of 256, 128 hidden units, and still uses GELU as the activation function, followed by a Batch Normalization layer to stabilize the output distribution. This module mines deep correlations between features through nonlinear transformation. The output layer is a 1-layer fully connected layer (input dimension = 128, output dimension = 128), which maps the output vector values to the [0,1] interval using the Sigmoid activation function, ultimately generating a 128-dimensional comprehensive feature representation.
[0047] After obtaining the comprehensive feature vector, a pre-built joint classifier is used for processing to generate prediction results. In some embodiments, the process of using the pre-built joint classifier to process the comprehensive feature vector and generate sentiment category prediction results and intent type prediction results includes: The integrated feature vector is input into the joint classifier to calculate the joint probability distribution of sentiment category and intent type; The sentiment category and intent type with the highest joint probability are selected through a soft voting mechanism and used as the prediction results for the sentiment category and intent type, respectively.
[0048] In this embodiment, the joint classifier can adopt a deep neural network-based architecture, consisting of an input layer, multiple fully connected hidden layers, and an output layer. The input layer receives a comprehensive feature vector of fixed dimensions. The hidden layer contains 2-3 fully connected network layers (each with 256 and 128 hidden units respectively, and ReLU activation function is used to enhance nonlinear mapping capability). The output layer uses the softmax activation function (this function can transform the network output into a probability value in the range of 0-1, and the sum of all output probabilities is 1, which is suitable for multi-classification problems). At the same time, the classifier uses the cross-entropy loss function as the optimization objective and is trained on a labeled dataset of electricity marketing customer service scenarios (including sentiment and intent labels for scenarios such as complaints, inquiries, and service requests). It is optimized for industry-specific emotions such as "excitement during complaints" and specific intents such as "electricity bill disputes" to ensure the adaptability of fine-grained classification.
[0049] After the comprehensive feature vector is input into the joint classifier, it is first standardized by the input layer, and then nonlinearly mapped through multiple fully connected hidden layers to transform the features into raw scores in the joint classification space of sentiment and intent. Finally, it is transformed into a joint probability distribution by the softmax activation function of the output layer. Here, the joint probability distribution refers to the probability set corresponding to all combinations of sentiment category and intent type. For example, if the sentiment categories include "anger," "calm," and "anxiety," and the intent types include "complaining about abnormal electricity bills," "inquiring about electricity consumption," and "requesting electricity connection," then there will be 9 combinations. Each combination corresponds to a probability value between 0 and 1, and the sum of the probabilities of all combinations is 1. This distribution intuitively reflects the credibility of each sentiment-intent combination.
[0050] For example, in the joint probability distribution corresponding to a customer's comprehensive feature vector, the probability of the combination "anger + complaint about abnormal electricity bills" is 0.78, "anxiety + complaint about abnormal electricity bills" is 0.12, "calm + consultation about electricity consumption" is 0.08, and the probabilities of the remaining combinations are all below 0.02. In this case, the soft voting mechanism will select the combination "anger + complaint about abnormal electricity bills" with the highest joint probability as the prediction result of sentiment category and intent type. Moreover, the mechanism will integrate the contributions of multimodal features through a weighted average method, such as assigning higher weights to voice sentiment features, to highlight the impact of customer emotions on the classification results.
[0051] In addition, when the highest joint probability in the joint probability distribution is lower than a preset threshold (which is usually set to 0.7 and is determined by optimizing historical interaction data in the power marketing scenario), for example, if the probability of the combination of "anxiety + consultation measurement failure" in a customer's joint probability distribution is 0.63, which is lower than the threshold of 0.7, then the result is marked as a suspected emotional intent result that needs to be manually verified. The marked result is stored in JSON format, which includes the suspected emotional category, intent type and corresponding probability value, and is pushed to the customer service personnel's workbench for manual verification, so as to avoid misjudging vague customer expressions as definite intent.
[0052] In some embodiments, after the step of processing the integrated feature vector using a pre-built joint classifier to generate sentiment category prediction results and intent type prediction results, the method further includes: Obtain the sentiment category prediction result and intent type prediction result of at least one other electricity customer. For each electricity customer, generate a corresponding importance score based on the corresponding voice sentiment feature vector, text intent feature vector, and behavioral context feature vector; and generate the association strength based on the corresponding cross-modal association feature vector. Based on the importance scores and correlation strengths of each item, calculate the priority score of each electricity customer's request. Using a pre-defined clustering algorithm, the priority scores of each request are processed to obtain the priority category of the target electricity customer's request; Based on the cross-modal association feature vectors of the target electricity customer, emotional intent association relationships are generated. The step of responding to the target electricity customer based on the emotion category prediction result and the intent type prediction result includes: responding to the target electricity customer according to the emotion category prediction result and the intent type prediction result, the emotion intent correlation, and the request priority category.
[0053] In this embodiment, specifically, when generating the corresponding importance score, firstly, for each electricity customer's voice emotion feature vector, extract core indicators related to emotion intensity (such as the volume peak and speech rate fluctuation amplitude corresponding to anger), and generate a voice emotion feature importance score through quantification scoring in the 0-1 range; based on the intent category in the text intent feature vector, allocate quantified scores according to the business urgency rules of the electricity marketing scenario (such as complaint > consultation > routine inquiry), and generate a text intent feature importance score; for the behavior context feature vector, count the frequency of customer-related behaviors (such as the number of times electricity bills have been checked in the past 3 days) or operation density, convert it into a quantified result in the 0-1 range, and generate a behavior context feature importance score; when generating the association strength, calculate the semantic similarity or association matching degree between the cross-modal association feature vector and the core business entities in the knowledge graph of the electricity marketing domain (such as electricity bill disputes, fault repair process, and electricity installation service), and obtain the association strength after quantification processing.
[0054] During implementation, the priority score can be calculated using the following formula:
[0055] Where S represents the priority score, , , These represent the importance scores for speech emotion features, text intent features, and behavioral context features, respectively. This represents the strength of the cross-modal correlation feature vector, where λ, μ, and ν are adjustable, preset weight coefficients.
[0056] Geometric mean in the formula ,Right now and The square root of the product, with an exponent of 0.5, balances the contributions of sentiment and intent features, emphasizing their interaction; for example, the strong correlation between anger and complaint intent boosts priority. Furthermore, the geometric mean ensures that when any feature score is low (e.g., weak sentiment intensity), the score is not dominated by a single high-scoring feature, thus reflecting the overall importance of customer interactions. The weighting coefficients λ, μ, and ν are optimized using validation data from electricity marketing scenarios, for example, assigning higher weights to complaint scenarios.
[0057] After generating priority scores for each electricity customer's requests, the density-based DBSCAN clustering algorithm can be used to process these scores. First, all priority scores are used as input data. Combined with historical interaction data from the electricity marketing scenario, a neighborhood radius (e.g., 0.15) and a minimum sample size (e.g., 5) are preset. The algorithm will automatically identify the density distribution clusters of the scores. Those clustered in the high score range (e.g., 0.8-1.0) are classified as high priority categories (corresponding to complaints + negative emotion requests), those in the medium score range (0.5-0.8) are classified as medium priority categories (e.g., requests for electricity bill calculation consultation), and those in the low score range (0-0.5) with low density are classified as low priority categories (e.g., requests for querying business processing channels). At the same time, the clustering threshold is optimized through scenario data training to ensure that the category division fits the actual business needs.
[0058] Based on the cross-modal association feature vectors of target electricity customers, sentiment and intent association relationships are generated. Specifically, this can be achieved by using an association analysis module to leverage the electricity marketing domain-specific relationships of sentiment and intent encoded in the vectors (such as voice sentiment, text intent, behavioral context, and association logic with electricity business entities). Rule-based reasoning and semantic matching techniques are employed, combined with business rules in the electricity marketing domain (such as complaint handling procedures and electricity bill consultation standards), to mine the interactive relationships between sentiment and intent. For example, the causal relationship between "anger" and "intent to complain about abnormal electricity bills" and the association logic between "anxiety" and "intent to inquire about electricity connection" are identified. Finally, structured sentiment and intent association relationship data (in JSON format) is generated, presented in formats such as "anger → complaint about abnormal electricity bills → electricity bill dispute" and "anxiety → inquiry about electricity connection → connection process query," clearly reflecting the association logic between sentiment, intent, and corresponding electricity business entities.
[0059] The predicted results of emotion categories and intent types, the correlation between these emotions and intents, and the priority categories of the requests constitute a structured identification result list. Specifically, the structured identification result list is generated in JSON format and includes: fine-grained emotion categories (e.g., anger, anxiety), intent types (e.g., complaining about electricity bills, requesting electricity connection), priority categories (e.g., high priority), and correlations between emotions and intents (e.g., "anger → complaining about electricity bills → electricity bill dispute"). The result list is stored in a database and can be queried by customer service personnel through an interactive workbench. The visualization module presents the results through charts (e.g., bar charts showing priority distribution) or dashboards; for example, high-priority complaint cases are highlighted in red to facilitate quick response by service personnel. The visualization interface can be optimized using interactive data from electricity marketing scenarios to ensure it meets business needs. Understandably, this step, through structured and visualized output, enhances the practicality and operability of the identification results in customer service.
[0060] During implementation, semantic consistency verification can be performed on the recognition results within each priority category to eliminate abnormal results. Specifically, the semantic consistency verification module analyzes the recognition results within each priority category to verify whether the sentiment category, intent type, and interaction relationship conform to the business logic of the electricity marketing scenario. For example, the results of the "anger" sentiment and the "complain about electricity bills" intent in the high-priority category are compared with business rules in the knowledge graph (such as the electricity bill dispute handling process) to confirm consistency. If abnormal results are found, such as the "satisfaction" sentiment being incorrectly associated with the "complaint" intent, they are eliminated through semantic similarity calculation. The verification module uses a pre-trained language model to perform semantic embedding comparison on the results and optimizes it on the labeled data of the electricity marketing scenario to improve adaptability to industry-specific logic. Understandably, this step ensures the reliability of the recognition results through semantic verification and reduces the impact of misjudgments on customer service.
[0061] This embodiment generates importance scores and association strengths by integrating relevant data from other electricity customers. It then calculates accurate priority scores for requests based on the characteristics of the target customer, clarifies priority categories through clustering algorithms, and mines sentiment intent relationships. Finally, it responds to the target customer by combining sentiment intent prediction results, association relationships, and priority categories. This not only improves the comprehensiveness and rationality of request priority determination but also makes responses more aligned with the logic of electricity marketing business and the real needs of customers. It helps service personnel respond efficiently in a tiered manner, further improving the accuracy, response efficiency, and user experience of electricity customer service.
[0062] In some embodiments, after the step of processing the integrated feature vector using a pre-built joint classifier to generate sentiment category prediction results and intent type prediction results, the method further includes: If the predicted emotion category is negative and / or the predicted intent type is high-risk intent, then the warning signal and the predicted emotion category and / or the predicted intent type will be pushed to the customer service terminal to assist service personnel in responding quickly.
[0063] In this embodiment, a real-time warning signal is triggered when negative emotions or high-risk intentions are detected. Here, high-risk intentions can be preset, and this embodiment does not limit this.
[0064] Specifically, the relationship between emotion intensity and intent confidence level and thresholds can be compared (e.g., emotion intensity > 0.7 or intent confidence level > 0.8), triggering a real-time alert signal when the conditions are met. For example, an angry tone in a customer's voice combined with a complaint intent about "abnormal electricity bills" might trigger an alert. The alert signal is generated in JSON format, containing the emotion category, intent type, and triggering reason, and stored in the log system for future reference. The system optimizes threshold settings using verification data from electricity marketing scenarios to improve the accuracy of alerts. Understandably, this step ensures the timely detection of potential service problems by rapidly detecting high-risk interactions.
[0065] Furthermore, the warning signals and related sentiment analysis results are pushed to customer service personnel's terminals to assist them in responding quickly. Specifically, the warning signals and sentiment analysis results are pushed to the customer service personnel's workstation terminals in JSON format via a message queue, displayed as pop-up notifications or dashboard highlights. For example, a warning for "anger + electricity bill complaint" is marked in red, with accompanying analysis results such as "Emotion: Anger (Intensity 0.85), Intent: Electricity bill complaint (Confidence 0.90)". The push system supports multi-terminal compatibility (such as PCs and mobile devices) and provides interactive functions, allowing service personnel to view detailed analysis or initiate manual intervention. The system optimizes push priority through interaction logs from electricity marketing scenarios to ensure that high-risk warnings arrive first.
[0066] This embodiment detects negative emotions in the emotion category prediction results or high-risk intentions in the intention type prediction results, triggers early warnings in a timely manner, and pushes the relevant prediction results to the customer service terminal. This allows service personnel to quickly capture potential risks and urgent requests in customer service, assisting them in prioritizing responses and handling them accurately, effectively reducing the probability of escalating customer dissatisfaction, and improving the emergency response efficiency and service quality of power customer service.
[0067] Those skilled in the art will understand that in the above-described method of the specific embodiments, the order in which the steps are written does not imply a strict execution order, but constitutes no limitation on the implementation process. The specific execution order of each step should be determined by its function and possible internal logic.
[0068] It should be noted that in practical applications, all the above possible implementation methods can be arbitrarily combined to form possible embodiments of this application, and will not be described in detail here. The information (including but not limited to device information, user information, etc.) and data (including but not limited to data used for analysis, storage, and display) involved in this application are all information and data authorized by the user or fully authorized by all parties. Furthermore, the code snippets in the above steps are only examples illustrating the specific implementation process of page navigation and are not intended to limit this embodiment. The software tools or components appearing in the embodiments of this application are merely illustrative and do not represent actual use.
[0069] Based on the same concept, this application also provides an electric customer service device, which corresponds one-to-one with the electric customer service method in the above embodiments. Figure 2 A schematic diagram of the structure of the power customer service device provided in this application embodiment is shown. See also: Figure 2 As shown, the power customer service device 200 provided in this application embodiment includes: The acquisition module 201 is used to acquire multimodal interaction data of the target electricity customer, the multimodal interaction data including: voice call records, text interaction records and customer behavior data; The first feature extraction module 202 is used to process the multimodal interaction data based on a pre-built multimodal feature extraction model to generate speech emotion feature vectors, text intent feature vectors, and behavioral context feature vectors. The second feature extraction module 203 is used to generate a cross-modal association representation vector based on a pre-built knowledge graph in the field of electricity marketing, using the multimodal interaction data, the voice emotion feature vector, the text intent feature vector, and the behavioral context feature vector; The generation module 204 is used to generate a comprehensive feature vector based on the speech emotion feature vector, the text intent feature vector, the behavioral context feature vector, and the cross-modal association representation vector; The prediction module 205 is used to process the comprehensive feature vector using a pre-built joint classifier to generate sentiment category prediction results and intent type prediction results, so as to reply to the target electricity customer based on the sentiment category prediction results and intent type prediction results.
[0070] In some embodiments, in the above-described apparatus, the first feature extraction module 202, when processing the multimodal interaction data based on a pre-built multimodal feature extraction model to generate a speech emotion feature vector, is used for: The voice call record is processed using a preset acoustic feature extraction algorithm to generate a pitch change feature vector, a speech rate fluctuation feature vector, and a volume feature vector. Using a pre-trained sentiment decomposition model, the tone change feature vector, speech rate fluctuation feature vector, and volume feature vector are processed to generate basic sentiment feature vectors and composite sentiment feature vectors. Using a pre-trained emotion change analysis model, the tone change feature vector, speech rate fluctuation feature vector, and volume feature vector are processed to generate an emotion change feature vector. Using a pre-trained emotion intensity quantification model, the tone change feature vector, speech rate fluctuation feature vector, and volume feature vector are processed to generate an emotion intensity feature vector; Based on the pitch change feature vector, speech rate fluctuation feature vector, volume feature vector, basic emotion feature vector, composite emotion feature vector, emotion change feature vector, and emotion intensity feature vector, the speech emotion feature vector is generated.
[0071] In some embodiments, in the above-described apparatus, the first feature extraction module 202, when processing the multimodal interaction data based on a pre-built multimodal feature extraction model to generate a text intent feature vector, is used to: Sentence-level and word-level semantic segmentation is performed on the text interaction record to obtain segmentation results; and natural language processing technology is used to process the segmentation results to generate text semantic embedding vectors. Based on a pre-built hierarchical intent classifier, the segmentation results are processed to generate multi-level intent categories and confidence scores, wherein the multi-level intent categories include intent graphs and sub-intents; The text intent feature vector is generated based on the text semantic embedding vector, multi-level intent categories, and confidence level.
[0072] In some embodiments, in the above-described apparatus, the second feature extraction module 203, when generating a cross-modal association representation vector based on a pre-built knowledge graph in the power marketing domain, using the multimodal interaction data, the voice emotion feature vector, the text intent feature vector, and the behavioral context feature vector, is used for: Based on the multimodal interaction data, multiple related entities are extracted from the pre-built knowledge graph of the power marketing domain; Using the speech emotion feature vector, the text intent feature vector, and the behavior context feature vector, a pre-trained graph neural network is used to calculate the association strength between the multimodal interaction data and each of the associated entities; The correlation strengths of each modality are encoded to generate a cross-modal correlation representation vector.
[0073] In some embodiments, in the above-described apparatus, the prediction module 205 is specifically used for: The integrated feature vector is input into the joint classifier to calculate the joint probability distribution of sentiment category and intent type; The sentiment category and intent type with the highest joint probability are selected through a soft voting mechanism and used as the prediction results for the sentiment category and intent type, respectively.
[0074] In some embodiments, the apparatus further includes a generation module, which, after the step of processing the integrated feature vector using a pre-built joint classifier to generate sentiment category prediction results and intent type prediction results, is used to: Obtain the sentiment category prediction result and intent type prediction result of at least one other electricity customer. For each electricity customer, generate a corresponding importance score based on the corresponding voice sentiment feature vector, text intent feature vector, and behavioral context feature vector; and generate the association strength based on the corresponding cross-modal association feature vector. Based on the importance scores and correlation strengths of each item, calculate the priority score of each electricity customer's request. Using a pre-defined clustering algorithm, the priority scores of each request are processed to obtain the priority category of the target electricity customer's request; Based on the cross-modal association feature vectors of the target electricity customer, emotional intent association relationships are generated. The step of responding to the target electricity customer based on the emotion category prediction result and the intent type prediction result includes: responding to the target electricity customer according to the emotion category prediction result and the intent type prediction result, the emotion intent correlation, and the request priority category.
[0075] In some embodiments, the apparatus further includes an early warning module, which, after the step of processing the integrated feature vector using a pre-built joint classifier to generate sentiment category prediction results and intent type prediction results, is used to: If the predicted emotion category is negative and / or the predicted intent type is high-risk intent, then the warning signal and the predicted emotion category and / or the predicted intent type will be pushed to the customer service terminal to assist service personnel in responding quickly.
[0076] This invention provides a power customer service device. First, it acquires multimodal interaction data comprised of voice call records, text interaction records, and customer behavior data, breaking the limitations of existing technologies that rely on single-modal data and comprehensively covering various key information in customer interactions. Then, based on a multimodal feature extraction model, it generates fine-grained voice emotion feature vectors, text intent feature vectors, and behavioral context feature vectors, accurately capturing customers' emotional expressions, core needs, and behavioral tendencies, thus overcoming the coarseness of single-feature extraction. Next, it combines a knowledge graph from the power marketing field to generate cross-modal association representation vectors, deeply binding multimodal features with industry-specific business logic and entity relationships, avoiding recognition biases of general models in power marketing scenarios. By fusing the four types of feature vectors to generate a comprehensive feature vector, it retains the core information of each modality while strengthening the contribution of key information through dynamic weight allocation, ensuring the comprehensiveness and relevance of feature representation. Finally, by utilizing a joint classifier to output emotion category and intent type prediction results, it achieves accurate joint recognition of emotion and intent, effectively reducing misjudgments of core needs, assisting service personnel in quickly grasping customers' true needs and providing accurate responses, ultimately significantly improving the intelligence level, response efficiency, and user experience of power customer service.
[0077] Specific limitations regarding the electricity customer service device can be found in the limitations of the electricity customer service method described above, and will not be repeated here. Each module in the aforementioned electricity customer service device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device in hardware form, or stored in the memory of a computer device in software form, so that the processor can call and execute the operations corresponding to each module.
[0078] Figure 3 A schematic diagram of the structure of an electronic device provided in an embodiment of this application is shown. Figure 3 As shown, at the hardware level, this electronic device includes a processor, and optionally also includes an internal bus, a network interface, and memory. The memory may include main memory, such as high-speed random-access memory (RAM), or it may include non-volatile memory, such as at least one disk drive. Of course, this electronic device may also include other hardware required for other business operations.
[0079] The processor, network interface, and memory can be interconnected via an internal bus, which can be an ISA (Industry Standard Architecture) bus, a PCI (Peripheral Component Interconnect) bus, or an EISA (Extended Industry Standard Architecture) bus, etc. The bus can be divided into address bus, data bus, control bus, etc. For ease of representation, Figure 3 The symbol is represented by a single double-headed arrow, but this does not mean that there is only one bus or one type of bus.
[0080] Memory is used to store programs. Specifically, programs may include program code, which includes computer operation instructions. Memory may include main memory and non-volatile memory, and provides instructions and data to the processor.
[0081] The processor reads the corresponding computer program from non-volatile memory into main memory and then runs it, forming a power customer service device at the logical level. The processor executes the program stored in memory and specifically performs the aforementioned methods.
[0082] The processor may be an integrated circuit chip with signal processing capabilities. In implementation, each step of the above method can be completed by integrated logic circuits in the processor's hardware or by instructions in software form. The processor can be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it can also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this application can be directly embodied in the execution of a hardware decoding processor, or executed by a combination of hardware and software modules in the decoding processor. The software modules can reside in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. The storage medium is located in the memory, and the processor reads the information in the memory and, in conjunction with its hardware, completes the steps of the above method.
[0083] The electronic device can execute the electricity customer service method provided in several embodiments of this application, and can be implemented as an electricity customer service device. Figure 2 The functions of the embodiments shown are not described in detail here.
[0084] This application also proposes a computer-readable storage medium that stores one or more programs, the programs including instructions that, when executed by an electronic device including multiple applications, enable the electronic device to perform the power customer service methods provided in various embodiments of this application.
[0085] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0086] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0087] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0088] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0089] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0090] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.
[0091] Computer-readable media include both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.
[0092] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0093] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0094] The above are merely embodiments of this application and are not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.
Claims
1. A method for providing electricity customer service, characterized in that, The method includes: Acquire multimodal interaction data of the target electricity customer, the multimodal interaction data including: voice call records, text interaction records and customer behavior data; Based on a pre-built multimodal feature extraction model, the multimodal interaction data is processed to generate speech emotion feature vectors, text intent feature vectors, and behavioral context feature vectors. Based on a pre-built knowledge graph in the field of electricity marketing, cross-modal association representation vectors are generated using the multimodal interaction data, the voice emotion feature vector, the text intent feature vector, and the behavioral context feature vector. A comprehensive feature vector is generated based on the speech emotion feature vector, text intent feature vector, behavioral context feature vector, and cross-modal association representation vector. The comprehensive feature vector is processed using a pre-built joint classifier to generate sentiment category prediction results and intent type prediction results, and the target electricity customer is responded to based on the sentiment category prediction results and intent type prediction results.
2. The method according to claim 1, characterized in that, Based on a pre-built multimodal feature extraction model, the multimodal interaction data is processed to generate a speech emotion feature vector, including: The voice call record is processed using a preset acoustic feature extraction algorithm to generate a pitch change feature vector, a speech rate fluctuation feature vector, and a volume feature vector. Using a pre-trained sentiment decomposition model, the tone change feature vector, speech rate fluctuation feature vector, and volume feature vector are processed to generate basic sentiment feature vectors and composite sentiment feature vectors. Using a pre-trained emotion change analysis model, the tone change feature vector, speech rate fluctuation feature vector, and volume feature vector are processed to generate an emotion change feature vector. Using a pre-trained emotion intensity quantification model, the tone change feature vector, speech rate fluctuation feature vector, and volume feature vector are processed to generate an emotion intensity feature vector; Based on the pitch change feature vector, speech rate fluctuation feature vector, volume feature vector, basic emotion feature vector, composite emotion feature vector, emotion change feature vector, and emotion intensity feature vector, the speech emotion feature vector is generated.
3. The method according to claim 1, characterized in that, Based on a pre-built multimodal feature extraction model, the multimodal interaction data is processed to generate a text intent feature vector, including: Sentence-level and word-level semantic segmentation is performed on the text interaction record to obtain segmentation results; and natural language processing technology is used to process the segmentation results to generate text semantic embedding vectors. Based on a pre-built hierarchical intent classifier, the segmentation results are processed to generate multi-level intent categories and confidence scores, wherein the multi-level intent categories include intent graphs and sub-intents; The text intent feature vector is generated based on the text semantic embedding vector, multi-level intent categories, and confidence level.
4. The method according to claim 1, characterized in that, The pre-built knowledge graph for the power marketing domain, utilizing the multimodal interaction data, the voice emotion feature vector, the text intent feature vector, and the behavioral context feature vector, generates a cross-modal association representation vector, including: Based on the multimodal interaction data, multiple related entities are extracted from the pre-built knowledge graph of the power marketing domain; Using the speech emotion feature vector, the text intent feature vector, and the behavior context feature vector, a pre-trained graph neural network is used to calculate the association strength between the multimodal interaction data and each of the associated entities; The correlation strengths of each modality are encoded to generate a cross-modal correlation representation vector.
5. The method according to claim 1, characterized in that, The process of using a pre-built joint classifier to process the comprehensive feature vector and generate sentiment category prediction results and intent type prediction results includes: The integrated feature vector is input into the joint classifier to calculate the joint probability distribution of sentiment category and intent type; The sentiment category and intent type with the highest joint probability are selected through a soft voting mechanism and used as the prediction results for the sentiment category and intent type, respectively.
6. The method according to any one of claims 1-5, characterized in that, After the step of processing the comprehensive feature vector using a pre-built joint classifier to generate sentiment category prediction results and intent type prediction results, the method further includes: Obtain the sentiment category prediction result and intent type prediction result of at least one other electricity customer. For each electricity customer, generate a corresponding importance score based on the corresponding voice sentiment feature vector, text intent feature vector, and behavioral context feature vector; and generate the association strength based on the corresponding cross-modal association feature vector. Based on the importance scores and correlation strengths of each item, calculate the priority score of each electricity customer's request. Using a pre-defined clustering algorithm, the priority scores of each request are processed to obtain the priority category of the target electricity customer's request; Based on the cross-modal association feature vectors of the target electricity customer, emotional intent association relationships are generated. The step of responding to the target electricity customer based on the emotion category prediction result and the intent type prediction result includes: responding to the target electricity customer according to the emotion category prediction result and the intent type prediction result, the emotion intent correlation, and the request priority category.
7. The method according to claim 1, characterized in that, After the step of processing the comprehensive feature vector using a pre-built joint classifier to generate sentiment category prediction results and intent type prediction results, the method further includes: If the predicted emotion category is negative and / or the predicted intent type is high-risk intent, then the warning signal and the predicted emotion category and / or the predicted intent type will be pushed to the customer service terminal to assist service personnel in responding quickly.
8. A power customer service device, characterized in that, The device includes: The acquisition module is used to acquire multimodal interaction data of the target electricity customer, including: voice call records, text interaction records, and customer behavior data; The first feature extraction module is used to process the multimodal interaction data based on a pre-built multimodal feature extraction model to generate voice emotion feature vectors, text intent feature vectors, and behavioral context feature vectors. The second feature extraction module is used to generate a cross-modal association representation vector based on a pre-built knowledge graph in the field of electricity marketing, using the multimodal interaction data, the voice emotion feature vector, the text intent feature vector, and the behavioral context feature vector. The generation module is used to generate a comprehensive feature vector based on the speech emotion feature vector, text intent feature vector, behavioral context feature vector, and cross-modal association representation vector; The prediction module is used to process the comprehensive feature vector using a pre-built joint classifier to generate sentiment category prediction results and intent type prediction results, and to respond to the target electricity customer based on the sentiment category prediction results and intent type prediction results.
9. An electronic device, comprising: processor; as well as A memory configured to store computer-executable instructions, characterized in that, when executed, the executable instructions cause the processor to perform the steps of the method as described in any one of claims 1-7.
10. A computer-readable storage medium storing one or more programs, characterized in that, When the one or more programs are executed by an electronic device including multiple applications, the electronic device causes the electronic device to perform the steps of the method as described in any one of claims 1-7.