Intelligent customer service interaction method and related device

By acquiring and enhancing users' facial, voice, and text features, identifying inconsistent emotional features, and generating targeted responses for intelligent customer service, the problem of a single modality being unable to meet user needs is solved, thereby improving user satisfaction.

CN120639899APending Publication Date: 2025-09-12中国农业银行股份有限公司宁波市分行
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510832942.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-20
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

Traditional intelligent customer service interaction methods rely on a single modality and are unable to fully capture user emotions, making it difficult to generate customer service responses that meet user needs, affecting user satisfaction.

Method used

Acquire the user's facial data, voice data, and text data during the interaction process, extract features separately, identify inconsistent emotional features between facial features, voice features, and text features, perform feature enhancement, and generate target response content for intelligent customer service.

Benefits of technology

By enhancing the features of multimodal data, we can generate responses that better meet the user's real needs, thereby improving the user's intelligent customer service interaction experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120639899A_ABST
    Figure CN120639899A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent customer service interaction method and a related device, and relates to the technical field of artificial intelligence, and the method comprises the steps: obtaining face data, voice data and text data of a user in the current round of interaction in the process of interaction with an intelligent customer service, extracting features of the face data, the voice data and the text data, and storing the features in a database; the method comprises the steps of obtaining facial features, voice features and text features, recognizing non-consistent emotion features among the facial features, the voice features and the text features, and performing feature enhancement on the facial features, the voice features and the text features according to the non-consistent emotion features to obtain facial enhancement features, voice enhancement features and text enhancement features; and according to the face enhancement feature, the voice enhancement feature and the text enhancement feature, generating a target reply content of the intelligent customer service in the current round of interaction. According to the method and the device, the modal features are enhanced by adopting the non-consistent emotion features among the multi-modal features, so that the target reply content better meets the user requirements, and the user experience is better.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and in particular to an intelligent customer service interaction method and related devices. Background Art

[0002] Traditional intelligent customer service interaction methods rely on a single modality, such as voice or video, for analysis to generate customer service responses. However, a single modality often fails to fully capture user emotions, making it difficult to generate customer service responses that meet user needs, affecting user satisfaction. Summary of the Invention

[0003] In view of the above problems, this application provides an intelligent customer service interaction method and related devices to solve the problem that the existing technology uses a single modality to generate customer service responses that meet user needs. The specific solution is as follows:

[0004] The first aspect of the present application provides an intelligent customer service interaction method, comprising:

[0005] Acquire facial data, voice data, and text data of the user in the current round of interaction with the intelligent customer service. Text data refers to the text form of voice data.

[0006] Extract features from facial data, voice data, and text data respectively to obtain facial features, voice features, and text features;

[0007] Identify inconsistent emotional features among facial features, voice features, and text features, and enhance the facial features, voice features, and text features according to the inconsistent emotional features to obtain facial enhancement features, voice enhancement features, and text enhancement features;

[0008] Based on facial enhancement features, voice enhancement features, and text enhancement features, the target response content of the intelligent customer service in the current round of interaction is generated.

[0009] In one possible implementation, the process of determining facial features includes:

[0010] Perform face recognition on facial data to obtain the face area;

[0011] Extracting multiple key points in the face area and obtaining coordinate values ​​of the multiple key points based on the facial data, wherein the multiple key points are feature points reflecting the user's expression and emotion;

[0012] Calculate geometric features and extract texture features based on the coordinate values ​​of multiple key points and facial data to obtain the user's geometric features and texture features;

[0013] Generate facial features based on geometric features and texture features.

[0014] In one possible implementation, the process of determining the speech features includes:

[0015] Preprocessing the voice data to obtain preprocessed voice data;

[0016] Extracting prosodic and spectral features from preprocessed speech data;

[0017] Generate speech features based on prosodic and spectral features.

[0018] In one possible implementation, the process of determining text features includes:

[0019] Split the text data into multiple logically complete sentences as multiple candidate sentences;

[0020] The target topic sentence is selected from multiple candidate sentences using the word frequency inverse document frequency method and text sorting method;

[0021] Process the target topic sentence into global semantic features;

[0022] Perform sentiment analysis and importance analysis on the target topic sentence to obtain the sentiment intensity and weight of the target topic sentence, where the sentiment intensity reflects the intensity of the negative sentiment in the target topic sentence;

[0023] According to the sentiment intensity and weight of the target topic sentence and the global semantic features, the target semantic features that integrate the sentiment information are generated as text features.

[0024] In one possible implementation, the target response content of the intelligent customer service in the current round of interaction is generated based on the facial enhancement features, voice enhancement features, and text enhancement features, including:

[0025] Perform feature fusion on facial enhancement features, voice enhancement features, and text enhancement features to obtain the user's psychological portrait features in the current round of interaction;

[0026] Determine the user's multiple psychological state indicator values ​​in the current round of interaction based on the psychological portrait characteristics;

[0027] Generate target reply content based on multiple psychological state indicator values ​​of the user in the current round of interaction.

[0028] In one possible implementation, multiple psychological state indicator values ​​of the user in the current round of interaction are determined based on the psychological profile characteristics, including:

[0029] The psychological portrait features are input into the psychological index prediction model to obtain multiple psychological state index values ​​of the user in the current round of interaction, wherein the psychological index prediction model includes regression heads corresponding to the multiple psychological state index values.

[0030] In one possible implementation, target response content is generated based on multiple psychological state indicator values ​​of the user in the current round of interaction, including:

[0031] Generate the initial response content of the intelligent customer service in the current round of interaction based on the user's multiple psychological state indicator values ​​in the current round of interaction;

[0032] Obtain multiple psychological state indicator values ​​of users in multiple rounds of historical interactions;

[0033] Determining the user's psychological state fluctuation data based on multiple psychological state indicator values ​​of the user in multiple rounds of historical interactions and multiple psychological state indicator values ​​of the user in the current round of interactions;

[0034] The initial response content is adjusted according to the psychological state fluctuation data to obtain the target response content.

[0035] In a possible implementation, the multiple psychological state indicator values ​​include an emotion activation value, a cognitive load value, and a trust value.

[0036] A second aspect of the present application provides an intelligent customer service interaction device, comprising:

[0037] The data acquisition module is used to obtain the user's facial data, voice data, and text data in the current round of interaction with the intelligent customer service. The text data refers to the text form of the voice data.

[0038] A feature extraction module is used to extract features from facial data, voice data and text data respectively to obtain facial features, voice features and text features;

[0039] A feature enhancement module is used to identify inconsistent emotional features among facial features, voice features, and text features, and enhance the facial features, voice features, and text features according to the inconsistent emotional features to obtain facial enhancement features, voice enhancement features, and text enhancement features;

[0040] The customer service response generation module is used to generate the target response content of the intelligent customer service in the current round of interaction based on facial enhancement features, voice enhancement features and text enhancement features.

[0041] The third aspect of the present application provides a computer program product, including computer-readable instructions. When the computer-readable instructions are executed on an electronic device, the electronic device implements the intelligent customer service interaction method of the above-mentioned first aspect or any implementation method of the first aspect.

[0042] A fourth aspect of the present application provides an electronic device, comprising at least one processor and a memory connected to the processor, wherein:

[0043] Memory is used to store computer programs;

[0044] The processor is used to execute a computer program so that the electronic device can implement the intelligent customer service interaction method of the above-mentioned first aspect or any implementation method of the first aspect.

[0045] In a fifth aspect, the present application provides a computer storage medium, which carries one or more computer programs. When the one or more computer programs are executed by an electronic device, the electronic device can implement the intelligent customer service interaction method of the above-mentioned first aspect or any implementation method of the first aspect.

[0046] With the help of the above-mentioned technical solution, the intelligent customer service interaction method provided by this application takes into account that the facial data, voice data and text data of the user in the process of interacting with the intelligent customer service can reflect the user's current emotional and psychological state. Therefore, this application can obtain the facial data, voice data and text data of the user in the current round of interaction with the intelligent customer service, and extract features of the facial data, voice data and text data respectively to obtain facial features, voice features and text features.

[0047] Since the multimodal data of users in the process of interacting with intelligent customer service may show inconsistent emotional expressions, or even contradictory emotional expressions, for example, the user shows a frown of dissatisfaction, but the voice rhythm is steady and the emotion is weakened, and this inconsistent emotional expression can often better reflect the user's most real needs. Based on this, the present application can identify inconsistent emotional features between facial features, voice features and text features, and enhance the facial features, voice features and text features according to the inconsistent emotional features to obtain facial enhancement features, voice enhancement features and text enhancement features, and then generate the target reply content of the intelligent customer service in the current round of interaction based on the facial enhancement features, voice enhancement features and text enhancement features. Since the present application can identify inconsistent emotional features in multimodal data, the target reply content can be made more appropriate and more in line with user needs, thereby improving the user's intelligent customer service interaction experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] The above and other features, advantages, and aspects of the various embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and that the originals and elements are not necessarily drawn to scale.

[0049] Figure 1 A flowchart of an intelligent customer service interaction method provided in this application;

[0050] Figure 2 Schematic diagram of the feature fusion process based on feature fusion pyramid network;

[0051] Figure 3 A schematic diagram of the structure of an intelligent customer service interaction device provided in this application;

[0052] Figure 4 This is a schematic diagram of the structure of an electronic device provided in this application. DETAILED DESCRIPTION

[0053] The following describes the embodiments of the present application in conjunction with the accompanying drawings. The terms used in the implementation methods of the present application are only used to explain the specific embodiments of the present application and are not intended to limit the present application.

[0054] The embodiments of the present application are described below in conjunction with the accompanying drawings. Those skilled in the art will appreciate that, with the development of technology and the emergence of new scenarios, the technical solutions provided in the embodiments of the present application are also applicable to similar technical problems.

[0055] The terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequential order. It should be understood that the terms used in this way can be interchangeable under appropriate circumstances, and this is merely a way of distinguishing the objects of the same attributes when describing them in the embodiments of the present application. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, so that the process, method, system, product or equipment comprising a series of units need not be limited to those units, but may include other units that are not clearly listed or inherent to these processes, methods, products or equipment.

[0056] This application provides an intelligent customer service interaction method, which can be applied to an intelligent customer service system.

[0057] Optionally, the intelligent customer service interaction method can be applied to scenarios where users interact with intelligent customer service in an intelligent customer service system. In this scenario, the application can collect the user's personal information (including the following facial data, voice data, and text data) and analyze it to generate target reply content that meets the user's needs.

[0058] It is understandable that before using the technical solutions disclosed in the embodiments of this application, the type, scope of use, usage scenarios, etc. of the personal information involved in this application should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.

[0059] For example, when a user initiates a voice call through the intelligent customer service system, a prompt message is sent to the user to clearly inform the user that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the electronic device, application, server, storage medium, or other software or hardware that performs the operations of the technical solution of the present invention based on the prompt message.

[0060] As an optional but non-limiting implementation, when a user initiates voice data through the intelligent customer service system, a prompt message may be sent to the user in the form of a pop-up window, in which the prompt message may be presented in text form. Furthermore, the pop-up window may also contain a selection control for the user to select "agree" or "disagree" to provide personal information to the electronic device.

[0061] It is understandable that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this application. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this application.

[0062] It is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and relevant provisions.

[0063] In order to enable those skilled in the art to better understand the present application, the intelligent customer service interaction method of an embodiment of the present application is described in detail below with reference to the accompanying drawings.

[0064] Reference Figure 1 , Figure 1 This is a flow chart of an intelligent customer service interaction method provided in an embodiment of the present application, as shown in FIG. Figure 1 As shown, the intelligent customer service interaction method may include:

[0065] Step S101: Acquire facial data, voice data, and text data of the user in the current round of interaction with the intelligent customer service.

[0066] It is understood that the interaction between users and intelligent customer service is typically implemented on a terminal device (such as a mobile phone, computer, tablet, or learning machine). In one possible implementation, this application can use a camera installed on the terminal to collect the user's facial data, such as video and images, and use a sound receiving device to collect the user's voice data, i.e., the user's conversation data, when the user interacts with the intelligent customer service. Facial data refers to data containing the user's facial area, for example, image data of the user's upper body.

[0067] For example, a user talks to an intelligent customer service representative through a smartphone, uses the phone's camera to capture video data containing the user's face, and uses a microphone device to collect the user's voice data.

[0068] The above-mentioned text data refers to the text form of voice data, which can be obtained by performing voice recognition on the voice data.

[0069] It should be noted that the above process of obtaining facial data, voice data and text data is only an example and is not intended to limit this application.

[0070] Step S102: extract features from the facial data, voice data, and text data respectively to obtain facial features, voice features, and text features.

[0071] In this embodiment, facial data, voice data, and text data can be deeply characterized using high-dimensional features in their respective channels to obtain more refined feature data that better reflects the user's emotional state. Specifically, features can be extracted from facial data to obtain facial features that reflect the user's true expression, features can be extracted from voice data to obtain voice features that reflect the user's true emotion, and features can be extracted from text data to obtain text features that reflect the user's true expression and emotion.

[0072] Step S103: Identify inconsistent emotional features among facial features, voice features, and text features, and enhance the facial features, voice features, and text features according to the inconsistent emotional features to obtain facial enhancement features, voice enhancement features, and text enhancement features.

[0073] Considering that the multimodal data of users in the process of interacting with intelligent customer service may show inconsistent or even contradictory emotional expressions, for example, the user may show dissatisfaction with a frown, but the voice rhythm is steady and the emotion is weakened. This kind of emotional expression with steady voice but dissatisfied expression indicates that the user is in a state of emotional depression. If the user problem cannot be solved in a timely and effective manner, it may cause the user to have an emotional outburst and even cause user churn.

[0074] Based on this, this embodiment can identify inconsistent emotional features between facial features, voice features and text features, which represent the inconsistent or even potentially contradictory emotional expression information presented by the user between at least two data among facial data, voice data and text data.

[0075] Furthermore, the facial features, speech features and text features can be enhanced respectively according to the inconsistent emotion features to obtain facial enhancement features, speech enhancement features and text enhancement features. That is, the inconsistent emotion features are integrated into the facial features, speech features and text features respectively, so that the obtained facial enhancement features, speech enhancement features and text enhancement features not only contain their own original emotion expression information, but also contain the emotion expression information of other features, thereby improving the accuracy and comprehensiveness of the emotion expression represented by the facial enhancement features, speech enhancement features and text enhancement features respectively.

[0076] In one possible implementation, this step can be implemented through a cross-modal attention module, that is, facial features, speech features, and text features are input into the cross-modal attention module to obtain facial enhancement features, speech enhancement features, and text enhancement features output by the cross-modal attention module.

[0077] Here, the cross-modal attention module adopts a Transformer encoder structure. Each modal feature can serve as a query vector and participate in the key-value pair matching operation, thus establishing a bidirectional connection between all modalities. For example, when facial features are used as the query, emotional features inconsistent with facial features in speech and text features are incorporated into the facial features; when speech features are used as the query, emotional features inconsistent with speech features in facial and text features (such as features representing nervous expressions in facial features or features representing hesitation in text features) are incorporated into speech features, and so on. Because each modal feature is integrated with information from other modalities after fusion, a more context-aware representation is formed.

[0078] In the cross-modal attention module, the query-key / value matching process is implemented using an attention mechanism. This attention mechanism has dynamic modeling capabilities, meaning it can automatically adjust weights based on each round of input, eliminating the need to manually set the importance of the modalities. For example, if the speech modality is clear and emotionally intense (such as high pitch and high energy) at the current moment, the attention mechanism will naturally increase the weight of speech features in the final fused representation. However, if the speech rhythm is steady and the emotion is weakened, while subtle facial expressions such as frowning and evasive gaze appear, the attention mechanism will automatically shift attention to facial features. This "attention reallocation" capability enables the cross-modal attention module to capture asymmetric and potentially contradictory emotional expressions across multiple modalities and integrate them into the corresponding features.

[0079] Considering that facial features, voice features and text features are features extracted from their respective channels, there may be a problem of inconsistent dimensions. Therefore, optionally, this embodiment can first use a small multi-layer perceptron MLP to perform dimensionally unified mapping of facial features, voice features and text features before identifying inconsistent emotional features, that is, projecting facial features, voice features and text features into the same shared representation space to facilitate effective interaction between subsequent modalities.

[0080] Step S104: Generate target reply content for the intelligent customer service in the current round of interaction based on the facial enhancement features, voice enhancement features, and text enhancement features.

[0081] In this embodiment, facial enhancement features, voice enhancement features, and text enhancement features can be used to more accurately reflect the user's true emotional expression during the interaction with intelligent customer service, thereby generating target reply content that meets the user's real needs.

[0082] The intelligent customer service interaction method provided in this application takes into account that the facial data, voice data and text data of the user in the process of interacting with the intelligent customer service can reflect the user's current emotional and psychological state. Therefore, this application can obtain the facial data, voice data and text data of the user in the current round of interaction with the intelligent customer service, and extract features from the facial data, voice data and text data respectively to obtain facial features, voice features and text features.

[0083] Since the multimodal data of users in the process of interacting with intelligent customer service may show inconsistent emotional expressions, or even contradictory emotional expressions, for example, the user shows a frown of dissatisfaction, but the voice rhythm is steady and the emotion is weakened, and this inconsistent emotional expression can often better reflect the user's most real needs. Based on this, the present application can identify inconsistent emotional features between facial features, voice features and text features, and enhance the facial features, voice features and text features according to the inconsistent emotional features to obtain facial enhancement features, voice enhancement features and text enhancement features, and then generate the target reply content of the intelligent customer service in the current round of interaction based on the facial enhancement features, voice enhancement features and text enhancement features. Since the present application can identify inconsistent emotional features in multimodal data, the target reply content can be made more appropriate and more in line with user needs, thereby improving the user's intelligent customer service interaction experience.

[0084] In some embodiments of the present application, the process of the aforementioned step S102 of "extracting features from facial data, voice data, and text data respectively to obtain facial features, voice features, and text features" is introduced in detail.

[0085] In an optional embodiment, the process of determining facial features may include: performing face positioning and recognition on facial data to obtain a facial area; extracting multiple key points in the facial area, and obtaining coordinate values ​​of the multiple key points based on the facial data, wherein the multiple key points are feature points reflecting the user's expression and emotion; performing geometric feature calculation and texture feature extraction based on the coordinate values ​​of the multiple key points and the facial data to obtain the user's geometric features and texture features; and generating facial features based on the geometric features and texture features.

[0086] The following is a detailed explanation of the process of determining facial features.

[0087] Optionally, this embodiment may use a RetinaFace model to perform face location recognition on facial data to obtain a face area.

[0088] In the RetinaFace model, in the bottom-up stage, standard neural networks (such as ResNet50, a convolutional neural network model in deep learning that solves the vanishing gradient problem in deep neural network training through residual connections and is often used for image feature extraction) are used to extract features from the input facial data and generate feature maps at multiple scales. A Feature Pyramid Network (FPN) is then used to extract rich semantic information at different scales while maintaining high resolution. Specifically, in the top-down stage, low-level feature maps are upsampled and then fused with adjacent high-level feature maps.

[0089] This feature pyramid network can improve the detection and classification capabilities of faces of different sizes to a certain extent. However, considering that in this application, users may appear in different positions on the video screen, and the size and blur level of the face are not fixed, in order to improve the ability to extract fine-grained features and obtain more accurate face regions, this embodiment can improve the feature pyramid network of the RetinaFace model to obtain a Feature Fusion Pyramid Network (FFPN), and then obtain an improved RetinaFace model. The improved RetinaFace model can be used to perform face location recognition on facial data to obtain face regions.

[0090] See also Figure 2 As shown in FIG, a schematic diagram of the process of feature fusion based on the feature fusion pyramid network is shown. The feature maps of multiple scales obtained in the bottom-up stage are represented by layer1 feature map, layer2 feature map, layer3 feature map and layer4 feature map.

[0091] See also Figure 2 In the lower right corner of the FFPN part, the layer4 feature map can be directly output in the FFPN.

[0092] See also Figure 2 In the figure in the lower left corner of the FFPN part, this embodiment can use a trainable weight kernel (i.e., layer4 upsampling weight) to use a one-dimensional convolution kernel to upsample the layer4 feature map by a factor of 2. The obtained feature map is then fused with the layer3 feature map to obtain an updated layer3 feature map and output it.

[0093] See also Figure 2 In the figure in the upper right corner of the FFPN part, this embodiment can use a trainable weight kernel (i.e., layer4 upsampling weight) to use a one-dimensional convolution kernel to upsample the layer4 feature map by 4 times, and use another trainable weight kernel (i.e., layer3 upsampling weight) to use a one-dimensional convolution kernel to upsample the layer3 feature map by 2 times. The feature maps obtained by the two upsampling steps are then fused with the layer2 feature map to obtain the updated layer2 feature map and output it.

[0094] See also Figure 2 In the figure in the upper left corner of the FFPN part, this embodiment can use a trainable weight kernel (i.e., layer4 upsampling weight) to use a one-dimensional convolution kernel to upsample the layer4 feature map by 8 times, and use another trainable weight kernel (i.e., layer3 upsampling weight) to use a one-dimensional convolution kernel to upsample the layer3 feature map by 4 times, and use another trainable weight kernel (i.e., layer2 upsampling weight) to use a one-dimensional convolution kernel to upsample the layer2 feature map by 2 times. The feature maps obtained by the three upsampling steps are then fused with the layer1 feature map to obtain the updated layer1 feature map and output it.

[0095] In the FFPN module, through layer-by-layer upsampling and trainable weight kernels, each low-level feature map can be fused into the high-level feature map, thereby improving the ability of FFPN to extract fine-grained features. In this way, FFPN can obtain rich feature information at different levels and can effectively extract and utilize fine-grained feature information, thereby improving the performance of face positioning tasks.

[0096] After obtaining high-precision facial data in the above manner, this embodiment can extract multiple key points in the facial data to accurately identify the user's expression and emotional understanding through multiple key points.

[0097] Optionally, you can use the Dlib algorithm to extract key points from facial data and obtain multiple key points. Here, Dlib uses a facial key point detection algorithm based on a regression tree ensemble. By training a series of regression trees to predict the locations of facial key points, it can accurately detect 68 key points on the face, including the eyes, eyebrows, nose, mouth, and other parts.

[0098] Furthermore, this embodiment can extract geometric features based on the coordinate values ​​of multiple key points, that is, by calculating the distance, angle or area shape changes between key points to capture the facial expression information (such as the degree of openness of the left and right eyes, the degree of raising of the eyebrows, the degree of openness of the mouth, the degree of inclination of the corners of the mouth, etc.) and obtain geometric features.

[0099] For example, the vertical distance between the left eyebrow and the center of the left eye is calculated based on the coordinate values ​​of the corresponding key points, and / or the vertical distance between the right eyebrow and the center of the right eye is calculated to identify whether the user has an expression of eyebrows being too high. The above vertical distances can then form a feature vector as a geometric feature.

[0100] This embodiment can also determine various sub-regions of the face, such as the mouth region, the eye region, etc., based on the coordinate values ​​of multiple key points, and then perform texture feature extraction based on the facial data within the sub-regions. Taking the mouth region as an example, the mouth region can be cropped out, and then the LBP algorithm can be used to extract features of the mouth region to obtain the texture features of the mouth region.

[0101] Finally, facial features can be generated based on the geometric features and texture features obtained above. For example, the geometric features and texture features can be spliced ​​together to obtain facial features.

[0102] In the above embodiment, the user's expression and emotion can be recognized from two different perspectives through geometric features and texture features, with higher accuracy.

[0103] In another optional embodiment, the process of determining speech features may include: preprocessing speech data to obtain preprocessed speech data; extracting prosodic features and spectral features from the preprocessed speech data; and generating speech features based on the prosodic features and spectral features.

[0104] The following is a detailed explanation of the process of determining speech features.

[0105] Optionally, this embodiment can preprocess the speech data by removing silence segments, pre-emphasizing, and performing frame-by-frame windowing to improve the stability and accuracy of subsequent feature extraction. Furthermore, prosodic and spectral features can be extracted from each frame of preprocessed speech data. These extracted prosodic and spectral features from each frame of speech data can then be concatenated to obtain speech features.

[0106] Here, prosodic features refer to suprasegmental features in speech data that transcend basic speech units (such as phonemes and syllables). They reflect the dynamic changes in speech data over time and carry non-semantic information such as the user's (i.e., speaker's) emotion, intonation, rhythm, and stress. Unlike segmental features such as phonemes and syllables (such as pronunciation and initials and finals), prosodic features focus on the "melody" and "rhythm" of speech and are important non-content information in speech data.

[0107] Optionally, the prosodic features include fundamental frequency (F0), energy, speaking rate and other features. For example, when angry, the fundamental frequency and energy are usually higher, and when sad, the speaking rate is slower and the standard deviation of the fundamental frequency is reduced.

[0108] Spectral features are a series of characteristic parameters extracted through spectral analysis of speech data. They are used to describe the energy distribution, frequency components, and their variation patterns in the frequency domain. Spectral features reflect the frequency domain characteristics of speech data and are important non-semantic information in speech signal processing.

[0109] Optionally, spectral features include Mel-Frequency Cepstral Coefficients (MFCCs), spectral flux, harmonic-to-noise ratio, and other features. MFCCs are used to capture changes in vocal tract shape, spectral flux is used to measure the difference in the spectra of adjacent frames (for example, spectral flux increases when surprised), and harmonic-to-noise ratio is used to quantify the intensity of periodic components (for example, the harmonic-to-noise ratio decreases when tired).

[0110] It should be noted that the above-mentioned process of determining speech features is only an example and is not intended to limit the present application.

[0111] This embodiment can accurately and comprehensively extract non-content information from the voice data by performing feature extraction on the voice data, and can accurately determine the emotional state and psychological state of the user when speaking based on the non-content information.

[0112] In another optional embodiment, the text feature extraction process may include: splitting the text data into multiple logically complete sentences as multiple candidate sentences; using the word frequency inverse document frequency method and the text sorting method to screen out the target topic sentence from the multiple candidate sentences; processing the target topic sentence as a global semantic feature; performing sentiment analysis and importance analysis on the target topic sentence to obtain the sentiment intensity and weight of the target topic sentence, wherein the sentiment intensity reflects the intensity of the negative sentiment in the target topic sentence; generating a target semantic feature that integrates sentiment information as a text feature based on the sentiment intensity and weight of the target topic sentence and the global semantic feature.

[0113] Taking the text data "Hello, I transferred money to a friend using the mobile banking app yesterday. The operation clearly indicated that the transaction was successful, but my friend hasn't received the money. I checked the transfer record and the status is 'processing', and it's been over 24 hours. Is there a problem with your app? Will the handling fee be refunded in this case? I need this money urgently, can you please help me to expedite it?" as an example, this article introduces an optional determination process for text features.

[0114] First, this embodiment can split the text data into multiple logically complete sentences through preprocessing and sentence segmentation. For example, the above text data can be split into the following five sentences.

[0115] Sentence 1 (hereinafter referred to as S1): Hello, I used the mobile banking app to transfer money to a friend yesterday. It clearly indicated that the operation was successful, but my friend has not received the money.

[0116] Sentence 2 (hereinafter referred to as S2): I checked the transfer record and the status is also 'processing', which is more than 24 hours ago.

[0117] Sentence 3 (hereinafter referred to as S3): Is there something wrong with your APP?

[0118] Sentence 4 (hereinafter referred to as S4): Will the handling fee be refunded in this case?

[0119] Sentence 5 (hereinafter referred to as S5): I need this money urgently now, can you help me to urge it?

[0120] In this embodiment, the five sentences mentioned above may be used as five candidate sentences, and the term frequency-inverse document frequency (TF-IDF) method and the text ranking (TextRank) method are used to screen out the target topic sentence from the five candidate sentences.

[0121] Among them, in the TF-IDF channel (focusing on detecting global high-frequency business keywords and non-high-frequency business keywords), first, based on the customer service historical dialogue corpus (containing a large number of business terms such as "transfer", "failure", "query", and "handling fee", which are defined as business keywords in this embodiment), the word frequency inverse document frequency value, i.e., the TF-IDF value, of the business keywords contained in each candidate sentence is calculated. Then, based on the TF-IDF value calculated for each candidate sentence, a preset operation (such as averaging or weighted summing) is performed to obtain the TF-IDF score value of each candidate sentence.

[0122] For example, S1 contains high-frequency business keywords such as "transfer", "success", and "receive money", and has a high TF-IDF score.

[0123] S2: Contains high-frequency business keywords "transfer record", "processing", and "24 hours", and has a high TF-IDF score.

[0124] S3: Contains the business keywords "APP" and "problem", and has a low TF-IDF score.

[0125] S4: Contains the business keywords "handling fee" and "refund" and includes specific business points. The TF-IDF score is medium.

[0126] S5: Contains the business keywords "urgent" and "urgent", as well as user demands, with a medium TF-IDF score.

[0127] Based on this, the TF-IDF scores of the five candidate sentences are sorted, and the following ranking results are obtained: S2 > S1 > S4 ≈ S5 > S3, indicating that the text data emphasizes "abnormal transfer status" and "time limit timeout".

[0128] In the TextRank channel (which focuses on contextual semantic associations and detecting core pivot sentences), we first construct a sentence similarity graph with five sentences as five nodes and the similarity between every two sentences (for example, the similarity based on word overlap Jaccard or the cosine similarity based on word vectors) as the edge weight. For example, by calculating the similarity between sentences, we can obtain: S1 and S2 are highly similar, both describing the failure of transfers to arrive and abnormal status; S1 is related to S3, and S2 is related to S3, and both describe the cause of the incident due to APP problems; S1 is related to S4, S1 is related to S5, S2 is related to S4, and S2 is related to S5, and both describe the consequences and demands.

[0129] After constructing the sentence similarity graph, the TextRank algorithm can be used to iteratively calculate the TextRank score value of each candidate sentence.

[0130] For example, S1 and S2: describe the core issues, i.e., transfer failure, abnormal status, and timeout. They are the central nodes of the semantic network and have the highest TextRank score.

[0131] S5: Expresses urgent demands, i.e., urging for action, is strongly associated with the core issue, and has the second highest TextRank score.

[0132] S4: Inquiring about the refund of handling fees is a specific question derived from the core question, and the TextRank score is medium.

[0133] S3: Questioning the APP problem is a possible cause speculation, and the TextRank score value is relatively low (probably not as important as the sentence describing the facts).

[0134] Based on this, the five candidate sentences are sorted according to their TextRank scores, and the following ranking results are obtained: S2 > S1 > S5 > S4 > S3, indicating that the text data emphasizes "problem description" and "urgent appeal".

[0135] Based on the above sorting, this embodiment needs to find a set of candidate sentences containing a preset topic from the five candidate sentences as the target topic sentence.

[0136] Optionally, the preset topics can be: core user issues (e.g., transfer failure), direct requests (e.g., urging for processing), and key questions (e.g., fee refund).

[0137] In a possible implementation, the top n (n ≥ 1, for example, n is 3) candidate sentences ranked in the TF-IDF ranking results and the TextRank ranking results can be directly used as target topic sentences, namely S1, S2, S5 and S4.

[0138] In another possible implementation, overlapping candidate sentences among the top n candidate sentences in the TF-IDF ranking results and the TextRank ranking results may be first used as target topic sentences, namely, S1 and S2.

[0139] In another possible implementation, it is considered that the method of selecting overlapping candidate sentences may lose some candidate sentences containing important information, thereby affecting the accuracy of text features.

[0140] To this end, this embodiment can perform complementary completion processing on the basis of the overlapping candidate sentences found in the previous text, that is, for the preset topics not covered by the overlapping candidate sentences, further select candidate sentences containing the uncovered preset topics according to the TF-IDF ranking results and the TextRank ranking results, and use the selected candidate sentences and the overlapping candidate sentences as the target topic sentences.

[0141] For example, overlapping candidate sentences S1 and S2 contain the user's core question topic, but do not contain the two topics of direct appeal and key question.

[0142] Since the TextRank channel is better at capturing the semantic relevance of appeals (such as the mention of "urge it") in S5, S5 in the TextRank ranking results is used as the target topic sentence.

[0143] Since TF-IDF can better capture specific business points (such as the mention of "Can the handling fee be refunded?" in S4), S4 in the TF-IDF ranking results is used as the target topic sentence.

[0144] Therefore, the target topic sentences finally obtained include: S2, S1, S5 and S4.

[0145] Next, this embodiment can input the target topic sentence into the BERT model to obtain the corresponding global semantic features for the target topic sentence, namely, the global semantic features for S2, S1, S5, and S4. BERT (Bidirectional Encoder Representations from Transformers) is a pre-trained language model based on the Transformer architecture. It can output a 768-dimensional BERT vector labeled [CLS] as the global semantic feature.

[0146] Furthermore, the sentiment analysis model fine-tuned for customer service scenarios is used to perform sentiment analysis on the target topic sentence to obtain the sentiment intensity of the target topic sentence.

[0147] For example, the target topic sentence 1 is "A bank card was stolen and 5,000 yuan was used to pay for it", with an emotional intensity of -0.92, indicating a negative intensity of 0.92; the target topic sentence 2 is "Why didn't your risk control system intercept it?", with an emotional intensity of -0.87, indicating a negative intensity of 0.87 (containing an accusatory tone); the target topic sentence 3 is "The account must be frozen immediately", with an emotional intensity of -0.78, indicating a negative intensity of 0.78 (reflecting an urgent requirement); the target topic sentence 4 is "Otherwise I will complain to you", with an emotional intensity of -0.98, indicating a negative intensity of 0.98 (containing a threatening tone).

[0148] Since different target topic sentences have different representations of the overall sentiment, it is also necessary to perform importance analysis on the target topic sentences to assign weights to each target topic sentence.

[0149] Optionally, weight distribution can take into account the following three factors:

[0150] 1. Sentence length weight: Generally, longer sentences contain more detailed information and have higher weight.

[0151] 2. Keyword density: Sentences containing more business keywords (such as "fraudulent credit card," "freeze," "complaint," etc.) have higher weights.

[0152] 3. Position weight: Appeal sentences that appear later in the text data (such as threats or urgent requests) are often more important and have higher weights.

[0153] Finally, the emotional characteristic index of each target topic sentence can be obtained as: emotional intensity × weight.

[0154] Finally, the global semantic features of all target topic sentences can be formed into a matrix T1, and the emotional feature indexes of all target topic sentences can be formed into a matrix T2. The self-attention method based on the transformer architecture is used to obtain the target semantic features that integrate emotional information as text features.

[0155] This embodiment performs dual-channel independent processing on text data, which can filter out target topic sentences that contain more important information in the text data, and then perform sentiment tendency quantitative analysis on the target topic sentences, which can more accurately capture the sentiment tendency information in the target topic sentences, so that the target semantic features of the fused sentiment information are more comprehensive and accurate.

[0156] In other embodiments of the present application, the process of "generating target reply content of the intelligent customer service in the current round of interaction based on facial enhancement features, voice enhancement features and text enhancement features" in the previous step S104 is introduced.

[0157] In this embodiment, the facial enhancement features, voice enhancement features, and text enhancement features may be first fused to obtain the psychological portrait features of the user in the current round of interaction.

[0158] Then, based on the psychological portrait characteristics, multiple psychological state indicator values ​​of the user in the current round of interaction can be determined.

[0159] Optionally, the multiple psychological state index values ​​include an emotional activation value, a cognitive load value, and a trust value. Of course, the psychological state index value can also be other, and this application does not make specific limitations.

[0160] In one possible implementation, this embodiment can explicitly predict multiple psychological state indicator values ​​using independent regression heads. Specifically, a psychological indicator prediction model containing multiple regression heads is used to predict multiple psychological state indicator values. Here, the multiple regression heads correspond one-to-one to the multiple psychological state indicator values. This independent regression head approach allows the psychological indicator prediction model to focus on the different psychological cues that the corresponding psychological state indicators rely on during learning, thereby enhancing the accuracy and interpretability of the predictions.

[0161] Optionally, each regression head is a lightweight prediction sub-network composed of several fully connected layers, which, combined with a nonlinear activation function, can accurately predict the value of the psychological state indicator.

[0162] Based on this, this embodiment can input the psychological portrait features into the psychological index prediction model to obtain multiple psychological state index values ​​of the user in the current round of interaction.

[0163] Taking emotion activation value, cognitive load value and trust value as examples, the psychological index prediction model includes a prediction branch for emotion activation value, a prediction branch for cognitive load value and a prediction branch for trust value.

[0164] Among them, the prediction branch of the emotion activation value can learn to recognize the raised eyebrows in the face, the volume fluctuations in the audio, and the emotional words in the language, so as to judge the user's emotional arousal level.

[0165] The cognitive load value prediction branch focuses more on clues such as speaking rate, pauses, spectral changes in voice data, and syntactic complexity in text data to infer whether the user is engaged in high-intensity thinking or information processing.

[0166] The trust value prediction branch tends to comprehensively evaluate the user's current acceptance and trust in the intelligent customer service system based on features such as the softness of the tone, the degree of relaxation of the expression, and the presence of questioning, complaining, or positive feedback words in the text data.

[0167] All three prediction paths above use psychological profile features as input. Due to the separate network structures, each can learn the most useful modal feature combinations, thereby improving prediction performance. Ultimately, this embodiment outputs three values ​​ranging from 0 to 1, representing the user's emotional activation value, cognitive load value, and trust value in the current interaction round.

[0168] Finally, the target reply content is generated based on the user's multiple psychological state indicator values ​​in the current round of interaction.

[0169] Optionally, this embodiment can use a preset threshold to divide each mental state index value into three levels: high, medium, and low, and generate different target reply content according to the different levels. Here, the preset threshold can be set according to the actual scenario and is not limited here.

[0170] Still taking the emotion activation value, cognitive load value, and trust value as an example, optionally, the target reply content can be generated according to the following strategy.

[0171] When the emotional activation value is high and the cognitive load value is high, there is a first strategy as follows:

[0172] High trustworthiness: Maintain clear step-by-step instructions, take emotional counseling into account, and promote quick resolution;

[0173] Medium trust value: Strengthen emotional comfort, simplify information, and actively express understanding and concern;

[0174] Low trust value: Focus on calming emotions, increasing transparency, frequently confirming needs, and alleviating distrust.

[0175] When the emotional activation value is high and the cognitive load value is medium, there is a second strategy as follows:

[0176] High trustworthiness: respond quickly and in a steady tone, ensuring information transparency and efficient resolution;

[0177] Medium trust: Use concise language to soothe emotions, slow down the pace appropriately, and build communication trust;

[0178] Low trust value: Patiently appease, proactively commit to follow up, and enhance communication transparency.

[0179] When the emotional activation value is high and the cognitive load value is low, there is a third strategy as follows:

[0180] High trust value: respond quickly, maintain positive feedback, and prevent emotional escalation;

[0181] Medium trustworthiness: Speak softly, keep information open, and reduce intense conflict;

[0182] Low trust value: Continuous reassurance, enhanced sense of commitment, and reduced communication barriers.

[0183] When the emotional activation value is medium and the cognitive load value is high, there is a fourth strategy as follows:

[0184] High trustworthiness: Patiently and meticulously explain, pay attention to emotional fluctuations, and prevent cognitive burden from increasing;

[0185] Medium Trust: Maintain a steady pace, provide emotional support when appropriate, and enhance trust;

[0186] Low trust: Slow down the pace, simplify the instructions, and work hard to restore trust.

[0187] When the Emotional Activation value is Medium and the Cognitive Load value is Medium, there is a fifth strategy as follows:

[0188] High trust value: moderate advancement, maintaining good communication and trust;

[0189] Trust value: balance information depth and emotional support to maintain the quality of interaction;

[0190] Low trust value: Confirm user needs more often, proactively respond to doubts, and alleviate lack of trust.

[0191] When the emotional activation value is medium and the cognitive load value is low, there is a sixth strategy as follows:

[0192] High trust value: concise and efficient responses that meet user needs;

[0193] Trustworthiness: Avoid over-explanation, be concise and patient;

[0194] Low trust value: Be patient, ensure smooth communication, and increase trust building.

[0195] When the emotional activation value is low and the cognitive load value is high, there is a seventh strategy as follows:

[0196] High trustworthiness: Provide detailed guidance in a gentle tone, avoiding interrupting the user's train of thought.

[0197] Medium trust value: patiently guide, provide necessary help, and maintain a solid relationship;

[0198] Low trust value: Simplify instructions, emphasize transparency and commitment, and reduce user anxiety.

[0199] When the emotional activation value is low and the cognitive load value is medium, there is an eighth strategy as follows:

[0200] High trust value: Maintain stable communication and respond quickly to user needs;

[0201] Medium trust value: moderate guidance, maintaining smooth communication and emotional connection;

[0202] Low trust value: Increase commitment, confirm frequently, and reduce communication friction.

[0203] When the emotional activation value is low and the cognitive load value is low, there is a ninth strategy as follows:

[0204] High trustworthiness: Be concise and stay positive;

[0205] Medium Trust: Maintaining patience and transparency to ensure user comfort;

[0206] Low trust value: Continue to appease and actively build an atmosphere of trust.

[0207] For example, if a user enters "I've been waiting for this result for two hours, and there's been no response," and the emotional activation value is 0.85 (high), the cognitive load value is 0.30 (low), and the trust value is 0.4 (low), we can identify that the user is emotionally agitated but cognitively clear, and has significant distrust of the system. Therefore, we can employ the third strategy described above, "Continuous reassurance, enhanced commitment, and reduced communication barriers," to generate a targeted response that is soothing, transparent, and proactive in restoring trust. For example, "We apologize for the delay. We are urgently addressing your issue and are currently reviewing it. Thank you for your patience. We will provide a clear response within 15 minutes. If you have any questions, I will follow up with you throughout the process."

[0208] Optionally, in order to enhance the user's sense of interaction, a three-dimensional virtual image may be provided in the intelligent customer service system, and the three-dimensional virtual image may be made to synchronously perform expressions or actions to soothe the user.

[0209] It should be noted that the above strategies are merely examples and are not intended to limit this application. Other implementations are possible. For example, when conflicting signals are detected (e.g., high emotional activation values ​​and high trust values), a targeted response can be generated for the user to choose from, such as "You are currently quite agitated but still trust us. Should you prioritize calming your emotions or continuing to resolve the issue?" When a preset extreme situation is detected (e.g., cognitive overload and a low trust value persisting for more than 10 seconds), an alert is automatically triggered, forcing the human agent to be connected and pushing the conversation context to ensure that the human customer service representative can continue to provide services to the user.

[0210] In one possible implementation, considering the method of generating target response content based only on the conversation data of the current round of interaction, there may be a problem that the target response content is not accurate due to incomplete conversation data information. To this end, this embodiment also provides the following continuously optimized response strategy.

[0211] Optionally, the process of "generating target reply content based on multiple psychological state indicator values ​​of the user in the current round of interaction" may include: generating the initial reply content of the intelligent customer service in the current round of interaction based on multiple psychological state indicator values ​​of the user in the current round of interaction; obtaining multiple psychological state indicator values ​​of the user in historical multiple rounds of interaction; determining the user's psychological state fluctuation data based on the user's multiple psychological state indicator values ​​in historical multiple rounds of interaction and the user's multiple psychological state indicator values ​​in the current round of interaction; adjusting the initial reply content based on the psychological state fluctuation data to obtain the target reply content.

[0212] For example, this embodiment can generate initial reply content based on multiple psychological state indicator values ​​of the user in the current round of interaction using the strategy described above, and then perform statistical analysis on the changing trend of the user's psychological state based on historical data to obtain psychological state fluctuation data. Finally, the initial reply content is optimized based on the psychological state fluctuation data to obtain target reply content that better meets user needs.

[0213] For example, historical multi-round interactions refer to interactions within the past 5 seconds. This embodiment can adjust the initial reply content based on the user's psychological state fluctuation data within the past 5 seconds. For example, when the emotional activation value continues to be high, the length of the voice response is shortened and open-ended questions are increased. After the trust value recovers, the frequency of displaying the chain of evidence is reduced. If the cognitive load value drops suddenly, related information is automatically supplemented, etc.

[0214] Adopting this closed-loop mechanism can ensure that the interaction between intelligent customer service and users always matches the user's real-time psychological portrait, forming an adaptive communication rhythm and improving the user's interactive experience.

[0215] In summary, this embodiment integrates the user's real-time psychological state to implement emotional soothing, information transparency, and trust restoration strategies in intelligent customer service responses, thereby improving user satisfaction and service experience. Detailed response strategies are designed for different combinations of emotional activation values, cognitive load values, and trust values ​​to ensure effective responses even under complex and changing user conditions. For abnormal conditions, an early warning mechanism is automatically triggered, rapidly adjusting the response strategy to implement emotional soothing or redirect users to manual service, ensuring service stability and security, effectively preventing service interruptions or negative impacts caused by emotional outbursts, and improving the overall customer experience and risk control capabilities.

[0216] The above describes an intelligent customer service interaction method provided by an embodiment of the present application. The following describes an apparatus for executing the above intelligent customer service interaction method.

[0217] See also Figure 3 , is a structural diagram of an intelligent customer service interaction device provided in an embodiment of the present application, such as Figure 3 As shown, the intelligent customer service interaction device may include:

[0218] The data acquisition module 301 is used to acquire facial data, voice data, and text data of the user in the current round of interaction with the intelligent customer service. The text data refers to the text form of the voice data.

[0219] A feature extraction module 302 is used to extract features from facial data, voice data, and text data to obtain facial features, voice features, and text features;

[0220] A feature enhancement module 303 is configured to identify inconsistent emotional features among facial features, speech features, and text features, and to enhance the facial features, speech features, and text features according to the inconsistent emotional features to obtain facial enhancement features, speech enhancement features, and text enhancement features.

[0221] The customer service response generation module 304 is used to generate the target response content of the intelligent customer service in the current round of interaction based on the facial enhancement features, voice enhancement features and text enhancement features.

[0222] In one possible implementation, when determining facial features, the feature extraction module may be used to:

[0223] Perform face recognition on facial data to obtain the face area;

[0224] Extracting multiple key points in the face area and obtaining coordinate values ​​of the multiple key points based on the facial data, wherein the multiple key points are feature points reflecting the user's expression and emotion;

[0225] Calculate geometric features and extract texture features based on the coordinate values ​​of multiple key points and facial data to obtain the user's geometric features and texture features;

[0226] Generate facial features based on geometric features and texture features.

[0227] In one possible implementation, when determining speech features, the feature extraction module may be specifically used to:

[0228] Preprocessing the voice data to obtain preprocessed voice data;

[0229] Extracting prosodic and spectral features from preprocessed speech data;

[0230] Generate speech features based on prosodic and spectral features.

[0231] In one possible implementation, when determining text features, the feature extraction module may be used to:

[0232] Split the text data into multiple logically complete sentences as multiple candidate sentences;

[0233] The target topic sentence is selected from multiple candidate sentences using the word frequency inverse document frequency method and text sorting method;

[0234] Process the target topic sentence into global semantic features;

[0235] Perform sentiment analysis and importance analysis on the target topic sentence to obtain the sentiment intensity and weight of the target topic sentence, where the sentiment intensity reflects the intensity of the negative sentiment in the target topic sentence;

[0236] According to the sentiment intensity and weight of the target topic sentence and the global semantic features, the target semantic features that integrate the sentiment information are generated as text features.

[0237] In one possible implementation, when the customer service response generation module generates the target response content of the intelligent customer service in the current round of interaction based on the facial enhancement features, voice enhancement features, and text enhancement features, it can be specifically used to:

[0238] Perform feature fusion on facial enhancement features, voice enhancement features, and text enhancement features to obtain the user's psychological portrait features in the current round of interaction;

[0239] Determine the user's multiple psychological state indicator values ​​in the current round of interaction based on the psychological portrait characteristics;

[0240] Generate target reply content based on multiple psychological state indicator values ​​of the user in the current round of interaction.

[0241] In one possible implementation, when the customer service response generation module determines multiple psychological state indicator values ​​of the user in the current round of interaction based on the psychological profile characteristics, it can be specifically used to:

[0242] The psychological portrait features are input into the psychological index prediction model to obtain multiple psychological state index values ​​of the user in the current round of interaction, wherein the psychological index prediction model includes regression heads corresponding to the multiple psychological state index values.

[0243] In one possible implementation, when the customer service response generation module generates target response content based on multiple psychological state indicator values ​​of the user in the current round of interaction, it can be specifically used to:

[0244] Generate the initial response content of the intelligent customer service in the current round of interaction based on the user's multiple psychological state indicator values ​​in the current round of interaction;

[0245] Obtain multiple psychological state indicator values ​​of users in multiple rounds of historical interactions;

[0246] Determining the user's psychological state fluctuation data based on multiple psychological state indicator values ​​of the user in multiple rounds of historical interactions and multiple psychological state indicator values ​​of the user in the current round of interactions;

[0247] The initial response content is adjusted according to the psychological state fluctuation data to obtain the target response content.

[0248] In a possible implementation, the multiple psychological state indicator values ​​include an emotion activation value, a cognitive load value, and a trust value.

[0249] Each module in the intelligent customer service interaction device described above can be implemented in whole or in part through software, hardware, or a combination thereof. Each module can be embedded in or independent of a processor in a computer device in hardware form, or can be stored in a computer device memory in software form, so that the processor can call and execute the corresponding operations of each module.

[0250] An embodiment of the present application further provides an electronic device, which may include at least one processor and a memory connected to the processor, wherein:

[0251] Memory is used to store computer programs;

[0252] The processor is used to execute a computer program so that the electronic device can implement any intelligent customer service interaction method provided in the embodiments of the present application.

[0253] refer to Figure 4, which shows a schematic diagram of the structure of an electronic device suitable for implementing the embodiments of the present application. The electronic device in the embodiments of the present application may include, but is not limited to, fixed terminals such as mobile phones, laptops, PDAs (personal digital assistants), PADs (tablet computers), desktop computers, etc. Figure 4 The electronic device shown is merely an example and should not limit the functions and scope of use of the embodiments of the present application.

[0254] like Figure 4 As shown, the electronic device may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 401, which can perform various appropriate actions and processes based on programs stored in a read-only memory (ROM) 402 or programs loaded from a storage device 408 into a random access memory (RAM) 403. When the electronic device is powered on, the RAM 403 also stores various programs and data required for the operation of the electronic device. The processing device 401, ROM 402, and RAM 403 are interconnected via a bus 404. An input / output (I / O) interface 405 is also connected to the bus 404.

[0255] Typically, the following devices may be connected to the I / O interface 405: an input device 406 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 407 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 408 including, for example, a memory card, a hard disk, etc.; and a communication device 409. The communication device 409 may allow the electronic device to communicate with other devices wirelessly or by wire to exchange data. Figure 4 The electronic device is shown with various devices, but it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed instead.

[0256] An embodiment of the present application also provides a computer program product including computer-readable instructions. When the computer-readable instructions are executed on an electronic device, the electronic device implements any one of the intelligent customer service interaction methods provided in the embodiments of the present application.

[0257] A computer-readable storage medium is also provided in an embodiment of the present application. The storage medium carries one or more computer programs. When one or more computer programs are executed by an electronic device, the electronic device can implement any intelligent customer service interaction method provided in an embodiment of the present application.

[0258] It should also be noted that the device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separate, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed across multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the present embodiment. In addition, in the drawings of the device embodiments provided in this application, the connection relationship between the modules indicates that there is a communication connection between them, which can be specifically implemented as one or more communication buses or signal lines.

[0259] Through the description of the above embodiments, those skilled in the art can clearly understand that the present application can be implemented by means of software plus necessary general hardware, and of course can also be implemented by special hardware including application-specific integrated circuits, special CPUs, special memories, special components, etc. In general, all functions performed by computer programs can be easily implemented with corresponding hardware, and the specific hardware structures used to implement the same function can also be diverse, such as analog circuits, digital circuits or special circuits, etc. However, for the present application, software program implementation is a better implementation method in most cases. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which is stored in a readable storage medium, such as a computer's floppy disk, USB flash drive, mobile hard disk, ROM, RAM, magnetic disk or optical disk, etc., and includes a number of instructions to enable a computer device (which can be a personal computer, training equipment, or network equipment, etc.) to execute the methods described in each embodiment of the present application.

[0260] In the above embodiments, all or part of the embodiments may be implemented by software, hardware, firmware, or any combination thereof. When implemented by software, all or part of the embodiments may be implemented in the form of a computer program product.

[0261] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website, a computer, a training device or a data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) mode to another website, a computer, a training device or a data center. The computer-readable storage medium can be any available medium that a computer can store or a data storage device such as a training device, a data center, etc. that includes one or more available media integrations. The available medium can be a magnetic medium, (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive (SSD)).

Claims

1. An intelligent customer service interaction method, characterized in that: include: Acquire facial data, voice data, and text data of the user in the current round of interaction with the intelligent customer service, where the text data refers to the text form of the voice data; Extracting features from the facial data, the voice data, and the text data to obtain facial features, voice features, and text features; Identifying inconsistent emotional features among the facial features, the voice features, and the text features, and performing feature enhancement on the facial features, the voice features, and the text features according to the inconsistent emotional features to obtain facial enhancement features, voice enhancement features, and text enhancement features; Generate target reply content of the intelligent customer service in the current round of interaction based on the facial enhancement features, the voice enhancement features, and the text enhancement features.

2. The intelligent customer service interaction method according to claim 1, characterized in that: The process of determining the facial features includes: Performing face recognition on the facial data to obtain a face area; Extracting a plurality of key points in the face region, and obtaining coordinate values ​​of the plurality of key points based on the facial data, wherein the plurality of key points are feature points reflecting the expression and emotion of the user; Calculating geometric features and extracting texture features based on the coordinate values ​​of the multiple key points and the facial data to obtain geometric features and texture features of the user; The facial features are generated according to the geometric features and the texture features.

3. The intelligent customer service interaction method according to claim 1, characterized in that: The process of determining the speech features includes: Preprocessing the voice data to obtain preprocessed voice data; extracting prosodic features and spectral features from the preprocessed speech data; The speech feature is generated according to the prosodic feature and the spectral feature.

4. The intelligent customer service interaction method according to claim 1, characterized in that: The process of determining the text features includes: Splitting the text data into a plurality of logically complete sentences as a plurality of candidate sentences; Selecting a target topic sentence from the plurality of candidate sentences by using a word frequency inverse document frequency method and a text sorting method; processing the target topic sentence into a global semantic feature; Performing sentiment analysis and importance analysis on the target topic sentence to obtain the sentiment intensity and weight of the target topic sentence, wherein the sentiment intensity reflects the intensity of the negative sentiment in the target topic sentence; According to the sentiment intensity and weight of the target topic sentence and the global semantic feature, a target semantic feature integrating sentiment information is generated as the text feature.

5. The intelligent customer service interaction method according to any one of claims 1 to 4, characterized in that: Generating target reply content of the intelligent customer service in the current round of interaction based on the facial enhancement feature, the voice enhancement feature, and the text enhancement feature includes: Performing feature fusion on the facial enhancement feature, the voice enhancement feature, and the text enhancement feature to obtain a psychological portrait feature of the user in the current round of interaction; Determining, based on the psychological profile features, multiple psychological state indicator values ​​of the user in the current round of interaction; The target reply content is generated according to the multiple psychological state indicator values ​​of the user in the current round of interaction.

6. The intelligent customer service interaction method according to claim 5, characterized in that: Determining, based on the psychological portrait features, multiple psychological state indicator values ​​of the user in the current round of interaction includes: The psychological portrait features are input into a psychological index prediction model to obtain multiple psychological state index values ​​of the user in the current round of interaction, wherein the psychological index prediction model includes regression heads corresponding to the multiple psychological state index values ​​respectively.

7. The intelligent customer service interaction method according to claim 5, characterized in that: Generating the target reply content according to the multiple psychological state indicator values ​​of the user in the current round of interaction includes: generating, based on the multiple psychological state indicator values ​​of the user in the current round of interaction, an initial reply content of the intelligent customer service in the current round of interaction; Obtaining multiple psychological state indicator values ​​of the user in multiple rounds of historical interactions; Determining the user's psychological state fluctuation data based on multiple psychological state indicator values ​​of the user in multiple rounds of historical interactions and multiple psychological state indicator values ​​of the user in the current round of interactions; The initial reply content is adjusted according to the mental state fluctuation data to obtain the target reply content.

8. The intelligent customer service interaction method according to claim 5, characterized in that: The multiple psychological state index values ​​include an emotion activation value, a cognitive load value, and a trust value.

9. An intelligent customer service interaction device, characterized in that: include: A data acquisition module is used to acquire facial data, voice data, and text data of the user in the current round of interaction with the intelligent customer service, where the text data refers to the text form of the voice data; A feature extraction module, configured to extract features from the facial data, the voice data, and the text data to obtain facial features, voice features, and text features; a feature enhancement module, configured to identify inconsistent emotional features among the facial features, the voice features, and the text features, and to enhance the facial features, the voice features, and the text features according to the inconsistent emotional features to obtain facial enhancement features, voice enhancement features, and text enhancement features; The customer service response generation module is used to generate the target response content of the intelligent customer service in the current round of interaction based on the facial enhancement features, the voice enhancement features and the text enhancement features.

10. An electronic device, characterized in that: comprising at least one processor and a memory connected to the processor, wherein: The memory is used to store computer programs; The processor is used to execute the computer program so that the electronic device can implement the intelligent customer service interaction method as described in any one of claims 1 to 8.