Customer service question answering method and device, equipment, storage medium and computer program product

By generating digital human expressions and facial animations and providing voice broadcasts, the problem of low efficiency in text-based responses from intelligent customer service has been solved, enabling more intelligent and vivid user interaction and improving the convenience and adaptability of online banking services.

CN114694224BActive Publication Date: 2025-12-30INDUSTRIAL AND COMMERCIAL BANK OF CHINA
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202210323198.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-30
Publication Date
2025-12-30
Estimated Expiration
2042-03-30

AI Technical Summary

Technical Problem

Existing intelligent customer service systems are inefficient at providing text responses via chat boxes, reducing the convenience for customers to conduct banking transactions online.

Method used

By acquiring user questions, matching emotional responses with question responses, and generating digital human expressions and facial animations, and then synthesizing them with voice broadcasts, intelligent and vivid interaction with users can be achieved.

Benefits of technology

It improves the convenience and flexibility of online banking for users, enhances the intelligence and adaptability of customer service Q&A, and simulates the experience of face-to-face conversations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114694224B_ABST
    Figure CN114694224B_ABST
Patent Text Reader

Abstract

The application relates to a customer service question answering method and device, computer equipment, a storage medium and a computer program product, relates to the field of artificial intelligence, and can be used in the field of digital finance or other fields. The method comprises the following steps: acquiring reply data matched with a question input by a user on a question answering interface; the reply data comprises emotion reply data and question reply data; inputting the reply data into a preset animation model for processing to generate an expression and a facial animation of a digital person; synthesizing the expression of the digital person and the facial animation of the digital person to generate an animation of the digital person, and displaying the animation of the digital person on the question answering interface; the animation of the digital person is used for displaying an expression corresponding to the expression data and performing voice broadcast on the question reply data. The method can improve the convenience of online bank business handling of customers.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular to a customer service question-and-answer method, apparatus, computer equipment, storage medium, and computer program product. Background Technology

[0002] With the proliferation of online banking services, the importance of online customer service has become increasingly apparent. When customers encounter business-related issues, they expect immediate assistance. As transaction volume grows, intelligent customer service technologies have been introduced to handle these issues. Specifically, customers can input their questions through a dialog box, and the intelligent customer service system matches the input with relevant answers, providing the solutions directly to the customer.

[0003] However, existing intelligent customer service systems only respond to customer inquiries by providing text replies through chat boxes. Text-based responses are inefficient and reduce the convenience for customers to conduct banking transactions online. Summary of the Invention

[0004] Therefore, it is necessary to provide a customer service Q&A method, device, computer equipment, computer-readable storage medium, and computer program product that can improve the convenience of customers conducting online banking business, in response to the above-mentioned technical problems.

[0005] Firstly, this application provides a customer service question-and-answer method. The method includes:

[0006] Acquire response data that matches the questions entered by the user on the Q&A interface; the response data includes emotional response data and question response data; input the response data into a preset animation model for processing to generate the digital human's expressions and facial animations; synthesize the digital human's expressions and facial animations to generate the digital human's animation, and display the digital human's animation on the Q&A interface; the digital human's animation is used to display the expressions corresponding to the expression data, and to provide voice broadcast of the question response data.

[0007] In one embodiment, the facial animation includes multiple sets of mouth image frames; the response data is input into a preset animation model for processing to generate the digital human's expressions and facial animations, including:

[0008] The system matches the expressions of the digital human corresponding to the emotion response data from a preset expression library. The preset expression library stores the correspondence between the emotion response data and the expressions of the digital human. The question response data is input into a preset animation model for processing to generate multiple sets of mouth bone point graphics of the digital human corresponding to the question response data. The mouth bone point graphics include multiple mouth bone points at the target positions. Based on the multiple sets of mouth bone point graphics of the digital human, multiple mouth image frames of the digital human are generated.

[0009] In one embodiment, if the question response data is text data, the question response data is input into a preset animation model for processing to generate multiple sets of mouth bone point graphics of the digital human corresponding to the question response data, including:

[0010] The question response data is formatted and converted to generate audio data corresponding to the question response data; the audio data is divided into multiple audio data units sequentially using preset audio segmentation rules; for each audio data unit, the target positions of multiple mouth bone points of the digital human are determined when the audio data unit is emitted; based on the multiple mouth bone points at the target positions, multiple sets of mouth bone point graphics are generated.

[0011] In one embodiment, the digital human's facial expressions are synthesized with the digital human's facial animation to generate an animation of the digital human, including:

[0012] Based on the chronological order of the appearance of multiple audio data units in the audio data, multiple mouth image frames of the digital human corresponding to the audio data units are arranged to generate the mouth animation of the digital human corresponding to the question response data; the digital human's facial expressions and the digital human's mouth animation are then combined to generate the digital human's animation.

[0013] In one embodiment, obtaining response data that matches the question entered by the user on the question-and-answer interface includes:

[0014] The process involves inputting the user's question into a pre-defined emotion recognition model for emotion recognition, generating the user's emotion recognition result, and determining the corresponding emotion response data. Based on the emotion response data and the user's question on the Q&A interface, question response data is generated. Finally, based on the emotion response data and the question response data, response data is generated.

[0015] In one embodiment, the question entered by the user on the question-and-answer interface is input into a preset emotion recognition model for processing, generating the user's emotion recognition result, including:

[0016] The questions entered by users on the question-and-answer interface are fed into sentence-level semantic feature extraction models and word-level semantic feature extraction models for feature extraction, generating sentence-level and word-level semantic features of the questions; the user's first sentiment data is generated based on the sentence-level semantic features, and the user's second sentiment data is generated based on the multi-word phrase semantic features; the user's sentiment recognition result is determined from the first and second sentiment data using preset filtering rules.

[0017] In one embodiment, a preset filtering rule is used to determine the user's current emotion recognition result from the first emotion data and the second emotion data, including:

[0018] Determine whether the first sentiment data belongs to the positive type and whether the second sentiment data belongs to the negative type; if not, then determine the first sentiment data as the user's current sentiment recognition result.

[0019] In one embodiment, the method further includes:

[0020] If the first sentiment data is positive and the second sentiment data is negative, then the second sentiment data is determined as the user's current sentiment recognition result.

[0021] Secondly, this application also provides a customer service Q&A device. The device includes:

[0022] The acquisition module is used to acquire response data that matches the questions entered by the user on the Q&A interface; the response data includes sentiment response data and question response data;

[0023] The first generation module is used to input the response data into a preset animation model for processing, and generate the digital human's expressions and facial animations.

[0024] The second generation module is used to synthesize the digital human's facial expressions and facial animations to generate the digital human's animation, which is then displayed on the question-and-answer interface. The digital human's animation is used to display the facial expressions corresponding to the expression data and to provide voice broadcasts of the question and answer data.

[0025] Thirdly, this application also provides a computer device. The computer device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the method steps of any of the embodiments of the first aspect described above.

[0026] Fourthly, this application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program thereon, which, when executed by a processor, implements the method steps of any of the embodiments of the first aspect described above.

[0027] Fifthly, this application also provides a computer program product. The computer program product includes a computer program that, when executed by a processor, implements the method steps described in any of the embodiments of the first aspect.

[0028] The aforementioned customer service Q&A method, device, computer equipment, storage medium, and computer program product are implemented. In the technical solution provided in this application embodiment, response data matching the questions entered by the user on the Q&A interface is obtained; the response data is input into a preset animation model for processing to generate digital human expressions and facial animations; the digital human expressions and facial animations are synthesized to generate digital human animations, which are then displayed on the Q&A interface; the response data includes emotional response data and question response data; the digital human animation is used to display expressions corresponding to the expression data and to provide voice broadcasts of the question response data. Compared with existing technologies, this approach does not rely on traditional dialog boxes to respond to user questions. Using a digital human allows for voice broadcasts of response data matching the user's questions, and the display is achieved through digital human animations. This enables more intelligent and vivid interaction with users, providing a sense of immersion similar to face-to-face conversations, thereby improving the convenience and flexibility of online banking transactions. Furthermore, matching corresponding expressions to the digital human based on the user's input questions better adapts to different user emotions, further enhancing the intelligence and flexibility of customer service Q&A. Attached Figure Description

[0029] Figure 1 This is an internal structural diagram of a computer device in one embodiment;

[0030] Figure 2 This is a flowchart illustrating a customer service Q&A method in one embodiment;

[0031] Figure 3 This is a flowchart illustrating the process of generating expressions and facial animations for a digital human in one embodiment.

[0032] Figure 4 This is a schematic diagram of the process for generating multiple sets of mouth bone point graphics for a digital human in one embodiment;

[0033] Figure 5 This is a schematic diagram of the process for generating an animation of a digital human in one embodiment;

[0034] Figure 6 This is a schematic diagram of the process for obtaining response data in one embodiment;

[0035] Figure 7 This is a flowchart illustrating the process of collecting sentiment data from native users in one embodiment.

[0036] Figure 8 This is a flowchart illustrating the process of generating a user's emotion recognition result in one embodiment;

[0037] Figure 9 This is a flowchart illustrating the customer service Q&A method in yet another embodiment;

[0038] Figure 10 This is a structural block diagram of a customer service Q&A device in one embodiment. Detailed Implementation

[0039] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0040] The customer service Q&A method provided in this application can be applied to computer devices, which can be servers or terminals. The server can be a single server or a server cluster composed of multiple servers. This application does not specifically limit this. The terminal can be, but is not limited to, various personal computers, laptops, smartphones, tablets, and portable wearable devices.

[0041] Taking a computer device as an example, Figure 1 A block diagram of a server is shown, such as Figure 1 As shown, the computer device includes a processor, memory, and a network interface connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The database stores customer service question-and-answer data. The network interface communicates with external terminals via a network connection. When executed by the processor, the computer program implements a customer service question-and-answer method.

[0042] Those skilled in the art will understand that Figure 1 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the server to which the present application is applied. Optionally, the server may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements.

[0043] It should be noted that the execution subject of this application embodiment can be a computer device or a customer service question and answer device. The following method embodiment will be described with a computer device as the execution subject.

[0044] In one embodiment, such as Figure 2 As shown, a flowchart of a customer service Q&A method provided in an embodiment of this application is illustrated. The method may include the following steps:

[0045] Step 220: Obtain response data that matches the questions entered by the user on the Q&A interface; the response data includes sentiment response data and question response data.

[0046] With the increasing prevalence of online banking services, users can more conveniently conduct various transactions online. When customers encounter business-related issues, they can input their questions via voice, text, or other means on the Q&A interface and obtain corresponding response data. By processing and analyzing the questions entered by users on the Q&A interface, matching response data can be obtained, including sentiment-based response data and question-based response data.

[0047] Emotional response data is used to address a user's current emotion. For example, if a user is angry, the emotional response data could be apologetic voice messages, text messages, emojis, or gestures to comfort the user. Question response data is used to answer users' business-related questions. For example, if a user asks how to set up fingerprint payment for mobile banking, the question response data could be the steps to set up fingerprint payment for mobile banking, provided through voice, text, or images.

[0048] Step 240: Input the response data into the preset animation model for processing to generate the digital human's expressions and facial animations.

[0049] The preset animation model is trained based on historical response data and reference facial expressions and animations of the corresponding digital human. Historical response data is input into the initial animation model for processing, generating predicted facial expressions and animations corresponding to the historical responses. These predicted and reference facial expressions and animations are then fed into a preset loss function for calculation, updating the model parameters of the initial animation model based on the loss function value until a preset convergence condition is met. Finally, the preset animation model is generated based on the updated model parameters. In practical use, the acquired response data is input into the preset animation model for processing, generating the digital human's facial expressions and animations. The digital human's expressions can include facial expressions, speech tone expressions, and body posture expressions; for example, a facial expression can be a smile. Facial animation is the dynamic change process of the digital human's face when uttering speech; for example, the digital human's face changes accordingly when uttering different speech sounds.

[0050] Digital humans are virtual beings based on artificial intelligence technologies such as image recognition, speech recognition and synthesis, semantic understanding, and human modeling. They possess the ability to perceive, recognize, and express the physical world, and interact with humans through devices such as electronic screens and VR. Digital humans offer new intelligent customer services to industries such as finance, broadcasting, education, marketing, healthcare, retail, and gaming, reducing labor costs and improving service quality and efficiency. By constructing virtual human figures or cartoon characters using computer technology, digital humans can interact naturally with users, providing warm and personalized service. Through a visual dimension, they not only enrich the user experience and interaction but also allow for a deeper understanding of users during the service process, leading to better customer service.

[0051] Step 260: Combine the digital human's facial expressions with the digital human's facial animation to generate the digital human's animation, and display the digital human's animation on the question-and-answer interface; the digital human's animation is used to display the facial expressions corresponding to the facial expression data, and to provide voice broadcast of the question and answer data.

[0052] The digital human's facial animation is composed of multiple consecutive image frames. The images corresponding to the digital human's expressions can be overlaid with each frame of the facial animation to synthesize the digital human's expressions and facial animation. Other synthesis methods can also be used, but this embodiment does not specifically limit them. The resulting digital human animation simultaneously includes both expressions and facial animation. The facial animation allows for voice broadcasting of question responses, with corresponding facial expressions accompanying the voice broadcast. By displaying the digital human's animation on the question-and-answer interface, users can obtain a response to their inquiry.

[0053] In this embodiment, response data matching the questions entered by the user on the Q&A interface is acquired; the response data is input into a preset animation model for processing to generate digital human expressions and facial animations; the digital human expressions and facial animations are synthesized to generate digital human animation, which is then displayed on the Q&A interface; the response data includes emotional response data and question response data; the digital human animation is used to display expressions corresponding to the expression data and to provide voice broadcast of the question response data. Compared with existing technologies, this method does not rely on traditional dialog boxes to reply to user questions. Using a digital human, the response data matching the user's question can be broadcast voice-over and displayed through digital human animation, enabling a more intelligent and vivid interaction with the user. This achieves a sense of immersion similar to face-to-face conversations, thereby improving the convenience and flexibility of online banking transactions for users. Furthermore, matching corresponding expressions to the digital human based on the user's input questions better adapts to different user emotions, further enhancing the intelligence and flexibility of customer service Q&A.

[0054] In one embodiment, such as Figure 3 The diagram illustrates a flowchart of a customer service Q&A process provided in an embodiment of this application. Specifically, it relates to a possible process for generating digital human expressions and facial animations. This method may include the following steps:

[0055] Step 320: Match the digital human's expression to the emotion response data from the preset expression library; the preset expression library stores the correspondence between emotion response data and digital human expressions in advance.

[0056] The system includes a pre-stored emoji library that maps emotional response data to digital human expressions. This mapping can be set based on human experience or derived from analyzing large amounts of historical data. The system can then search the library for the corresponding digital human expression based on the emotional response data. For example, if the emotional response data is "Dear user, I'm sorry, we are working hard to improve, and we apologize for the poor experience," the matched digital human expression should be something like "polite" or "apology." The emotional response data can have a one-to-one correspondence with the digital human's expressions, or it can correspond to multiple expressions. Users can choose any expression from multiple options or select multiple expressions simultaneously and switch between them. This embodiment does not impose specific limitations on this approach.

[0057] Step 340: Input the question response data into the preset animation model for processing to generate multiple sets of mouth bone point graphics corresponding to the question response data; the mouth bone point graphics include multiple mouth bone points at the target position.

[0058] The digital human's facial animation can include multiple sets of mouth image frames. By inputting question response data into a preset animation model for processing, multiple sets of mouth skeletal point graphics corresponding to the question response data are generated for the digital human. Each mouth skeletal point graphic can include multiple mouth skeletal points at target positions. The mouth skeletal point graphics can be formed based on the movement of mouth skeletal points corresponding to different pronunciations. The more mouth skeletal points corresponding to different pronunciations in the mouth skeletal point graphics, the better, resulting in more delicate mouth movements during pronunciation and a more accurate match to the pronunciation.

[0059] Step 360: Generate multiple mouth image frames of the digital human based on multiple sets of mouth skeletal point graphics.

[0060] In this process, after sorting and rendering multiple sets of mouth skeletal point graphics of the digital human, corresponding images are formed. Furthermore, image processing operations can be performed on the generated images to generate multiple mouth image frames of the final digital human. Image processing operations may include, but are not limited to, image smoothing, noise reduction, and size scaling.

[0061] In this embodiment, the expressions of the digital human corresponding to the emotion response data are matched from a preset expression library; the question response data is input into a preset animation model for processing to generate multiple sets of mouth bone point graphics of the digital human corresponding to the question response data; and multiple mouth image frames of the digital human are generated based on the multiple sets of mouth bone point graphics, thereby enabling more accurate and faster acquisition of the digital human's expressions and facial animations.

[0062] In one embodiment, such as Figure 4 As shown, it illustrates a flowchart of a customer service Q&A process provided in an embodiment of this application, specifically involving a possible process for generating multiple sets of mouth bone point graphics of a digital human. This method may include the following steps:

[0063] Step 420: Convert the format of the question response data to generate audio data corresponding to the question response data.

[0064] If the question response data is text data, it can be processed using a text-to-speech tool to obtain the corresponding audio data. When processing with the text-to-speech tool, the conversion can be performed either in real-time by acquiring the user's input text data or after the user has finished inputting all the text data; this embodiment does not specify a particular method.

[0065] Step 440: Divide the audio data into multiple audio data units sequentially using preset audio segmentation rules.

[0066] The audio data includes the audio data corresponding to each character in the question-and-answer data. For each character's audio data, a preset audio segmentation rule can be used to divide the audio data into multiple audio data units. The preset audio segmentation rule can be set based on pronunciation time. For example, if the total pronunciation time for a character is 0.5 seconds, then that 0.5 seconds can be divided into three time segments, each corresponding to a different audio data unit. It should be noted that the pronunciation time for each character can be divided into different numbers of time segments, or the same number of time segments, depending on actual needs; this embodiment does not impose a specific limitation on this.

[0067] Step 460: For each audio data unit, determine the target positions of multiple mouth bone points of the digital human when the audio data unit is emitted.

[0068] Based on the correspondence between audio data units and the target positions of mouth bone points, for each audio data unit, the target positions of multiple mouth bone points of the digital human when emitting the audio data unit can be matched. Different audio data units correspond to different target positions of mouth bone points, and the target positions of multiple mouth bone points when emitting the audio data unit can be represented by two-dimensional coordinate information.

[0069] Step 480: Generate multiple sets of mouth bone point graphics based on multiple mouth bone points located at the target position.

[0070] This allows for the determination of multiple mouth bone points at the target position when the audio data unit is emitted, based on the two-dimensional coordinate information of the mouth bone points. This enables the generation of a mouth bone point graphic for each character's audio data. By integrating the mouth bone point graphics of all characters in the audio data, multiple sets of mouth bone point graphics are obtained.

[0071] In this embodiment, the question response data is formatted to generate corresponding audio data. A preset audio segmentation rule is used to divide the audio data into multiple audio data units. For each audio data unit, the target positions of multiple mouth bone points of the digital human are determined when the audio data unit is emitted. Based on the multiple mouth bone points at the target positions, multiple sets of mouth bone point graphics are generated. By segmenting the audio data, the position of the mouth bone points corresponding to each audio data unit can be accurately obtained, thus making the final generated mouth bone point graphics more accurate.

[0072] In one embodiment, such as Figure 5 As shown, it illustrates a flowchart of a customer service Q&A process provided in an embodiment of this application, specifically relating to a possible process for generating animations of digital humans. This method may include the following steps:

[0073] Step 520: Arrange multiple mouth image frames of the digital human corresponding to the audio data units according to the order of their appearance in the audio data, and generate a mouth animation of the digital human corresponding to the question response data.

[0074] When converting question-and-answer data into audio data, the conversion must follow the order of the text data to ensure the audio data follows the same order. First, multiple mouth skeletal point graphics corresponding to the audio data units are obtained. These graphics are then processed to generate multiple mouth image frames for the corresponding audio data units. These mouth image frames are then arranged according to the chronological order of their appearance in the audio data. Finally, playing these arranged mouth image frames generates the mouth animation of the digital human corresponding to the question-and-answer data.

[0075] Step 540: Combine the digital human's facial expressions with the digital human's mouth animation to generate the digital human's animation.

[0076] In this embodiment, the digital human's facial expression is synthesized by overlaying the image corresponding to the digital human's facial expression with each mouth image frame in the mouth animation. Other synthesis methods can also be used, but this embodiment does not specifically limit them.

[0077] In this embodiment, multiple mouth image frames of the digital human corresponding to each audio data unit are arranged according to the chronological order of their appearance in the audio data to generate a mouth animation of the digital human corresponding to the question response data. The digital human's facial expression and mouth animation are then combined to generate the digital human's animation. Because the mouth animation of the digital human can be generated according to the order of the audio data units, the accuracy of the voice broadcast of the question response data is ensured.

[0078] In one embodiment, such as Figure 6 As shown, it illustrates a flowchart of a customer service Q&A process provided in an embodiment of this application, specifically involving a possible process for obtaining response data. This method may include the following steps:

[0079] Step 620: Input the question entered by the user on the question-and-answer interface into the preset emotion recognition model for emotion recognition, generate the user's emotion recognition result, and determine the emotion response data corresponding to the user's emotion recognition result.

[0080] The preset sentiment recognition model is trained based on historical question data and corresponding reference sentiment recognition results. Historical question data is input into the initial sentiment recognition model for processing, generating predicted sentiment recognition results corresponding to the historical question data. These predicted and reference sentiment recognition results are then fed into a preset loss function for calculation, updating the model parameters of the initial sentiment recognition model based on the loss function value until a preset convergence condition is met. Finally, the preset sentiment recognition model is generated based on the updated model parameters. In practical use, the user's question input on the question-and-answer interface is fed into the preset sentiment recognition model for sentiment recognition, generating the user's sentiment recognition result. Emotional response data corresponding to the user's sentiment recognition result is matched from a preset question database, or emotional response data can be generated in real-time based on the user's sentiment recognition result.

[0081] Step 640: Generate question response data based on the emotion response data and the questions entered by the user on the question-and-answer interface.

[0082] Specifically, the system obtains data on the business questions that users ask by matching the questions they enter on the Q&A interface. This data can be obtained from a pre-set question database. By extracting keywords from the user's questions and matching them against the pre-set question database, the system finds answers that correspond to the keywords. These answers are pre-organized and reviewed by experts.

[0083] This process combines emotional response data with data related to answering users' business questions into a complete sentence, which serves as the answer to the question, thus generating question response data. For example, if a user enters, "What kind of garbage service is this? Why can't I select the person I want to transfer money to?", the resulting emotional response data would be, "Dear user, we are constantly working to improve. We sincerely apologize for the poor experience." The data for answering the user's business question would be, "Relevant business information and operating methods for money transfers." These two parts of data are then combined into a complete sentence to generate question response data.

[0084] Step 660: Generate response data based on emotion response data and question response data.

[0085] Specifically, based on the emotion response data, digital human expressions and facial animations can be generated. Based on the digital human's expressions and facial animations, along with the question response data, the final response data is then generated. The specific processing procedure is similar to the above embodiments and will not be repeated here.

[0086] In this embodiment, the user's question input on the Q&A interface is fed into a preset emotion recognition model for emotion recognition, generating the user's emotion recognition result and determining the corresponding emotional response data. Based on the emotional response data and the user's question input on the Q&A interface, question response data is generated. Finally, based on the emotional response data and the question response data, overall response data is generated. By dividing the question response data into two parts, a more intelligent and vivid interaction with the user can be achieved, and the wording when responding to the user is more flexible and accurate.

[0087] In one embodiment, such as Figure 7 As shown, it illustrates a flowchart of a customer service Q&A method provided in an embodiment of this application, specifically involving a possible process for generating user sentiment data. This method may include the following steps:

[0088] Step 720: Input the questions entered by the user on the question-and-answer interface into the sentence-level semantic feature extraction model and the word-level semantic feature extraction model respectively for feature extraction, and generate the sentence-level semantic features and word-level semantic features of the questions.

[0089] The sentence-level semantic feature extraction model is a simple word vector-based model called SWEM-aver. Specifically, it extracts sentence-level semantic features by averaging the word vectors element-wise using average pooling, thus obtaining sentence-level semantic features from user speech. The word-level semantic feature extraction model is an improved CNN model. Specifically, the traditional CNN model extracts n-gram semantic features, where n is the size of the convolutional window, which can be set to 2, 3, or 4. For each window size, 14 convolutional kernels are set to extract rich n-gram semantic information from the original word vector matrix.

[0090] By incorporating specific hyperparameters during CNN model training and adding Dropout to fully connected layers to randomly discard some connections, overfitting is effectively prevented. Furthermore, the pooling layer is modified to use k-Max pooling to retain more features, resulting in an improved CNN model. By inputting user-defined questions into sentence-level and word-level semantic feature extraction models respectively, sentence-level and word-level semantic features of the questions are extracted.

[0091] Step 740: Generate the user's first sentiment data based on sentence-level semantic features, and generate the user's second sentiment data based on multi-word phrase semantic features.

[0092] The extracted features are categorized to generate the user's first sentiment data based on sentence-level semantic features and the user's second sentiment data based on multi-word phrase semantic features. Specifically, sentence-level semantic features can be input into the classifier of the sentence-level semantic feature extraction model to generate the user's first sentiment data; word-level semantic features can be input into the classifier of the word-level semantic feature extraction model to generate the user's second sentiment data.

[0093] Step 760: Using preset filtering rules, determine the user's emotion recognition result from the first emotion data and the second emotion data.

[0094] The preset filtering rules can be pre-set according to actual needs. By analyzing and comparing the first and second emotional data, the user's emotional recognition result can be determined from the first and second emotional data.

[0095] In this embodiment, the questions entered by the user on the question-and-answer interface are fed into a sentence-level semantic feature extraction model and a word-level semantic feature extraction model for feature extraction, generating sentence-level and word-level semantic features for the questions. First sentiment data of the user is generated based on the sentence-level semantic features, and second sentiment data is generated based on multi-word phrase semantic features. A preset filtering rule is used to determine the user's sentiment recognition result from the first and second sentiment data. By using different models to generate the user's sentiment data, the advantages of different models can be combined to filter the final user's sentiment recognition result from the first and second sentiment data, making the obtained user sentiment recognition result more accurate.

[0096] In one embodiment, such as Figure 8 As shown, it illustrates a flowchart of a customer service Q&A method provided in an embodiment of this application, specifically involving a possible process for generating user emotion recognition results. This method may include the following steps:

[0097] Step 820: Determine whether the first sentiment data belongs to the positive type and whether the second sentiment data belongs to the negative type.

[0098] Experiments have shown that SWEM's classification accuracy is comparable to or slightly higher than that of the CNN model. Therefore, the classification results of the SWEM model are prioritized. Based on this, after obtaining the first and second emotion data, the user's current emotion recognition result can be determined by judging the types of the first and second emotion data. Specifically, it is determined whether the first emotion data belongs to a positive type and whether the second emotion data belongs to a negative type. Positive types include, but are not limited to, calm, happy, and excited types, while negative types include, but are not limited to, angry, nervous, and afraid types.

[0099] Step 840: If not, then the first emotion data is determined as the user's current emotion recognition result.

[0100] The exceptions include: the first sentiment data is not positive, and the second sentiment data is not negative; the first sentiment data is not positive, and the second sentiment data is negative; the first sentiment data is positive, and the second sentiment data is not negative. In all these cases, the first sentiment data is determined as the user's current sentiment recognition result.

[0101] For example, if the first emotional data is tension and the second emotional data is happiness, then the user's current emotional recognition result is tension; if both the first and second emotional data are tension, then the user's current emotional recognition result is tension; if the first emotional data is happiness and the second emotional data is neutral, then the user's current emotional recognition result is happiness. This example is only used to explain the process of determining the user's current emotional recognition result, and will not be illustrated in detail here.

[0102] Step 860: If the first sentiment data is positive and the second sentiment data is negative, then the second sentiment data is determined as the user's current sentiment recognition result.

[0103] Specifically, the CNN model result is used only when the SWEM classification result is positive and the CNN result is negative. This ensures that the customer's emotions are soothed in as many cases as possible. Specifically, if the first emotion data is positive and the second emotion data is negative, the second emotion data is determined as the user's current emotion identification result. For example, if the first emotion data is neutral and the second emotion data is angry, the user's current emotion identification result is angry.

[0104] In this embodiment, the system determines whether the first emotional data is positive and whether the second emotional data is negative. If neither is true, the first emotional data is identified as the user's current emotional recognition result. If both are positive and negative, the second emotional data is identified as the user's current emotional recognition result. By analyzing and judging both types of emotional data, the resulting current emotional recognition result is more accurate. Furthermore, identifying the second emotional data as the user's current emotional recognition result when both are positive and negative can soothe the customer's emotions in as many situations as possible, further improving the intelligence and flexibility of customer service Q&A.

[0105] In one embodiment, if a user is detected expressing prolonged negative emotions such as dissatisfaction, a warning message can be sent to a human customer service representative, prompting intervention from a dedicated hotline to effectively resolve the customer's issue. The conversation in question is recorded as a problem that the intelligent customer service system struggles to handle and submitted to experts for analysis, thereby optimizing the customer service Q&A process. Furthermore, if the user's emotional identification result, based on a preset emotion recognition model, indicates tension or fear, a warning message can be directly sent to a human customer service representative, indicating that the user may be in a dangerous situation. Other warning methods can also be used; this embodiment does not specify a particular method. Moreover, after each use of the intelligent customer service system, the user can be asked to rate it. Low ratings and related conversations are recorded and submitted to experts for analysis to identify any anomalies and upgrade the intelligent customer service system accordingly.

[0106] In this embodiment, the intelligent customer service is assisted by human customer service representatives, thereby improving the user experience. Furthermore, it can issue warnings when a user is in danger, further enhancing the intelligence of the customer service Q&A. Moreover, the intelligent customer service can be continuously upgraded and improved, increasing the accuracy and reliability of customer service Q&A.

[0107] In one embodiment, such as Figure 9 As shown, a flowchart of a customer service Q&A method provided in an embodiment of this application is illustrated. The method may include the following steps:

[0108] Step 901: Record the user's question via voice.

[0109] Obtain the audio data of the question entered by the user.

[0110] Step 902: Perform speech recognition and convert it into text.

[0111] The audio data was subjected to speech recognition to obtain text data.

[0112] Step 903: User sentiment detection.

[0113] The user's question on the question-and-answer interface is input into a preset emotion recognition model for emotion recognition, and the emotion recognition result of the user is generated.

[0114] Step 904: Matching questions and answers.

[0115] Retrieve response data that matches the questions entered by the user on the Q&A interface.

[0116] Step 905: Warning of Abnormal Emotions.

[0117] If the user's emotional recognition result, obtained from the preset emotion recognition model, is that they are nervous or afraid, a warning message can be sent directly to customer service, indicating that the user may be in a dangerous situation.

[0118] Step 906: Determine if there are multiple abnormalities.

[0119] Detect whether a user has been expressing negative emotions such as dissatisfaction for an extended period of time.

[0120] Step 907: If so, then human customer service will intervene.

[0121] Send alerts to customer service representatives so that a dedicated hotline can intervene and effectively resolve customer issues.

[0122] Step 908: Answer processing.

[0123] The response data is generated based on the emotion response data and the question response data.

[0124] Step 909: Digital Human Broadcasts.

[0125] The response data is input into a preset animation model for processing to generate the digital human's expressions and facial animations; the digital human's expressions and facial animations are combined to generate the digital human's animation, which is then displayed on the question-and-answer interface; the digital human's animation is used to display expressions corresponding to the expression data and to provide voice broadcasts of the question response data.

[0126] Step 910: Satisfaction survey.

[0127] Each time a user uses the intelligent customer service, they can be asked to rate the service. Low ratings and related conversations are recorded and sent to experts for analysis to check for any anomalies and upgrade or improve the intelligent customer service.

[0128] In this embodiment, compared with the prior art, it does not rely on traditional dialog boxes to reply to user questions. Instead, a digital human can broadcast the corresponding response data to the user's question via voice and display it through digital human animation. This enables a more intelligent and vivid interaction with the user, achieving a sense of immersion similar to face-to-face conversations, thereby improving the convenience and flexibility of online banking transactions for users. Furthermore, by matching corresponding expressions to the digital human based on the user's input questions, it can better adapt to different user emotions, further enhancing the intelligence and flexibility of customer service Q&A.

[0129] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.

[0130] Based on the same inventive concept, this application also provides a customer service question-and-answer device for implementing the aforementioned customer service question-and-answer method. The solution provided by this device is similar to the solution described in the above method; therefore, the specific limitations of the one or more customer service question-and-answer device embodiments provided below can be found in the limitations of the customer service question-and-answer method described above, and will not be repeated here.

[0131] In one embodiment, such as Figure 10 As shown, a 1000-fold device is provided, comprising: an acquisition module 1002, a first generation module 1004, and a second generation module 1006, wherein:

[0132] The acquisition module 1002 is used to acquire response data that matches the questions entered by the user on the Q&A interface; the response data includes sentiment response data and question response data.

[0133] The first generation module 1004 is used to input the response data into a preset animation model for processing, and generate the digital human's expressions and facial animations.

[0134] The second generation module 1006 is used to synthesize the digital human's facial expressions and facial animations to generate the digital human's animation, and to display the digital human's animation on the question-and-answer interface; the digital human's animation is used to display facial expressions corresponding to the facial expression data, and to provide voice broadcast of the question response data.

[0135] In one embodiment, the facial animation includes multiple sets of mouth image frames; the first generation module 1004 is specifically used to match the expressions of the digital human corresponding to the emotion response data from a preset expression library; the preset expression library stores the correspondence between the emotion response data and the expressions of the digital human in advance; the question response data is input into a preset animation model for processing to generate multiple sets of mouth bone point graphics of the digital human corresponding to the question response data; the mouth bone point graphics include multiple mouth bone points at target positions; multiple mouth image frames of the digital human are generated based on the multiple sets of mouth bone point graphics of the digital human.

[0136] In one embodiment, if the question response data is text data, the first generation module 1004 is further configured to convert the question response data into a format and generate audio data corresponding to the question response data; divide the audio data into multiple audio data units sequentially using a preset audio segmentation rule; determine the target position of multiple mouth bone points of the digital human when the audio data unit is emitted for each audio data unit; and generate the multiple sets of mouth bone point graphics based on the multiple mouth bone points located at the target positions.

[0137] In one embodiment, the second generation module 1006 is specifically used to arrange multiple mouth image frames of the digital human corresponding to the audio data units according to the order of their appearance times in the audio data, thereby generating a mouth animation of the digital human corresponding to the question response data; and to synthesize the digital human's facial expression with the digital human's mouth animation to generate the animation of the digital human.

[0138] In one embodiment, the acquisition module 1002 is specifically used to input the question entered by the user on the question-and-answer interface into a preset emotion recognition model for emotion recognition, generate the user's emotion recognition result, determine the emotion response data corresponding to the user's emotion recognition result; generate question response data based on the emotion response data and the question entered by the user on the question-and-answer interface; and generate the response data based on the emotion response data and the question response data.

[0139] In one embodiment, the acquisition module 1002 is further configured to input the questions entered by the user on the question-and-answer interface into a sentence-level semantic feature extraction model and a word-level semantic feature extraction model for feature extraction, thereby generating sentence-level semantic features and word-level semantic features of the questions; generate the user's first sentiment data based on the sentence-level semantic features, and generate the user's second sentiment data based on the multi-word phrase semantic features; and determine the user's sentiment recognition result from the first sentiment data and the second sentiment data using a preset filtering rule.

[0140] In one embodiment, the acquisition module 1002 is further configured to determine whether the first emotion data belongs to the positive type and whether the second emotion data belongs to the negative type; if not, the first emotion data is determined as the current emotion recognition result of the user.

[0141] In one embodiment, the acquisition module 1002 is further configured to determine the second emotion data as the user's current emotion recognition result if the first emotion data is of a positive type and the second emotion data is of a negative type.

[0142] The modules in the aforementioned customer service Q&A device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in the processor of a computer device in hardware form or independent of it, or stored in the memory of a computer device in software form, so that the processor can call and execute the corresponding operations of each module.

[0143] In one embodiment, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to perform the following steps:

[0144] Acquire response data that matches the questions entered by the user on the Q&A interface; the response data includes emotional response data and question response data; input the response data into a preset animation model for processing to generate the digital human's expressions and facial animations; synthesize the digital human's expressions and facial animations to generate the digital human's animation, and display the digital human's animation on the Q&A interface; the digital human's animation is used to display the expressions corresponding to the expression data, and to provide voice broadcast of the question response data.

[0145] In one embodiment, facial animation includes multiple sets of mouth image frames;

[0146] When a processor executes a computer program, it also performs the following steps:

[0147] The system matches the expressions of the digital human corresponding to the emotion response data from a preset expression library. The preset expression library stores the correspondence between the emotion response data and the expressions of the digital human. The question response data is input into a preset animation model for processing to generate multiple sets of mouth bone point graphics of the digital human corresponding to the question response data. The mouth bone point graphics include multiple mouth bone points at the target positions. Based on the multiple sets of mouth bone point graphics of the digital human, multiple mouth image frames of the digital human are generated.

[0148] In one embodiment, the question response data is text data;

[0149] When a processor executes a computer program, it also performs the following steps:

[0150] The question response data is formatted and converted to generate audio data corresponding to the question response data; the audio data is divided into multiple audio data units sequentially using preset audio segmentation rules; for each audio data unit, the target positions of multiple mouth bone points of the digital human are determined when the audio data unit is emitted; based on the multiple mouth bone points at the target positions, multiple sets of mouth bone point graphics are generated.

[0151] In one embodiment, the processor, when executing a computer program, also performs the following steps:

[0152] Based on the chronological order of the appearance of multiple audio data units in the audio data, multiple mouth image frames of the digital human corresponding to the audio data units are arranged to generate the mouth animation of the digital human corresponding to the question response data; the digital human's facial expressions and the digital human's mouth animation are then combined to generate the digital human's animation.

[0153] In one embodiment, the processor, when executing a computer program, also performs the following steps:

[0154] The process involves inputting the user's question into a pre-defined emotion recognition model for emotion recognition, generating the user's emotion recognition result, and determining the corresponding emotion response data. Based on the emotion response data and the user's question on the Q&A interface, question response data is generated. Finally, based on the emotion response data and the question response data, response data is generated.

[0155] In one embodiment, the processor, when executing a computer program, also performs the following steps:

[0156] The questions entered by users on the question-and-answer interface are fed into sentence-level semantic feature extraction models and word-level semantic feature extraction models for feature extraction, generating sentence-level and word-level semantic features of the questions; the user's first sentiment data is generated based on the sentence-level semantic features, and the user's second sentiment data is generated based on the multi-word phrase semantic features; the user's sentiment recognition result is determined from the first and second sentiment data using preset filtering rules.

[0157] In one embodiment, the processor, when executing a computer program, also performs the following steps:

[0158] Determine whether the first sentiment data belongs to the positive type and whether the second sentiment data belongs to the negative type; if not, then determine the first sentiment data as the user's current sentiment recognition result.

[0159] In one embodiment, the processor, when executing a computer program, also performs the following steps:

[0160] If the first sentiment data is positive and the second sentiment data is negative, then the second sentiment data is determined as the user's current sentiment recognition result.

[0161] The computer device provided in this application embodiment has a similar implementation principle and technical effect to the above method embodiment, and will not be described again here.

[0162] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, the computer program performing the following steps when executed by a processor:

[0163] Acquire response data that matches the questions entered by the user on the Q&A interface; the response data includes emotional response data and question response data; input the response data into a preset animation model for processing to generate the digital human's expressions and facial animations; synthesize the digital human's expressions and facial animations to generate the digital human's animation, and display the digital human's animation on the Q&A interface; the digital human's animation is used to display the expressions corresponding to the expression data, and to provide voice broadcast of the question response data.

[0164] In one embodiment, the question response data is text data;

[0165] When a computer program is executed by a processor, it also performs the following steps:

[0166] The question response data is formatted and converted to generate audio data corresponding to the question response data; the audio data is divided into multiple audio data units sequentially using preset audio segmentation rules; for each audio data unit, the target positions of multiple mouth bone points of the digital human are determined when the audio data unit is emitted; based on the multiple mouth bone points at the target positions, multiple sets of mouth bone point graphics are generated.

[0167] In one embodiment, when the computer program is executed by a processor, it also performs the following steps:

[0168] Based on the chronological order of the appearance of multiple audio data units in the audio data, multiple mouth image frames of the digital human corresponding to the audio data units are arranged to generate the mouth animation of the digital human corresponding to the question response data; the digital human's facial expressions and the digital human's mouth animation are then combined to generate the digital human's animation.

[0169] In one embodiment, when the computer program is executed by a processor, it also performs the following steps:

[0170] The process involves inputting the user's question into a pre-defined emotion recognition model for emotion recognition, generating the user's emotion recognition result, and determining the corresponding emotion response data. Based on the emotion response data and the user's question on the Q&A interface, question response data is generated. Finally, based on the emotion response data and the question response data, response data is generated.

[0171] In one embodiment, when the computer program is executed by a processor, it also performs the following steps:

[0172] The questions entered by users on the question-and-answer interface are fed into sentence-level semantic feature extraction models and word-level semantic feature extraction models for feature extraction, generating sentence-level and word-level semantic features of the questions; the user's first sentiment data is generated based on the sentence-level semantic features, and the user's second sentiment data is generated based on the multi-word phrase semantic features; the user's sentiment recognition result is determined from the first and second sentiment data using preset filtering rules.

[0173] In one embodiment, when the computer program is executed by a processor, it also performs the following steps:

[0174] Determine whether the first sentiment data belongs to the positive type and whether the second sentiment data belongs to the negative type; if not, then determine the first sentiment data as the user's current sentiment recognition result.

[0175] In one embodiment, when the computer program is executed by a processor, it also performs the following steps:

[0176] If the first sentiment data is positive and the second sentiment data is negative, then the second sentiment data is determined as the user's current sentiment recognition result.

[0177] The computer-readable storage medium provided in this embodiment is similar in principle and technical effect to the method embodiment described above, and will not be repeated here.

[0178] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, performs the following steps:

[0179] Acquire response data that matches the questions entered by the user on the Q&A interface; the response data includes emotional response data and question response data; input the response data into a preset animation model for processing to generate the digital human's expressions and facial animations; synthesize the digital human's expressions and facial animations to generate the digital human's animation, and display the digital human's animation on the Q&A interface; the digital human's animation is used to display the expressions corresponding to the expression data, and to provide voice broadcast of the question response data.

[0180] In one embodiment, the question response data is text data;

[0181] When a computer program is executed by a processor, it also performs the following steps:

[0182] The question response data is formatted and converted to generate audio data corresponding to the question response data; the audio data is divided into multiple audio data units sequentially using preset audio segmentation rules; for each audio data unit, the target positions of multiple mouth bone points of the digital human are determined when the audio data unit is emitted; based on the multiple mouth bone points at the target positions, multiple sets of mouth bone point graphics are generated.

[0183] In one embodiment, when the computer program is executed by a processor, it also performs the following steps:

[0184] Based on the chronological order of the appearance of multiple audio data units in the audio data, multiple mouth image frames of the digital human corresponding to the audio data units are arranged to generate the mouth animation of the digital human corresponding to the question response data; the digital human's facial expressions and the digital human's mouth animation are then combined to generate the digital human's animation.

[0185] In one embodiment, when the computer program is executed by a processor, it also performs the following steps:

[0186] The process involves inputting the user's question into a pre-defined emotion recognition model for emotion recognition, generating the user's emotion recognition result, and determining the corresponding emotion response data. Based on the emotion response data and the user's question on the Q&A interface, question response data is generated. Finally, based on the emotion response data and the question response data, response data is generated.

[0187] In one embodiment, when the computer program is executed by a processor, it also performs the following steps:

[0188] The questions entered by users on the question-and-answer interface are fed into sentence-level semantic feature extraction models and word-level semantic feature extraction models for feature extraction, generating sentence-level and word-level semantic features of the questions; the user's first sentiment data is generated based on the sentence-level semantic features, and the user's second sentiment data is generated based on the multi-word phrase semantic features; the user's sentiment recognition result is determined from the first and second sentiment data using preset filtering rules.

[0189] In one embodiment, when the computer program is executed by a processor, it also performs the following steps:

[0190] Determine whether the first sentiment data belongs to the positive type and whether the second sentiment data belongs to the negative type; if not, then determine the first sentiment data as the user's current sentiment recognition result.

[0191] In one embodiment, when the computer program is executed by a processor, it also performs the following steps:

[0192] If the first sentiment data is positive and the second sentiment data is negative, then the second sentiment data is determined as the user's current sentiment recognition result.

[0193] The computer program product provided in this embodiment has a similar implementation principle and technical effect to the method embodiment described above, and will not be repeated here.

[0194] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties.

[0195] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.

[0196] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0197] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.

Claims

1. A method of customer care question answering, the method comprising: The method comprises: obtaining reply data matched with a question input by a user on a question-answering interface; the reply data comprises emotional reply data and question reply data; inputting the reply data into a preset animation model for processing to generate an expression and a facial animation of a digital person; synthesizing the expression of the digital person and the facial animation of the digital person to generate an animation of the digital person, and displaying the animation of the digital person on the question-answering interface; the animation of the digital person is used to display an expression corresponding to the expression data and to voice broadcast the question reply data; the obtaining of the reply data matched with the question input by the user on the question-answering interface comprises: inputting the question input by the user on the question-answering interface into a sentence-level semantic feature extraction model and a word-level semantic feature extraction model respectively for feature extraction; the sentence-level semantic feature extraction model extracts sentence-level semantic features by averaging word vectors element by element using average pooling; the word-level semantic feature extraction model extracts word-level semantic features using an improved CNN model; generating first emotional data of the user based on the sentence-level semantic features and generating second emotional data of the user based on multi-element word group semantic features; determining whether the first emotional data belongs to a positive type and whether the second emotional data belongs to a negative type; if not, determining the first emotional data as a current emotional recognition result of the user, and determining emotional reply data corresponding to the emotional recognition result of the user; generating question reply data according to the emotional reply data and the question input by the user on the question-answering interface; generating the reply data based on the emotional reply data and the question reply data.

2. The method of claim 1, wherein, The facial animation comprises a plurality of mouth image frames; the inputting of the reply data into the preset animation model for processing to generate the expression and the facial animation of the digital person comprises: matching an expression of the digital person corresponding to the emotional reply data from a preset expression library; the preset expression library pre-stores a corresponding relationship between emotional reply data and expressions of the digital person; inputting the question reply data into the preset animation model for processing to generate a plurality of mouth skeleton point patterns of the digital person corresponding to the question reply data; the mouth skeleton point patterns comprise a plurality of mouth skeleton points at target positions; generating a plurality of mouth image frames of the digital person according to the plurality of mouth skeleton point patterns of the digital person.

3. The method of claim 2, wherein, If the question reply data is text data, the inputting of the question reply data into the preset animation model for processing to generate a plurality of mouth skeleton point patterns of the digital person corresponding to the question reply data comprises: performing format conversion on the question reply data to generate audio data corresponding to the question reply data; dividing the audio data into a plurality of audio data units in sequence using a preset audio segmentation rule; determining target positions of a plurality of mouth skeleton points of the digital person when the audio data units are emitted for each audio data unit; Generate the plurality of sets of mouth skeleton point graphics based on a plurality of mouth skeleton points in the target position.

4. The method of claim 3, wherein, The synthesizing of the expression of the digital person and the facial animation of the digital person generates the animation of the digital person, including: According to the order of the occurrence time points of the plurality of audio data units in the audio data, arrange a plurality of mouth image frames of the digital person corresponding to the audio data units to generate a mouth animation of the digital person corresponding to the question reply data; Synthesize the expression of the digital person and the mouth animation of the digital person to generate the animation of the digital person.

5. The method of claim 1, wherein, The method further includes: If the first emotional data belongs to a positive type and the second emotional data belongs to a negative type, the second emotional data is determined as the current emotional recognition result of the user.

6. A customer care answering apparatus characterized by comprising: The device includes: An acquisition module is configured to acquire reply data matched with a question input by a user on a question and answer interface; the reply data includes emotional reply data and question reply data; A first generation module is configured to input the reply data into a preset animation model for processing to generate an expression and a facial animation of a digital person; A second generation module is configured to synthesize the expression of the digital person and the facial animation of the digital person to generate an animation of the digital person, and display the animation of the digital person on the question and answer interface; the animation of the digital person is used to display an expression corresponding to the expression data and voice broadcast of the question reply data; The acquisition module is further configured to input a question input by a user on a question and answer interface into a sentence-level semantic feature extraction model and a word-level semantic feature extraction model respectively for feature extraction; the sentence-level semantic feature extraction model extracts sentence-level semantic features by averaging word vectors element by element using average pooling; the word-level semantic feature extraction model extracts word-level semantic features using an improved CNN model; first emotional data of the user is generated based on the sentence-level semantic features, and second emotional data of the user is generated based on multi-element word group semantic features; it is determined whether the first emotional data belongs to a positive type and the second emotional data belongs to a negative type; if not, the first emotional data is determined as the current emotional recognition result of the user, emotional reply data corresponding to the emotional recognition result of the user is determined; question reply data is generated according to the emotional reply data and the question input by the user on the question and answer interface; the reply data is generated based on the emotional reply data and the question reply data. 7.A computer device, comprising a memory and a processor, wherein the memory stores a computer program, and the computer device is configured to perform the method according to any one of claims 1-6 when the computer program is executed by the processor. The processor executes the computer program to implement the steps of the method of any one of claims 1 to 5.

8. A computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to implement the steps of the method of any one of claims 1 to 5.

9. A computer program product comprising a computer program, characterized in that, The computer program is executed by the processor to implement the steps of the method of any one of claims 1 to 5. The computer program is executed by the processor to implement the steps of the method of any one of claims 1 to 5.

Citation Information

Patent Citations

  • Video processing method, device and system, terminal equipment and storage medium

    CN110688911A

  • Semantic emotion recognition method, device and equipment and storage medium

    CN112613324A

  • Information output method and device, computer equipment and storage medium

    CN116303967A

  • Question and answer processing method and device, computer equipment and storage medium

    CN118568239A