User data processing method and device based on AI customer service
By acquiring multimodal features and emotion recognition results from AI customer service user dialogue data, the problem of insufficient intelligent interaction by AI customer service after the current round of dialogue with the user is solved, realizing efficient utilization of dialogue data and intelligent prompts.
Patent Information
- Application Number
- CN202511045193.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-29
- Publication Date
- 2025-11-11
AI Technical Summary
Existing AI customer service does not perform sufficient intelligent interaction processing after the user's current conversation, resulting in low utilization of conversation data.
By acquiring multimodal features and emotion recognition results from user dialogue data, and combining them with a pre-set modal feature library and a pre-trained emotion recognition model, user prompts are generated and sent.
It improves the utilization rate of dialogue data and enables intelligent analysis and timely generation of prompts in environments with high user privacy.
Smart Images

Figure CN120930183A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing technology, and in particular to a method and apparatus for processing user data based on AI customer service. Background Technology
[0002] With the development of artificial intelligence (AI) technology, AI customer service has been widely applied in various fields. Current AI customer service, after acquiring the dialogue data of visiting users, combines various AI models in the corresponding backend server to understand the user's dialogue data and then provides corresponding responses based on a knowledge base. However, current AI customer service focuses more on the period between the current conversation and the user's next conversation, without performing much intelligent interactive processing, only ensuring the secure storage of data. This results in low data utilization of the current conversation. Summary of the Invention
[0003] This invention provides a user data processing method and apparatus based on AI customer service, aiming to solve the problem that in the prior art, AI customer service focuses more on the period after the current conversation with the user and before the user starts the next conversation, without performing more intelligent interactive processing, and only performing data security storage, resulting in low data utilization of the current conversation data.
[0004] In a first aspect, embodiments of the present invention provide a user data processing method based on AI customer service, comprising:
[0005] In response to a dialogue access command, acquire the AI customer service agent corresponding to the dialogue access command and the user authorization record command, as well as the current user dialogue data;
[0006] Obtain the current data type of the current user dialogue data, and identify the current user dialogue data based on the recognition model corresponding to the current data type to obtain the current recognition result;
[0007] The text features, speech features, and behavioral features in the current user dialogue data are obtained to form the current multimodal features corresponding to the current recognition result;
[0008] The current feature type corresponding to the current multimodal feature is obtained based on a preset modal feature library;
[0009] The current emotion recognition result is obtained by performing emotion recognition based on the pre-trained emotion recognition model.
[0010] If it is determined that the current feature type belongs to a preset feature type set and the current emotion recognition result belongs to a preset emotion recognition result set, then the corresponding user prompt information is obtained and sent to the user terminal corresponding to the dialogue access command.
[0011] Secondly, embodiments of the present invention also provide a user data processing device based on AI customer service, comprising:
[0012] The dialogue data acquisition unit is used to respond to the dialogue access command and acquire the AI customer service agent and the current user dialogue data corresponding to the dialogue access command and the user authorization record command.
[0013] The recognition result acquisition unit is used to acquire the current data type of the current user dialogue data, and to recognize the current user dialogue data based on the recognition model corresponding to the current data type to obtain the current recognition result;
[0014] A multimodal feature acquisition unit is used to acquire text features, speech features, and behavioral features in the current user dialogue data to form current multimodal features corresponding to the current recognition result;
[0015] The feature type acquisition unit is used to acquire the current feature type corresponding to the current multimodal feature based on a preset modal feature library;
[0016] An emotion recognition result acquisition unit is used to perform emotion recognition on the current recognition result based on a pre-trained emotion recognition model to obtain the current emotion recognition result;
[0017] The user prompt information sending unit is used to obtain the corresponding user prompt information and send it to the user terminal corresponding to the dialogue access instruction if it is determined that the current feature type belongs to a preset feature type set and the current emotion recognition result belongs to a preset emotion recognition result set.
[0018] Thirdly, embodiments of the present invention also provide a computer device, which includes a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the method described in the first aspect above.
[0019] Fourthly, embodiments of the present invention also provide a computer-readable storage medium storing a computer program, the computer program including program instructions that, when executed by a processor, can implement the method described in the first aspect above.
[0020] This invention provides a user data processing method and apparatus based on AI customer service. The method includes: responding to a dialogue access command, acquiring current user dialogue data corresponding to the dialogue access command and the user authorization record command; acquiring the current data type of the current user dialogue data, and recognizing the current user dialogue data based on the recognition model corresponding to the current data type to obtain a current recognition result; acquiring text features, voice features, and behavioral features in the current user dialogue data to form current multimodal features corresponding to the current recognition result; acquiring the current feature type corresponding to the current multimodal features based on a preset modality feature library; performing emotion recognition on the current recognition result based on a pre-trained emotion recognition model to obtain a current emotion recognition result; if it is determined that the current feature type belongs to a preset feature type set and the current emotion recognition result belongs to a preset emotion recognition result set, then acquiring the corresponding user prompt information and sending it to the user terminal corresponding to the dialogue access command. This invention can intelligently acquire communication content features including multimodal features and emotion recognition results when AI customer service engages in highly private user dialogues with users, perform intelligent analysis, and generate prompt information in a timely manner and automatically send it to the user terminal when relevant conditions are met. Attached Figure Description
[0021] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0022] Figure 1 This is a schematic diagram illustrating an application scenario of the AI-based customer service user data processing method provided in an embodiment of the present invention.
[0023] Figure 2 A flowchart illustrating the user data processing method based on AI customer service provided in an embodiment of the present invention;
[0024] Figure 3 A schematic diagram of a sub-process of the user data processing method based on AI customer service provided in an embodiment of the present invention;
[0025] Figure 4 This is another sub-process diagram of the user data processing method based on AI customer service provided in an embodiment of the present invention;
[0026] Figure 5 This is another sub-process diagram of the user data processing method based on AI customer service provided in an embodiment of the present invention;
[0027] Figure 6 A schematic block diagram of a user data processing device based on AI customer service provided in an embodiment of the present invention;
[0028] Figure 7 A schematic block diagram of a computer device provided for an embodiment of the present invention. Detailed Implementation
[0029] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0030] It should be understood that, when used in this specification and the appended claims, the terms "comprising" and "including" indicate the presence of the described features, integrals, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or collections thereof.
[0031] It should also be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the invention. As used in this specification and the appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms unless the context clearly indicates otherwise.
[0032] It should also be further understood that the term "and / or" as used in this specification and the appended claims refers to any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.
[0033] Please also refer to Figure 1 and Figure 2 ,in Figure 1 This is a schematic diagram illustrating a scenario of the user data processing method based on AI customer service according to an embodiment of the present invention. Figure 2 This is a flowchart illustrating the user data processing method based on AI customer service provided in an embodiment of the present invention. Figure 1 As shown, the user data processing method based on AI customer service provided in this embodiment of the invention is applied to server 10, and server 10 is connected to user terminal 20.
[0034] like Figure 2 As shown, the method includes the following steps S110-S160.
[0035] S110. In response to the dialogue access command, obtain the AI customer service agent corresponding to the dialogue access command and the user authorization record command, as well as the current user dialogue data.
[0036] In this embodiment, the technical solution is described with the server as the executing entity. An intelligent communication platform is deployed on the server, integrating AI customer service (i.e., artificial intelligence customer service). When a user logs into the intelligent communication platform using a smart terminal (such as a smartphone, tablet, etc.), a chat box is displayed on the user interface of the user terminal. This chat box corresponds to an AI customer service agent, and the user can send any form of dialogue data to the AI customer service agent in this chat box, such as text, images, audio, and video. Of course, for high security of the dialogue data in the above communication process, it is necessary to obtain whether the user has authorized the uploading of dialogue data to the server for secure interaction before the start of the current dialogue. Only after the server obtains the dialogue access instruction and user authorization record instruction from the user terminal can it obtain the dialogue data entered by the user in the chat box and upload it to the server for subsequent intelligent dialogue. Furthermore, to more accurately open the AI customer service agent for chatting with the user, the server can either re-open the AI customer service agent based on the one the user has previously interacted with, or the server can first recommend 3-5 AI agents related to the user's user tags to the smart terminal, and then open the corresponding chat box of the AI customer service agent after the user selects one, and conduct subsequent intelligent interaction with the user. Different AI agents on the server specialize in different areas of intelligent interaction; for example, different AI agents may correspond to fields such as clinical psychology, counseling psychology, health psychology, sports psychology, and forensic psychology.
[0037] Since the above interaction process involves a private dialogue between the user and the AI customer service, the current user dialogue data can be encrypted and sent to the server. After the encrypted dialogue data is decrypted and processed on the server, it can be re-encrypted and stored on the server or periodically deleted to ensure the security of the dialogue data. Furthermore, users can conduct the dialogue in highly private locations (such as their residences) without face-to-face human interaction, facilitating smoother communication. To further intelligently confirm that the user is in a highly private location (such as their residence), the user's location data can be obtained before step S110, provided the user authorizes it. If the user's location data is determined to belong to a preset residential area location dataset, the user will be prompted that they are in a safe and private communication location.
[0038] S120. Obtain the current data type of the current user dialogue data, and identify the current user dialogue data based on the recognition model corresponding to the current data type to obtain the current recognition result.
[0039] In this embodiment, when the user has completed at least one round of conversation (such as including at least one sentence sent by the user and the corresponding reply data sent by the AI customer service) and sends it to the server as the current user conversation data, the server will promptly determine the data type of the conversation data to obtain the current data type. After obtaining the current data type, the corresponding recognition model is called to recognize the current user conversation data to obtain the current recognition result.
[0040] In one embodiment, as Figure 3 shown, step S120 includes:
[0041] S121. If it is determined that the current data type is text type, the corresponding language fine-tuning model is obtained and the current user conversation data is recognized to obtain the current recognition result;
[0042] S122. If it is determined that the current data type is voice type, the corresponding voice recognition model is obtained and the current user conversation data is recognized to obtain the current recognition result.
[0043] In this embodiment, multiple types of recognition models are deployed in the server. For example, a language fine-tuning model for text fine-tuning and correction of text-type conversation content, and a voice recognition model for voice recognition of voice-type conversations. If it is determined that the user interacts with the AI customer service in text form, it is necessary to fully consider the situation where the user may enter some characters incorrectly. To facilitate more accurate text recognition, the current user conversation data can be recognized based on the language fine-tuning model to obtain the current recognition result. For example, the language fine-tuning model is a BERT model (a pre-trained language model based on bidirectional encoders and deep learning) and other models. A two-tower structure of "error detection - correction" is constructed through the BERT model. The detection model is used to determine the error position in the input text (such as when the current user conversation data is: [CLS] I'm a bit out of cotton recently [SEP], the detection model outputs the error position as 7); then the error position is input into the correction model to generate candidate words, and the correct vocabulary is predicted through MLM (masked language model) (such as replacing "cotton" with "sleep"). After the language fine-tuning model performs intelligent correction of the incorrect characters in the text-type current user conversation data, the current recognition result can be obtained. Of course, if there are no incorrect characters in the current user conversation data, the current user conversation data will not be corrected after being input into the language fine-tuning model, but will be directly input as the original text as the current recognition result.
[0044] If it is determined that the user is interacting with the AI customer service via voice, then the current data type is determined to be voice. At this point, a speech recognition model is used to recognize the current user dialogue data to obtain the current recognition result. Specifically, the speech recognition model adopts the Wav2Vec model, which extracts features from speech data through self-supervised learning and can learn the acoustic and semantic information of speech. The Wav2Vec model includes a feature extractor (which can be composed of multiple layers of 1D convolutional neural networks, directly processing the raw waveform in the current user dialogue data and converting the time-domain signal into a feature sequence), a context network (which can be constructed using a bidirectional long short-term memory network, used to perform temporal modeling of the feature extractor's output, capturing long-distance dependencies to output latent representations, and including rich acoustic and semantic information in the latent representations), and a contrastive learning objective (such as multiple network layers including fully connected layers, convolutional layers, etc.) processed by multiple network layers. By recognizing the current user dialogue data through the above speech recognition model, the current recognition result can be obtained. As can be seen, the server can accurately identify various types of user dialogue data and obtain the current identification results for subsequent intelligent interaction.
[0045] In one embodiment, step S120 further includes:
[0046] If it is determined that the current data type is a mixed type, and the mixed type includes text type and speech type, then the language fine-tuning model and the speech recognition model are obtained;
[0047] The language fine-tuning model is used to identify the text content in the current user dialogue data to obtain a first current identification result;
[0048] The speech content in the current user dialogue data is identified by the speech recognition model to obtain a second current recognition result;
[0049] The first current recognition result and the second current recognition result are concatenated according to the dialogue sequence to obtain the current recognition result.
[0050] In this embodiment, if the current data type is determined to be a mixed type, and the mixed type includes both text and voice types, it means that the user and AI customer service have both text and voice dialogue data in a complete round of conversation. Since the sending time of each sentence in the intelligent interaction between the user and AI customer service is known, the above example can be used as a reference. Specifically, the language fine-tuning model is used to identify the text content in the current user dialogue data to obtain a first current recognition result; and the voice recognition model is used to identify the voice content in the current user dialogue data to obtain a second current recognition result.
[0051] For example, the current user dialogue data includes N1 text-type dialogues (N1 being a positive integer) and N2 speech-type dialogues (N2 being a positive integer). The server knows the sending time of each of the N1 text-type dialogues and the sending time of each of the N2 speech-type dialogues. After recognizing each of the N1 text-type dialogues using a language fine-tuning model to obtain a first recognition sub-result (N1 first recognition sub-results constitute a first current recognition result), and after recognizing each of the N2 speech-type dialogues using a speech recognition model to obtain a second recognition sub-result (N2 second recognition sub-results constitute a second current recognition result), the N1 first recognition sub-results in the first current recognition result and the N2 second recognition sub-results in the second current recognition result are concatenated according to the dialogue sequence to obtain the current recognition result. The obtained current recognition result is text data, facilitating subsequent semantic understanding by the server. Of course, if the current data type is determined to be a mixed type, and the mixed type includes more types such as text, voice, video, and image, then the language fine-tuning model recognizes the text content in the current user dialogue data, and the voice recognition model recognizes the voice content in the current user dialogue data. In addition, the image recognition model and the voice recognition model are also needed to recognize the dialogue data of video and / or image types to obtain the recognition results (mainly the recognition results corresponding to the text content and voice content in the video and / or image).
[0052] S130. Obtain the text features, speech features, and behavioral features from the current user dialogue data, and form the current multimodal features corresponding to the current recognition result.
[0053] In this embodiment, the server can also extract text features (such as sentence structure, emotional tendency, etc.), speech features (such as speech rate, duration of pauses, number of pauses in a sentence, number of pauses in the whole conversation, etc.) and behavioral features (such as the number of interruptions in dialogue input, etc.) from the original dialogue data of the current user dialogue data, thereby forming the current multimodal features corresponding to the current user dialogue data.
[0054] In one embodiment, such as Figure 4 As shown, step S130 includes:
[0055] S131. Obtain the sentence structure statistics in the current user dialogue data as the text feature;
[0056] S132. Obtain the voice pause count statistics in the current user dialogue data as the voice feature;
[0057] S133. Obtain the input interruption count statistics from the current user dialogue data as the behavioral feature.
[0058] In this embodiment, the text features are obtained by using a Natural Language Processing (NLP) model to acquire statistical information on sentence structure in the current user dialogue data (e.g., the sentence structures mainly include interrogative, declarative, imperative, and exclamatory sentences, and the total number of each sentence structure and its proportion in the complete dialogue round). Specifically, the speech features are obtained by using the aforementioned Wav2Vec model combined with Mel-frequency cepstral coefficients (MFCCs). The speech features are obtained by using the Wav2Vec model for text recognition of the current user dialogue data, and obtaining MFCCs from the corresponding audio signal to further analyze the audio features. The speech pause statistics are obtained by statistically analyzing the interruption times and cumulative terminal duration in the audio features. Furthermore, the input interruption statistics are obtained from the embedded components in the chat box when the user communicates with the AI customer service. After acquiring these three features, subsequent intelligent dialogue data analysis can be performed.
[0059] S140. Obtain the current feature type corresponding to the current multimodal feature based on the preset modal feature library.
[0060] In this embodiment, after acquiring the current multimodal features of the current user dialogue data, the current feature type corresponding to the current multimodal features can be further determined by combining the modality feature library, so as to analyze the psychological state, emotional state, etc. of the dialogue between the user and the AI customer service based on the current feature type.
[0061] In one embodiment, such as Figure 5 As shown, step S140 includes:
[0062] S141. Obtain the current feature vector composed of the text features, the speech features, and the behavior features;
[0063] S142. Among the feature vectors corresponding to each multimodal feature in the modal feature library, determine the feature vector with the maximum vector similarity to the current feature vector and use it as the target feature vector;
[0064] S143. Obtain the feature type corresponding to the target feature vector and use it as the current feature type corresponding to the current multimodal feature.
[0065] The feature type is either a pause / hesitation type or a normal dialogue type.
[0066] In this embodiment, after obtaining the text features as sentence structure statistics, the speech features as speech pause count statistics, and the behavior features as input interruption count statistics, the following steps are taken: the first proportion of interrogative sentences in the sentence structure statistics within the complete dialogue round; the second proportion between the total number of speech pauses in the complete dialogue round and the total number of sentences in the complete dialogue round in the speech pause count statistics; and the third proportion between the total number of input interruptions in the complete dialogue round and the total number of sentences in the complete dialogue round in the input interruption count statistics. Finally, the first proportion, the second proportion, and the third proportion are used to form the current feature vector.
[0067] Subsequently, since feature vectors have been pre-configured for each multimodal feature in the modal feature library, the feature vector similarity between the feature vectors corresponding to each multimodal feature in the modal feature library and the current feature vector can be calculated (e.g., the cosine similarity between the feature vector and the current feature vector is used as the feature vector similarity). The feature vector with the maximum vector similarity to the current feature vector is then selected as the target feature vector.
[0068] Finally, since feature types are pre-configured for each multimodal feature in the modal feature library, the feature type corresponding to the target feature vector is used as the current feature type corresponding to the current multimodal feature. For example, the feature type could be either a pause / hesitation type or a normal dialogue type. Of course, in specific implementations, the feature type is not limited to the two types in the above example; it can also include other types such as request for assistance. The types of feature special cases can be expanded according to the actual needs of the user.
[0069] S150. Based on the pre-trained emotion recognition model, perform emotion recognition on the current recognition result to obtain the current emotion recognition result.
[0070] In this embodiment, after analyzing and obtaining the current feature type in the current user dialogue data, it can be further combined with a pre-trained emotion recognition model on the server, such as the BiLSTM-Attention model, which is a bidirectional long short-term memory network combined with an attention mechanism model. This model mainly includes an input layer, an embedding layer, a bidirectional long short-term memory layer (BiLSTM layer), an attention layer, a pooling layer, a fully connected layer, and an output layer connected in sequence. Through this emotion recognition model, emotion recognition can be performed on the current recognition result in text form to obtain the current emotion recognition result. Furthermore, the obtained current emotion recognition result can be combined with the obtained current feature type to form multi-dimensional user data to be analyzed.
[0071] In one embodiment, step S150 includes:
[0072] Obtain the current semantic features corresponding to the current recognition result, and input the current semantic features into the emotion recognition model to obtain the current emotion recognition result.
[0073] In this embodiment, if the emotion recognition model adopts the BiLSTM-Attention model as shown in the example above, the current recognition result needs to be converted into the corresponding current semantic features based on the input layer and the embedding layer, and then processed sequentially through the bidirectional long short-term memory layer, attention layer, pooling layer, fully connected layer and output layer to obtain the current emotion recognition result.
[0074] S160. If it is determined that the current feature type belongs to a preset feature type set and the current emotion recognition result belongs to a preset emotion recognition result set, then obtain the corresponding user prompt information and send it to the user terminal corresponding to the dialogue access instruction.
[0075] In this embodiment, if it is determined that the current feature type belongs to a preset feature type set (e.g., the preset feature type set includes at least the pause and hesitation type) and the current emotion recognition result belongs to a preset emotion recognition result set (e.g., the preset emotion recognition result set includes at least the emotion recognition results such as dissatisfaction, anxiety, worry, frustration, and disappointment), then it can be determined that the AI customer service needs to promptly issue a user prompt message to the user after the end of this round of dialogue, such as "After this dialogue ends, we feel that your recent state is not very good. Do you need a professional to communicate with you via voice?", and send the user prompt message to the user terminal corresponding to the dialogue access command. Moreover, in addition to text content, the user prompt message may also include two virtual buttons, namely "Agree" and "Reject". After receiving the above user prompt message, the user terminal can choose to click one of the two virtual buttons or ignore the message without making any click.
[0076] In one embodiment, the method further includes the following after step S160:
[0077] The system acquires the dialogue access instruction and the target storage space corresponding to the user terminal, and stores the current user dialogue data in the target storage space; wherein the target storage space is an encrypted storage space.
[0078] If it is determined that the feedback interval of the user feedback information in response to the user prompt exceeds the preset feedback duration threshold, a manual follow-up request will be sent to the corresponding receiving terminal to prompt the receiving terminal to conduct a manual follow-up communication with the user terminal.
[0079] In this embodiment, after the user completes the current round of communication with the AI customer service and completes the above-mentioned processes such as obtaining dialogue feature types and emotion recognition, in order to provide security for the dialogue data, the target storage space corresponding to the dialogue access instruction (mainly the user's unique identifier included in it) and the user terminal can be obtained in the server. Then, the target storage space is used as an encrypted storage space to store all the dialogue data between the user and the AI customer service, achieving a technical effect similar to "AI tree hole" and ensuring the security of the transmission and storage of dialogue data.
[0080] If the sending time of the user prompt and the sending time of the user feedback have been obtained, the interval between the sending time of the feedback and the sending time of the prompt is used as the information feedback interval. If the above information feedback interval exceeds a preset feedback interval threshold, it indicates that the user has not viewed the user prompt and provided corresponding feedback in a timely manner (such as agreeing to receive subsequent expert communication or refusing to receive subsequent expert communication). In this case, a manual follow-up request can be sent to the corresponding receiving terminal to prompt the receiving terminal to conduct a manual follow-up communication with the user terminal, so that the user can still receive the manual follow-up from relevant experts in a timely manner.
[0081] Furthermore, when performing emotion recognition on the current recognition result based on the pre-trained emotion recognition model in the previous steps, it can also combine the historical emotion recognition results corresponding to the user's recent (e.g., the last three days, the last week, the last half month, the last month, the last quarter, etc.) historical dialogue data in the target storage space to draw an emotion state change curve. The historical dialogue data stored in the target storage space can also serve as historical context data for the user's dialogue with AI customer service, enabling AI customer service to provide users with more accurate response data.
[0082] As can be seen, the implementation of this method can intelligently acquire communication content features, including multimodal features and emotion recognition results, when AI customer service engages in highly private user conversations with users, and perform intelligent analysis. When relevant conditions are met, it can promptly generate prompt information and automatically send it to the user's terminal.
[0083] Figure 6 This is a schematic block diagram of a user data processing device based on AI customer service provided in an embodiment of the present invention. Figure 6 As shown, corresponding to the above-described AI-based customer service user data processing method, the present invention also provides an AI-based customer service user data processing apparatus 100. This AI-based customer service user data processing apparatus 100 includes a unit for executing the above-described AI-based customer service user data processing method. Please refer to... Figure 6The AI-based customer service user data processing device 100 is configured on a server and includes: a dialogue data acquisition unit 110, a recognition result acquisition unit 120, a multimodal feature acquisition unit 130, a feature type acquisition unit 140, an emotion recognition result acquisition unit 150, and a user prompt information sending unit 160.
[0084] The dialogue data acquisition unit 110 is used to respond to the dialogue access command and acquire the AI customer service agent and the current user dialogue data corresponding to the dialogue access command and the user authorization record command.
[0085] In this embodiment, the technical solution is described with the server as the executing entity. An intelligent communication platform is deployed on the server, integrating AI customer service (i.e., artificial intelligence customer service). When a user logs into the intelligent communication platform using a smart terminal (such as a smartphone, tablet, etc.), a chat box is displayed on the user interface of the user terminal. This chat box corresponds to an AI customer service agent, and the user can send any form of dialogue data to the AI customer service agent in this chat box, such as text, images, audio, and video. Of course, for high security of the dialogue data in the above communication process, it is necessary to obtain whether the user has authorized the uploading of dialogue data to the server for secure interaction before the start of the current dialogue. Only after the server obtains the dialogue access instruction and user authorization record instruction from the user terminal can it obtain the dialogue data entered by the user in the chat box and upload it to the server for subsequent intelligent dialogue. Furthermore, to more accurately open the AI customer service agent for chatting with the user, the server can either re-open the AI customer service agent based on the one the user has previously interacted with, or the server can first recommend 3-5 AI agents related to the user's user tags to the smart terminal, and then open the corresponding chat box of the AI customer service agent after the user selects one, and conduct subsequent intelligent interaction with the user. Different AI agents on the server specialize in different areas of intelligent interaction; for example, different AI agents may correspond to fields such as clinical psychology, counseling psychology, health psychology, sports psychology, and forensic psychology.
[0086] Since the above interaction process involves a private dialogue between the user and the AI customer service, the current user dialogue data can be encrypted and sent to the server. After the encrypted dialogue data is decrypted and processed on the server, it can be re-encrypted and stored on the server or periodically deleted to ensure the security of the dialogue data. Furthermore, users can conduct the dialogue in highly private locations (such as their residences) without face-to-face human interaction, facilitating smoother communication. To further intelligently confirm that the user is in a highly private location (such as their residence), the user's location data can be obtained before step S110, provided the user authorizes it. If the user's location data is determined to belong to a preset residential area location dataset, the user will be prompted that they are in a safe and private communication location.
[0087] The recognition result acquisition unit 120 is used to acquire the current data type of the current user dialogue data, and to recognize the current user dialogue data based on the recognition model corresponding to the current data type to obtain the current recognition result.
[0088] In this embodiment, after a user completes at least one round of dialogue (e.g., including at least one sentence initiated by the user and a corresponding reply from the AI customer service) and sends the current user dialogue data to the server, the server promptly determines the data type of the dialogue data to obtain the current data type. After obtaining the current data type, the server calls the corresponding recognition model to recognize the current user dialogue data and obtains the current recognition result.
[0089] In one embodiment, the recognition result acquisition unit 120 is specifically used for:
[0090] If the current data type is determined to be text, then the corresponding language fine-tuning model is obtained and the current user dialogue data is identified to obtain the current identification result;
[0091] If the current data type is determined to be speech, then the corresponding speech recognition model is obtained and the current user dialogue data is recognized to obtain the current recognition result.
[0092] In this embodiment, various types of recognition models are deployed in the server. For example, a language fine-tuning model for text fine-tuning and correction of text-type conversation content, and for another example, a speech recognition model for speech recognition of speech-type conversations. If it is determined that the user interacts with the AI customer service in text form, it is necessary to fully consider the situation where the user may enter some characters incorrectly. To facilitate more accurate text recognition, the current user conversation data can be recognized based on the language fine-tuning model to obtain the current recognition result. For example, the language fine-tuning model is a BERT model (a pre-trained language model based on bidirectional encoders and deep learning) and other models. A two-tower structure of "error detection - correction" is constructed through the BERT model. The detection model is used to determine the error position in the input text (for example, if the current user conversation data is: [CLS] I've been feeling a bit "shimian" lately [SEP], the detection model outputs the error position as 7); then the error position is input into the correction model to generate candidate words, and the correct vocabulary is predicted through the MLM (Masked Language Model) (for example, replacing "shimian" with "sleepless"). When the language fine-tuning model performs intelligent correction of incorrect characters on the current user conversation data of text type, the current recognition result can be obtained. Of course, if there are no incorrect characters in the current user conversation data, the current user conversation data will not be corrected after being input into the language fine-tuning model, but will be directly input as the original text as the current recognition result.
[0093] If it is determined that the user interacts with the AI customer service in voice form, then it is determined that the current data type is voice type. At this time, the current user conversation data is recognized using the speech recognition model to obtain the current recognition result. Specifically, the speech recognition model specifically uses the Wav2Vec model, which extracts features from speech data through self-supervised learning and can learn the acoustic and semantic information of speech. The Wav2Vec model includes a feature extractor (which can be composed of multiple layers of 1D convolutional neural networks, directly processing the original waveform in the current user conversation data and converting the time-domain signal into a feature sequence), a context network (which can be constructed using a bidirectional long short-term memory network, used to perform temporal modeling on the output of the feature extractor, capture long-range dependencies, and output a latent representation, and the latent representation includes rich acoustic and semantic information), and a contrastive learning objective jointly processed by multiple network layers (such as multiple network layers including fully connected layers, convolutional layers, etc.). By using the above speech recognition model to recognize the current user conversation data, the current recognition result can be obtained. It can be seen that the server can accurately recognize various types of user conversation data and obtain the current recognition result for subsequent intelligent interaction.
[0094] In one embodiment, the recognition result acquisition unit 120 is further specifically configured to:
[0095] If it is determined that the current data type is a mixed type, and the mixed type includes text type and speech type, then the language fine-tuning model and the speech recognition model are obtained;
[0096] The language fine-tuning model is used to identify the text content in the current user dialogue data to obtain a first current identification result;
[0097] The speech content in the current user dialogue data is identified by the speech recognition model to obtain a second current recognition result;
[0098] The first current recognition result and the second current recognition result are concatenated according to the dialogue sequence to obtain the current recognition result.
[0099] In this embodiment, if the current data type is determined to be a mixed type, and the mixed type includes both text and voice types, it means that the user and AI customer service have both text and voice dialogue data in a complete round of conversation. Since the sending time of each sentence in the intelligent interaction between the user and AI customer service is known, the above example can be used as a reference. Specifically, the language fine-tuning model is used to identify the text content in the current user dialogue data to obtain a first current recognition result; and the voice recognition model is used to identify the voice content in the current user dialogue data to obtain a second current recognition result.
[0100] For example, the current user dialogue data includes N1 text-type dialogues (N1 being a positive integer) and N2 speech-type dialogues (N2 being a positive integer). The server knows the sending time of each of the N1 text-type dialogues and the sending time of each of the N2 speech-type dialogues. After recognizing each of the N1 text-type dialogues using a language fine-tuning model to obtain a first recognition sub-result (N1 first recognition sub-results constitute a first current recognition result), and after recognizing each of the N2 speech-type dialogues using a speech recognition model to obtain a second recognition sub-result (N2 second recognition sub-results constitute a second current recognition result), the N1 first recognition sub-results in the first current recognition result and the N2 second recognition sub-results in the second current recognition result are concatenated according to the dialogue sequence to obtain the current recognition result. The obtained current recognition result is text data, facilitating subsequent semantic understanding by the server. Of course, if the current data type is determined to be a mixed type, and the mixed type includes more types such as text, voice, video, and image, then the language fine-tuning model recognizes the text content in the current user dialogue data, and the voice recognition model recognizes the voice content in the current user dialogue data. In addition, the image recognition model and the voice recognition model are also needed to recognize the dialogue data of video and / or image types to obtain the recognition results (mainly the recognition results corresponding to the text content and voice content in the video and / or image).
[0101] The multimodal feature acquisition unit 130 is used to acquire text features, speech features and behavioral features in the current user dialogue data, and form current multimodal features corresponding to the current recognition result.
[0102] In this embodiment, the server can also extract text features (such as sentence structure, emotional tendency, etc.), speech features (such as speech rate, duration of pauses, number of pauses in a sentence, number of pauses in the whole conversation, etc.) and behavioral features (such as the number of interruptions in dialogue input, etc.) from the original dialogue data of the current user dialogue data, thereby forming the current multimodal features corresponding to the current user dialogue data.
[0103] In one embodiment, the multimodal feature acquisition unit 130 is specifically used for:
[0104] The sentence structure statistics in the current user dialogue data are obtained as the text features;
[0105] The statistical information of the number of speech pauses in the current user dialogue data is obtained as the speech feature;
[0106] The statistical information on the number of input interruptions in the current user dialogue data is obtained as the behavioral feature.
[0107] In this embodiment, the text features are obtained by using a Natural Language Processing (NLP) model to acquire statistical information on sentence structure in the current user dialogue data (e.g., the sentence structures mainly include interrogative, declarative, imperative, and exclamatory sentences, and the total number of each sentence structure and its proportion in the complete dialogue round). Specifically, the speech features are obtained by using the aforementioned Wav2Vec model combined with Mel-frequency cepstral coefficients (MFCCs). The speech features are obtained by using the Wav2Vec model for text recognition of the current user dialogue data, and obtaining MFCCs from the corresponding audio signal to further analyze the audio features. The speech pause statistics are obtained by statistically analyzing the interruption times and cumulative terminal duration in the audio features. Furthermore, the input interruption statistics are obtained from the embedded components in the chat box when the user communicates with the AI customer service. After acquiring these three features, subsequent intelligent dialogue data analysis can be performed.
[0108] The feature type acquisition unit 140 is used to acquire the current feature type corresponding to the current multimodal feature based on a preset modal feature library.
[0109] In this embodiment, after acquiring the current multimodal features of the current user dialogue data, the current feature type corresponding to the current multimodal features can be further determined by combining the modality feature library, so as to analyze the psychological state, emotional state, etc. of the dialogue between the user and the AI customer service based on the current feature type.
[0110] In one embodiment, the feature type acquisition unit 140 is specifically used for:
[0111] Obtain the current feature vector composed of the text features, the speech features, and the behavior features;
[0112] Among the feature vectors corresponding to each multimodal feature in the modal feature library, the feature vector with the maximum vector similarity to the current feature vector is determined and used as the target feature vector;
[0113] Obtain the feature type corresponding to the target feature vector and use it as the current feature type corresponding to the current multimodal feature.
[0114] The feature type is either a pause / hesitation type or a normal dialogue type.
[0115] In this embodiment, after obtaining the text features as sentence structure statistics, the speech features as speech pause count statistics, and the behavior features as input interruption count statistics, the following steps are taken: the first proportion of interrogative sentences in the sentence structure statistics within the complete dialogue round; the second proportion between the total number of speech pauses in the complete dialogue round and the total number of sentences in the complete dialogue round in the speech pause count statistics; and the third proportion between the total number of input interruptions in the complete dialogue round and the total number of sentences in the complete dialogue round in the input interruption count statistics. Finally, the first proportion, the second proportion, and the third proportion are used to form the current feature vector.
[0116] Subsequently, since feature vectors have been pre-configured for each multimodal feature in the modal feature library, the feature vector similarity between the feature vectors corresponding to each multimodal feature in the modal feature library and the current feature vector can be calculated (e.g., the cosine similarity between the feature vector and the current feature vector is used as the feature vector similarity). The feature vector with the maximum vector similarity to the current feature vector is then selected as the target feature vector.
[0117] Finally, since feature types are pre-configured for each multimodal feature in the modal feature library, the feature type corresponding to the target feature vector is used as the current feature type corresponding to the current multimodal feature. For example, the feature type could be either a pause / hesitation type or a normal dialogue type. Of course, in specific implementations, the feature type is not limited to the two types in the above example; it can also include other types such as request for assistance. The types of feature special cases can be expanded according to the actual needs of the user.
[0118] The emotion recognition result acquisition unit 150 is used to perform emotion recognition on the current recognition result based on the pre-trained emotion recognition model to obtain the current emotion recognition result.
[0119] In this embodiment, after analyzing and obtaining the current feature type in the current user dialogue data, it can be further combined with a pre-trained emotion recognition model on the server, such as the BiLSTM-Attention model, which is a bidirectional long short-term memory network combined with an attention mechanism model. This model mainly includes an input layer, an embedding layer, a bidirectional long short-term memory layer (BiLSTM layer), an attention layer, a pooling layer, a fully connected layer, and an output layer connected in sequence. Through this emotion recognition model, emotion recognition can be performed on the current recognition result in text form to obtain the current emotion recognition result. Furthermore, the obtained current emotion recognition result can be combined with the obtained current feature type to form multi-dimensional user data to be analyzed.
[0120] In one embodiment, the emotion recognition result acquisition unit 150 is specifically used for:
[0121] Obtain the current semantic features corresponding to the current recognition result, and input the current semantic features into the emotion recognition model to obtain the current emotion recognition result.
[0122] In this embodiment, if the emotion recognition model adopts the BiLSTM-Attention model as shown in the example above, the current recognition result needs to be converted into the corresponding current semantic features based on the input layer and the embedding layer, and then processed sequentially through the bidirectional long short-term memory layer, attention layer, pooling layer, fully connected layer and output layer to obtain the current emotion recognition result.
[0123] The user prompt information sending unit 160 is used to obtain the corresponding user prompt information and send it to the user terminal corresponding to the dialogue access instruction if it is determined that the current feature type belongs to a preset feature type set and the current emotion recognition result belongs to a preset emotion recognition result set.
[0124] In this embodiment, if it is determined that the current feature type belongs to a preset feature type set (e.g., the preset feature type set includes at least the pause and hesitation type) and the current emotion recognition result belongs to a preset emotion recognition result set (e.g., the preset emotion recognition result set includes at least the emotion recognition results such as dissatisfaction, anxiety, worry, frustration, and disappointment), then it can be determined that the AI customer service needs to promptly issue a user prompt message to the user after the end of this round of dialogue, such as "After this dialogue ends, we feel that your recent state is not very good. Do you need a professional to communicate with you via voice?", and send the user prompt message to the user terminal corresponding to the dialogue access command. Moreover, in addition to text content, the user prompt message may also include two virtual buttons, namely "Agree" and "Reject". After receiving the above user prompt message, the user terminal can choose to click one of the two virtual buttons or ignore the message without making any click.
[0125] In one embodiment, the AI-based customer service user data processing device 100 further includes:
[0126] A target storage space location unit is used to obtain the dialogue access instruction and the target storage space corresponding to the user terminal, and to store the current user dialogue data in the target storage space; wherein, the target storage space is an encrypted storage space;
[0127] The manual follow-up request sending unit is used to send a manual follow-up request to the corresponding receiving terminal if it is determined that the information feedback interval of the user feedback information in response to the user prompt information exceeds a preset feedback duration threshold, so as to prompt the receiving terminal to conduct a manual follow-up communication with the user terminal.
[0128] In this embodiment, after the user completes the current round of communication with the AI customer service and completes the above-mentioned processes such as obtaining dialogue feature types and emotion recognition, in order to provide security for the dialogue data, the target storage space corresponding to the dialogue access instruction (mainly the user's unique identifier included in it) and the user terminal can be obtained in the server. Then, the target storage space is used as an encrypted storage space to store all the dialogue data between the user and the AI customer service, achieving a technical effect similar to "AI tree hole" and ensuring the security of the transmission and storage of dialogue data.
[0129] If the sending time of the user prompt and the sending time of the user feedback have been obtained, the interval between the sending time of the feedback and the sending time of the prompt is used as the information feedback interval. If the above information feedback interval exceeds a preset feedback interval threshold, it indicates that the user has not viewed the user prompt and provided corresponding feedback in a timely manner (such as agreeing to receive subsequent expert communication or refusing to receive subsequent expert communication). In this case, a manual follow-up request can be sent to the corresponding receiving terminal to prompt the receiving terminal to conduct a manual follow-up communication with the user terminal, so that the user can still receive the manual follow-up from relevant experts in a timely manner.
[0130] Furthermore, when performing emotion recognition on the current recognition result based on the pre-trained emotion recognition model in the previous steps, it can also combine the historical emotion recognition results corresponding to the user's recent (e.g., the last three days, the last week, the last half month, the last month, the last quarter, etc.) historical dialogue data in the target storage space to draw an emotion state change curve. The historical dialogue data stored in the target storage space can also serve as historical context data for the user's dialogue with AI customer service, enabling AI customer service to provide users with more accurate response data.
[0131] It is evident that the implementation of this device can intelligently acquire communication content features, including multimodal features and emotion recognition results, when AI customer service engages in highly private user conversations with users, and perform intelligent analysis. When relevant conditions are met, it can promptly generate prompt information and automatically send it to the user's terminal.
[0132] The aforementioned AI-based customer service user data processing device can be implemented as a computer program, which can, for example... Figure 7 It runs on the computer device shown.
[0133] Please see Figure 7 , Figure 7This is a schematic block diagram of a computer device provided in an embodiment of the present invention. This computer device integrates any of the AI-based customer service user data processing devices provided in the embodiments of the present invention.
[0134] See Figure 7 The computer device 400 includes a processor 402, a memory, and a network interface 405 connected via a system bus 401. The memory may include a storage medium 403 and internal memory 404.
[0135] The storage medium 403 may store an operating system 4031 and a computer program 4032. The computer program 4032 includes program instructions that, when executed, cause the processor 402 to perform a user data processing method based on AI customer service.
[0136] The processor 402 provides computing and control capabilities to support the operation of the entire computer device.
[0137] The internal memory 404 provides an environment for the computer program 4032 in the storage medium 403 to run. When the computer program 4032 is executed by the processor 402, the processor 402 can execute the above-mentioned user data processing method based on AI customer service.
[0138] This network interface 405 is used for network communication with other devices. Those skilled in the art will understand that... Figure 7 The structure shown is merely a block diagram of a portion of the structure related to the present invention and does not constitute a limitation on the computer device to which the present invention is applied. A specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0139] The processor 402 is used to run a computer program 4032 stored in the memory to implement the user data processing method based on AI customer service as described above.
[0140] It should be understood that, in this embodiment of the invention, the processor 402 may be a Central Processing Unit (CPU), or it may be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor.
[0141] It will be understood by those skilled in the art that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program includes program instructions and can be stored in a storage medium, which is a computer-readable storage medium. The program instructions are executed by at least one processor in the computer system to implement the process steps of the embodiments of the above methods.
[0142] Therefore, the present invention also provides a computer-readable storage medium. This computer-readable storage medium stores a computer program, wherein the computer program includes program instructions. When executed by a processor, the program instructions cause the processor to perform the user data processing method based on AI customer service as described above.
[0143] The storage medium can be any computer-readable storage medium that can store program code, such as a USB flash drive, external hard drive, read-only memory (ROM), magnetic disk, or optical disk.
[0144] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.
[0145] In the several embodiments provided by this invention, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative. For example, the division of each unit is merely a logical functional division, and there may be other division methods in actual implementation. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed.
[0146] The steps in the method of this invention can be adjusted, merged, or reduced in order according to actual needs. The units in the device of this invention can be merged, divided, or reduced according to actual needs. Furthermore, the functional units in the various embodiments of this invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0147] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, a terminal, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention.
[0148] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and these modifications or substitutions should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A user data processing method based on AI customer service, characterized in that, include: In response to a dialogue access command, acquire the AI customer service agent corresponding to the dialogue access command and the user authorization record command, as well as the current user dialogue data; Obtain the current data type of the current user dialogue data, and identify the current user dialogue data based on the recognition model corresponding to the current data type to obtain the current recognition result; The text features, speech features, and behavioral features in the current user dialogue data are obtained to form the current multimodal features corresponding to the current recognition result; The current feature type corresponding to the current multimodal feature is obtained based on a preset modal feature library; The current emotion recognition result is obtained by performing emotion recognition based on the pre-trained emotion recognition model. If it is determined that the current feature type belongs to a preset feature type set and the current emotion recognition result belongs to a preset emotion recognition result set, then the corresponding user prompt information is obtained and sent to the user terminal corresponding to the dialogue access command.
2. The method according to claim 1, characterized in that, The step of identifying the current user dialogue data based on the recognition model corresponding to the current data type to obtain the current recognition result includes: If the current data type is determined to be text, then the corresponding language fine-tuning model is obtained and the current user dialogue data is identified to obtain the current identification result; If the current data type is determined to be speech, then the corresponding speech recognition model is obtained and the current user dialogue data is recognized to obtain the current recognition result.
3. The method according to claim 2, characterized in that, The step of identifying the current user dialogue data based on the recognition model corresponding to the current data type to obtain the current recognition result further includes: If it is determined that the current data type is a mixed type, and the mixed type includes text type and speech type, then the language fine-tuning model and the speech recognition model are obtained; The language fine-tuning model is used to identify the text content in the current user dialogue data to obtain a first current identification result; The speech content in the current user dialogue data is identified by the speech recognition model to obtain a second current recognition result; The first current recognition result and the second current recognition result are concatenated according to the dialogue sequence to obtain the current recognition result.
4. The method according to claim 1, characterized in that, The acquisition of text features, voice features, and behavioral features from the current user dialogue data includes: The sentence structure statistics in the current user dialogue data are obtained as the text features; The statistical information of the number of speech pauses in the current user dialogue data is obtained as the speech feature; The statistical information on the number of input interruptions in the current user dialogue data is obtained as the behavioral feature.
5. The method according to claim 4, characterized in that, The process of obtaining the current feature type corresponding to the current multimodal feature based on a preset modal feature library includes: Obtain the current feature vector composed of the text features, the speech features, and the behavior features; Among the feature vectors corresponding to each multimodal feature in the modal feature library, the feature vector with the maximum vector similarity to the current feature vector is determined and used as the target feature vector; Obtain the feature type corresponding to the target feature vector and use it as the current feature type corresponding to the current multimodal feature; wherein, the feature type is either a pause / hesitation type or a normal dialogue type.
6. The method according to claim 1, characterized in that, The pre-trained emotion recognition model performs emotion recognition on the current recognition result to obtain the current emotion recognition result, including: Obtain the current semantic features corresponding to the current recognition result, and input the current semantic features into the emotion recognition model to obtain the current emotion recognition result.
7. The method according to claim 1, characterized in that, After the step of obtaining the corresponding user prompt information and sending it to the user terminal corresponding to the dialogue access command if it is determined that the current feature type belongs to a preset feature type set and the current emotion recognition result belongs to a preset emotion recognition result set, the method further includes: The system acquires the dialogue access instruction and the target storage space corresponding to the user terminal, and stores the current user dialogue data in the target storage space; wherein the target storage space is an encrypted storage space. If it is determined that the feedback interval of the user feedback information in response to the user prompt exceeds the preset feedback duration threshold, a manual follow-up request will be sent to the corresponding receiving terminal to prompt the receiving terminal to conduct a manual follow-up communication with the user terminal.
8. A user data processing device based on AI customer service, characterized in that, include: The dialogue data acquisition unit is used to respond to the dialogue access command and acquire the AI customer service agent and the current user dialogue data corresponding to the dialogue access command and the user authorization record command. The recognition result acquisition unit is used to acquire the current data type of the current user dialogue data, and to recognize the current user dialogue data based on the recognition model corresponding to the current data type to obtain the current recognition result; A multimodal feature acquisition unit is used to acquire text features, speech features, and behavioral features in the current user dialogue data to form current multimodal features corresponding to the current recognition result; The feature type acquisition unit is used to acquire the current feature type corresponding to the current multimodal feature based on a preset modal feature library; An emotion recognition result acquisition unit is used to perform emotion recognition on the current recognition result based on a pre-trained emotion recognition model to obtain the current emotion recognition result; The user prompt information sending unit is used to obtain the corresponding user prompt information and send it to the user terminal corresponding to the dialogue access instruction if it is determined that the current feature type belongs to a preset feature type set and the current emotion recognition result belongs to a preset emotion recognition result set.
9. A computer device, characterized in that, The computer device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the user data processing method based on AI customer service as described in any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, the computer program including program instructions, which, when executed by a processor, can implement the user data processing method based on AI customer service as described in any one of claims 1-7.