Human-computer interaction methods, devices, electronic equipment and media
By acquiring user voice data from human-computer interaction, extracting emotional and personality features, and using a pre-set model to predict and fuse emotional and personality types, the problem of misjudgment in emotion recognition in existing technologies is solved, achieving more accurate emotion guidance and improving user experience.
Patent Information
- Application Number
- CN202411974845.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-30
- Publication Date
- 2025-12-02
- Estimated Expiration
- 2044-12-30
AI Technical Summary
In existing technologies, AI-based emotion recognition struggles to accurately capture and understand the emotional states of users with different language habits, leading to misjudgments of the user's current emotional state and making it difficult to provide timely emotional guidance, thus affecting the user experience.
By acquiring user voice data from AI dialogues in human-computer interaction, extracting emotional and personality features, using a preset emotion and personality recognition model to predict the probability of emotion and personality types, and performing fusion processing, the emotion and personality fusion type with the highest probability is selected to obtain a matching guidance plan for interaction.
It improves the accuracy of judging the user's current emotional state, enabling timely and personalized emotional guidance and enhancing the user experience.
Smart Images

Figure CN119993157B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of artificial intelligence technology, specifically relating to a human-computer interaction method, device, electronic device, and medium. Background Technology
[0002] User emotion recognition is a common method used to improve the quality of AI (Artificial Intelligence) dialogue and human-like emotional expression. However, emotion is a subjective psychological experience, and individuals differ in their emotional expression and cognitive habits. Different users have different language habits and ways of expressing emotions.
[0003] Currently, in human-computer interaction, AI-based emotion recognition typically judges a user's current emotional state based on the content of their voice interaction or the fluctuations in their tone. This approach struggles to accurately capture and understand the emotional states of users with different language habits. For example, some users habitually express dissatisfaction humorously, and current emotion recognition models may misjudge their current emotional state. Consequently, it becomes difficult to provide timely and appropriate emotional guidance based on the user's current state, thus hindering a better user experience during human-computer interaction. Summary of the Invention
[0004] Based on the current state of human-computer interaction, the present invention provides a human-computer interaction method, device, electronic device and medium to overcome at least one technical problem existing in the prior art.
[0005] To achieve the above objectives, the present invention provides a human-computer interaction method, comprising:
[0006] Obtain current user voice data from AI dialogues in human-computer interaction;
[0007] Extracting the user's current emotional features from the user's voice data; and extracting the user's personality features from the text data by converting the user's voice data into text data;
[0008] By inputting the user's current emotional characteristics into a preset emotion recognition model, probability data of the current user's emotion type is obtained; and by inputting the user's personality characteristics into a preset personality recognition model, probability data of the user's personality type is obtained; wherein, the probability data of the current user's emotion type includes the emotion type and the emotion probability corresponding to the emotion type; the probability data of the user's personality type includes the personality type and the personality probability corresponding to the personality type;
[0009] The current user's emotion type probability data and the user's personality type probability data are fused together to obtain the emotion and personality fusion type and the corresponding fusion type probability.
[0010] The emotional personality fusion type corresponding to the fusion type with the highest probability value is selected as the user's current emotional personality fusion type.
[0011] The system retrieves a user emotion guidance scheme from a preset user emotion guidance scheme library that matches the user's current emotion and personality fusion type, and then conducts human-computer interaction with the user based on the user emotion guidance scheme.
[0012] To address the above problems, the present invention also provides a human-computer interaction device, the device comprising:
[0013] The voice acquisition module is used to acquire current user voice data from AI dialogues in human-computer interaction;
[0014] A multimodal feature extraction module is used to extract the user's current emotional features from the user's voice data; and to extract the user's personality features from the text data by converting the user's voice data into text data.
[0015] The probability estimation module is used to obtain current user emotion type probability data by inputting the user's current emotional characteristics into a preset emotion recognition model; and to obtain user personality type probability data by inputting the user's personality characteristics into a preset personality recognition model; wherein, the current user emotion type probability data includes emotion type and emotion probability corresponding to the emotion type; and the user personality type probability data includes personality type and personality probability corresponding to the personality type.
[0016] The fusion module is used to fuse the current user's emotion type probability data with the user's personality type probability data to obtain the emotion and personality fusion type and the corresponding fusion type probability.
[0017] The fusion type determination module is used to select the emotional personality fusion type corresponding to the fusion type with the highest probability value as the user's current emotional personality fusion type.
[0018] The human-computer interaction execution module is used to obtain a user emotion guidance scheme that matches the user's current emotion and personality fusion type from a preset user emotion guidance scheme library, and to conduct human-computer interaction with the user based on the user emotion guidance scheme.
[0019] To address the above problems, the present invention also provides an electronic device, the electronic device comprising:
[0020] At least one processor; and,
[0021] A memory communicatively connected to the at least one processor; wherein,
[0022] The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the steps in the human-computer interaction method as described above.
[0023] To address the aforementioned problems, the present invention also provides a computer-readable storage medium storing at least one instruction, which, when executed by a processor in an electronic device, implements the aforementioned human-computer interaction method.
[0024] The human-computer interaction method, device, electronic device, and medium provided by this invention acquire current user voice data from AI dialogue in human-computer interaction. The current emotional features extracted from the user's voice data and the user's personality features extracted from the text data converted from the user's voice data are respectively input into a preset emotion recognition model and a preset personality recognition model. This predicts the probability of the user currently experiencing different types of emotions and the probability of the user exhibiting different personality types, respectively, thus obtaining current user emotion type probability data and user personality type probability data. The user personality type probability data and the current user emotion type probability data are then fused to obtain an emotion-personality fusion type and a probability of fusion type. The emotion-personality fusion type with the highest probability value is used as the result of judging the user's current emotion. Based on this result, the most suitable emotion guidance scheme is obtained in a timely manner. By incorporating individual personality differences into the analysis of the user's current emotion, the judgment of the user's current emotional state is more accurate, enabling timely and accurate corresponding emotional guidance based on individual emotional changes, thereby improving the user experience. Attached Figure Description
[0025] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0026] Figure 1 This is a flowchart illustrating a human-computer interaction method provided in an embodiment of the present invention;
[0027] Figure 2 This is a schematic diagram of a human-computer interaction device provided in an embodiment of the present invention;
[0028] Figure 3 This is a schematic diagram of the internal structure of an electronic device that implements a human-computer interaction method according to an embodiment of the present invention.
[0029] Figure 4This is a schematic diagram of a human-computer interaction method provided in an embodiment of the present invention.
[0030] The realization of the objective, functional features and advantages of the present invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0031] It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.
[0032] Based on the problems existing in the prior art, the present invention mainly provides a human-computer interaction method, device, electronic device and medium. Its main purpose is to solve the problem that in the prior art, emotion recognition based on AI dialogue generally judges the user's current emotional state from the user's voice interaction content or tone fluctuations. It is difficult to completely and accurately capture and understand the emotional state of users with different language habits, which may lead to misjudgment of the user's current emotional state, and thus make it difficult to provide corresponding emotional guidance in a timely manner according to the user's current state.
[0033] Figure 1 This is a flowchart illustrating a human-computer interaction method provided in an embodiment of the present invention; Figure 4 This is a schematic diagram illustrating the principle of a human-computer interaction method according to an embodiment of the present invention. The method can be executed by a device, which can be implemented in software and / or hardware.
[0034] Figure 1 Combination Figure 4 A comprehensive description of human-computer interaction methods is provided. For example... Figure 1 Combination Figure 4 As shown in the figure, in this embodiment, the human-computer interaction method includes steps S110 to S160.
[0035] Step S110: Obtain the current user voice data from the AI dialogue of human-computer interaction.
[0036] Specifically, during human-computer interaction, voice data of the user is collected through voice collection devices, such as miniature microphones, and this voice data is used for subsequent extraction of emotional and personality characteristics.
[0037] As an optional embodiment of the present invention, obtaining current user voice data from AI dialogue in human-computer interaction includes:
[0038] Obtain raw user voice data within a preset time range from AI dialogues in human-computer interaction;
[0039] The high-frequency components of the original user voice data are pre-emphasized to obtain emphasized voice data.
[0040] The stressed speech data is subjected to frame-segmentation and windowing to obtain time-domain speech data;
[0041] Endpoint detection is performed on the time-domain speech data to select valid speech segments from the time-domain speech data, and the valid speech segments are used as the current user speech data.
[0042] Specifically, in the field of artificial intelligence, companion robots, voice assistants, and service robots in various scenarios generally require AI dialogue during human-computer interaction. Since changes in human emotions do not occur suddenly, to accurately and promptly detect these changes, voice data is collected within a preset time range using voice acquisition devices during human-computer interaction. For example, voice data collected within 5 minutes, half an hour, or 1 hour (as needed) serves as the basis for judging the user's current emotion, allowing for timely assessment. The user voice data acquired by the voice acquisition device is raw voice data; therefore, preprocessing is required before feature extraction to improve subsequent feature extraction processing.
[0043] The high-frequency components of the original user speech data are emphasized to remove the influence of lip radiation and glottal excitation during vocalization. This invention preferably, but is not limited to, using a first-order FIR high-pass digital filter (finite impulse response digital filter) to perform pre-emphasis processing on the original user speech data. After pre-emphasis processing, the speech signal yields a flatter spectrum, which is more beneficial for spectral analysis.
[0044] Speech signals typically exhibit relatively stable characteristics within a short timeframe of 10–30 ms, meaning they possess short-term stationarity. Therefore, frame-segmentation and windowing processing are necessary for emphasized speech data. Frame segmentation divides the speech signal into segments, with frames as the unit. Windowing emphasizes the speech waveform near the sampled area while attenuating the rest of the waveform. This embodiment preferably uses, but is not limited to, a Hamming window, which has lower frequency resolution, lower sidelobes, and less spectral leakage.
[0045] The time-domain speech data obtained after frame-segmentation and windowing may contain invalid speech segments. Therefore, it is necessary to select valid speech segments from the time-domain speech data to use them as the current user speech data. Endpoint detection is an effective method for detecting valid speech segments from a continuous speech stream. Endpoint detection identifies the start point (front end) and end point (back end) of valid speech. This embodiment preferably uses, but is not limited to, a dual-threshold method, combining an energy threshold and a zero-crossing rate threshold to detect speech endpoints.
[0046] The operation is as follows: First, potential speech segments are initially screened using an energy threshold. Then, within these segments, a zero-crossing rate threshold is used to further refine the endpoint detection, thereby improving the accuracy of endpoint detection. For example, a higher energy threshold T can be set. E1 and a lower energy threshold T E2 and a zero-crossing rate threshold T Z When the signal energy E(n) < T E1 When the signal energy E(n) < T, mark the possible start point of the speech; E2 At this point, the location where the speech might end is marked. Then, within these possible speech segments, further judgment is made based on the zero-crossing rate Z(n). When Z(n) > T... Z When the time is right, it is determined to be a valid speech segment.
[0047] Step S120: Extract the user's current emotional features from the user's voice data; and extract the user's personality features from the text data by converting the user's voice data into text data.
[0048] Specifically, user's current emotional characteristics can be extracted from user voice data, and user's personality characteristics can be extracted from dialogue content, i.e., text data converted from speech. The user's current emotional characteristics and personality characteristics extracted from the dual modality of auditory and text can be analyzed from different aspects to understand the user's current emotional state and avoid misjudgment of the user's current emotional state due to individual personality differences.
[0049] As an optional embodiment of the present invention, extracting the user's current emotional features from the user's voice data includes:
[0050] Perform Fourier transform processing on each frame of the speech signal in the user's speech data to obtain the user's speech amplitude spectrum;
[0051] The amplitude spectrum of the user's speech is converted by a Mel frequency bandpass filter bank to obtain the Mel spectrum;
[0052] Based on the Mel spectrum, calculate the logarithmic energy of the Mel spectrum obtained by each Mel frequency bandpass filter;
[0053] The discrete cosine transform is performed on all logarithmic energies to obtain the Mel frequency cepstral coefficient parameters, which are then used as the user's current sentiment characteristics.
[0054] Specifically, Mel Frequency Cepstral Coefficients (MFCCs) are features based on the frequency spectrum. Using frequency-domain-based feature parameters for emotion recognition achieves better performance because it is derived from the characteristics of the human auditory system and can simulate the human ear's perception of different frequencies of speech. First, a Discrete Fourier Transform (DFT) is used to convert the time-domain signal to the frequency-domain signal. The DFT is performed on each frame of the user's speech signal after windowing to obtain the user's speech amplitude spectrum. Then, the linear frequencies of the user's speech amplitude spectrum are converted to Mel frequencies. Next, triangular filters (preferably, but not limited to, 30 triangular filters) are uniformly divided in the Mel frequency domain. The energy of the signal on these filters is calculated, and the logarithm is taken. Finally, a Discrete Cosine Transform (DTC) is performed on the logarithmic result to obtain the Mel frequency cepstral coefficient parameters. These parameters are used as the user's current emotion feature.
[0055] As an optional embodiment of the present invention, user personality traits are extracted from the text data by converting user voice data into text data, including:
[0056] Convert user voice data into text data;
[0057] The text data is preprocessed to obtain semantically complete text statements with sentence lengths that conform to the preset sentence length.
[0058] Based on the Big Five personality domain dictionary, personality clue words are extracted from text statements to obtain a set of personality clue words;
[0059] The words in the personality clue word set are concatenated into personality clue sentences according to their positional relationship in the text data;
[0060] Calculate the context representation of each word in the text statement to obtain the text feature vector matrix; and calculate the context representation of each word in the personality cue statement to obtain the personality feature vector matrix.
[0061] By using conditional semantic fusion technology, word vectors in the personality feature vector matrix are fused into word vectors in the text feature vector matrix as external conditions, thereby obtaining a personality conditional fusion matrix.
[0062] The personality condition fusion matrix is used as the user's personality traits.
[0063] Specifically, user voice data is converted into text data using any existing speech-to-text conversion tool. The text data is then preprocessed to obtain semantically complete text statements with a preset sentence length. For example, preprocessing involves deleting and simplifying data with missing, repetitive, garbled, or meaningless characters. Data that is too short to effectively express personality or too long, resulting in redundant expression, is removed to form sentence C, i.e., the text statement.
[0064] Based on a Big Five personality domain dictionary and collected AI dialogue scenarios, user personality traits are labeled using an expert database with backgrounds in psychology and personality analysis. To ensure assessment accuracy, labeling is only performed when there is clear evidence in the text indicating the corresponding Big Five personality dimension. Based on the labeling, personality cue words (gerunds and / or adjectives) are extracted from the text statement (sentence C) to form a personality cue word set A = {A1, A2, ..., Ai}, where i represents the number of personality cue words. The words in the personality cue word set are concatenated according to their positional relationships in the text data to form a personality cue statement P. A pre-trained model BERT is used to vectorize sentence C and personality cue statement P to calculate the contextual representation of each word, resulting in the corresponding word vector matrix, i.e., the text feature vector matrix H. C and personality trait vector matrix H P By using conditional semantic fusion technology, the personality feature vector matrix H is... P The word vectors in the text are fused into the text feature vector matrix H as external conditions. C From the word vectors, the personality condition fusion matrix is obtained; the personality condition fusion matrix is used as the user's personality feature.
[0065] As an optional embodiment of the present invention, by using conditional semantic fusion technology, word vectors in the personality feature vector matrix are fused into the word vectors of the text feature vector matrix as external conditions, thereby obtaining a personality conditional fusion matrix, including:
[0066] The standard deviation of each word vector in the text feature vector matrix is normalized to obtain the text word vector matrix; the text feature vector matrix is represented as follows:
[0067]
[0068] Among them, H C Represents the text feature vector matrix. Represents the feature vector of text words. Represents n text word vectors, Represents the word separation vector of the text;
[0069] A conditional fusion function is used to conditionally fuse the personality cue feature vectors in the personality feature vector matrix with the text word vector matrix through semantic interaction, thereby obtaining the personality conditional fusion matrix; where the personality feature vector matrix is represented as:
[0070] in,
[0071] H P Represents a personality trait vector matrix. Represents the feature vector of personality cues. Represents n personality cue vectors, Represents the personality cue separation vector;
[0072] The formula for the conditional fusion function is:
[0073]
[0074] Among them, H m Let H represent the personality conditional fusion matrix, CLN represent the conditional fusion function, and H represent the personality conditional fusion matrix. normal c Represents the text word vector matrix. Represents the feature vector of personality cues, γ p H represents normal c The conditional gain vector, β p H represents normal c The conditional bias vector, W γ W represents the gain effect control matrix. β This represents the bias effect control matrix, b γ b represents the bias value of the gain effect control matrix. β This represents the bias value of the bias effect control matrix.
[0075] Specifically, the text feature vector matrix H C and personality trait vector matrix H P They are represented as follows:
[0076]
[0077] Among them, H C Represents the text feature vector matrix. Represents the feature vector of text words. Represents n text word vectors, H represents the text word separation vector; P Represents a personality trait vector matrix. Represents the feature vector of personality cues. Represents n personality cue vectors, Represents the personality cue separation vector;
[0078] The personality feature vector matrix H is obtained through conditional semantic fusion technology. P Vectors containing personality cues are fused into the text's feature vectors as external conditions. First, the text feature vector matrix H... C Each word vector in the dataset is normalized using its standard deviation, as shown in the following formula:
[0079]
[0080] Among them, H normal c Let represent the text word vector matrix, μ and σ represent the mean and variance of each word vector, i represent the index of the word vector in the matrix, and n represent the number of word vectors in the matrix.
[0081] The conditional fusion function CLN is used to integrate personality cue feature vectors. The text word vector matrix H after standard deviation normalization normal c Obtain the conditional fusion matrix H of dynamic semantic interaction between text and personality cues. m H normal c The conditional gain vector γ p and conditional bias vector β p Each is controlled by the gain effect control matrix W. γ and bias effect control matrix W β Characteristic vectors of personality cues Multiply and then multiply by their respective bias values b. γ and b β Adding them together, we get W. γ W β b γ and b β All of these can be obtained dynamically during model training and can be considered as known. The final result is the personality condition fusion matrix.
[0082] Step S130: By inputting the user's current emotional characteristics into a preset emotion recognition model, the current user's emotion type probability data is obtained; and by inputting the user's personality characteristics into a preset personality recognition model, the user's personality type probability data is obtained; wherein, the current user's emotion type probability data includes the emotion type and the emotion probability corresponding to the emotion type; the user's personality type probability data includes the personality type and the personality probability corresponding to the personality type.
[0083] Specifically, by inputting the user's current emotional characteristics into the preset emotion recognition model and the user's personality characteristics into the preset emotion recognition model and preset personality recognition model respectively, the probability of the user currently experiencing different emotions and the probability of the user having different personalities can be obtained, so as to facilitate the subsequent fusion of emotions and personalities.
[0084] As an optional embodiment of the present invention, by inputting the user's current emotional characteristics into a preset emotion recognition model to obtain probability data of the current user's emotion type, the method includes:
[0085] Input the user's current emotional characteristics into a preset emotion recognition model; the preset emotion recognition model contains a probability prediction module for different types of emotions;
[0086] The probability prediction modules for different types of emotions are used to calculate the probability of the corresponding emotion type for the user's current emotional characteristics, so as to obtain different emotion types and the corresponding emotion probabilities.
[0087] Different emotion types and the corresponding emotion probabilities are used as the current user's emotion type probability data.
[0088] Specifically, the preset emotion recognition model is preferably obtained through pre-training of a neural network. This preferably includes, but is not limited to, convolutional layers for acquiring local features, pooling layers for dimensionality reduction of the local features acquired by the convolutional layers, and fully connected layers for the user to synthesize the dimensionality-reduced local features from the pooling layers. Since the above constitutes the basic structure of a neural network, it will not be elaborated further. In this embodiment, the preset emotion recognition model internally includes probability prediction modules for different types of emotions. The activation function within each module, after being trained with sample data, can calculate and process the comprehensive features provided by the fully connected layer, thereby obtaining the probabilities of different emotion types. That is, after the user's current emotional features are input into the preset emotion recognition model, they sequentially pass through convolutional layers, pooling layers, and fully connected layers to obtain comprehensive features extracted based on the user's current emotional features. These comprehensive features are then used as input to the probability prediction modules for different types of emotions. After calculation by these modules, different emotion types and their corresponding probabilities are obtained, i.e., the current user's emotion type probability data. For example, after inputting the user's current emotional characteristics into a preset emotion recognition model, the output results are as follows: the probability of the user's current emotion being happy is 70%, the probability of being angry is 10%, the probability of being sad is 15%, and the probability of being grief is 13%, etc.
[0089] As an optional embodiment of the present invention, user personality characteristics are input into a preset personality recognition model to obtain user personality type probability data, including:
[0090] Input the user's personality traits into a preset personality recognition model; the preset personality recognition model includes a probability prediction module for different personality types.
[0091] The probability prediction modules for different personality types are used to calculate the probability of user personality traits corresponding to different personality types, so as to obtain different personality types and the personality probabilities corresponding to the personality types.
[0092] Different personality types and their corresponding probability are used as user personality type probability data.
[0093] Specifically, the preferred personality recognition model is obtained through pre-training via a neural network. The feature vector at the CLS position in the personality conditional fusion matrix is used as the global representation of the sentence. First, the conditional fusion matrix H... m The average fusion vector h is calculated based on the word vector dimension. e The input is fed into a fully connected network responsible for personality recognition and the output is a personality recognition vector h. o Then, the pooling function trained within the probability prediction module for different personality types is used to predict h. o The model calculates personality probabilities to obtain different personality types and their corresponding probabilities. For example, after inputting a user's personality traits into a preset emotion recognition model, the output results might show that the user's personality type has an 80% probability of being cheerful and a 10% probability of being melancholic.
[0094] Step S140: Merge the current user's emotion type probability data with the user's personality type probability data to obtain the emotion and personality fusion type and the corresponding fusion type probability.
[0095] Specifically, the probability data of the current user's emotion type is fused with the probability data of the user's personality type to obtain the personality-emotion combination type and its corresponding probability. This allows for the provision of corresponding interaction plans based on personality and emotion. For example, fusion of the user's current emotion of being happy with the user's personality type of being cheerful yields the fusion probability of the emotion of being happy and the probability of the personality type of being cheerful.
[0096] As an optional embodiment of the present invention, the current user's emotion type probability data and user's personality type probability data are fused to obtain the emotion-personality fusion type and the corresponding fusion type probability, including:
[0097] Based on Bayes' theorem, the probability of each emotion type in the current user's emotion type probability data is multiplied by the probability of each personality type in the user's personality type probability data, and then multiplied by the prior probability of the preset emotion recognition model and the prior probability of the preset personality recognition model to obtain the emotion and personality fusion type and the corresponding fusion type probability.
[0098] Specifically, using Bayes' theorem to fuse the prediction results of two models can more accurately reflect the relationship between personality traits and emotional states, thereby better enabling the fusion model to comprehensively judge the emotional state under the influence of individual personality. This leads to more reasonable AI dialogue content decisions and tone corrections.
[0099] Step S150: Select the emotional personality fusion type corresponding to the fusion type with the highest probability value as the user's current emotional personality fusion type.
[0100] Specifically, the emotional and personality fusion type corresponding to the fusion type with the highest probability value is the user's most accurate current emotional state. Therefore, a corresponding interaction plan can be given to the user based on this emotional and personality fusion type.
[0101] Step S160: Obtain a user emotion guidance plan that matches the user's current emotional personality fusion type from the preset user emotion guidance plan library, and conduct human-computer interaction with the user based on the user emotion guidance plan.
[0102] Specifically, a pre-set user emotion guidance scheme library is provided, which sets corresponding user emotion guidance schemes for each type of user personality under different emotions, including tone of voice during interaction, automatically provided music, etc.
[0103] like Figure 2 As shown, the present invention provides a human-computer interaction device 200, which can be installed in an electronic device. Depending on the functions implemented, the human-computer interaction device 200 may include: a voice acquisition module 210, a multimodal feature extraction module 220, a probability estimation module 230, a fusion module 240, a fusion type determination module 250, and a human-computer interaction execution module 260. The unit of the present invention can also be referred to as a module, which refers to a series of computer program segments that can be executed by the processor of an electronic device and can perform a fixed function, and are stored in the memory of the electronic device.
[0104] In this embodiment, the functions of each module / unit are as follows:
[0105] The voice acquisition module 210 is used to acquire current user voice data from AI dialogue in human-computer interaction;
[0106] The multimodal feature extraction module 220 is used to extract the user's current emotional features from the user's voice data; and to extract the user's personality features from the text data by converting the user's voice data into text data.
[0107] The probability estimation module 230 is used to obtain the probability data of the current user's emotion type by inputting the user's current emotional characteristics into a preset emotion recognition model; and to obtain the probability data of the user's personality type by inputting the user's personality characteristics into a preset personality recognition model; wherein, the probability data of the current user's emotion type includes the emotion type and the emotion probability corresponding to the emotion type; and the probability data of the user's personality type includes the personality type and the personality probability corresponding to the personality type.
[0108] The fusion module 240 is used to fuse the current user's emotion type probability data with the user's personality type probability data to obtain the emotion and personality fusion type and the corresponding fusion type probability.
[0109] The fusion type determination module 250 is used to select the emotional personality fusion type corresponding to the fusion type with the highest probability value as the user's current emotional personality fusion type.
[0110] The human-computer interaction execution module 260 is used to obtain a user emotion guidance scheme that matches the user's current emotion and personality fusion type from a preset user emotion guidance scheme library, and to conduct human-computer interaction with the user based on the user emotion guidance scheme.
[0111] The human-computer interaction device 200 of the present invention acquires current user voice data from AI dialogue in human-computer interaction, and inputs the current emotional features of the user extracted from the user voice data and the user personality features extracted from the text data converted from the user voice data into preset emotion recognition models and preset personality recognition models, respectively. This predicts the probability of the user currently experiencing different types of emotions and the probability of the user having different types of personalities, thus obtaining current user emotion type probability data and user personality type probability data. The user personality type probability data and the current user emotion type probability data are then fused to obtain an emotion-personality fusion type and a probability of fusion type. The emotion-personality fusion type with the highest probability value is taken as the result of judging the user's current emotion, and based on this result, the most suitable emotion guidance scheme is obtained in a timely manner. By incorporating individual personality differences into the analysis of the user's current emotion, the judgment of the user's current emotional state is more accurate, and corresponding emotional guidance can be provided in a timely and accurate manner according to the user's individual emotional changes, thereby improving the user experience.
[0112] like Figure 3 As shown, the present invention provides an electronic device 3 with a human-computer interaction method.
[0113] The electronic device 3 may include a processor 30, a memory 31, and a bus, and may also include a computer program, such as a human-computer interaction program 32, stored in the memory 31 and executable on the processor 30. The memory 31 may include both internal storage units of the human-computer interaction device and external storage devices. The memory 31 can be used not only to store application software and various types of data, such as the code of the human-computer interaction program, but also to temporarily store data that has been output or will be output.
[0114] The memory 31 includes at least one type of readable storage medium, such as flash memory, portable hard drive, multimedia card, card-type memory (e.g., SD or DX memory), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the memory 31 can be an internal storage unit of the electronic device 3, such as a portable hard drive. In other embodiments, the memory 31 can be an external storage device of the electronic device 3, such as a plug-in portable hard drive, smart media card (SMC), secure digital card (SD), flash card, etc. Furthermore, the memory 31 can include both internal and external storage units of the electronic device 3. The memory 31 can be used not only to store application software and various types of data installed on the electronic device 3, such as human-computer interaction method code, but also to temporarily store data that has been output or will be output.
[0115] In some embodiments, the processor 30 may be composed of integrated circuits, such as a single packaged integrated circuit or multiple integrated circuits with the same or different functions, including combinations of one or more central processing units (CPUs), microprocessors, digital processing chips, graphics processors, and various control chips. The processor 30 is the control unit of the electronic device, connecting various components of the entire electronic device through various interfaces and lines. It executes programs or modules (such as human-computer interaction programs) stored in the memory 31, and calls data stored in the memory 31 to perform various functions of the electronic device 3 and process data.
[0116] The bus can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This bus can be divided into an address bus, a data bus, a control bus, etc. The bus is configured to enable communication between the memory 31 and at least one processor 30, etc.
[0117] Figure 3 Only electronic devices with components are shown; it will be understood by those skilled in the art that... Figure 3 The structure shown does not constitute a limitation on the electronic device 3, and may include fewer or more components than shown, or combine certain components, or have different component arrangements.
[0118] For example, although not shown, the electronic device 3 may also include a power supply (such as a battery) to power various components. Preferably, the power supply can be logically connected to the at least one processor 30 through a power management system, thereby enabling functions such as charging management, discharging management, and power consumption management through the power management system. The power supply may also include one or more DC or AC power supplies, recharging systems, power fault detection circuits, power converters or inverters, power status indicators, and other arbitrary components. The electronic device 3 may also include various sensors, Bluetooth modules, Wi-Fi modules, etc., which will not be described in detail here.
[0119] Furthermore, the electronic device 3 may also include a network interface. Optionally, the network interface may include a wired interface and / or a wireless interface (such as a Wi-Fi interface, a Bluetooth interface, etc.), which is typically used to establish communication connections between the electronic device 3 and other electronic devices.
[0120] Optionally, the electronic device 3 may further include a user interface, which may be a display, an input unit (such as a keyboard), or a standard wired or wireless interface. Optionally, in some embodiments, the display may be an LED display, a liquid crystal display, a touch-sensitive liquid crystal display, or an OLED (Organic Light-Emitting Diode) touchscreen. The display may also be appropriately referred to as a screen or display unit, used to display information processed in the electronic device 3 and to display a visual user interface.
[0121] It should be understood that the embodiments described are for illustrative purposes only and are not limited to this structure in the scope of the invention application.
[0122] The human-computer interaction program 32 stored in the memory 31 of the electronic device 3 is a combination of multiple instructions, which, when run in the processor 30, can achieve the following:
[0123] Step S110: Obtain the current user voice data from the AI dialogue of human-computer interaction;
[0124] Step S120: Extract the user's current emotional features from the user's voice data; and extract the user's personality features from the text data by converting the user's voice data into text data.
[0125] Step S130: By inputting the user's current emotional characteristics into a preset emotion recognition model, the current user's emotion type probability data is obtained; and by inputting the user's personality characteristics into a preset personality recognition model, the user's personality type probability data is obtained; wherein, the current user's emotion type probability data includes the emotion type and the emotion probability corresponding to the emotion type; the user's personality type probability data includes the personality type and the personality probability corresponding to the personality type.
[0126] Step S140: Merge the current user's emotion type probability data with the user's personality type probability data to obtain the emotion-personality fusion type and the corresponding fusion type probability.
[0127] Step S150: Select the emotional personality fusion type corresponding to the fusion type with the highest probability value as the user's current emotional personality fusion type;
[0128] Step S160: Obtain a user emotion guidance plan that matches the user's current emotional personality fusion type from the preset user emotion guidance plan library, and conduct human-computer interaction with the user based on the user emotion guidance plan.
[0129] Specifically, the processor 30's implementation method for the above instructions can be found in [reference needed]. Figure 1 The descriptions of the relevant steps in the corresponding embodiments are not repeated here.
[0130] Furthermore, if the modules / units integrated in the electronic device 3 are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. The computer-readable medium may include: any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a portable hard drive, a magnetic disk, an optical disk, a computer memory, or a read-only memory (ROM).
[0131] This invention also provides a computer-readable storage medium, which may be non-volatile or volatile, and stores a computer program that is implemented when executed by a processor:
[0132] Step S110: Obtain the current user voice data from the AI dialogue of human-computer interaction;
[0133] Step S120: Extract the user's current emotional features from the user's voice data; and extract the user's personality features from the text data by converting the user's voice data into text data.
[0134] Step S130: By inputting the user's current emotional characteristics into a preset emotion recognition model, the current user's emotion type probability data is obtained; and by inputting the user's personality characteristics into a preset personality recognition model, the user's personality type probability data is obtained; wherein, the current user's emotion type probability data includes the emotion type and the emotion probability corresponding to the emotion type; the user's personality type probability data includes the personality type and the personality probability corresponding to the personality type.
[0135] Step S140: Merge the current user's emotion type probability data with the user's personality type probability data to obtain the emotion-personality fusion type and the corresponding fusion type probability.
[0136] Step S150: Select the emotional personality fusion type corresponding to the fusion type with the highest probability value as the user's current emotional personality fusion type;
[0137] Step S160: Obtain a user emotion guidance plan that matches the user's current emotional personality fusion type from the preset user emotion guidance plan library, and conduct human-computer interaction with the user based on the user emotion guidance plan.
[0138] Specifically, the specific implementation method of the computer program when executed by the processor can be referred to the description of the relevant steps in the human-computer interaction method of the embodiment, and will not be repeated here.
[0139] In the several embodiments provided by this invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules is only a logical functional division, and other division methods may be used in actual implementation.
[0140] The modules described as separate components may or may not be physically separate. The components shown as modules may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0141] Furthermore, the functional modules in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or in the form of hardware plus software functional modules.
[0142] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention.
[0143] Therefore, the embodiments should be considered exemplary and non-limiting in all respects, and the scope of the invention is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of equivalents of the claims are intended to be embraced within the invention. No appended diagram markings in the claims should be construed as limiting the scope of the claims.
[0144] Furthermore, it is clear that the word "comprising" does not exclude other units or steps, and the singular does not exclude the plural. Multiple units or devices recited in a device claim may also be implemented by a single unit or device through software or hardware. The term "second class" is used to indicate names and does not indicate any specific order.
[0145] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention.
Claims
1. A human-computer interaction method, characterized in that, Includes the following steps: Obtain current user voice data from AI dialogues in human-computer interaction; Extracting the user's current emotional features from the user's voice data; and extracting the user's personality features from the text data by converting the user's voice data into text data; By inputting the user's current emotional characteristics into a preset emotion recognition model, probability data of the current user's emotion type can be obtained. And by inputting the user's personality characteristics into a preset personality recognition model, user personality type probability data is obtained; wherein, the current user emotion type probability data includes emotion type and emotion probability corresponding to the emotion type; the user personality type probability data includes personality type and personality probability corresponding to the personality type; The current user's emotion type probability data and the user's personality type probability data are fused together to obtain the emotion and personality fusion type and the corresponding fusion type probability. The emotional personality fusion type corresponding to the fusion type with the highest probability value is selected as the user's current emotional personality fusion type. The system retrieves a user emotion guidance scheme from a preset user emotion guidance scheme library that matches the user's current emotion and personality fusion type, and then conducts human-computer interaction with the user based on the user emotion guidance scheme.
2. The human-computer interaction method according to claim 1, characterized in that, The acquisition of current user voice data from AI dialogue in human-computer interaction includes: Obtain raw user voice data within a preset time range from AI dialogues in human-computer interaction; The high-frequency components of the original user voice data are pre-emphasized to obtain emphasized voice data. The stressed speech data is subjected to frame-segmentation and windowing processing to obtain time-domain speech data; Endpoint detection is performed on the time-domain speech data to select valid speech segments from the time-domain speech data, and the valid speech segments are used as the current user speech data.
3. The human-computer interaction method according to claim 1, characterized in that, Extracting the user's current emotional features from the user's voice data includes: Perform Fourier transform processing on each frame of the speech signal in the user speech data to obtain the user speech amplitude spectrum; The amplitude spectrum of the user's speech is converted by a Mel frequency bandpass filter bank to obtain the Mel spectrum; Based on the Mel spectrum, calculate the logarithmic energy of the Mel spectrum obtained by each Mel frequency bandpass filter; The discrete cosine transform is performed on all the logarithmic energies to obtain the Mel frequency cepstral coefficient parameters, which are then used as the user's current emotional characteristics.
4. The human-computer interaction method according to claim 1, characterized in that, The step of converting the user's voice data into text data to extract user personality traits from the text data includes: Convert the user's voice data into text data; The text data is preprocessed to obtain semantically complete text statements with sentence lengths conforming to a preset length. Based on the Big Five personality domain dictionary, personality clue words are extracted from the text statements to obtain a set of personality clue words; The words in the personality clue word set are concatenated into personality clue sentences according to their positional relationship in the text data; Calculate the context representation of each word in the text statement to obtain a text feature vector matrix; and calculate the context representation of each word in the personality clue statement to obtain a personality feature vector matrix. By using conditional semantic fusion technology, the word vectors in the personality feature vector matrix are fused into the word vectors in the text feature vector matrix as external conditions, thereby obtaining a personality conditional fusion matrix. The personality condition fusion matrix is used as the user's personality characteristics.
5. The human-computer interaction method according to claim 4, characterized in that, The step involves using conditional semantic fusion technology to fuse word vectors from the personality feature vector matrix as external conditions into the word vectors of the text feature vector matrix, thereby obtaining a personality conditional fusion matrix, including: The standard deviation of each word vector in the text feature vector matrix is normalized to obtain the text word vector matrix; the text feature vector matrix is represented as follows: ; in, Represents the text feature vector matrix. Represents the feature vector of text words. Represents n text word vectors, Represents the word separation vector of the text; A conditional fusion function is used to conditionally fuse the personality cue feature vectors in the personality feature vector matrix with the text word vector matrix through semantic interaction, thereby obtaining a personality conditional fusion matrix; wherein, the personality feature vector matrix is represented as follows: ;in, Represents a personality trait vector matrix. Represents the feature vector of personality cues. Represents n personality cue vectors, Represents the personality cue separation vector; The formula for the conditional fusion function is: in, This represents the personality condition fusion matrix. CLN This represents the conditional fusion function. Represents the text word vector matrix. Represents the feature vector of personality cues. express The conditional gain vector, express The conditional bias vector. This represents the gain effect control matrix. This represents the bias effect control matrix. This represents the bias value of the gain effect control matrix. This represents the bias value of the bias effect control matrix.
6. The human-computer interaction method according to claim 1, characterized in that, The step of inputting the user's current emotional characteristics into a preset emotion recognition model to obtain probability data of the current user's emotion type includes: The user's current emotional characteristics are input into a preset emotion recognition model; wherein, the preset emotion recognition model is equipped with a probability prediction module for different types of emotions; The probability prediction modules for different types of emotions are used to calculate the probability of the corresponding emotion type for the user's current emotional characteristics, so as to obtain different emotion types and the emotion probability corresponding to the emotion type. The different emotion types and the corresponding emotion probabilities are used as the current user's emotion type probability data.
7. The human-computer interaction method according to claim 1, characterized in that, The step of inputting the user's personality traits into a preset personality recognition model to obtain user personality type probability data includes: The user's personality traits are input into a preset personality recognition model; wherein, the preset personality recognition model is equipped with a probability prediction module for different personality types; The probability prediction modules for different personality types are used to calculate the probability of the user's personality traits corresponding to the personality types, so as to obtain different personality types and the personality probabilities corresponding to the personality types. The different personality types and the corresponding personality probabilities are used as user personality type probability data.
8. The human-computer interaction method according to claim 1, characterized in that, The step of fusing the current user's emotion type probability data with the user's personality type probability data to obtain the emotion / personality fusion type and the corresponding fusion type probability includes: Based on Bayes' theorem, the emotion probability corresponding to each type of emotion in the current user emotion type probability data is multiplied by the personality probability corresponding to each type of personality in the user personality type probability data, and then multiplied by the prior probability of the preset emotion recognition model and the prior probability of the preset personality recognition model to obtain the emotion and personality fusion type and the corresponding fusion type probability.
9. A human-computer interaction device, characterized in that, The device includes: The voice acquisition module is used to acquire current user voice data from AI dialogues in human-computer interaction; A multimodal feature extraction module is used to extract the user's current emotional features from the user's voice data; and to extract the user's personality features from the text data by converting the user's voice data into text data. The probability estimation module is used to obtain current user emotion type probability data by inputting the user's current emotional characteristics into a preset emotion recognition model; and to obtain user personality type probability data by inputting the user's personality characteristics into a preset personality recognition model; wherein, the current user emotion type probability data includes emotion type and emotion probability corresponding to the emotion type; and the user personality type probability data includes personality type and personality probability corresponding to the personality type. The fusion module is used to fuse the current user's emotion type probability data with the user's personality type probability data to obtain the emotion and personality fusion type and the corresponding fusion type probability. The fusion type determination module is used to select the emotional personality fusion type corresponding to the fusion type with the highest probability value as the user's current emotional personality fusion type. The human-computer interaction execution module is used to obtain a user emotion guidance scheme that matches the user's current emotion and personality fusion type from a preset user emotion guidance scheme library, and to conduct human-computer interaction with the user based on the user emotion guidance scheme.
10. An electronic device, characterized in that, The electronic device includes: At least one processor; and, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the steps in the human-computer interaction method as described in any one of claims 1 to 8.
11. A computer-readable storage medium storing at least one instruction, characterized in that, When the at least one instruction is executed by a processor in an electronic device, it implements the human-computer interaction method as described in any one of claims 1 to 8.
Citation Information
Patent Citations
Multi-mode intelligent emotion sensing system
CN107220591A
Speech emotion recognition method and device and related equipment
CN110751943A