Man-machine interaction method and device, electronic equipment and medium

By extracting emotional and personality traits in user voice data in human-computer interaction and integrating the output of emotional and personality recognition models, the problem of misjudgment of emotions recognition in the existing technology is solved, more accurate judgment of emotional state and personalized guidance are achieved, and user experience is improved.

CN119993157AActive Publication Date: 2025-05-13GOERTEK INC
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202411974845.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-30
Publication Date
2025-05-13
Estimated Expiration
2044-12-30

AI Technical Summary

Technical Problem

The prior art is difficult to accurately capture and understand the user's emotional state of different language habits during human-computer interaction, which may lead to misjudgment of emotions and thus difficult to provide timely emotional guidance.

Method used

By obtaining user voice data from AI dialogues of human-computer interaction, extracting the user's current emotional characteristics and personality characteristics, and inputting preset emotions recognition models and personality recognition models respectively, integrating emotions type probability data and personality type probability data to obtain emotional personality fusion types and corresponding probability.

Benefits of technology

It improves the accurate judgment of the user's current emotional state, and can provide timely corresponding emotional guidance based on the user's individual emotional changes to improve the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119993157A_ABST
    Figure CN119993157A_ABST
Patent Text Reader

Abstract

The invention provides a man-machine interaction method and device, electronic equipment and a medium, and belongs to the technical field of artificial intelligence. Inputting the extracted current emotion features of the user and the character features of the user extracted from the text data converted from the current emotion features into a preset emotion recognition model and a preset character recognition model respectively to obtain current user emotion type probability data and user character type probability data; and obtaining an emotional character fusion type and fusion type probabilities, taking the emotional character fusion type corresponding to the maximum fusion type probability as a current emotion judgment result of the user, obtaining a most matched emotion guide scheme based on the result, and fusing individual character difference factors of the user to an emotion analysis process of the user. The current emotional state of the user is judged more accurately, corresponding sexual emotion guidance can be given timely and accurately according to the individual emotion change of the user, and the user experience is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of artificial intelligence technology, and specifically relates to a human-computer interaction method, device, electronic equipment and medium. Background Art

[0002] User emotion recognition is currently a common way to improve the quality of AI (Artificial Intelligence) conversations and anthropomorphic emotional expression, but emotion is a subjective psychological experience, and there are differences in individual emotional expression and cognitive habits. Different users have different language habits and emotional expression methods.

[0003] At present, in the process of human-computer interaction, emotion recognition based on AI dialogue generally judges the user's current emotional state from the user's voice interaction content or tone fluctuations. It is difficult to fully and accurately capture and understand the emotional state of users with different language habits. For example, some users are accustomed to expressing dissatisfaction in a humorous way, and the existing emotion recognition model may misjudge their current emotional state. As a result, it is difficult to give corresponding emotional guidance in time according to the user's current state, so that the user can get a better experience in the process of human-computer interaction. Summary of the invention

[0004] Based on the above current status of human-computer interaction, the present invention provides a human-computer interaction method, device, electronic device and medium to overcome at least one technical problem existing in the prior art.

[0005] To achieve the above object, the present invention provides a human-computer interaction method, comprising:

[0006] Obtain current user voice data from the AI ​​dialogue of human-computer interaction;

[0007] Extracting the user's current emotional features from the user's voice data; and extracting the user's personality features from the text data by converting the user's voice data into text data;

[0008] By inputting the current emotion characteristics of the user into a preset emotion recognition model, the current user emotion type probability data is obtained; and by inputting the user personality characteristics into a preset personality recognition model, the user personality type probability data is obtained; wherein the current user emotion type probability data includes an emotion type and an emotion probability corresponding to the emotion type; the user personality type probability data includes a personality type and a personality probability corresponding to the personality type;

[0009] Fusing the current user emotion type probability data with the user personality type probability data to obtain an emotion-personality fusion type and a corresponding fusion type probability;

[0010] Select the emotion-personality fusion type corresponding to the fusion type probability with the largest probability value as the user's current emotion-personality fusion type;

[0011] A user emotion guidance scheme matching the current emotion-personality fusion type of the user is obtained from a preset user emotion guidance scheme library, and human-computer interaction is performed with the user based on the user emotion guidance scheme.

[0012] In order to solve the above problems, the present invention further provides a human-computer interaction device, the device comprising:

[0013] The voice acquisition module is used to obtain the current user voice data from the AI ​​dialogue of human-computer interaction;

[0014] A multimodal feature extraction module, used to extract the user's current emotional features from the user's voice data; and to extract the user's personality features from the text data by converting the user's voice data into text data;

[0015] A probability estimation module, configured to obtain current user emotion type probability data by inputting the current user emotion characteristics into a preset emotion recognition model; and to obtain user personality type probability data by inputting the user personality characteristics into a preset personality recognition model; wherein the current user emotion type probability data includes an emotion type and an emotion probability corresponding to the emotion type; and the user personality type probability data includes a personality type and a personality probability corresponding to the personality type;

[0016] A fusion module, used for fusing the current user emotion type probability data with the user personality type probability data to obtain an emotion personality fusion type and a corresponding fusion type probability;

[0017] A fusion type determination module is used to select the emotion and personality fusion type corresponding to the fusion type probability with the largest probability value as the user's current emotion and personality fusion type;

[0018] The human-computer interaction execution module is used to obtain a user emotion guidance scheme that matches the current emotion and personality fusion type of the user from a preset user emotion guidance scheme library, and perform human-computer interaction with the user based on the user emotion guidance scheme.

[0019] In order to solve the above problem, the present invention further provides an electronic device, the electronic device comprising:

[0020] at least one processor; and,

[0021] a memory communicatively connected to the at least one processor; wherein,

[0022] The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the steps in the human-computer interaction method as described above.

[0023] In order to solve the above problem, the present invention also provides a computer-readable storage medium, in which at least one instruction is stored. When the at least one instruction is executed by a processor in an electronic device, the above human-computer interaction method is implemented.

[0024] The human-computer interaction method, device, electronic device and medium provided by the present invention obtain the current user voice data from the AI ​​dialogue of human-computer interaction, and input the user's current emotional features extracted from the user voice data and the user's personality features extracted from the text data converted from the user voice data into a preset emotion recognition model and a preset personality recognition model respectively, so as to respectively predict the probability of the user currently being in different types of emotions and the probability of predicting the user being in different types of personality, so as to respectively obtain the current user emotion type probability data and the user personality type probability data, and then fuse the user personality type probability data with the current user emotion type probability data to obtain an emotion-personality fusion type and a fusion type probability, and use the emotion-personality fusion type corresponding to the fusion type probability with the largest probability value as the result of judging the user's current emotion, and timely obtain the most matching emotion guidance scheme based on the result, and fuse the user's individual personality difference factors into the current emotion analysis of the user, so as to make the judgment of the user's current emotional state more accurate, and can timely and accurately give corresponding emotional guidance according to the individual emotional changes of the user, thereby improving the user experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0026] Figure 1 A schematic diagram of a flow chart of a human-computer interaction method provided by an embodiment of the present invention;

[0027] Figure 2 A schematic diagram of a module of a human-computer interaction device provided by an embodiment of the present invention;

[0028] Figure 3 A schematic diagram of the internal structure of an electronic device for implementing a human-computer interaction method provided by an embodiment of the present invention;

[0029] Figure 4A principle block diagram of a human-computer interaction method provided in one embodiment of the present invention.

[0030] The realization of the purpose, functional features and advantages of the present invention will be further explained in conjunction with embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION

[0031] It should be understood that the specific embodiments described herein are only used to explain the present invention, and are not used to limit the present invention.

[0032] Based on the problems existing in the above-mentioned prior art, the present invention mainly provides a human-computer interaction method, device, electronic device and medium, whose main purpose is to solve the problem in the prior art that emotion recognition based on AI dialogue generally judges the user's current emotional state from the user's voice interaction content or tone fluctuations, and it is difficult to fully and accurately capture and understand the emotional state of users with different language habits, and may misjudge the user's current emotional state, making it difficult to provide corresponding emotional guidance in time according to the user's current state.

[0033] Figure 1 A schematic diagram of a flow chart of a human-computer interaction method provided by an embodiment of the present invention; Figure 4 This is a principle block diagram of a human-computer interaction method provided by an embodiment of the present invention. The method can be executed by a device, and the device can be implemented by software and / or hardware.

[0034] Figure 1 Combination Figure 4 The human-computer interaction method is described in general. Figure 1 Combination Figure 4 As shown together, in this embodiment, the human-computer interaction method includes steps S110 to S160.

[0035] Step S110: Acquire current user voice data from the AI ​​dialogue of human-computer interaction.

[0036] Specifically, during the human-computer interaction process, the current user voice data is collected through a voice collection device, such as a miniature microphone, and the user voice data is used for subsequent emotional feature and personality feature extraction.

[0037] As an optional embodiment of the present invention, obtaining current user voice data from an AI dialogue of human-computer interaction includes:

[0038] Obtaining original user voice data within a preset time range from the AI ​​dialogue of human-computer interaction;

[0039] Pre-emphasize the high frequency part of the original user voice data to obtain emphasized voice data;

[0040] Performing frame division and windowing processing on the aggravated speech data to obtain time domain speech data;

[0041] Endpoint detection is performed on the time-domain voice data to select a valid voice segment from the time-domain voice data, and the valid voice segment is used as the current user voice data.

[0042] Specifically, in the field of artificial intelligence, companion robots, voice assistants, and service robots in various scenarios generally require AI dialogues during human-computer interaction. Human emotional changes do not occur suddenly. In order to accurately and timely detect user emotional changes, during human-computer interaction, user voice data within a preset time range is collected through voice collection equipment. For example, user voice data collected within 5 minutes, half an hour, or 1 hour (set as needed) is used as a basis for judging the user's current emotions, so as to timely judge the user's current emotions. The user voice data obtained by the voice collection device is original voice data. In order to improve the subsequent feature extraction processing, it is necessary to pre-process the original user voice data before feature extraction.

[0043] The high frequency part of the original user voice data is emphasized to remove the influence of lip radiation and glottal excitation during the human voice production process. The embodiment of the present invention preferably, but not limited to, uses a first-order FIR high-pass digital filter (finite impulse response digital filter) to complete the pre-emphasis processing of the original user voice data. After the voice signal is pre-emphasized, a flatter spectrum can be obtained, which is more conducive to spectrum analysis.

[0044] The characteristics of speech signals are usually relatively stable within a short time range of 10 to 30 ms, that is, speech signals have short-term stability. Therefore, it is necessary to perform frame and window processing on the aggravated speech data. Framing is to segment the speech signal in units of frames. Windowing can emphasize the speech waveform near the sampling and weaken the rest of the waveform. This embodiment preferably but not limited to the use of Hamming window, which has lower frequency resolution, lower side lobes and less spectrum leakage.

[0045] There may be invalid speech segments in the time domain speech data obtained by frame segmentation and windowing processing. Therefore, it is necessary to select valid speech segments from the time domain speech data to use the valid speech segments as the current user speech data. Among them, endpoint detection is an effective way to detect valid speech segments from a continuous speech stream. The starting point (front end point) of the valid speech and the end point (rear end point) of the valid speech are detected by endpoint detection. This embodiment preferably but not limited to uses a double threshold method, that is, combining an energy threshold and a zero crossing rate threshold to detect speech endpoints.

[0046] The operation is as follows: first, the possible speech segments are initially screened out by the energy threshold, and then the endpoints are further accurately determined within these speech segments using the zero-crossing rate threshold to improve the accuracy of endpoint detection. For example, a higher energy threshold T can be set. E1 and a lower energy threshold T E2 , and a zero-crossing rate threshold T Z When the signal energy E(n)<T E1 When the signal energy E(n) < T E2 When , it is marked as the possible end position of the speech. Then, in these possible speech segments, further judgment is made based on the zero crossing rate Z(n). When Z(n)>T Z , it is determined to be a valid speech segment.

[0047] Step S120, extracting the user's current emotional features from the user's voice data; and extracting the user's personality features from the text data by converting the user's voice data into text data.

[0048] Specifically, the user's current emotional characteristics can be extracted from the user's voice data, and the user's personality characteristics can be extracted through the conversation content, that is, the text data converted from speech. The user's current emotional characteristics and user personality characteristics extracted from the dual modalities of hearing and text can be used to analyze and understand the user's current emotional state from different aspects, avoiding misjudgment of the user's current emotional state due to individual personality differences.

[0049] As an optional embodiment of the present invention, extracting the user's current emotion feature from the user's voice data includes:

[0050] Performing Fourier transform processing on each frame of the voice signal in the user voice data to obtain the user voice amplitude spectrum;

[0051] Perform spectrum conversion processing on the amplitude spectrum of the user's speech through a Mel frequency bandpass filter group to obtain a Mel spectrum;

[0052] According to the Mel spectrum, the logarithmic energy of the Mel spectrum obtained by each Mel frequency bandpass filter is calculated;

[0053] All logarithmic energies are processed by discrete cosine transform to obtain Mel-frequency cepstral coefficient parameters, which are used as the user's current emotional features.

[0054] Specifically, Mel Frequency Cepstral Coefficients (MFCC) are features based on the spectrum. Using frequency-domain-based feature parameters for emotion recognition can achieve better performance. It is proposed based on the characteristics of the human auditory system and can simulate the human ear's perception of speech of different frequencies. First, use discrete Fourier transform (DFT) to convert time domain signals to frequency domain signals, perform DFT on each frame of the voice signal in the user voice data after windowing, obtain the user voice amplitude spectrum, and then convert the linear frequency of the user voice amplitude spectrum into Mel frequency. Then, evenly divide the triangular filters (preferably but not limited to 30 triangular filters) in the Mel frequency domain, calculate the energy of the signal on these filters, and then perform logarithmic operations. Finally, perform discrete cosine transform (DTC) on the above logarithmic operation results to obtain Mel frequency cepstral coefficient parameters, and use the Mel frequency cepstral coefficient parameters as the user's current emotional features.

[0055] As an optional embodiment of the present invention, the user voice data is converted into text data to extract the user personality characteristics from the text data, including:

[0056] Convert user voice data into text data;

[0057] Preprocess the text data to obtain text sentences with complete semantics and a sentence length that meets the preset sentence length;

[0058] Based on the Big Five personality domain dictionary, personality clue words are extracted from text sentences to obtain a personality clue word set;

[0059] The words in the character clue word set are spliced ​​into character clue sentences according to their positional relationship in the text data;

[0060] Calculating the context representation of each word in the text sentence to obtain a text feature vector matrix; and calculating the context representation of each word in the character clue sentence to obtain a character feature vector matrix;

[0061] Through the conditional semantic fusion technology, the word vectors in the personality feature vector matrix are fused into the word vectors in the text feature vector matrix as external conditions, thereby obtaining the personality condition fusion matrix;

[0062] The personality condition fusion matrix is ​​used as the user personality feature.

[0063] Specifically, the user's voice data is converted into text data by any speech-to-text conversion tool in the prior art, and then the text data is preprocessed to obtain a text sentence with complete semantics and a sentence length that meets the preset sentence length. For example, data with missing, repeated, garbled, meaningless, etc. are deleted and simplified, and data with a sentence length that is too short to effectively express personality and a sentence length that is too long to cause redundant expression is removed to form a sentence C, that is, a text sentence.

[0064] Based on the Big Five personality domain dictionary combined with the collected AI dialogue scenes, the user's personality traits are annotated through an expert database with a background in psychology and personality analysis research. To ensure the accuracy of the evaluation, annotation is performed only when there is clear evidence in the text indicating the corresponding Big Five personality dimensions. Based on the annotation, personality clue words (gerunds and / or adjectives) are extracted from the text sentence (sentence C) to form a personality clue word set A = {A1, A2, ..., Ai}, where i represents the number of personality clue words. The words in the personality clue word set are spliced ​​into personality clue sentences P according to their positional relationship in the text data. The pre-trained model BERT can be used to vectorize sentences C and personality clue sentences P respectively to calculate the context representation of each word and obtain the corresponding word vector matrix, that is, the text feature vector matrix H C and the personality trait vector matrix H P Through conditional semantic fusion technology, the personality feature vector matrix H P The word vectors in are fused into the text feature vector matrix H as external conditions C ’s word vector to obtain the personality condition fusion matrix; the personality condition fusion matrix is ​​used as the user’s personality feature.

[0065] As an optional embodiment of the present invention, the word vectors in the personality feature vector matrix are fused into the word vectors in the text feature vector matrix as external conditions through conditional semantic fusion technology, thereby obtaining a personality condition fusion matrix, including:

[0066] The standard deviation of each word vector in the text feature vector matrix is ​​normalized to obtain the text word vector matrix; the text feature vector matrix is ​​expressed as:

[0067]

[0068] Among them, H C represents the text feature vector matrix, represents the text word feature vector, represents n text word vectors, Represents the text word separation vector;

[0069] The conditional fusion function is used to conditionally fuse the personality clue feature vector in the personality feature vector matrix with the text word vector matrix in a semantically interactive manner, thereby obtaining a personality conditional fusion matrix; wherein the personality feature vector matrix is ​​expressed as:

[0070] in,

[0071] H P represents the personality trait vector matrix, represents the personality clue feature vector, represents n personality clue vectors, represents the personality clue separation vector;

[0072] The formula of the conditional fusion function is:

[0073]

[0074] Among them, H m represents the character condition fusion matrix, CLN represents the conditional fusion function, H normal c represents the text word vector matrix, represents the personality clue feature vector, γ p Indicates H normal c The conditional gain vector, β p Indicates H normal c The conditional bias vector, W γ Represents the gain effect control matrix, W β represents the bias effect control matrix, b γ Represents the bias value of the gain effect control matrix, b β Indicates the bias value of the bias effect control matrix.

[0075] Specifically, the text feature vector matrix H C and the personality trait vector matrix H P Respectively expressed as:

[0076]

[0077] Among them, H C represents the text feature vector matrix, represents the text word feature vector, represents n text word vectors, Represents the text word separation vector; H P represents the personality trait vector matrix, represents the personality clue feature vector, represents n personality clue vectors, represents the personality clue separation vector;

[0078] The personality feature vector matrix H is transformed into P The vector containing personality clues is used as an external condition to merge into the feature vector of the text. First, the text feature vector matrix H C Each word vector in is normalized by standard deviation, and the formula is as follows:

[0079]

[0080] Among them, H normal c Represents the text word vector matrix, μ and σ represent the mean and variance of each word vector, i represents the ordinal number of the word vector in the matrix, and n represents the number of word vectors in the matrix.

[0081] The conditional fusion function CLN is used to transform the personality clue feature vector And the text word vector matrix H after standard deviation normalization normal c Get the conditional fusion matrix H of the dynamic semantic interaction between text and personality clues m .H normal c The conditional gain vector γ p and the conditional bias vector β p , respectively, by the gain effect control matrix W γ and the bias effect control matrix W β Characteristic clues vector Multiply them together and add them to their respective bias values ​​b γ and b β Add them together to get, where W γ , W β 、b γ and b β All of them can be obtained in the dynamic learning during model training and can be regarded as known. Finally, the personality condition fusion matrix is ​​obtained.

[0082] Step S130: obtaining the current user's emotion type probability data by inputting the user's current emotion characteristics into a preset emotion recognition model; and obtaining the user's personality type probability data by inputting the user's personality characteristics into a preset personality recognition model; wherein the current user's emotion type probability data includes the emotion type and the emotion probability corresponding to the emotion type; and the user's personality type probability data includes the personality type and the personality probability corresponding to the personality type.

[0083] Specifically, by inputting the user's current emotional characteristics into the preset emotion recognition model and the user's personality characteristics into the preset emotion recognition model and the preset personality recognition model respectively, the probability that the user currently has different emotions and the probability that the user has different personalities can be obtained, so as to facilitate the subsequent fusion of emotions and personalities.

[0084] As an optional embodiment of the present invention, the current emotion characteristics of the user are input into a preset emotion recognition model to obtain the current user emotion type probability data, including:

[0085] Inputting the user's current emotional characteristics into a preset emotion recognition model; wherein the preset emotion recognition model is provided with probability prediction modules of different types of emotions;

[0086] The probability prediction modules of different types of emotions are used to calculate the probability of the corresponding emotion type for the user's current emotion characteristics, and different emotion types and emotion probabilities corresponding to the emotion types are obtained;

[0087] Different emotion types and emotion probabilities corresponding to the emotion types are used as the current user emotion type probability data.

[0088] Specifically, the preset emotion recognition model is preferably obtained by pre-training a neural network, which preferably includes, but is not limited to, a convolution layer for obtaining local features, a pooling layer for dimensionality reduction processing of local features obtained by the convolution layer, and a fully connected layer for users to integrate local features after dimensionality reduction processing of the pooling layer, etc. Since the above is the basic structure of the neural network, it will not be repeated here. The preset emotion recognition model in the embodiment of the present invention is internally provided with probability prediction modules of different types of emotions. After the activation function in each module is trained with sample data, the comprehensive features provided by the fully connected layer can be calculated and processed to obtain the probabilities of different emotion types. That is, after the user's current emotional features are input into the preset emotion recognition model, the convolution layer, the pooling layer and the fully connected layer are sequentially passed to obtain the comprehensive features extracted based on the user's current emotional features, and the comprehensive features are used as the input of the probability prediction modules of different types of emotions. After the probability prediction modules of different types of emotions are calculated, different emotion types and emotion probabilities corresponding to the emotion types are obtained, that is, the current user emotion type probability data. For example, after inputting the user's current emotional characteristics into the preset emotion recognition model, the output result is that the probability that the user's current emotion is happy is 70%, the probability of anger is 10%, the probability of sadness is 15%, the probability of sorrow is 13%, etc.

[0089] As an optional embodiment of the present invention, the user's personality characteristics are input into a preset personality recognition model to obtain the user's personality type probability data, including:

[0090] Inputting the user's personality characteristics into a preset personality recognition model; wherein the preset personality recognition model is provided with probability prediction modules of different personality types;

[0091] The probability prediction modules of different personality types are used to calculate the probability of the corresponding personality types of the user's personality characteristics, and different personality types and the personality probabilities corresponding to the personality types are obtained;

[0092] Different personality types and personality probabilities corresponding to the personality types are used as user personality type probability data.

[0093] Specifically, the preset personality recognition model is preferably obtained by pre-training a neural network. The feature vector at the CLS position in the personality condition fusion matrix is ​​used as the global representation of the sentence. First, the conditional fusion matrix H m Take the average of the word vector dimensions to form the conditional fusion vector h e , input to the fully connected network responsible for personality recognition and output personality recognition vector h o , and then use the pooling function trained in the probability prediction module of different personality types to h o The personality probability calculation is performed to obtain different personality types and the probabilities corresponding to the personality types. For example, after inputting the user's personality characteristics into the preset emotion recognition model, the output result obtained is that the probability that the user's personality is cheerful is 80%, the probability that the user's personality is melancholic is 10%, etc.

[0094] Step S140: fusing the current user's emotion type probability data with the user's personality type probability data to obtain an emotion-personality fusion type and a corresponding fusion type probability.

[0095] Specifically, the current user emotion type probability data is fused with the user personality type probability data to obtain a personality-emotion combination type and a corresponding probability, so that a corresponding interaction plan can be given based on personality and emotion. For example, the user's current emotion is happy and the user's personality is cheerful, and the corresponding emotion probability of being happy and the probability of being cheerful are obtained.

[0096] As an optional embodiment of the present invention, the current user emotion type probability data and the user personality type probability data are fused to obtain the emotion personality fusion type and the corresponding fusion type probability, including:

[0097] Based on Bayes' theorem, the emotion probability corresponding to each type of emotion in the current user's emotion type probability data is multiplied by the personality probability corresponding to each type of personality in the user's personality type probability data, and then multiplied by the prior probability of the preset emotion recognition model and the prior probability of the preset personality recognition model to obtain the emotion-personality fusion type and the corresponding fusion type probability.

[0098] Specifically, using Bayesian theorem to fuse the prediction results of the two models will more accurately reflect the relationship between personality traits and emotional states, thereby better realizing the fusion model's comprehensive judgment of emotional states under the influence of individual personality. This will lead to more reasonable AI dialogue content decisions and tone corrections.

[0099] Step S150: Select the emotion and personality fusion type corresponding to the fusion type probability with the largest probability value as the user's current emotion and personality fusion type.

[0100] Specifically, the emotion-personality fusion type corresponding to the fusion type probability with the largest probability value is the most accurate current emotional state of the user, so the user can be given a corresponding interaction plan based on the emotion-personality fusion type.

[0101] Step S160: Obtain a user emotion guidance scheme that matches the user's current emotion-personality fusion type from a preset user emotion guidance scheme library, and perform human-computer interaction with the user based on the user emotion guidance scheme.

[0102] Specifically, a user emotion guidance program library is preset, and corresponding user emotion guidance programs are set for each personality type of user under different emotions, including tone of voice during interaction, automatically provided music, etc.

[0103] like Figure 2 As shown, the present invention provides a human-computer interaction device 200, which can be installed in an electronic device. According to the functions to be implemented, the human-computer interaction device 200 may include: a speech acquisition module 210, a multimodal feature extraction module 220, a probability estimation module 230, a fusion module 240, a fusion type determination module 250 and a human-computer interaction execution module 260. The unit of the present invention can also be called a module, which refers to a series of computer program segments that can be executed by an electronic device processor and can complete fixed functions, which are stored in the memory of the electronic device.

[0104] In this embodiment, the functions of each module / unit are as follows:

[0105] The voice acquisition module 210 is used to acquire the current user voice data from the AI ​​dialogue of human-computer interaction;

[0106] The multimodal feature extraction module 220 is used to extract the user's current emotional features from the user's voice data; and to extract the user's personality features from the text data by converting the user's voice data into text data;

[0107] The probability estimation module 230 is used to obtain the current user emotion type probability data by inputting the user's current emotion characteristics into a preset emotion recognition model; and to obtain the user's personality type probability data by inputting the user's personality characteristics into a preset personality recognition model; wherein the current user emotion type probability data includes the emotion type and the emotion probability corresponding to the emotion type; the user's personality type probability data includes the personality type and the personality probability corresponding to the personality type;

[0108] A fusion module 240 is used to fuse the current user emotion type probability data with the user personality type probability data to obtain an emotion personality fusion type and a corresponding fusion type probability;

[0109] A fusion type determination module 250 is used to select the emotion and personality fusion type corresponding to the fusion type probability with the largest probability value as the user's current emotion and personality fusion type;

[0110] The human-computer interaction execution module 260 is used to obtain a user emotion guidance scheme that matches the user's current emotion-character fusion type from a preset user emotion guidance scheme library, and perform human-computer interaction with the user based on the user emotion guidance scheme.

[0111] The human-computer interaction device 200 of the present invention obtains the current user voice data from the AI ​​dialogue of human-computer interaction, and inputs the user's current emotional features extracted from the user voice data and the user's personality features extracted from the text data converted from the user voice data into a preset emotion recognition model and a preset personality recognition model respectively, thereby respectively predicting the probability of the user currently being in different types of emotions and predicting the probability of the user being in different types of personalities, so as to respectively obtain the current user emotion type probability data and the user personality type probability data, and fuse the user personality type probability data with the current user emotion type probability data to obtain an emotion-personality fusion type and a fusion type probability, and use the emotion-personality fusion type corresponding to the fusion type probability with the largest probability value as the result of the user's current emotion judgment, and timely obtain the most matching emotion guidance plan based on the result, and fuse the user's individual personality difference factors into the user's current emotion analysis, so as to make the judgment of the user's current emotional state more accurate, and can timely and accurately give corresponding emotional guidance according to the user's individual emotional changes, thereby improving the user experience.

[0112] like Figure 3 As shown, the present invention provides an electronic device 3 for a human-computer interaction method.

[0113] The electronic device 3 may include a processor 30, a memory 31 and a bus, and may also include a computer program stored in the memory 31 and executable on the processor 30, such as a human-computer interaction program 32. The memory 31 may also include both an internal storage unit of the human-computer interaction device and an external storage device. The memory 31 may be used not only to store application software and various data installed, such as codes of the human-computer interaction program, but also to temporarily store data that has been output or is to be output.

[0114] The memory 31 includes at least one type of readable storage medium, and the readable storage medium includes a flash memory, a mobile hard disk, a multimedia card, a card-type memory (e.g., SD or DX memory, etc.), a magnetic memory, a disk, an optical disk, etc. The memory 31 may be an internal storage unit of the electronic device 3 in some embodiments, such as a mobile hard disk of the electronic device 3. The memory 31 may also be an external storage device of the electronic device 3 in other embodiments, such as a plug-in mobile hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (SecureDigital, SD) card, a flash card (Flash Card), etc. equipped on the electronic device 3. Further, the memory 31 may also include both an internal storage unit of the electronic device 3 and an external storage device. The memory 31 may not only be used to store application software and various types of data installed in the electronic device 3, such as a human-computer interaction method code, etc., but may also be used to temporarily store data that has been output or is to be output.

[0115] The processor 30 may be composed of an integrated circuit in some embodiments, for example, a single packaged integrated circuit, or a plurality of packaged integrated circuits with the same or different functions, including one or more central processing units (CPUs), microprocessors, digital processing chips, graphics processors, and combinations of various control chips, etc. The processor 30 is the control core (Control Unit) of the electronic device, and uses various interfaces and lines to connect various components of the entire electronic device, and executes or executes programs or modules (such as human-computer interaction programs, etc.) stored in the memory 31, and calls data stored in the memory 31 to execute various functions of the electronic device 3 and process data.

[0116] The bus may be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus, etc. The bus may be divided into an address bus, a data bus, a control bus, etc. The bus is configured to realize connection and communication between the memory 31 and at least one processor 30, etc.

[0117] Figure 3 Only an electronic device with components is shown, and those skilled in the art will understand that Figure 3 The structure shown does not constitute a limitation on the electronic device 3, and may include fewer or more components than shown in the figure, or combine certain components, or arrange the components differently.

[0118] For example, although not shown, the electronic device 3 may also include a power source (such as a battery) for supplying power to each component. Preferably, the power source may be logically connected to the at least one processor 30 through a power management system, so that the power management system can realize functions such as charging management, discharging management, and power consumption management. The power source may also include any components such as one or more DC or AC power sources, recharging systems, power failure detection circuits, power converters or inverters, and power status indicators. The electronic device 3 may also include a variety of sensors, Bluetooth modules, Wi-Fi modules, etc., which will not be repeated here.

[0119] Furthermore, the electronic device 3 may also include a network interface. Optionally, the network interface may include a wired interface and / or a wireless interface (such as a WI-FI interface, a Bluetooth interface, etc.), which is generally used to establish a communication connection between the electronic device 3 and other electronic devices.

[0120] Optionally, the electronic device 3 may further include a user interface, which may be a display, an input unit (such as a keyboard), or a standard wired interface or a wireless interface. Optionally, in some embodiments, the display may be an LED display, a liquid crystal display, a touch-sensitive liquid crystal display, and an OLED (Organic Light-Emitting Diode) touch device. The display may also be appropriately referred to as a display screen or a display unit, which is used to display information processed in the electronic device 3 and to display a visual user interface.

[0121] It should be understood that the embodiments are for illustration purposes only and the scope of the invention application is not limited by this structure.

[0122] The human-computer interaction program 32 stored in the memory 31 of the electronic device 3 is a combination of multiple instructions. When running in the processor 30, it can achieve:

[0123] Step S110: obtaining current user voice data from the AI ​​dialogue of human-computer interaction;

[0124] Step S120, extracting the user's current emotional features from the user's voice data; and converting the user's voice data into text data to extract the user's personality features from the text data;

[0125] Step S130, inputting the current emotion characteristics of the user into a preset emotion recognition model to obtain the current user emotion type probability data; and inputting the user personality characteristics into a preset personality recognition model to obtain the user personality type probability data; wherein the current user emotion type probability data includes the emotion type and the emotion probability corresponding to the emotion type; the user personality type probability data includes the personality type and the personality probability corresponding to the personality type;

[0126] Step S140, fusing the current user's emotion type probability data with the user's personality type probability data to obtain an emotion-personality fusion type and a corresponding fusion type probability;

[0127] Step S150, selecting the emotion and personality fusion type corresponding to the fusion type probability with the largest probability value as the user's current emotion and personality fusion type;

[0128] Step S160: Obtain a user emotion guidance scheme that matches the user's current emotion-personality fusion type from a preset user emotion guidance scheme library, and perform human-computer interaction with the user based on the user emotion guidance scheme.

[0129] Specifically, the specific implementation method of the processor 30 for the above instructions can refer to Figure 1 The description of the relevant steps in the corresponding embodiments will not be repeated here.

[0130] Furthermore, if the module / unit integrated in the electronic device 3 is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. The computer-readable medium may include: any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a mobile hard disk, a magnetic disk, an optical disk, a computer memory, and a read-only memory (ROM).

[0131] An embodiment of the present invention further provides a computer-readable storage medium, which may be non-volatile or volatile, and stores a computer program, which, when executed by a processor, implements:

[0132] Step S110: obtaining current user voice data from the AI ​​dialogue of human-computer interaction;

[0133] Step S120, extracting the user's current emotional features from the user's voice data; and converting the user's voice data into text data to extract the user's personality features from the text data;

[0134] Step S130, inputting the current emotion characteristics of the user into a preset emotion recognition model to obtain the current user emotion type probability data; and inputting the user personality characteristics into a preset personality recognition model to obtain the user personality type probability data; wherein the current user emotion type probability data includes the emotion type and the emotion probability corresponding to the emotion type; the user personality type probability data includes the personality type and the personality probability corresponding to the personality type;

[0135] Step S140, fusing the current user's emotion type probability data with the user's personality type probability data to obtain an emotion-personality fusion type and a corresponding fusion type probability;

[0136] Step S150, selecting the emotion and personality fusion type corresponding to the fusion type probability with the largest probability value as the user's current emotion and personality fusion type;

[0137] Step S160: Obtain a user emotion guidance scheme that matches the user's current emotion-personality fusion type from a preset user emotion guidance scheme library, and perform human-computer interaction with the user based on the user emotion guidance scheme.

[0138] Specifically, the specific implementation method when the computer program is executed by the processor can refer to the description of the relevant steps in the human-computer interaction method in the embodiment, which will not be repeated here.

[0139] In the several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses and methods can be implemented in other ways. For example, the device embodiments described above are only schematic, for example, the division of the modules is only a logical function division, and there may be other division methods in actual implementation.

[0140] The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0141] In addition, each functional module in each embodiment of the present invention may be integrated into one processing unit, each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of hardware plus software functional modules.

[0142] It is obvious to those skilled in the art that the present invention is not limited to the details of the above exemplary embodiments, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention.

[0143] Therefore, no matter from which point of view, the embodiments should be regarded as illustrative and non-restrictive, and the scope of the present invention is limited by the appended claims rather than the above description, so it is intended that all changes falling within the meaning and scope of the equivalent elements of the claims are included in the present invention. Any attached figure mark in the claims should not be regarded as limiting the claims involved.

[0144] In addition, it is clear that the word "comprising" does not exclude other units or steps, and the singular does not exclude the plural. Multiple units or devices stated in a device claim can also be implemented by one unit or device through software or hardware. The second and other words are used to indicate names, but not to indicate any particular order.

[0145] Finally, it should be noted that the above embodiments are only used to illustrate the technical solution of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solution of the present invention can be modified or replaced by equivalents without departing from the spirit and scope of the technical solution of the present invention.

Claims

1. A human-computer interaction method, characterized in that: The steps include: Obtain current user voice data from the AI ​​dialogue of human-computer interaction; Extracting the user's current emotional features from the user's voice data; and extracting the user's personality features from the text data by converting the user's voice data into text data; By inputting the current emotion characteristics of the user into a preset emotion recognition model, the current user emotion type probability data is obtained; and obtaining user personality type probability data by inputting the user personality characteristics into a preset personality recognition model; wherein the current user emotion type probability data includes an emotion type and an emotion probability corresponding to the emotion type; and the user personality type probability data includes a personality type and a personality probability corresponding to the personality type; Fusing the current user emotion type probability data with the user personality type probability data to obtain an emotion-personality fusion type and a corresponding fusion type probability; Select the emotion-personality fusion type corresponding to the fusion type probability with the largest probability value as the user's current emotion-personality fusion type; A user emotion guidance scheme matching the current emotion-personality fusion type of the user is obtained from a preset user emotion guidance scheme library, and human-computer interaction is performed with the user based on the user emotion guidance scheme.

2. The human-computer interaction method according to claim 1, characterized in that: The obtaining of current user voice data from the AI ​​dialogue of human-computer interaction includes: Obtaining original user voice data within a preset time range from the AI ​​dialogue of human-computer interaction; Pre-emphasize the high frequency part of the original user voice data to obtain emphasized voice data; Performing frame division and windowing processing on the aggravated speech data to obtain time domain speech data; Endpoint detection is performed on the time-domain voice data to select a valid voice segment from the time-domain voice data, and the valid voice segment is used as current user voice data.

3. The human-computer interaction method according to claim 1, characterized in that: The step of extracting the user's current emotion feature from the user's voice data comprises: Performing Fourier transform processing on each frame of the voice signal in the user voice data to obtain a user voice amplitude spectrum; Performing spectrum conversion processing on the amplitude spectrum of the user's speech through a Mel frequency bandpass filter group to obtain a Mel spectrum; According to the Mel spectrum, calculating the logarithmic energy of the Mel spectrum obtained by each Mel frequency bandpass filter; All the logarithmic energies are processed by discrete cosine transformation to obtain Mel-frequency cepstral coefficient parameters, and the Mel-frequency cepstral coefficient parameters are used as the current emotion features of the user.

4. The human-computer interaction method according to claim 1, characterized in that: The method of converting the user voice data into text data to extract the user personality characteristics from the text data includes: Converting the user voice data into text data; Preprocessing the text data to obtain a text sentence with complete semantics and a sentence length that meets a preset sentence length; Based on the Big Five personality domain dictionary, extract personality clue words from the text sentence to obtain a personality clue word set; splicing the words in the character clue word set into character clue sentences according to their positional relationships in the text data; Calculating the contextual representation of each word in the text sentence in the text sentence to obtain a text feature vector matrix; and calculating the contextual representation of each word in the character clue sentence in the character clue sentence to obtain a character feature vector matrix; By using conditional semantic fusion technology, the word vectors in the personality feature vector matrix are fused into the word vectors in the text feature vector matrix as external conditions, thereby obtaining a personality condition fusion matrix; The personality condition fusion matrix is ​​used as the user personality feature.

5. The human-computer interaction method according to claim 4, characterized in that: The conditional semantic fusion technology is used to fuse the word vectors in the personality feature vector matrix as external conditions into the word vectors in the text feature vector matrix, thereby obtaining a personality condition fusion matrix, including: Each word vector in the text feature vector matrix is ​​subjected to standard deviation normalization to obtain a text word vector matrix; the text feature vector matrix is ​​expressed as: Among them, H C represents the text feature vector matrix, represents the text word feature vector, represents n text word vectors, Represents the text word separation vector; The conditional fusion function is used to conditionally fuse the character clue feature vector in the character feature vector matrix with the text word vector matrix in a semantically interactive manner, thereby obtaining a character conditional fusion matrix; wherein the character feature vector matrix is ​​represented as: in, H P represents the personality trait vector matrix, represents the personality clue feature vector, represents n personality clue vectors, represents the personality clue separation vector; The formula of the conditional fusion function is: Among them, H m represents the character condition fusion matrix, CLN represents the conditional fusion function, H normal c represents the text word vector matrix, represents the personality clue feature vector, γ p Indicates H normal c The conditional gain vector, β p Indicates H normal c The conditional bias vector, W γ Represents the gain effect control matrix, W β represents the bias effect control matrix, b γ Represents the bias value of the gain effect control matrix, b β Indicates the bias value of the bias effect control matrix.

6. The human-computer interaction method according to claim 1, characterized in that: The step of inputting the current emotion characteristics of the user into a preset emotion recognition model to obtain the current user emotion type probability data comprises: Inputting the current emotional characteristics of the user into a preset emotion recognition model; wherein the preset emotion recognition model is provided with probability prediction modules of different types of emotions; The probability prediction modules of different types of emotions are used to calculate the probability of the corresponding emotion types for the current emotion characteristics of the user, so as to obtain different emotion types and emotion probabilities corresponding to the emotion types; The different emotion types and the emotion probabilities corresponding to the emotion types are used as current user emotion type probability data.

7. The human-computer interaction method according to claim 1, characterized in that: The step of inputting the user's personality characteristics into a preset personality recognition model to obtain the user's personality type probability data comprises: Inputting the user's personality characteristics into a preset personality recognition model; wherein the preset personality recognition model is provided with probability prediction modules of different personality types; The probability prediction modules of different personality types are used to calculate the probability of the corresponding personality types of the personality characteristics of the user, so as to obtain different personality types and personality probabilities corresponding to the personality types; The different personality types and the personality probabilities corresponding to the personality types are used as user personality type probability data.

8. The human-computer interaction method according to claim 1, characterized in that: The fusing the current user emotion type probability data with the user personality type probability data to obtain an emotion personality fusion type and a corresponding fusion type probability includes: Based on Bayes' theorem, the emotion probability corresponding to each type of emotion in the current user's emotion type probability data is multiplied by the personality probability corresponding to each type of personality in the user's personality type probability data, and then multiplied by the prior probability of the preset emotion recognition model and the prior probability of the preset personality recognition model to obtain the emotion-personality fusion type and the corresponding fusion type probability.

9. A human-computer interaction device, characterized in that: The device comprises: The voice acquisition module is used to obtain the current user voice data from the AI ​​dialogue of human-computer interaction; A multimodal feature extraction module, used to extract the user's current emotional features from the user's voice data; and to extract the user's personality features from the text data by converting the user's voice data into text data; A probability estimation module, configured to obtain current user emotion type probability data by inputting the current user emotion characteristics into a preset emotion recognition model; and to obtain user personality type probability data by inputting the user personality characteristics into a preset personality recognition model; wherein the current user emotion type probability data includes an emotion type and an emotion probability corresponding to the emotion type; and the user personality type probability data includes a personality type and a personality probability corresponding to the personality type; A fusion module, used for fusing the current user emotion type probability data with the user personality type probability data to obtain an emotion personality fusion type and a corresponding fusion type probability; A fusion type determination module is used to select the emotion and personality fusion type corresponding to the fusion type probability with the largest probability value as the user's current emotion and personality fusion type; The human-computer interaction execution module is used to obtain a user emotion guidance scheme that matches the current emotion and personality fusion type of the user from a preset user emotion guidance scheme library, and perform human-computer interaction with the user based on the user emotion guidance scheme.

10. An electronic device, characterized in that: The electronic device comprises: at least one processor; and, a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the steps in the human-computer interaction method as described in any one of claims 1 to 8.

11. A computer-readable storage medium storing at least one instruction, characterized in that: When the at least one instruction is executed by a processor in an electronic device, the human-computer interaction method as described in any one of claims 1 to 8 is implemented.

Citation Information

Patent Citations

  • Multi-mode intelligent emotion sensing system

    CN107220591A

  • Speech emotion recognition method and device and related equipment

    CN110751943A

  • Emotion recognition method and device, equipment and medium

    CN115631745A

  • AI-based voice emotion recognition model training method

    CN117524262A

  • Intelligent outbound method, device and equipment based on emotion recognition and storage medium

    CN117690436A