Emotion prompting method and device of vehicle-mounted system, electronic equipment, storage medium and vehicle

By obtaining image and voice information in the on-board system, emotional recognition and fusion are performed, and voice prompts are output, the problem of driver emotions affecting driving safety is solved, and driving safety and emotional judgment are improved.

CN120014608APending Publication Date: 2025-05-16CHONGQING WUTONG CAR LINK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411881943.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-19
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

In the prior art, the driver's mood changes affect driving safety, and the prior art is difficult to accurately identify the driver's emotions.

Method used

By obtaining the image acquisition information and voice acquisition information of the on-board system, emotional recognition is performed separately, image emotion vectors and voice emotion vectors are obtained, and emotional fusion parameters are determined based on vector distance, and image and voice emotion information are fused, and voice prompt information is output.

Benefits of technology

It realizes accurate recognition and integration of driver emotions, timely outputs voice prompts, improves driving safety, and reduces the error in emotional judgment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014608A_ABST
    Figure CN120014608A_ABST
Patent Text Reader

Abstract

The invention relates to an emotion prompting method and device for a vehicle-mounted system, electronic equipment, a storage medium and a vehicle, and the method comprises the steps: obtaining image collection information and voice collection information of the vehicle-mounted system, determining the image information and voice information of a vehicle driver according to the image collection information and the voice collection information, and carrying out the emotion prompting of the vehicle-mounted system; the method comprises the steps of acquiring image information and voice information, performing emotion recognition according to the image information and the voice information to obtain an image emotion vector and a voice emotion vector, determining an emotion fusion parameter according to a vector distance between the image emotion vector and the voice emotion vector, and fusing the image emotion vector and the voice emotion vector based on the emotion fusion parameter to obtain an emotion fusion result. Target emotion information corresponding to the vehicle driver is obtained, and then voice prompt information is output based on the target emotion information; therefore, voice prompt can be given to the vehicle driver in time, the problem that driving safety is affected due to the fact that the driver is affected by emotion in the prior art is solved, and the driving safety can be effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of automobile technology, and in particular to an emotion prompting method, device, electronic equipment, storage medium and vehicle for an in-vehicle system. Background Art

[0002] With the development of vehicles, the scenes and frequency of people driving are also increasing. Vehicle drivers will be affected by many factors during driving, which will lead to emotional changes, and different emotions will also have different degrees of impact on vehicle drivers' driving behavior. For example, when a driver drives for a long time and becomes tired, he will lose concentration and reduce alertness, which will affect driving safety. Or when the driver is angry for some reason, his judgment ability will decrease and he will easily have excessive driving behavior, which will affect driving safety. It can be seen that when the driver is in certain emotions, his mind will be distracted, which will reduce his alertness and judgment ability during driving, leading to problems that affect driving safety. Summary of the invention

[0003] The present application provides an emotion prompting method, device, electronic device, storage medium and vehicle for an in-vehicle system to solve the problem in the existing related technologies that driving safety is affected because the driver is affected by emotions.

[0004] In a first aspect, the present application provides an emotion prompting method for an in-vehicle system, comprising:

[0005] Obtain image acquisition information and voice acquisition information from the vehicle-mounted system;

[0006] Determining the image information of the vehicle driver and the voice information of the vehicle driver according to the image acquisition information and the voice acquisition information;

[0007] Performing emotion recognition based on the image information and the voice information respectively to obtain an image emotion vector and a voice emotion vector;

[0008] Determining an emotion fusion parameter according to a vector distance between the image emotion vector and the speech emotion vector;

[0009] Based on the emotion fusion parameter, the image emotion vector and the voice emotion vector are fused to obtain target emotion information corresponding to the vehicle driver;

[0010] Output voice prompt information based on the target emotion information.

[0011] Optionally, determining the emotion fusion parameter according to the vector distance between the image emotion vector and the speech emotion vector includes:

[0012] Determining a first fusion parameter according to a vector distance between the image emotion vector and the speech emotion vector, wherein the first fusion parameter is exponentially related to the vector distance;

[0013] Based on the first fusion parameter, a parameter is calculated in combination with a preset fusion constant threshold to obtain a second fusion parameter;

[0014] The emotion fusion parameter is determined according to the first fusion parameter and the second fusion parameter.

[0015] Optionally, the fusing the image emotion vector and the voice emotion vector based on the emotion fusion parameter to obtain target emotion information corresponding to the vehicle driver includes:

[0016] Extracting a first fusion parameter and a second fusion parameter from the emotion fusion parameters;

[0017] Calculating the speech emotion vector and the first fusion parameter to obtain a first fusion vector;

[0018] Calculating the image emotion vector and the second fusion parameter to obtain a second fusion vector;

[0019] Combining the first fusion vector and the second fusion vector to obtain a target fusion vector;

[0020] The target emotion information is determined according to the target fusion vector.

[0021] Optionally, combining the first fusion vector and the second fusion vector to obtain a target fusion vector includes:

[0022] extracting at least one first emotion vector from the first fusion vector, and determining a first emotion category corresponding to the first emotion category vector;

[0023] extracting at least one second emotion vector from the second fusion vector, and determining a second emotion category corresponding to the second emotion category vector;

[0024] In the case where the first emotion type matches the second emotion type, combining the first emotion vector and the second emotion vector to obtain a target emotion vector;

[0025] The target fusion vector is obtained according to the target emotion vector.

[0026] Optionally, after performing emotion recognition based on the image information and the voice information to obtain an image emotion vector and a voice emotion vector, the method further includes:

[0027] Determining an image emotion identifier corresponding to the image emotion vector;

[0028] Determining a speech emotion identifier corresponding to the speech emotion vector;

[0029] In the case where the image emotion identifier and the voice emotion identifier do not belong to preset opposing emotions, the step of determining the emotion fusion parameter based on the vector distance between the image emotion vector and the voice emotion vector is performed.

[0030] Optionally, after determining the speech emotion identifier corresponding to the speech emotion vector, the method further includes:

[0031] When the image emotion identifier and the voice emotion identifier belong to preset opposing emotions, obtaining preset neutral emotion information;

[0032] The neutral emotion information is determined as the target emotion information.

[0033] Optionally, outputting voice prompt information based on the target emotion information includes:

[0034] Determine the emotion prompt text and emotion level corresponding to the target emotion information;

[0035] Obtaining tone information corresponding to the emotion level;

[0036] Combining the tone information with the emotion prompt text to generate the voice prompt information;

[0037] Output is performed based on the voice prompt information.

[0038] Optionally, outputting voice prompt information based on the target emotion information includes:

[0039] Get vehicle status information;

[0040] When the vehicle status information belongs to a preset status, pausing the output of the voice prompt information;

[0041] When the vehicle status information does not belong to a preset status, the voice prompt information is continuously outputted.

[0042] In a second aspect, the present application provides an emotion prompting device for a vehicle-mounted system, comprising:

[0043] An acquisition module, used to acquire image acquisition information and voice acquisition information of the vehicle-mounted system;

[0044] A determination module, used to determine the image information of the vehicle driver and the voice information of the vehicle driver according to the image acquisition information and the voice acquisition information;

[0045] An emotion recognition module, used to perform emotion recognition based on the image information and the voice information to obtain an image emotion vector and a voice emotion vector;

[0046] A parameter module, used for determining an emotion fusion parameter according to a vector distance between the image emotion vector and the speech emotion vector;

[0047] A fusion module, configured to fuse the image emotion vector and the speech emotion vector based on the emotion fusion parameter to obtain target emotion information corresponding to the vehicle driver;

[0048] An output module is used to output voice prompt information based on the target emotion information.

[0049] In a third aspect, an electronic device is provided, comprising a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other via the communication bus;

[0050] Memory, used to store computer programs;

[0051] The processor is used to implement the emotion prompting method of the vehicle-mounted system described in any one of the first aspects when executing the program stored in the memory.

[0052] In a fourth aspect, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the emotion prompting method of the vehicle-mounted system as described in any one of the first aspects is implemented.

[0053] In a fifth aspect, a vehicle is provided, comprising the emotion prompting device of the vehicle-mounted system described in the second aspect.

[0054] The embodiment of the present application obtains image acquisition information and voice acquisition information of the vehicle system, determines the image information and voice information of the vehicle driver based on the image acquisition information and the voice acquisition information, and performs emotion recognition based on the image information and the voice information respectively to obtain the image emotion vector and the voice emotion vector, determines the emotion fusion parameter based on the vector distance between the image emotion vector and the voice emotion vector, fuses the image emotion vector and the voice emotion vector based on the emotion fusion parameter, obtains the target emotion information corresponding to the vehicle driver, and then outputs the voice prompt information based on the target emotion information; thereby, the vehicle driver can be given a voice prompt in time, thereby solving the problem in the existing related technology that the driving safety is affected by the driver's emotions, and can effectively improve the driving safety. BRIEF DESCRIPTION OF THE DRAWINGS

[0055] Figure 1 A schematic diagram of a flow chart of an emotion prompting method for an in-vehicle system provided in an embodiment of the present application;

[0056] Figure 2 A schematic diagram of an image processing scenario in an emotion prompting method for an in-vehicle system provided in an embodiment of the present application;

[0057] Figure 3 A schematic diagram of a speech processing scenario in an emotion prompting method for an in-vehicle system provided in an embodiment of the present application;

[0058] Figure 4 A schematic diagram of another scenario of voice processing in an emotion prompting method for an in-vehicle system provided in an embodiment of the present application;

[0059] Figure 5 A schematic diagram of another scenario of voice processing in an emotion prompting method for an in-vehicle system provided in an embodiment of the present application;

[0060] Figure 6 Another schematic diagram of a flow chart of an emotion prompting method for an in-vehicle system provided in an embodiment of the present application;

[0061] Figure 7 Another flowchart of an emotion prompting method for an in-vehicle system provided in an embodiment of the present application;

[0062] Figure 8 A schematic diagram of an application scenario of an emotion prompting method for an in-vehicle system provided in an embodiment of the present application;

[0063] Fig. 9 A schematic diagram of the structure of an emotion prompting device for a vehicle-mounted system provided in an embodiment of the present application;

[0064] Fig.10 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0065] The following will describe the embodiments of the present invention with reference to the accompanying drawings and preferred embodiments. Those skilled in the art can easily understand other advantages and effects of the present invention from the contents disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present invention. It should be understood that the preferred embodiments are only for illustrating the present invention, not for limiting the scope of protection of the present invention.

[0066] In order to improve vehicle driving safety, vehicles nowadays usually use statistical driving time to determine whether the user is driving fatigued or collect driving behavior and road conditions to determine whether the road conditions are safe. Among them, statistical driving time can usually only collect the driver's driving time, and it is difficult to confirm whether the driver is fatigued. Analyzing driving behavior and road conditions can only determine whether there is an impact on driving safety in the external environment, but cannot determine whether there is an impact on the driver's driving safety.

[0067] In the existing related technologies, in order to determine whether the driver has driving safety issues, user images are usually collected for emotion analysis. For example, the user's facial features such as blinking frequency, eye opening range, facial expression movements, etc. are analyzed through user image recognition, and the driver's current emotion is determined by facial features. Adaptive reminders are then given based on the current emotion to avoid driving safety issues. However, each driver's driving habits and facial expression habits are different, that is, relying on the collected user images for emotion judgment will result in low accuracy of emotion judgment.

[0068] In order to solve the problem of low accuracy of emotion judgment in the existing related technologies, the present application provides an emotion prompt method, device, electronic device, storage medium and vehicle for an in-vehicle system, by obtaining image acquisition information and voice acquisition information of the in-vehicle system, determining the image information and voice information of the vehicle driver based on the image acquisition information and the voice acquisition information, and performing emotion recognition based on the image information and the voice information respectively to obtain an image emotion vector and a voice emotion vector, and determining an emotion fusion parameter based on the vector distance between the image emotion vector and the voice emotion vector, and fusing the image emotion vector and the voice emotion vector based on the emotion fusion parameter to obtain target emotion information corresponding to the vehicle driver, and then outputting voice prompt information based on the target emotion information; thereby, voice prompts can be given to the vehicle driver in a timely manner, thereby solving the problem in the existing related technologies that driving safety is affected by the driver being affected by emotions, and while effectively improving driving safety, the fusion of image emotion vectors and voice emotion vectors can also effectively improve the accuracy of emotion judgment.

[0069] Figure 1A flow chart of an emotion prompting method for an in-vehicle system provided in an embodiment of the present application. This method can be applied to one or more electronic devices such as vehicles, in-vehicle systems, and clients. In addition, the execution subject of this method can be hardware or software. When the above-mentioned execution subject is hardware, the execution subject can be one or more of the above-mentioned electronic devices. For example, a single electronic device can execute this method, or multiple electronic devices can cooperate with each other to execute this method. When the above-mentioned execution subject is software, this method can be implemented as multiple software or software modules, or as a single software or software module. No specific limitation is made here.

[0070] like Figure 1 As shown, an emotion prompting method for a vehicle-mounted system provided in an embodiment of the present application may specifically include the following steps:

[0071] Step S110: Acquire image acquisition information and voice acquisition information of the vehicle-mounted system.

[0072] Among them, the vehicle-mounted system can represent a system configured in the vehicle for vehicle control, the image acquisition information can represent information collected by the vehicle-mounted system through an image acquisition device configured in the vehicle, and the image acquisition device can be a camera, a scanner, etc., and the voice acquisition information can represent information collected by the vehicle-mounted system through a voice acquisition device configured in the vehicle, and the voice acquisition device can be a microphone, an audio interface, etc.

[0073] Step S120: Determine the image information of the vehicle driver and the voice information of the vehicle driver based on the image acquisition information and the voice acquisition information.

[0074] Specifically, after obtaining the image acquisition information and the voice acquisition information, the image acquisition information can be processed to obtain image information corresponding to the vehicle driver, and the image information corresponding to the vehicle driver can represent an image containing only the vehicle driver; the voice acquisition information can also be subjected to voice recognition and extraction to obtain voice information corresponding to the vehicle driver, and the voice information corresponding to the vehicle driver can represent an image containing only the vehicle driver.

[0075] It should be noted that, in the process of performing image processing on the image acquisition information to obtain image information corresponding to the vehicle driver in the present embodiment, the specific image processing method may be to perform vehicle driver feature recognition on the image acquisition information, and extract image information containing only the vehicle driver image from the image acquisition information; or it may be to perform vehicle driving area recognition on the image acquisition information, and extract image information containing the vehicle driver image within the vehicle driving area from the image acquisition information; of course, it may also be other image processing methods, which are not specifically limited in the present embodiment.

[0076] In addition, in the present embodiment, when performing voice recognition and extraction on the voice collection information to obtain the voice information corresponding to the vehicle driver, the specific voice recognition and extraction method may be to perform vehicle driver voiceprint recognition on the voice collection information, and extract voice information matching the vehicle driver's voiceprint from the voice collection information; or it may be to identify the sound emission area of ​​the voice collection information, and extract voice information whose sound emission area is the vehicle driving area from the image collection information; of course, it may also be other voice recognition and extraction methods, which are not specifically limited in the present embodiment.

[0077] Specifically, in the data collection stage, that is, the process of collecting image acquisition information and voice acquisition information, the driver's facial expression image can be continuously captured as image acquisition information by the vehicle-mounted camera. The camera has a high resolution to capture subtle changes in the face, and the camera can automatically adjust the focus and exposure to ensure the clarity and stability of the image acquisition information. The microphone can be positioned near the driver to record the driver's voice in real time, and can also use its noise suppression function to eliminate the interference of background noise on voice recognition. At the same time, the microphone also has a sound enhancement function, which can ensure that the voice acquisition information with high clarity and strong stability is recorded. In addition, the data collection process includes a synchronization verification step, which is used to synchronize the image acquisition information of the vehicle driver and the voice acquisition information of the vehicle driver on the time axis to ensure that the image acquisition information and the voice acquisition information are at the same time point, so as to facilitate subsequent data fusion and emotion recognition. Of course, the above is only an example for illustration, and this embodiment does not make specific limitations on this.

[0078] In the data preprocessing stage, that is, in the process of determining the image information of the vehicle driver and the voice information of the vehicle driver, the process of determining the image information may include face detection and alignment, face cropping, graying and normalization of the image acquisition information. Face detection and alignment is to detect the driver's face and align it with a predefined template by using a feature cascade classifier or a deep learning method such as a multi-task convolutional neural network. Face cropping is to crop the detected and aligned facial area from the original image to obtain an image containing only the face. Graying is to convert the facial image into a grayscale image to reduce the computational complexity and reduce the influence of illumination changes. Normalization is to normalize the grayscale image to eliminate the influence of illumination and color, while improving robustness, and then obtain the image information of the vehicle driver. The process of determining the voice information may include noise reduction, framing, window function processing and fast Fourier transform processing to obtain the voice information of the vehicle driver. Among them, noise reduction is to remove background noise from the voice acquisition information by using a noise suppression algorithm such as spectral subtraction to obtain a denoised voice signal. Framing and window function processing is to divide the denoised speech signal into short frames, and then perform window function processing on the speech signal of each frame to obtain the speech signal after window function processing. Feature extraction can be to perform feature extraction on the speech signal after window function processing to obtain the speech information of the vehicle driver. Of course, the above is only an example for illustration, and this embodiment does not make specific limitations on this.

[0079] Step S130: performing emotion recognition based on the image information and the voice information respectively to obtain an image emotion vector and a voice emotion vector.

[0080] Specifically, after determining the image information and voice information of the vehicle driver, emotion recognition can be performed on the image information to obtain an image emotion vector, which can represent the emotion feature vector corresponding to the vehicle driver extracted from the image information, and the specific emotion recognition can be carried out using a neural network model or a decision tree model, etc., which is not specifically limited in this embodiment; emotion recognition can also be performed on the voice information to obtain a voice emotion vector, which can represent the emotion feature vector corresponding to the vehicle driver extracted from the voice information, and the specific emotion recognition can be carried out using a support vector machine (SVM) or a neural network model, etc., which is not specifically limited in this embodiment.

[0081] Specifically, when using a neural network model to perform emotion recognition on image information to obtain an image emotion vector, each facial punctuation can be modeled as a part, and a global mixture can be used to capture changes in topological punctuation. Since a punctuation may only be visible in certain views, in order to reduce the complexity of multi-punctuation modeling, different mixture models can be used to share punctuation models. Subsequently, the model is discriminatively trained in a maximum edge framework. The model can use a Histogram of Oriented Gradient (HOG) descriptor as a representation of each punctuation, such as every 15 discretes of the head posture (yaw angle) from -90 to +90 result in 13 different punctuations, and each such punctuation is represented by a tree. In order to simplify the basic concept, a frontal face representative tree T = (V, E) can be selected, where V represents the punctuation and E represents the edge. For the input image information i, let l i =(x i ,y i ) is the punctuation mark V i The scoring function scores the punctuation points located at L on image i.

[0082] Score(I,L)=Appearance(I,L)+Shape(L)+α (1)

[0083] Appearance(I,L)=∑w i ·Φ(I,l i ) (2)

[0084] Shape(L)=∑a ij dx+b ij dx 2 +c ij dy+d ij dy 2 (3)

[0085] Among them, the preset function (2) gives the template w i With HOG feature Φ(I,L i ) in l i The local matching score calculated for a given specific location. The preset function (3) gives the shape score, which for each pair of components V located at L i and V j Perform spatial constraints, where dx = x i -x j and dy = y i y j can be interpreted as a spring that helps deform facial punctuation points to explain the elastic deformation of the face; the parameter a in the preset function (3) ij , b ij 、cij and d ij represents the static position and stiffness of the spring, which are obtained by training a set of images; the bias term α is associated with the corresponding punctuation of the face. The corresponding inference of this model is the value of L when Score(I,L) is maximized, as shown in Equation 4

[0086] Score * (I,L)=max L [Score(I,L)] (4)

[0087] With the help of dynamic programming, the internal maximization of the tree T = (V, E) can be efficiently achieved; assuming that L * The L with the largest Score(I,L) is the Score * (I,L)=Score(I,L * ). If for the spatial arrangement of parts L * ,Score(I,L * )≥θ 1 , where θ 1 If is a predefined threshold, a face is considered detected; the extracted facial region is used for facial expression recognition to identify the driver's emotions. Therefore, the extracted facial region and punctuation can be used as input to a pre-trained 16-layer Visual Geometry Group 16-layer network (VGG16) to classify emotions; in this method, a pre-trained convolutional neural network can be used to extract features. VGG16 is trained on the ImageNet dataset consisting of a large number of images and can have 1000 classes; in this framework, such as Figure 2 As shown, two VGG16 networks can be fine-tuned on the region of interest (ROI) image and facial landmarks to learn face and punctuation features, thereby improving recognition accuracy. In addition, fine-tuning a pre-trained network using a small number of images is usually much faster than training the network from scratch, saving time and improving recognition accuracy.

[0088] In the VGG16 network, the last three layers are configured with 1000 classes and need to be configured for the new classification problem. The last three layers in both networks can be replaced with a fully connected layer, a softmax layer, and a classification output layer. The pre-trained model is fine-tuned for 1000 epochs with a learning rate of 0.01 and a stochastic gradient descent with momentum optimizer. After fine-tuning, the pre-trained model can be used to classify facial expressions. In the preprocessing step, the face and landmark points in each frame can be detected first; then the ROI image of the face detection and the facial landmarks are passed to two independent VGG16 networks respectively; then, the top-level features of the two networks are integrated and classified using the weighted sum shown in the preset function (5).

[0089]

[0090] Where n = 1, 2, ..., E, E is the total number of emotion classes. Parameters α and β are between 0 and 1, and the values ​​of α and β depend on the performance of each network model. Through experience, the values ​​of α and β can be determined to be 0.6 and 0.4 respectively. vgg16 and V vgg16 They are the output of VGG16 for face ROI and FLP data, respectively. integrated is the final classification vector, that is, the image emotion vector obtained by performing emotion recognition on the image information. Of course, the above is only an example for illustration, and this embodiment does not make any specific limitation on this.

[0091] In addition, in the process of obtaining speech emotion vectors through emotion recognition of speech information, MFCC and LPC can be used to extract features, and then the SVM training data set can be used to identify emotions, such as Figure 3 shown.

[0092] Support Vector Machine (SVM) is a supervised machine learning technique used for classification and regression. It can classify data by finding suitable hyperplanes that can separate the data with the largest margin. New values ​​are separated and analyzed based on the training set.

[0093] The purpose of emotion recognition of speech information is to identify human emotions based on the input speech information. First, the speech signal is extracted, and then the Mel Frequency Cepstrum Coefficients (MFCC) and Linear Predictive Cepstral Coefficients (LPCC) are used to extract features, which contain emotional information, speech patterns and coefficients, as input for further analysis by the classification system. Specifically, it can include four modules: input speech signal, feature extraction using MFCC and LPC, classification based on SVM and output.

[0094] When extracting features, the following information is mainly collected: pitch frequency, MFCC, LPCC, energy, and speaking rate.

[0095] Pitch frequency: Tone signal is one of the important features in speech emotion recognition. Pitch frequency is defined as the vibration frequency of sound. It contains all the information about emotions, because the emotion depends on the pressure in the vocal cords and the air pressure under the glottis. Therefore, under different basic emotional states, the average pitch value in the sample, the range of variation and the profile of the sample are all different.

[0096] MFCC: It is used to calculate the characteristics of the human ear. The human ear has nonlinear frequency units, which makes it similar to the human auditory system. The calculation process of MFCC is as follows Figure 4 shown.

[0097] LPCC: An envelope used to represent a compressed form of the input speech signal. The input signal can be analyzed by estimating the formant or enhanced frequency band, removing their effects from the signal, and estimating the strength and frequency of the remaining signal. The calculation process of LPCC is as follows: Figure 5 shown.

[0098] Energy: Energy calculation is one of the important features of speech analysis. In order to obtain complete statistics of all energy features, some short-term functions can be used to extract the energy value in each speech frame. Using the above information, the overall statistics of the energy of the entire speech sample can be obtained by calculating the mean, maximum value, and range of variation in the speech sample.

[0099] Speech rate: Speech rate shows how fast a person speaks. It is strongly correlated with any emotion such as happiness, sadness, fear, and anger. The time domain waveform of the speech signal is processed, such as using MATLAB functions to calculate the speech.

[0100] Support vector machine is a simple and effective classification algorithm used for classification and pattern recognition. The main purpose of this algorithm is to obtain a function that constructs hyperplanes or boundaries. These hyperplanes are used to separate different categories of input data points, where SVM uses binary classification.

[0101] A support vector machine is a system that uses a hyperplane in a high-dimensional feature space to distinguish values ​​according to specific specifications. The hyperplane is trained using a specific algorithm for statistical learning. The classification method of a support vector machine is similar to supervised learning, involving feature extraction and generating desired outputs. There are generally two types of support vector machine classifiers: linear and nonlinear. A support vector machine also has a kernel function that can be used to monitor the environment. In the training phase, a radial basis kernel function can be used because it limits the training data set to a specified boundary. Then, the LIBSVM tool can be used for SVM classification. During the training process of the signal, the eigenvalue of the speech signal is extracted from the speech signal, and its classification label is Happy, Angry, Sad, Fear, and sent to LIBSVM. Using the extracted emotional state eigenvalue, an SVM model for each emotional state is obtained. After the training model is trained, the trained SVM model can be used to extract features from the speech signal, that is, to perform emotion recognition on the speech information to obtain a speech emotion vector. With the help of the SVM model value generated by the training model, the emotion can be automatically classified into vectors such as Happy, Angry, Sad or Fear. Of course, the above is only an example, and this embodiment does not specifically limit this.

[0102] Step S140: determining an emotion fusion parameter according to the vector distance between the image emotion vector and the speech emotion vector.

[0103] Specifically, after obtaining the image emotion vector and the speech emotion vector, vector distance calculation can be performed on the image emotion vector and the speech emotion vector to obtain the vector distance between the image emotion vector and the speech emotion vector, wherein the vector distance calculation is a way to measure the difference between two vectors in a multidimensional space, and the vector distance can represent the difference between the image emotion vector and the speech emotion vector in a multidimensional space; the emotion fusion parameter can also be determined through the vector distance; wherein the emotion fusion parameter can represent the parameter for fusing the image emotion vector and the speech emotion vector, and the specific determination method can be using a preset function, a neural network model, etc., and this embodiment does not make any specific limitation on this.

[0104] Step S150: Based on the emotion fusion parameter, the image emotion vector and the speech emotion vector are fused to obtain the target emotion information corresponding to the vehicle driver.

[0105] Specifically, after determining the emotion fusion parameter, the emotion fusion parameter can be used to fuse the image emotion vector and the speech emotion vector to obtain the target emotion information corresponding to the vehicle driver, wherein the target emotion information can represent the specific emotion of the vehicle driver. The image emotion vector and the speech emotion vector can be fused by using a preset function, a neural network model, etc., which is not specifically limited in this embodiment.

[0106] Step S160: outputting voice prompt information based on the target emotion information.

[0107] Specifically, after obtaining the target emotion information, different voice prompt information can be output according to different target emotion information, thereby playing the role of prompting the vehicle driver immediately, avoiding the vehicle driver being affected by emotions and affecting driving safety, and can effectively improve driving safety.

[0108] It can be seen that this embodiment can obtain the image acquisition information and voice acquisition information of the vehicle system, determine the image information and the voice information of the vehicle driver based on the image acquisition information and the voice acquisition information, and perform emotion recognition based on the image information and the voice information respectively to obtain the image emotion vector and the voice emotion vector, and determine the emotion fusion parameter based on the vector distance between the image emotion vector and the voice emotion vector, and fuse the image emotion vector and the voice emotion vector based on the emotion fusion parameter to obtain the target emotion information corresponding to the vehicle driver, and then output the voice prompt information based on the target emotion information; thereby, the vehicle driver can be given a voice prompt in time, which solves the problem in the existing related technology that the driving safety is affected by the driver's emotions, and can effectively improve the driving safety. At the same time, due to the fusion of the image emotion vector and the voice emotion vector, it can also effectively improve the accuracy of emotion judgment.

[0109] In an optional embodiment of the present application, step S140 determines the emotion fusion parameter based on the vector distance between the image emotion vector and the speech emotion vector, and may specifically include the following sub-steps: determining a first fusion parameter based on the vector distance between the image emotion vector and the speech emotion vector, the first fusion parameter being exponentially related to the vector distance; performing parameter calculation based on the first fusion parameter in combination with a preset fusion constant threshold to obtain a second fusion parameter; determining the emotion fusion parameter based on the first fusion parameter and the second fusion parameter.

[0110] After obtaining the vector distance between the image emotion vector and the speech emotion vector, this embodiment can determine the first fusion parameter based on the vector distance, wherein the first fusion parameter can be a parameter having a preset relationship with the vector distance, and the preset relationship can be a functional relationship, such as the first fusion parameter and the vector distance are exponentially related; a preset fusion constant threshold can also be obtained, and the fusion constant threshold can represent a parameter threshold pre-configured for parameter calculation, so that the first fusion parameter can be combined with the preset fusion constant threshold to perform parameter calculation to obtain the second fusion parameter; and the first fusion parameter and the second fusion parameter can be integrated to obtain the emotion fusion parameter.

[0111] Specifically, the following function can be used for parameter calculation:

[0112] α=1-β

[0113] Where β is the first fusion parameter, α is the second fusion parameter, and 1 is the preset fusion constant threshold; the first fusion parameter β can be calculated using the following function:

[0114] β=0.4e -d / 100

[0115] Wherein, d is the vector distance, e is the natural logarithm, and 0.4 and 100 are constants. Of course, the above is only an example for illustration, and this embodiment does not make any specific limitation to this.

[0116] It can be seen that the first fusion parameter and the second fusion parameter in the emotion fusion parameters of this embodiment are both related to the vector distance, that is, different vector distances can determine different emotion fusion parameters, and the vector distance is determined by the image emotion vector corresponding to the image acquisition information and the voice emotion vector corresponding to the voice acquisition information. Therefore, when this embodiment uses the emotion fusion parameters to fuse the image emotion vector and the voice emotion vector to obtain the target emotion information corresponding to the vehicle driver, it determines the fused emotion fusion parameters based on the specific distance between the image emotion vector and the voice emotion vector, thereby achieving a fusion effect based on the difference between the voice emotion vector and the image emotion vector, that is, each time the process of determining the target emotion information will dynamically adjust the specific emotion fusion parameters, which can effectively improve the accuracy and flexibility of emotion determination.

[0117] In an optional embodiment of the present application, step S150 fuses the image emotion vector and the voice emotion vector based on the emotion fusion parameter to obtain the target emotion information corresponding to the vehicle driver, which may specifically include the following steps: extracting the first fusion parameter and the second fusion parameter from the emotion fusion parameter; calculating the voice emotion vector and the first fusion parameter to obtain the first fusion vector; calculating the image emotion vector and the second fusion parameter to obtain the second fusion vector; combining the first fusion vector and the second fusion vector to obtain the target fusion vector; and determining the target emotion information based on the target fusion vector.

[0118] After obtaining the emotion fusion parameter, this embodiment can extract the first fusion parameter and the second fusion parameter from the emotion fusion parameter, and calculate the image emotion vector with the first fusion parameter to obtain the first fusion vector, wherein the specific calculation method can be to perform weight calculation on the image emotion vector according to the weight corresponding to the first fusion parameter to obtain the first fusion vector. The voice emotion vector can also be calculated with the second fusion parameter to obtain the second fusion vector, wherein the specific calculation method can be to perform weight calculation on the voice emotion vector according to the weight corresponding to the second fusion parameter to obtain the first fusion vector; thus, the first fusion vector and the second fusion vector can be combined to obtain the target fusion vector, wherein the combination method of the first fusion vector and the second fusion vector can be to superimpose the vector values, or to superimpose the vector values ​​of the first fusion vector and the second fusion vector according to the preset weights, and of course, other combination methods can also be used, which are not specifically limited in this embodiment. Then, the emotion corresponding to the target fusion vector can be determined as the target emotion information.

[0119] Specifically, the process of obtaining the target fusion vector can be calculated using the following formula:

[0120]

[0121] Among them, E N is the speech emotion vector, K N is the image emotion vector, β is the first fusion parameter, α is the second fusion parameter, Z N is the target fusion vector. Since the first fusion parameter is β = 0.4e -d / 100 , where d is the vector distance, and the exponent has an explosive growth characteristic. It can be seen that when the vector distance is large, a smaller proportion can be assigned to the speech emotion vector for fusion. When the emotion distance d is 0, the initial fusion factor is 0.4; when the emotion distance d is 100, β shows an explosive downward trend, thereby improving the assignment of the image emotion vector, achieving the purpose of determining the target fusion vector with the image emotion vector as the main and the speech emotion vector as the auxiliary. Of course, the above is only an example for illustration, and this embodiment does not make specific limitations on this.

[0122] In an optional embodiment of the present application, the first fusion vector and the second fusion vector are combined to obtain a target fusion vector, which may specifically include the following steps: extracting at least one first emotion vector from the first fusion vector, and determining the first emotion category corresponding to the first emotion category vector; extracting at least one second emotion vector from the second fusion vector, and determining the second emotion category corresponding to the second emotion category vector; when the first emotion type matches the second emotion type, combining the first emotion vector and the second emotion vector to obtain a target emotion vector; and obtaining a target fusion vector based on the target emotion vector.

[0123] In the process of combining the first fusion vector and the second fusion vector to obtain the target fusion vector in this embodiment, the first fusion vector may include a fusion vector corresponding to the speech emotion vector representing the probability distribution of each emotion type; thereby, at least one first emotion vector can be extracted from the first fusion vector, and the first emotion category corresponding to the first emotion category vector can be determined, and at least one second emotion vector can be extracted from the second fusion vector, and the second emotion category corresponding to the second emotion category vector can be determined; then, it can be determined whether each first emotion category matches the second emotion category. When the first emotion type matches the second emotion type, the first emotion vector and the second emotion vector can be combined to obtain the target emotion vector, that is, one or more target emotion vectors can be obtained, and at this time, each target emotion vector can be integrated to obtain the target fusion vector.

[0124] In one example, when the target fusion vector is calculated using the following formula:

[0125]

[0126] From the first fusion vector αE N Extract at least one first emotion vector αE 1 , αE 2 ....αE N , and determine the first emotion category corresponding to the first emotion category vector, such as the first emotion vector αE 1 The corresponding first emotion category is the fear emotion category, and the first emotion vector αE 2 The corresponding first emotion category is the surprise emotion category, and the first emotion vector αE N The corresponding first emotion category is the alert emotion category, etc. From the second fusion vector αE N Extract at least one second emotion vector βK 1 , βK 2 ....βK N , and determine the second emotion category corresponding to the second emotion category vector, such as the second emotion vector βK 1The corresponding second emotion category is the fear emotion category, and the second emotion vector βK 2 The corresponding second emotion category is the surprise emotion category, and the second emotion vector βK N The corresponding second emotion category is the alert emotion category. When the first emotion type matches the second emotion type, the first emotion vector and the second emotion vector are combined. For example, the first emotion category and the second emotion category are both fear emotion categories. 1 and βK 1 Combined, the target fusion vector is Z 1 =αE 1 +βK 1 , where the first emotion category and the second emotion category are both surprise emotion categories 2 and βK 2 Combined, the target fusion vector is Z 2 =αE 2 +βK 2 , the first emotion category and the second emotion category are both vigilance emotion categories αE N and βK N Combined, the target fusion vector is Z N =αE N +βK N . Each target fusion vector can then be added together to obtain the target fusion vector Of course, the above is only an example to illustrate the effect, and this embodiment does not make any specific limitation to this.

[0127] like Figure 6 As shown, in an optional embodiment of the present application, after performing emotion recognition based on the image information and the voice information to obtain the image emotion vector and the voice emotion vector, step S130 may further include the following steps:

[0128] Step S131: determining the image emotion identifier corresponding to the image emotion vector;

[0129] Step S132: determining the speech emotion identifier corresponding to the speech emotion vector;

[0130] Step S133: When the image emotion identifier and the speech emotion identifier do not belong to the preset opposing emotions, a step of determining an emotion fusion parameter according to a vector distance between the image emotion vector and the speech emotion vector is performed.

[0131] After obtaining the image emotion vector and the speech emotion vector, this embodiment can determine the image emotion identifier corresponding to the image emotion vector and determine the speech emotion identifier corresponding to the speech emotion vector, wherein the image emotion identifier can represent the identifier of the emotion corresponding to the image emotion vector, and the speech emotion identifier can represent the identifier of the emotion corresponding to the speech emotion vector, so that according to the pre-configured corresponding relationship, it can be judged whether the image emotion identifier and the speech emotion identifier are preset opposing emotions; wherein the preset opposing emotions can represent pre-configured mutually opposing emotions, such as happiness and sadness, trust and disgust, etc., and this embodiment does not make specific limitations on this. In the case where the image emotion identifier and the speech emotion identifier do not belong to the preset opposing emotions, it means that the emotions represented by the image emotion vector and the speech emotion vector are not opposite, and at this time, step S140 can be executed to determine the emotion fusion parameter based on the vector distance between the image emotion vector and the speech emotion vector.

[0132] It can be seen that the premise for executing steps S140-S160 for emotion fusion in this embodiment is that the emotions represented by both the image emotion vector and the speech emotion vector are non-opposing emotions. That is, only when the image emotion recognition and the speech emotion recognition respectively identify non-opposing emotions, will the image emotion vector and the speech emotion vector obtained by the two emotion recognitions be fused, which can effectively improve the accuracy of emotion fusion and improve the fusion efficiency.

[0133] like Figure 6 As shown, in an optional embodiment of the present application, after determining the speech emotion identifier corresponding to the speech emotion vector, the following steps may be specifically included:

[0134] Step S134: when the image emotion identifier and the voice emotion identifier belong to preset opposing emotions, obtaining preset neutral emotion information;

[0135] Step S135: determining the neutral emotion information as the target emotion information.

[0136] After determining the voice emotion identifier corresponding to the voice emotion vector in this embodiment, if the image emotion identifier and the voice emotion identifier belong to preset opposing emotions, it means that the emotions represented by the image emotion vector and the voice emotion vector are not opposite. For example, the image emotion identifier and the voice emotion identifier respectively represent opposite emotions such as happiness and sadness, trust and disgust. At this time, the preset neutral emotion information can be obtained. The neutral emotion information is pre-configured emotion information indicating that the emotion is in a neutral state; thereby, the neutral emotion information can be determined as the target emotion information.

[0137] It can be seen that in this embodiment, when the image emotion marker and the voice emotion marker respectively represent opposing emotions, the preset neutral emotion information is directly determined as the target emotion information, which can achieve at least two effects. On the one hand, it can avoid the erroneous results caused by image recognition or voice recognition deviations. On the other hand, it can save the time consumed by re-emotion recognition and quickly enter the next cycle of emotion recognition, thereby effectively improving the efficiency of emotion recognition.

[0138] In an optional embodiment of the present application, step S160 outputs voice prompt information based on the target emotional information, which may specifically include the following steps: determining the emotional prompt text and emotional level corresponding to the target emotional information; obtaining the tone information corresponding to the emotional level; combining the tone information with the emotional prompt text to generate voice prompt information; and outputting based on the voice prompt information.

[0139] After obtaining the target emotion information, the present embodiment can determine the emotion prompt text and emotion level corresponding to the target emotion information. The emotion prompt text can represent the text pre-configured for prompting the target emotion, and the emotion level can represent the level used to represent the emotion degree in the target emotion information, such as level one alert, level two alert, level three alert, etc. Different emotion levels represent that the current vehicle driver is in a state corresponding to the target emotion to different degrees; thereby, the tone information corresponding to the emotion level can be obtained, and the tone information can represent the tone situation pre-configured corresponding to the emotion level. The tone information can include one or more tone features, and the tone features can include but are not limited to intonation features, emotional color features, attitude features, and urgency features, etc., which are used to represent different tones; and then the tone information can be combined with the emotion prompt text to generate voice prompt information, and the voice prompt information can represent information used to output and prompt the user, and the combination method can be to generate voice from the tone features in the tone information and the emotion prompt text, and then output based on the voice prompt information, specifically, the voice prompt information can be played through a speaker configured in the vehicle, and the present embodiment does not make specific restrictions on this.

[0140] In one example, a preset voice text prompt library can be pre-established. The voice text prompt may include emotional prompt texts corresponding to various target emotional information, such as when the target emotional information is fatigue, the corresponding emotional prompt text is "You look a little tired, it's time to stop and take a rest"; when the target emotional information is anxiety / tension, the corresponding emotional prompt text is "Please take a deep breath and stay calm; remember, safety is the most important"; when the target emotional information is anger / excitement, the corresponding emotional prompt text is "It seems that you are a little excited, please calm down and drive safely"; when the target emotional information is happy / excited, the corresponding emotional prompt text is "You look in a good mood, but please note that safe driving is always the first priority"; when the target emotional information is distraction, the corresponding emotional prompt text is "It seems that you are a little distracted, pay attention to returning to driving, and stay focused". Of course, the emotional prompt text in the voice text prompt library can be a default generated text or a user-defined input text, which is not specifically limited in this embodiment.

[0141] In addition, if it is determined that the emotional prompt text corresponding to the target emotional information is "Please take a deep breath and stay calm. Remember, safety is the most important thing" and the emotional level is secondary anxiety, the tone information corresponding to the emotional level of secondary anxiety is a medium tone feature and a low-urgency feature, so that the emotional prompt text "Please take a deep breath and stay calm. Remember, safety is the most important thing" can be used for voice generation with the medium tone feature and the low-urgency feature contained in the tone information to obtain a voice prompt information of "Please take a deep breath and stay calm. Remember, safety is the most important thing" in a medium tone and low urgency; that is, when the vehicle outputs the voice prompt information, the user can hear the text "Please take a deep breath and stay calm. Remember, safety is the most important thing" in a medium tone and low urgency, so as to not only broadcast the text content to prompt the vehicle driver's current emotions, but also adapt and comfort the vehicle driver in the playback tone.

[0142] like Figure 7 As shown, in an optional embodiment of the present application, step S160 outputs voice prompt information based on the target emotion information, which may specifically include the following steps:

[0143] Step S161: Obtain vehicle status information;

[0144] Step S162: When the vehicle status information belongs to a preset status, suspend outputting the voice prompt information;

[0145] Step S163: When the vehicle status information does not belong to the preset status, continuously output the voice prompt information.

[0146] After obtaining the target emotion information, the present embodiment can also obtain vehicle status information. The vehicle status information can represent relevant information of the current driving status of the vehicle, and can include but not be limited to the vehicle driving status such as straight driving, merging, turning, and the navigation playback status in the cockpit such as navigation being broadcast, navigation not being broadcast, etc.; thereby, it can be determined whether the vehicle status information belongs to a preset state, wherein the preset state can represent a pre-configured state that is not suitable for outputting voice prompt information; that is, when the vehicle status information belongs to the preset state, it means that it is not suitable for outputting voice prompt information. At this time, the output of voice prompt information can be suspended, and continuous monitoring can be carried out to determine whether the vehicle status information is in the preset state; and when the vehicle status information does not belong to the preset state, the voice prompt information can be continuously output.

[0147] This embodiment determines whether it is suitable to output voice prompt information through the preset state, which plays a role in avoiding outputting voice prompt information in unsuitable scenarios, and causing unnecessary interference to the vehicle driver and affecting the vehicle operation. For example, the current vehicle state information indicates that the vehicle is in a lane-merging state, and the lane-merging state belongs to the preset state. Since the vehicle driver needs to concentrate more attention to observe vehicles in multiple directions when merging, and complete the lane-merging operation in time, if the voice prompt information is suddenly played and output at this time, the vehicle driver may be disrupted by the changed voice prompt information when operating, thereby affecting the vehicle driver's lane-merging operation; at this time, configuring the lane-merging state as a preset operation can avoid the above-mentioned problem. The preset state can also include the navigation being broadcast state. Since the navigation information is related to the specific driving direction and route of the vehicle driver, if the voice prompt information is output and played in the navigation being broadcast state, it will cause the vehicle driver to be unable to hear or ignore the navigation information, and it will also cause interference to the vehicle driver, and affect the subsequent vehicle driving route; at this time, configuring the navigation being broadcast state as a preset operation can avoid the above-mentioned problem. Of course, the preset state can also include other states that are not suitable for outputting voice prompt information, which will not be described one by one in this embodiment. In addition, the preset status may also be added or removed according to the status configuration information input by the user, which is not specifically limited in this embodiment.

[0148] It can be seen that Figure 8As shown, this embodiment can obtain image acquisition information and voice acquisition information of the vehicle system through the vehicle's data collection modules such as the vehicle-mounted camera and the vehicle-mounted microphone, and perform image preprocessing and voice preprocessing respectively according to the image acquisition information and the voice acquisition information through the data preprocessing module to obtain the image information of the vehicle driver and the voice information of the vehicle driver. It can also perform emotion recognition according to the image information and the voice information respectively through the emotion recognition module to obtain the image emotion vector and the voice emotion vector, and determine the emotion fusion parameter according to the vector distance between the image emotion vector and the voice emotion vector, so as to fuse the image emotion vector and the voice emotion vector based on the emotion fusion parameter to obtain the target emotion information corresponding to the vehicle driver, and then output the voice prompt information based on the target emotion information; thereby, the vehicle driver can be given a voice prompt in time, so as to solve the problem of driving safety affected by the driver's emotions in the existing related technology, and can effectively improve driving safety. At the same time, due to the fusion of the image emotion vector and the voice emotion vector, the accuracy of emotion judgment can also be effectively improved.

[0149] like Fig. 9 As shown, the present application also discloses an embodiment, providing an emotion prompting device of a vehicle-mounted system, comprising:

[0150] An acquisition module 910 is used to acquire image acquisition information and voice acquisition information of the vehicle-mounted system;

[0151] A determination module 920, configured to determine the image information of the vehicle driver and the voice information of the vehicle driver according to the image acquisition information and the voice acquisition information;

[0152] An emotion recognition module 930 is used to perform emotion recognition based on the image information and the voice information to obtain an image emotion vector and a voice emotion vector;

[0153] A parameter module 940, configured to determine an emotion fusion parameter according to a vector distance between the image emotion vector and the speech emotion vector;

[0154] A fusion module 950 is used to fuse the image emotion vector and the speech emotion vector based on the emotion fusion parameter to obtain target emotion information corresponding to the vehicle driver;

[0155] The output module 960 is used to output voice prompt information based on the target emotion information.

[0156] In an optional embodiment of the present application, the parameter module 940 may include:

[0157] A first determining unit, configured to determine a first fusion parameter according to a vector distance between the image emotion vector and the speech emotion vector, wherein the first fusion parameter is exponentially related to the vector distance;

[0158] a parameter calculation unit, configured to perform parameter calculation based on the first fusion parameter and a preset fusion constant threshold to obtain a second fusion parameter;

[0159] The second determining unit is configured to determine the emotion fusion parameter according to the first fusion parameter and the second fusion parameter.

[0160] In an optional embodiment of the present application, the fusion module 950 may include:

[0161] A first extraction unit, configured to extract a first fusion parameter and a second fusion parameter from the emotion fusion parameter;

[0162] A first calculation unit, configured to calculate the speech emotion vector and the first fusion parameter to obtain a first fusion vector;

[0163] A second calculation unit, configured to calculate the image emotion vector and the second fusion parameter to obtain a second fusion vector;

[0164] A first combining unit, configured to combine the first fusion vector and the second fusion vector to obtain a target fusion vector;

[0165] The third determining unit is used to determine the target emotion information according to the target fusion vector.

[0166] In an optional embodiment of the present application, the first combining unit may include:

[0167] A first determining subunit, configured to extract at least one first emotion vector from the first fusion vector, and determine a first emotion category corresponding to the first emotion category vector;

[0168] A second determining subunit, configured to extract at least one second emotion vector from the second fusion vector, and determine a second emotion category corresponding to the second emotion category vector;

[0169] A first combining subunit, configured to combine the first emotion vector and the second emotion vector to obtain a target emotion vector when the first emotion type matches the second emotion type;

[0170] The third determining subunit is used to obtain the target fusion vector according to the target emotion vector.

[0171] In an optional embodiment of the present application, the emotion prompting device of the vehicle-mounted system may further include:

[0172] A first emotion identification module, used to determine the image emotion identification corresponding to the image emotion vector;

[0173] A second emotion identification module, used to determine the speech emotion identification corresponding to the speech emotion vector;

[0174] The execution module is used to execute the step of determining the emotion fusion parameter based on the vector distance between the image emotion vector and the speech emotion vector when the image emotion identifier and the speech emotion identifier do not belong to the preset opposing emotions.

[0175] In an optional embodiment of the present application, the emotion prompting device of the vehicle-mounted system may further include:

[0176] A neutral emotion module, used for obtaining preset neutral emotion information when the image emotion identifier and the voice emotion identifier belong to preset opposing emotions;

[0177] A target module is used to determine the neutral emotion information as the target emotion information.

[0178] In an optional embodiment of the present application, the output module 960 may include:

[0179] A fourth determining unit, configured to determine an emotion prompt text and an emotion level corresponding to the target emotion information;

[0180] A first acquisition unit, configured to acquire tone information corresponding to the emotion level;

[0181] A generating unit, used for combining the tone information with the emotion prompt text to generate the voice prompt information;

[0182] An output unit is used to output based on the voice prompt information.

[0183] In an optional embodiment of the present application, the output module 960 may include:

[0184] A second acquisition unit, used to acquire vehicle status information;

[0185] A pause output unit, used for pausing the output of the voice prompt information when the vehicle status information belongs to a preset status;

[0186] The continuous output unit is used to continuously output the voice prompt information when the vehicle status information does not belong to a preset status.

[0187] The implementation process of the functions and effects of each module in the above-mentioned device is specifically described in the implementation process of the corresponding steps in the above-mentioned method, which will not be repeated here.

[0188] like Fig.10 As shown, an embodiment of the present application provides an electronic device, including a processor 1010, a communication interface 1020, a memory 1010 and a communication bus 1040, wherein the processor 1010, the communication interface 1020, and the memory 1010 communicate with each other through the communication bus 1040.

[0189] Memory 1010, used for storing computer programs;

[0190] In one embodiment of the present application, the processor 1010 is used to execute the program stored in the memory 1010 to implement the emotion prompt method of the vehicle system provided by any of the aforementioned method embodiments, by acquiring image acquisition information and voice acquisition information of the vehicle system, determining the image information of the vehicle driver and the voice information of the vehicle driver based on the image acquisition information and the voice acquisition information, and performing emotion recognition based on the image information and the voice information respectively to obtain the image emotion vector and the voice emotion vector, and determining the emotion fusion parameter based on the vector distance between the image emotion vector and the voice emotion vector, so as to fuse the image emotion vector and the voice emotion vector based on the emotion fusion parameter to obtain the target emotion information corresponding to the vehicle driver, and then output the voice prompt information based on the target emotion information; thereby, the vehicle driver can be given a voice prompt in a timely manner, thereby solving the problem of driving safety affected by the driver being affected by emotions in the existing related technologies, and effectively improving driving safety.

[0191] The embodiment of the present application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the emotion prompt method of the vehicle system provided in any of the aforementioned method embodiments are implemented, by acquiring image acquisition information and voice acquisition information of the vehicle system, determining the image information of the vehicle driver and the voice information of the vehicle driver based on the image acquisition information and the voice acquisition information, and performing emotion recognition based on the image information and the voice information respectively to obtain an image emotion vector and a voice emotion vector, and determining an emotion fusion parameter based on the vector distance between the image emotion vector and the voice emotion vector, so as to fuse the image emotion vector and the voice emotion vector based on the emotion fusion parameter to obtain target emotion information corresponding to the vehicle driver, and then outputting voice prompt information based on the target emotion information; thereby, voice prompts can be given to the vehicle driver in a timely manner, solving the problem of driving safety affected by the driver being affected by emotions in the existing related technologies, and effectively improving driving safety.

[0192] An embodiment of the present application also provides a vehicle, which may include the emotion prompting device of the vehicle-mounted system described in any of the aforementioned embodiments.

[0193] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0194] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a general hardware platform, and of course, by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the relevant technology can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0195] It should be understood that the terms used herein are only for the purpose of describing specific example embodiments and are not intended to be limiting. Unless the context clearly indicates otherwise, the singular forms "one", "an" and "said" as used herein may also be meant to include plural forms. The terms "include", "comprise", "contain", and "have" are inclusive, and therefore specify the existence of stated features, steps, operations, elements and / or parts, but do not exclude the existence or addition of one or more other features, steps, operations, elements, parts, and / or combinations thereof. The method steps, processes, and operations described herein are not interpreted as necessarily requiring them to be performed in the specific order described or illustrated, unless the execution order is clearly indicated. It should also be understood that additional or alternative steps may be used.

[0196] The foregoing is merely a specific embodiment of the present invention, which enables those skilled in the art to understand or implement the present invention. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but rather to the widest scope consistent with the principles and novel features claimed herein.

Claims

1. An emotion prompting method for a vehicle-mounted system, characterized in that: include: Obtain image acquisition information and voice acquisition information from the vehicle-mounted system; Determining the image information of the vehicle driver and the voice information of the vehicle driver according to the image acquisition information and the voice acquisition information; Performing emotion recognition based on the image information and the voice information respectively to obtain an image emotion vector and a voice emotion vector; Determining an emotion fusion parameter according to a vector distance between the image emotion vector and the speech emotion vector; Based on the emotion fusion parameter, the image emotion vector and the voice emotion vector are fused to obtain target emotion information corresponding to the vehicle driver; Output voice prompt information based on the target emotion information.

2. The emotion prompting method of the vehicle-mounted system according to claim 1, characterized in that: The step of determining the emotion fusion parameter according to the vector distance between the image emotion vector and the speech emotion vector comprises: Determining a first fusion parameter according to a vector distance between the image emotion vector and the speech emotion vector, wherein the first fusion parameter is exponentially related to the vector distance; Based on the first fusion parameter, a parameter is calculated in combination with a preset fusion constant threshold to obtain a second fusion parameter; The emotion fusion parameter is determined according to the first fusion parameter and the second fusion parameter.

3. The emotion prompting method of the vehicle-mounted system according to claim 1, characterized in that: The step of fusing the image emotion vector and the voice emotion vector based on the emotion fusion parameter to obtain target emotion information corresponding to the vehicle driver includes: Extracting a first fusion parameter and a second fusion parameter from the emotion fusion parameters; Calculating the speech emotion vector and the first fusion parameter to obtain a first fusion vector; Calculating the image emotion vector and the second fusion parameter to obtain a second fusion vector; Combining the first fusion vector and the second fusion vector to obtain a target fusion vector; The target emotion information is determined according to the target fusion vector.

4. The emotion prompting method of the vehicle-mounted system according to claim 3, characterized in that: The combining the first fusion vector and the second fusion vector to obtain a target fusion vector includes: extracting at least one first emotion vector from the first fusion vector, and determining a first emotion category corresponding to the first emotion category vector; extracting at least one second emotion vector from the second fusion vector, and determining a second emotion category corresponding to the second emotion category vector; In the case where the first emotion type matches the second emotion type, combining the first emotion vector and the second emotion vector to obtain a target emotion vector; The target fusion vector is obtained according to the target emotion vector.

5. The emotion prompting method of the vehicle-mounted system according to claim 1, characterized in that: After performing emotion recognition according to the image information and the voice information to obtain an image emotion vector and a voice emotion vector, the method further includes: Determining an image emotion identifier corresponding to the image emotion vector; Determining a speech emotion identifier corresponding to the speech emotion vector; In the case where the image emotion identifier and the voice emotion identifier do not belong to preset opposing emotions, the step of determining the emotion fusion parameter based on the vector distance between the image emotion vector and the voice emotion vector is performed.

6. The emotion prompting method of the vehicle-mounted system according to claim 5, characterized in that: After determining the speech emotion identifier corresponding to the speech emotion vector, the method further includes: When the image emotion identifier and the voice emotion identifier belong to preset opposing emotions, obtaining preset neutral emotion information; The neutral emotion information is determined as the target emotion information.

7. The emotion prompting method of the vehicle-mounted system according to any one of claims 1 to 6, characterized in that: The outputting voice prompt information based on the target emotion information comprises: Determine the emotion prompt text and emotion level corresponding to the target emotion information; Obtaining tone information corresponding to the emotion level; Combining the tone information with the emotion prompt text to generate the voice prompt information; Output is performed based on the voice prompt information.

8. The emotion prompting method of the vehicle-mounted system according to any one of claims 1 to 6, characterized in that: The outputting voice prompt information based on the target emotion information comprises: Get vehicle status information; When the vehicle status information belongs to a preset status, pausing the output of the voice prompt information; When the vehicle status information does not belong to a preset status, the voice prompt information is continuously outputted.

9. An emotion prompting device for a vehicle-mounted system, characterized in that: include: An acquisition module, used to acquire image acquisition information and voice acquisition information of the vehicle-mounted system; A determination module, used to determine the image information of the vehicle driver and the voice information of the vehicle driver according to the image acquisition information and the voice acquisition information; An emotion recognition module, used to perform emotion recognition based on the image information and the voice information to obtain an image emotion vector and a voice emotion vector; A parameter module, used for determining an emotion fusion parameter according to a vector distance between the image emotion vector and the speech emotion vector; A fusion module, configured to fuse the image emotion vector and the speech emotion vector based on the emotion fusion parameter to obtain target emotion information corresponding to the vehicle driver; An output module is used to output voice prompt information based on the target emotion information.

10. An electronic device, characterized in that: It includes a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other through the communication bus; Memory, used to store computer programs; The processor is used to implement the emotion prompting method of the vehicle-mounted system as described in any one of claims 1-8 when executing the program stored in the memory.

11. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the emotion prompting method of the vehicle-mounted system as described in any one of claims 1-8 is implemented.

12. A vehicle, characterized in that: The vehicle comprises the emotion prompting device of the vehicle-mounted system according to claim 9.