Multi-modal emotion recognition method and device, electronic equipment, storage medium and product
By segmenting and feature engineering audio and video, and combining audio and video recognition results, the problem of inaccurate emotion recognition in a single recognition method is solved, achieving more accurate and reliable emotion recognition, and making it easier for users to quickly locate emotion results.
Patent Information
- Application Number
- CN202510906574.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-01
- Publication Date
- 2025-09-16
AI Technical Summary
Existing emotion recognition methods rely on single image or audio recognition, resulting in inaccurate emotion recognition results, incomplete information, and low accuracy and credibility.
By obtaining the audio and video to be recognized, segmenting the audio stream and inputting it into the audio recognition model, combining it with the video stream for feature engineering, analyzing the audio and video recognition results separately, and finally determining the target emotion result, the confidence level of the audio and video recognition results is used for weighting or selection, and the emotion result is marked on the audio and video segment.
The accuracy and reliability of emotion recognition have been improved, and users can quickly locate the audio and video positions corresponding to the emotion results, improving the playback experience.
Smart Images

Figure CN120656489A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of emotion recognition technology. Specifically, the present application relates to a multimodal emotion recognition method, device, electronic device, computer-readable storage medium, and computer program product. Background Art
[0002] With the continuous breakthroughs in deep learning technology, emotion recognition technology has made great progress in recent years. Deep learning-based models, such as recurrent neural networks and convolutional neural networks, can recognize multimodal information such as facial expressions, body language, and head posture from video images, and perform emotion recognition and judgment in audio and video.
[0003] Existing emotion recognition methods mostly use single image recognition or sound recognition, inputting images and sounds into the same model to obtain recognition results. However, relying solely on audio or video has the limitation of incomplete information, resulting in emotion recognition errors, low credibility and accuracy. Summary of the Invention
[0004] The embodiments of the present application provide a multimodal emotion recognition method, apparatus, electronic device, computer-readable storage medium, and computer program product, which aim to solve the technical problem that emotion results obtained by single image or audio recognition are not accurate enough.
[0005] In a first aspect, a multimodal emotion recognition method is provided, the method comprising:
[0006] Obtain the audio and video to be identified; the audio and video to be identified include audio stream and video stream;
[0007] Segmenting the audio stream to obtain at least one audio segment;
[0008] Input each audio segment into the audio recognition model to obtain the audio recognition result;
[0009] Determine a corresponding video segment in the video stream according to a target audio segment for which the audio recognition result is an emotion result;
[0010] Input the video segment into the video recognition model to obtain the video recognition result;
[0011] Based on the audio recognition results and the video recognition results, the target emotion result of the audio and video to be recognized is determined.
[0012] Optionally, each audio segment is input into an audio recognition model to obtain audio recognition results, including:
[0013] Perform feature engineering on each audio segment to obtain the audio feature vector corresponding to each audio segment; the audio feature vector includes the frequency change, amplitude change, pitch and timbre of the audio segment;
[0014] The audio feature vectors corresponding to each audio segment are input into the audio recognition model one by one to obtain the audio recognition results corresponding to each audio segment.
[0015] Optionally, feature engineering is performed on each audio segment to obtain an audio feature vector corresponding to each audio segment, including:
[0016] Frame each audio segment to obtain multiple audio frames;
[0017] Perform frequency analysis on each audio frame to obtain a frequency feature vector; the frequency feature vector includes spectrum centroid data, spectrum dispersion data, spectrum slope data, and spectrum rolling data;
[0018] Perform amplitude analysis on each audio frame to obtain an amplitude feature vector; the amplitude feature vector includes instantaneous amplitude, amplitude envelope, energy data, zero-crossing rate data, and loudness data;
[0019] Perform pitch analysis on each audio frame to obtain a pitch feature vector; the pitch feature vector includes fundamental frequency data, fundamental frequency time series data and harmonic wave data;
[0020] Performing timbre analysis on each audio frame to obtain a timbre feature vector; the timbre feature vector includes Mel-frequency cepstral coefficient data, chroma feature data, and linear prediction coefficient data;
[0021] The frequency feature vector, amplitude feature vector, pitch feature vector and timbre feature vector corresponding to each audio frame are aggregated to obtain the audio feature vector corresponding to each audio segment.
[0022] Optionally, the video segment is input into a video recognition model to obtain video recognition results, including:
[0023] Perform feature engineering on each video segment to obtain the corresponding video feature vector of each video segment; the video feature vector includes the face change features, action change features and scene change features of the video segment;
[0024] The video feature vectors corresponding to each video segment are input into the video recognition model one by one to obtain the video recognition results corresponding to each video segment.
[0025] Optionally, the video segment includes multiple image frames;
[0026] After inputting the video segment into the video recognition model and obtaining the video recognition result, it also includes:
[0027] Input each image frame into the image recognition model to obtain the image recognition result and the corresponding confidence level;
[0028] Based on the image recognition result and the corresponding confidence level, obtaining the image recognition result and the corresponding image recognition confidence level of the video segment;
[0029] Get the video recognition confidence of the video recognition result;
[0030] If the image recognition confidence is greater than the video recognition confidence, the video recognition result is updated based on the image recognition result.
[0031] Optionally, each image frame is input into an image recognition model to obtain an image recognition result, including:
[0032] Perform feature engineering on each image frame to obtain an image feature vector corresponding to each image frame; the image feature vector includes facial features, action features, and scene features of the image frame;
[0033] The image feature vectors corresponding to each image frame are input into the image recognition model one by one to obtain the image recognition results.
[0034] Optionally, based on the audio recognition results and the image recognition results, a target emotion result of the audio or video to be recognized is determined, including:
[0035] determining a first confidence level of the audio recognition result and a second confidence level of the video recognition result;
[0036] If the audio recognition result is inconsistent with the video recognition result, the first confidence level and the second confidence level are compared, and the recognition result with the higher confidence level is used as the target emotion result.
[0037] Optionally, determining a corresponding video segment in the video stream according to the target audio segment for which the audio recognition result is an emotion result includes:
[0038] Determine a target audio segment for which the audio recognition result is an emotion result;
[0039] The video stream is segmented according to the target audio segment to obtain corresponding video segments; the start and end times of the video segments correspond to the start and end times of the audio segments.
[0040] Optionally, the method further includes:
[0041] Based on the target emotion result, determining a tagging style for the corresponding audio segment and / or video segment;
[0042] Based on the corresponding marking style, the target emotion result is marked at the corresponding position of the audio segment and / or video segment to be identified.
[0043] Optionally, the method further includes:
[0044] In response to a user's request to view the marked audio or video, a marked list of audio segments or a marked list of video segments of the marked audio or video is displayed; the marked list includes target emotion result marks sorted by time;
[0045] In response to the user's selection of a tag for a target emotion result, the corresponding audio segment and / or video segment is played.
[0046] In a second aspect, a multimodal emotion recognition device is provided, the device comprising:
[0047] The audio and video acquisition module is used to acquire the audio and video to be identified; the audio and video to be identified include audio stream and video stream;
[0048] A segmentation module, configured to segment the audio stream to obtain at least one audio segment;
[0049] The first recognition module is used to input each audio segment into the audio recognition model to obtain an audio recognition result;
[0050] A video segment determination module is used to determine a corresponding video segment in a video stream according to a target audio segment for which the audio recognition result is an emotion result;
[0051] The second recognition module is used to input the video segment into the video recognition model to obtain the video recognition result;
[0052] The result determination module is used to determine the target emotion result of the audio and video to be recognized based on the audio recognition result and the video recognition result.
[0053] According to a third aspect, an electronic device is provided, comprising:
[0054] A memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the steps of any method in the first aspect of the present application.
[0055] In a fourth aspect, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the multimodal emotion recognition method shown in any one of the first aspects of the present application is implemented.
[0056] In a fifth aspect, a computer program product is provided, comprising a computer program, which implements the steps of any one of the methods in the first aspect of the present application when the computer program is executed by a processor.
[0057] The beneficial effects of the technical solution provided by the embodiments of the present application are:
[0058] The multimodal emotion recognition method provided by the present application obtains the audio and video to be recognized, performs audio emotion recognition on the audio stream in the audio and video to be recognized, and if the audio recognition result is the presence of emotion, determines the corresponding video segment in the video stream, and performs emotion recognition on the video segment to obtain a video recognition result. Setting corresponding recognition models for audio and video recognition respectively can better fit the characteristics of the specific audio or video, so that the audio recognition results and video recognition results obtained can also be more accurate, and the target emotion results obtained based on the audio recognition results and video recognition results are more reliable and accurate. BRIEF DESCRIPTION OF THE DRAWINGS
[0059] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for describing the embodiments of the present application.
[0060] Figure 1 A schematic diagram of an application scenario of a multimodal emotion recognition method provided in an embodiment of the present application;
[0061] Figure 2 A schematic diagram of a multimodal emotion recognition method provided in an embodiment of the present application;
[0062] Figure 3 A schematic diagram of a process for obtaining audio feature vectors in a multimodal emotion recognition method provided in an embodiment of the present application;
[0063] Figure 4 A flowchart illustrating an example of a multimodal emotion recognition method provided in an embodiment of the present application;
[0064] Figure 5 A schematic diagram of the structure of a multimodal emotion recognition device provided in an embodiment of the present application;
[0065] Figure 6 A schematic diagram of the structure of an electronic device applicable to the multimodal emotion recognition method provided in an embodiment of the present application. DETAILED DESCRIPTION
[0066] The following describes the embodiments of the present application in conjunction with the accompanying drawings. It should be understood that the embodiments described below in conjunction with the accompanying drawings are exemplary descriptions for explaining the technical solutions of the embodiments of the present application and do not constitute a limitation on the technical solutions of the embodiments of the present application.
[0067] Those skilled in the art will understand that, unless otherwise stated, the singular forms "a", "an", "said", and "the" used herein may also include plural forms. It should be further understood that the terms "including" and "comprising" used in the embodiments of the present application mean that the corresponding features can be implemented as the presented features, information, data, steps, operations, elements, and / or components, but do not exclude implementation as other features, information, data, steps, operations, elements, components, and / or combinations thereof supported by the present technical field. It should be understood that when we say that an element is "connected" or "coupled" to another element, the element can be directly connected or coupled to the other element, or it can refer to the element and the other element establishing a connection relationship through an intermediate element. In addition, the "connection" or "coupling" used herein may include wireless connection or wireless coupling. The terms "or", "and / or", "including at least one of the following", etc. used in this application may be interpreted as inclusive, or mean any one or any combination.
[0068] In the embodiments of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program that has a predetermined function and works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as processing circuits or memories) or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be part of an overall module or unit that includes the function of the module or unit.
[0069] In the specific implementation of this application, any data related to an object, such as data involved in the object's use of an application, is required. When the embodiments of this application are applied to specific products or technologies, the permission or consent of the object must be obtained, and the collection, use, and processing of the relevant data must comply with the relevant laws, regulations, and standards of the relevant countries and regions. In other words, if any of the above-mentioned data related to an object is involved in the embodiments of this application, such data must be obtained with the authorization and consent of the object and in compliance with the relevant laws, regulations, and standards of the relevant countries and regions.
[0070] In order to make the objectives, technical solutions and advantages of this application clearer, the implementation methods of this application will be further described in detail below with reference to the accompanying drawings.
[0071] In existing technologies, relying solely on audio or video has the limitation of incomplete information. Audio has difficulty capturing visual clues, and video may miss key emotional information in speech. Inputting audio and video into a large model for recognition means that individual differences, cultural backgrounds, diversity of expression habits, as well as factors such as lighting changes, occlusions, and background interference in real scenes, all seriously restrict the robustness and generalization ability of recognition.
[0072] The multimodal emotion recognition method, device, electronic device, computer-readable storage medium, and computer program product provided in this application are intended to solve at least one of the above technical problems in the prior art.
[0073] In response to at least one of the above-mentioned technical problems or areas that need improvement in the relevant technology, the present application proposes a multimodal emotion recognition method, device, electronic device, computer-readable storage medium and computer program product. The multimodal emotion recognition method provided by this solution obtains the audio and video to be recognized, and performs audio emotion recognition on the audio stream in the audio and video to be recognized. If the audio recognition result is that there is emotion, the corresponding video segment in the video stream is determined, and emotion recognition is performed on the video segment to obtain a video recognition result. Corresponding recognition models are set for audio and video recognition respectively, which can be more in line with the characteristics of the specific audio or video. In this way, the audio recognition results and video recognition results obtained can also be more accurate, and the target emotion results obtained based on the audio recognition results and video recognition results are more reliable and accurate.
[0074] Furthermore, based on the target emotional result, the marking style of the corresponding audio segment and / or video segment can be determined, and based on the corresponding marking style, the target emotional result can be marked at the corresponding position of the audio segment and / or video segment to be identified, so that the user can quickly locate the audio and video position corresponding to the emotional result based on the mark, which is convenient for the user to watch back and effectively improves the watching back experience.
[0075] The following describes several exemplary embodiments to illustrate the technical solutions of the embodiments of the present application and the technical effects produced by the technical solutions of the present application. It should be noted that the following embodiments can refer to, draw on, or combine with each other, and the same terms, similar features, and similar implementation steps in different embodiments will not be repeated.
[0076] Figure 1 Schematic diagram of an application scenario of the multimodal emotion recognition method provided in an embodiment of the present application, wherein the application environment may include a camera device 100 and a server 200, and the camera device 100 and the server 200 are connected via a network; a multimodal emotion recognition system may be configured in the server 200.
[0077] Specifically, the audio and video to be identified are collected by the camera device 100, and the audio and video to be marked are sent to the server 200. The server 200 obtains the audio and video to be identified, which includes an audio stream and a video stream. The audio stream is segmented to obtain at least one audio segment. Each audio segment is input into the audio recognition model to obtain an audio recognition result. According to the target audio segment whose audio recognition result is an emotion result, the corresponding video segment is determined in the video stream, and the video segment is input into the video recognition model to obtain a video recognition result. Based on the audio recognition result and the video recognition result, the target emotion result of the audio and video to be identified is determined. The above-mentioned multimodal emotion recognition method can also be executed by the multimodal emotion recognition system in the server 200.
[0078] In a specific implementation process, the server 200 may also be a terminal that can process audio and video and identify target emotion results, and a multimodal emotion recognition system may be configured on the terminal.
[0079] During the specific implementation process, when the camera device 100 is a monitoring video device, the monitoring video device performs a monitoring task. Once it captures emotional expressions such as happiness, sadness or anger in the environment, its built-in algorithm system will be triggered to quickly analyze and process the audio and video, accurately judge the characteristics and intensity of the audio, and accurately identify the emotional results.
[0080] Among them, camera equipment is a device that can capture audio and video, such as: professional cameras (including movie cameras, broadcast-level cameras), portable camcorders, smartphones (usually equipped with front and rear cameras and microphone arrays), tablets, laptops (equipped with cameras and microphones), professional recording equipment (such as voice recorders, portable audio interfaces, microphones), webcams (commonly used for video conferencing and live broadcasts), sports cameras, pan-tilt camera systems mounted on drones, virtual reality (VR) or augmented reality (AR) headsets (equipped with environment capture cameras and microphones), live streaming equipment, audio interfaces and microphones for digital audio workstations (DAWs), security surveillance cameras (some with two-way voice functions), driving recorders (some with recording functions), medical imaging equipment (such as endoscope systems, which may include audio and video capture), panoramic cameras, 3D scanners (some with audio and video recording functions), etc. These devices can be connected directly or indirectly via wired (such as USB, SDI, XLR, optical fiber, network cable) or wireless (such as Wi-Fi, Bluetooth, 5G / 4G, dedicated wireless audio and video transmission modules) communication methods to achieve synchronous or separate capture, preview, recording, transmission, or control of audio and video signals. They can work independently or as part of a larger system (such as film and television production systems, live broadcast systems, remote conferencing systems, security monitoring systems, virtual production systems), but are not limited to these systems.
[0081] The above application scenario is only an example and does not limit the application scenario of the audio and video emotion labeling method of this application.
[0082] Those skilled in the art will appreciate that a server may include a server that is equipped with a computer capable of processing database operations. The server may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It may also be a cloud server or server cluster that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. The embodiments of the present application can be applied to various scenarios, including but not limited to cloud technology, artificial intelligence, smart transportation, assisted driving, etc. The specific requirements can also be determined based on the actual application scenario and are not limited here.
[0083] The terminal can be a smart phone (such as an Android phone, an iOS phone, etc.), a tablet computer, a laptop computer, a digital broadcast receiver, a MID (Mobile Internet Devices), a PDA (Personal Digital Assistant), a desktop computer, a smart home appliance, a vehicle-mounted terminal (such as a vehicle-mounted navigation terminal, a vehicle-mounted computer, etc.), a smart speaker, a smart watch, etc. The terminal and the server can be directly or indirectly connected via wired or wireless communication, but are not limited to this.
[0084] In some possible implementations, taking the execution subject as a multimodal emotion recognition system as an example, the embodiment of the present application provides a multimodal emotion recognition method, such as Figure 2 As shown, the following steps may be included:
[0085] S210: Acquire audio and video to be identified.
[0086] The audio and video to be identified include audio stream and video stream.
[0087] Specifically, the audio and video can be collected by a camera device, which collects audio and video data simultaneously, and aligns the audio stream data and the video stream data in time to obtain the audio and video to be recognized.
[0088] During the specific implementation process, the audio and video acquisition equipment can be installed in different scenes, so the audio and video to be identified is the audio and video of the corresponding scene. For example, in the home entertainment scene, the audio and video to be obtained is the home entertainment audio and video, in the public entertainment venue, the audio and video to be obtained is the entertainment venue audio and video, and in the film and television shooting scene, the audio and video to be obtained is the shooting audio and video. The solution of this application can obtain audio data and video data in different scenes for analysis, so as to accurately identify emotional results, and has high versatility.
[0089] S220: Segment the audio stream to obtain at least one audio segment.
[0090] The segments may include fixed-length segments, content segments, or external event segments. External events may include a user clicking save or detecting a specific event.
[0091] Specifically, the audio stream can be evenly divided into segments of fixed length based on a fixed duration, for example, every 5 seconds or 10 seconds as an audio segment; it can also be segmented based on content: detect silent segments in the audio, use the silent segments as segmentation points, and use the non-silent part between two silent segments as an audio segment. If the audio is speech, it can be segmented based on semantic pauses, such as the end of a sentence or a topic change, or it can be accurately segmented based on recognized word boundaries. For music audio, it can be segmented based on changes in beats and paragraphs; it can also be segmented based on external events: the audio stream is collected in real time, and segmentation can be triggered by external events or internal state changes. For example, the user clicks the Save Current Segment button, or segmentation is performed after detecting specific sound events, such as a phone ringing or a door opening.
[0092] During the specific implementation process, if the audio stream is a real-time audio stream, the audio data is continuously read to obtain audio frames, and these data are accumulated in the memory until the segmentation conditions are met. For example, a fixed duration is accumulated, or the start or end of a silent segment is detected. The accumulated audio data is processed or stored as an audio segment, the accumulation buffer is reset, and subsequent data is continued to be read to generate audio segments.
[0093] S230: Input each audio segment into an audio recognition model to obtain an audio recognition result.
[0094] The audio recognition result may indicate whether the corresponding audio frame segment has emotion or not.
[0095] Specifically, each audio segment is input into the audio recognition model one by one in chronological order to obtain the emotion result corresponding to each audio segment, and it is judged whether each audio segment is an audio segment with emotion. Then, further emotion recognition is performed on the audio segments that may meet the emotional characteristics. Each audio segment is recognized, and background noise or non-emotion-related voice fluctuations are effectively isolated, thereby improving the quality of audio features and obtaining more accurate audio recognition results.
[0096] During the specific implementation process, the audio features of each audio segment can be obtained, and the audio segments can be input into the audio recognition model in turn to obtain the audio recognition results output by the model. The audio recognition results can indicate whether the audio segment corresponding to the audio feature is an emotion segment, and can also include information such as the emotion category and emotion intensity corresponding to each emotion segment. After obtaining the audio recognition results through the audio recognition model, the adjacent audio segments of the corresponding audio segment can be obtained. By simply analyzing the context information in the adjacent audio segments, the corresponding audio recognition results can be corrected to avoid recognition errors caused by short audio time. The context information can include text information, music information, time or space information in the audio.
[0097] In the specific implementation process, the audio recognition model may include a multidimensional convolution layer, a timing processing layer and a fully connected layer. Each audio segment is input into the audio recognition model to obtain an audio recognition result, including: performing feature engineering on each audio segment to obtain an audio feature vector, inputting the audio feature vector corresponding to each audio segment into the multidimensional convolution layer to obtain the corresponding convolution vector, inputting the convolution vector into the timing processing layer to obtain a timing feature vector, inputting the timing feature vector into the fully connected layer to map it to the corresponding emotion space, and obtaining the emotion result corresponding to each audio segment.
[0098] During the specific implementation process, multiple audio segment samples are obtained, and each audio segment sample is input into the initial recognition model one by one to obtain the emotion recognition result. The parameters of the initial recognition model are updated according to the emotion recognition result and the emotion result label until the predetermined end condition is reached, and the training is terminated to obtain a trained audio recognition model, wherein the predetermined end condition may include the loss function convergence or the accuracy rate being greater than the accuracy rate threshold. Specifically, the corresponding loss function is determined based on the emotion recognition result and the emotion result label, and the parameters of the initial recognition model are updated until the loss function converges, and the training is terminated to obtain a trained audio recognition model, or the emotion recognition result and the emotion result label are compared to determine the accuracy rate of the emotion recognition result, and the parameters of the initial recognition model are updated until the accuracy rate of the emotion recognition result is greater than the accuracy rate threshold, and the training is terminated to obtain a trained audio recognition model.
[0099] In the specific implementation process, the loss function can be a cross entropy loss function. Suppose we have an emotion recognition classification task, where y i is the encoding vector corresponding to the emotion label, indicating the true emotion category of the i-th sample, and the model's predicted output f γ (x i ) is a probability distribution, indicating that the model believes that the input x i The probability of belonging to each emotion category, the cross entropy loss function L(f θ (x i ), yi ) is as follows:
[0100]
[0101] Where C is the total number of emotion categories, θ is the model parameter, i represents the i-th sample in the dataset, j is an index variable representing the emotion category, and x i represents the i-th input sample in the dataset, y i Represents the i-th input sample x i The corresponding true emotion label, y ij is the true label of the jth category of the i-th sample, usually 0 or 1, f θ (x i ) j is the probability that the model predicts that the i-th sample belongs to the j-th category.
[0102] For the entire dataset D, the objective function J(θ) is the average of all sample losses, as shown below:
[0103]
[0104] Where N is the size of the dataset, θ is the model parameter, i represents the i-th sample in the dataset, j is an index variable representing the emotion category, and x i represents the i-th input sample in the dataset, y i Represents the i-th input sample x i The corresponding true emotion label, L(f θ (x i ), y i ) is the loss of the i-th sample.
[0105] During training, we update the model parameters θ by gradient descent to minimize the objective function J(θ):
[0106]
[0107] Where α is the learning rate, k is the kth iteration, It is the gradient of the objective function J(θ) with respect to the parameter θ. By iteratively updating the parameter θ until the loss function J(θ) converges or reaches the preset accuracy threshold, we can finally obtain a trained audio recognition model.
[0108] In the specific implementation process, the specific steps of training the audio recognition model may include: data preparation and preprocessing steps: collect audio files containing different emotional expressions (such as happiness, sadness, anger, surprise, neutrality, etc.), which can be public data sets or self-recorded speech, etc., assign accurate emotion labels to each audio segment, which can be manually labeled or labeled through a large model to ensure the accuracy of the labels, and then divide the continuous audio stream into audio segments of fixed length. These audio segments can overlap with each other to retain more contextual information. Each segment can be represented as a time domain signal or a more commonly used frequency domain feature, such as a Mel spectrogram. Then, some transformations can be performed on the audio segments, such as adding noise, changing volume, changing speed and pitch, etc., to generate more training samples. Then, the processed audio segments and their corresponding labels are divided into training set, validation set and test set, among which the training set is used for model learning, the validation set is used to adjust hyperparameters and monitor overfitting, and the test set is used to finally evaluate the model performance. Model initialization step: Select or design an initial audio recognition model architecture, which can be based on a convolutional neural network (CNN) or a recurrent neural network (RNN). Use random initialization or an initialization method based on a pre-trained model to initialize the corresponding model parameters, such as weights and bias values. Model training steps: read audio segment samples from the training set in batches, each batch contains multiple audio segments and their corresponding labels, input the audio segments of the current batch into the initial recognition model, and after processing these segments, the model outputs the probability distribution of the emotion category corresponding to each segment, for example, the probability of happiness is 0.7, sadness is 0.2, and anger is 0.1. Then, compare the emotion probability distribution output by the model with the true emotion label corresponding to the batch, and calculate the loss value. The loss function can be a cross-entropy loss function, which is used to measure the difference between the predicted distribution and the true distribution. Then, based on the calculated loss value, the gradient of the model parameters to the loss is calculated through the back propagation algorithm, and the optimization algorithm is used to update the model parameters according to the calculated gradient and the preset learning rate to reduce the loss value; after each training cycle, that is, after the entire training set is processed, the performance of the model is evaluated using the validation set, for example, by evaluating the performance through accuracy, score, etc., and judging whether the model converges or whether overfitting occurs. Termination condition judgment step: Repeat the above training cycle continuously to determine whether the predetermined termination conditions are met. The predetermined termination conditions may include: reaching a preset number of training rounds, the loss or performance indicators on the validation set not significantly improving after multiple consecutive trainings, or reaching a preset performance standard. When the termination conditions are met, the training process is stopped and the current best-performing model parameters (weights and biases) are saved to a file to obtain a trained audio recognition model.
[0109] S240 , determining a corresponding video segment in the video stream according to the target audio segment for which the audio recognition result is an emotion result.
[0110] Specifically, the target audio segment whose audio recognition result is the emotion result is determined, the video stream is segmented according to the target audio segment to obtain the corresponding video segment, the time point of the target audio segment can be mapped to the video stream, and the video stream can be divided into multiple video segments according to the mapped time point. The corresponding video segment is found based on the target audio segment with detected emotion, which can effectively improve the efficiency of video processing and emotion recognition and enhance the user experience.
[0111] In the specific implementation process, after obtaining the target audio segment, it is necessary to synchronize the target audio segment with the video stream to ensure that the time point of the segmentation accurately corresponds to the video content. The time axis of the two can be synchronized by comparing the timestamps of the target audio segment and the video stream. The similarity between the audio and image frames can also be calculated to automatically adjust the position of the image frame to align it with the audio feature points of the target audio segment.
[0112] In the specific implementation process, when performing audio recognition, the audio can be divided into multiple audio frames, and each audio frame is input into the audio recognition model one by one to obtain the results corresponding to each audio frame, and the target audio frame whose recognition result is the emotional result is determined, and the continuous target audio frames are merged to obtain the target audio segment, and the video stream is segmented according to the target audio segment, and the video segment corresponding to the start and end time is determined in the video stream. Among them, how to process the audio frame can be selected based on specific needs, such as performing feature engineering on the audio frame, obtaining audio frame vectors and inputting them into the audio recognition model one by one, etc.
[0113] S250: Input the video segment into a video recognition model to obtain a video recognition result.
[0114] The video recognition model may include a convolutional layer, a pooling layer, and at least one fully connected layer.
[0115] Specifically, feature engineering is performed on the video segment to obtain a video feature vector, the video feature vector is input into the convolution layer to obtain a video convolution vector, the video convolution vector is input into the pooling layer for average pooling or maximum pooling to obtain a video pooling vector, the video pooling vector is input into at least one fully connected layer, mapped to the corresponding emotion space, and the video recognition result is obtained.
[0116] During the specific implementation process, multiple video segment samples are obtained, and each video segment sample is input into the initial recognition model one by one to obtain an emotion recognition result. The parameters of the initial recognition model are updated according to the emotion recognition result and the video emotion result label until a predetermined end condition is reached, and the training is terminated to obtain a trained video recognition model. The predetermined end condition may include the loss function convergence or the accuracy rate being greater than the accuracy rate threshold. Specifically, the corresponding loss function is determined based on the emotion recognition result and the video emotion result label, and the parameters of the initial recognition model are updated until the loss function converges, and the training is terminated to obtain a trained video recognition model. Alternatively, the emotion recognition result and the video emotion result label are compared to determine the accuracy rate of the emotion recognition result, and the parameters of the initial recognition model are updated until the accuracy rate of the emotion recognition result is greater than the accuracy rate threshold, and the training is terminated to obtain a trained video recognition model.
[0117] In the specific implementation process, the specific steps for training the video recognition model may include: data preparation and preprocessing steps: Collect video files containing different emotional expressions, such as happiness, sadness, anger, and surprise, which can be public data sets or self-recorded voices, etc., assign accurate emotional labels to each video file or clip, which can be manually labeled or labeled through a large model to ensure the accuracy of the labels, and then divide the continuous video stream into video segments of fixed length. These video segments can overlap with each other to retain more contextual information. Each frame can be represented as a time domain signal or a more commonly used frequency domain feature. Then, some transformations can be performed on the video segments, such as adding noise, changing volume, changing speed and pitch, etc., to generate more training samples. Then, the processed video segments and their corresponding labels are divided into training set, validation set and test set, for example 70% training set, 15% validation set and 15% test set. Among them, the training set is used for model learning, the validation set is used to adjust hyperparameters and monitor overfitting, and the test set is used to finally evaluate the model performance. Model initialization: Select a model based on a convolutional neural network (CNN), a recurrent neural network (RNN), or a Transformer. Initialize the corresponding model parameters, such as weights and biases, using random initialization or initialization based on a pretrained model. Model training: Read video segments from the training set in batches. Each batch contains multiple video segments and their corresponding labels. Input the current batch of video segments into the initial recognition model. After processing these frames, the model outputs a probability distribution of the emotion category corresponding to each frame (for example, the probability of happiness is 0.7, sadness is 0.2, and anger is 0.1). The emotion probability distribution output by the model is then compared with the true emotion labels corresponding to the batch, and a loss value is calculated. The loss function can be a cross-entropy loss function, which measures the difference between the predicted distribution and the true distribution. Based on the calculated loss value, the gradient of the model parameters with respect to the loss is calculated using the backpropagation algorithm. Using the calculated gradient and the preset learning rate, the optimization algorithm updates the model parameters to reduce the loss value. After each training cycle, that is, after the entire training set has been processed, the model performance is evaluated on the validation set. For example, performance is evaluated through accuracy and scores to determine whether the model has converged or is overfitting. Termination condition judgment step: Repeat the above training cycle continuously to determine whether the predetermined termination conditions are met. The predetermined termination conditions may include: reaching a preset number of training rounds, the loss or performance indicators on the validation set have not significantly improved in multiple consecutive trainings, or reaching a preset performance standard, etc. When the termination conditions are met, stop the training process, save the current best performance model parameters (weights and biases) to a file, and obtain a trained video recognition model.
[0118] In the specific implementation process, the loss function can be a cross entropy loss function. Suppose we have an emotion recognition classification task, where y iis the encoding vector corresponding to the emotion label, indicating the true emotion category of the i-th sample, and the model's predicted output f θ (x i ) is a probability distribution, indicating that the model believes that the input x i The probability of belonging to each emotion category, the cross entropy loss function L(f θ (x i ), y i ) is as follows:
[0119]
[0120] Where C is the total number of emotion categories, θ is the model parameter, i represents the i-th sample in the dataset, j is an index variable representing the emotion category, and x i represents the i-th input sample in the dataset, y i Represents the i-th input sample x i The corresponding true emotion label, y ij is the true label of the jth category of the i-th sample, usually 0 or 1, f θ (x i ) j is the probability that the model predicts that the i-th sample belongs to the j-th category.
[0121] For the entire dataset D, the objective function J(θ) is the average of all sample losses, as shown below:
[0122]
[0123] Where N is the size of the dataset, θ is the model parameter, i represents the i-th sample in the dataset, j is an index variable representing the emotion category, and x i represents the i-th input sample in the dataset, y i Represents the i-th input sample x i The corresponding true emotion label, L(f θ (x i ), y i ) is the loss of the i-th sample.
[0124] During training, we update the model parameters θ by gradient descent to minimize the objective function J(θ):
[0125]
[0126] Where α is the learning rate, k is the kth iteration, It is the gradient of the objective function J(θ) with respect to the parameter θ. By iteratively updating the parameter θ until the loss function J(θ) converges or reaches the preset accuracy threshold, we can finally obtain a trained video recognition model.
[0127] S260: Determine a target emotion result of the audio and video to be recognized based on the audio recognition result and the video recognition result.
[0128] Among them, both the audio recognition results and the video recognition results include corresponding emotion types.
[0129] Specifically, determine the first emotion type corresponding to the audio recognition result, and determine the second emotion type corresponding to the video recognition result. If the first emotion type is inconsistent with the second emotion type, weighted average or select a target emotion result with higher confidence based on the confidence levels output by the two models. For example, if the first emotion type recognized by the audio is happy 0.6, and the second emotion type corresponding to the expression recognized by the video is sad, with a confidence level of 0.4, then happy emotion is determined to be the target emotion result.
[0130] During the specific implementation process, when the first emotion type is inconsistent with the second emotion type and the target emotion result cannot be selected based on the confidence level, the corresponding audio and video segment is determined in the audio and video to be identified based on the audio segment, and the audio and video segment is marked as an unknown emotion or mixed emotion, avoiding giving a single result that may be wrong or misleading. The corresponding result can also be displayed to the user, and the user can customize the corresponding emotional content to facilitate subsequent review of the audio and video segment.
[0131] In some possible implementations, in the above steps, inputting each audio segment into an audio recognition model to obtain an audio recognition result includes:
[0132] Perform feature engineering on each audio segment to obtain the audio feature vector corresponding to each audio segment;
[0133] The audio feature vectors corresponding to each audio segment are input into the audio recognition model one by one to obtain the audio recognition results corresponding to each audio segment.
[0134] The audio feature vector includes the frequency change, amplitude change, pitch and timbre of the audio segment.
[0135] Specifically, feature engineering can be performed on each audio segment to obtain an audio feature vector corresponding to each audio segment, and the audio feature vectors corresponding to each audio segment can be input into the audio recognition model one by one to obtain an audio recognition result corresponding to each audio segment; or each audio segment can be divided into multiple audio frames, and feature engineering can be performed on each audio frame to obtain an audio feature vector corresponding to each audio frame, and each audio feature vector can be input into the audio recognition model one by one to obtain an audio recognition result corresponding to each audio frame, and the audio frames within an audio segment can be clustered according to the time sequence and the audio recognition result to obtain a clustering result, wherein the clustering result includes multiple cluster clusters, and different cluster clusters correspond to different emotion types. Based on the clustering result, the audio segment can be re-divided into at least one sub-audio segment to obtain an audio recognition result for the sub-audio segment.
[0136] During the specific implementation process, feature engineering is performed on each audio frame to obtain an audio feature vector corresponding to each audio frame, including: obtaining multiple audio frames, using a window function to perform weighted processing on each frame of data to reduce spectral leakage and improve the accuracy of frequency domain analysis, and then performing high-pass filtering on the signal to enhance the high-frequency part and compensate for the attenuation of high-frequency energy in the speech signal. Representative and discriminative feature vectors are then extracted from the preprocessed audio frames, and the most useful features for the task are selected based on statistical indicators or model evaluation to reduce the dimension. Methods such as principal component analysis (PCA) or linear discriminant analysis (LDA) are used to reduce the feature dimension while retaining the main information. The features are standardized to have the same scale to avoid certain features dominating the model training due to their large values. The data volume of the audio feature vector obtained in this way is smaller than that of ordinary features, but it is more efficient in expressing data features, effectively improving the accuracy of emotion recognition.
[0137] In the specific implementation process, the emotion recognition model includes a multi-dimensional convolution layer, a time series processing layer and a fully connected layer. The audio feature vectors corresponding to each audio frame are input into the emotion recognition model one by one to obtain the emotion results corresponding to each audio frame. Specifically, the audio feature vectors corresponding to each audio frame are input into the multi-dimensional convolution layer to obtain the corresponding convolution vector; the convolution vector is input into the time series processing layer to obtain the time series feature vector; the time series feature vector is input into the fully connected layer to map it to the corresponding emotion category space to obtain the emotion results corresponding to each audio frame. Specifically, the audio feature vectors corresponding to each audio frame are input into the multi-dimensional convolution layer, and the convolution layer is used to find patterns and extract more complex features on the audio feature vector to obtain the corresponding convolution vector, and the convolution vector is input into The time series processing layer uses the spatial local features captured by the convolutional layer as input, and then uses the time series processing layer to analyze and understand the changes and relationships of these features in the time dimension, thereby achieving a deeper understanding of the spatiotemporal data and obtaining a time series feature vector. The time series feature vector is input into the fully connected layer, and the fully connected layer maps the time series feature vector into another vector. The dimension and meaning of this new vector correspond to the emotion category space. It determines whether the mapped vector is in the emotion category space and which emotion category space it is in, thereby obtaining the emotion result of the corresponding audio frame; for example, if we want to distinguish 10 different emotion categories, the mapped vector may be 10-dimensional, and each dimension represents a score or probability of the category.
[0138] In the specific implementation process, the audio frames containing emotions can be determined based on the audio recognition model, the adjacent frames of the audio frames with emotions can be determined, the emotion results of the adjacent frames can be compared with the emotion result of the audio frame, and the adjacent audio frames with matching emotion results can be merged to obtain audio segments with the same emotion and continuous emotional state. The audio segments marked in this way can contain complete emotion fragments. Since the audio frames are input into the audio recognition model one by one in chronological order, when the emotion result of an audio frame is that it contains emotions, the audio frame is used as the starting frame, and the emotion results of the audio frames output next are compared with the starting frame in turn. The audio frames with matching emotion results are used as frame sets until it is detected that the emotion result of an audio frame does not match, and the obtained frame sets are merged in chronological order to generate a new audio segment.
[0139] For example, if the emotional result of any audio frame is a happy emotion, the emotional result of the adjacent audio frame is determined. If the emotional result of the adjacent audio frame is also a happy emotion, and the similarity is greater than the similarity threshold, the audio frame can be merged with the adjacent audio frame. If the emotional result of the adjacent audio frame is a happy emotion, but the similarity is less than the similarity threshold, or the emotional result of the adjacent audio frame is no emotion or other emotions, the adjacent audio frame is discarded; if the similarity of the adjacent audio frames does not meet the similarity threshold, the audio frame can be marked or discarded. The specific setting can be made according to actual needs.
[0140] During the specific implementation process, the context of the audio stream, such as the words just said, background sound, and even other modal information such as facial expressions in the video, can be combined to infer the scene where laughter is most likely to occur. The same emotion category corresponds to different audio and video segment lengths in different scenes. For example, these differences are reflected in the following aspects: In a relaxed and enjoyable social gathering, such as watching TV and chatting, laughter is usually triggered by humorous jokes, interesting experiences or common topics. This kind of laughter is often more natural, loud, and lasts longer. In embarrassing, tense or trying to ease the atmosphere, laughter may be shorter, suppressed, and hesitant. Therefore, it makes sense to adjust the audio frames based on the scene, which can more accurately understand and divide emotional segments. Including scene information in the training stage of the audio recognition model can further improve the accuracy of the audio recognition results.
[0141] In some possible implementations, such as Figure 3 As shown, in the above steps, feature engineering is performed on each audio segment to obtain the audio feature vector corresponding to each audio segment, including:
[0142] S310, dividing each audio segment into frames to obtain multiple audio frames;
[0143] S320, performing frequency analysis on each audio frame to obtain a frequency feature vector;
[0144] S330, performing amplitude analysis on each audio frame to obtain an amplitude feature vector;
[0145] S340, performing pitch analysis on each audio frame to obtain a pitch feature vector;
[0146] S350, performing timbre analysis on each audio frame to obtain a timbre feature vector;
[0147] S360: Aggregate the frequency feature vector, amplitude feature vector, pitch feature vector, and timbre feature vector corresponding to each audio frame to obtain an audio feature vector corresponding to each audio segment.
[0148] Among them, the frequency feature vector includes spectrum centroid data, spectrum dispersion data, spectrum slope data and spectrum rolling data; the amplitude feature vector includes instantaneous amplitude, amplitude envelope, energy data, zero-crossing rate data and loudness data; the pitch feature vector includes fundamental frequency data, fundamental frequency time series data and harmonic data; the timbre feature vector includes Mel-frequency cepstral coefficient data, chroma feature data and linear prediction coefficient data.
[0149] Specifically, preset frame configuration information is obtained, and the audio segment is divided into frames based on the frame configuration information. The frame configuration information may include frame length and inter-frame movement distance. The audio segment is divided based on the frame length and inter-frame movement distance to obtain multiple audio frames.
[0150] During the specific implementation process, frame configuration information is obtained, and a loop or vectorized operation is used. Starting from the starting position of the audio segment, the system moves with a step size of the inter-frame moving distance, and each time a piece of audio data with a length of the frame is extracted as a frame until the end of the audio segment. Emotion recognition is performed through frames, focusing on local sounds, effectively isolating background noise or non-emotion-related voice fluctuations, making emotion recognition more accurate.
[0151] During the specific implementation process, frequency feature vectors, amplitude feature vectors, pitch feature vectors and timbre feature vectors are obtained, and the frequency feature vectors, amplitude feature vectors, pitch feature vectors and timbre feature vectors are spliced to obtain a spliced vector. According to the adjacent audio frames in each time, the time series features of each spliced vector are enhanced, and the spliced vectors are nonlinearly transformed to obtain the audio feature vectors corresponding to each audio frame. The vectors are spliced together, and the obtained high-dimensional vector can represent the content of the original audio frame more comprehensively and finely, providing strong data support for the subsequent audio frame emotion recognition task.
[0152] In the specific implementation process, the spectral centroid data measures the mass center or average frequency of the audio spectrum. A high spectral centroid usually means that the sound sounds brighter, sharper and thinner; a low spectral centroid means that the sound is darker, deeper and thicker. Excited and happy emotions correspond to higher spectral centroids, while sad and calm emotions correspond to lower spectral centroids. The spectral dispersion data measures the distribution width or dispersion of spectral energy around the spectral centroid. High spectral dispersion means that the spectral energy is widely distributed, and the sound sounds rougher, noisier or contains more overtones; low spectral dispersion means that the spectral energy is concentrated, and the sound sounds purer, clearer or thinner. The spectral slope data describes the rate at which spectral energy decays or increases with increasing frequency, that is, the slope of the spectral curve. A negative slope diagonally downward means that high-frequency energy decays faster than low-frequency energy, and the sound sounds dimmer; a positive slope diagonally upward means that high-frequency energy is relatively stronger; the spectral roll data refers to the percentage of the sum of the energy of all frequency components below a certain frequency in the spectrum to the total energy. The specific frequency is the roll point. A lower roll point means that most of the energy is concentrated in the low frequency, and the sound sounds deeper; a higher roll point means that It means that the energy distribution is wider or biased towards high frequencies, and the sound sounds brighter; instantaneous amplitude refers to the instantaneous value of the audio signal at a specific time point; the amplitude envelope describes the overall profile or trend of the audio signal amplitude changing over time, capturing the macroscopic pattern of sound intensity changes and ignoring rapid oscillations; energy data refers to the sum (or average) of the squares of the audio signal amplitude within a time window, that is, the signal strength or power; zero-crossing rate data refers to the number of times the audio signal crosses the zero level (from positive to negative or from negative to positive) per unit time; loudness data is to evaluate the difference between loudness and softness heard by the human ear; basic Frequency data refers to the lowest frequency component of the sound waveform, that is, the fundamental frequency of the vocal cord vibration, which determines the speaker's pitch; fundamental frequency time series data provides dynamic information about pitch changes, for example, it can analyze the fluctuations, trends, and change rates of pitch; harmonic data refers to the frequency components that are integer multiples of the fundamental frequency; Mel-frequency cepstral coefficient data simulates the human ear's nonlinear perception of sound frequency (Mel scale), and then extracts the key coefficients that can represent the shape of the spectrum through cepstral transformation; chromatic feature data is used to identify music in audio, focusing on the pitch category and intensity of the music; linear prediction coefficient data is used to evaluate timbre.
[0153] In some possible implementations, the above steps of inputting the video segment into the video recognition model to obtain the video recognition result include:
[0154] Perform feature engineering on each video segment to obtain the video feature vector corresponding to each video segment;
[0155] The video feature vectors corresponding to each video segment are input into the video recognition model one by one to obtain the video recognition results corresponding to each video segment.
[0156] The video feature vector includes the face change features, action change features and scene change features of the video segment.
[0157] Specifically, each video segment can be divided into frames to obtain multiple image frames, and feature engineering can be performed on each image frame to obtain a video feature vector corresponding to each image frame. The video feature vector corresponding to each image frame can be input into the video recognition model one by one to obtain an image recognition result corresponding to each image frame, thereby determining the image recognition result corresponding to the corresponding video segment.
[0158] In the specific implementation process, the video feature vectors corresponding to each image frame are input into the video recognition model one by one to obtain the image recognition results corresponding to each image frame, specifically including: inputting the video feature vectors corresponding to each image frame into the multidimensional convolution layer to obtain the corresponding convolution vector; inputting the convolution vector into the time series processing layer to obtain the time series feature vector; inputting the time series feature vector into the fully connected layer to map it to the corresponding emotion category space to obtain the image recognition results corresponding to each image frame; inputting the video feature vectors corresponding to each image frame into the multidimensional convolution layer, finding patterns on the video feature vectors through the convolution layer, extracting more complex features, and obtaining the corresponding emotion category space. The corresponding convolution vector is input into the time series processing layer, and the spatial local features captured by the convolution layer are used as input. Then, the time series processing layer is used to analyze and understand the changes and relationships of these features in the time dimension, so as to achieve a deeper understanding of the spatiotemporal data and obtain the time series feature vector. The time series feature vector is input into the fully connected layer, and the fully connected layer maps the time series feature vector into another vector. The dimension and meaning of this new vector correspond to the emotion category space. It is judged whether the mapped vector is in the emotion category space and which emotion category space it is in, so as to obtain the image recognition result of the corresponding image frame.
[0159] During the specific implementation process, the video feature vector may also include facial expression features, head posture features and eye tracking data to represent the facial expressions, nodding and shaking of the head and eye movement status of the characters in the video.
[0160] In some possible implementations, after inputting the video segment into the video recognition model and obtaining the video recognition result in the above steps, the following steps may be further included:
[0161] Input each image frame into the image recognition model to obtain the image recognition result and the corresponding confidence level;
[0162] Based on the image recognition result and the corresponding confidence level, obtaining the image recognition result and the corresponding image recognition confidence level of the video segment;
[0163] Get the video recognition confidence of the video recognition result;
[0164] If the image recognition confidence is greater than the video recognition confidence, the video recognition result is updated based on the image recognition result.
[0165] The video segment includes multiple image frames.
[0166] Specifically, each image frame is input into the image recognition model to obtain the image recognition result and the corresponding confidence. Based on the image recognition result and the corresponding confidence of each image frame, the image recognition result and the corresponding image recognition confidence of the video segment are obtained, and the video recognition confidence of the video recognition result is obtained. If the image recognition confidence is greater than the video recognition confidence, the video recognition result is updated based on the image recognition result, that is, a result is obtained based on the video segment, and a result is obtained by analyzing the image frames in the video segment. The result with the highest confidence is selected as the video recognition result, which can improve the accuracy and credibility of the video recognition result.
[0167] In some possible implementations, the steps of inputting each image frame into an image recognition model to obtain an image recognition result include:
[0168] Perform feature engineering on each image frame to obtain the image feature vector corresponding to each image frame;
[0169] The image feature vectors corresponding to each image frame are input into the image recognition model one by one to obtain the image recognition results.
[0170] The image feature vector includes the facial features, action features and scene features of the image frame.
[0171] Specifically, feature engineering is performed on each image frame to obtain an image feature vector corresponding to each image frame, and the image feature vector corresponding to each image frame is input into the image recognition model one by one to obtain an image recognition result, wherein the video feature vector corresponding to each image frame is input into the video recognition model one by one to obtain an image recognition result corresponding to each image frame, specifically including: inputting the video feature vector corresponding to each image frame into the multidimensional convolution layer to obtain the corresponding convolution vector; inputting the convolution vector into the time series processing layer to obtain the time series feature vector; inputting the time series feature vector into the fully connected layer to map it to the corresponding emotion category space to obtain the image recognition result corresponding to each image frame, and inputting the video feature vector corresponding to each image frame into the multidimensional convolution layer , find patterns on the video feature vector through the convolution layer, extract more complex features, and obtain the corresponding convolution vector. The convolution vector is input into the time series processing layer, and the spatial local features captured by the convolution layer are used as input. Then, the time series processing layer is used to analyze and understand the changes and relationships of these features in the time dimension, thereby achieving a deeper understanding of spatiotemporal data and obtaining a time series feature vector. The time series feature vector is input into the fully connected layer, and the fully connected layer maps the time series feature vector into another vector. The dimension and meaning of this new vector correspond to the emotion category space. It is judged whether the mapped vector is in the emotion category space and which emotion category space it is in, thereby obtaining the image recognition result of the corresponding image frame.
[0172] During the specific implementation process, image features can be obtained based on the scene combined with sensor information other than audio. For example, in a video scene, facial expression recognition can be used to identify information such as raised corners of the mouth and wrinkles around the eyes. Head posture recognition can be used to identify nodding, shaking, or tilting the head. Eye tracking recognition can be used to identify whether the eyes are focused or looking around. In the wearable device scene, facial muscle activity can also be detected through electromyography signals, and emotional state can be reflected through heart rate variability.
[0173] In some possible implementations, the above steps of determining the target emotion result of the audio or video to be recognized based on the audio recognition result and the image recognition result include:
[0174] determining a first confidence level of the audio recognition result and a second confidence level of the video recognition result;
[0175] If the audio recognition result is inconsistent with the video recognition result, the first confidence level and the second confidence level are compared, and the recognition result with the higher confidence level is used as the target emotion result.
[0176] Specifically, when the audio recognition model outputs the audio recognition result, it can also synchronously output the first confidence level of the recognition result. When the video recognition model outputs the video recognition result, it can also input the second confidence level of the recognition result. If the audio recognition result is inconsistent with the video recognition result, the confidence levels corresponding to the two recognition results are compared, and the recognition result with higher confidence level is used as the target emotion result.
[0177] For example, in real-time detection of laughter emotions in audio and video, if suspected laughter emotions are detected in the audio, but no laughing expression or action is recognized in the video, the position of the camera device can be adjusted based on the loudness of the sound to obtain effective video recognition results, thereby obtaining more accurate target emotion results.
[0178] During the specific implementation process, the confidence level may appear in the form of a probability value. The model outputs a probability value for each possible emotion category, such as happiness, sadness, anger, etc. The category with the highest confidence level is the corresponding recognition result. The highest confidence levels of two different models are compared to obtain the target emotion result. A confidence threshold can be set for each model. If the confidence level of an emotion category is lower than the threshold, it may be judged as unrecognizable or uncertain. During the specific implementation process, if the audio recognition result is inconsistent with the video recognition result, it is determined whether the video recognition result is that no person or face is detected. That is, the confidence level of the prediction result of the emotion recognition model is compared with the confidence threshold. If it is determined to be unrecognizable or uncertain, the person or face is not detected. At this time, the person may be outside the video, with his back blocked for most of the view, or blurred and difficult to identify, resulting in the video being unable to detect valid data. At this time, the audio recognition result can be selected as the target emotion result, or the position of the camera device can be adjusted to collect valid data before making a judgment.
[0179] In some possible implementations, the above steps of determining a corresponding video segment in the video stream based on the target audio segment for which the audio recognition result is an emotion result include:
[0180] Determine a target audio segment for which the audio recognition result is an emotion result;
[0181] The video stream is segmented according to the target audio segment to obtain the corresponding video segment.
[0182] The start and end times of the video segment correspond to the start and end times of the audio segment. The audio recognition results can include emotion results or other results; emotion results are further divided into emotions such as happiness, sadness, anger, and fear.
[0183] Specifically, determine the target audio segment whose audio recognition result is the emotion result. After obtaining the target audio segment, synchronize the target audio segment with the video stream to ensure that the time point of the segmentation accurately corresponds to the video content. The time axis of the target audio segment and the video stream can be synchronized by comparing the timestamps of the target audio segment and the video stream. The position of the image frame can also be automatically adjusted by calculating the similarity between the audio and image frames to align it with the audio feature points of the target audio segment. Then, segment the video stream according to the target audio segment to obtain the corresponding video segment. The time point of the target audio segment can be mapped to the video stream. According to the mapped time point, the video stream can be divided into multiple video segments. Finding the corresponding video segment based on the target audio segment with detected emotions can effectively improve the efficiency of video processing and emotion recognition and enhance the user experience.
[0184] In the specific implementation process, when performing audio recognition, the audio can be divided into multiple audio frames, and each audio frame is input into the audio recognition model one by one to obtain the results corresponding to each audio frame, and the target audio frame whose recognition result is the emotional result is determined, and the continuous target audio frames are merged to obtain the target audio segment, and the video stream is segmented according to the target audio segment, and the video segment corresponding to the start and end time is determined in the video stream. Among them, how to process the audio frame can be selected based on specific needs, such as performing feature engineering on the audio frame, obtaining audio frame vectors and inputting them into the audio recognition model one by one, etc.
[0185] In some possible implementations, the above method further includes:
[0186] Based on the target emotion result, determining a tagging style for the corresponding audio segment and / or video segment;
[0187] Based on the corresponding marking style, the target emotion result is marked at the corresponding position of the audio segment and / or video segment to be identified.
[0188] Specifically, you can mark the corresponding positions of the audio segment or video segment respectively, determine the corresponding marking style based on the preset marking style table according to the target emotion result, and mark the target emotion result at the corresponding position of the audio segment and / or video segment to be identified with the corresponding marking style; you can also determine the corresponding audio and video segment in the audio and video to be identified based on the audio segment or video segment, determine the marking style corresponding to the audio and video segment, and mark the target emotion result in the audio and video, mark each audio and video segment, so that users can quickly locate the audio and video segment they want to view, improve the efficiency of video playback positioning, and enhance user experience.
[0189] For example, taking the recognition of laughter as an example, the camera equipment simultaneously collects audio and video, including time-synchronized video and audio. When laughter is recognized based on audio, the corresponding timestamp information can be quickly determined, and based on the received timestamp, the location of the laughter is accurately marked on the time axis corresponding to the audio and video. At the same time, the marking information will be properly stored to ensure that it can be displayed stably and accurately during video playback.
[0190] During the specific implementation process, the emotional results can include emotional categories. Different audio segments, video segments or audio and video segments correspond to different emotional categories, and can be distinguished and marked in different styles. For example, the audio and video segments with happy emotions are marked with yellow highlights, and the corresponding emotional results are marked with yellow words. The audio and video segments with sad emotions are marked with blue, and the corresponding emotional results are marked with black words. It can be set according to user preferences, so that users can intuitively see different types of audio and video segments in the audio and video, so as to quickly find the segments they want to view.
[0191] During the specific implementation process, a playback timeline can be set for audio and video, and the audio or video to be marked can be segmented according to the recognition results, including the audio or video segments of the emotions, to obtain the corresponding audio and video segments, and the marking style corresponding to the audio and video segments can be determined. The audio and video segments are marked with the corresponding marking style on the timeline segment, and then the position of the frames of the audio and video segments on the timeline is determined, and the corresponding emotional results are marked with the corresponding style at the corresponding position on the timeline. The user can click on the timeline to jump directly to the corresponding audio and video segment for playback.
[0192] In some possible implementations, the above method further includes:
[0193] In response to a user's request to view the marked audio or video, displaying a marked list of audio segments or a marked list of video segments of the marked audio or video;
[0194] In response to the user's selection of a tag for a target emotion result, the corresponding audio segment and / or video segment is played.
[0195] The tag list includes target emotion result tags sorted by time; the viewing request may include clicking a view button, searching for keywords, or filtering emotion fields.
[0196] Specifically, in response to a user's request to view the marked audio and video, a marked summary is displayed. The marked summary may include a list of audio segment marks or a list of video segment marks of the audio and video, and may also include a list of audio and video segment marks corresponding to the audio segment or video segment in the audio and video. In response to the user's selection operation on any mark in the mark list, the corresponding audio segment and / or video segment is played, wherein the audio and video segment may be a combination of audio segments and video segments with the same start and end time.
[0197] Specifically, the system receives a user's request to view audio or video and displays a list of tagged audio or video segments. Users can click a tagged entry in the list to jump to the corresponding audio or video segment for playback, allowing them to quickly locate the segment of interest and meet user needs. After obtaining the tagged audio or video, the system can extract the audio or video segment corresponding to the desired emotion based on the tags in the audio or video, and use it for video editing or repeated playback.
[0198] In the specific implementation process, emotional result tags can have different functions in different scenarios. For example, in family scenarios, the camera equipment installed in the home can continuously record the happy moments of family members, which is convenient for reviewing and editing these interesting contents into wonderful family short videos; in public entertainment scenarios, the wonderful moments of live performances and interactive sessions can be recorded, which is convenient for editing wonderful promotional videos; in film and television recording scenarios, this function can be used to quickly locate the wonderful clips during the actor's performance, so that when editing the promotional video later, the material can be efficiently screened, effectively improving the creative efficiency and quality of the work.
[0199] In the above embodiment, by obtaining the audio and video to be identified, audio emotion recognition is performed on the audio stream in the audio and video to be identified. If the audio recognition result is the presence of emotion, the corresponding video segment in the video stream is determined, and emotion recognition is performed on the video segment to obtain a video recognition result. Setting corresponding recognition models for audio and video recognition respectively can better fit the characteristics of the specific audio or video. In this way, the audio recognition results and video recognition results obtained can also be more accurate, and the target emotion results obtained based on the audio recognition results and video recognition results are more reliable and accurate.
[0200] Furthermore, based on the target emotional result, the marking style of the corresponding audio segment and / or video segment can be determined, and based on the corresponding marking style, the target emotional result can be marked at the corresponding position of the audio segment and / or video segment to be identified, so that the user can quickly locate the audio and video position corresponding to the emotional result based on the mark, which is convenient for the user to watch back and effectively improves the watching back experience.
[0201] In one example, the multimodal emotion recognition method of the present application, such as Figure 4 As shown, this may include:
[0202] S410: Acquire audio and video to be identified.
[0203] The audio and video to be identified include audio stream and video stream.
[0204] S420, segmenting the audio stream to obtain at least one audio segment;
[0205] S430, inputting each audio segment into an audio recognition model to obtain an audio recognition result;
[0206] S440, determining a corresponding video segment in the video stream according to the target audio segment for which the audio recognition result is an emotion result;
[0207] S450, inputting the video segment into a video recognition model to obtain a video recognition result;
[0208] S460, determining a first confidence level of the audio recognition result and a second confidence level of the video recognition result;
[0209] S470: If the audio recognition result is inconsistent with the video recognition result, the first confidence level and the second confidence level are compared, and the recognition result with the higher confidence level is used as the target emotion result.
[0210] The above-mentioned multimodal emotion recognition method obtains the audio and video to be recognized, performs audio emotion recognition on the audio stream in the audio and video to be recognized, and if the audio recognition result is the presence of emotion, determines the corresponding video segment in the video stream, and performs emotion recognition on the video segment to obtain the video recognition result. Setting corresponding recognition models for audio and video recognition respectively can better fit the characteristics of the specific audio or video, so that the audio recognition results and video recognition results obtained can also be more accurate, and the target emotion results obtained based on the audio recognition results and video recognition results are more reliable and accurate.
[0211] Furthermore, based on the target emotional result, the marking style of the corresponding audio segment and / or video segment can be determined, and based on the corresponding marking style, the target emotional result can be marked at the corresponding position of the audio segment and / or video segment to be identified, so that the user can quickly locate the audio and video position corresponding to the emotional result based on the mark, which is convenient for the user to watch back and effectively improves the watching back experience.
[0212] The present application embodiment provides a multimodal emotion recognition device, such as Figure 5 As shown, the multimodal emotion recognition device 50 may include: an audio and video acquisition module 510, a segmentation module 520, a first recognition module 530, a video segment determination module 540, a second recognition module 550 and a result determination module 560, wherein:
[0213] The audio and video acquisition module 510 is used to acquire the audio and video to be identified; the audio and video to be identified include audio streams and video streams;
[0214] A segmentation module 520, configured to segment the audio stream to obtain at least one audio segment;
[0215] A first recognition module 530 is configured to input each audio segment into an audio recognition model to obtain an audio recognition result;
[0216] A video segment determination module 540 is configured to determine a corresponding video segment in a video stream according to a target audio segment for which the audio recognition result is an emotion result;
[0217] The second recognition module 550 is used to input the video segment into the video recognition model to obtain a video recognition result;
[0218] The result determination module 560 is used to determine the target emotion result of the audio and video to be recognized based on the audio recognition result and the video recognition result.
[0219] As an optional embodiment, in the device, the first identification module 530 is specifically configured to:
[0220] Perform feature engineering on each audio segment to obtain the audio feature vector corresponding to each audio segment; the audio feature vector includes the frequency change, amplitude change, pitch and timbre of the audio segment;
[0221] The audio feature vectors corresponding to each audio segment are input into the audio recognition model one by one to obtain the audio recognition results corresponding to each audio segment.
[0222] As an optional embodiment, in the device, the first identification module 530 is specifically configured to:
[0223] Frame each audio segment to obtain multiple audio frames;
[0224] Perform frequency analysis on each audio frame to obtain a frequency feature vector; the frequency feature vector includes spectrum centroid data, spectrum dispersion data, spectrum slope data, and spectrum rolling data;
[0225] Perform amplitude analysis on each audio frame to obtain an amplitude feature vector; the amplitude feature vector includes instantaneous amplitude, amplitude envelope, energy data, zero-crossing rate data, and loudness data;
[0226] Perform pitch analysis on each audio frame to obtain a pitch feature vector; the pitch feature vector includes fundamental frequency data, fundamental frequency time series data and harmonic wave data;
[0227] Performing timbre analysis on each audio frame to obtain a timbre feature vector; the timbre feature vector includes Mel-frequency cepstral coefficient data, chroma feature data, and linear prediction coefficient data;
[0228] The frequency feature vector, amplitude feature vector, pitch feature vector and timbre feature vector corresponding to each audio frame are aggregated to obtain the audio feature vector corresponding to each audio segment.
[0229] As an optional embodiment, in the device, the second identification module 550 is specifically configured to:
[0230] Perform feature engineering on each video segment to obtain the corresponding video feature vector of each video segment; the video feature vector includes the face change features, action change features and scene change features of the video segment;
[0231] The video feature vectors corresponding to each video segment are input into the video recognition model one by one to obtain the video recognition results corresponding to each video segment.
[0232] As an optional embodiment, in the device, the second identification module 550 is specifically configured to:
[0233] Input each image frame into the image recognition model to obtain the image recognition result and the corresponding confidence level;
[0234] Based on the image recognition result and the corresponding confidence level, obtaining the image recognition result and the corresponding image recognition confidence level of the video segment;
[0235] Get the video recognition confidence of the video recognition result;
[0236] If the image recognition confidence is greater than the video recognition confidence, the video recognition result is updated based on the image recognition result.
[0237] As an optional embodiment, in the device, the second identification module 550 is specifically configured to:
[0238] Perform feature engineering on each image frame to obtain an image feature vector corresponding to each image frame; the image feature vector includes facial features, action features, and scene features of the image frame;
[0239] The image feature vectors corresponding to each image frame are input into the image recognition model one by one to obtain the image recognition results.
[0240] As an optional embodiment, in the device, the result determination module 560 is specifically configured to:
[0241] determining a first confidence level of the audio recognition result and a second confidence level of the video recognition result;
[0242] If the audio recognition result is inconsistent with the video recognition result, the first confidence level and the second confidence level are compared, and the recognition result with the higher confidence level is used as the target emotion result.
[0243] As an optional embodiment, in the apparatus, the video segment determination module 540 is specifically configured to:
[0244] Determine a target audio segment for which the audio recognition result is an emotion result;
[0245] The video stream is segmented according to the target audio segment to obtain corresponding video segments; the start and end times of the video segments correspond to the start and end times of the audio segments.
[0246] As an optional embodiment, the device further includes a marking module, which is specifically configured to:
[0247] Based on the target emotion result, determining a tagging style for the corresponding audio segment and / or video segment;
[0248] Based on the corresponding marking style, the target emotion result is marked at the corresponding position of the audio segment and / or video segment to be identified.
[0249] As an optional embodiment, the device further includes a tag viewing module, which is specifically configured to:
[0250] In response to a user's request to view the marked audio or video, a marked list of audio segments or a marked list of video segments of the marked audio or video is displayed; the marked list includes target emotion result marks sorted by time;
[0251] In response to the user's selection of a tag for a target emotion result, the corresponding audio segment and / or video segment is played.
[0252] The multimodal emotion recognition device provided by the present application obtains the audio and video to be recognized, performs audio emotion recognition on the audio stream in the audio and video to be recognized, and if the audio recognition result is the presence of emotion, determines the corresponding video segment in the video stream, and performs emotion recognition on the video segment to obtain a video recognition result. Setting corresponding recognition models for audio and video recognition respectively can better fit the characteristics of the specific audio or video, so that the audio recognition results and video recognition results obtained can also be more accurate, and the target emotion results obtained based on the audio recognition results and video recognition results are more reliable and accurate.
[0253] Furthermore, based on the target emotional result, the marking style of the corresponding audio segment and / or video segment can be determined, and based on the corresponding marking style, the target emotional result can be marked at the corresponding position of the audio segment and / or video segment to be identified, so that the user can quickly locate the audio and video position corresponding to the emotional result based on the mark, which is convenient for the user to watch back and effectively improves the watching back experience.
[0254] The devices of the embodiments of the present application can execute the methods provided in the embodiments of the present application, and their implementation principles are similar and have corresponding technical effects. The actions performed by each module in the devices of the embodiments of the present application correspond to the steps in the methods of the embodiments of the present application. For detailed functional descriptions of each module of the device, please refer to the descriptions of the corresponding methods shown above, and will not be repeated here.
[0255] In an embodiment of the present application, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory, and the processor executes the above-mentioned computer program to implement the steps of the method provided in any optional embodiment of the present application. Compared with the prior art, it can be achieved: by obtaining the audio and video to be identified, audio emotion recognition is performed on the audio stream in the audio and video to be identified, if the audio recognition result is that there is emotion, the corresponding video segment in the video stream is determined, and emotion recognition is performed on the video segment to obtain a video recognition result, and corresponding recognition models are set for audio and video recognition respectively, which can be more in line with the characteristics of the specific audio or video, so that the audio recognition results and video recognition results obtained in this way can also be more accurate, and the target emotion results obtained based on the audio recognition results and video recognition results are more reliable and accurate.
[0256] In an alternative embodiment, an electronic device is provided, such as Figure 6 As shown, Figure 6 The electronic device shown in FIG. 1 may be a device capable of implementing the above-mentioned audio and video emotion tagging method, and its internal structure diagram may be as shown in FIG. Figure 6 As shown. The electronic device includes a processor, a memory, an input / output interface, a communication interface, a display unit, an audio and video acquisition unit and an input device. The processor, the memory and the input / output interface are connected via a system bus, and the communication interface, the display unit, the receiving frame and the input device are connected to the system bus via the input / output interface. The processor of the electronic device is used to provide computing and control capabilities. The memory of the electronic device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The input / output interface of the electronic device is used to exchange information between the processor and an external device. The communication interface of the electronic device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be implemented through WIFI, a mobile cellular network, NFC (near field communication) or other technologies. The display unit of the electronic device is used to form a visually visible picture, which can be a display screen, a projection device or a virtual reality imaging device. The audio and video acquisition unit of the electronic device is used to obtain audio and video to complete the audio and video emotion tagging method of the present application. The display screen can be a liquid crystal display screen or an electronic ink display screen, and the input device of the electronic device can be a touch layer covering the display screen, or a button, trackball or touchpad set on the electronic device casing, or an external keyboard, touchpad, mouse, air mouse or remote control, etc.
[0257] Those skilled in the art will understand that Figure 6The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.
[0258] An embodiment of the present application provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the steps and corresponding contents of the aforementioned method embodiment can be implemented.
[0259] It should be noted that the computer-readable storage medium mentioned above in the present application may be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or component, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0260] An embodiment of the present application also provides a computer program product, including a computer program, which can implement the steps and corresponding contents of the aforementioned method embodiment when executed by a processor.
[0261] The terms "first," "second," "third," "fourth," "1," "2," and the like (if any) in the specification and claims of this application and the accompanying drawings are used to distinguish similar objects and are not necessarily used to describe a particular order or sequential sequence. It should be understood that the terms used in this manner are interchangeable where appropriate, so that the embodiments of the application described herein can be implemented in an order other than that shown or described in the drawings.
[0262] The flow charts and block diagrams in the accompanying drawings illustrate the possible architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present application. In this regard, each box in the flow chart or block diagram can represent a module, program segment or a part of code, and the module, program segment or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be implemented by a dedicated hardware-based system that performs the specified function or operation, or can be implemented by a combination of dedicated hardware and computer instructions.
[0263] It should be understood that, although each operation step is indicated by arrows in the flowchart of the embodiment of the present application, the order of implementation of these steps is not limited to the order indicated by the arrows. Unless otherwise clearly stated herein, in some implementation scenarios of the embodiment of the present application, the implementation steps in each flowchart can be performed in other orders according to demand. In addition, some or all of the steps in each flowchart can include multiple sub-steps or multiple stages based on actual implementation scenarios. Some or all of these sub-steps or stages can be executed at the same time, and each sub-step or stage in these sub-steps or stages can also be executed at different times respectively. Under different scenarios at the execution time, the execution order of these sub-steps or stages can be flexibly configured according to demand, and the embodiment of the present application does not limit this.
[0264] The above description is only an optional implementation method for some implementation scenarios of this application. It should be pointed out that for ordinary technicians in this technical field, without departing from the technical concept of the solution of this application, the use of other similar implementation methods based on the technical ideas of this application also falls within the protection scope of the embodiments of this application.
Claims
1. A multimodal emotion recognition method, characterized in that: include: Get the audio and video to be recognized; The audio and video to be identified include audio stream and video stream; Segmenting the audio stream to obtain at least one audio segment; Inputting each of the audio segments into an audio recognition model to obtain an audio recognition result; Determining a corresponding video segment in the video stream according to a target audio segment for which the audio recognition result is an emotion result; Inputting the video segment into a video recognition model to obtain a video recognition result; Based on the audio recognition result and the video recognition result, a target emotion result of the audio and video to be recognized is determined.
2. The multimodal emotion recognition method according to claim 1, characterized in that Inputting each of the audio segments into an audio recognition model to obtain an audio recognition result includes: Performing feature engineering on each of the audio segments to obtain an audio feature vector corresponding to each of the audio segments; the audio feature vector includes frequency change, amplitude change, pitch, and timbre of the audio segment; The audio feature vectors corresponding to the audio segments are input into the audio recognition model one by one to obtain the audio recognition results corresponding to the audio segments.
3. The multimodal emotion recognition method according to claim 2, characterized in that: The performing feature engineering on each of the audio segments to obtain an audio feature vector corresponding to each of the audio segments includes: framing each of the audio segments to obtain a plurality of audio frames; Performing frequency analysis on each of the audio frames to obtain a frequency feature vector; the frequency feature vector includes spectrum centroid data, spectrum dispersion data, spectrum slope data, and spectrum rolling data; Performing amplitude analysis on each of the audio frames to obtain an amplitude feature vector; the amplitude feature vector includes instantaneous amplitude, amplitude envelope, energy data, zero-crossing rate data, and loudness data; Performing pitch analysis on each of the audio frames to obtain a pitch feature vector; the pitch feature vector includes fundamental frequency data, fundamental frequency time series data, and harmonic wave data; Performing timbre analysis on each of the audio frames to obtain a timbre feature vector; the timbre feature vector includes Mel-frequency cepstral coefficient data, chroma feature data, and linear prediction coefficient data; The frequency feature vector, the amplitude feature vector, the pitch feature vector, and the timbre feature vector of each audio frame are aggregated to obtain an audio feature vector corresponding to each audio segment.
4. The multimodal emotion recognition method according to claim 1, wherein: Inputting the video segment into a video recognition model to obtain a video recognition result includes: Performing feature engineering on each of the video segments to obtain a video feature vector corresponding to each of the video segments; the video feature vector includes facial change features, action change features, and scene change features of the video segment; The video feature vectors corresponding to the video segments are input into the video recognition model one by one to obtain the video recognition results corresponding to the video segments.
5. The multimodal emotion recognition method according to claim 1, wherein: The video segment includes a plurality of image frames; After inputting the video segment into the video recognition model to obtain the video recognition result, the method further includes: Inputting each of the image frames into an image recognition model to obtain an image recognition result and a corresponding confidence level; Based on the image recognition result and the corresponding confidence level, obtaining the image recognition result and the corresponding image recognition confidence level of the video segment; Obtaining a video recognition confidence level of the video recognition result; If the image recognition confidence is greater than the video recognition confidence, the video recognition result is updated based on the image recognition result.
6. The multimodal emotion recognition method according to claim 5, characterized in that: Inputting each of the image frames into an image recognition model to obtain an image recognition result includes: Performing feature engineering on each of the image frames to obtain an image feature vector corresponding to each of the image frames; the image feature vector includes facial features, action features, and scene features of the image frame; The image feature vectors corresponding to the image frames are input into the image recognition model one by one to obtain image recognition results.
7. The multimodal emotion recognition method according to claim 1, characterized in that: The determining, based on the audio recognition result and the image recognition result, a target emotion result of the audio and video to be recognized includes: Determining a first confidence level of the audio recognition result and a second confidence level of the video recognition result; If the audio recognition result is inconsistent with the video recognition result, the first confidence level and the second confidence level are compared, and the recognition result with the higher confidence level is used as the target emotion result.
8. A multimodal emotion recognition device, characterized in that: include: Audio and video acquisition module, used to obtain the audio and video to be identified; The audio and video to be identified include audio stream and video stream; A segmentation module, configured to segment the audio stream to obtain at least one audio segment; A first recognition module, configured to input each of the audio segments into an audio recognition model to obtain an audio recognition result; A video segment determination module, configured to determine a corresponding video segment in the video stream according to a target audio segment whose audio recognition result is an emotion result; A second recognition module is used to input the video segment into a video recognition model to obtain a video recognition result; A result determination module is used to determine a target emotion result of the audio and video to be recognized based on the audio recognition result and the video recognition result.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory, characterized in that: The processor executes the computer program to implement the steps of the multimodal emotion recognition method according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the multimodal emotion recognition method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Emotion recognition method, device and equipment based on audio and video
CN115376559A
Cited By
Interaction method, device, medium, product and chip
CN121919328A
An interaction method, device, medium, product, and chip
CN121919328B