Audio and video emotion marking method and device, electronic equipment, storage medium and product
By framing and feature engineering the audio stream data and combining it with an emotion recognition model, the problem of inaccurate audio and video emotion labeling in existing technologies is solved, achieving more accurate emotion judgment and a better user experience.
Patent Information
- Application Number
- CN202510906612.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-01
- Publication Date
- 2025-09-23
AI Technical Summary
Existing audio and video emotion tagging methods rely on video image recognition, which is prone to recognition errors, resulting in inaccurate emotion tagging and affecting the user's viewing experience.
By framing the audio stream data, obtaining audio frames and performing feature engineering, the emotion recognition model is used to identify the emotion results of the audio frames, and adjacent frames are merged to generate the target audio segment and mark it in the audio and video segment.
The robustness and accuracy of emotion judgment are improved, and users can quickly locate the audio and video clips corresponding to the emotion results, improving the review experience.
Smart Images

Figure CN120690233A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of audio and video emotion tagging. Specifically, the present application relates to an audio and video emotion tagging method, device, electronic device, computer-readable storage medium and computer program product. Background Art
[0002] With the rapid development of machine learning, especially deep learning technology, it has become possible to extract and interpret emotional information from voice intonation, facial expressions, and body language. Recognizing emotions in audio and video is an increasingly important research direction in the field of artificial intelligence. Extracting features from longer audio and video, such as facial expressions, body movements, duration, and surrounding environment, and identifying character emotions, can enable users to quickly understand the emotional content corresponding to audio and video clips when watching them again, enriching the viewing experience.
[0003] Current emotion tagging methods use facial expression recognition based on images in the video to mark expressions and emotions in the video. However, determining emotions based on image recognition or expression recognition is prone to recognition errors, resulting in inaccurate emotion tagging and affecting the user's viewing experience. Summary of the Invention
[0004] The embodiments of the present application provide a method, device, electronic device, computer-readable storage medium, and computer program product for tagging audio and video emotions, which aim to solve the technical problem of inaccurate emotion recognition in audio and video, resulting in inaccurate tagging and poor playback experience.
[0005] In a first aspect, a method for labeling audio and video emotions is provided, the method comprising:
[0006] Get the audio stream data of the audio and video to be marked;
[0007] Performing frame processing on the audio stream data to obtain multiple audio frames;
[0008] Input each audio frame into the emotion recognition model one by one according to time to obtain the emotion result corresponding to each audio frame;
[0009] Merge adjacent audio frames that match the emotion results to obtain the corresponding target audio segment;
[0010] The audio and video to be marked are segmented according to the target audio segment, and the corresponding emotion results are marked in the corresponding audio and video segments.
[0011] In some possible implementations, each audio frame is input into the emotion recognition model one by one according to time, and the emotion result corresponding to each audio frame is obtained, specifically including:
[0012] Perform feature engineering on each audio frame to obtain the corresponding embedded vector of each audio frame. The embedded vector contains the frequency, amplitude and timbre of the audio frame.
[0013] The embedded vectors corresponding to each audio frame are input into the emotion recognition model one by one to obtain the emotion results corresponding to each audio frame.
[0014] In some possible implementations, feature engineering is performed on each audio frame to obtain an embedded vector corresponding to each audio frame, specifically including:
[0015] Extract amplitude features from each audio frame to obtain a corresponding multi-dimensional amplitude vector, which includes root mean square amplitude data, peak amplitude data, and zero crossing rate data;
[0016] Frequency features are extracted for each audio frame to obtain the corresponding multidimensional frequency vector, which includes Mel spectrum data, MFCC data, spectrum centroid data, and spectrum bandwidth data;
[0017] Extracting timbre features from each audio frame to obtain a corresponding multidimensional timbre vector, the multidimensional timbre vector including spectrum roll-off point data and harmonic distortion ratio data;
[0018] splicing the multi-dimensional amplitude vector, the multi-dimensional frequency vector and the multi-dimensional timbre vector to obtain a spliced vector;
[0019] According to each adjacent audio frame, the temporal features of each splicing vector are enhanced;
[0020] Perform nonlinear transformation on the concatenated vector to obtain the embedded vector corresponding to each audio frame.
[0021] In some possible implementations, the emotion recognition model includes a multi-dimensional convolutional layer, a temporal processing layer, and a fully connected layer;
[0022] The embedded vectors corresponding to each audio frame are input into the emotion recognition model one by one to obtain the emotion results corresponding to each audio frame, including:
[0023] Input the embedded vector corresponding to each audio frame into the multidimensional convolution layer to obtain the corresponding convolution vector;
[0024] Input the convolution vector into the time series processing layer to obtain the time series feature vector;
[0025] The time series feature vector is input into the fully connected layer and mapped to the corresponding emotion category space to obtain the emotion result corresponding to each audio frame.
[0026] In some possible implementations, segmenting the audio and video to be marked according to the target audio segment and marking the corresponding emotion results in the corresponding audio and video segments specifically includes:
[0027] Segment the audio and video to be marked according to the target audio segment to obtain corresponding audio and video segments, where the start and end times of the audio and video segments correspond to the target audio segment;
[0028] The corresponding emotion result is marked in the target frame of the corresponding audio and video segment.
[0029] In some possible implementations, marking the corresponding emotion result in the corresponding audio or video segment specifically includes:
[0030] According to the emotion results, determine the corresponding marking style of each audio and video segment;
[0031] Based on the corresponding marking style, each audio and video segment is displayed to the user in the audio and video to be marked.
[0032] In some possible implementations, merging adjacent audio frames matching emotion results to obtain a corresponding target audio segment specifically includes:
[0033] Traversing the emotion result list to determine the similarity between the emotion result corresponding to each audio frame and the emotion results corresponding to adjacent audio frames;
[0034] If the similarity is greater than a predetermined similarity threshold, the audio frame is merged with the adjacent audio frame to obtain a corresponding target audio segment.
[0035] In some possible implementations, the audio and video to be marked includes video stream data;
[0036] Merge adjacent audio frames that match the emotion results to obtain the corresponding target audio segment, including:
[0037] Merge adjacent audio frames that match the emotion results to obtain the corresponding initial audio segment;
[0038] Based on the initial audio segment, determining a corresponding video segment from the video stream data;
[0039] Based on the video segment, obtain the corresponding scene information;
[0040] Based on the scene information, the audio frames in the initial audio segment are adjusted to obtain the corresponding target audio segment.
[0041] In some possible implementations, the method further includes:
[0042] In response to a user's request to view the marked target audio and video, a segment mark list of the target audio and video is displayed, where the segment mark list includes emotion result marks corresponding to each audio and video segment sorted by time;
[0043] In response to the user's selection of the emotion result tag, the corresponding audio and video segment is played.
[0044] In some possible implementations, the method further includes:
[0045] Acquire multiple audio frame samples; the audio frame samples are marked with corresponding emotion result labels;
[0046] Input each audio frame sample into the initial recognition model one by one to obtain the emotion recognition result;
[0047] According to the emotion recognition results and emotion result labels, the parameters of the initial recognition model are updated until the predetermined end conditions are met, and the training is terminated to obtain a trained emotion recognition model.
[0048] In a second aspect, a device for labeling audio and video emotions is provided, the device comprising:
[0049] A data acquisition module is used to obtain the audio stream data of the audio and video to be marked;
[0050] The frame acquisition module is used to perform frame processing on the audio stream data to obtain multiple audio frames;
[0051] The recognition module is used to input each audio frame into the emotion recognition model one by one according to time to obtain the emotion result corresponding to each audio frame;
[0052] A merging module is used to merge adjacent audio frames that match the emotion results to obtain the corresponding target audio segment;
[0053] The marking module is used to segment the audio and video to be marked according to the target audio segment, and mark the corresponding emotion results in the corresponding audio and video segments.
[0054] According to a third aspect, an electronic device is provided, comprising:
[0055] A memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the steps of any method in the first aspect of the present application.
[0056] In a fourth aspect, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the audio and video emotion labeling method shown in any one of the first aspects of the present application is implemented.
[0057] In a fifth aspect, a computer program product is provided, comprising a computer program, which implements the steps of any one of the methods in the first aspect of the present application when the computer program is executed by a processor.
[0058] The beneficial effects of the technical solution provided by the embodiments of the present application are:
[0059] The audio and video emotion tagging method provided by the present application obtains the audio stream data of the audio and video to be tagged, divides the audio stream data into frames to obtain multiple audio frames, and performs emotion recognition on the audio frames, so that the algorithm focuses on local audio and effectively isolates background noise or non-emotion-related voice fluctuations. The audio frames are input into the emotion recognition model in chronological order to obtain the emotion results corresponding to each audio frame, and the adjacent audio frames that match the emotion results are merged to obtain the target audio segment. The audio segment is used to express the continuous emotional state, which can reduce the incorrect recognition caused by individual frame misjudgment or noise interference, and improve the robustness and accuracy of emotion judgment. The audio and video segment is then determined according to the target audio segment, and the corresponding emotion result is marked in the audio and video segment, so that the user can quickly locate the audio and video clip corresponding to the emotion result based on the mark, effectively improving the review experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0060] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for describing the embodiments of the present application.
[0061] Figure 1 A schematic diagram of an application scenario of an audio and video emotion tagging method provided in an embodiment of the present application;
[0062] Figure 2 A flowchart of a method for labeling audio and video emotions provided in an embodiment of the present application;
[0063] Figure 3 A flowchart of a method for labeling audio and video emotions provided in an embodiment of the present application;
[0064] Figure 4 A flowchart illustrating an example of an audio and video emotion tagging method provided in an embodiment of the present application;
[0065] Figure 5 A schematic diagram of the structure of an audio and video emotion tagging device provided in an embodiment of the present application;
[0066] Figure 6 A schematic structural diagram of an electronic device applicable to the audio and video emotion labeling method provided in an embodiment of the present application. DETAILED DESCRIPTION
[0067] The following describes the embodiments of the present application in conjunction with the accompanying drawings. It should be understood that the embodiments described below in conjunction with the accompanying drawings are exemplary descriptions for explaining the technical solutions of the embodiments of the present application and do not constitute a limitation on the technical solutions of the embodiments of the present application.
[0068] Those skilled in the art will understand that, unless otherwise stated, the singular forms "a", "an", "said", and "the" used herein may also include plural forms. It should be further understood that the terms "including" and "comprising" used in the embodiments of the present application mean that the corresponding features can be implemented as the presented features, information, data, steps, operations, elements, and / or components, but do not exclude implementation as other features, information, data, steps, operations, elements, components, and / or combinations thereof supported by the present technical field. It should be understood that when we say that an element is "connected" or "coupled" to another element, the element can be directly connected or coupled to the other element, or it can refer to the element and the other element establishing a connection relationship through an intermediate element. In addition, the "connection" or "coupling" used herein may include wireless connection or wireless coupling. The terms "or", "and / or", "including at least one of the following", etc. used in this application may be interpreted as inclusive, or mean any one or any combination.
[0069] In the embodiments of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program that has a predetermined function and works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as processing circuits or memories) or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be part of an overall module or unit that includes the function of the module or unit.
[0070] In the specific implementation of this application, any data related to an object, such as data involved in the object's use of an application, is required. When the embodiments of this application are applied to specific products or technologies, the permission or consent of the object must be obtained, and the collection, use, and processing of the relevant data must comply with the relevant laws, regulations, and standards of the relevant countries and regions. In other words, if any of the above-mentioned data related to an object is involved in the embodiments of this application, such data must be obtained with the authorization and consent of the object and in compliance with the relevant laws, regulations, and standards of the relevant countries and regions.
[0071] In order to make the objectives, technical solutions and advantages of this application clearer, the implementation methods of this application will be further described in detail below with reference to the accompanying drawings.
[0072] First, the technical terms involved in this application are introduced and explained:
[0073] Feature engineering: refers to the process of extracting, selecting, and constructing features that are useful for machine learning models from raw data. The goal is to convert raw data into a format suitable for model processing, thereby optimizing the model's training effect and predictive ability.
[0074] Embedded vector: is a technique or result that maps non-numeric data (such as words, sentences, images, users, items, etc.) into a low-dimensional, continuous numerical vector space.
[0075] Currently, people often watch complete recorded videos to find clips of interest. If the video being viewed is long, the viewing efficiency is low, and it is easy to miss short, subtle emotional expressions, and it is difficult for people to maintain concentration for a long time. In existing technologies, image recognition or facial expression recognition can be used to determine user actions and expressions, and mark them for easy review. However, image or facial recognition has difficulty dealing with occlusion, lighting, subtle expressions, etc., and lacks understanding of the context. The recognition is not accurate enough and it is easy to mark meaningless clips. Users still need to watch manually, and the review experience is poor.
[0076] The audio and video emotion tagging method, device, electronic device, computer-readable storage medium and computer program product provided in this application are intended to solve at least one of the above technical problems in the prior art.
[0077] In response to at least one of the above-mentioned technical problems or areas that need improvement in the related art, the present application proposes a method, device, electronic device, computer-readable storage medium and computer program product for audio and video emotion tagging. The audio and video emotion tagging method provided by the scheme obtains audio stream data of the audio and video to be tagged, frames the audio stream data to obtain multiple audio frames, and performs emotion recognition on the audio frames, so that the algorithm focuses on local sound events and effectively isolates background noise or non-emotion-related voice fluctuations. The audio frames are input into the emotion recognition model in chronological order to obtain the emotion results corresponding to each audio frame, and the adjacent audio frames with matching emotion results are merged to obtain the target audio segment. The audio segment is used to express the continuous emotional state, which can reduce the erroneous recognition caused by individual frame misjudgment or noise interference, and improve the robustness and accuracy of emotion judgment. Then, the audio and video segment is determined based on the target audio segment, and the corresponding emotion result is marked in the audio and video segment, so that the user can quickly locate the audio and video clip corresponding to the emotion result based on the mark, effectively improving the playback experience.
[0078] The following describes several exemplary embodiments to illustrate the technical solutions of the embodiments of the present application and the technical effects produced by the technical solutions of the present application. It should be noted that the following embodiments can refer to, draw on, or combine with each other, and the same terms, similar features, and similar implementation steps in different embodiments will not be repeated.
[0079] Figure 1A schematic diagram of an application scenario of the audio and video emotion tagging method provided in an embodiment of the present application, wherein the application environment may include a camera device 100 and a server 200, and the camera device 100 and the server 200 are connected via a network; the server 200 may be configured with an audio and video emotion tagging system.
[0080] Specifically, the audio and video to be marked are collected by the camera device 100, and the audio and video to be marked are sent to the server 200. The server 200 obtains the audio stream data of the audio and video to be marked, and performs frame processing on the audio stream data to obtain multiple audio frames. Each audio frame is input into the emotion recognition model one by one according to time to obtain the emotion result corresponding to each audio frame, and the adjacent audio frames that match the emotion result are merged to obtain the corresponding target audio segment. The audio and video to be marked are segmented according to the target audio segment, and the corresponding emotion result is marked in the corresponding audio and video segment. The above-mentioned audio and video emotion marking method can also be executed by the audio and video emotion marking system in the server 200.
[0081] In a specific implementation process, the server 200 may also be a terminal capable of processing audio and video and marking emotion results, and an audio and video emotion marking system may be configured on the terminal.
[0082] During the specific implementation process, when the camera device 100 is a monitoring video device, the monitoring video device performs the monitoring task. Once it captures emotional expressions such as laughter in the environment, its built-in algorithm system will be triggered to quickly analyze and process the audio signal, accurately judge the characteristics and intensity of the audio, and after confirming that it is a valid emotion, the system will automatically mark a special mark on the corresponding monitoring video timeline, so that when the user reviews the video, there is no need to browse the lengthy video minute by minute. With the help of these clear marks, they can quickly locate those video clips full of emotional atmosphere, which greatly improves the efficiency of users in obtaining specific and interesting content and enriches the video viewing experience.
[0083] Among them, camera equipment is a device that can capture audio and video, such as: professional cameras (including movie cameras, broadcast-level cameras), portable camcorders, smartphones (usually equipped with front and rear cameras and microphone arrays), tablets, laptops (equipped with cameras and microphones), professional recording equipment (such as voice recorders, portable audio interfaces, microphones), webcams (commonly used for video conferencing and live broadcasts), sports cameras, pan-tilt camera systems mounted on drones, virtual reality (VR) or augmented reality (AR) headsets (equipped with environment capture cameras and microphones), live streaming equipment, audio interfaces and microphones for digital audio workstations (DAWs), security surveillance cameras (some with two-way voice functions), driving recorders (some with recording functions), medical imaging equipment (such as endoscope systems, which may include audio and video capture), panoramic cameras, 3D scanners (some with audio and video recording functions), etc. These devices can be connected directly or indirectly via wired (such as USB, SDI, XLR, optical fiber, network cable) or wireless (such as Wi-Fi, Bluetooth, 5G / 4G, dedicated wireless audio and video transmission modules) communication methods to achieve synchronous or separate capture, preview, recording, transmission, or control of audio and video signals. They can work independently or as part of a larger system (such as film and television production systems, live broadcast systems, remote conferencing systems, security monitoring systems, virtual production systems), but are not limited to these systems.
[0084] The above application scenario is only an example and does not limit the application scenario of the audio and video emotion labeling method of this application.
[0085] Those skilled in the art will appreciate that a server may include a server that is equipped with a computer capable of processing database operations. The server may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It may also be a cloud server or server cluster that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. The embodiments of the present application can be applied to various scenarios, including but not limited to cloud technology, artificial intelligence, smart transportation, assisted driving, etc. The specific requirements can also be determined based on the actual application scenario and are not limited here.
[0086] The terminal can be a smart phone (such as an Android phone, an iOS phone, etc.), a tablet computer, a laptop computer, a digital broadcast receiver, a MID (Mobile Internet Devices), a PDA (Personal Digital Assistant), a desktop computer, a smart home appliance, a vehicle-mounted terminal (such as a vehicle-mounted navigation terminal, a vehicle-mounted computer, etc.), a smart speaker, a smart watch, etc. The terminal and the server can be directly or indirectly connected via wired or wireless communication, but are not limited to this.
[0087] In some possible implementations, taking the execution subject as an audio and video emotion tagging system as an example, the embodiment of the present application provides an audio and video emotion tagging method, such as Figure 2 As shown, the following steps may be included:
[0088] S210: Acquire audio stream data of the audio and video to be marked.
[0089] The audio and video to be marked include audio stream data and video stream data.
[0090] Specifically, the audio and video can be collected by a camera device, which collects audio and video data simultaneously, and aligns the audio stream data and the video stream data in time to obtain the audio and video to be marked.
[0091] During the specific implementation process, the audio and video acquisition equipment can be installed in different scenes, so the audio and video to be marked obtained are the audio and video of the corresponding scene. For example, in the home entertainment scene, the audio and video obtained is the home entertainment audio and video, in the public entertainment venue, the audio and video obtained is the entertainment venue audio and video, and in the film and television shooting scene, the audio and video obtained is the shooting audio and video. The solution of this application can obtain audio data and video data in different scenes for analysis, thereby identifying emotional results and completing the emotional marking of audio and video, and has high versatility.
[0092] S220: Perform frame processing on the audio stream data to obtain multiple audio frames.
[0093] Specifically, preset frame configuration information is obtained, and the audio stream data is framed based on the frame configuration information. The frame configuration information may include frame length and inter-frame movement distance. The audio stream data is divided based on the frame length and inter-frame movement distance to obtain multiple audio frames.
[0094] During the specific implementation process, frame configuration information is obtained, and a loop or vectorized operation is used. Starting from the starting position of the audio stream data, the system moves with a step size of the inter-frame moving distance, and each time a piece of audio data with a length of the frame length is extracted as a frame until the end of the audio stream data. Emotion recognition is performed through frames, focusing on local sounds, effectively isolating background noise or non-emotion-related voice fluctuations, and making emotion recognition more accurate.
[0095] During the specific implementation process, the collected audio signal is framed and processed to extract the characteristic parameters of each frame of audio, such as frequency, amplitude, timbre, etc. These characteristic parameters are combined with the emotion recognition model to perform emotion recognition on each audio frame. Sounds with different emotions usually have specific frequency ranges and rhythm changes. The algorithm can recognize these characteristics and distinguish them from other environmental sounds, making emotion recognition more accurate.
[0096] S230: Input each audio frame into the emotion recognition model one by one according to time to obtain the emotion result corresponding to each audio frame.
[0097] The emotion result may indicate whether the corresponding audio frame has emotion or not.
[0098] Specifically, each audio frame is input into the emotion recognition model one by one in chronological order to obtain the emotion result corresponding to each audio frame, and it is judged whether each frame is an emotional frame, and then further emotion recognition is performed on the audio frames that may meet the emotional characteristics. Identifying each audio frame can enable the algorithm to focus on local sound features, effectively isolate background noise or non-emotion-related voice fluctuations, improve the quality of audio features, and thus obtain more accurate emotion results.
[0099] During the specific implementation process, the audio features of each audio frame can be obtained, and the audio frames can be input into the emotion recognition model in turn to obtain the emotion recognition results output by the model. The emotion recognition results can indicate whether the audio frame corresponding to the audio features is an emotion frame, and can also include information such as the emotion category and emotion intensity corresponding to each emotion frame.
[0100] S240: Merge adjacent audio frames whose emotion results match to obtain a corresponding target audio segment.
[0101] Specifically, based on the emotion recognition model, the audio frames containing emotions are determined, the adjacent frames of the audio frames containing emotions are determined, the emotion results of the adjacent frames are compared with the emotion results of the audio frame, and the adjacent audio frames with matching emotion results are merged to obtain the target audio segments with the same emotion and continuous emotional state. In this way, the marked audio segments can contain complete emotion fragments, which is more in line with the user's viewing habits and improves the user experience.
[0102] During the specific implementation process, since the audio frames are input into the emotion recognition model one by one in chronological order, when the emotion result of an audio frame is that it contains emotion, the audio frame is used as the starting frame, and the emotion results of the next output audio frame are compared with the starting frame in turn. The audio frames with matching emotion results are used as frame sets until it is detected that the emotion result of an audio frame does not match. The obtained frame sets are merged in chronological order to generate the corresponding target audio segment.
[0103] S250 , segmenting the audio and video to be marked according to the target audio segment, and marking the corresponding emotion results in the corresponding audio and video segments.
[0104] Specifically, based on the time and duration corresponding to the target audio segment, the marked audio and video are segmented, and audio and video segments of corresponding duration are divided according to the time in the audio and video. In order to enable users to quickly locate the audio and video segments when viewing the audio and video, each audio and video segment is marked, which facilitates users to quickly locate the audio and video segments they want to view, improves the efficiency of video playback positioning, and enhances the user experience.
[0105] During the specific implementation process, the audio and video emotion tagging system can be set inside the camera equipment. The audio and video emotion tagging system can include a video recording system and an audio processing system. Taking the identification of laughter emotions as an example: the video recording system and the audio processing system achieve high-precision time synchronization. The video recording system records video (including audio). When the audio processing system recognizes laughter, it can quickly pass the corresponding timestamp information to the video recording system. The video recording system accurately marks the location of the laughter on the video timeline based on the received timestamp. At the same time, the marking information will be properly stored to ensure that it can be displayed stably and accurately during video playback.
[0106] In some possible implementations, in the above steps, each audio frame is input into the emotion recognition model one by one according to time to obtain the emotion result corresponding to each audio frame, specifically including:
[0107] Perform feature engineering on each audio frame to obtain the embedded vector corresponding to each audio frame;
[0108] The embedded vectors corresponding to each audio frame are input into the emotion recognition model one by one to obtain the emotion results corresponding to each audio frame.
[0109] The embedded vector contains the frequency, amplitude and timbre of the audio frame.
[0110] Specifically, feature engineering is performed on each audio frame, and the features corresponding to each audio frame are extracted to obtain an embedded vector. The embedded vector corresponding to each video frame is input into the emotion recognition model to obtain the emotion result of the corresponding video frame. Compared with directly inputting the original features, the advantage of inputting the model through the embedded vector is that the embedded vector can compress the high-dimensional and sparse original data into a low-dimensional, dense representation that contains rich semantic information. This enables the model to understand the input more accurately and robustly, improve recognition accuracy, enhance generalization ability, reduce computational complexity, and reduce manual participation costs.
[0111] In a specific implementation process, the steps of obtaining the embedded vector corresponding to the audio frame may include: obtaining multiple audio frames, using a window function to perform weighted processing on each frame of data to reduce spectral leakage and improve the accuracy of frequency domain analysis, and then performing high-pass filtering on the signal to enhance the high-frequency part and compensate for the attenuation of high-frequency energy in the speech signal, and then extracting representative and discriminative feature vectors from the preprocessed audio frame, selecting the most useful features for the task based on statistical indicators or model evaluation, reducing the dimension, using principal component analysis (PCA) or linear discriminant analysis (LDA) and other methods to reduce the feature dimension while retaining the main information, and standardizing the features to make them have the same scale to avoid certain features dominating the model training due to their large values. The data volume of the embedded vector obtained in this way is smaller than that of ordinary features, but it is more efficient in expressing data features, effectively improving the accuracy of emotion recognition.
[0112] In some possible implementations, the above steps of performing feature engineering on each audio frame to obtain an embedded vector corresponding to each audio frame specifically include:
[0113] Extract amplitude features of each audio frame to obtain the corresponding multi-dimensional amplitude vector;
[0114] Extract frequency features of each audio frame to obtain the corresponding multi-dimensional frequency vector;
[0115] Extract timbre features from each audio frame to obtain a corresponding multi-dimensional timbre vector;
[0116] splicing the multi-dimensional amplitude vector, the multi-dimensional frequency vector and the multi-dimensional timbre vector to obtain a spliced vector;
[0117] According to each adjacent audio frame, the temporal features of each splicing vector are enhanced;
[0118] Perform nonlinear transformation on the concatenated vector to obtain the embedded vector corresponding to each audio frame.
[0119] Among them, the amplitude vector includes root mean square amplitude data, peak amplitude data and zero crossing rate data, the multidimensional frequency vector includes Mel spectrum data, MFCC data, spectrum centroid data and spectrum bandwidth data, and the multidimensional timbre vector includes spectrum roll-off point data and harmonic distortion ratio data.
[0120] Specifically, amplitude features are extracted for each audio frame to obtain a corresponding multidimensional amplitude vector, frequency features are extracted for each audio frame to obtain a corresponding multidimensional frequency vector, timbre features are extracted for each audio frame to obtain a corresponding multidimensional timbre vector, the multidimensional amplitude vector, multidimensional frequency vector and multidimensional timbre vector are spliced to obtain a spliced vector, time series features of each spliced vector are enhanced according to adjacent audio frames, the spliced vector is nonlinearly transformed to obtain an embedded vector corresponding to each audio frame, these three vectors are spliced together to form a comprehensive spliced vector that integrates loudness, spectrum and timbre information. This high-dimensional vector can represent the content of the original audio frame more comprehensively and finely, providing strong data support for subsequent audio frame emotion recognition tasks.
[0121] During the specific implementation process, vectors corresponding to the three features of amplitude, frequency and timbre in the audio frame are extracted to obtain more comprehensive information. The amplitude reflects the strength of the sound, the frequency reveals the high and low pitch of the sound, and the timbre distinguishes the unique texture of different sound sources. Combining these three is like seeing the color, shape and texture of an object at the same time. It can more accurately distinguish and identify various sounds. When applied to emotion recognition, the accuracy and robustness will be greatly improved.
[0122] In practice, the root mean square amplitude (RMS) is calculated by taking the square root of the average of the squares of the audio signal's amplitude over a period of time. This can be understood as the average energy or loudness of the sound. A larger value indicates a louder sound, which is important for distinguishing between loud and soft sounds. Peak amplitude refers to the maximum value reached by the audio signal during a specific time period. This may be the absolute value of the highest or lowest point. It reflects the maximum intensity or impact of the sound. For example, a sudden explosion will have a high peak amplitude, while a steady voice will have a relatively low peak amplitude. The zero-crossing rate (ZCR) refers to the number of times the audio signal crosses zero per unit time, i.e., the number of times it changes from positive to negative or vice versa. A high ZCR indicates a sharper or hissing sound, such as a grating sound or high-frequency noise, while a low ZCR may indicate a deeper or smoother sound. This helps distinguish between speech and music, or different types of timbre. The Mel spectrum simulates the frequency scale of auditory perception. It converts the audio signal to the Mel scale and displays the energy distribution of each frequency at different time points in Mels. This effectively reflects the pitch and timbre perceived by the human ear. MFCC Mel-Frequency Cepstral Coefficients (MFCCs) are a set of coefficients extracted from the Mel-Frequency Cepstral Coefficients (MFCCs). They simulate how the human auditory system processes sound, emphasizing the spectral shape of the sound rather than specific frequency peaks. This makes them very effective for distinguishing different speakers, different speech content, or different musical instruments. The spectral centroid can be understood as the center of gravity or average frequency of the audio signal's spectrum. It is calculated by multiplying the energy of all frequency components by their respective energies and dividing the result by the total energy. A higher spectral centroid indicates a higher concentration of high-frequency components, potentially making the sound brighter or sharper. A lower spectral centroid indicates a higher concentration of low-frequency components, potentially making the sound darker or more muffled. Spectral bandwidth measures the width or dispersion of spectral energy around the spectral centroid. A wider bandwidth means the energy is spread across a wider frequency range, potentially making the sound more complex or fuzzy. A narrower bandwidth means the energy is more concentrated, potentially making the sound cleaner or sharper. The spectral roll-off point is a frequency value below which frequency components contain a certain percentage of the total energy. It measures where the spectral energy is primarily concentrated. A lower roll-off point indicates a higher proportion of low-frequency energy; a higher roll-off point indicates a higher proportion of high-frequency energy. Harmonic Distortion Ratio: There are some harmonic and non-harmonic components (noise) in the sound. It is used to measure the ratio of harmonic energy to noise energy. A high ratio value indicates a more regular sound, while a low ratio value indicates a high noise component and the sound may sound rougher or distorted.
[0123] In some possible implementations, such as Figure 3 As shown, in the above steps, the embedded vectors corresponding to each audio frame are input into the emotion recognition model one by one to obtain the emotion results corresponding to each audio frame, which specifically include:
[0124] S310, inputting the embedded vector corresponding to each audio frame into the multi-dimensional convolution layer to obtain the corresponding convolution vector;
[0125] S320, input the convolution vector into the time series processing layer to obtain a time series feature vector;
[0126] S330: Input the time series feature vector into the fully connected layer and map it to the corresponding emotion category space to obtain the emotion result corresponding to each audio frame.
[0127] Among them, the emotion recognition model includes multi-dimensional convolutional layer, time series processing layer and fully connected layer.
[0128] Specifically, the embedded vector corresponding to each audio frame is input into the multidimensional convolution layer. The convolution layer finds patterns on the embedded vector and extracts more complex features to obtain the corresponding convolution vector. The convolution vector is input into the time series processing layer, and the spatial local features captured by the convolution layer are used as input. Then, the time series processing layer is used to analyze and understand the changes and relationships of these features in the time dimension, thereby achieving a deeper understanding of the spatiotemporal data and obtaining a time series feature vector. The time series feature vector is input into the fully connected layer. The fully connected layer maps the time series feature vector into another vector. The dimension and meaning of this new vector correspond to the emotion category space. It is judged whether the mapped vector is in the emotion category space and which emotion category space it is in, thereby obtaining the emotion result of the corresponding audio frame. For example, if we want to distinguish 10 different emotion categories, the mapped vector may be 10-dimensional, and each dimension represents a score or probability of the category, thereby realizing the classification of the original time series data.
[0129] In some possible implementations, the above steps of segmenting the audio and video to be marked according to the target audio segment and marking the corresponding emotion results in the corresponding audio and video segments specifically include:
[0130] Segment the marked audio and video according to the target audio segment to obtain the corresponding audio and video segments;
[0131] The corresponding emotion result is marked in the target frame of the corresponding audio and video segment.
[0132] The start and end time of the audio and video segment corresponds to the target audio segment; the target frame may include a start frame, an end frame or a specified frame.
[0133] Specifically, according to the start and end time of the target audio segment, the audio and video segments corresponding to the audio and video to be marked are determined and segmented. Based on the pre-set target frame, the corresponding emotional result is marked in the target frame of the corresponding audio and video segment. For example, if the target frame is the starting frame, the emotional result is marked on the starting frame of the audio and video segment. If the target frame is the middle frame, the emotional result is marked on the middle frame of the audio and video segment. If the target frame is the starting frame and the ending frame, the emotional result can be marked on the starting frame and the ending frame.
[0134] During the specific implementation process, the audio and video may include a playback timeline, segmenting the audio and video to be marked according to the target audio segment to obtain the corresponding audio and video segments, marking the audio and video segments on the timeline segments, determining the position of the target frame of the audio and video segment on the timeline, and marking the corresponding emotional results at the corresponding position on the timeline. Users can click on the timeline to jump directly to the corresponding audio and video segment for playback.
[0135] In some possible implementations, the above steps of marking the corresponding emotion results in the corresponding audio and video segments specifically include:
[0136] According to the emotion results, determine the corresponding marking style of each audio and video segment;
[0137] Based on the corresponding marking style, each audio and video segment is displayed to the user in the audio and video to be marked.
[0138] Among them, the marking styles can include different colors, different fonts, special markings, etc.
[0139] Specifically, according to the emotional results, the marking style corresponding to each audio and video segment is determined, and based on the corresponding marking style, the audio and video segment is marked in the audio and video, and the emotional result is marked at the specified position of the audio and video segment. Then, each audio and video segment and the corresponding emotional result mark can be displayed to the user in the audio and video to be marked. The marking style of the marked audio and video segment and the marking style of the emotional result can be different. For example, the audio and video segment is set to be highlighted, and the emotional result marked on the audio and video segment is set to red.
[0140] During the specific implementation process, the emotional results can include emotional categories. Different audio and video segments correspond to different emotional categories, and can be distinguished and marked in different styles. For example, the audio and video segments with happy emotions are marked with yellow highlights, and the corresponding emotional results are marked with yellow words. The audio and video segments with sad emotions are marked with blue, and the corresponding emotional results are marked with black words. It can be set according to user preferences, so that users can intuitively see different types of audio and video segments in the audio and video, so as to quickly find the segments they want to view.
[0141] During the specific implementation process, the audio and video may include a playback timeline, segment the audio and video to be marked according to the target audio segment, obtain the corresponding audio and video segments, determine the marking style corresponding to the audio and video segment, mark the audio and video segment with the corresponding marking style on the timeline segment, and then determine the position of the target frame of the audio and video segment on the timeline, mark the corresponding emotional result with the corresponding style at the corresponding position on the timeline, and the user can click on the timeline to jump directly to the corresponding audio and video segment for playback.
[0142] In some possible implementations, the steps of merging adjacent audio frames matching the emotion results to obtain the corresponding target audio segment specifically include:
[0143] Traversing the emotion result list to determine the similarity between the emotion result corresponding to each audio frame and the emotion results corresponding to adjacent audio frames;
[0144] If the similarity is greater than a predetermined similarity threshold, the audio frame is merged with the adjacent audio frame to obtain a corresponding target audio segment.
[0145] The emotion result list is a list of emotion results marked in the audio and video to be marked sorted by time.
[0146] Specifically, the emotion result list is traversed, and the audio frame is compared with the adjacent audio frames to determine the similarity between the emotion result corresponding to each audio frame and the emotion result corresponding to the adjacent audio frame. If the similarity is greater than the predetermined similarity threshold, the audio frame is merged with the adjacent audio frame to obtain the corresponding target audio segment. The continuous audio frames can be connected to obtain a continuous emotional state. The marked fragments are more in line with the user's viewing habits, thereby improving the user experience.
[0147] During the specific implementation process, if the emotional result of the first audio frame is a happy emotion, the emotional result of the adjacent audio frame is determined. If the emotional result of the adjacent audio frame is also a happy emotion, and the similarity is greater than the similarity threshold, the first audio frame can be merged with the adjacent audio frame. If the emotional result of the adjacent audio frame is a happy emotion, but the similarity is less than the similarity threshold, or the emotional result of the adjacent audio frame is no emotion or other emotions, the adjacent audio frame is discarded; if the similarity of the adjacent audio frames does not meet the similarity threshold, the first audio frame can be marked, or the first audio frame can be discarded. The specific settings can be made according to actual needs to avoid marking errors affecting the user viewing experience.
[0148] In some possible implementations, merging adjacent audio frames matching the emotion results in the above step to obtain a corresponding target audio segment includes:
[0149] Merge adjacent audio frames that match the emotion results to obtain the corresponding initial audio segment;
[0150] Based on the initial audio segment, determining a corresponding video segment from the video stream data;
[0151] Based on the video segment, obtain the corresponding scene information;
[0152] Based on the scene information, the audio frames in the initial audio segment are adjusted to obtain the corresponding target audio segment.
[0153] The audio and video to be marked include video stream data.
[0154] Specifically, adjacent audio frames that match the emotion results are merged to obtain a corresponding initial audio segment. Based on the initial audio segment, a corresponding video segment is determined from the video stream data. Based on the video segment, corresponding scene information is determined. Based on the scene information, the number or other attributes of the audio frames in the initial audio segment are adjusted to obtain a corresponding target audio segment.
[0155] During implementation, the lengths of audio and video segments corresponding to the same emotion category vary in different scenarios. For example, these differences may manifest themselves in the following aspects: In relaxed and enjoyable social gatherings, such as watching TV and chatting, laughter is usually triggered by humorous jokes, interesting experiences, or common topics. This type of laughter is often more natural, loud, and lasts longer. In awkward, tense situations, or when trying to lighten the atmosphere, laughter may be shorter, suppressed, and hesitant. Therefore, it makes sense to adjust the audio frames based on the scenario, which can more accurately understand and divide emotional segments.
[0156] During implementation, the context of the initial audio stream, such as what was just said, background sounds, and even other modal information such as facial expressions in the video, can be combined to infer the most likely scenarios in which laughter occurs, such as jokes between friends, forced laughter after hearing sad news, and nervousness during a public speech.
[0157] In some possible implementations, the above method further includes:
[0158] In response to a user's request to view the marked target audio or video, display a segment mark list of the target audio or video;
[0159] In response to the user's selection of the emotion result tag, the corresponding audio and video segment is played.
[0160] The segment mark list includes emotion result marks corresponding to each audio and video segment sorted by time; the viewing request may include clicking a view button, searching for keywords, or filtering emotion fields.
[0161] Specifically, the system receives a user's request to view the target audio and video, and displays a segment mark list of the marked target audio and video. The user can click on the mark entry in the segment mark list to jump to the corresponding audio and video segment for playback, which can quickly locate the segment of interest and meet user needs.
[0162] In the specific implementation process, after obtaining the marked target audio and video, the audio and video segments corresponding to the required emotions can be extracted according to the marks in the target audio and video according to actual needs, and used for video editing or repeated playback.
[0163] In the specific implementation process, emotional result tags can have different functions in different scenarios. For example, in family entertainment scenarios, in family gatherings, parent-child activities and other occasions, the camera equipment installed in the home can continuously record. When you need to review the video, you can quickly find the wonderful moments when everyone laughed by identifying happy emotions, such as children's funny performances, and the laughter moments caused by humorous stories told by elders. After confirming these audio and video clips, you can also play them repeatedly to add more happy memories to the family, and it is also convenient to edit these interesting contents into wonderful family short videos and share them on social platforms; for example, in public entertainment scenarios, you can install camera equipment with happy or sad emotion recognition function, which can not only record the excitement of live performances and interactive sessions It can not only record colorful moments but also warm moments. When users need to edit promotional videos, they can quickly filter out clips that best show the happy or sad atmosphere of the venue based on the tags, and use them to make attractive videos to attract more potential customers. Customers can also purchase video clips of their own laughing moments in the venue as a unique consumption souvenir. For example, in auxiliary scenes of film and television production, at the film and television shooting site, the director or crew members can use this function to quickly locate the clips that cause natural laughter of the crew members during the actor's performance, and assist in judging the comedy effect of the performance. During post-editing, the editor can use the tags of happy emotions to efficiently filter out materials with comedy potential, facilitate the creation of comedy film and television works, and improve creative efficiency and work quality.
[0164] In some possible implementations, the above method further includes:
[0165] Get multiple audio frame samples;
[0166] Input each audio frame sample into the initial recognition model one by one to obtain the emotion recognition result;
[0167] According to the emotion recognition results and emotion result labels, the parameters of the initial recognition model are updated until the predetermined end conditions are met, and the training is terminated to obtain a trained emotion recognition model.
[0168] Among them, the audio frame samples are marked with corresponding emotion result labels.
[0169] Specifically, multiple audio frame samples are obtained, and each audio frame sample is input into the initial recognition model one by one to obtain the emotion recognition result. According to the emotion recognition result and the emotion result label, the parameters of the initial recognition model are updated until the predetermined end condition is reached, the training is ended, and a trained emotion recognition model is obtained.
[0170] In the specific implementation process, the predetermined end condition may include the convergence of the loss function or the accuracy being greater than the accuracy threshold. Specifically, the corresponding loss function is determined based on the emotion recognition result and the emotion result label, and the parameters of the initial recognition model are updated until the loss function converges, and the training is terminated to obtain a trained emotion recognition model. Alternatively, the emotion recognition result and the emotion result label are compared to determine the accuracy of the emotion recognition result, and the parameters of the initial recognition model are updated until the accuracy of the emotion recognition result is greater than the accuracy threshold, and the training is terminated to obtain a trained emotion recognition model.
[0171] In the specific implementation process, the specific steps for training the emotion recognition model may include: data preparation and preprocessing steps: collecting audio files containing different emotional expressions (such as happiness, sadness, anger, surprise, neutrality, etc.), which can be public datasets or self-recorded speech, etc., assigning accurate emotion labels to each audio file or its fragment, which can be manually annotated or annotated using a large model to ensure the accuracy of the labels. Then, the continuous audio stream is divided into audio frames of fixed length. These audio frames can overlap with each other to retain more contextual information. Each frame can be represented as a time domain signal or a more commonly used frequency domain feature (such as a Mel-spectrogram). Then, some transformations can be performed on the audio frames, such as adding noise, changing volume, changing speed and pitch, to generate more training samples. Then, the processed audio frames and their corresponding labels are divided into training, validation and test sets, for example, 70% training set, 15% validation set and 15% test set. The training set is used for model learning, the validation set is used to adjust hyperparameters and monitor overfitting, and the test set is used to finally evaluate the model performance. Model initialization step: Select or design an initial emotion recognition model architecture, which can be a model based on a convolutional neural network (CNN), a recurrent neural network (RNN), or a Transformer. Use random initialization or initialization based on a pre-trained model to initialize the corresponding model parameters, such as weights and bias values. Model training steps: read audio frame samples from the training set in batches, each batch contains multiple audio frames and their corresponding labels, input the audio frames of the current batch into the initial recognition model, and after processing these frames, the model outputs the emotion category probability distribution corresponding to each frame, for example, the probability of happiness is 0.7, sadness is 0.2, and anger is 0.1. Then, compare the emotion probability distribution output by the model with the true emotion label corresponding to the batch, and calculate the loss value. The loss function can be a cross-entropy loss function, which is used to measure the difference between the predicted distribution and the true distribution. Then, based on the calculated loss value, the gradient of the model parameters to the loss is calculated through the back propagation algorithm, and the optimization algorithm is used to update the model parameters to reduce the loss value based on the calculated gradient and the preset learning rate. After each training cycle, that is, after the entire training set is processed, the validation set is used to evaluate the performance of the model, for example, by evaluating the performance through accuracy, score, etc., and to determine whether the model has converged or whether overfitting occurs. Termination condition judgment step: Repeat the above training cycle continuously to determine whether the predetermined termination conditions are met. The predetermined termination conditions may include: reaching a preset number of training rounds, the loss or performance indicators on the validation set have not significantly improved after multiple consecutive trainings, or reaching a preset performance standard, etc. When the termination conditions are met, stop the training process, save the current best performance model parameters (weights and biases) to a file, and obtain a trained emotion recognition model.
[0172] For example, the emotion label of the original audio file can be in text form, such as "happy", "angry" and "neutral", or a digital code, such as 0 (neutral), 1 (happy), 2 (sad) and 3 (angry). The format of the preprocessed audio frame can be a time domain waveform fragment (a one-dimensional array) or a two-dimensional Mel spectrogram (matrix), which represents the energy distribution of different frequencies in the frame. For example, a 128x80 matrix is obtained, where 128 represents the number of Mel filter banks (frequency dimension) and 80 represents the number of sampling points in the time frame or the time domain window of the short-time Fourier transform. The data type of the audio frame can be a floating point number (float32) ranging from 0 to 1 or -1 to 1. A batch used in training may contain 32 audio frame samples, each sample is a Mel spectrogram matrix, and the corresponding label is a list or tensor containing 32 emotion labels. The model parameters are the weight matrix and bias vector that are continuously updated during the training process, and are finally saved as a file in a specific format.
[0173] In the above embodiment, by obtaining the audio stream data of the audio and video to be marked, the audio stream data is framed to obtain multiple audio frames, and emotion recognition is performed on the audio frames, so that the algorithm focuses on local sound events, effectively isolating background noise or non-emotion-related voice fluctuations, and inputting the audio frames into the emotion recognition model in chronological order to obtain the emotion results corresponding to each audio frame, and merging the adjacent audio frames that match the emotion results to obtain the target audio segment. Using the audio segment to express the continuous emotional state can reduce the erroneous recognition caused by individual frame misjudgment or noise interference, and improve the robustness and accuracy of emotion judgment. Then, the audio and video segment is determined according to the target audio segment, and the corresponding emotion result is marked in the audio and video segment, so that the user can quickly locate the audio and video clip corresponding to the emotion result based on the mark, effectively improving the review experience.
[0174] Furthermore, the emotion results obtained by the emotion recognition model can not only indicate whether each audio frame contains emotion, but also include the emotion category corresponding to the audio frame containing emotion. The adjacent audio frames that match the emotion results are merged to obtain the corresponding audio segment, and the audio and video segment corresponding to the audio segment in the audio and video is determined. When marking the emotion results in the audio and video segment, the corresponding marking style can also be determined based on different emotion categories, and the corresponding emotion category can be marked in the corresponding audio and video segment using different marking styles, so that users can quickly locate the relevant audio and video clips of the emotion category they want to view, effectively improving the efficiency of audio and video playback positioning and improving user experience.
[0175] In addition, the audio frames containing emotions can be compared with adjacent audio frames to determine the degree of similarity with each adjacent audio frame, and consecutive frames with a similarity greater than a similarity threshold can be merged to obtain audio segments, so as to obtain a more continuous emotional state, determine the video segment corresponding to the audio segment time in the audio and video, and adjust the audio frames in the audio segment based on the scene shown in the video segment to obtain the target audio segment. The duration of emotions corresponding to different scenes may be different. By flexibly adjusting the audio segments through scene information, the audio segments can be accurately divided, making the marking of emotional results more accurate and effectively improving the user's review efficiency.
[0176] In one example, the audio and video emotion tagging method of the present application is as follows: Figure 4 As shown, this may include:
[0177] S410, obtaining audio stream data of the audio and video to be marked;
[0178] S420, performing frame processing on the audio stream data to obtain multiple audio frames;
[0179] S430, inputting each audio frame into the emotion recognition model one by one according to time, and obtaining the emotion result corresponding to each audio frame;
[0180] S440, merging adjacent audio frames matching the emotion results to obtain a corresponding target audio segment;
[0181] S450, segmenting the audio and video to be marked according to the target audio segment to obtain corresponding audio and video segments;
[0182] The start and end times of the audio and video segments correspond to the target audio segment.
[0183] S460, marking the corresponding emotion result in the target frame of the corresponding audio or video segment;
[0184] S470 , in response to a user's request to view the marked target audio or video, display a segment mark list of the target audio or video;
[0185] The segment mark list includes emotion result marks corresponding to each audio and video segment sorted by time.
[0186] S480: In response to the user's selection of the emotion result tag, play the corresponding audio and video segment.
[0187] The above-mentioned audio and video emotion tagging method obtains the audio stream data of the audio and video to be tagged, divides the audio stream data into frames to obtain multiple audio frames, and performs emotion recognition on the audio frames, so that the algorithm focuses on local sound events and effectively isolates background noise or non-emotion-related voice fluctuations. The audio frames are input into the emotion recognition model in chronological order to obtain the emotion results corresponding to each audio frame, and the adjacent audio frames that match the emotion results are merged to obtain the target audio segment. The audio segment is used to express the continuous emotional state, which can reduce the incorrect recognition caused by individual frame misjudgment or noise interference, and improve the robustness and accuracy of emotion judgment. The audio and video segment is then determined according to the target audio segment, and the corresponding emotion result is marked in the audio and video segment, so that the user can quickly locate the audio and video clip corresponding to the emotion result based on the mark, effectively improving the playback experience.
[0188] Furthermore, the emotion results obtained by the emotion recognition model can not only indicate whether each audio frame contains emotion, but also include the emotion category corresponding to the audio frame containing emotion. The adjacent audio frames that match the emotion results are merged to obtain the corresponding audio segment, and the audio and video segment corresponding to the audio segment in the audio and video is determined. When marking the emotion results in the audio and video segment, the corresponding marking style can also be determined based on different emotion categories, and the corresponding emotion category can be marked in the corresponding audio and video segment using different marking styles, so that users can quickly locate the relevant audio and video clips of the emotion category they want to view, effectively improving the efficiency of audio and video playback positioning and improving user experience.
[0189] In addition, the audio frames containing emotions can be compared with adjacent audio frames to determine the degree of similarity with each adjacent audio frame, and consecutive frames with a similarity greater than a similarity threshold can be merged to obtain audio segments, so as to obtain a more continuous emotional state, determine the video segment corresponding to the audio segment time in the audio and video, and adjust the audio frames in the audio segment based on the scene shown in the video segment to obtain the target audio segment. The duration of emotions corresponding to different scenes may be different. By flexibly adjusting the audio segments through scene information, the audio segments can be accurately divided, making the marking of emotional results more accurate and effectively improving the user's review efficiency.
[0190] The embodiment of the present application provides an audio and video emotion marking device, such as Figure 5 As shown, the audio and video emotion tagging device 50 may include: a data acquisition module 510, a frame acquisition module 520, a recognition module 530, a merging module 540 and a tagging module 550, wherein:
[0191] The data acquisition module 510 is used to acquire the audio stream data of the audio and video to be marked;
[0192] The frame acquisition module 520 is used to perform frame processing on the audio stream data to obtain multiple audio frames;
[0193] The recognition module 530 is used to input each audio frame into the emotion recognition model one by one according to time to obtain the emotion result corresponding to each audio frame;
[0194] A merging module 540 is configured to merge adjacent audio frames that match the emotion results to obtain a corresponding target audio segment;
[0195] The marking module 550 is used to segment the audio and video to be marked according to the target audio segment, and mark the corresponding emotion results in the corresponding audio and video segments.
[0196] As an optional embodiment, in the device, the identification module 530 is specifically configured to:
[0197] Perform feature engineering on each audio frame to obtain the corresponding embedded vector of each audio frame. The embedded vector contains the frequency, amplitude and timbre of the audio frame.
[0198] The embedded vectors corresponding to each audio frame are input into the emotion recognition model one by one to obtain the emotion results corresponding to each audio frame.
[0199] As an optional embodiment, in the device, the identification module 530 is specifically configured to:
[0200] Extract amplitude features from each audio frame to obtain a corresponding multi-dimensional amplitude vector, which includes root mean square amplitude data, peak amplitude data, and zero crossing rate data;
[0201] Frequency features are extracted for each audio frame to obtain the corresponding multidimensional frequency vector, which includes Mel spectrum data, MFCC data, spectrum centroid data, and spectrum bandwidth data;
[0202] Extracting timbre features from each audio frame to obtain a corresponding multidimensional timbre vector, the multidimensional timbre vector including spectrum roll-off point data and harmonic distortion ratio data;
[0203] splicing the multi-dimensional amplitude vector, the multi-dimensional frequency vector and the multi-dimensional timbre vector to obtain a spliced vector;
[0204] According to each adjacent audio frame, the temporal features of each splicing vector are enhanced;
[0205] Perform nonlinear transformation on the concatenated vector to obtain the embedded vector corresponding to each audio frame.
[0206] As an optional embodiment, in the device, the identification module 530 is specifically configured to:
[0207] The emotion recognition model includes a multi-dimensional convolutional layer, a temporal processing layer, and a fully connected layer;
[0208] Input the embedded vector corresponding to each audio frame into the multidimensional convolution layer to obtain the corresponding convolution vector; input the convolution vector into the time series processing layer to obtain the time series feature vector;
[0209] The time series feature vector is input into the fully connected layer and mapped to the corresponding emotion category space to obtain the emotion result corresponding to each audio frame.
[0210] As an optional embodiment, in the device, the marking module 550 is specifically configured to:
[0211] Segment the audio and video to be marked according to the target audio segment to obtain corresponding audio and video segments, where the start and end times of the audio and video segments correspond to the target audio segment;
[0212] The corresponding emotion result is marked in the target frame of the corresponding audio and video segment.
[0213] As an optional embodiment, in the device, the marking module 550 is specifically configured to:
[0214] According to the emotion results, determine the corresponding marking style of each audio and video segment;
[0215] Based on the corresponding marking style, each audio and video segment is displayed to the user in the audio and video to be marked.
[0216] As an optional embodiment, in the device, the merging module 540 is specifically configured to:
[0217] Traversing the emotion result list to determine the similarity between the emotion result corresponding to each audio frame and the emotion results corresponding to adjacent audio frames;
[0218] If the similarity is greater than a predetermined similarity threshold, the audio frame is merged with the adjacent audio frame to obtain a corresponding target audio segment.
[0219] As an optional embodiment, in the device, the merging module 540 is specifically configured to:
[0220] The audio and video to be marked include video stream data;
[0221] Merge adjacent audio frames that match the emotion results to obtain the corresponding initial audio segment;
[0222] Based on the initial audio segment, determining a corresponding video segment from the video stream data;
[0223] Based on the video segment, obtain the corresponding scene information;
[0224] Based on the scene information, the audio frames in the initial audio segment are adjusted to obtain the corresponding target audio segment.
[0225] As an optional embodiment, the device further includes a viewing module, specifically configured to:
[0226] In response to a user's request to view the marked target audio and video, a segment mark list of the target audio and video is displayed, where the segment mark list includes emotion result marks corresponding to each audio and video segment sorted by time;
[0227] In response to the user's selection of the emotion result tag, the corresponding audio and video segment is played.
[0228] As an optional embodiment, the device further includes a model training module, which is specifically used to:
[0229] Acquire multiple audio frame samples; the audio frame samples are marked with corresponding emotion result labels;
[0230] Input each audio frame sample into the initial recognition model one by one to obtain the emotion recognition result;
[0231] According to the emotion recognition results and emotion result labels, the parameters of the initial recognition model are updated until the predetermined end conditions are met, and the training is terminated to obtain a trained emotion recognition model.
[0232] The audio and video emotion tagging device provided by the present application obtains the audio stream data of the audio and video to be tagged, frames the audio stream data to obtain multiple audio frames, and performs emotion recognition on the audio frames, so that the algorithm focuses on local sound events and effectively isolates background noise or non-emotion-related voice fluctuations. The audio frames are input into the emotion recognition model in chronological order to obtain the emotion results corresponding to each audio frame, and the adjacent audio frames that match the emotion results are merged to obtain the target audio segment. The audio segment is used to express the continuous emotional state, which can reduce the incorrect recognition caused by individual frame misjudgment or noise interference, and improve the robustness and accuracy of emotion judgment. The audio and video segment is then determined according to the target audio segment, and the corresponding emotion result is marked in the audio and video segment, so that the user can quickly locate the audio and video clip corresponding to the emotion result based on the mark, effectively improving the review experience.
[0233] Furthermore, the emotion results obtained by the emotion recognition model can not only indicate whether each audio frame contains emotion, but also include the emotion category corresponding to the audio frame containing emotion. The adjacent audio frames that match the emotion results are merged to obtain the corresponding audio segment, and the audio and video segment corresponding to the audio segment in the audio and video is determined. When marking the emotion results in the audio and video segment, the corresponding marking style can also be determined based on different emotion categories, and the corresponding emotion category can be marked in the corresponding audio and video segment using different marking styles, so that users can quickly locate the relevant audio and video clips of the emotion category they want to view, effectively improving the efficiency of audio and video playback positioning and improving user experience.
[0234] In addition, the audio frames containing emotions can be compared with adjacent audio frames to determine the degree of similarity with each adjacent audio frame, and consecutive frames with a similarity greater than a similarity threshold can be merged to obtain audio segments, so as to obtain a more continuous emotional state, determine the video segment corresponding to the audio segment time in the audio and video, and adjust the audio frames in the audio segment based on the scene shown in the video segment to obtain the target audio segment. The duration of emotions corresponding to different scenes may be different. By flexibly adjusting the audio segments through scene information, the audio segments can be accurately divided, making the marking of emotional results more accurate and effectively improving the user's review efficiency.
[0235] The devices of the embodiments of the present application can execute the methods provided in the embodiments of the present application, and their implementation principles are similar and have corresponding technical effects. The actions performed by each module in the devices of the embodiments of the present application correspond to the steps in the methods of the embodiments of the present application. For detailed functional descriptions of each module of the device, please refer to the descriptions of the corresponding methods shown above, and will not be repeated here.
[0236] In an embodiment of the present application, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory, and the processor executes the above-mentioned computer program to implement the steps of the method provided in any optional embodiment of the present application. Compared with the prior art, it can be achieved: by obtaining audio stream data of the audio and video to be marked, framing the audio stream data to obtain multiple audio frames, performing emotion recognition on the audio frames, so that the algorithm focuses on local sound events, effectively isolating background noise or non-emotion-related voice fluctuations, inputting the audio frames into the emotion recognition model in chronological order, obtaining the emotion results corresponding to each audio frame, merging adjacent audio frames that match the emotion results, obtaining a target audio segment, and using the audio segment to express a continuous emotional state, which can reduce the erroneous recognition caused by individual frame misjudgment or noise interference, and improve the robustness and accuracy of emotion judgment, and then determining the audio and video segment according to the target audio segment, and marking the corresponding emotion result in the audio and video segment, so that the user can quickly locate the audio and video clip corresponding to the emotion result based on the mark, effectively improving the review experience.
[0237] In an alternative embodiment, an electronic device is provided, such as Figure 6 As shown, Figure 6 The electronic device shown in FIG. 1 may be a device capable of implementing the above-mentioned audio and video emotion tagging method, and its internal structure diagram may be as shown in FIG. Figure 6As shown. The electronic device includes a processor, a memory, an input / output interface, a communication interface, a display unit, an audio and video acquisition unit and an input device. The processor, the memory and the input / output interface are connected via a system bus, and the communication interface, the display unit, the receiving frame and the input device are connected to the system bus via the input / output interface. The processor of the electronic device is used to provide computing and control capabilities. The memory of the electronic device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The input / output interface of the electronic device is used to exchange information between the processor and an external device. The communication interface of the electronic device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be implemented through WIFI, a mobile cellular network, NFC (near field communication) or other technologies. The display unit of the electronic device is used to form a visually visible picture, which can be a display screen, a projection device or a virtual reality imaging device. The audio and video acquisition unit of the electronic device is used to obtain audio and video to complete the audio and video emotion tagging method of the present application. The display screen can be a liquid crystal display screen or an electronic ink display screen, and the input device of the electronic device can be a touch layer covering the display screen, or a button, trackball or touchpad set on the electronic device casing, or an external keyboard, touchpad, mouse, air mouse or remote control, etc.
[0238] Those skilled in the art will understand that Figure 6 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.
[0239] An embodiment of the present application provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the steps and corresponding contents of the aforementioned method embodiment can be implemented.
[0240] It should be noted that the computer-readable storage medium mentioned above in the present application may be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or component, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0241] An embodiment of the present application also provides a computer program product, including a computer program, which can implement the steps and corresponding contents of the aforementioned method embodiment when executed by a processor.
[0242] The terms "first," "second," "third," "fourth," "1," "2," and the like (if any) in the specification and claims of this application and the accompanying drawings are used to distinguish similar objects and are not necessarily used to describe a particular order or sequential sequence. It should be understood that the terms used in this manner are interchangeable where appropriate, so that the embodiments of the application described herein can be implemented in an order other than that shown or described in the drawings.
[0243] The flow charts and block diagrams in the accompanying drawings illustrate the possible architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present application. In this regard, each box in the flow chart or block diagram can represent a module, program segment or a part of code, and the module, program segment or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be implemented by a dedicated hardware-based system that performs the specified function or operation, or can be implemented by a combination of dedicated hardware and computer instructions.
[0244] It should be understood that, although each operation step is indicated by arrows in the flowchart of the embodiment of the present application, the order of implementation of these steps is not limited to the order indicated by the arrows. Unless otherwise clearly stated herein, in some implementation scenarios of the embodiment of the present application, the implementation steps in each flowchart can be performed in other orders according to demand. In addition, some or all of the steps in each flowchart can include multiple sub-steps or multiple stages based on actual implementation scenarios. Some or all of these sub-steps or stages can be executed at the same time, and each sub-step or stage in these sub-steps or stages can also be executed at different times respectively. Under different scenarios at the execution time, the execution order of these sub-steps or stages can be flexibly configured according to demand, and the embodiment of the present application does not limit this.
[0245] The above description is only an optional implementation method for some implementation scenarios of this application. It should be pointed out that for ordinary technicians in this technical field, without departing from the technical concept of the solution of this application, the use of other similar implementation methods based on the technical ideas of this application also falls within the protection scope of the embodiments of this application.
Claims
1. A method for labeling audio and video emotions, characterized in that: include: Get the audio stream data of the audio and video to be marked; Performing frame processing on the audio stream data to obtain multiple audio frames; Inputting each of the audio frames into the emotion recognition model one by one according to time to obtain the emotion result corresponding to each of the audio frames; Merge adjacent audio frames that match the emotion results to obtain the corresponding target audio segment; The audio and video to be marked are segmented according to the target audio segment, and the corresponding emotion results are marked in the corresponding audio and video segments.
2. The audio and video emotion labeling method according to claim 1, wherein: Inputting each of the audio frames into the emotion recognition model one by one according to time to obtain the emotion result corresponding to each of the audio frames specifically includes: Performing feature engineering on each of the audio frames to obtain an embedded vector corresponding to each of the audio frames, wherein the embedded vector includes the frequency, amplitude, and timbre of the audio frame; The embedded vectors corresponding to the audio frames are input into the emotion recognition model one by one to obtain the emotion results corresponding to the audio frames.
3. The audio and video emotion labeling method according to claim 2, wherein: The performing feature engineering on each of the audio frames to obtain an embedded vector corresponding to each of the audio frames specifically includes: Extracting amplitude features from each of the audio frames to obtain a corresponding multi-dimensional amplitude vector, wherein the amplitude vector includes root mean square amplitude data, peak amplitude data, and zero-crossing rate data; Performing frequency feature extraction on each of the audio frames to obtain a corresponding multidimensional frequency vector, wherein the multidimensional frequency vector includes Mel spectrum data, MFCC data, spectrum centroid data, and spectrum bandwidth data; Extracting timbre features from each of the audio frames to obtain a corresponding multidimensional timbre vector, the multidimensional timbre vector including spectrum roll-off point data and harmonic distortion ratio data; splicing the multidimensional amplitude vector, the multidimensional frequency vector, and the multidimensional timbre vector to obtain a spliced vector; According to each adjacent audio frame, performing temporal feature enhancement on each of the splicing vectors; A nonlinear transformation is performed on the splicing vector to obtain an embedded vector corresponding to each of the audio frames.
4. The audio and video emotion labeling method according to claim 2, wherein: The emotion recognition model includes a multi-dimensional convolutional layer, a temporal processing layer and a fully connected layer; Inputting the embedded vectors corresponding to the audio frames into the emotion recognition model one by one to obtain the emotion results corresponding to the audio frames specifically includes: Inputting the embedded vector corresponding to each audio frame into the multi-dimensional convolution layer to obtain a corresponding convolution vector; Inputting the convolution vector into the time series processing layer to obtain a time series feature vector; The time series feature vector is input into the fully connected layer and mapped to the corresponding emotion category space to obtain the emotion result corresponding to each audio frame.
5. The audio and video emotion labeling method according to claim 1, wherein: The step of segmenting the audio and video to be marked according to the target audio segment and marking the corresponding emotion result in the corresponding audio and video segment specifically includes: Segmenting the audio and video to be marked according to the target audio segment to obtain corresponding audio and video segments, where the start and end times of the audio and video segments correspond to the target audio segment; The corresponding emotion result is marked in the target frame of the corresponding audio and video segment.
6. The audio and video emotion labeling method according to claim 5, characterized in that: The step of marking the corresponding emotion result in the corresponding audio and video segment specifically includes: Determining a marking style corresponding to each audio and video segment according to the emotion result; Based on the corresponding marking style, each of the audio and video segments is displayed to the user in the audio and video to be marked.
7. The audio and video emotion labeling method according to claim 1, characterized in that: Merging adjacent audio frames matching the emotion results to obtain a corresponding target audio segment specifically includes: Traversing the emotion result list to determine the similarity between the emotion result corresponding to each audio frame and the emotion results corresponding to adjacent audio frames; If the similarity is greater than a predetermined similarity threshold, the audio frame is merged with the adjacent audio frame to obtain a corresponding target audio segment.
8. An audio and video emotion tagging device, characterized in that: include: A data acquisition module is used to obtain the audio stream data of the audio and video to be marked; A frame acquisition module, configured to perform frame processing on the audio stream data to obtain multiple audio frames; A recognition module, configured to input each of the audio frames into an emotion recognition model one by one according to time, to obtain an emotion result corresponding to each of the audio frames; A merging module is used to merge adjacent audio frames that match the emotion results to obtain the corresponding target audio segment; The marking module is used to segment the audio and video to be marked according to the target audio segment, and mark the corresponding emotion result in the corresponding audio and video segment.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory, characterized in that: The processor executes the computer program to implement the steps of the audio and video emotion labeling method according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the audio and video emotion labeling method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Speech emotion recognition method and device, and storage medium
CN110556130A
Speech emotion recognition method and device based on adjacent frame similarity fusion
CN120148559A