Method and electronic device for identifying emotion in video content
By extracting video and audio features and performing multi-layer fusion processing, the limitations and inaccuracy of emotion recognition in the prior art are solved, and more effective identification and more accurate classification of human emotions in video content are achieved.
Patent Information
- Application Number
- CN202380078649.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-11-21
- Filing Date
- 2023-07-06
- Publication Date
- 2025-06-27
AI Technical Summary
The prior art has limitations in identifying emotions contained in video content, usually limited to identifying a small number of core emotions and relying on information from a single modality, resulting in inaccurate results.
By obtaining video sequences of multiple video frames and audio data, video features associated with faces in video frames and audio features associated with audio data, and processing using trained machine learning models, multi-layer fusion of different subsets of video features and audio features is performed to identify human emotions in the video sequence.
It realizes more effective identification of human emotions in video content, supports multimodal understanding, and can identify larger or more detailed emotional classifications, improving the accuracy of recognition.
Smart Images

Figure CN120226015A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure generally relates to machine learning systems. More specifically, the present disclosure relates to methods and electronic devices for identifying at least one emotion in video content. Background Art
[0002] The ability to accurately identify the various emotions contained in video content can be useful in various applications, but current methods for identifying the emotions contained in video content have many drawbacks. For example, these methods are typically limited to identifying a small number of core emotions (such as anger, disgust, fear, happiness, sadness, and surprise), which may limit the effectiveness of these methods. In addition, these methods typically use single-modal information (such as only audio data associated with the video content) to identify the emotions in the video content. Therefore, these methods typically produce inaccurate results. As a specific example, the crowd noise in a stadium during a sports event can typically be classified as associated with a positive emotion of "happiness". However, the actual emotion of the stadium crowd may be "excited" (such as during a specific game period) or "sad" (such as in response to a game or attempt by a particular team to lose). Summary of the Invention
[0003] Technical Solution
[0004] In an embodiment, a method includes obtaining a video sequence having a plurality of video frames and audio data. The method further includes extracting video features associated with at least one face in the video frames and audio features associated with the audio data. The method further includes using a trained machine learning model to process the video features and the audio features. The trained machine learning model performs multi-layer fusion of different subsets of the video features and the audio features to identify at least one emotion expressed by at least one person in the video sequence.
[0005] In an embodiment, an electronic device includes at least one memory configured to store a video sequence having a plurality of video frames and audio data. The electronic device further includes at least one processor configured to extract video features associated with at least one face in the video frames and audio features associated with the audio data, and to use a trained machine learning model to process the video features and the audio features. The trained machine learning model is configured to perform multi-layer fusion of different subsets of the video features and the audio features to identify at least one emotion expressed by at least one person in the video sequence.
[0006] In an embodiment, a computer-readable medium includes instructions that, when executed, cause at least one processor to obtain a video sequence having a plurality of video frames and audio data. The computer-readable medium further includes instructions that, when executed, cause the at least one processor to extract video features associated with at least one face in the video frames and audio features associated with the audio data. The computer-readable medium further includes instructions that, when executed, cause the at least one processor to process the video features and the audio features using a trained machine learning model. The trained machine learning model is configured to perform multi-layer fusion of different subsets of the video features and the audio features in order to identify at least one emotion expressed by at least one person in the video sequence. BRIEF DESCRIPTION OF THE DRAWINGS
[0007] To more fully understand the present disclosure and its advantages, reference is now made to the following description taken in conjunction with the accompanying drawings, in which like reference numerals represent like parts:
[0008] Figure 1 FIG. 1 illustrates an example network configuration including an electronic device in accordance with the present disclosure;
[0009] Figure 2 FIG. 2 illustrates an example architecture in accordance with the present disclosure that supports multimodal understanding of emotions in video content;
[0010] Figure 3 FIG. 3 illustrates an example machine learning model in an architecture in accordance with the present disclosure for Figure 2 ;
[0011] Figure 4 FIG. 4 illustrates an example machine learning model in an architecture in accordance with the present disclosure for Figure 2 ;
[0012] Figure 5 FIG. 5 illustrates an example machine learning model in an architecture in accordance with the present disclosure for Figure 2 ; and
[0013] Figure 6 FIG. 6 illustrates an example method for multimodal understanding of emotions in video content in accordance with the present disclosure. DETAILED DESCRIPTION
[0014] In the present disclosure, the terms "send", "receive", and "communicate" and their derivatives cover both direct and indirect communication. The terms "comprise" and "include" and their derivatives mean including but not limited to. The term "or" is inclusive in meaning and / or. The phrase "associated with" and its derivatives mean including, being included within, being interconnected with, containing, being contained within, being connected to or connected with, being coupled to or coupled with, being communicable with, cooperating with, being interleaved, being juxtaposed, being close to, being bound to or bound with, having, having the property of, having a relationship to or with, etc.
[0015] In addition, the various functions described below can be implemented or supported by one or more computer programs, each computer program being formed of computer-readable program code and embodied in a computer-readable medium. The terms "application" and "program" refer to one or more computer programs, software components, instruction sets, procedures, functions, objects, classes, instances, related data, or a portion thereof that are adapted to be implemented in suitable computer-readable program code. The phrase "computer-readable program code" includes any type of computer code, including source code, object code, and executable code. The phrase "computer-readable medium" includes any type of medium that can be accessed by a computer, such as read-only memory (ROM), random access memory (RAM), hard disk drive, compact disc (CD), digital video disc (DVD), or any other type of memory. A computer-readable medium does not include a wired, wireless, optical, or other communication link that transmits transient electrical signals or other signals. A computer-readable medium includes a medium in which data can be permanently stored and a medium in which data can be stored and later rewritten, such as a rewritable optical disc or an erasable memory device.
[0016] As used herein, terms and phrases such as "have", "may have", "include", or "may include" a feature (such as a number, a function, an operation, or a component such as a part) indicate the presence of the feature without precluding the presence of other features. In addition, as used herein, the phrase "A or B", "at least one of A and / or B", or "one or more of A and / or B" may include all possible combinations of A and B. For example, "A or B", "at least one of A and B", and "at least one of A or B" may indicate any of the following: (1) including at least one A, (2) including at least one B, or (3) including at least one A and at least one B. In addition, as used herein, the terms "first" and "second" may modify various components regardless of importance and do not limit the components. These terms are only used to distinguish one component from another. For example, a first user device and a second user device may indicate different user devices from each other, regardless of the order or importance of the devices. Without departing from the scope of the present disclosure, the first component may be represented as the second component, and vice versa.
[0017] It should be understood that when an element (such as a first element) is referred to as being “coupled” / “coupled to” or “connected to” another element (such as a second element) (operatively or communicatively), it can be coupled or connected to the other element directly or via a third element. In contrast, it will be understood that when an element (such as a first element) is referred to as being “directly coupled” or “directly connected” to another element (such as a second element), no other element (such as a third element) intervenes between the element and the other element.
[0018] As used herein, the phrase “configured (or set) to” can be used interchangeably with the phrases “suitable for”, “capable of”, “designed to”, “adapted to”, “manufactured to”, or “able to” as the context may require. The phrase “configured (or set) to” does not essentially mean “specially designed in hardware”. Instead, the phrase “configured to” can mean that a device can perform an operation in conjunction with another device or part. For example, the phrase “a processor configured (or set) to perform A, B, and C” can mean a general-purpose processor (such as a CPU or an application processor) that can perform the operation by executing one or more software programs stored in a memory device, or a dedicated processor (such as an embedded processor) for performing the operation.
[0019] The terms and phrases used herein are for describing some embodiments of the present disclosure only and do not limit the scope of other embodiments of the present disclosure. It should be understood that unless the context clearly indicates otherwise, the singular forms “a”, “an”, and “the” include plural referents. All terms and phrases used herein (including technical and scientific terms and phrases) have the same meaning as commonly understood by one of ordinary skill in the art to which the embodiments of the present disclosure pertain. It will be further understood that terms and phrases (such as those defined in a common dictionary) should be interpreted as having a meaning consistent with their meaning in the context of the relevant art and will not be interpreted in an idealized or overly formal sense unless expressly so defined herein. In an embodiment, the terms and phrases defined herein may be interpreted to exclude embodiments of the present disclosure.
[0020] Examples of an "electronic device" according to an embodiment of the present disclosure may include at least one of a smart phone, a tablet personal computer (PC), a mobile phone, a video phone, an e-book reader, a desktop PC, a laptop computer, a netbook computer, a workstation, a personal digital assistant (PDA), a portable multimedia player (PMP), an MP3 player, a mobile medical device, a camera, or a wearable device (such as smart glasses, a head-mounted device (HMD), electronic clothing, an electronic bracelet, an electronic necklace, an electronic accessory, an electronic tattoo, a smart mirror, or a smart watch). Other examples of the electronic device include smart home appliances. Examples of the smart home appliances may include at least one of a television, a digital video disc (DVD) player, an audio player, a refrigerator, an air conditioner, a vacuum cleaner, an oven, a microwave oven, a washing machine, a dryer, an air purifier, a set-top box, a home automation control panel, a security control panel, a TV box (such as SAMSUNG HOMESYNC, APPLE TV, or GOOGLE TV), a smart speaker or a speaker having an integrated digital assistant (such as SAMSUNG GALAXY HOME, APPLE HOMEPOD, or AMAZON ECHO), a game console (such as XBOX, PLAYSTATION, or NINTENDO), an electronic dictionary, an electronic key, a camera, or an electronic photo frame. Other examples of the electronic device include at least one of the following: various medical devices (such as various portable medical measurement devices (e.g., a blood glucose measurement device, a heart rate measurement device, or a body temperature measurement device), a magnetic resource angiography (MRA) device, a magnetic resource imaging (MRI) device, a computed tomography (CT) device, an imaging device, or an ultrasonic device), a navigation device, a global positioning system (GPS) receiver, an event data recorder (EDR), a flight data recorder (FDR), an in-vehicle infotainment device, marine electronic devices (such as a marine navigation device or a gyrocompass), avionics, a security device, an in-vehicle head unit, an industrial or home robot, an automated teller machine (ATM), a point of sale (POS) device, or an Internet of Things (IoT) device (such as a light bulb, various sensors, a water meter, an electricity meter, or a gas meter, a sprinkler, a fire alarm, a thermostat, a street lamp, an oven, a fitness device, a hot water tank, a heater, or a boiler). Other examples of the electronic device include at least one of a piece of furniture or at least a part of a building / structure, an electronic board, an electronic signature receiving device, a projector, or various measurement devices (such as a device for measuring water, electricity, gas, or electromagnetic waves). Note that, according to various embodiments of the present disclosure, the electronic device may be one or a combination of the devices listed above. According to an embodiment of the present disclosure, the electronic device may be a flexible electronic device. The electronic device disclosed herein is not limited to the devices listed above and may include new electronic devices depending on technological development.
[0021] In the following description, according to various embodiments of the present disclosure, an electronic device is described with reference to the accompanying drawings. As used herein, the term "user" may refer to a person using the electronic device or another device (such as an artificial intelligence electronic device).
[0022] Throughout this patent document, definitions of certain other words and phrases may be provided. Those of ordinary skill in the art should understand that in many, if not most, instances, such definitions apply to the prior and future use of the words and phrases so defined.
[0023] The present disclosure relates to a method and an electronic device for multimodal understanding of emotions in video content. The present disclosure relates to a method and an electronic device for identifying at least one emotion in video content.
[0024] The following discussion is described with reference to the accompanying drawings Figures 1 to 6 and various embodiments of the present disclosure. However, it should be understood that the present disclosure is not limited to these embodiments, and all changes and / or equivalents or substitutions thereof also fall within the scope of the present disclosure. Throughout the specification and the drawings, the same or similar reference numerals may be used to refer to the same or similar elements.
[0025] As described above, the ability to accurately identify various emotions contained in video content may be useful in various applications, but current methods for identifying emotions in video content have many drawbacks. For example, these methods are typically limited to identifying a small number of core emotions (such as anger, disgust, fear, happiness, sadness, and surprise), which may limit the effectiveness of these methods. In addition, these methods typically use single-modal information (such as only audio data associated with the video content) to identify emotions in the video content. Therefore, these methods typically produce inaccurate results. As a specific example, the crowd noise in a stadium during a sports event can typically be classified as being associated with a positive emotion of "happiness". However, the actual emotion of the stadium crowd may be "excited" (such as during a particular game period) or "sad" (such as in response to a game in which a particular team loses).
[0026] The present disclosure provides a method and an electronic device (or technology) for multimodal understanding of emotions in video content. The emotion in the video content may refer to at least one emotion (or at least one human emotion) expressed by at least one person in the video sequence. As described in more detail below, a video sequence can be obtained, where the video sequence includes (i) a plurality of video frames and (ii) audio data. At least one of the video frames in the video sequence can capture the face of at least one person. Video features associated with at least one face in the video frames are extracted, and audio features associated with the audio data are extracted. In an embodiment, the video features can be extracted by dividing the video frames into multiple sets, performing face detection using the sets, and processing the sets based on the face detection results to identify the video features associated with at least one face. Additionally, in an embodiment, the audio features can be extracted by processing the audio data to identify a first subset of audio features (such as features associated with the original audio waveform of the audio data) and a second subset of audio features (such as features determined using a pre-trained audio model and based on the audio data).
[0027] The video features and the audio features are provided to a trained machine learning model and processed using the trained machine learning model, where the trained machine learning model (which includes) performs multi-layer fusion of different subsets of the video features and the audio features to identify at least one emotion expressed by at least one person in the video sequence. For example, the trained machine learning model can perform (i) a first fusion of the video features and a first subset of the audio features and (ii) a second fusion of the processed features and a second subset of the audio features, where the processed features are based on the first fusion. In an embodiment, the trained machine learning model can include (i) at least one cross-modal transformer encoder layer that receives and fuses the video features and the first subset of the audio features and generates multi-modal features, (ii) at least one fusion encoder layer that combines the multi-modal features, and (iii) a multi-layer perceptron decoder layer that decodes the output of the fusion encoder layer fused with the second subset of the audio features.
[0028] In this way, the described technology enables more effective recognition of human emotions contained in video content. Additionally, the described technology supports the use of multiple modalities since the machine learning model can process features associated with both video data and audio data of the video content. Human emotions contained in the video content can be better inferred based on an effective fusion of information from multiple modalities (such as audio and face) present in the video content, and the machine learning model can be effectively trained to utilize both audio and visual modalities to identify emotions in the video content. Furthermore, the described technology can support a larger or more elaborate classification of human emotions, which allows for the detection of subtle emotions in the video content. As a specific example, in video content showing a sports event, the described technology can detect emotions such as positive happiness (usually) or positive happiness (excitement) when the crowd expects a score, or negative shock or surprise (usually) when an athlete faces an injury (rather than just "happy" or "sad"). Additionally, a large emotion video dataset can be used to train the machine learning model, and the large emotion video dataset can incorporate various levels of emotion annotations as well as modalities. This helps improve the training of the machine learning model and increases the overall accuracy of the machine learning model.
[0029] Figure 1 An example network configuration 100 including an electronic device 101 in accordance with the present disclosure is shown. Figure 1 The illustrated embodiment of the network configuration 100 is for illustrative purposes only. Other embodiments of the network configuration 100 may be used without departing from the scope of the present disclosure.
[0030] According to an embodiment of the present disclosure, the electronic device 101 is included in the network configuration 100. The electronic device 101 may include at least one of a bus 110, a processor 120, a memory 130, an input / output (I / O) interface 150, a display 160, a communication interface 170, or a sensor 180. In an embodiment, the electronic device 101 may exclude at least one of these components, or may add at least one other component. The bus 110 includes circuitry for connecting the components 120 - 180 to each other and for transmitting communications (such as control messages and / or data) between the components.
[0031] The processor 120 includes one or more processing devices, such as one or more microprocessors, microcontrollers, digital signal processors (DSPs), application specific integrated circuits (ASICs), or field programmable gate arrays (FPGAs). In an embodiment, the processor 120 includes one or more of a central processing unit (CPU), an application processor (AP), a communication processor (CP), or a graphics processing unit (GPU). The processor 120 is capable of performing control and / or executing operations or data processing related to communication or other functions on at least one of the other components of the electronic device 101. As described below, the processor 120 can be used to process video sequences using a feature extractor and a multimodal machine learning model to identify the emotions of one or more persons in each video sequence. The processor 120 can also or alternatively be used to perform one or more actions or initiate the execution of one or more actions based on or in response to at least one emotion identified in at least one video sequence.
[0032] The memory 130 may include volatile and / or non-volatile memory. For example, the memory 130 may store commands or data related to at least one of the other components of the electronic device 101. According to an embodiment of the present disclosure, the memory 130 may store software and / or a program 140. The program 140 includes, for example, a kernel 141, middleware 143, an application programming interface (API) 145, and / or an application program (or “app”) 147. At least a portion of the kernel 141, middleware 143, or API 145 may be represented as an operating system (OS).
[0033] The kernel 141 may control or manage system resources (such as the bus 110, the processor 120, or the memory 130) for performing operations or functions implemented in other programs (such as middleware 143, API 145, or application 147). The kernel 141 provides an interface that allows the middleware 143, API 145, or application 147 to access various components of the electronic device 101 to control or manage system resources. The application 147 may include one or more applications for supporting or using multimodal understanding of emotions in video content. These functions may be performed by a single application or multiple applications, with each application performing one or more of these functions. For example, the middleware 143 may act as a repeater to allow the API 145 or application 147 to communicate data with the kernel 141. In an embodiment, multiple applications 147 may be provided. The middleware 143 is capable of controlling work requests received from the application 147, such as by assigning priorities for using system resources of the electronic device 101 (such as the bus 110, the processor 120, or the memory 130) to at least one of the multiple applications 147. The API 145 is an interface that allows the application 147 to control functions provided by the kernel 141 or the middleware 143. For example, the API 145 includes at least one interface or function (such as a command) for archive control, window control, image processing, or text control.
[0034] The I / O interface 150 serves as an interface that can transfer, for example, commands or data input from a user or other external device to other components of the electronic device 101. The I / O interface 150 may also output commands or data received from other components of the electronic device 101 to the user or other external device.
[0035] The display 160 includes at least one of, for example, a liquid crystal display (LCD), a light-emitting diode (LED) display, an organic light-emitting diode (OLED) display, a quantum dot light-emitting diode (QLED) display, a microelectromechanical systems (MEMS) display, or an electronic paper display. The display 160 may also be a depth perception display, such as a multi-focus display. The display 160 is capable of displaying various contents (such as text, images, videos, icons, or symbols) to the user. The display 160 may include a touch screen and may receive, for example, touch, gesture, proximity, or hover inputs using an electronic pen or a user's body part.
[0036] For example, the communication interface 170 is capable of establishing communication between the electronic device 101 and an external electronic device (such as the first electronic device 102, the second electronic device 104, or the server 106). For example, the communication interface 170 may be connected to the network 162 or 164 through wireless or wired communication to communicate with the external electronic device. The communication interface 170 may be a wired or wireless transceiver or any other component for transmitting and receiving signals (such as images).
[0037] The electronic device 101 also includes one or more sensors 180, which may measure physical quantities or detect the activation state of the electronic device 101 and convert the measured or detected information into an electrical signal. For example, the one or more sensors 180 may include one or more cameras or other imaging sensors, which may be used to capture images of a scene. In an embodiment, the one or more cameras or other imaging sensors may capture consecutive video frames. The time interval (or frame rate) of each of the video frames may be predetermined according to the settings of the manufacturer or the user. The sensor 180 may also include one or more buttons for touch input, one or more microphones, a gesture sensor, a gyroscope or gyroscope sensor, a barometric pressure sensor, a magnetic sensor or magnetometer, an acceleration sensor or accelerometer, a grip sensor, a proximity sensor, a color sensor (such as an RGB sensor), a biophysical sensor, a temperature sensor, a humidity sensor, an illuminance sensor, an ultraviolet (UV) sensor, an electromyogram (EMG) sensor, an electroencephalogram (EEG) sensor, an electrocardiogram (ECG) sensor, an infrared (IR) sensor, an ultrasonic sensor, an iris sensor, or a fingerprint sensor. The (one or more) sensors 180 may also include an inertial measurement unit, which may include one or more accelerometers, gyroscopes, and other components. Additionally, the (one or more) sensors 180 may include a control circuit for controlling at least one of the sensors included herein. Any one of these sensors 180 may be located within the electronic device 101.
[0038] The first electronic device 102 or the second electronic device 104 may be a wearable device or a wearable device mountable to an electronic device (such as a head-mounted display (HMD)). When the electronic device 101 is mounted in the first electronic device 102 (such as an HMD), the electronic device 101 may communicate with the first electronic device 102 through the communication interface 170. The electronic device 101 may be directly connected to the first electronic device 102 to communicate with the first electronic device 102 without involving a separate network. The electronic device 101 may also be an augmented reality wearable device including one or more cameras, such as glasses.
[0039] Wireless communication can use at least one of, for example, Long Term Evolution (LTE), Long Term Evolution - Advanced (LTE - A), 5th Generation Wireless System (5G), millimeter - wave or 60 GHz wireless communication, Wireless USB, Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Universal Mobile Telecommunications System (UMTS), Wireless Broadband (WiBro), or Global System for Mobile Communications (GSM) as a cellular communication protocol. The wired connection can include at least one of, for example, Universal Serial Bus (USB), High - Definition Multimedia Interface (HDMI), Recommended Standard 232 (RS - 232), or Plain Old Telephone Service (POTS). Network 162 includes at least one communication network, such as a computer network (e.g., Local Area Network (LAN) or Wide Area Network (WAN)), the Internet, or a telephone network.
[0040] The first electronic device 102, the second electronic device 104, and the server 106 can each be a device of the same or different type as the electronic device 101. According to an embodiment of the present disclosure, the server 106 includes a group of one or more servers. In addition, according to an embodiment of the present disclosure, all or some of the operations performed on the electronic device 101 can be performed on another or multiple other electronic devices (such as the first electronic device 102, the second electronic device 104, or the server 106). In addition, according to an embodiment of the present disclosure, when the electronic device 101 should automatically or upon request perform some function or service, the electronic device 101 can request another device (such as the first electronic device 102, the second electronic device 104, or the server 106) to perform at least some functions associated therewith, rather than performing the function or service itself, or additionally performing the function or service. Another electronic device (such as the first electronic device 102, the second electronic device 104, or the server 106) can perform the requested function or additional functions and transmit the result of the execution to the electronic device 101. The electronic device 101 can provide the requested function or service by processing the received result as is or additionally. For this purpose, for example, cloud computing, distributed computing, or client - server computing technologies can be used. Although Figure 1 The electronic device 101 is shown to include a communication interface 170 that communicates with the second electronic device 104 or the server 106 via the network 162, but according to an embodiment of the present disclosure, the electronic device 101 can operate independently without a separate communication function.
[0041] Server 106 may include components that are the same as or similar to those of electronic device 101 (or a suitable subset thereof). Server 106 may support driving electronic device 101 by performing at least one of the operations (or functions) implemented on electronic device 101. For example, server 106 may include a processing module or processor that may support processor 120 implemented in electronic device 101. As described below, server 106 may be used to process video sequences using a feature extractor and a multimodal machine learning model to identify the emotions of one or more persons in each video sequence. Server 106 may also or alternatively be used to perform one or more actions or initiate the execution of one or more actions based on or in response to at least one emotion identified in at least one video sequence. Based on the video frames and audio data included in the video sequence, the at least one identified emotion may be referred to as one or more predicted emotions or one or more estimated emotions.
[0042] Although Figure 1 an example of network configuration 100 including electronic device 101 is shown, various changes may be made to Figure 1 it. For example, network configuration 100 may include any number of each component in any suitable arrangement. Generally, computing and communication systems have a wide variety of configurations, and Figure 1 the scope of the present disclosure is not limited to any particular configuration. Additionally, although Figure 1 an operating environment in which various features of the present disclosure may be used is shown, these features may be used in any other suitable system.
[0043] Figure 2 An example architecture 200 that supports multimodal understanding of emotions in video content according to the present disclosure is shown. For ease of explanation, Figure 2 the architecture 200 shown is described as being implemented on or supported by electronic device 101 in Figure 1 network configuration 100. However, Figure 2 the architecture 200 shown in may be used with any other suitable device and may be used in any other suitable system, such as when the architecture 200 is implemented on or supported by server 106.
[0044] As Figure 2As shown, architecture 200 generally receives and processes video sequence 202. Each video sequence 202 includes a plurality of video frames 204 and audio data 206. The plurality of video frames may be referred to as a plurality of video frames. Each video sequence 202 can be obtained from any suitable source. In an embodiment, for example, each video sequence 202 can represent video content provided from server 106 or other source to electronic device 101 for presentation by electronic device 101 to one or more viewers. The video sequence 202 can represent any suitable video content, such as a movie, a television program, a home video, a long or short video clip available for downloading or viewing over the Internet, or any other suitable video content. Each video sequence 202 can have any suitable length, possibly ranging from a relatively short video clip to a lengthy movie or other video.
[0045] Each video frame 204 can represent one of the images included in the sequence of images in the associated video sequence 202. Each video frame 204 can have any suitable format and resolution, possibly up to and including 4K or 8K resolution or even higher. A variety of processes can be used to capture each video frame 204 from the associated video sequence 202. By way of example, the various processes can include processes using a video player, video editor software, or video capture software. Architecture 200 can capture one or more faces of one or more persons from at least some of the video frames 204 in each video sequence 202. The audio data 206 represents an audio waveform that can be reproduced during the presentation of the associated video sequence 202 to provide audible sound to one or more viewers. For example, the audio data 206 can include one or more voices of one or more persons speaking, cheering, or otherwise making sounds in the associated video sequence 202, music or sound effects in the associated video sequence 202, or sounds in the scene captured in the associated video sequence 202. The audio data 206 can have any suitable format and any suitable resolution, possibly up to and including 24-bit or 32-bit resolution or even higher. The audio data 206 can also have any suitable data rate, possibly up to and including 256 kbit / sec or even higher. The audio data 206 can be captured, obtained, or detected from the associated video sequence 202 using audio analysis software or audio capture software. In an embodiment, the video frame 204 can be referred to as a video frame acquisition processor, video frame acquisition function, video frame acquisition module, or video frame acquisition unit for obtaining the video frame 204 from each video sequence 202. In an embodiment, the audio data 206 can be referred to as an audio acquisition processor, audio acquisition function, audio acquisition module, or audio acquisition unit for obtaining the audio data 206 from each video sequence 202.
[0046] In an embodiment, the video frames 204 of each video sequence 202 can be provided to and processed using a face detection and video feature extraction function 208. The face detection and video feature extraction function 208 generally operates to identify the locations of the faces of one or more persons in the video frame 204 and to extract or otherwise identify features of the video frame 204. The identified features of the video frame 204 include (at least) features related to the identified faces within the video frame 204, meaning that the identified features of the video frame 204 include the facial features of one or more persons captured in the video frame 204. In an embodiment, the face detection and video feature extraction function 208 can process the video frames 204 individually. In an embodiment, the face detection and video feature extraction function 208 can process a set of video frames 204, such as when the video frames 204 are grouped into sets of relatively short durations (such as sets each of approximately six seconds in length).
[0047] The face detection and video feature extraction function 208 can use any suitable technique to perform face detection and extract facial features or other video features of the video frames 204 in one or more video sequences 202. The face detection and video feature extraction function 208 can be referred to as a face detection and video feature extraction processor, a face detection and video feature extraction module, a face detection and video feature extraction section, or a face detection and video feature extraction unit. In an embodiment, the face detection and video feature extraction function 208 can be implemented using one or more machine learning models that have been trained to perform face detection and video feature extraction. In an embodiment, the face detection portion of the face detection and video feature extraction function 208 can be implemented using a Multi-task Cascaded Convolutional Network (MTCNN), and the video feature extraction portion of the face detection and video feature extraction function 208 can be implemented using a Self-Correcting Network (SCN). The MTCNN or other machine learning models can be used to support face detection (which involves identifying the locations of faces in the video frame 204) and facial landmark alignment (which involves identifying the locations of a person's eyes, nose, mouth, or other facial landmarks in the video frame 204). The SCN or other machine learning models can be used to identify features associated with the facial expressions of persons in the video frame 204. In an embodiment, the SCN includes five frozen layers and supports up to 512 classes related to facial features.
[0048] In the present disclosure, the audio data 206 of each video sequence 202 can be provided to at least one audio feature extraction function 210 and processed using the at least one audio feature extraction function 210. The audio feature extraction function 210 generally operates to identify multiple subsets of audio features associated with the audio data 206. For example, the audio feature extraction function 210 can identify a first subset of audio features of the audio data 206, where these features can include general features based on the original audio waveform of the audio data 206. Examples of features that can be used in the first subset of audio features can include energy-related features, spectrum-related features, or other features of the audio data 206. The audio feature extraction function 210 can also identify a second subset of audio features of the audio data 206, where these features can be generated using a pre-trained audio model and are based on the audio data 206.
[0049] The audio feature extraction function 210 can use any suitable technique to extract the audio features of the audio data 206 in one or more video sequences 202. In an embodiment, the audio feature extraction function 210 can be implemented at least in part using one or more machine learning models that have been trained to perform audio feature extraction. In an embodiment, the audio feature extraction function 210 can be implemented using PyAudio analysis to identify the first subset of audio features of each instance of the audio data 206, and using the pre-training, sampling, labeling, and aggregation (PSLA) process (or aggregation) in the PSLA model to identify the second subset of audio features of each instance of the audio data 206. PyAudio analysis is the process of analyzing audio in Python. In the PSLA model, the aggregation process can extract the second subset of audio features of each instance of the audio data 206, label the second subset of audio features, and aggregate the labeled second subset of audio features. In an embodiment, the PSLA model can be implemented using an EfficientNet B2 pre-trained model with four attention heads. The audio feature extraction function 210 can be referred to as an audio feature extraction processor, an audio feature extraction module, an audio feature extraction section, or an audio feature extraction unit.
[0050] The extracted video features and the extracted audio features of each video sequence 202 are provided to a trained machine learning model 212, which generally operates to process the extracted features and generate one or more predicted emotions 214 expressed by at least one person in each video sequence 202. For example, the trained machine learning model 212 can be trained to perform multi-layer fusion of different subsets of the video features and the audio features during the identification of one or more predicted emotions 214 expressed by at least one person in each video sequence 202. The trained machine learning model 212 can have any suitable machine learning-based structure configured to process the video and audio features and estimate the emotions contained in the video content. In an embodiment, a trained machine learning model 212 can be implemented using an attention-based transformer architecture, which can be used to support multi-modal feature fusion techniques. In an embodiment, a trained machine learning model 212 can be implemented using a single-layer transformer with four attention heads for audio and eight attention heads for video. The following describes an example implementation of the trained machine learning model 212 with respect to Figures 3 to 5 Describe an example implementation of the trained machine learning model 212.
[0051] The trained machine learning model 212 here can be trained to effectively identify human emotions in video content using both the audio and visual modalities present in the video content. As described below, the unimodal features of each video sequence 202 (i.e., the first subset of the video / facial features and the audio features) can be fused, such as via concatenation or other suitable fusion techniques, and processed using one or more cross-modal transformer encoder layers and one or more fusion encoder layers of a multi-modal transformer in the trained machine learning model 212. The final output of the multi-modal transformer is then fused with the other audio features of each video sequence 202 (i.e., the second subset of the audio features), and the fused features are processed using the decoder of the trained machine learning model 212 to produce one or more predicted emotions 214 for each video sequence 202.
[0052] The trained machine learning model 212 can be trained in any suitable manner. For example, in an embodiment, training data and ground truth data can be obtained, where (i) the training data can include a video sequence 202 containing faces of various people showing various emotions, and (ii) the ground truth data can include the correct output generated by the trained machine learning model 212 using the video sequence 202 of the training data. The training data can be provided to the face detection and video feature extraction function 208 and the audio feature extraction function 210 to extract video and audio features from the video sequence 202 in the training data. The video and audio features can be provided to the trained machine learning model 212, and the trained machine learning model 212 can be used to generate one or more predicted emotions 214 for the video sequence 202 in the training data. The one or more predicted emotions 214 can be compared with the ground truth data, and the differences or errors between the one or more predicted emotions 214 and the ground truth data can be identified and used to calculate the overall error or loss of the trained machine learning model 212. Any suitable loss function can be used here to calculate the loss, such as the focal loss with a learning rate of 0.0001 and a dropout rate of 0.2. If the calculated loss exceeds a threshold, the parameters of the trained machine learning model 212 can be adjusted, and the updated trained machine learning model 212 can be used to process the video sequence 202 in the training data again (or a new video sequence 202 in different training data can be processed) to generate new one or more predicted emotions 214, which can be compared with the ground truth data to calculate the updated loss. The calculated loss decreases over time and eventually falls below the threshold, indicating that the trained machine learning model 212 has been trained to accurately (at least within the desired accuracy represented by the threshold) predict the emotions of the people in the video sequence 202.
[0053] The trained machine learning model 212 here can be trained to recognize a variety of human emotions in video content. In an embodiment, for example, the trained machine learning model 212 can be trained to recognize emotions arranged in a hierarchy, where the two root categories of the hierarchy can include positive emotions and negative emotions. In an embodiment, the trained machine learning model 212 can be trained to recognize positive emotions such as happiness, surprise or wonder, love, hope, and curiosity. For positive happiness, the trained machine learning model 212 can be trained to recognize specific types of happiness, such as humor or comedy, excitement, pride, and relief. For positive love, the trained machine learning model 212 can be trained to recognize specific types of love, such as romantic and platonic. For positive hope, the trained machine learning model 212 can be trained to recognize specific types of hope, such as faith and confidence. For example, the trained machine learning model 212 can be trained to recognize negative emotions such as sadness, anger, surprise or shock, arrogance, disgust, fear, jealousy, and doubt. For negative sadness, the trained machine learning model 212 can be trained to recognize specific types of sadness, such as embarrassment, grief, and regret. For negative anger, the trained machine learning model 212 can be trained to recognize specific types of anger, such as frustration, argument, and rage. For negative fear, the trained machine learning model 212 can be trained to recognize specific types of fear, such as nervousness and horror. Fuzzy classification can also be supported to recognize emotions that may be positive or negative, such as surprise or startle, confusion, and irony.
[0054] To support this or other types of hierarchical arrangements of possible emotions, a sufficient amount of training data can be used during the training process so as to effectively train the trained machine learning model 212 on how to recognize these emotions and distinguish between similar emotions. This may involve using a very large training data set, possibly including a training data set having hundreds of thousands or more video sequence training samples and associated ground truth data. This type of training data set can be obtained in any suitable manner. For example, this type of training data set can be obtained by automatically extracting video sequences containing facial expressions (such as from video channels owned and operated by SAMSUNG TV PLUS or “O&O” or the 8M video data set or other public / private data sets) and having humans manually annotate the extracted video sequences with ground truth labels identifying the actual emotions in the extracted video sequences. Additionally, for example, each of the video sequences 202 in the training data set can be relatively short, such as when each of the video sequences 202 has a length of about six seconds or less (although other suitable durations can be used).
[0055] It should be noted that Figure 2 shown in or with respect to Figure 2The described functionality may be implemented in any suitable manner in the electronic device 101, the server 106, or other devices. For example, in an embodiment, one or more software applications or other software instructions executed by a processor (or at least one processor) 120 of the electronic device 101, the server 106, or other devices may be used to implement or support Figure 2 shown in or with respect to Figure 2 at least some of the functionality described. In an embodiment, dedicated hardware components may be used to implement or support at least some of the functionality shown in or with respect to Figure 2 shown in or with respect to Figure 2 at least some of the functionality described. Generally, any suitable hardware or any suitable combination of hardware and software / firmware instructions may be used to perform Figure 2 shown in or with respect to Figure 2 the functionality described. Moreover, a single device or multiple devices may be used to perform Figure 2 shown in or with respect to Figure 2 the functionality described.
[0056] Although Figure 2 an example of an architecture 200 that supports multimodal understanding of emotions in video content is shown, various changes may be made to Figure 2 it. For example, Figure 2 the various components and functions in it may be combined, further subdivided, replicated, or rearranged according to specific needs. In addition, one or more additional components and functions may be included if needed or desired.
[0057] Figures 3 to 5 shown is a machine learning model 212 trained according to an example in the architecture 200 for Figure 2 . As Figures 3 to 5 shown, in an embodiment, the audio feature extraction function 210 may be implemented using separate audio feature extraction functions 210a - 210b. The audio feature extraction function 210a may be used to identify general features of the original audio waveform in the audio data 206, such as energy-related features, spectrum-related features, or other features of the audio data 206. For example, the audio feature extraction function 210a may be implemented using PyAudio analysis. The audio feature extraction function 210a may be referred to as a waveform-based audio feature extraction function or a first audio feature extraction function. The audio feature extraction function 210b may be used to identify features of the audio data 206 using a pre-trained audio model. For example, the audio feature extraction function 210b may be implemented using the PSLA model. The audio feature extraction function 210b may be referred to as a model-based audio feature extraction function or a second audio feature extraction function.
[0058] In embodiments incorporating the MTCNN / SCN or the face detection and video feature extraction function 208, the architecture 200 may use a combination of pre-trained machine learning models in the respective domains to perform efficient feature extraction for multiple modalities (i.e., visual (face) and audio). In an embodiment, the face detection and video feature extraction function 208 may generate a feature vector with 512 dimensions, the audio feature extraction function 210a may generate a feature vector with 68 dimensions, and the audio feature extraction function 210b may generate a feature vector with 527 dimensions. However, note that these values are only examples and may vary according to needs or expectations.
[0059] In Figures 3 to 5 the example of, a trained machine learning model 212 may be implemented using an attention-based transformer architecture. Here, the trained machine learning model 212 may include a multimodal transformer 302, which is implemented using one or more cross-modal transformer encoder layers 304 and one or more fusion encoder layers 306. One or more cross-modal transformer encoder layers 304 may receive and fuse (among other things) video features determined by the face detection and video feature extraction function 208 and a first subset of audio features determined by the audio feature extraction function 210a. For example, the video features and the first subset of audio features may be concatenated or otherwise combined, and the resulting combined features may be processed by one or more cross-modal transformer encoder layers 304. One or more cross-modal transformer encoder layers 304 are responsible for encoding the combined video and audio features into a suitable form for subsequent processing. As a result, one or more cross-modal transformer encoder layers 304 may produce processed features that are based on the fusion of the video features and the first subset of audio features. In an embodiment, each of the one or more cross-modal transformer encoder layers 304 may include a self-attention mechanism and a feed-forward network.
[0060] The processed feature representations from one or more cross-modal transformer encoder layers 304 represent multimodal features because they are formed using a combination of video and audio features. Here, the multimodal features are provided to one or more fusion encoder layers 306, which may encode the multimodal features from the cross-modal transformer encoder layers 304. For example, one or more fusion encoder layers 306 may be used to learn the relationships between various multimodal features when identifying emotions in video content. This allows the trained machine learning model 212 to learn how to relate various multimodal features based on video and audio features to each other.
[0061] The output from one or more fusion encoder layers 306 can represent the final output of the multimodal transformer 302. As shown here, these outputs can be referred to as the transformer outputs of the multimodal transformer 302. These outputs (or transformer outputs) are fused (such as via concatenation) with a second subset of audio features (or spectral features) determined by the audio feature extraction function 210b via the fusion function 308. The resulting fused values are provided to the decoder 310, which decodes the fused features produced by the fusion function 308. For example, the decoder 310 can represent at least a portion of a machine learning model that has been trained to combine the fused features in a manner that results in one or more predicted emotions 214 for generating the video sequence 202. For example, the decoder 310 can be implemented as a multi-layer perceptron (MLP) layer that can include an input layer that receives the fused features, a hidden layer that jointly processes the fused features using (among other things) a non-linear activation function, and an output layer that provides one or more predicted emotions 214 based on the output from the last hidden layer. As a specific example, the decoder 310 can be implemented as a multi-layer perceptron layer with one hundred hidden layers.
[0062] Generally speaking, the illustrated embodiment of the trained machine learning model 212 can adjust the transformer model to learn temporal relationships in unimodal features by leveraging cross-modal inputs and incorporating feature fusion techniques. Figures 3 to 5 The embodiment of the trained machine learning model 212 shown in differs in how the video and audio features are fused. For example, as Figure 3 shown, there can be a single cross-modal transformer encoder layer 304 that receives and combines a first subset of video and audio features, and multiple (such as three) fusion encoder layers 306 that process the output of the cross-modal transformer encoder layer 304. This type of approach can be said to represent an "early-late" fusion technique because the video and audio features are fused earlier in the layers of the multimodal transformer 302 and later before the decoder 310. As Figure 4 shown, there can be multiple (such as two) cross-modal transformer encoder layers 304 that receive and combine a first subset of video and audio features, and multiple (such as two) fusion encoder layers 306 that process the output of the cross-modal transformer encoder layer 304. This type of approach can be said to represent a "mid-late" fusion technique because the video and audio features are fused closer to the middle layers of the multimodal transformer 302 and later before the decoder 310. As Figure 5As shown, there can be multiple (such as three) cross-modal transformer encoder layers 304 that receive and combine a first subset of video and audio features, and a single fusion encoder layer 306 that processes the output of the cross-modal transformer encoder layers 304. This type of approach can be said to represent a "late-late" fusion technique because the video and audio features are fused relatively late in the layers of the multi-modal transformer 302 and relatively late before the decoder 310.
[0063] Although Figures 3 to 5 illustrates an example of a trained machine learning model 212 for use in the Figure 2 architecture 200, various changes can be made to Figures 3 to 5 it. For example, Figures 3 to 5 the various components and functions in each of Figures 3 to 5 can be combined, further subdivided, replicated, or rearranged according to specific needs. Additionally, if desired or needed, one or more additional components and functions can be included in each of
[0064] An architecture that supports the ability to effectively recognize emotions in video content can be used to support any number of possible applications. The following represent example use cases where the ability to effectively recognize emotions in video content can be used. These use cases can include social platform engagement, "living" art generation, advertising targeting, metaverse emotion understanding, and content recommendation. However, note that these use cases are only examples, and the architecture 200 can be used for any other suitable purpose in any other suitable way.
[0065] Regarding social platform engagement, an application can use the ability to recognize emotions in video content to create an "emotion-based" sentence generator, which can be used to interact with viewers of the video content. For example, the sentence generator can use the recognition of emotions in the video content presented to the viewer to generate coherent sentences and engage in a conversation with the viewer. For example, during an exciting sports event, the architecture 200 can be used to sense "tension", and the sentence generator can produce engaging sentences such as "This is a thrilling moment!"
[0066] Regarding "living" art generation, an application can be to generate a "living" artwork that changes based on the emotions of the people viewing the artwork. For example, one or more cameras can be used to capture a video sequence of audience members viewing the artwork, and the architecture 200 can be used to recognize the emotions of the audience members. The content of the artwork can be changed based on the detected emotions, such as by changing the artwork based on the emotions of the viewer in front of or closest to the artwork.
[0067] Regarding advertisement targeting, an application can show relevant advertisements to appropriate users and based on the viewing history of the users' video content. For example, an advertiser can target the interests of specific users based on a taste map, and the emotional categories of the video content watched or preferred by the users can be included in the taste map. This may help increase the precision of advertisement targeting. Thus, the architecture 200 can be used to identify the emotions in the video content watched by a specific user, and the preferences of the user in terms of the emotional content in the watched videos can be included in the taste map. The architecture 200 can also be used to identify the emotions in a specific advertisement. This allows for a fine-grained category matching between the advertisement content and the user history, which allows emotions to be a useful component when determining an advertisement relevance score and providing a specific advertisement to a specific user.
[0068] Regarding metaverse emotion understanding, the ability to understand the emotions displayed by real human avatars in the metaverse can be used in various applications. For example, the architecture 200 can be used to identify the emotions in the video content that a user is watching and project those emotions onto the face of the human avatar (such as by projecting the emotions of the content onto the facial key points of the avatar), thereby making the avatar behave more human-like. The architecture 200 can be used to detect the emotions of human-like avatars in the metaverse, and these detected emotions can be used to model the sentence generation of metaverse conversations. The architecture 200 can be used to moderate extreme emotions in the metaverse of human-like avatars, such as when restricting or otherwise modulating the actions of a human-like avatar in response to detecting "anger" or other specific emotions associated with that avatar (in order to avoid extreme behavior).
[0069] Regarding content recommendation, an application can recommend video content to a viewer based on, for example, the current video that the viewer is watching and the viewer's viewing history or user profile. For example, the architecture 200 can be used to create an emotion profile of the viewer based on the video content that the viewer has watched over time. Some videos with high emotions that the viewer does not like can be removed from the recommendations for those viewers. For example, if someone tends to only like positive and happy shows or movies, the recommendation system can avoid recommending videos with "negative" as the main emotion to that viewer.
[0070] Figure 6 An example method 600 for multimodal understanding of emotions in video content according to the present disclosure is shown. For ease of explanation, Figure 6 the method 600 shown is described as being executed by Figure 1 the electronic device 101 in the network configuration 100. However, Figure 6 the method 600 shown in
[0071] As Figure 6 shown, a video sequence is obtained at step 602. This can include, for example, the processor 120 (or at least one processor) of the electronic device 101 obtaining the video sequence 202 from any suitable source. The video sequence 202 includes a plurality of video frames 204 and audio data 206, such as data defining an audio waveform. At step 604, face detection is performed and video features are extracted using the video frames. This can include, for example, the processor 120 of the electronic device 101 performing face detection and video feature extraction functions 208 to identify the locations in the video frames 204 that contain a person's face, and extracting facial features or other video features associated with the video frames 204 based on the face detection results. For example, the video frames 204 can be divided into sets of video frames 204, such as six-second sets of video frames 204 or other sets, and each set of video frames 204 can be processed by the face detection and video feature extraction functions 208. Additionally, for example, face detection can be performed using MTCNN, and feature extraction can be performed using SCN.
[0072] At step 606, a first subset of audio features is extracted from the audio data, and at step 608, a second subset of audio features is extracted from the audio data. This can include, for example, the processor 120 of the electronic device 101 performing at least one audio feature extraction function 210, 210a - 210b to extract audio features associated with the audio data 206. For example, the first subset of audio features can be based on the audio waveform defined by the audio data 206, and these features can be identified using PyAudio analysis or other suitable signal analysis. Additionally, for example, a pre-trained audio model (such as the PSLA model) can be used to determine the second subset of audio features.
[0073] At step 610, the extracted video features and the subsets of the extracted audio features are provided to a trained machine learning model 212. This can include, for example, the processor 120 of the electronic device 101 providing the video features and the first subset of audio features to the multimodal transformer 302, and providing the second subset of audio features to the fusion function 308. At step 612, the trained machine learning model 212 is used to perform the fusion of the video features and the first subset of audio features. This can include, for example, the processor 120 of the electronic device 101 connecting or otherwise combining the video features and the first subset of audio features, and processing the fused features using one or more cross-modal transformer encoder layers 304 and one or more fusion encoder layers 306. The final output of the multimodal transformer 302 can represent the processed features based on the video features and the first subset of audio features.
[0074] In step 614, a fusion of the processed features and the second subset of audio features is performed. This can include, for example, the processor 120 of the electronic device 101 executing a fusion function 308 to concatenate or otherwise combine the processed features and the second subset of audio features. The output of the fusion function 308 can represent an encoded output based on the two subsets of video and audio features. In step 616, the encoded output is decoded to generate one or more predicted emotions 214 of at least one person included in the video sequence. The one or more predicted emotions can be referred to as an estimate of at least one emotion of at least one person, since the at least one emotion is estimated from the video frames 204 and audio data 206 included in the video sequence 202. This can include, for example, the processor 120 of the electronic device 101 processing the encoded output using a decoder 310 (such as an MLP decoder). The decoder 310 can use the fusion output from the fusion function 308 to produce one or more predicted emotions 214 associated with the video sequence 202. The one or more predicted emotions 214 can be referred to as at least one emotion, since the at least one emotion is predicted from the video frames 204 and audio data 206 included in the video sequence 202.
[0075] In step 618, the one or more predicted emotions 214 can be stored, output, or used in some way. The exact use of the one or more predicted emotions 214 can vary based on the situation. Example applications (such as social platform engagement, "living" art generation, advertising targeting, metaverse emotion understanding, and content recommendation) have been described above, but the one or more predicted emotions 214 can be used for any other suitable purpose in any other suitable way.
[0076] Although Figure 6 illustrates one example of a method 600 for multimodal understanding of emotions in video content, various changes can be made to Figure 6 For example, although shown as a series of steps, Figure 6 the various steps in
[0077] In the present disclosure, the method 600 includes obtaining 602 a video sequence 202 including a plurality of video frames 204 and audio data 206, extracting 604 video features associated with at least one face in the plurality of video frames 204 and audio features associated with the audio data 206, and processing the video features and audio features using a trained machine learning model 212 that performs multi-layer fusion of different subsets of the video features and audio features to identify at least one emotion 214 expressed by at least one person in the video sequence 202.
[0078] In the present disclosure, the extraction 604 of video features and audio features includes extracting 604 video features, the extraction of video features including (i) dividing a plurality of video frames 204 into a plurality of sets of video frames, (ii) performing face detection in the plurality of sets of video frames, and (iii) processing the plurality of sets of video frames based on the results of the face detection to identify video features associated with at least one face, and extracting 606, 608 audio features. The extraction of audio features includes (i) processing audio data 206 to identify a first subset of audio features associated with the waveform of the audio data 206, and (ii) using a pre-trained audio model to process the audio data 206 to identify a second subset of audio features.
[0079] In the present disclosure, processing the plurality of sets of video frames to identify video features includes processing the plurality of sets of video frames using a self-healing network (SCN), and using a pre-trained audio model to process the audio data 206 includes processing the audio data 206 using a pre-trained, sampled, labeled, and aggregated (PSLA) model.
[0080] In the present disclosure, the trained machine learning model 212 includes at least one cross-modal transformer encoder layer 304 configured to receive and fuse the first subset of video features and audio features and generate multimodal features, at least one fusion encoder layer 306 configured to combine the multimodal features, and a multi-layer perceptron (MLP) decoder layer 310 configured to decode the output of at least one fusion encoder layer 306 fused with the second subset of audio features.
[0081] In the present disclosure, the trained machine learning model 212 includes a multimodal transformer 302, the multimodal transformer 302 including one or more cross-modal transformer encoder layers 304 and one or more fusion encoder layers 306, the output of the multimodal transformer 302 being fused with the second subset of audio features, and the first subset of video features and audio features being fused by one of the following: an earlier layer in the multimodal transformer 302 for supporting early-late fusion of video features and audio features, a later layer in the multimodal transformer 302 for supporting late-late fusion of video features and audio features, or a layer between the earlier layer and the later layer in the multimodal transformer 302 for supporting mid-late fusion of video features and audio features.
[0082] In the present disclosure, the multi-layer fusion of video features and audio features includes a first fusion of the first subset of video features and audio features, and a second fusion of the processed features and the second subset of audio features, the processed features being based on the first fusion.
[0083] In the present disclosure, a trained machine learning model 212 is trained to recognize a plurality of emotions arranged in a hierarchy, and two root categories of the hierarchy include positive emotions and negative emotions.
[0084] In the present disclosure, an electronic device 101 includes at least one memory 130 and at least one processor 120. The at least one memory 130 is configured to store a video sequence 202 including a plurality of video frames 204 and audio data 206. The at least one processor 120 is configured to extract video features associated with at least one face in the plurality of video frames and audio features associated with the audio data 206, and process the video features and the audio features using the trained machine learning model 212. The trained machine learning model 212 is configured to perform multi-layer fusion of different subsets of the video features and the audio features to identify at least one emotion 214 expressed by at least one person in the video sequence 202.
[0085] In the present disclosure, in order to extract video features, the at least one processor 120 is configured to (i) divide the plurality of video frames into a plurality of sets of video frames, (ii) perform face detection in the plurality of sets of video frames, and (iii) process the plurality of sets of video frames based on the results of the face detection to identify video features associated with at least one face. And in order to extract audio features, the at least one processor 120 is configured to (i) process the audio data 206 to identify a first subset of audio features associated with the waveform of the audio data 206, and (ii) use a pre-trained audio model to process the audio data 206 to identify a second subset of audio features.
[0086] In the present disclosure, in order to process the plurality of sets of video frames, the at least one processor 120 is configured to use a self-healing network (SCN), and in order to process the audio data using a pre-trained audio model, the at least one processor 120 is configured to use a pre-trained, sampled, labeled, and aggregated (PSLA) model to process the audio data.
[0087] In the present disclosure, the trained machine learning model 212 includes at least one cross-modal transformer encoder layer 304, at least one fusion encoder layer 306, and a multi-layer perceptron (MLP) decoder layer 310. The at least one cross-modal transformer encoder layer 304 is configured to receive and fuse a first subset of the video features and the audio features and generate multi-modal features. The at least one fusion encoder layer 306 is configured to combine the multi-modal features. And the multi-layer perceptron (MLP) decoder layer 310 is configured to decode the output of at least one fusion encoder layer 306 fused with the second subset of the audio features.
[0088] In the present disclosure, the trained machine learning model 212 includes a multimodal transformer 302, the multimodal transformer 302 includes one or more cross-modal transformer encoder layers 304 and one or more fusion encoder layers 306, and the trained machine learning model 212 is further configured to fuse the output of the multimodal transformer 302 with a second subset of audio features, and fuse the first subset of video features and audio features in one of the following: an earlier layer in the multimodal transformer 302 for supporting early-late fusion of video features and audio features, a later layer in the multimodal transformer 302 for supporting late-late fusion of video features and audio features, or a layer between the earlier layer and the later layer in the multimodal transformer 302 for supporting mid-late fusion of video features and audio features.
[0089] In the present disclosure, the multi-layer fusion of video features and audio features includes a first fusion of a first subset of video features and audio features, and a second fusion of the processed features and a second subset of audio features, the processed features being based on the first fusion.
[0090] In the present disclosure, the trained machine learning model 212 is trained to identify a plurality of emotions arranged in a hierarchy, the two root categories of the hierarchy including positive emotions and negative emotions.
[0091] In the present disclosure, a computer-readable medium containing instructions that, when executed, cause at least one processor 120 to implement a method of obtaining 602 a video sequence 202 including a plurality of video frames 204 and audio data 206, the method including extracting 604 video features associated with at least one face in the plurality of video frames 204 and audio features associated with the audio data 206, and processing the video features and audio features using the trained machine learning model 212, the trained machine learning model 212 performing multi-layer fusion of different subsets of the video features and audio features to identify at least one emotion 214 expressed by at least one person in the video sequence 202.
[0092] In the present disclosure, the instructions that, when executed, cause at least one processor 120 to extract video features include instructions that, when executed, cause at least one processor 120 to (i) divide a plurality of video frames 204 into a plurality of sets of video frames, (ii) perform face detection on the plurality of sets of video frames, and (iii) process the plurality of sets of video frames based on the results of the face detection to identify video features associated with at least one face, and the instructions that, when executed, cause at least one processor 120 to extract audio features include instructions that, when executed, perform the following operations: cause at least one processor 120 to (i) process audio data 206 to identify a first subset of audio features associated with the waveform of the audio data 206, and (ii) use a pre-trained audio model to process the audio data 206 to identify a second subset of audio features.
[0093] In the present disclosure, the instructions that, when executed, cause at least one processor 120 to process a plurality of sets of video frames include instructions that, when executed, cause at least one processor 120 to use a self-healing network (SCN), and the instructions that, when executed, cause at least one processor 120 to process audio data 206 using a pre-trained audio model include instructions that, when executed, cause at least one processor 120 to process audio data 206 using a pre-trained, sampled, labeled, and aggregated (PSLA) model.
[0094] Although the present disclosure has been described with reference to various example embodiments, various changes and modifications can be suggested to those skilled in the art. The present disclosure is intended to cover such changes and modifications that fall within the scope of the appended claims.
Claims
1. A method 600, comprising: Obtaining 602 a video sequence 202 including a plurality of video frames 204 and audio data 206; Extracting 604 video features associated with at least one face in the plurality of video frames 204 and audio features associated with the audio data 206; And Using a trained machine learning model 212 to process the video features and audio features, the trained machine learning model 212 performing multi-layer fusion of different subsets of the video features and audio features to identify at least one emotion 214 expressed by at least one person in the video sequence 202.
2. The method according to claim 1, wherein, Extracting 604 the video features and the audio features includes: Extracting 604 video features, including (i) dividing the plurality of video frames 204 into a plurality of sets of video frames, (ii) performing face detection in the plurality of sets of video frames, and (iii) processing the plurality of sets of video frames based on the results of the face detection to identify video features associated with at least one face; and Extracting 606, 608 audio features, including (i) processing the audio data 206 to identify a first subset of audio features associated with the waveform of the audio data 206, and (ii) using a pre-trained audio model to process the audio data 206 to identify a second subset of audio features.
3. The method according to claim 2, wherein: Processing the plurality of sets of video frames to identify the video features includes using a self-healing network (SCN) to process the plurality of sets of video frames; and Using a pre-trained audio model to process the audio data 206 includes using a pre-trained, sampled, labeled, and aggregated (PSLA) model to process the audio data 206.
4. The method according to any one of claims 1 to 3, wherein, The trained machine learning model 212 includes: At least one cross-modal transformer encoder layer 304 configured to receive and fuse a first subset of video features and audio features and generate multi-modal features; At least one fusion encoder layer 306 configured to combine the multi-modal features; and A multi-layer perceptron (MLP) decoder 310 layer configured to decode the output of at least one fusion encoder layer 306 fused with a second subset of audio features.
5. The method according to any one of claims 1 to 3, wherein: The trained machine learning model 212 includes a multi-modal transformer 302, the multi-modal transformer 302 including one or more cross-modal transformer encoder layers 304 and one or more fusion encoder layers 306; The output of the multi-modal transformer 302 is fused with a second subset of audio features; and The first subset of the video features and the audio features is fused by one of the following: An earlier layer in the multi-modal transformer 302 for supporting early-late fusion of video features and audio features; A later layer in the multi-modal transformer 302 for supporting late-late fusion of video features and audio features; or A layer between an earlier layer and a later layer in the multi-modal transformer 302 for supporting mid-late fusion of video features and audio features.
6. The method according to any one of claims 1 to 5, wherein The multi-layer fusion of the video features and the audio features includes: a first fusion of the video features and a first subset of the audio features; and a second fusion of the processed features and a second subset of the audio features, the processed features being based on the first fusion.
7. The method according to any one of claims 1 to 6, wherein, The trained machine learning model 212 is trained to identify a plurality of emotions arranged in a hierarchy, two root categories of the hierarchy including positive emotions and negative emotions.
8. An electronic device 101, comprising:[[]] at least one memory 130 configured to store a video sequence 202 including a plurality of video frames 204 and audio data 206; and at least one processor 120 configured to:[[]] extract video features associated with at least one face in a plurality of video frames 204 and audio features associated with audio data 206; and use a trained machine learning model 212 to process the video features and the audio features, the trained machine learning model 212 being configured to perform multi-layer fusion of different subsets of the video features and the audio features to identify at least one emotion 214 expressed by at least one person in the video sequence 202.
9. The electronic device 101 according to claim 8, wherein:[[]] To extract video features, the at least one processor 120 is further configured to (i) divide the plurality of video frames 204 into a plurality of sets of video frames, (ii) perform face detection in the plurality of sets of video frames, and (iii) process the plurality of sets of video frames based on the results of the face detection to identify video features associated with at least one face; and To extract audio features, the at least one processor 120 is further configured to (i) process the audio data 206 to identify a first subset of audio features associated with the waveform of the audio data 206, and (ii) use a pre-trained audio model to process the audio data 206 to identify a second subset of audio features.
10. The electronic device 101 according to claim 9, wherein:[[]] To process the plurality of sets of video frames, the at least one processor 120 is further configured to use a self-healing network (SCN); and To use a pre-trained audio model to process the audio data, the at least one processor 120 is further configured to use a pre-trained, sampled, labeled, and aggregated (PSLA) model to process the audio data.
11. The electronic device 101 according to any one of claims 8 to 10, wherein, The trained machine learning model 212 includes:[[]] at least one cross-modal transformer encoder layer 304 configured to receive and fuse a first subset of video features and audio features and generate multi-modal features; at least one fusion encoder layer 306 configured to combine the multi-modal features; and a multi-layer perceptron (MLP) decoder layer 310 configured to decode the output of at least one fusion encoder layer 306 fused with a second subset of audio features.
12. The electronic device 101 according to any one of claims 8 to 11, wherein:[[]] The trained machine learning model 212 includes a multi-modal transformer 302, the multi-modal transformer 302 including one or more cross-modal transformer encoder layers 304 and one or more fusion encoder layers 306; The trained machine learning model 212 is also configured to fuse the output of the multimodal transformer 302 with a second subset of audio features; and The trained machine learning model 212 is also configured to fuse the first subset of video features and audio features in one of the following: An earlier layer in the multimodal transformer 302 for supporting early-late fusion of video features and audio features; A later layer in the multimodal transformer 302 for supporting late-late fusion of video features and audio features; or A layer between the earlier layer and the later layer in the multimodal transformer 302 for supporting mid-late fusion of video features and audio features.
13. The electronic device 101 according to any one of claims 8 to 12, wherein, The multi-layer fusion of the video features and the audio features includes: A first fusion of the first subset of the video features and the audio features; and A second fusion of the processed features and the second subset of the audio features, the processed features being based on the first fusion.
14. The electronic device 101 according to any one of claims 8 to 12, wherein, The trained machine learning model 212 is trained to identify a plurality of emotions arranged in a hierarchy, two root categories of the hierarchy including positive emotions and negative emotions.
15. A computer-readable medium containing instructions that, when executed, cause at least one processor 120 to implement the method according to any one of claims 1 to 7.