Face emotion detection method and system, terminal and storage medium

By constructing a category feature matrix and feature fusion technology, face emotions are directly detected from the video stream, solving the problem of high computing resources consumption and achieving efficient and accurate face emotions detection.

CN120340082APending Publication Date: 2025-07-18上海蜜度蜜巢智能科技有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510243221.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-03
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

In the prior art, the face emotion detection method consumes a lot of computing resources in the video stream, resulting in high computing capabilities requirements for hardware equipment, affecting the real-time and response speed of the system.

Method used

A category feature matrix is constructed, and visual feature extraction and position feature extraction of multi-frame face images in the video stream, combined with a large language model and a Transformer encoder for feature fusion, and directly detect face emotion categories from the video stream.

Benefits of technology

The detection process can be simplified without face detection, improved detection efficiency, saved computing resources, and accurately captured the subtle fluctuations and dynamic characteristics of facial emotions, realizing the detection of open-categorized emotions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120340082A_ABST
    Figure CN120340082A_ABST
Patent Text Reader

Abstract

The invention provides a face emotion detection method and system, a terminal and a storage medium. The method comprises the following steps: constructing a category feature matrix including at least one face emotion category; performing visual feature extraction on multiple frames of face images in the video stream to obtain visual feature vectors; acquiring a position feature vector of the to-be-detected frame; fusing the category feature matrix, the visual feature vector and the position feature vector to obtain a fused feature matrix; and based on the fusion feature matrix, determining a face emotion category in the to-be-detected frame. According to the method, the face emotion can be detected without face detection, the detection process is simplified, the detection efficiency is improved, and computing resources are saved; the time sequence information of the continuous frames of the video stream can be fully utilized, and the fine fluctuation and dynamic characteristics of the face emotion can be accurately captured; the method is not limited to limitation of limited emotion categories, and detection of open category emotions is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the technical field of image detection, and relates to a face emotion detection method, system, terminal and storage medium. Background Art

[0002] In the field of image detection, face emotion detection has gradually become an important research topic. By detecting the emotions in facial expression images, the current mental state of the user can be obtained. Based on this, face emotion detection shows great potential application value in the fields of psychology, intelligent robots, intelligent monitoring, virtual reality and synthetic animation.

[0003] Traditional face emotion detection methods need to perform face detection on each frame of the video stream to accurately locate and extract the face region in the facial expression image. Subsequently, the pre-trained emotion classification model is used to classify the emotions of the intercepted face region, so as to determine the emotion category expressed by the face in this frame of image.

[0004] However, with the continuous increase in the number of video frames, this repeated operation of face detection will cause a significant increase in the consumption of computing resources. This not only places higher requirements on the computing power of hardware devices, but may also cause problems of low efficiency in actual applications. Especially when processing long videos or real-time video streams, the bottleneck of computing resources may seriously affect the real-time performance and response speed of the system. Summary of the Invention

[0005] The purpose of this application is to provide a face emotion detection method, system, terminal and storage medium, which are used to solve the technical problem of large consumption of computing resources in the prior art.

[0006] In a first aspect, this application provides a face emotion detection method, including: constructing a category feature matrix including at least one face emotion category; extracting visual features from multiple frames of face images in a video stream to obtain visual feature vectors; obtaining a position feature vector of a frame to be detected; the frame to be detected is any one of the multiple frames of face images; the position feature vector is used to represent the frame sequence number of the frame to be detected in the video stream; fusing the category feature matrix, the visual feature vector and the position feature vector to obtain a fused feature matrix; based on the fused feature matrix, determining the face emotion category in the frame to be detected.

[0007] In one implementation manner of the first aspect, constructing a category feature matrix including at least one face emotion category includes: setting an emotion category set; the emotion category set includes at least one of the face emotion categories; extracting features from each of the face emotion categories to obtain feature vectors of the face emotion categories; combining the feature vectors of all the face emotion categories to obtain the category feature matrix.

[0008] In one implementation of the first aspect, visual feature extraction is performed on multiple frames of face images in a video stream, and the obtained visual feature vector includes: splitting the video stream based on a preset frame rate to obtain multiple frames of independent face images; performing visual feature extraction on each frame of the face image to obtain the feature vector of the face image; performing frame-level feature fusion on the feature vectors of each frame of the face image to obtain a visual feature matrix; performing global average pooling on the visual feature matrix to obtain the visual feature vector.

[0009] In one implementation of the first aspect, a large language model is used to perform frame-level feature fusion on the feature vectors of each frame of the face image to obtain a visual feature matrix.

[0010] In one implementation of the first aspect, fusing the category feature matrix, the visual feature vector, and the position feature vector to obtain a fusion feature matrix includes: concatenating the category feature matrix, the visual feature vector, and the position feature vector in a specific order to obtain a concatenated matrix; performing modality fusion processing on the concatenated matrix to obtain the fusion feature matrix.

[0011] In one implementation of the first aspect, the specific order is the visual feature vector, the position feature vector, and the category feature matrix.

[0012] In one implementation of the first aspect, based on the fusion feature matrix, determining the face emotion category in the frame to be detected includes: obtaining the vector at the same position as the position feature vector from the fusion feature matrix as the frame feature vector; obtaining the matrix at the same position as the category feature matrix from the fusion feature matrix as the emotion feature matrix; the emotion feature matrix includes multiple emotion feature vectors, and each emotion feature vector corresponds to a face emotion category; calculating the similarity between the frame feature vector and each emotion feature vector in the emotion feature matrix to obtain similarity scores; selecting the face emotion category corresponding to the maximum similarity score as the face emotion category in the frame to be detected.

[0013] Second aspect, the present application provides a face emotion detection system, including: a matrix construction module for constructing a category feature matrix including at least one face emotion category; a feature extraction module for performing visual feature extraction on multiple frames of face images in a video stream to obtain visual feature vectors; a vector acquisition module for acquiring a position feature vector of a frame to be detected, where the frame to be detected is any one of the multiple frames of face images, and the position feature vector is used to represent the frame sequence number of the frame to be detected in the video stream; a feature fusion module for fusing the category feature matrix, the visual feature vector, and the position feature vector to obtain a fused feature matrix;

[0014] a category determination module for determining the face emotion category in the frame to be detected based on the fused feature matrix.

[0015] Third aspect, the present application provides a face emotion detection terminal, including: a processor and a memory; the memory is used for storing a computer program; the processor is used for executing the computer program stored in the memory so that the face emotion detection terminal executes the method described in any one of the above.

[0016] Fourth aspect, the present application provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the method described in any one of the above is implemented.

[0017] As described above, the face emotion detection method, system, terminal, and storage medium of the present application have the following beneficial effects:

[0018] (1) It is possible to detect face emotions without performing face detection, simplifies the detection process, improves the detection efficiency, and saves computing resources;

[0019] (2) It can make full use of the temporal information of consecutive frames of the video stream to accurately capture the subtle fluctuations and dynamic features of face emotions;

[0020] (3) It is not limited to the limitations of a finite number of emotion categories and realizes the detection of open-category emotions. Description of the Drawings

[0021] Figure 1 It shows a structural schematic diagram of the mobile terminal in an embodiment of the present application.

[0022] Figure 2 It shows a flowchart of the face emotion detection method in an embodiment of the present application.

[0023] Figure 3 It shows a flowchart of the face emotion detection method in another embodiment of the present application.

[0024] Figure 4 It shows a flowchart of the face emotion detection method described in this application in yet another embodiment.

[0025] Figure 5 It shows a flowchart of the face emotion detection method described in this application in yet another embodiment.

[0026] Figure 6 It shows a schematic structural diagram of the face emotion detection system described in this application in an embodiment.

[0027] Figure 7 It shows an architecture diagram of the face emotion detection system described in this application in an embodiment.

[0028] Figure 8 It shows a schematic structural diagram of the face emotion detection terminal described in this application in an embodiment. Detailed implementation manners

[0029] The following uses specific specific examples to illustrate the implementation manners of this application. Those skilled in the art can easily understand other advantages and effects of this application from the content disclosed in this specification. This application can also be implemented or applied through other different specific implementation manners. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of this application. It should be noted that, without conflict, the following embodiments and the features in the embodiments can be combined with each other.

[0030] It should be noted that the diagrams provided in the following embodiments only illustrate the basic concept of this application in a schematic manner. Therefore, only the components related to this application are shown in the diagrams, rather than being drawn according to the number, shape, and size of the components in actual implementation. The type, quantity, and proportion of each component in actual implementation can be an arbitrary change, and the component layout type may also be more complex.

[0031] In addition, in this application, descriptions such as "first" and "second" are only for descriptive purposes, and cannot be understood as indicating or implying their relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of these features. In addition, the technical solutions between various embodiments can be combined with each other, but it must be based on the fact that those skilled in the art can implement them. When the combination of technical solutions appears to be contradictory or unable to be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the protection scope required by this application.

[0032] The following embodiments of the present application provide a face emotion detection method, system, terminal and storage medium. The present application can detect face emotions without performing face detection, simplifies the detection process, improves the detection efficiency, and saves computing resources; can make full use of the temporal information of consecutive frames of the video stream to accurately capture the subtle fluctuations and dynamic characteristics of face emotions; is not limited to the limitations of a finite number of emotion categories, and realizes the detection of open-category emotions.

[0033] Please refer to Figure 1 As shown, the retrieval feature extraction method provided by the embodiments of the present application can run on similar devices such as mobile terminals and computer terminals. Taking running on the mobile terminal as an example, Figure 1 is the hardware structure block diagram of the mobile terminal, as Figure 1 shown, the mobile terminal may include: a processor and a memory. The processor may be a central processing unit, and the memory is used to store data. Figure 1 The mobile terminal in

[0034] is only for illustration and does not limit the specific structure of the mobile terminal.

[0035] Optionally, the memory may be used to store computer programs, such as software programs and modules of application software. The memory may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memories. In some instances, the memory may further include memories remotely provided with respect to the processor, and these remote memories may be connected to the mobile terminal through a network. Examples of the above networks include, but are not limited to, the Internet, enterprise intranets, local area networks, mobile communication networks, and combinations thereof.

[0036] Optionally, the communication transmission device may be used to receive or send data via a network, and the network may include a wireless network provided by a communication provider of the mobile terminal. The communication transmission device may include a NIC (Network Interface Controller), which may be connected to other network devices through a base station and thus communicate with the Internet.

[0037] Next, the technical solutions in the embodiments of the present application will be described in detail with reference to the accompanying drawings in the embodiments of the present application.

[0038] Please refer to Figure 2 which shows the flowchart of the face emotion detection method described in the present application in an embodiment. As Figure 2 shown, the embodiments of the present application provide a face emotion detection method, including the following steps S100 to step S500.

[0039] In step S100, a category feature matrix including at least one facial emotion category is constructed.

[0040] Please refer to Figure 3 , which shows a flowchart of the facial emotion detection method described in this application in another embodiment.

[0041] As Figure 3 shown, step S100 of constructing a category feature matrix including at least one facial emotion category may include the following steps S101 to S103.

[0042] In step S101, an emotion category set is set.

[0043] In an embodiment of the present application, the emotion category set includes at least one of the facial emotion categories.

[0044] Specifically, the emotion category set may include any emotion category that the user wants to detect. Whether it is common emotions such as happiness and sadness, or other more subtle or complex emotional states, they can all be included, with high flexibility and scalability.

[0045] The elements in the emotion category set may be descriptive texts, keywords, or representative sentence fragments related to specific emotion categories, etc.

[0046] The emotion category set of this application supports the detection of open-category emotions, which means that it can capture and identify those emotion types that are not explicitly listed, greatly broadening the application scope of emotion detection. For example, the emotion category set can be represented as an array containing M elements, [happiness, sadness,..., NO], where "NO" is a special automatically added item, and this item represents the following two possible situations:

[0047] (1) Emotions that exist in the video but do not appear in the emotion category set;

[0048] (2) There is no face in the video frame, that is, there is no emotion information.

[0049] Such a design makes the emotion category set more perfect and can handle various complex actual situations.

[0050] Existing emotion classification models usually can only classify a limited number of preset emotion categories. Although common emotion categories such as happiness, sadness, and anger can be recognized well, human emotion expressions are rich, diverse, and complex. There may be many emotional states in actual scenarios that are difficult to fully cover with a limited number of categories. This makes it possible that when the emotion classification model faces some relatively special or subtle emotion expressions, it may not be able to accurately classify them, thus limiting the universality and accuracy of its application.

[0051] In this implementation manner, the set of emotion categories is not limited to a finite number of emotion categories, and the detection of open-category emotions is realized. This feature enables this application to cover human emotion expressions more comprehensively, providing a broader and deeper perspective for emotion research and applications.

[0052] In step S102, feature extraction is performed on each of the face emotion categories to obtain the feature vectors of the face emotion categories.

[0053] In an embodiment of this application, a text encoder is used to perform text feature extraction on each of the emotion categories to obtain the feature vectors corresponding to the face emotion categories.

[0054] Specifically, the text encoder can be a Contrastive Language-Image Pretrained text encoder (CLIP text encoder). As a powerful text encoding tool, the CLIP text encoder has excellent language understanding and feature extraction capabilities. By learning and analyzing a large amount of text data, it can deeply mine the emotional semantic information contained in the text and convert it into a numerical form that can be processed by a computer.

[0055] For example, when processing an array containing M elements, the CLIP text encoder will use its deep learning architecture and pre-trained language model to perform a detailed analysis and encoding on each element, thereby generating a corresponding feature vector for each element. What this application finally obtains are M feature vectors, where the dimension of each feature vector is 1×512, which can comprehensively and accurately capture the key features of the corresponding emotion category.

[0056] In step S103, the feature vectors of all the face emotion categories are combined to obtain the category feature matrix.

[0057] Specifically, when performing the combination operation, it can be arranged and combined in an orderly manner according to certain rules and logic. For example, each 1×512-dimensional feature vector can be used as a row of the matrix and arranged in sequence to obtain a category feature matrix with a dimension of M×512. Each row of the category feature matrix represents the feature information of a face emotion category, and each column can be regarded as a set with the same semantics or feature dimension in different emotion categories.

[0058] In step S200, visual feature extraction is performed on multiple frames of face images in the video stream to obtain visual feature vectors.

[0059] It should be noted that this application uses a video stream instead of a single-frame face image, which is based on the consideration that the video stream can more comprehensively reflect global information. Specifically, a video stream is a dynamic sequence composed of a series of consecutive frames in chronological order, which contains rich spatio-temporal information. Compared with a single-frame image, a video stream can provide information in more dimensions. A single-frame image is just an instantaneous picture in the video, which can only capture the facial expression features at a specific moment and cannot reflect the changing trend and dynamic process of emotions.

[0060] Please refer to Figure 4 , which shows the flowchart of the face emotion detection method described in this application in another embodiment.

[0061] As Figure 4 shown, the steps of extracting visual features from multiple-frame face images in the video stream to obtain visual feature vectors include the following steps S201 to S204.

[0062] In step S201, the video stream is split into frames based on a preset frame rate to obtain multiple independent face images.

[0063] In the embodiment of this application, the preset frame rate is provided by the user, and the user can flexibly set it according to different application scenarios and requirements.

[0064] It should be noted that this application currently only considers the application scenario where a single face is included in the face image. If multiple faces are included in the face image, it is not within the protection scope of this application. This limitation makes the entire emotion detection process more targeted and controllable, and can avoid complex interference factors caused by the simultaneous appearance of multiple faces, enabling the system to focus more on accurately identifying the emotion state of a single face. In this case, each independent face image can provide clear and definite feature information for emotion analysis, which helps to improve the accuracy and reliability of emotion recognition.

[0065] In step S202, visual features are extracted from each frame of the face image to obtain the feature vector of the face image.

[0066] In an embodiment of this application, the ViT (Vision Transformer) algorithm is used to extract visual features from each frame of the face image.

[0067] The ViT algorithm has powerful image processing and feature extraction capabilities. It abandons the local receptive field limitation of the traditional convolutional neural network (CNN) and can comprehensively capture the global dependence and context information in the image through the self-attention mechanism.

[0068] When the ViT algorithm is applied to N frames of face images, the algorithm will conduct a comprehensive analysis of each area of the face image, taking into account the correlation between pixels and the overall semantic information. Finally, N 1×512-dimensional feature vectors are obtained, each of which represents the global features of the corresponding frame of face image as a whole. For example, it contains rich information such as the overall expression, posture, facial contour, and relative position of various organs of the face.

[0069] In this implementation, the visual feature extraction method based on global features will not ignore the overall context due to excessive focus on local details, and can grasp the emotional information conveyed by the face image more comprehensively and accurately.

[0070] In step S203, feature fusion is performed on the feature vectors of the facial images of each frame at the frame level to obtain a visual feature matrix.

[0071] In one embodiment of the present application, a large language model (LLM) is used to perform frame-level feature fusion on the feature vectors of the facial image of each frame to obtain a visual feature matrix.

[0072] With its powerful language understanding and generation capabilities and rich knowledge reserves, the large language model can deeply explore the semantic information behind the data and determine the correlation between adjacent frames when performing fusion operations, which involves some crucial timing information and weight information.

[0073] In actual application scenarios, the changes in facial emotions are not isolated and static, but present continuous and dynamic characteristics. For example, in a video, the emotions of a character may gradually undergo subtle changes as the plot develops, the conversation proceeds, or the environment changes. This kind of time series information is a key element in understanding the evolution of emotions and is essential for accurately capturing the instantaneous changes in emotions. The large language model takes advantage of its own advantages to analyze and process the feature vectors of different time series frames, reasonably assign weights, and thus highlight the features of key frames and better capture the trajectory of emotional changes. In this way, the large language model can fuse N 1×512-dimensional feature vectors into an organic whole to obtain a visual feature matrix with a dimension of N×512.

[0074] In this implementation, the timing information between consecutive frames of the video stream can be fully utilized to accurately capture the subtle fluctuations and dynamic characteristics of emotions, thereby providing more accurate and comprehensive data support for facial emotion detection.

[0075] In step S204, global average pooling is performed on the visual feature matrix to obtain the visual feature vector.

[0076] The global average pooling process calculates the average of all feature vectors in the visual feature matrix. This step can effectively reduce the data dimension while retaining the most important feature information. During this process, the feature vectors corresponding to each frame of the face image are included in the calculation scope, ensuring that the finally obtained feature vector can comprehensively reflect the content of the entire video.

[0077] For example, the steps of global average pooling include accumulating the N×512 eigenvalues in the visual feature matrix and then dividing by N to obtain the average value of each feature dimension. After such processing, the visual feature matrix with multiple feature vectors is compressed into a single feature vector with a dimension of 1×512.

[0078] The visual feature vector with a feature dimension of 1×512 obtained after global average pooling is essentially a highly condensed information carrier. It does not simply represent the face image information of a certain frame or a local area, but comprehensively summarizes the face emotion-related information in the entire video as a whole.

[0079] In step S300, obtain the position feature vector of the frame to be detected.

[0080] In the embodiment of the present application, the frame to be detected is any one of the multiple frames of face images. This flexibility enables the system to perform accurate emotion analysis on specific frames in the video stream, especially suitable for the study of key frames or representative frames.

[0081] The position feature vector is used to represent the frame number of the frame to be detected in the video stream. The frame number reflects the position of the frame on the entire video timeline.

[0082] In an embodiment of the present application, a contrastive language-image pre-trained image encoder (CLIP image encoder) is used to obtain the position feature vector of the frame to be detected.

[0083] As an advanced image encoder, CLIP image encoder has powerful image processing and feature learning capabilities. Through the analysis of images, it can effectively extract various features related to the image content, including the position features that can represent the frame number information. During this process, CLIP image encoder will deeply mine the image data of the frame to be detected and convert the implicit information related to the frame number into an explicit feature vector form. In the embodiment of the present application, the finally obtained position feature vector has a dimension of 1×512.

[0084] In step S400, fuse the class feature matrix, the visual feature vector, and the position feature vector to obtain a fused feature matrix.

[0085] In an embodiment of the present application, fusing the class feature matrix, the visual feature vector, and the position feature vector to obtain a fused feature matrix includes the following steps S401 to S402.

[0086] In step S401, the class feature matrix, the visual feature vector, and the position feature vector are concatenated in a specific order to obtain a concatenated matrix.

[0087] In an embodiment of the present application, the specific order is the visual feature vector, the position feature vector, and the class feature matrix.

[0088] In the embodiment of the present application, the concatenation order reflects a causal relationship and thus has uniqueness and shall not be randomly swapped.

[0089] Combined with the embodiments in steps S100 to S300, the video feature vector (1×512 dimensions) represents the global information of the entire video and belongs to a relatively broad feature representation method; the position feature vector (1×512 dimensions) focuses on the position of a specific frame in the video stream; the class feature matrix (M×512 dimensions) describes different classes in detail. After concatenation, a concatenated matrix with a dimension of (M + 2)×512 dimensions can be obtained. The concatenated matrix combines the global information of the video, the position information of a specific frame, and the detailed class information, thereby forming a complete feature representation.

[0090] In step S402, the concatenated matrix is subjected to modality fusion processing to obtain the fused feature matrix.

[0091] In an embodiment of the present application, a Transformer encoder is used to perform modality fusion processing on the concatenated matrix to obtain the fused feature matrix.

[0092] The video feature vector (1×512 dimensions), the position feature vector (1×512 dimensions), and the category feature matrix (M×512 dimensions) belong to three types of modal information. The Transformer encoder can effectively fuse these three modalities. Its working principle is based on the self-attention mechanism. By calculating the attention scores between different modal feature vectors, it determines the degree of mutual correlation between them, thus organically combining the information of different modalities. For example, when processing the video feature vector and the position feature vector, the Transformer encoder can, according to the relationship between the two, better understand the connection between the video content at a specific position and the overall emotional expression; at the same time, when fusing with the category feature matrix, it can, based on the existing emotional category information, more accurately interpret the emotional cues contained in the video features and position features.

[0093] After the modal fusion processing, the obtained fused feature matrix has the same dimension as the concatenation matrix, that is, it still maintains the size of (M + 2)×512, ensuring the integrity and continuity of the information.

[0094] In this implementation, the Transformer encoder, by integrating the global information of the video stream, the information of the specified frame, and the information of different categories, gives full play to the advantages of various information, and through its powerful learning ability and processing ability, provides strong support for improving the accuracy of emotion recognition in the subsequent stage.

[0095] In step S500, based on the fused feature matrix, determine the facial emotion category in the frame to be detected.

[0096] Please refer to Figure 5 , which shows the flowchart of the facial emotion detection method described in this application in another embodiment. As Figure 5 shown, based on the fused feature matrix, determining the facial emotion category in the frame to be detected includes the following steps S501 to S504.

[0097] In step S501, obtain the vector at the same position as the position feature vector from the fused feature matrix as the frame feature vector.

[0098] Specifically, when the position feature vector is in the second row (or column) of the concatenation matrix, the frame feature vector is correspondingly taken from the second row (or column) of the fused feature matrix.

[0099] In step S502, obtain the matrix at the same position as the category feature matrix from the fused feature matrix as the emotion feature matrix. The emotion feature matrix includes multiple emotion feature vectors, and each emotion feature vector corresponds to a facial emotion category.

[0100] Specifically, when the category feature matrix is located in the 3rd to (M + 2)th rows (or columns) of the splicing matrix, the emotion feature matrix is correspondingly taken from the 3rd to (M + 2)th rows (or columns) of the fusion feature matrix.

[0101] In step S503, calculate the similarity between the frame feature vector and each emotion feature vector in the emotion feature matrix to obtain similarity scores.

[0102] In an embodiment of the present application, calculating the similarity between the frame feature vector and each emotion feature vector in the emotion feature matrix includes performing matrix multiplication on the frame feature vector (1×512 dimensions) and the emotion feature matrix (M×512 dimensions) to obtain M similarity scores. The M similarity scores intuitively represent the similarity degree between the frame feature vector and M categories.

[0103] In step S504, select the face emotion category corresponding to the maximum similarity score as the face emotion category in the frame to be detected.

[0104] In this implementation, it is possible to detect the face emotion without performing face detection, which simplifies the detection process, improves the detection efficiency, and saves computing resources.

[0105] It should be noted that the protection scope of the face emotion detection method described in the embodiments of the present application is not limited to the execution order of the steps listed in this embodiment. Any solution achieved by adding or subtracting steps of the prior art and replacing steps according to the principle of the present application is included in the protection scope of the present application.

[0106] Please refer to Figure 6 , which shows the structural schematic diagram of the face emotion detection system described in the present application in an embodiment.

[0107] As Figure 6 shown, the embodiments of the present application provide a face emotion detection system, including a matrix construction module, a feature extraction module, a vector acquisition module, a feature fusion module, and a category determination module.

[0108] The matrix construction module is used to construct a category feature matrix including at least one face emotion category.

[0109] The feature extraction module is used to extract visual features from multiple frames of face images in the video stream to obtain visual feature vectors.

[0110] The vector acquisition module is used to obtain the position feature vector of the frame to be detected; the frame to be detected is any one of the multiple frames of face images; the position feature vector is used to represent the frame number of the frame to be detected in the video stream.

[0111] The feature fusion module is used to fuse the category feature matrix, the visual feature vector, and the position feature vector to obtain a fused feature matrix.

[0112] The category determination module is used to determine the facial emotion category in the frame to be detected based on the fused feature matrix.

[0113] Please refer to Figure 7 , which shows the architecture diagram of the facial emotion detection system described in this application in an embodiment.

[0114] It should be noted that the structures and principles of the matrix construction module, the feature extraction module, the vector acquisition module, the feature fusion module, and the category determination module described in the embodiments of this application correspond one by one to the steps in the above facial emotion detection method, so they will not be elaborated here.

[0115] The facial emotion detection system provided in the embodiments of this application can implement the facial emotion detection method described in this application. However, the implementation devices of the facial emotion detection method described in this application include, but are not limited to, the structures of the facial emotion detection systems listed in this embodiment. Any structural deformation and replacement of the prior art made according to the principles of this application are included in the protection scope of this application.

[0116] Please refer to Figure 8 , which shows the structural schematic diagram of the facial emotion detection terminal described in this application in an embodiment. As Figure 8 shown, the embodiments of this application provide a facial emotion detection terminal, including: a processor and a memory.

[0117] The memory is used to store a computer program.

[0118] The processor is used to execute the computer program stored in the memory, so that the facial emotion detection terminal executes the method described in any one of the above.

[0119] In an embodiment of the present application, the processor may be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it may also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The memory may be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disk.

[0120] This embodiment further includes one or more of a multimedia component, an input / output (I / O) interface, and a communication component.

[0121] The multimedia component may include a screen and an audio component. The screen may be, for example, a touch screen, and the audio component is used to output and / or input audio signals. For example, the audio component may include a microphone for receiving external audio signals. The received audio signals may be further stored in the memory or sent through the communication component. The audio component further includes at least one speaker for outputting audio signals. The I / O interface provides an interface between the processor and other interface modules, and the other interface modules may be a keyboard, a mouse, buttons, etc. These buttons may be virtual buttons or physical buttons. The communication component is used for wired or wireless communication between this timer and other devices. Wireless communication, such as Wi-Fi, Bluetooth, near field communication (NFC), 2G, 3G, or 4G, or a combination of one or more of them. Accordingly, the communication component may include: a Wi-Fi module, a Bluetooth module, and an NFC module.

[0122] In several embodiments provided by the present application, it should be understood that the disclosed system, device or method can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of modules / units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple modules or units can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of devices or modules or units can be in electrical, mechanical or other forms.

[0123] The modules / units described as separate components may or may not be physically separated. The components shown as modules / units may or may not be physical modules, that is, they can be located in one place or distributed to multiple network units. Some or all of the modules / units can be selected according to actual needs to achieve the purpose of the embodiments of the present application. For example, in each embodiment of the present application, the functional modules / units can be integrated in a processing module, or each module / unit can exist physically alone, or two or more modules / units can be integrated in one module / unit.

[0124] Those of ordinary skill in the art should also further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0125] The embodiments of the present application also provide a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, the method described in any one of the above is implemented. Those of ordinary skill in the art can understand that all or part of the steps in the method of the above embodiments can be completed by instructing a processor through a program. The program can be stored in a computer-readable storage medium. The storage medium is a non-transitory medium, such as random access memory, read-only memory, flash memory, hard disk, solid state drive, magnetic tape, floppy disk, optical disc, and any combination thereof. The above storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or a data center that integrates one or more available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a digital video disc (DVD)), or a semiconductor medium (such as a solid state disk (SSD)).

[0126] The embodiments of the present application can also provide a computer program product, which includes one or more computer instructions. When the computer instructions are loaded and executed on a computing device, the processes or functions described in the embodiments of the present application are fully or partially generated. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website, computer, or data center to another website, computer, or data center in a wired manner (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or a wireless manner (such as infrared, wireless, microwave, etc.).

[0127] When the computer program product is executed by a computer, the computer executes the method described in the foregoing method embodiments. The computer program product can be a software installation package. In the case where the foregoing method is required, the computer program product can be downloaded and executed on the computer.

[0128] The descriptions of the processes or structures corresponding to the above respective drawings each have their own focuses. For parts not detailed in a certain process or structure, reference can be made to the relevant descriptions of other processes or structures.

[0129] The above embodiments are only illustrative of the principles and effects of the present application and are not intended to limit the present application. Any person familiar with this technology can modify or change the above embodiments without departing from the spirit and scope of the present application. Therefore, all equivalent modifications or changes made by those with ordinary knowledge in the technical field without departing from the spirit and technical idea disclosed by the present application should still be covered by the claims of the present application.

Claims

1. A method for detecting human face emotions, characterized in that, Including: Construct a category feature matrix including at least one facial emotion category; Extract visual features from multiple frames of facial images in the video stream to obtain visual feature vectors; Obtain the position feature vector of the frame to be detected; the frame to be detected is any one of the multiple frames of facial images; the position feature vector is used to represent the frame sequence number of the frame to be detected in the video stream; Fuse the category feature matrix, the visual feature vector, and the position feature vector to obtain a fused feature matrix; Based on the fused feature matrix, determine the facial emotion category in the frame to be detected.

2. The method according to claim 1, wherein Constructing a category feature matrix including at least one facial emotion category includes: Set an emotion category set; the emotion category set includes at least one of the facial emotion categories; Extract features for each of the facial emotion categories to obtain the feature vectors of the facial emotion categories; Combine the feature vectors of all the facial emotion categories to obtain the category feature matrix.

3. The method according to claim 1, characterized in that, Extracting visual features from multiple frames of facial images in the video stream to obtain visual feature vectors includes: Split the video stream based on a preset frame rate to obtain multiple independent facial images; Extract visual features from each frame of the facial image to obtain the feature vector of the facial image; Perform frame-level feature fusion on the feature vectors of each frame of the facial image to obtain a visual feature matrix; Perform global average pooling on the visual feature matrix to obtain the visual feature vector.

4. The method according to claim 3, wherein Use a large language model to perform frame-level feature fusion on the feature vectors of each frame of the facial image to obtain a visual feature matrix.

5. The method according to claim 1, wherein Fusing the category feature matrix, the visual feature vector, and the position feature vector to obtain a fused feature matrix includes: Concatenate the category feature matrix, the visual feature vector, and the position feature vector in a specific order to obtain a concatenated matrix; Perform modal fusion processing on the concatenated matrix to obtain the fused feature matrix.

6. The method according to claim 5, characterized in that, The specific order is the visual feature vector, the position feature vector, and the category feature matrix.

7. The method according to claim 1, characterized in that, Based on the fused feature matrix, determining the facial emotion category in the frame to be detected includes: Obtain the vector at the same position as the position feature vector from the fused feature matrix as the frame feature vector; Obtain the matrix at the same position as the category feature matrix from the fused feature matrix as the emotion feature matrix; the emotion feature matrix includes multiple emotion feature vectors, and each emotion feature vector corresponds to a facial emotion category; Calculate the similarity between the frame feature vector and each emotion feature vector in the emotion feature matrix to obtain similarity scores; Select the facial emotion category corresponding to the maximum similarity score as the facial emotion category in the frame to be detected.

8. A face emotion detection system, characterized in that, Including: A matrix construction module for constructing a category feature matrix including at least one facial emotion category; A feature extraction module for extracting visual features from multiple frames of facial images in the video stream to obtain visual feature vectors; A vector acquisition module, configured to acquire a position feature vector of a frame to be detected; the frame to be detected is any one of the multiple face images; the position feature vector is used to represent the frame sequence number of the frame to be detected in the video stream; A feature fusion module, configured to fuse the category feature matrix, the visual feature vector and the position feature vector to obtain a fused feature matrix; A category determination module, configured to determine the face emotion category in the frame to be detected based on the fused feature matrix.

9. A face emotion detection terminal, characterized in that, Comprising: A processor and a memory; The memory is used to store a computer program; The processor is configured to execute the computer program stored in the memory, so that the face emotion detection terminal executes the method according to any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, the method according to any one of claims 1 to 7 is implemented.