Emotion Recognition Method, Device, Storage Medium, and Terminal
By constructing the pixel information matrix of continuous image subsequence of video objects and using the Swin Transformer model, the problem of low-emotion recognition efficiency in videos is solved, and efficient and accurate emotion recognition is achieved.
Patent Information
- Application Number
- CN202111510500.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-10
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2041-12-10
AI Technical Summary
The prior art lacks methods that can accurately and efficiently identify the emotions of objects such as characters in videos.
By obtaining the continuous image subsequence of each object in the pending video, an object pixel information matrix is constructed, and input it to the trained Swin Transformer model for feature extraction, and emotional category determination is performed by combining the scene content information matrix.
It improves the efficiency and accuracy of object emotion recognition in video, reduces the computational complexity, and enhances the effect of emotion recognition.
Smart Images

Figure CN114255420B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer vision, and in particular, to a method and device for emotion recognition, a storage medium, and a terminal. Background Art
[0002] In the task of emotion recognition, visual information is essential. Emotion recognition helps to more effectively improve the efficiency and accuracy of staff in multiple scenarios such as intelligent monitoring and psychological counseling. In the prior art, usually, the emotion of a person in a single image is recognized, and there is a lack of a method that can accurately and efficiently recognize the emotion of an object such as a person in a video.
[0003] Therefore, there is an urgent need for an emotion recognition method that can accurately and efficiently recognize the emotion of an object such as a person in a video. Summary of the Invention
[0004] The technical problem solved by the present invention is to provide an emotion recognition method that can accurately and efficiently recognize the emotion of an object such as a person in a video.
[0005] To solve the above technical problem, an embodiment of the present invention provides an emotion recognition method, the method includes: obtaining a video to be processed, the video to be processed including the images of at least one object; determining a continuous image subsequence of each object in the video to be processed, wherein, the continuous image subsequence of each object includes a plurality of consecutive frames of images and each frame of image includes the image of the object; for the continuous image subsequence of each object, constructing an object pixel information matrix of the object according to the pixel values of the object region in each frame of image, the object region being the region containing the object; inputting the object pixel information matrix of each object into a first feature extraction model to obtain first feature information of the object, wherein, the first feature extraction model is obtained by training a first preset model with first training data, the first training data including: the object pixel information matrix of a sample object in a sample video and the pre-annotated emotion category of the sample object; determining the emotion category of the object in the video to be processed at least according to the first feature information of each object.
[0006] Optionally, the first preset model is a Swin Transformer model.
[0007] Optionally, determining the continuous image subsequence of each object in the video to be processed includes: if there are multiple continuous image subsequences of any object in the video to be processed, selecting the continuous image subsequence containing the largest number of images.
[0008] Optionally, for a continuous image subsequence of each object, constructing an object pixel information matrix of the object according to the pixel values of the object region in each frame of the image includes: for a continuous image subsequence of each object, determining a first pixel information matrix of the object according to the pixel values of the object region in each frame of the image, where the first pixel information matrix is a two-dimensional matrix, and at least a part of the row vectors or column vectors in the first pixel information matrix correspond one by one to the images in the continuous image subsequence, and the elements of the at least a part of the row vectors or column vectors are the pixel values of the object region in the image corresponding to them; copying the first pixel information matrix of each object to obtain an object pixel information matrix of the object, where the object pixel information matrix is a three-dimensional matrix, and the object pixel information matrix includes a plurality of first pixel information matrices.
[0009] Optionally, for a continuous image subsequence of each object, determining a first pixel information matrix of the object according to the pixel values of the object region in each frame of the image includes: for a continuous image subsequence of each object, determining a plurality of first initial pixel matrices of the object according to the pixel values of the object region in each frame of the image, where the first initial pixel matrix corresponds one by one to each frame of the image in the continuous image subsequence, and the first initial pixel matrix is a multi-dimensional matrix; stretching the plurality of first initial pixel matrices of each object to obtain a plurality of first intermediate pixel matrices of the object, where the first intermediate pixel matrix is a one-dimensional matrix; combining the plurality of first intermediate pixel matrices of the object in the order of multiple frames of the continuous image subsequence of each object to obtain the first pixel information matrix of the object.
[0010] Optionally, before determining the emotion category of each object in the video to be processed at least according to the first feature information of each object, the method further includes: constructing a scene content information matrix of each object at least according to the pixel values of each frame of the image in the continuous image subsequence corresponding to each object; inputting the scene content information matrix of each object into a second feature extraction model to obtain second feature information of the object, where the second feature extraction model is obtained by training a second preset model with second training data, and the second training data includes: the scene content information matrix of the sample object in the sample video and the pre-annotated emotion category of the sample object.
[0011] Optionally, determining the emotion category of each object in the video to be processed at least according to the first feature information of each object includes: performing a fusion process on the first feature information and the second feature information of each object to obtain fused feature information; determining the emotion category of each object in the video to be processed according to the fused feature information of each object.
[0012] Optionally, the second preset model is a Swin Transformer model.
[0013] Optionally, determining the scene content information matrix of each object based on at least the pixel values of each frame of image in the continuous image subsequence of each object includes: obtaining audio data synchronized with the time of each frame of image in the continuous image subsequence corresponding to each object; determining the scene content information matrix of each object according to the pixel values of each frame of image in the continuous image subsequence of each object and the audio data corresponding to each frame of image, where the audio data corresponding to each frame of image is the audio data synchronized with its time.
[0014] Optionally, determining the scene content information matrix of each object according to the pixel values of each frame of image in the continuous image subsequence of each object and the audio data corresponding to each frame of image includes: determining the second pixel information matrix of each object according to the pixel values of each frame of image in the continuous image subsequence of each object and the audio data corresponding to each frame of image, where the second pixel information matrix is a two-dimensional matrix, and at least a part of the row vectors or column vectors in the second pixel information matrix correspond one by one to the images in the continuous image subsequence, and the elements of the at least a part of the row vectors or column vectors include the pixel values of the corresponding images and the corresponding audio data; copying the second pixel information matrix of each object to obtain the scene content information matrix of each object, where the scene content information matrix is a three-dimensional matrix, and the scene content information matrix includes a plurality of second pixel information matrices.
[0015] To solve the above technical problems, an embodiment of the present invention further provides an emotion recognition device, where the device includes: an acquisition module, configured to acquire a video to be processed, where the video to be processed includes images of at least one object; a preprocessing module, configured to determine a continuous image subsequence of each object in the video to be processed, where the continuous image subsequence of each object includes a continuous plurality of frames of images and each frame of image includes the image of the object; a first matrix construction module, configured to, for the continuous image subsequence of each object, construct an object pixel information matrix of the object according to the pixel values of the object region in each frame of image, where the object region is a region including the object; a first feature extraction module, configured to input the object pixel information matrix of each object into a first feature extraction model to obtain first feature information of the object, where the first feature extraction model is obtained by training a first preset model with first training data, and the first training data includes: an object pixel information matrix of a sample object in a sample video and a pre-annotated emotion category of the sample object; and an identification module, configured to determine the emotion category of each object in the video to be processed based on at least the first feature information of each object.
[0016] An embodiment of the present invention also provides a storage medium, on which a computer program is stored. When the computer program is run by a processor, the steps of the above-mentioned emotion recognition method are executed.
[0017] An embodiment of the present invention also provides a terminal, including a memory and a processor. A computer program that can run on the processor is stored on the memory. When the processor runs the computer program, the steps of the above-mentioned emotion recognition method are executed.
[0018] Compared with the prior art, the technical solution of the embodiment of the present invention has the following beneficial effects:
[0019] In the solution of the embodiment of the present invention, a continuous image subsequence of each object in the video to be processed is determined, and an object pixel information matrix of the object is constructed according to the continuous image subsequence of each object. Since the object pixel information matrix is constructed based on the pixel values of the object region in each frame of the image and contains the pixel information of the object region in the image, the object pixel information matrix can be used to represent information such as the posture of the object in the video to be processed. Further, the object pixel information matrix of each object is input into the first feature extraction model to obtain the first feature information of the object. Since the first feature extraction model is obtained by training a first preset model with first training data, and the first training data includes the object pixel information matrix of the sample object in the sample video and the pre-annotated emotion category of the sample object, the trained first feature extraction model can extract the feature information concerned by the emotion recognition task based on the object pixel information matrix. Further still, the first feature information of each object can be used to determine the emotion category of the object in the video to be processed. Compared with the prior art, the solution of the embodiment of the present invention does not need to input each frame of the image in the video to be processed into the model for feature extraction and then determine the emotion category. Instead, the continuous image subsequence of each object in the video to be processed is used as a whole to construct the object pixel information matrix. The constructed object pixel information matrix can represent the information of the object in the video to be processed. Only by inputting the object pixel information matrix of the object into the first preset model can the feature information of the object in the video to be processed be extracted, and further the emotion category of the object in the video to be processed can be obtained. Adopting such a solution is beneficial to more efficiently determining the emotion category of each object in the video.
[0020] Further, in the solution of the embodiment of the present invention, the first preset model is a Swin Transformer model. Adopting such a solution, since the number of layers of the Swin Transformer model is small, it is beneficial to further improve the efficiency of feature extraction.
[0021] Further, in the solution of the embodiment of the present invention, according to the pixel values of each frame of image in the continuous image subsequence of each object, a scene content information matrix of the object is constructed, and the scene content information matrix of each object is input into the second feature extraction model to obtain the second feature information of the object; the first feature information and the second feature information of each object are fused to obtain the fused feature information; according to the fused feature information of each object, the emotion category of the object in the video to be processed is determined. By adopting such a solution, for each object in the video to be processed, the constructed scene content information matrix can represent the scene information associated with the object in the video to be processed, and the emotion category is determined according to the scene information associated with the object, which is beneficial to improving the accuracy of emotion recognition.
[0022] Further, in the solution of the embodiment of the present invention, the scene content information matrix is obtained according to the pixel values of each frame of image and the audio data synchronized with each frame of image in the continuous image subsequence. Thus, the scene content information matrix can more comprehensively represent the scene information associated with the object, which is beneficial to further improving the accuracy of emotion recognition. Description of the Drawings
[0023] Figure 1 is a schematic flowchart of an emotion recognition method in an embodiment of the present invention;
[0024] Figure 2 is Figure 1 a schematic flowchart of a specific implementation manner of step S103 in;
[0025] Figure 3 is a partial schematic flowchart of another emotion recognition method in an embodiment of the present invention;
[0026] Figure 4 is a schematic structural diagram of an emotion recognition device in an embodiment of the present invention;
[0027] Figure 5 is a schematic structural diagram of another emotion recognition device in an embodiment of the present invention. Detailed Embodiments
[0028] As described in the background art, there is an urgent need for an emotion recognition method that can efficiently and accurately recognize the emotions of objects in a video.
[0029] The inventors of the present invention have found that in the prior art, usually each frame of image in a video is respectively input into a trained model for feature extraction, then the emotion category corresponding to each frame of image is determined based on the feature information of each frame of image, and finally the emotion category of the object in the video is determined by synthesizing the emotion categories of multiple frames of images. By adopting such a solution, since feature extraction and recognition need to be performed on each frame of image, the calculation process is complex and the efficiency is low.
[0030] To solve the above technical problems, an embodiment of the present invention provides an emotion recognition method. In the solution of the embodiment of the present invention, a continuous image subsequence of each object in the video to be processed is determined, and an object pixel information matrix of the object is constructed according to the continuous image subsequence of each object. Since the object pixel information matrix is constructed based on the pixel values of the object region in each frame of the image and contains the pixel information of the object region in the image, the object pixel information matrix can be used to represent information such as the posture of the object in the video to be processed. Further, the object pixel information matrix of each object is input into the first feature extraction model to obtain the first feature information of the object. Since the first feature extraction model is trained by using the first training data for the first preset model, and the first training data includes the object pixel information matrix of the sample object in the sample video and the pre-annotated emotion category of the sample object, the trained first feature extraction model can extract the feature information concerned by the emotion recognition task based on the object pixel information matrix. Furthermore, the first feature information of each object can be used to determine the emotion category of the object in the video to be processed. Compared with the prior art, the solution of the embodiment of the present invention does not need to input each frame of the image in the video to be processed into the model for feature extraction and then determine the emotion category. Instead, the continuous image subsequence of each object in the video to be processed is used as a whole to construct the object pixel information matrix. The constructed object pixel information matrix can represent the information of the object in the video to be processed. Only by inputting the object pixel information matrix of the object into the first preset model can the feature information of the object in the video to be processed be extracted, and further the emotion category of the object in the video to be processed can be obtained. By adopting such a solution, it is beneficial to more efficiently determine the emotion category of each object in the video.
[0031] To make the above objects, features, and beneficial effects of the present invention more obvious and understandable, the following will describe the specific embodiments of the present invention in detail with reference to the accompanying drawings.
[0032] Refer to Figure 1 , Figure 1 which is a schematic flowchart of an emotion recognition method in an embodiment of the present invention. The emotion recognition method can be executed by a terminal, and the terminal can be various devices with data receiving and data processing functions. For example, it can be a mobile phone, a computer, a tablet computer, a wearable device, a server, etc., but is not limited thereto. Through the solution in the embodiment of the present invention, the emotion category of the object in the video to be processed can be efficiently recognized.
[0033] The emotion recognition method provided by the embodiment of the present invention can be applied to a variety of fields or scenarios. For example, distance education, assisted driving, intelligent monitoring, etc. The following only exemplarily and non-restrictively describes the application scenarios of the embodiment of the present invention.
[0034] In the first application scenario, during the teaching process, the videos of students can be collected, and the emotions of the students in the videos can be recognized through the solution of this embodiment to improve the teaching effect.
[0035] In the second application scenario, during the work process of customer service staff, the videos of the customer service staff can be collected, and the emotions of the customer service staff during the work process can be recognized through the solution of this embodiment to improve the service quality provided by the customer service staff.
[0036] In the third application scenario, during the driving process, the videos of the driver can be collected, and it can be recognized whether the emotions of the driver belong to the pre-set emotions related to dangerous driving through the solution of this embodiment, so as to perform corresponding processing in time and improve driving safety.
[0037] It should be noted that the emotion recognition method in the embodiments of the present invention can also be applied to other fields or scenarios, and this embodiment does not impose any restrictions.
[0038] Figure 1 The shown emotion recognition method may include the following steps:
[0039] Step S101: Obtain the video to be processed;
[0040] Step S102: Determine the continuous image subsequence of each object in the video to be processed;
[0041] Step S103: For the continuous image subsequence of each object, construct the object pixel information matrix of the object according to the pixel values of the object area in each frame of the image;
[0042] Step S104: Input the object pixel information matrix of each object into the first feature extraction model to obtain the first feature information of the object;
[0043] Step S105: Determine the emotion category of the object in the video to be processed at least according to the first feature information of each object.
[0044] It can be understood that in specific implementation, the method can be implemented in the form of a software program, and the software program runs in a processor integrated inside a chip or a chip module; alternatively, the method can be implemented in a hardware or a combination of hardware and software manner.
[0045] In the specific implementation of step S101, the video to be processed can be stored locally in the terminal or obtained from the outside, and this embodiment of the present invention does not impose any restrictions on this.
[0046] Further, the video to be processed may include images of one or more objects, where the objects may be living beings, such as humans, animals, etc., but are not limited thereto. Specifically, the video to be processed may include multiple frames of images, and each frame of image has a timestamp. Among them, the images of the objects included in the multiple frames of images may be different. For example, the first frame of image may include images of object A and object B, the second frame of image may include an image of object C, may also include an image of object A, or may not include images of any object, etc., and the embodiments of the present invention do not limit this.
[0047] In a specific example, an input video may be obtained and segmented according to a preset duration to obtain multiple video segments. Further, each video segment may be sequentially processed as the video to be processed in chronological order to obtain the emotion categories of each object in each video segment. In other words, the solution of the embodiments of the present invention can be used to identify the emotion categories of objects in an offline video.
[0048] In another specific example, whenever a video of a preset duration is captured, the video of the preset duration may be processed as the video to be processed to obtain the emotion category of the object in the video. In other words, the solution of the embodiments of the present invention can be used to identify the emotion categories of objects in an online video.
[0049] In the specific implementation of step S102, a continuous image subsequence of each object in the video to be processed may be determined. Among them, the continuous image subsequence of each object includes multiple consecutive frames of images and each frame of image contains an image of the object. More specifically, the timestamps of the multiple frames of images in the continuous image subsequence are continuous and the multiple frames of images are arranged in chronological order.
[0050] Specifically, object detection may be performed on each frame of image in the video to be processed to obtain the object region of each frame of image, where the object region refers to the region containing the object. Among them, if the image contains images of multiple objects, the object regions of each object may be obtained. In a non-limiting example, when the type of the object is a person, a human detection algorithm may be used to detect each frame of image to obtain the human frame of each person in each frame of image, and the region indicated by the human frame is the object region.
[0051] It should be noted that in the solution of this embodiment, the object region is the region where the whole object in the image is located. Taking the type of the object being a person as an example, the object region is the region where the whole person is located, not only referring to the face region. Even if the image does not contain a face region, it can still be used to construct the object pixel information matrix and identify the emotion category of the person.
[0052] Furthermore, a tracking algorithm can be adopted to determine continuous image subsequences of each object in the video to be processed according to the object regions of each frame of image. It should be noted that for the same frame of image, if it contains images of multiple objects, then this frame of image can belong to the continuous image subsequences of each object. For example, if an image contains both the image of object A and the image of object B, then this image belongs to both the continuous image subsequence of object A and the image subsequence of object B.
[0053] Furthermore, for any object in the video to be processed, if there are multiple continuous image subsequences of this object, then the continuous image subsequence with the largest number of images can be selected to construct the object pixel information matrix of this object. By adopting such a solution, the pixel information of the object region included in the object pixel information matrix of the object can be as much as possible, enabling the object pixel information matrix to better represent information such as the pose of the object in the video to be processed, thereby facilitating the improvement of the accuracy of emotion recognition.
[0054] In the specific implementation of step S103, for the continuous image subsequence of each object, an object pixel information matrix of this object can be constructed according to the pixel values of the object region in each frame of the image.
[0055] Specifically, the object pixel matrix can include at least one first pixel information matrix. The first pixel information matrix can be a two-dimensional matrix. The pixel values of the object regions in each frame of the continuous image subsequence are arranged in order to form at least a part of the row vectors or column vectors in the first pixel information matrix. More specifically, for at least a part of the row vectors in the above-mentioned first pixel information matrix, each row vector corresponds one-to-one with each image in the continuous image subsequence, and the arrangement order of the row vectors is the same as the order of the images. For example, each row vector is constructed based on a corresponding frame of image, and the next row vector is constructed based on the next frame of image.
[0056] It can be understood that the height and / or width of the object regions in each frame of the continuous image subsequence may be different. Among them, the height can be the number of pixel points in the column direction, and the width can be the number of pixel points in the row direction. Therefore, before constructing the object pixel information matrix, the object regions in each frame of the continuous image subsequence can be processed first to adjust the height of the object region to a preset height and the width to a preset width, where the preset height and preset width are pre-set.
[0057] Further, considering that the number of images in the consecutive image subsequences of multiple objects is usually different, in the solution of the embodiments of the present invention, the number of row vectors of the first pixel information matrix constructed can be the same as the number of images in the video to be processed, and the row vectors of the first pixel information matrix can correspond one by one to the images in the video to be processed. Among them, the elements of the row vectors corresponding to the images in the consecutive image subsequences are obtained according to the pixel values of the object regions in the corresponding images, and the element values of the row vectors corresponding to the images other than the consecutive image subsequences in the video to be processed are all 0. By adopting such a solution, the sizes of the object pixel information matrices of each object in the video to be processed can be made consistent, which is convenient for subsequent processing.
[0058] Referring to Figure 2 , Figure 2 is Figure 1 a schematic flowchart of a specific implementation manner of step S103 in Figure 2 The steps shown in step S103 may include the following steps:
[0059] Step S201: For the consecutive image subsequences of each object, determine a plurality of first initial pixel matrices of the object according to the pixel values of the object regions in each frame of the image;
[0060] Step S202: Stretch the plurality of first initial pixel matrices of each object to obtain a plurality of first intermediate pixel matrices of the object;
[0061] Step S203: Combine the plurality of first intermediate pixel matrices of the object in the order of multiple frames of images in the consecutive image subsequence corresponding to the object to obtain the first pixel information matrix of the object;
[0062] Step S204: Copy the first pixel information matrix of each object to obtain the object pixel information matrix of the object.
[0063] In the specific implementation of step S201, for the consecutive image subsequences of each object, the height of the object region can be adjusted to a preset height first, and the width can be adjusted to a preset width. More specifically, for any object region with a size of h×w, it can be transformed into an object region with a size of H×W. Wherein, H is the preset height and W is the preset width.
[0064] Further, a plurality of first initial pixel matrices can be determined according to the pixel values of the object regions. Specifically, the first initial pixel matrices and the images in the consecutive image subsequences are in one-to-one correspondence, and the element values in the first initial pixel matrices are the pixel values of each pixel point in the corresponding images. More specifically, the image can be a three-channel image, and the first initial pixel matrix can be an H×W×3 matrix.
[0065] In the specific implementation of step S202, multiple first initial pixel matrices of each object can be stretched to obtain multiple first intermediate pixel matrices of the object. That is, each first initial pixel matrix is stretched to obtain a corresponding first intermediate pixel matrix. The first intermediate pixel matrix is a one-dimensional matrix. Thus, the first intermediate pixel matrix and the images in the continuous image subsequence also correspond one by one. More specifically, the H×W×3 matrix can be stretched into a 1×(H×W×3) matrix or a (H×W×3)×1 matrix.
[0066] In the specific implementation of step S203, according to the order of multiple frames of images in the continuous image subsequence of each object, multiple first intermediate pixel matrices of the object are combined to obtain a first pixel information matrix corresponding to the object. The first pixel information matrix is a two-dimensional matrix. It can be understood that the first pixel information matrix and the object in the video to be processed correspond one by one. That is, the first pixel information matrix can contain the pixel information of the object area in multiple frames of images in the corresponding continuous image subsequence. In other words, the first pixel information matrix of each object can include the pixel information of the object in the video to be processed.
[0067] Specifically, if the first intermediate pixel matrix is a 1×(H×W×3) matrix, the combined first pixel information matrix can be an f_n×(H×W×3) matrix; if the first intermediate pixel matrix is a (H×W×3)×1 matrix, the combined first pixel information matrix can be a (H×W×3)×f_n matrix, where f_n is the number of images in the continuous image subsequence, and f_n is a positive integer greater than 1.
[0068] In the specific implementation, considering that the number of images in the continuous image subsequences of multiple objects is usually different, the first pixel information matrix can be filled with zeros, and the obtained first pixel information matrix is an F_n×(H×W×3) matrix. Where F_n is the number of images in the video to be processed, F_n > f_n, and F_n is a positive integer. Among them, the first row to the f_n-th row of the first pixel information matrix correspond one by one to the images in the continuous image subsequence, and the element value of the i-th row is obtained according to the pixel value of the object area in the image corresponding to it. The elements of the (f_n + 1)-th row to the F_n-th row are all 0. Where 1 ≤ i ≤ f_n, and i is a positive integer.
[0069] In the specific implementation of step S204, the first pixel information matrix can be copied to obtain an object pixel information matrix, and the object pixel information matrix can be a three-dimensional matrix. Among them, the object pixel information matrix can be an F_n×F_n×(H×W×3) matrix. In other words, the object pixel information matrix can include multiple first pixel information matrices, and the number of first pixel information matrices can be the same as the number of images in the video to be processed. More specifically, the first pixel information matrix can be copied to obtain an F_n×(H×W×3)×F_n matrix, and then the F_n×(H×W×3)×F_n matrix can be transposed to obtain an F_n×F_n×(H×W×3) matrix, so that the subsequent first feature extraction model can obtain the mutual correlation relationship between F_n images. Thus, in the solution of this embodiment, the first pixel information matrix can be used to represent the information of F_n images. By copying the first pixel information matrix F_n times to obtain an object pixel information matrix of F_n×F_n×(H×W×3), the subsequent first feature extraction model can obtain the mutual correlation information between F_n images, which is beneficial to improving the accuracy of emotion recognition.
[0070] It should be noted that in the solutions of other embodiments, the number of first pixel information matrices in the object pixel information matrix can be any other value, and this embodiment does not limit this. For example, the first pixel information matrix can be directly used as the object pixel information matrix, etc.
[0071] It can be understood that the first pixel information matrix is copied to obtain an object pixel information matrix, and then the first feature information is extracted according to the object pixel information matrix. Compared with directly using the first pixel information as the object pixel information matrix for feature extraction, adopting such a solution can increase the pixel information of the object included in the object pixel information matrix in the video to be processed, which is beneficial to improving the accuracy of subsequent emotion recognition.
[0072] In a specific example, the object pixel information matrix can be input into the fully connected layer module to adjust the number of channels of the object pixel information matrix to a preset number. Specifically, if the object pixel information matrix is an F_n×F_n×(H×W×3) matrix, the input of the fully connected layer module is an F_n×F_n×(H×W×3) matrix, and the output of the fully connected layer module is an F_n×F_n×C matrix. Among them, C is the preset number of channels.
[0073] In another specific example, before performing step S203, the one-dimensional first intermediate pixel matrix can be input into the fully connected layer module to adjust the number of elements in the first intermediate pixel matrix to a preset number. For example, the input of the fully connected layer module can be 1×(H×W×3), and the output of the fully connected layer module can be 1×C. Further, the 1×C matrix is combined and replicated to obtain an F_n×F_n×C matrix.
[0074] Continuing to refer to Figure 1 , before performing step S104, the first preset model can be trained using the first training data to obtain the first feature extraction model. Specifically, multiple sample videos can be obtained. Each sample video includes the images of at least one sample object. The object pixel information matrix of each sample object can be constructed respectively, and each sample object has a pre-annotated emotion category. The emotion category can include one or more of the following: angry, disgusted, afraid, happy, sad, surprised, and neutral, etc., but is not limited thereto. The specific process of constructing the object pixel information matrix of the sample object can refer to the relevant description above and will not be elaborated here.
[0075] Further, the object pixel information matrix and the emotion category of the sample object in the sample video can be used as the first training data to train the first preset model. When the preset training conditions are met, the first feature extraction model can be obtained. It should be noted that the embodiment of the present invention does not limit the process of training the first preset network model using the first training data, and it can be an existing appropriate model training method. For example, it can be the gradient descent method, etc., but is not limited thereto.
[0076] Thus, the first feature extraction model can learn the information that the emotion recognition task focuses on the object pixel information matrix. Among them, the first preset model can be various existing network models with learning capabilities. For example, it can be a Convolutional Neural Networks (CNN) model, or a Donvolutional Neural Networks (DNN) model, etc., but is not limited thereto.
[0077] In a non-limiting example, the first preset model can be a Swin Transformer model.
[0078] Further, the object pixel information matrix of each object in the video to be processed can be input into the first feature extraction model, and the first feature information output by the first feature extraction model can be obtained.
[0079] In a non - restrictive example, the first preset model is the Swin Transformer model. The input of the first feature extraction model can be an object pixel matrix of F_n×F_n×C, and the output is the first feature matrix. Wherein, n is a positive integer, and the value of n depends on the number of encoding units in the Swin Transformer model. In a specific example, the number of encoding units is 4, and the value of n is 4.
[0080] In the specific implementation of step S105, the emotion category of the object can be determined at least according to the first feature information of the object.
[0081] In a specific example, the first feature information of each object can be input into a pre - trained classification model. The pre - trained classification model has learned the mapping relationship between the feature information and the emotion category. Thus, the probability value of the object on each preset emotion category can be obtained, and the emotion category with the largest probability value is used as the emotion category of the object.
[0082] Refer to Figure 3 , Figure 3 which is a partial process schematic diagram of another emotion recognition method in the embodiments of the present invention. Figure 3 The shown emotion recognition method may include the following steps:
[0083] Step S301: Determine the scene content information matrix for constructing the object at least according to the pixel values of each frame of the continuous image subsequence corresponding to each object.
[0084] Step S302: Input the scene content information matrix of each object into the second feature extraction model to obtain the second feature information of the object.
[0085] Step S303: Perform a fusion process on the first feature information and the second feature information of each object to obtain the fused feature information.
[0086] Step S304: Determine the emotion category of the object in the video to be processed according to the fused feature information of each object.
[0087] It should be noted that before executing step S303, steps S101 to S104 can also be executed to extract the first feature information of the object in the video to be processed.
[0088] It should also be noted that only the differences Figure 3 between the shown emotion recognition method and the above - mentioned emotion recognition method will be described below.
[0089] In the specific implementation of step S301, for the continuous image subsequence of each object, a scene content information matrix of the object can be constructed.
[0090] Specifically, the scene content information matrix can include at least one second pixel information matrix, and the second pixel information matrix can be a two-dimensional matrix.
[0091] In a specific example, the pixel values of each image in the continuous image subsequence are arranged in order to form row vectors or column vectors of at least a part of the second pixel information matrix. More specifically, for at least a part of the row vectors in the above-mentioned second pixel information matrix, each row vector corresponds one-to-one with each image in the continuous image subsequence, and the arrangement order of the row vectors is the same as the arrangement order of the images. For example, each row vector is constructed based on a corresponding frame of image, and the next row vector is constructed based on the next frame of image. It should be noted that the row vectors in the first pixel information matrix are constructed based on the pixel values of the object region in a corresponding frame of image, while the row vectors in the second pixel information matrix are constructed based on the pixel values of the whole corresponding frame of image.
[0092] Specifically, the height of each frame of image in the continuous image subsequence can be adjusted to the above-mentioned preset height H, and the width can be adjusted to the above-mentioned preset width W.
[0093] Furthermore, multiple second initial pixel matrices of the object can be determined according to the pixel values of each frame of image in the continuous image subsequence of each object. More specifically, the image can be a three-channel image, and the second initial pixel matrix can be an H×W×3 matrix. It should be noted that the element values in the first initial pixel matrix are the pixel values of the object region, while the element values in the second initial pixel matrix are the pixel values of the whole image.
[0094] Furthermore, the multiple second initial pixel matrices of each object are stretched to obtain multiple second intermediate pixel matrices of the object. That is, the second intermediate pixel matrix is a one-dimensional matrix and corresponds one-to-one with the images in the continuous image subsequence. More specifically, the H×W×3 matrix can be stretched into a 1×(H×W×3) matrix or an (H×W×3)×1 matrix.
[0095] In a non-limiting example, audio data synchronized with each frame of image can also be obtained, and the audio data is added to the second intermediate pixel matrix. Thus, the second intermediate pixel matrix can fuse the pixel information and audio information of the scene, can more comprehensively represent the scene information associated with the object, and is beneficial to further improving the accuracy of emotion recognition subsequently. Thus, the second intermediate pixel matrix can be a 1×(H×W×3 + 1) matrix or an (H×W×3 + 1)×1 matrix. Among them, the audio information can be a fundamental frequency signal, etc., but is not limited thereto.
[0096] Further, according to the order of multiple frames of images in the continuous image subsequence corresponding to each object, multiple second intermediate pixel matrices of the object can be combined to obtain a second pixel information matrix of the object, and the second pixel information matrix is a two-dimensional matrix. It can be understood that the second pixel information matrix and the object in the video to be processed are in one-to-one correspondence, and the second pixel information matrix can include the scene information associated with the object in the video to be processed.
[0097] In a specific implementation, if the second intermediate pixel matrix does not fuse audio information, the second pixel information matrix can be an F_n×(H×W×3) matrix; if the second intermediate pixel matrix fuses audio information, the second pixel information matrix can be an F_n×(H×W×3 + 1) matrix.
[0098] Further, the second pixel information matrix of each object is copied to obtain a scene content information matrix of the object, and the scene content information matrix can be a three-dimensional matrix. In other words, the scene content information matrix can include multiple second pixel information matrices. Among them, the number of second pixel information matrices can be the same as the number of images in the video to be processed.
[0099] It should be noted that more content about the specific process of constructing the scene content information matrix can refer to the relevant description about constructing the object pixel information matrix above, and will not be elaborated here.
[0100] It should also be noted that the size of the constructed scene content information matrix is the same as the size of the object pixel information matrix.
[0101] In the specific implementation of step S302, the scene content information matrix of each object can be input into the second feature extraction model to obtain the second feature information of the object. Among them, the second feature extraction model is trained by using the second training data for the second preset model, and the second training data can be the scene content information matrix and the emotion category of the sample object in the sample video. Among them, the type and structure of the second preset model can be the same as the type and structure of the first preset model. More specifically, the second preset model can also be a Swin Transformer model.
[0102] It should be noted that more content about step S302 can refer to the relevant description of step S104, and will not be elaborated here.
[0103] In the specific implementation of step S303, the first feature information and the second feature information of each object can be concatenated to obtain the fused feature information.
[0104] In a non-limiting example, both the first feature information and the second feature information are For a matrix, the fused feature information is a matrix.
[0105] In the specific implementation of step S304, the fused feature information of each object can be input into a pre-trained classification model. Since the pre-trained classification model has learned the mapping relationship between the feature information and the emotion categories, the probability value of each object on each preset emotion category can be obtained, and the emotion category with the largest probability value is used as the emotion category of the object.
[0106] Regarding Figure 3 For more content about the emotion recognition method shown, reference can be made to the relevant descriptions above, which will not be elaborated here.
[0107] Refer to Figure 4 , Figure 4 FIG. is a schematic structural diagram of an emotion recognition device in an embodiment of the present invention, Figure 4 The emotion recognition device shown may include:
[0108] An acquisition module 41, configured to acquire a video to be processed, where the video to be processed includes images of at least one object;
[0109] A preprocessing module 42, configured to determine a continuous image subsequence of each object in the video to be processed, where the continuous image subsequence of each object includes a plurality of consecutive frames of images and each frame of image includes an image of the object;
[0110] A first matrix construction module 43, configured to, for the continuous image subsequence of each object, construct an object pixel information matrix of the object according to the pixel values of the object region in each frame of image, where the object region is a region including the object;
[0111] A first feature extraction module 44, configured to input the object pixel information matrix of each object into a first feature extraction model to obtain first feature information of the object, where the first feature extraction model is obtained by training a first preset model with first training data, and the first training data includes: an object pixel information matrix of a sample object in a sample video and a pre-annotated emotion category of the sample object;
[0112] An identification module 45, configured to determine the emotion category of the object in the video to be processed at least according to the first feature information of each object.
[0113] Refer to Figure 5 , Figure 5 FIG. is a schematic structural diagram of another emotion recognition device in an embodiment of the present invention. Compared with the emotion recognition device shown in Figure 4 the emotion recognition device shown, Figure 5 the emotion recognition model shown further includes:
[0114] A second matrix construction module 53, configured to construct a scene content information matrix of each object based on at least pixel values of each frame of image in a continuous image subsequence of each object.
[0115] A second feature extraction module 54, configured to input the scene content information matrix of each object into a second feature extraction model to obtain second feature information of the object, where the second feature extraction model is obtained by training a second preset model using second training data, and the second training data includes: the scene content information matrix of a sample object in the sample video and the pre-annotated emotion category of the sample object.
[0116] For more content such as the working principle, working method, and beneficial effects of the emotion recognition device in the embodiments of the present invention, reference may be made to the relevant descriptions of the emotion recognition method above, and details are not described herein again.
[0117] Embodiments of the present invention further provide a storage medium, on which a computer program is stored. When the computer program is run by a processor, the steps of the above-mentioned emotion recognition method are executed. The storage medium may include a ROM, a RAM, a magnetic disk, an optical disc, etc. The storage medium may also include a non-volatile memory or a non-transitory memory, etc.
[0118] Embodiments of the present invention further provide a terminal, including a memory and a processor. A computer program that can run on the processor is stored on the memory. When the processor runs the computer program, the steps of the above-mentioned emotion recognition method are executed. The terminal includes, but is not limited to, terminal devices such as mobile phones, computers, and tablet computers. In a specific example, the terminal may be an in-vehicle terminal. In another specific example, the terminal may be a server.
[0119] It should be understood that in the embodiments of the present application, the processor may be a central processing unit (CPU for short), and the processor may also be other general-purpose processors, digital signal processors (DSP for short), application specific integrated circuits (ASIC for short), field programmable gate arrays (FPGA for short), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0120] It should also be understood that the memory in the embodiments of the present application can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of random access memory (RAM) are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchlink dynamic random access memory (SLDRAM), and direct rambus random access memory (DR RAM).
[0121] The above embodiments can be implemented in whole or in part by software, hardware, firmware, or any other combination. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer program can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer program can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired or wireless manner.
[0122] In several embodiments provided in this application, it should be understood that the disclosed methods, devices, and systems can be implemented in other ways. For example, the device embodiments described above are merely illustrative; for example, the division of the units is only a logical function division, and there can be other division methods in actual implementation; for example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0123] In addition, each functional unit in various embodiments of the present invention can be integrated in a processing unit, or each unit can be physically included separately, or two or more units can be integrated in one unit. The above integrated unit can be implemented in the form of hardware or in the form of a hardware plus software functional unit. For example, for each device and product applied to or integrated into a chip, each module / unit included therein can be implemented in the form of hardware such as circuits, or at least some modules / units can be implemented in the form of software programs that run on a processor integrated inside the chip, and the remaining (if any) part of the modules / units can be implemented in the form of hardware such as circuits; for each device and product applied to or integrated into a chip module, each module / unit included therein can be implemented in the form of hardware such as circuits, and different modules / units can be located in the same component (such as a chip, a circuit module, etc.) or different components of the chip module, or at least some modules / units can be implemented in the form of software programs that run on a processor integrated inside the chip module, and the remaining (if any) part of the modules / units can be implemented in the form of hardware such as circuits; for each device and product applied to or integrated into a terminal, each module / unit included therein can be implemented in the form of hardware such as circuits, and different modules / units can be located in the same component (such as a chip, a circuit module, etc.) or different components inside the terminal, or at least some modules / units can be implemented in the form of software programs that run on a processor integrated inside the terminal, and the remaining (if any) part of the modules / units can be implemented in the form of hardware such as circuits.
[0124] It should be understood that the term "and / or" in this article is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this article indicates that the associated objects before and after are in an "or" relationship.
[0125] In the embodiments of the present application, "a plurality of" means two or more. The first, second, etc. descriptions that appear in the embodiments of the present application are only used for illustration and to distinguish the described objects, without an order, and do not represent a special limitation on the number of devices in the embodiments of the present application, and cannot constitute any limitation to the embodiments of the present application.
[0126] Although the present invention is disclosed as above, the present invention is not limited thereto. Any person skilled in the art can make various changes and modifications without departing from the spirit and scope of the present invention. Therefore, the protection scope of the present invention should be subject to the scope defined by the claims.
Claims
1. A method for emotion recognition, characterized in that, The method includes: Obtaining a video to be processed, where the video to be processed includes images of at least one object; Determining a continuous image subsequence of each object in the video to be processed, where the continuous image subsequence of each object includes multiple consecutive frames of images and each frame of image contains an image of the object; For the continuous image subsequence of each object, constructing an object pixel information matrix of the object according to the pixel values of the object region in each frame of the image, where the object region is the region containing the object; Inputting the object pixel information matrix of each object into a first feature extraction model to obtain first feature information of the object, where the first feature extraction model is obtained by training a first preset model with first training data, and the first training data includes: the object pixel information matrix of a sample object in a sample video and the pre-annotated emotion category of the sample object; Determining the emotion category of the object in the video to be processed at least according to the first feature information of each object; Wherein, for the continuous image subsequence of each object, constructing the object pixel information matrix of the object according to the pixel values of the object region in each frame of the image includes: For each consecutive image subsequence of an object, a first pixel information matrix of the object is determined according to the pixel values of the object region in each frame of the images, where the first pixel information matrix is a two-dimensional matrix, at least a part of the row vectors or column vectors in the first pixel information matrix correspond one by one to the images in the consecutive image subsequence, and the elements of the at least a part of the row vectors or column vectors are obtained according to the pixel values of the object region in the images corresponding thereto, where the first pixel information matrix is a matrix, and the first pixel information matrix is used to characterize the information of M images, H H is a preset height, W W is a preset width; Copy the first pixel information matrix to obtain a matrix, and then transpose the matrix to obtain the object pixel information matrix.
2. The emotion recognition method according to claim 1, wherein The first preset model is a Swin Transformer model.
3. The emotion recognition method according to claim 1, characterized in that Determining the continuous image subsequence of each object in the video to be processed includes: If there are multiple continuous image subsequences of any object in the video to be processed, select the continuous image subsequence with the largest number of images.
4. The emotion recognition method according to claim 1, wherein For the continuous image subsequence of each object, determining the first pixel information matrix of the object according to the pixel values of the object region in each frame of the image includes: For the continuous image subsequence of each object, determining a plurality of first initial pixel matrices of the object according to the pixel values of the object region in each frame of the image, where the first initial pixel matrix corresponds to each frame of the continuous image subsequence one by one, and the first initial pixel matrix is a multi-dimensional matrix; Stretching the plurality of first initial pixel matrices of each object to obtain a plurality of first intermediate pixel matrices of the object, where the first intermediate pixel matrix is a one-dimensional matrix; Combining the plurality of first intermediate pixel matrices of the object in the order of multiple frames of the continuous image subsequence of each object to obtain the first pixel information matrix of the object.
5. The emotion recognition method according to claim 1, wherein Before determining the emotion category of the object in the video to be processed at least according to the first feature information of each object, the method further includes: Constructing a scene content information matrix of the object at least according to the pixel values of each frame of the continuous image subsequence of each object; Inputting the scene content information matrix of each object into a second feature extraction model to obtain second feature information of the object, where the second feature extraction model is obtained by training a second preset model with second training data, and the second training data includes: the scene content information matrix of the sample object in the sample video and the emotion category of the sample object.
6. The emotion recognition method according to claim 5, characterized in that Determining the emotion category of the object in the video to be processed at least according to the first feature information of each object includes: Fuse the first feature information and the second feature information of each object to obtain the fused feature information; Determine the emotion category of each object in the video to be processed according to the fused feature information of each object.
7. The emotion recognition method according to claim 5, characterized in that The second preset model is the Swin Transformer model.
8. The emotion recognition method according to claim 5, wherein Determine the scene content information matrix of each object at least according to the pixel values of each frame of image in the continuous image subsequence of each object, including: For each frame of image in the continuous image subsequence of each object, obtain the audio data synchronized with its time; Determine the scene content information matrix of each object according to the pixel values of each frame of image and the audio data corresponding to each frame of image in the continuous image subsequence of each object, where the audio data corresponding to each frame of image is the audio data synchronized with its time.
9. The emotion recognition method according to claim 8, characterized in that Determine the scene content information matrix of each object according to the pixel values of each frame of image and the audio data corresponding to each frame of image in the continuous image subsequence of each object, including: Determine the second pixel information matrix of each object according to the pixel values of each frame of image and the audio data corresponding to each frame of image in the continuous image subsequence of each object, where the second pixel information matrix is a two-dimensional matrix, and at least a part of the row vectors or column vectors in the second pixel information matrix correspond one by one to the images in the continuous image subsequence, and the elements of the at least a part of the row vectors or column vectors include the pixel values of the corresponding images and the corresponding audio data; Copy the second pixel information matrix of each object to obtain the scene content information matrix of the object, where the scene content information matrix is a three-dimensional matrix, and the scene content information matrix includes multiple second pixel information matrices.
10. An emotion recognition device, characterized in that, The device includes: An acquisition module, configured to acquire a video to be processed, where the video to be processed includes the images of at least one object; A preprocessing module, configured to determine the continuous image subsequence of each object in the video to be processed, where the continuous image subsequence of each object includes multiple consecutive frames of images and each frame of image includes the image of the object; A first matrix construction module, configured to, for the continuous image subsequence of each object, construct the object pixel information matrix of the object according to the pixel values of the object region in each frame of image, where the object region is the region containing the object A first feature extraction module, configured to input the object pixel information matrix of each object into a first feature extraction model to obtain the first feature information of the object, where the first feature extraction model is obtained by training a first preset model with first training data, and the first training data includes: the object pixel information matrix of the sample object in the sample video and the pre-annotated emotion category of the sample object; An identification module, configured to determine the emotion category of each object in the video to be processed at least according to the first feature information of each object; Wherein, the first matrix construction module includes: A sub-module for determining a first pixel information matrix of an object according to pixel values of an object region in each frame of a continuous image subsequence for each object, wherein the first pixel information matrix is a two-dimensional matrix, at least a part of row vectors or column vectors in the first pixel information matrix correspond one by one to images in the continuous image subsequence, and elements of the at least a part of row vectors or column vectors are obtained according to pixel values of the object region in the images corresponding thereto, wherein the first pixel information matrix is a matrix for characterizing information of H a preset height, W and a preset width; For copying the first pixel information matrix to obtain a matrix, and then transposing the matrix to obtain the sub-module of the object pixel information matrix.
11. A storage medium, on which a computer program is stored, characterized in that, When the computer program is run by a processor, it executes the steps of the emotion recognition method according to any one of claims 1 to 9.
12. A terminal, comprising a memory and a processor, wherein a computer program that can run on the processor is stored on the memory, characterized in that, When the processor runs the computer program, it executes the steps of the emotion recognition method according to any one of claims 1 to 9.
Citation Information
Patent Citations
Driver emotion recognition algorithm design based on face image and vehicle operation information
CN110516658A
Video-based micro-expression recognition method and device, equipment and storage medium
CN113435330A
Character emotion recognition method, system, medium and electronic equipment
CN113536999A