Training Method, Recognition Method, System, Device and Medium of Emotion Recognition Model

By using the improved Transformer encoding structure in the emotion recognition model, combining the visual features and correlation information in the video sample, the problem in the prior art is difficult to accurately identify the emotions of the target object in real time in real dynamic scenarios, and the emotion recognition and real-time tracking and recognition of multi-target objects are achieved.

CN114596606BActive Publication Date: 2025-06-27HUA DATA TECH (SHANGHAI) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111371693.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-18
Publication Date
2025-06-27
Estimated Expiration
2041-11-18

AI Technical Summary

Technical Problem

The prior art is difficult to accurately identify the emotions of the target object in real time in real dynamic scenarios.

Method used

By obtaining the video samples of the target object, the visual feature information of each frame of the image and the correlation information between adjacent images are extracted, and the emotional recognition model is constructed using the improved Transformer encoding structure.

Benefits of technology

It realizes emotional recognition of multiple target objects in real scenes, and can track and recognize emotional information in real time, improving recognition accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114596606B_ABST
    Figure CN114596606B_ABST
Patent Text Reader

Abstract

The present invention discloses a training method, recognition method, system, device and medium for an emotion recognition model. The training method includes: obtaining a training set of a target object; for each video sample, obtaining the target detection information of the video sample, where the target detection information includes the visual feature information of each frame of image and the association information between adjacent images; using the target detection information as the input and the preset emotion information of the target object in each frame of image as the output to perform model training to construct an emotion recognition model. In the present invention, by training video samples, not only can the visual feature information of each frame of image be obtained, but also the association information between each image can be obtained. Based on the association information, the dynamic changes of the target object can be reflected, and the changes of emotions can be better reflected. Based on the association information, model training can be carried out more effectively to obtain an emotion recognition model that can accurately recognize the emotion information of the target object in a real scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of artificial intelligence information processing, and particularly relates to a training method, an identification method, a system, a device and a medium for an emotion recognition model. Background Art

[0002] Emotion recognition is an important part of the research on human-computer interaction and affective computing. Emotion recognition has a wide range of applications in various fields. For example, in the field of psychology, emotion recognition technology and other behavior analysis technologies can help researchers obtain more comprehensive and accurate data; in the field of children's education, emotion recognition technology can be used to evaluate teaching achievements and select suitable educational methods for children's different reactions; in the medical field, the pain level or depression level of patients can be evaluated through facial emotion recognition. In addition, emotion recognition can also be applied to the fields of nursing, traffic driving, letters and visits, and so on.

[0003] By recognizing emotions, it is possible to quickly make a judgment on an overly excited target object and issue a warning, thereby effectively preventing the occurrence of violent incidents and accelerating the response and processing speed of conflict incidents. However, in the prior art, emotion recognition is usually performed on static images or static scenes, and this method is difficult to be applied to real dynamic scenes and difficult to monitor the emotions of target objects in real time and accurately. Summary of the Invention

[0004] The technical problem to be solved by the present invention is to overcome the defect that it is difficult to identify the emotions of target objects in real time and accurately in the prior art, and to provide a training method, an identification method, a system, a device and a medium for an emotion recognition model that can identify the emotions of target objects in real time and accurately.

[0005] The present invention solves the above technical problem through the following technical solutions:

[0006] The present invention provides a training method for an emotion recognition model, and the training method includes:

[0007] Obtain a training set of a target object, where the training set includes a plurality of video samples, and each video sample includes at least two frames of images;

[0008] For each video sample, obtain the target detection information of the video sample, where the target detection information includes the visual feature information of each frame of image and the association information between adjacent images;

[0009] Use the target detection information as input and the preset emotion information of the target object in each frame of image as output for model training to construct the emotion recognition model.

[0010] Preferably, the obtaining of the target detection information of the video sample includes the following steps:

[0011] Input the video sample into a visual feature extraction algorithm to obtain the visual feature information of each frame of image;

[0012] Input the visual feature information of each frame of image into a target encoding structure to obtain the correlation information between adjacent images.

[0013] Preferably, the target encoding structure includes a Transformer (an encoding mechanism) encoding structure.

[0014] Preferably, the Transformer encoding structure includes a nested residual position encoding structure, and the nested residual position encoding structure includes a position encoding function and N cascaded dropout functions, N≥2, n = 1, 2…N, and n represents the nth dropout function;

[0015] The inputting of the visual feature information of each frame of image into the target encoding structure to obtain the correlation information between adjacent images includes the following steps:

[0016] For adjacent images, input the visual feature information of each frame of image into the position encoding function to obtain the first encoding information;

[0017] When n = 1, the input of the corresponding dropout function is the first encoding information and the visual feature information of each frame of image;

[0018] When n = 2,…N, the input of the corresponding dropout function is the output of the previous dropout function and the visual feature information of each frame of image, or the input of the corresponding dropout function is the outputs of all the previous dropout functions and the visual feature information of each frame of image;

[0019] When n = N, the output of the corresponding dropout function is the correlation information between the adjacent images.

[0020] Preferably, the obtaining of the target detection information of the video sample includes:

[0021] Input the video sample into a target detection algorithm to obtain first detection information, and the first detection information includes the box information of each frame of image, and the box information is used to represent the region of the target object in the image;

[0022] Obtain the target detection information according to the first detection information.

[0023] Preferably, the obtaining of the detection information according to the first detection information includes the following steps:

[0024] Input the first detection information into a target tracking algorithm to obtain second detection information, where the second detection information includes the motion trajectory of the target object;

[0025] Obtain target detection information based on the first detection information and the second detection information.

[0026] Preferably, the first detection information includes bounding box information of different target objects;

[0027] The obtaining of the target detection information based on the first detection information includes the following steps:

[0028] For each target object, intercept the region corresponding to the target object in each frame of the image according to the bounding box information;

[0029] Obtain target detection information based on the intercepted image.

[0030] Preferably, the target object specifically includes the facial expression of the target object and / or the body movements of the target object.

[0031] The present invention also provides an emotion recognition method, and the emotion recognition method includes the following steps:

[0032] Obtain a video to be detected of a target object;

[0033] Input the video to be detected into an emotion recognition model to obtain corresponding emotion information, where the emotion recognition model is a model obtained according to the training method of the emotion recognition model as described above.

[0034] The present invention also provides a training system for an emotion recognition model, and the training system includes: a training set acquisition module, a detection information acquisition module, and a model construction module;

[0035] The training set acquisition module is used to obtain a training set of a target object, where the training set includes a number of video samples, and each video sample includes at least two frames of images;

[0036] The detection information acquisition module is used to, for each video sample, obtain the target detection information of the video sample, where the target detection information includes the visual feature information of each frame of the image and the association information between adjacent images;

[0037] The model construction module is used to use the target detection information as input and the preset emotion information of the target object in each frame of the image as output for model training to construct the emotion recognition model.

[0038] The present invention also provides an emotion recognition system, and the emotion recognition system includes: a video acquisition module and an emotion recognition module;

[0039] The video acquisition module is used to acquire the video to be detected of the target object;

[0040] The emotion recognition module is used to input the video to be detected into an emotion recognition model to obtain corresponding emotion information, and the emotion recognition model is a model obtained according to the training system of the emotion recognition model as described above.

[0041] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the training method of the emotion recognition model as described above or the emotion recognition method as described above is implemented.

[0042] The present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the training method of the emotion recognition model as described above or the emotion recognition method as described above is implemented.

[0043] The positive and progressive effects of the present invention are as follows: In the present invention, through training video samples, not only can the visual feature information of each frame of image be obtained, but also the correlation information between each image can be obtained. Based on the correlation information, the dynamic changes of the target object can be reflected, the changes of emotions can be better reflected, and based on the correlation information, the model training can be more effectively carried out to obtain an emotion recognition model that can accurately identify the emotion information of the target object in a real scene.

[0044] Through the emotion recognition model of the present invention, not only can the visual feature information of each frame of image be obtained, but also the correlation information between each image can be obtained. Based on the correlation information, the dynamic changes of the target object can be reflected, the changes of emotions can be better reflected, and therefore the emotion information of the target object in a real scene can be accurately identified. On the one hand, the present invention overcomes the conditional limitations of existing emotion recognition algorithms that focus on specific scenarios or single individuals, can identify multiple target objects in a real scene, and further can identify the emotion information of multiple target objects, and can realize real-time dynamic tracking and recognition of the emotions of target objects in a real scene.

[0045] In the present invention, since an improved Transformer encoding structure is adopted in the emotion recognition model, the accuracy of emotion information recognition is further improved. Description of the Drawings

[0046] Figure 1 It is a flowchart of the training method of the emotion recognition model in Embodiment 1 of the present invention.

[0047] Figure 2 It is a flowchart of the implementation manner of obtaining target detection information in Embodiment 1 of the present invention.

[0048] Figure 3 Schematic diagram of partial modules of the Transformer encoding structure.

[0049] Figure 4 Schematic diagram of partial modules of the improved Transformer encoding structure.

[0050] Figure 5 Flowchart of the implementation method for obtaining target detection information in Embodiment 1 of the present invention.

[0051] Figure 6 Flowchart of the implementation method for step 1122 in Embodiment 1 of the present invention.

[0052] Figure 7 Flowchart of the overall algorithm in Embodiment 1 of the present invention.

[0053] Figure 8 Flowchart of partial algorithms in Embodiment 1 of the present invention.

[0054] Figure 9 Flowchart of the emotion recognition method in Embodiment 2 of the present invention.

[0055] Figure 10 Schematic diagram of modules of the training system of the emotion recognition model in Embodiment 3 of the present invention.

[0056] Figure 11 Schematic diagram of modules of the emotion recognition system in Embodiment 4 of the present invention.

[0057] Figure 12 Schematic diagram of modules of the electronic device in Embodiment 5 of the present invention. Detailed implementation manners

[0058] For ease of understanding, the terms that often appear in the embodiments are explained below:

[0059]

Definition of "including"

[0060]

Definition of "and / or"

[0061]

Definition of "first", "second", etc.

[0062] The present invention will be further described below by way of embodiments, but the present invention is not limited to the scope of the described embodiments.

[0063] Embodiment 1

[0064] This embodiment provides a training method for an emotion recognition model, as Figure 1 shown, the training method includes:

[0065] Step 101, obtain a training set of a target object.

[0066] Among them, the target object can be selected according to requirements, such as people, animals, etc. For the purpose of better describing this embodiment, people are used as the target object to illustrate the embodiment below:

[0067] Among them, the training set includes a number of video samples, and each video sample includes at least two frames of images.

[0068] For example, video data in a real scene can be collected and cut every 5 seconds, that is, each 5-second video is a video sample. The face recognition algorithm is used to perform face recognition on each video sample, and the video samples in which no face is detected in the video samples are deleted. In this way, data can be effectively screened to establish a training set.

[0069] Step 102, for each video sample, obtain the target detection information of the video sample.

[0070] Among them, the target detection information includes the visual feature information of each frame of image and the association information between adjacent images.

[0071] Specifically, the target detection information can be obtained through a pre-trained model.

[0072] Step 103: Use the target detection information as the input and the preset emotion information of the target object in each frame of the image as the output to train the model to construct an emotion recognition model.

[0073] In this embodiment, the implementation methods of different emotion information can be set according to the actual situation. In a specific implementation method, the preset emotion information is specifically the emotion excitement index, which can specifically include five categories: positive emotion, calm, first-level negative emotion, second-level negative emotion, and unable to obtain emotion. Among them, the first-level negative emotion indicates that the pedestrian is in a relatively excited state, the second-level negative emotion indicates that the pedestrian is in a very excited state, and unable to obtain emotion indicates that the pedestrian is in a special state such as being in the back or passing by in a flash in the video.

[0074] For each frame of the image, the emotion excitement index of the pedestrians in the image can be tagged in advance to form the preset emotion information.

[0075] Specifically in Step 103, the various target detection information can be concatenated together, and a fully connected layer is used to map the concatenated encoded features to the category space. Then, the softmax (a classification function) function is used to calculate the probability belonging to each emotion category, and the category with the maximum probability is used as the emotion excitement index level of the pedestrian. The calculated emotion excitement index level is compared with the preset emotion information used as the output, and then the loss error of the target loss function is adjusted. When the preset training stop condition is met, the emotion recognition model is successfully constructed.

[0076] In this embodiment, by training the video samples, not only the visual feature information of each frame of the image can be obtained, but also the correlation information between the images can be obtained. Based on the correlation information, the dynamic changes of the target object can be reflected, the changes of emotions can be better reflected, and based on the correlation information, the model training can be carried out more effectively to obtain an emotion recognition model that can accurately identify the emotion information of the target object in the real scene.

[0077] In a specific implementation method, as Figure 2 shown, the specific steps for obtaining the target detection information of the video sample in Step 102 can include the following steps:

[0078] Step 1021: Input the video sample into the visual feature extraction algorithm to obtain the visual feature information of each frame of the image.

[0079] Specifically, the visual feature information of each frame of the image can be obtained through a pre-trained visual feature extraction algorithm, such as the ResNet101 (a convolutional neural network) algorithm.

[0080] Step 1022: Input the visual feature information of each frame of the image into the target encoding structure to obtain the correlation information between adjacent images.

[0081] Among them, a Transformer encoding structure is specifically adopted as the target encoding structure, and the association information between adjacent images can be effectively obtained through the Transformer encoding structure.

[0082] In this embodiment, the visual feature information of each frame of image can be efficiently obtained through a pre-trained visual feature extraction algorithm, and the association information between adjacent images can be efficiently obtained through the target encoding structure. Based on this, the efficiency of model training can be improved.

[0083] In a specific implementation manner, the Transformer encoding structure includes a nested residual position encoding structure, and the nested residual position encoding structure includes a position encoding function and N cascaded dropout functions, N≥2, n = 1, 2... N, and n represents the nth dropout function.

[0084] Step 1022 may specifically include the following steps:

[0085] For adjacent images, the visual feature information of each frame of image is input into the position encoding function to obtain the first encoding information;

[0086] When n = 1, the input of the corresponding dropout function is the first encoding information and the visual feature information of each frame of image;

[0087] When n = 2,... N, the input of the corresponding dropout function is the output of the previous dropout function and the visual feature information of each frame of image;

[0088] When n = N, the output of the corresponding dropout function is the association information between adjacent images.

[0089] In other embodiments, when n = 2, the input of the corresponding dropout function may also be the outputs of all the previous dropout functions and the visual feature information of each frame of image. In this embodiment, it is preferably that the input of the corresponding dropout function is the output of the previous dropout function and the visual feature information of each frame of image to reduce the computational pressure and improve the speed of model training.

[0090] To better understand step 1022, the following takes N = 2 as an example to further illustrate step 1022:

[0091] Assume that the visual feature information of each frame of image is X.

[0092] Generally speaking, as Figure 3As shown, the position encoding module in the Transformer encoding structure includes a position encoding function and a dropout function. X is input into the position encoding function to obtain the first encoded information. The inputs of the dropout function are the first encoded information and X, and the output is Y1.

[0093] In this embodiment, the position encoding module is improved and changed into a nested residual position encoding structure, that is, as Figure 4 shown, the input of the second dropout function becomes X and the output Y1 of the first dropout function, and the output is Y2.

[0094] Since the role of the dropout function is to randomly discard a part of the data, therefore, in this embodiment, through multiple cascaded dropout functions, more original information can be extracted, that is, the visual feature information of each frame of image. While enhancing the original information, the generalization ability of the model can be improved through the dropout function, and further improve the accuracy of model training. After experiments, the improved transformer model can improve the recognition accuracy by about 4%.

[0095] It should be understood that in this embodiment, the encoding function is set at the entrance of the Transformer encoding structure to perform position encoding on the visual feature information of each frame of image, so as to obtain the position correlation information between images. The dropout function can be directly set after the encoding function, or indirectly set after the encoding function, such as being set in other modules of the Transformer encoding structure, such as being set in the Multi-head attention (multi-head attention mechanism), being set in the Add&Norm (a normalization processing mechanism), etc. Through the dropout function, the data in each module can be randomly discarded, improving the generalization ability of the model.

[0096] In a specific implementation manner, as Figure 5 shown, the obtaining of the target detection information of the video sample in step 102 may specifically further include:

[0097] Step 1121: Input the video sample into the target detection algorithm to obtain the first detection information.

[0098] Among them, the first detection information includes the box information of each frame of image, and the box information is used to represent the area of the target object in the image.

[0099] Specifically, a pre-trained YOLOV4 (a target detection algorithm) model can be used to detect pedestrians in each video sample, so as to obtain the box information of each frame of image.

[0100] Step 1122. Obtain the target detection information according to the first detection information.

[0101] Specifically, refer to Step 1021, input the target detection information into the visual feature extraction algorithm to obtain the visual feature information of each frame of image, and refer to Step 1022 to obtain the target detection information.

[0102] In this embodiment, the box information, that is, the region of the target object in the image, can be efficiently obtained through the target detection algorithm. Then, the features of the target object can be extracted specifically based on the box information, thereby improving the efficiency and accuracy of model training.

[0103] In a specific implementation manner, as Figure 6 shown, Step 1122 may specifically include the following steps:

[0104] Step 11221. Input the first detection information into the target tracking algorithm to obtain the second detection information.

[0105] Among them, the second detection information includes the motion trajectory of the target object.

[0106] Specifically, the pre-trained DeepSort (a multi-object tracking algorithm) algorithm can be used to track the target object to obtain the motion trajectory of the target object.

[0107] Step 11222. Obtain the target detection information according to the first detection information and the second detection information.

[0108] In this embodiment, the motion trajectory of the target object can be obtained, which can reflect the dynamic changes of the target object. Based on this dynamic change, the feature information that can reflect the emotional changes of the target object can be more effectively extracted, thus further improving the accuracy of model training.

[0109] In a specific implementation manner, the first detection information includes the box information of different target objects;

[0110] Step 1121 specifically includes the following steps:

[0111] For each target object, intercept the region corresponding to the target object in each frame of image according to the box information; obtain the target detection information according to the intercepted image.

[0112] In this embodiment, the region corresponding to the target object is intercepted and used as the target detection information, so that the feature information of the target object can be trained specifically. In addition, when multiple target objects are included, different target objects can be distinguished, and the corresponding features of each target object can be extracted for training, thereby further improving the accuracy of model training.

[0113] In a specific implementation, the target object specifically includes the facial expression of the target object and / or the body movements of the target object. This embodiment preferably includes a solution that simultaneously includes facial expression muscles and body movements.

[0114] Since both facial expressions and body movements can reflect the emotional information of the target object, in this embodiment, setting to simultaneously include facial expressions and body movements can perform feature extraction more comprehensively, and thus a model that can better reflect the emotional changes of the target object in the real scene can be obtained.

[0115] To better understand this embodiment, the following uses a specific example to illustrate this embodiment:

[0116] Figure 7 The overall algorithm flowchart of this embodiment is shown.

[0117] First, step 101 is executed to collect real-scene videos offline at a frame rate of 25. The video is segmented into a video sample set of 5 seconds, and each video sample is detected to remove samples without people. The video sample set specifically includes a training set and a test set. The training set is used to train the model, and the test set is used to test the trained model.

[0118] Next, step 102 is executed. For each video sample, it is input into the object detection algorithm to detect the bounding box information of each pedestrian to obtain the first detection information. Here, the object detection algorithm uses the pre-trained YOLOV4 model.

[0119] The pedestrians detected in each frame of the video sample are labeled, and the pedestrian emotional arousal index is set to five levels: positive emotion, calm, first-level negative emotion, second-level negative emotion, and unable to obtain emotion. Specifically, multiple annotators can perform the annotation separately. If the annotation results are different, the principle of the minority obeying the majority is adopted. If the annotation results of multiple annotators are different from each other, this sample is removed.

[0120] The first detection information, that is, the information including the bounding box information (the specific bounding box information in this embodiment is the pedestrian bounding box) is sent to the object tracking algorithm to obtain the movement trajectories of each pedestrian in the video sample, such as the trajectory bounding box of pedestrian No. 1, the trajectory bounding box of pedestrian No. 2... the trajectory bounding box of pedestrian No. N. Here, the object tracking algorithm uses the pre-trained DeepSort model.

[0121] According to the tracking bounding boxes of each pedestrian in each video sample, screenshots of each pedestrian in each frame of the video are intercepted. Then, the pedestrian screenshots are fed into a visual feature extraction algorithm (specifically, a pre-trained convolutional neural network Resnet101 model in this embodiment) to extract the visual features of the pedestrians. The visual features corresponding to each bounding box are 2048-dimensional. If there is no such pedestrian information in the frame image, the feature value is set to 0.

[0122] For each pedestrian, each video sample contains features of 125 * 2048, where 125 represents the number of frames in a 5-second video, and 2048 represents the dimension of the visual features in each frame. The features of each pedestrian are respectively fed into a Transformer encoding structure.

[0123] As Figure 8 shown, the left part is the Transformer encoding structure. Among them, the nested residual position encoding is the combined part of the position encoding function and the dropout function. After passing through the Transformer encoding structure, the correlation information between adjacent images will be obtained. The correlation information and the visual feature information are concatenated and mapped to the class space through a fully connected layer to obtain the probability on each class. During training, the cross-entropy loss function is used as the loss function. During testing, the class to which the maximum probability value belongs is used as the class of the emotional arousal index of the pedestrian.

[0124] In this embodiment, the dataset (i.e., the video sample set) is divided into 10 parts, one part is used as the test set, and nine parts are used as the training set to train the Transformer structure. During this process, the parameters in the object detection algorithm, object tracking algorithm, and visual feature extraction algorithm are controlled to remain unchanged.

[0125] During the training process of this embodiment, the dataset is trained for 100 rounds. Each round is fed into the model in batches, and each batch contains 32 data. After each round of training, it is tested once on the test set to obtain the recognition accuracy. After 100 rounds, the model with the highest accuracy on the test set is taken and saved as the finally trained emotion recognition model.

[0126] On the one hand, this embodiment overcomes the conditional limitations of existing emotion recognition algorithms that focus on specific scenarios or single individuals. It can train multiple target objects in a real scenario, and thus can recognize the emotion information of multiple target objects, and can realize the real-time tracking and recognition of the emotions of target objects in a real scenario. In addition, this embodiment improves the general Transformer encoding structure and further improves the accuracy of model training.

[0127] Embodiment 2

[0128] This embodiment provides an emotion recognition method, as Figure 9As shown in the figure, the emotion recognition method includes the following steps:

[0129] Step 201: Obtain the video to be detected of the target object;

[0130] Step 202: Input the video to be detected into the emotion recognition model to obtain the corresponding emotion information.

[0131] Among them, there can be multiple target objects, such as pedestrian 1, pedestrian 2, pedestrian 3, etc., and the output is the emotion information corresponding to each target object. The video to be detected can be a real-time monitoring video. Therefore, the real-time monitored and dynamically changing emotion information of the target object can be obtained.

[0132] The emotion recognition model is the model obtained according to the training method of the emotion recognition model described in Embodiment 1.

[0133] In this embodiment, through the emotion recognition model, not only the visual feature information of each frame of image can be obtained, but also the correlation information between each image can be obtained. Based on the correlation information, the dynamic changes of the target object can be reflected, and the changes of emotions can be better reflected. Therefore, the emotion information of the target object in the real scene can be accurately recognized. On the one hand, this embodiment overcomes the condition limitations of existing emotion recognition algorithms that focus on specific scenarios or single individuals. It can recognize multiple target objects in a real scene, and then can recognize the emotion information of multiple target objects, and can realize the real-time dynamic tracking and recognition of the emotions of the target object in the real scene. In addition, due to the improved Transformer encoding structure adopted in the emotion recognition model, the accuracy of emotion information recognition is further improved.

[0134] Embodiment 3

[0135] This embodiment provides a training system for an emotion recognition model, as Figure 10 shown, the training system includes: a training set acquisition module 301, a detection information acquisition module 302, and a model construction module 303.

[0136] The training set acquisition module 301 is used to acquire the training set of the target object. The training set includes a number of video samples, and each video sample includes at least two frames of images.

[0137] The detection information acquisition module 302 is used to obtain the target detection information of each video sample. The target detection information includes the visual feature information of each frame of image and the correlation information between adjacent images.

[0138] The model construction module 303 is used to use the target detection information as the input and the preset emotion information of the target object in each frame of image as the output for model training to construct an emotion recognition model.

[0139] In this embodiment, the specific implementation of each module can refer to the corresponding method in Embodiment 1, and will not be elaborated here.

[0140] In this embodiment, the model construction module can obtain not only the visual feature information of each frame of image but also the correlation information between each image by training video samples. Based on the correlation information, the dynamic changes of the target object can be reflected, the changes of emotions can be better reflected, and the model can be trained more effectively based on the correlation information to obtain an emotion recognition model that can accurately identify the emotion information of the target object in the real scene.

[0141] Embodiment 4

[0142] This embodiment provides an emotion recognition system, as Figure 11 shown, the emotion recognition system includes: a video acquisition module 401 and an emotion recognition module 402.

[0143] The video acquisition module 401 is used to acquire the video to be detected of the target object;

[0144] The emotion recognition module 402 is used to input the video to be detected into the emotion recognition model to obtain the corresponding emotion information, and the emotion recognition model is the model obtained by the training system of the emotion recognition model described in Embodiment 3.

[0145] In this embodiment, the specific implementation of each module can refer to the corresponding method in Embodiment 2, and will not be elaborated here.

[0146] In this embodiment, the emotion recognition module can obtain not only the visual feature information of each frame of image but also the correlation information between each image through the emotion recognition model. Based on the correlation information, the dynamic changes of the target object can be reflected, the changes of emotions can be better reflected, and thus the emotion information of the target object in the real scene can be accurately identified. On the one hand, this embodiment overcomes the conditional limitations of existing emotion recognition algorithms that focus on specific scenarios or single individuals, can identify multiple target objects in a real scene, and then can identify the emotion information of multiple target objects, and can realize real-time dynamic tracking and recognition of the emotions of target objects in the real scene. In addition, since the improved Transformer coding structure is adopted in the emotion recognition model, the accuracy of emotion information recognition is further improved.

[0147] Embodiment 5

[0148] This embodiment provides an electronic device, which can be presented in the form of a computing device (for example, it can be a server device), including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the training method of the emotion recognition model in Embodiment 1 can be implemented.

[0149] Figure 12 The schematic diagram of the hardware structure of this embodiment is shown, as Figure 12 shown, the electronic device 9 specifically includes:

[0150] At least one processor 91, at least one memory 92, and a bus 93 for connecting different system components (including the processor 91 and the memory 92), where:

[0151] The bus 93 includes a data bus, an address bus, and a control bus.

[0152] The memory 92 includes volatile memory, such as random access memory (RAM) 921 and / or cache memory 922, and may further include read-only memory (ROM) 923.

[0153] The memory 92 further includes a program / utilities 925 having a set (at least one) of program modules 924. Such program modules 924 include, but are not limited to: an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include the implementation of a network environment.

[0154] The processor 91 executes various functional applications and data processing by running the computer program stored in the memory 92, such as the training method of the emotion recognition model in Embodiment 1 of the present invention.

[0155] The electronic device 9 can further communicate with one or more external devices 94 (such as a keyboard, a pointing device, etc.). Such communication can be carried out through an input / output (I / O) interface 95. And, the electronic device 9 can also communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through a network adapter 96. The network adapter 96 communicates with other modules of the electronic device 9 through the bus 93. It should be understood that, although not shown in the figure, other hardware and / or software modules can be used in combination with the electronic device 9, including but not limited to: microcode, device drivers, redundant processors, external disk drive arrays, RAID (disk array) systems, tape drives, and data backup storage systems, etc.

[0156] It should be noted that although several units / modules or sub-units / modules of the electronic device are mentioned in the above detailed description, this division is merely exemplary and not mandatory. In fact, according to the embodiments of the present application, the features and functions of two or more units / modules described above can be embodied in one unit / modules. Conversely, the features and functions of one unit / modules described above can be further divided and embodied by multiple units / modules.

[0157] Embodiment 6

[0158] This embodiment provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the training method of the emotion recognition model in Embodiment 1.

[0159] Among them, the more specific forms that the readable storage medium can adopt can include but are not limited to: portable disks, hard disks, random access memories, read-only memories, erasable programmable read-only memories, optical storage devices, magnetic storage devices, or any suitable combination of the above.

[0160] In a possible implementation manner, the present invention can also be implemented in the form of a program product, which includes program code. When the program product runs on a terminal device, the program code is used to cause the terminal device to execute the training method of the emotion recognition model in Embodiment 1.

[0161] Among them, the program code for executing the present invention can be written in any combination of one or more programming languages. The program code can be executed entirely on the user device, partially on the user device, executed as an independent software package, partially on the user device and partially on a remote device, or entirely on a remote device.

[0162] Although the specific implementation manners of the present invention are described above, those skilled in the art should understand that this is only an example, and the protection scope of the present invention is defined by the appended claims. Those skilled in the art can make various changes or modifications to these implementation manners without departing from the principles and essences of the present invention, but these changes and modifications all fall within the protection scope of the present invention.

Claims

1. A training method for an emotion recognition model, characterized in that The training method includes: Obtaining a training set of a target object, where the training set includes a number of video samples, and each video sample includes at least two frames of images; For each video sample, obtaining target detection information of the video sample, where the target detection information includes visual feature information of each frame of image and association information between adjacent images; Using the target detection information as input and the preset emotion information of the target object in each frame of image as output to perform model training to construct the emotion recognition model; The obtaining of the target detection information of the video sample includes the following steps: Inputting the video sample into a visual feature extraction algorithm to obtain visual feature information of each frame of image; Inputting the visual feature information of each frame of image into a target encoding structure to obtain association information between adjacent images; The target encoding structure includes a Transformer encoding structure; The Transformer encoding structure includes a nested residual position encoding structure, and the nested residual position encoding structure includes a position encoding function and N cascaded dropout functions, N≥2, n = 1, 2…N, and n represents the nth dropout function; The inputting of the visual feature information of each frame of image into the target encoding structure to obtain association information between adjacent images includes the following steps: For adjacent images, inputting the visual feature information of each frame of image into the position encoding function to obtain first encoding information; When n = 1, the input of the corresponding dropout function is the first encoding information and the visual feature information of each frame of image; When n = 2,…N, the input of the corresponding dropout function is the output of the previous dropout function and the visual feature information of each frame of image, or the input of the corresponding dropout function is the output of all the previous dropout functions and the visual feature information of each frame of image; When n = N, the output of the corresponding dropout function is the association information between the adjacent images.

2. The training method of the emotion recognition model according to claim 1, characterized in that, The obtaining of the target detection information of the video sample includes: Inputting the video sample into a target detection algorithm to obtain first detection information, where the first detection information includes box information of each frame of image, and the box information is used to represent the region of the target object in the image; Obtaining target detection information according to the first detection information.

3. The training method of the emotion recognition model according to claim 2, wherein The obtaining of the detection information according to the first detection information includes the following steps: Inputting the first detection information into a target tracking algorithm to obtain second detection information, where the second detection information includes the motion trajectory of the target object; Obtaining target detection information according to the first detection information and the second detection information.

4. The training method of the emotion recognition model according to claim 2, wherein The first detection information includes box information of different target objects; The obtaining of the target detection information according to the first detection information includes the following steps: For each target object, intercepting the region corresponding to the target object in each frame of image according to the box information; Obtaining target detection information according to the intercepted image.

5. The training method of the emotion recognition model according to any one of claims 1-4, characterized in that The target object specifically includes the facial expression of the target object and / or the limb movement of the target object.

6. A method for emotion recognition, characterized in that, The emotion recognition method includes the following steps: Obtain the video to be detected of the target object; Input the video to be detected into the emotion recognition model to obtain corresponding emotion information, where the emotion recognition model is a model obtained according to the training method of the emotion recognition model described in any one of claims 1-5.

7. A training system for an emotion recognition model, characterized in that, The training system includes: a training set acquisition module, a detection information acquisition module, and a model construction module; The training set acquisition module is used to obtain the training set of the target object. The training set includes a number of video samples, and each video sample includes at least two frames of images; The detection information acquisition module is used to obtain the target detection information of each video sample. The target detection information includes the visual feature information of each frame of image and the correlation information between adjacent images; The model construction module is used to use the target detection information as input and the preset emotion information of the target object in each frame of image as output for model training to construct the emotion recognition model; The steps for the detection information acquisition module to obtain the target detection information of the video sample include: Input the video sample into the visual feature extraction algorithm to obtain the visual feature information of each frame of image; Input the visual feature information of each frame of image into the target encoding structure to obtain the correlation information between adjacent images; The target encoding structure includes a Transformer encoding structure; the Transformer encoding structure includes a nested residual position encoding structure, and the nested residual position encoding structure includes a position encoding function and N cascaded dropout functions, N≥2, n = 1, 2...N, where n represents the nth dropout function; The steps for inputting the visual feature information of each frame of image into the target encoding structure to obtain the correlation information between adjacent images include: For adjacent images, input the visual feature information of each frame of image into the position encoding function to obtain the first encoded information; When n = 1, the input of the corresponding dropout function is the first encoded information and the visual feature information of each frame of image; When n = 2,...N, the input of the corresponding dropout function is the output of the previous dropout function and the visual feature information of each frame of image, or the input of the corresponding dropout function is the outputs of all the previous dropout functions and the visual feature information of each frame of image; When n = N, the output of the corresponding dropout function is the correlation information between the adjacent images.

8. An emotion recognition system, characterized in that, The emotion recognition system includes: a video acquisition module and an emotion recognition module; The video acquisition module is used to obtain the video to be detected of the target object; The emotion recognition module is used to input the video to be detected into the emotion recognition model to obtain corresponding emotion information, where the emotion recognition model is a model obtained according to the training system of the emotion recognition model described in claim 7.

9. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the training method of the emotion recognition model described in any one of claims 1 to 5 or the emotion recognition method described in claim 6.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the training method of the emotion recognition model according to any one of claims 1 to 5 or the emotion recognition method according to claim 6.

Citation Information

Patent Citations

  • Image recognition method, device and related equipment

    CN108388876A

  • Continuous dimension emotion recognition method based on Transform encoder and multi-head multi-modal attention

    CN113269277A