A multimodal personality recognition method and system based on cross-modal attention mechanism
By adopting the cross-modal attention mechanism and LSTM+Attention layer in multimodal personality analysis, the timing relationship between audio, face and scene image features in video data is solved, and the problem of underutilization of the timing and importance of modal features in the prior art is improved, and the accuracy of personality recognition is improved.
Patent Information
- Application Number
- CN202210159056.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-02-21
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2042-02-21
AI Technical Summary
The existing multimodal personality analysis technology fails to fully consider the timing relationship of modal features and the importance of different modal features, resulting in insufficient feature extraction and fusion, which affects the model prediction effect.
The video data is preprocessed based on the cross-modal attention mechanism, and the audio, face and scene image features are extracted, and the deep feature extraction is performed through bidirectional GRU and cross-modal attention mechanisms. The timing features are extracted using the LSTM+Attention layer, and the personality score is finally calculated through weighted feature fusion.
Effectively extracting the interactive information and timing characteristics between modals improves the model's ability to identify personality characteristics, and solves the problem that modal complementarity and timing relationships are not fully utilized in the prior art.
Smart Images

Figure CN114549946B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of multimodal personality analysis, and more specifically, to a multimodal personality recognition method and system based on a cross-modal attention mechanism. Background Art
[0002] Personality can be defined as the psychological factors that influence individual behavior, thinking, and feeling patterns, thus distinguishing individuals from others. Traditional personality assessment generally requires evaluation through questionnaires, but this is time-consuming. With the development of social networks and multimedia, more and more people upload their videos to the Internet for sharing. Experts can use these videos to judge the personality of the characters in the video. In recent years, automatic personality recognition has gradually become an important research topic. Judging a person's personality through video requires judging through multiple modal information, which is multimodal personality analysis.
[0003] Currently, most multimodal personality analysis works focus on combining different features from different modalities using simple fusion techniques, which are then input into a classifier to obtain predicted personality traits.
[0004] In terms of feature extraction, these works ignore the interactivity between modalities. For example, when a person is very angry, the image modality may be represented by wide-eyed rage, and the sound may be an indignant sound. Other modalities can be used to assist the understanding of the main modality. At the same time, these works also ignore the temporal relationship between modal features, which makes the feature extraction of previous works not comprehensive, affecting the prediction effect of the model. In terms of feature fusion, previous works simply spliced the features of different modalities, ignoring the fact that the importance of different modalities is different.
[0005] The prior art discloses a multimodal emotion recognition method, including a data acquisition device, an output device, and an emotion analysis software system. The emotion analysis software system performs comprehensive analysis and reasoning on the data obtained by the data acquisition device, and finally outputs the result to the output device; the specific steps are: emotion recognition based on facial image expression, emotion recognition based on voice signal, emotion analysis based on text semantics, emotion recognition based on human posture, emotion recognition based on physiological signal, semantic understanding based on multi-round dialogue, and association judgment based on temporal multimodal emotion semantic fusion. Although the prior art is based on multimodality, it is aimed at emotion recognition and does not perform personality analysis. Summary of the invention
[0006] In order to overcome the defects in the above-mentioned prior art that the multimodal personality analysis does not take into account the temporal relationship of modal features and does not consider the different importance levels of modal features when fusing features, the present invention provides a multimodal personality recognition method and system based on a cross-modal attention mechanism.
[0007] The primary purpose of the present invention is to solve the above technical problems. The technical solution of the present invention is as follows:
[0008] The first aspect of the present invention provides a multimodal personality recognition method based on a cross-modal attention mechanism, comprising the following steps:
[0009] S1: Preprocess the video data to obtain the audio file in the video data and the face image and scene image in the video frame;
[0010] S2: Extract sound features from audio files;
[0011] S3: extract image features from face images and scene images respectively;
[0012] S4: Use cross-modal attention mechanism to perform deep feature extraction on the extracted sound features and image features;
[0013] S5: Perform weighted feature fusion on the deep features of different modalities, calculate the personality score using the preset fully connected layer, and obtain the personality result;
[0014] S6: Divide the pre-prepared video data into a training set, a validation set, and a test set, repeat steps S1-S5 for iterative training, use the validation set to validate the trained model, and save the model with the best validation effect for personality recognition.
[0015] Furthermore, the specific process of step S1 is as follows:
[0016] S101: Using a video editing tool to read video data and save the audio in the video in wav format;
[0017] S102: using an open source machine vision library to read each frame of the video, among all the frames read, at a fixed interval, randomly selecting one frame in each sub-interval as a scene image, and converting the obtained scene image into a preset size;
[0018] S103: using an open source face recognition model to identify a face image from the scene image, marking the face area, and converting the face image into a preset size.
[0019] Furthermore, the preset size is 112*112, and the scene image and the face image are both 3-channel images.
[0020] Furthermore, the specific process of step S4 is as follows:
[0021] S401: respectively applying the three modal features of sound feature, face image feature and scene image feature to obtain contextual feature representations of the three modal features through a bidirectional GRU;
[0022] S402: extracting features from the contextual feature representations of the three modal features using a cross-modal attention mechanism;
[0023] S403: Extracting temporal features from each modality feature extracted by the cross-modal attention mechanism through the LSTM+attention layer.
[0024] Furthermore, the three modal features of sound features, face image features, and scene image features are respectively expressed through a bidirectional GRU to obtain the contextual feature representation of the three modal features, as shown in the following expression:
[0025] X sence =BiGRU(s1,s2,s3,……,s t )
[0026] X face =BiGRU(f1,f2,f3,……,f t )
[0027] X audio =BiGRU(a1,a2,a3,……,a t )
[0028] Among them, BiGRU is a bidirectional GRU network, s1~s t 、f1~f t are the scene feature sequence and face feature sequence extracted by S3, a1~a t is the sound feature sequence extracted by S2, X sence , X face , X audio They are respectively scenes, faces and sound feature sequences represented in context.
[0029] Furthermore, the mathematical expression of the cross-modal attention mechanism is as follows:
[0030]
[0031] W f =γ·α m +(1-γ)·β a
[0032] W m =Softmax(W f )
[0033] X Att =Wm X m
[0034] Among them, α m represents the attention matrix of the main modality, β a represents the attention matrix of the auxiliary modality, W f represents the attention matrix after the hyperparameter r modulation, W m The weight matrix after Softmax activation, Q m and K m represents the characteristic sequence of the main mode, Q a and K a represents the feature sequence of the auxiliary modality, tanh represents the tangent function activation, γ represents the weight introduced by the auxiliary modality, X Att Represents the feature sequence obtained after the cross-modal attention mechanism.
[0035] Furthermore, the expression for extracting time series features through the LSTM+attention layer is as follows:
[0036] O t , H t =biLSTM(X att )
[0037]
[0038] W l =Softmax(W t )
[0039] Z=W l ·O t
[0040] Among them, O t , H t They are represented as the last layer output and all hidden layer outputs of LSTM, W t Represents the attention matrix of temporal features, W l It is represented by the weight corresponding to each hidden layer feature, and Z represents the weighted sequence feature, that is, the final feature extracted from each modality.
[0041] Furthermore, the specific process of step S5 is as follows:
[0042] S501: splicing the features of the three modes extracted in step S4;
[0043] S502: Obtain a weight vector through two layers of full connection activation and Softmax activation;
[0044] S503: Multiply the concatenated features and weight vectors and input them into a preset fully connected layer to output a predicted personality score.
[0045] Furthermore, the mathematical expressions included in step S5 are:
[0046] F=Cat[Z a ,Z f ,Z s ]
[0047] a=tanh(Vtanh(WF+b)+c)
[0048]
[0049] Among them, F represents the concatenated multimodal feature, a represents the weight vector of each dimension of the multimodal feature F, and Z a、 Z f、 Z s are the extracted modal features of sound, face, and scene, respectively. represents the weighted multimodal features, Represents the fused features after Softmax normalization.
[0050] A second aspect of the present invention provides a multimodal personality recognition system based on a cross-modal attention mechanism, the system comprising: a memory, a processor, the memory comprising a multimodal personality recognition method program based on a cross-modal attention mechanism, the multimodal personality recognition method program based on a cross-modal attention mechanism implementing the following steps when executed by the processor:
[0051] S1: Preprocess the video data to obtain the audio file in the video data and the face image and scene image in the video frame;
[0052] S2: Extract sound features from audio files;
[0053] S3: extract image features from face images and scene images respectively;
[0054] S4: Use cross-modal attention mechanism to perform deep feature extraction on the extracted sound features and image features;
[0055] S5: Perform weighted feature fusion on the deep features of different modalities, calculate the personality score using the preset fully connected layer, and obtain the personality result;
[0056] S6: Divide the pre-prepared video data into a training set, a validation set, and a test set, repeat steps S1-S5 for iterative training, use the validation set to validate the trained model, and save the model with the best validation effect for personality recognition.
[0057] Compared with the prior art, the technical solution of the present invention has the following beneficial effects:
[0058] The present invention utilizes the cross-modal attention mechanism when extracting modal features, fully extracts the interactive information between modalities, and helps the main modality to be understood through the auxiliary modality, thus solving the problem that predecessors did not consider the complementarity between modalities, making the model's single-modal extraction capability stronger. At the same time, the LSTM+Attention mechanism is used in the feature extraction part to effectively extract the temporal features within the modality, solving the problem that the prior art does not consider the temporal features of the modality. In addition, during feature fusion, the fusion network based on the attention mechanism is used to weight the features of different modalities, so that the model pays more attention to important modalities and makes full use of complementary information. BRIEF DESCRIPTION OF THE DRAWINGS
[0059] Figure 1 This is a flow chart of a multimodal personality recognition method based on a cross-modal attention mechanism of the present invention.
[0060] Figure 2 Schematic diagram of the structure of the multimodal personality recognition model corresponding to the method of the present invention.
[0061] Figure 3 2 is a structural diagram of the cross-modal attention mechanism in an embodiment of the present invention.
[0062] Figure 4 Schematic diagram of the timing feature extraction structure of an embodiment of the present invention.
[0063] Figure 5 It is a schematic diagram of the feature fusion structure of an embodiment of the present invention. DETAILED DESCRIPTION
[0064] In order to more clearly understand the above-mentioned purpose, features and advantages of the present invention, the present invention is further described in detail below in conjunction with the accompanying drawings and specific embodiments. It should be noted that the embodiments of the present application and the features in the embodiments can be combined with each other without conflict.
[0065] In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention may also be implemented in other ways different from those described herein. Therefore, the protection scope of the present invention is not limited to the specific embodiments disclosed below.
[0066] Example 1
[0067] like Figure 1-2 As shown, the first aspect of the present invention provides a multimodal personality recognition method based on a cross-modal attention mechanism, comprising the following steps:
[0068] S1: Preprocess the video data to obtain the audio file in the video data and the face image and scene image in the video frame;
[0069] It should be noted that in the present invention, the video data is first preprocessed, and the preprocessing includes separating the audio file and the image from the video data, and the image includes a face image and a scene image. The specific steps are:
[0070] S101: Using a video editing tool to read video data and save the audio in the video in wav format;
[0071] S102: using an open source machine vision library to read each frame of the video, among all the frames read, at a fixed interval, randomly selecting one frame in each sub-interval as a scene image, and converting the obtained scene image into a preset size;
[0072] S103: using an open source face recognition model to identify a face recognition image from the scene image, marking the face area, and converting the face image into a preset size.
[0073] It should be noted that in a specific embodiment, the moviepy tool can be used to read the video. In step S102, the open source machine vision library can be the Opencv vision library. Generally, 32 frames of scene images are selected, and the size of the 32 frames of scene images is converted to 112*112. Then, an open source face recognition model, such as the face recognition model of Opencv, is used to identify the face image from the scene image, and the face area is marked, and the face image size is converted to 112*112. The scene image and the face image are both 3-channel images.
[0074] S2: Extract sound features from audio files;
[0075] In a specific embodiment, the separated audio file is first divided into 32 sound segments, and wav2vec2.0 is used to extract features from each sound segment to extract 32 512-dimensional features.
[0076] S3: extract image features from face images and scene images respectively;
[0077] In a specific embodiment, a 512-dimensional feature is extracted from each acquired face image and scene image through Resnet34, and 32 512-dimensional features are extracted from both the face and scene modalities.
[0078] S4: Use cross-modal attention mechanism to perform deep feature extraction on the extracted sound features and image features;
[0079] The specific process is:
[0080] S401: respectively applying the three modal features of sound feature, face image feature and scene image feature to obtain contextual feature representations of the three modal features through a bidirectional GRU;
[0081] It should be noted that the mathematical expression of the contextual feature representation of the three modal features obtained by the bidirectional GRU is:
[0082] X sence =BiGRU(s1,s2,s3,……,s t )
[0083] X face =BiGRU(f1,f2,f3,……,f t )
[0084] X audio =BiGRU(a1,a2,a3,……,a t )
[0085] Among them, BiGRU is a bidirectional GRU network, s1~s t 、f1~f t are the extracted scene feature sequence and face feature sequence, a1~a t is the sound feature sequence extracted by S2, X sence , X face , X audio They are respectively scenes, faces and sound feature sequences represented in context.
[0086] like Figure 3 As shown, S402: representing the contextual features of the three modal features, and extracting features through a cross-modal attention mechanism;
[0087] The mathematical expression of the cross-modal attention mechanism is as follows:
[0088]
[0089] W f =γ·α m +(1-γ)·β a
[0090] W m =Softmax(W f )
[0091] X Att =W m X m
[0092] Among them, α m represents the attention matrix of the main modality, β a represents the attention matrix of the auxiliary modality, W f represents the attention matrix after the hyperparameter r modulation, W m The weight matrix after Softmax activation, Qm and K m represents the characteristic sequence of the main mode, Q a and K a represents the feature sequence of the auxiliary modality, tanh represents the tangent function activation, γ represents the weight introduced by the auxiliary modality, X Att Represents the feature sequence obtained after the cross-modal attention mechanism.
[0093] It should be noted that each modality has an auxiliary modality to help better understand it. For sound feature extraction, the auxiliary modality is face. For face and scene extraction, the auxiliary modality is sound.
[0094] like Figure 4 As shown, S403: extracting temporal features from each modality feature extracted by the cross-modal attention mechanism through the LSTM+attention layer.
[0095] The time series features are extracted through the LSTM+attention layer, and its mathematical expression is as follows:
[0096] O t , H t =biLSTM(X att )
[0097]
[0098] W l =Softmax(W t )
[0099] Z=W l ·O t
[0100] Among them, O t , H t They are represented as the last layer output and all hidden layer outputs of LSTM, W t Represents the attention matrix of temporal features, W l It is represented by the weight corresponding to each hidden layer feature, and Z represents the weighted sequence feature, that is, the final feature extracted from each modality.
[0101] like Figure 5 As shown, S5: weighted feature fusion is performed on the deep features of different modalities, and the personality score is calculated using the preset fully connected layer to obtain the personality result;
[0102] The specific process is:
[0103] S501: splicing the three modal features extracted in step S4;
[0104] S502: The three modal features are activated by two layers of full connection and Softmax to obtain weight vectors respectively;
[0105] S503: Multiply the concatenated features and the corresponding weight vector and input them into a preset fully connected layer to output the predicted personality score.
[0106] It should be noted that the three modal features are first concatenated using the Cat() function, and then the concatenated features are activated through two layers of full connection and Softmax activation to obtain a weight vector, the concatenated modal features are multiplied by the weight vector, and then input into the preset fully connected layer to output the predicted personality score, wherein the preset fully connected layer includes 5 fully connected layers, each fully connected layer outputs a personality score, that is, a total of 5 personality scores are output.
[0107] The formulas included in step S5 are:
[0108] F=Cat[Z a ,Z f ,Z s ]
[0109] a=tanh(Vtanh(WF+b)+c)
[0110]
[0111] Among them, F represents the concatenated multimodal feature, a represents the weight vector of each dimension of the multimodal feature F, and Z a、 Z f、 Z s are the extracted modal features of sound, face, and scene, respectively. represents the weighted multimodal features, Represents the fused features after Softmax normalization.
[0112] S6: Divide the pre-prepared video data into a training set, a validation set and a test set, repeat steps S1-S5 for iterative training, use the validation set to validate the trained model, save the model with the best validation effect for the bimodal classification task, and use the test set to test the validated model.
[0113] A second aspect of the present invention provides a multimodal personality recognition system based on a cross-modal attention mechanism, the system comprising: a memory, a processor, the memory comprising a multimodal personality recognition method program based on a cross-modal attention mechanism, the multimodal personality recognition method program based on a cross-modal attention mechanism implementing the following steps when executed by the processor:
[0114] S1: Preprocess the video data to obtain the audio file in the video data and the face image and scene image in the video frame;
[0115] It should be noted that in the present invention, the video data is first preprocessed, and the preprocessing includes separating the audio file and the image from the video data, and the image includes a face image and a scene image. The specific steps are:
[0116] S101: Using a video editing tool to read video data and save the audio in the video in wav format;
[0117] S102: using an open source machine vision library to read each frame of the video, among all the frames read, at a fixed interval, randomly selecting one frame in each sub-interval as a scene image, and converting the obtained scene image into a preset size;
[0118] S103: using an open source face recognition model to identify a face recognition image from the scene image, marking the face area, and converting the face image into a preset size.
[0119] It should be noted that in a specific embodiment, the moviepy tool can be used to read the video. In step S102, the open source machine vision library can be the Opencv vision library. Generally, 32 frames of scene images are selected, and the size of the 32 frames of scene images is converted to 112*112. Then, an open source face recognition model, such as the face recognition model of Opencv, is used to identify the face image from the scene image, and the face area is marked, and the face image size is converted to 112*112. The scene image and the face image are both 3-channel images.
[0120] S2: Extract sound features from audio files;
[0121] In a specific embodiment, the separated audio file is first divided into 32 sound segments, and wav2vec2.0 is used to extract features from each sound segment to extract 32 512-dimensional features.
[0122] S3: extract image features from face images and scene images respectively;
[0123] In a specific embodiment, a 512-dimensional feature is extracted from each acquired face image and scene image through Resnet34, and 32 512-dimensional features are extracted from both the face and scene modalities.
[0124] S4: Use cross-modal attention mechanism to perform deep feature extraction on the extracted sound features and image features;
[0125] The specific process is:
[0126] S401: respectively applying the three modal features of sound feature, face image feature and scene image feature to obtain contextual feature representations of the three modal features through a bidirectional GRU;
[0127] It should be noted that the mathematical expression of the contextual feature representation of the three modal features obtained by the bidirectional GRU is:
[0128] X sence =BiGRU(s1,s2,s3,……,s t )
[0129] X face =BiGRU(f1,f2,f3,……,f t )
[0130] X audio =BiGRU(a1,a2,a3,……,a t )
[0131] Among them, BiGRU is a bidirectional GRU network, s1~s t 、f1~f t are the extracted scene feature sequence and face feature sequence, a1~a t is the sound feature sequence extracted by S2, X sence , X face , X audio They are respectively scenes, faces and sound feature sequences represented in context.
[0132] S402: Represent the contextual features of the three modal features and extract features through a cross-modal attention mechanism;
[0133] The mathematical expression of the cross-modal attention mechanism is as follows:
[0134]
[0135] W f =γ·α m +(1-γ)·β a
[0136] W m =Softmax(W f )
[0137] X Att =W m X m
[0138] Among them, α m represents the attention matrix of the main modality, β a represents the attention matrix of the auxiliary modality, W frepresents the attention matrix after the hyperparameter γ modulation, W m The weight matrix after Softmax activation, Q m and K m represents the characteristic sequence of the main mode, Q a and K a represents the feature sequence of the auxiliary modality, tanh represents the tangent function activation, γ represents the weight introduced by the auxiliary modality, X Att Represents the feature sequence obtained after the cross-modal attention mechanism.
[0139] It should be noted that each modality has an auxiliary modality to help better understand it. For sound feature extraction, the auxiliary modality is face. For face and scene extraction, the auxiliary modality is sound.
[0140] S403: Extracting temporal features from each modality feature extracted by the cross-modal attention mechanism through the LSTM+attention layer.
[0141] The time series features are extracted through the LSTM+attention layer, and its mathematical expression is as follows:
[0142] O t , H t =biLSTM(X att )
[0143]
[0144] W l =Softmax(W t )
[0145]
[0146] Among them, O t , H t They are represented as the last layer output and all hidden layer outputs of LSTM, W t Indicates that W l It is represented by the weight corresponding to each hidden layer feature, and Z represents the weighted sequence feature, that is, the final feature extracted from each modality.
[0147] S5: Perform weighted feature fusion on the deep features of different modalities, calculate the personality score using the preset fully connected layer, and obtain the personality result;
[0148] The specific process is:
[0149] S501: splicing the features of the three modes extracted in step S4;
[0150] S502: The three modal features are activated by two layers of full connection and Softmax to obtain weight vectors respectively;
[0151] S503: Multiply the concatenated features and the corresponding weight vector and input them into a preset fully connected layer to output the predicted personality score.
[0152] It should be noted that the three modal features are first concatenated using the Cat() function, and then the concatenated features are activated through two layers of full connection and Softmax activation to obtain a weight vector, the concatenated modal features are multiplied by the weight vector, and then input into the preset fully connected layer to output the predicted personality score, wherein the preset fully connected layer includes 5 fully connected layers, each fully connected layer outputs a personality score, that is, a total of 5 personality scores are output.
[0153] The formulas included in step S5 are:
[0154] F=Cat[Z a ,Z f ,Z s ]
[0155] a=tanh(Vtanh(WF+b)+c)
[0156]
[0157] Among them, F represents the concatenated multimodal feature, a represents the weight vector of each dimension of the multimodal feature F, and Z a、 Z f、 Z s are the extracted modal features of sound, face, and scene, respectively. represents the weighted multimodal features, Represents the fused features after Softmax normalization.
[0158] S6: Divide the pre-prepared video data into a training set, a validation set, and a test set, repeat steps S1-S5 for iterative training, use the validation set to verify the trained model, and save the model with the best verification effect for the bimodal classification task.
[0159] Example 3
[0160] This example uses specific data for verification analysis. This example uses the Firstimpression dataset of Chalearn, which consists of a training set and a test set. The training set has 6,000 samples, and 600 samples are extracted from the training set as a verification set at a ratio of 9:1. The test set has 2,000 samples. Each sample consists of a 15-second video of a single person and a score of five personality traits, and the personality traits are manually annotated by Amazon Mechanical Turk (AMT).
[0161] First, the video data is preprocessed to separate and extract 32 frames of face images, scene images and 32 audio files. Then, the pre-trained model is used to extract features of each modality. The cross-modal attention mechanism is used to extract deep features. The LSTM is used to extract temporal features. Finally, weighted fusion is performed and linear layer regression is used to obtain the scores of 5 personalities. The details are as follows:
[0162] S1: Video data preprocessing. Video data preprocessing includes sound separation, image frame separation, extracting the audio file of the sound from the video, dividing the video into 32 time periods, randomly extracting a frame from each time period, and extracting the face image from each extracted frame using OpenCV.
[0163] S2: Preliminary extraction of sound features: For the sound data, the sound is first divided into 32 sound segments, and wav2vec2.0 is used to extract features for each sound segment, extracting 32 512-dimensional features.
[0164] S3: Preliminary extraction of image features. For each face image and scene image obtained from S1, a 512-dimensional feature is extracted through Resnet34, and finally 32 512-dimensional features are extracted from the two modalities respectively.
[0165] S4: Deeper extraction through cross-modal attention mechanism. The sequence features initially extracted from the three modalities (sound, face image, scene image) are sequentially passed through the bidirectional GRU to obtain better context feature representation, the interactive information between the modalities is obtained through the cross-modal attention mechanism, the temporal features are extracted through the LSTM+Attention mechanism, and the sequence features are integrated.
[0166] S5: The features extracted from different modalities are fused through an attention-based weighted fusion mechanism, and then the scores of the five personalities are calculated through a linear layer to obtain the personality results.
[0167] S6: Divide the processed data into a training set, a validation set, and a test set. Perform multiple iterations of training on the above model, and use the best model in the validation set for the bimodal classification task. Specifically, perform 100 epochs of iteration training on the model part of steps S2 to S5, record the best model in the validation set, and use it for testing the test set.
[0168] In order to compare with existing methods, the specific results are evaluated by absolute value error. In order to make the results appear as large as possible, the evaluation index is (average absolute value error of 1-5 personalities). The specific results are shown in the following table. Table 1 is a comparison table of personality recognition results of different methods.
[0169]
[0170] It can be seen from the above table that the present invention has obvious improvements over other methods and has reached a good level of the current data set.
[0171] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the present invention, and are not intended to limit the embodiments of the present invention. For those skilled in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to list all the embodiments here. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the protection scope of the claims of the present invention.
Claims
1. A multimodal personality recognition method based on a cross-modal attention mechanism, characterized in that: The following steps are involved: S1: Preprocess the video data to obtain the audio file in the video data and the face image and scene image in the video frame; S2: Extract sound features from audio files; S3: extract image features from face images and scene images respectively; S4: Deep feature extraction of extracted sound features and image features using cross-modal attention mechanism; The specific process in step S4 is: S401: respectively applying the three modal features of sound feature, face image feature and scene image feature to obtain contextual feature representations of the three modal features through a bidirectional GRU; S402: extracting features from the contextual feature representations of the three modal features using a cross-modal attention mechanism; S403: extracting temporal features from each modality feature extracted by the cross-modal attention mechanism through the LSTM+attention layer; The mathematical expression of the cross-modal attention mechanism is as follows: W f =c·a m +(1-c)·b a W m =Softmax(W f ) X Att =W m X m Among them, α m represents the attention matrix of the main modality, β a represents the attention matrix of the auxiliary modality, W f represents the attention matrix after the hyperparameter γ modulation, W m The weight matrix after Softmax activation, Q m and K m represents the characteristic sequence of the main mode, Q a and K a represents the feature sequence of the auxiliary modality, tanh represents the tangent function activation, γ represents the weight introduced by the auxiliary modality, X Att Represents the feature sequence obtained after the cross-modal attention mechanism; The mathematical expression for extracting time series features through the LSTM+attention layer is as follows: O t ,H t =biLSTM(X Att ) W l =Softmax(W t ) W=W l ·ABOUT t Among them, O t , H t They are represented as the last layer output and all hidden layer outputs of LSTM, W t Represents the attention matrix of temporal features, W l It is represented by the weight corresponding to each hidden layer feature, and Z represents the weighted sequence feature, that is, the final feature extracted from each modality; S5: Perform weighted feature fusion on the deep features of different modalities, calculate the personality score using the preset fully connected layer, and obtain the personality result; The specific process in step S5 is: S501: splicing the features of the three modes extracted in step S4; S502: Obtain a weight vector through two layers of full connection activation and Softmax activation; S503: multiplying the concatenated features and the weight vector and inputting the result into a preset fully connected layer, and outputting a predicted personality score; S6: Divide the pre-prepared video data into a training set, a validation set, and a test set, repeat steps S1-S5 for iterative training, use the validation set to validate the trained model, and save the model with the best validation effect for personality recognition.
2. According to claim 1, a multimodal personality recognition method based on a cross-modal attention mechanism is characterized in that: The specific process of step S1 is: S101: Using a video editing tool to read video data and save the audio in the video in wav format; S102: using an open source machine vision library to read each frame of the video, among all the frames read, at a fixed interval, randomly selecting one frame in each sub-interval as a scene image, and converting the obtained scene image into a preset size; S103: using an open source face recognition model to identify a face image from the scene image, marking the face area, and converting the face image into a preset size.
3. According to claim 2, a multimodal personality recognition method based on a cross-modal attention mechanism is characterized in that: The preset size is 112*112, and both the scene image and the face image are 3-channel images.
4. According to claim 1, a multimodal personality recognition method based on a cross-modal attention mechanism is characterized in that: The three modal features of sound features, face image features, and scene image features are respectively expressed through a bidirectional GRU to obtain the contextual feature representation of the three modal features. The expression is: X sence =BiGRU(s1,s2,s3,……,s t ) <h2 style=";text-align:left;direction:ltr">X<h2 style=";text-align:left;direction:ltr"> face <h2 style=";text-align:left;direction:ltr"> =BiGRU(f1,f2,f3,……,f<h2 style=";text-align:left;direction:ltr"> t <h2 style=";text-align:left;direction:ltr"> ) <h2 style=";text-align:left;direction:ltr">X<h2 style=";text-align:left;direction:ltr"> audio <h2 style=";text-align:left;direction:ltr"> =BiGRU(a1,a2,a3,……,a<h2 style=";text-align:left;direction:ltr"> t <h2 style=";text-align:left;direction:ltr"> ) Among them, BiGRU is a bidirectional GRU network, s1~s t 、f1~f t are the extracted scene feature sequence and face feature sequence, a1~a t is the sound feature sequence extracted by S2, X sence , X face , X audio They are respectively scenes, faces and sound feature sequences represented in context.
5. The multimodal personality recognition method based on cross-modal attention mechanism according to claim 1, characterized in that: The mathematical expressions included in step S5 are: F=Cat[Z a ,Z f ,Z s ] a=tanh(Vtanh(W·F+b)+c) Among them, F represents the concatenated multimodal feature, a represents the weight vector of each dimension of the multimodal feature F, and Z a、 Z f、 Z s are the extracted modal features of sound, face, and scene, respectively. represents the weighted multimodal features, Represents the fused features after Softmax normalization.
6. A multimodal personality recognition system based on a cross-modal attention mechanism, characterized in that: The system includes: a memory and a processor, wherein the memory includes a multimodal personality recognition method program based on a cross-modal attention mechanism, and when the multimodal personality recognition method program based on a cross-modal attention mechanism is executed by the processor, the following steps are implemented: S1: Preprocess the video data to obtain the audio file in the video data and the face image and scene image in the video frame; S2: Extract sound features from audio files; S3: extract image features from face images and scene images respectively; S4: Deep feature extraction of extracted sound features and image features using cross-modal attention mechanism; The specific process in step S4 is: S401: respectively applying the three modal features of sound feature, face image feature and scene image feature to obtain contextual feature representations of the three modal features through a bidirectional GRU; S402: extracting features from the contextual feature representations of the three modal features using a cross-modal attention mechanism; S403: extracting temporal features from each modality feature extracted by the cross-modal attention mechanism through the LSTM+attention layer; The mathematical expression of the cross-modal attention mechanism is as follows: W f =c·a m +(1-c)·b a W m =Softmax(W f ) X Att =W m X m Among them, α m represents the attention matrix of the main modality, β a represents the attention matrix of the auxiliary modality, W f represents the attention matrix after the hyperparameter γ modulation, W m The weight matrix after Softmax activation, Q m and K m represents the characteristic sequence of the main mode, Q a and K a represents the feature sequence of the auxiliary modality, tanh represents the tangent function activation, γ represents the weight introduced by the auxiliary modality, X Att Represents the feature sequence obtained after the cross-modal attention mechanism; The mathematical expression for extracting time series features through the LSTM+attention layer is as follows: O t ,H t =biLSTM(X Att ) W l =Softmax(W t ) W=W l ·ABOUT t Among them, O t , H t They are represented as the last layer output and all hidden layer outputs of LSTM, W t Represents the attention matrix of temporal features, W l It is represented by the weight corresponding to each hidden layer feature, and Z represents the weighted sequence feature, that is, the final feature extracted from each modality; S5: Perform weighted feature fusion on the deep features of different modalities, calculate the personality score using the preset fully connected layer, and obtain the personality result; The specific process in step S5 is: S501: splicing the features of the three modes extracted in step S4; S502: Obtain a weight vector through two layers of full connection activation and Softmax activation; S503: multiplying the concatenated features and the weight vector and inputting the result into a preset fully connected layer, and outputting a predicted personality score; S6: Divide the pre-prepared video data into a training set, a validation set, and a test set, repeat steps S1-S5 for iterative training, use the validation set to validate the trained model, and save the model with the best validation effect for personality recognition.
Citation Information
Patent Citations
Temporal semantic fusion association determining sub-system based on multimodal emotion recognition system
CN108805087A
Multi-modal emotion recognition method
CN112559835A