Alignment personality recognition model training method and device based on multi-modal scene
By extracting and aligning video frame images and face images in multi-modal scenes, and aligning and fusion of transcription text and audio features, an efficient alignment personality recognition model was trained, solving the problem of the complementarity of facial and panoramic features in the existing technology, and significantly improving the accuracy and efficiency of personality prediction.
Patent Information
- Application Number
- CN202510010605.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-03
- Publication Date
- 2025-05-06
AI Technical Summary
The prior art fails to effectively utilize the potential complementarity between facial features and panoramic features in personality prediction, and lacks an efficient modal fusion mechanism, resulting in the potential complementarity information between modals being underutilized.
A training method for aligned personality recognition model based on multimodal scenes is proposed. By extracting video frame images and face images from multiple scales, and aligning processing is performed, combining transcription text and audio features to train an efficient alignment personality recognition model.
The accuracy and efficiency of personality prediction are significantly improved, and the robustness of the model in complex data environments is enhanced by deeply understanding the unique personality traits of each modal and effectively fusion.
Smart Images

Figure CN119939250A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present disclosure relate to the field of computers, and in particular to a method and device for training an aligned personality recognition model based on a multimodal scenario. Background Art
[0002] In terms of visual information processing, current personality assessment research and strategies often fail to effectively tap into the potential complementarity between facial features and panoptic features. Facial features carry rich emotional expressions and subtle expression changes, while panoptic features reveal the behavioral patterns and interactive contexts of individuals in their environment. Although each of these two features provides a unique perspective for personality prediction, many studies focus only on the analysis of a single feature type, or when attempting to fuse these features, only basic feature cascades or simple weighting methods are adopted. Such methods lack in-depth exploration of the complex associations and complementarities between the two features. In addition, a common problem in the field of personality prediction research is the lack of an efficient modality fusion mechanism. Existing fusion methods fail to fully utilize the unique personality information provided by each modality, and fail to effectively identify and enhance the complementarity and synergy of personality information between different modalities. An ideal modality fusion framework should be able to deeply understand the unique personality traits carried by each modality and effectively integrate them into the overall personality assessment analysis.
[0003] However, most current methods fail to meet this requirement, resulting in the underutilization of potential complementary information between modalities. In addition to mining the unique information of each modality, effective modality fusion also requires identifying and optimizing the commonalities between different modalities. Current research generally fails to establish a balance mechanism to simultaneously emphasize the commonalities and individuality between modalities, which affects the accuracy and generalization ability of personality prediction results.
[0004] The above information disclosed in this Background section is only for enhancement of understanding of the background of the inventive concept and therefore it may contain information that does not form the prior art that is already known in this country to a person of ordinary skill in the art. Summary of the invention
[0005] The content of this disclosure is used to introduce concepts in a brief form, which will be described in detail in the detailed implementation section below. The content of this disclosure is not intended to identify the key features or essential features of the technical solution claimed for protection, nor is it intended to limit the scope of the technical solution claimed for protection.
[0006] Some embodiments of the present disclosure propose an aligned personality recognition model training method, device, electronic device, and computer-readable medium based on a multimodal scenario to solve one or more of the technical problems mentioned in the above background technology section.
[0007] In a first aspect, some embodiments of the present disclosure provide a method for training an aligned personality recognition model based on a multimodal scenario, the method comprising: performing video frame extraction on a user video to obtain a video frame image sequence; performing multi-scale feature extraction on each frame image in the video frame image sequence to generate multi-scale video frame image features to obtain a multi-scale video frame image feature sequence; performing face capture on a face in each video frame image in the video frame image sequence to generate a face image to obtain a face image sequence; performing multi-scale feature extraction on each face image in the face image sequence to generate multi-scale face image features to obtain a multi-scale face image feature sequence; aligning the multi-scale video frame image feature sequence with the multi-scale face image feature sequence to obtain a visual alignment information sequence, wherein the multi-scale video frame image features The multi-scale video frame image features in the feature sequence correspond to the multi-scale face image features in the multi-scale face image feature sequence; text feature extraction is performed on the transcribed text corresponding to the user video to obtain a transcribed text feature sequence; audio feature extraction is performed on the audio corresponding to the user video to obtain an audio feature sequence; the transcribed text feature sequence and the audio feature sequence are aligned and fused to obtain a non-visual alignment information sequence, wherein the transcribed text features in the transcribed text feature sequence correspond to the audio features in the audio feature sequence; an initial aligned personality recognition model is trained according to the visual alignment information sequence and the non-visual alignment information sequence to obtain a trained aligned personality recognition model, wherein the visual alignment information in the visual alignment information sequence corresponds to the non-visual alignment information in the non-visual alignment information sequence.
[0008] In a second aspect, some embodiments of the present disclosure provide an aligned personality recognition model training device based on a multimodal scenario, the device comprising: an extraction unit, configured to perform video frame extraction on a user video to obtain a video frame image sequence; a feature extraction unit, configured to perform multi-scale feature extraction on each frame image in the video frame image sequence to generate multi-scale video frame image features to obtain a multi-scale video frame image feature sequence; a capture unit, configured to perform face capture on a face in each video frame image in the video frame image sequence to generate a face image to obtain a face image sequence; a face extraction unit, configured to perform multi-scale feature extraction on each face image in the face image sequence to generate a multi-scale face image feature to obtain a multi-scale face image feature sequence; a first alignment unit, configured to perform alignment processing on the multi-scale video frame image feature sequence with the multi-scale face image feature sequence to obtain a visual alignment information sequence, wherein the multi-scale video The multi-scale video frame image features in the frame image feature sequence correspond to the multi-scale face image features in the multi-scale face image feature sequence; the text extraction unit is configured to perform text feature extraction on the transcribed text corresponding to the user video to obtain a transcribed text feature sequence; the audio extraction unit is configured to perform audio feature extraction on the audio corresponding to the user video to obtain an audio feature sequence; the second alignment unit is configured to perform alignment and fusion processing on the transcribed text feature sequence and the audio feature sequence to obtain a non-visual alignment information sequence, wherein the transcribed text features in the transcribed text feature sequence correspond to the audio features in the audio feature sequence; the training unit is configured to train the initial aligned personality recognition model according to the visual alignment information sequence and the non-visual alignment information sequence to obtain a trained aligned personality recognition model, wherein the visual alignment information in the visual alignment information sequence corresponds to the non-visual alignment information in the non-visual alignment information sequence.
[0009] In a third aspect, some embodiments of the present disclosure provide an electronic device comprising: one or more processors; a storage device on which one or more programs are stored, and when the one or more programs are executed by the one or more processors, the one or more processors implement the method described in any implementation manner of the above-mentioned first aspect.
[0010] In a fourth aspect, some embodiments of the present disclosure provide a computer-readable medium having a computer program stored thereon, wherein when the program is executed by a processor, the method described in any implementation manner of the above-mentioned first aspect is implemented.
[0011] The above-mentioned embodiments of the present disclosure have the following beneficial effects: through the aligned personality recognition model training method based on multimodal scenarios of some embodiments of the present disclosure, a novel aligned personality prediction network based on multimodal scenarios is proposed. This network design not only solves the limitations of the general network structure in processing specific modal features, but also optimizes the information fusion strategy between different modalities in a targeted manner, thereby significantly improving the accuracy and efficiency of personality prediction. Different from the previous method of using a single generalized network to simultaneously process multiple modal information, the triple alignment network structure of this study specifically designs customized processing strategies for different modalities. First, in terms of visual information processing, the network fully explores and utilizes the complementarity of these two types of visual information by carefully combining panoramic video sequences and facial frame information. In the processing of non-visual information, this study also adopts targeted feature extraction and alignment strategies for text modality and audio modality. Finally, this study effectively integrates visual and non-visual information through an innovative cross-modal interaction mechanism. This cross-modal interaction design not only improves the model's ability to process multimodal data, but also enhances the model's robustness in complex data environments. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] The above and other features, advantages and aspects of the embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the accompanying drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and that components and elements are not necessarily drawn to scale.
[0013] Figure 1 is a flowchart of some embodiments of the method for training an aligned personality recognition model based on a multimodal scenario according to the present disclosure;
[0014] Figure 2 It is a structural schematic diagram of some embodiments of the apparatus for training an aligned personality recognition model based on a multimodal scenario according to the present disclosure;
[0015] Figure 3 is a schematic diagram of the structure of an electronic device suitable for implementing some embodiments of the present disclosure;
[0016] Figure 4 is a model training schematic diagram of the method for training an aligned personality recognition model based on a multimodal scenario according to the present disclosure;
[0017] Figure 5 It is a structural schematic diagram of a visual modality alignment structure in the method for training an aligned personality recognition model based on a multimodal scenario according to the present disclosure;
[0018] Figure 6It is an alignment structure diagram of visual alignment in the alignment personality recognition model training method based on a multimodal scenario according to the present disclosure;
[0019] Figure 7 It is a structural schematic diagram of an interactive module feature enhancer in the aligned personality recognition model training method based on a multimodal scenario according to the present disclosure. DETAILED DESCRIPTION
[0020] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as being limited to the embodiments set forth herein. On the contrary, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not intended to limit the scope of protection of the present disclosure.
[0021] It should also be noted that, for ease of description, only the parts related to the invention are shown in the drawings. In the absence of conflict, the embodiments and features in the embodiments of the present disclosure can be combined with each other.
[0022] It should be noted that the concepts such as "first" and "second" mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.
[0023] It should be noted that the modifications of "one" and "plurality" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, it should be understood as "one or more".
[0024] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.
[0025] The present disclosure will be described in detail below with reference to the accompanying drawings and in conjunction with embodiments.
[0026] Figure 1 1 is a flowchart of some embodiments of the method for training an aligned personality recognition model based on a multimodal scenario according to the present disclosure. 100 of some embodiments of the method for training an aligned personality recognition model based on a multimodal scenario according to the present disclosure is shown. The method for training an aligned personality recognition model based on a multimodal scenario comprises the following steps:
[0027] Step 101: extract video frames from user videos to obtain video frame image sequences.
[0028] In some embodiments, the execution subject (e.g., computing device) of the method for training a personality recognition model based on alignment in a multimodal scenario can extract video frames from a user video to obtain a video frame image sequence. A user video may refer to a video in which each frame of the video image contains the user's face. The user video may be framed to obtain a video frame image sequence.
[0029] Step 102: Perform multi-scale feature extraction on each frame image in the video frame image sequence to generate multi-scale video frame image features and obtain a multi-scale video frame image feature sequence.
[0030] In some embodiments, the execution subject may perform multi-scale feature extraction on each frame image in the video frame image sequence to generate multi-scale video frame image features and obtain a multi-scale video frame image feature sequence. For example, a Transformer model may be used to perform multi-scale feature extraction on each frame image in the video frame image sequence to generate multi-scale video frame image features and obtain a multi-scale video frame image feature sequence.
[0031] For example, multi-scale feature extraction may be performed on each frame image in the video frame image sequence according to the formula to generate multi-scale video frame image features, thereby obtaining a multi-scale video frame image feature sequence:
[0032]
[0033] Among them, F patch and F MHSA They represent the operation of cutting the video frame image into blocks and performing multi-head attention branches, respectively. v_tran It represents the multi-scale video frame image features after the Transformer branch operation.
[0034] Step 103 , performing face capture on the face in each video frame image in the video frame image sequence to generate a face image and obtain a face image sequence.
[0035] In some embodiments, the execution subject may perform face capture on the face in each video frame image in the video frame image sequence to generate a face image and obtain a face image sequence. For example, an image of the face region in each video frame image may be captured to generate a face image and obtain a face image sequence.
[0036] Step 104: perform multi-scale feature extraction on each face image in the face image sequence to generate multi-scale face image features and obtain a multi-scale face image feature sequence.
[0037] In some embodiments, the execution subject may perform multi-scale feature extraction on each face image in the face image sequence to generate multi-scale face image features and obtain a multi-scale face image feature sequence. For example, a CNN convolutional neural network model may be used to perform multi-scale feature extraction on each face image in the face image sequence to generate multi-scale face image features and obtain a multi-scale face image feature sequence.
[0038] For example, multi-scale feature extraction may be performed on each face image in the face image sequence using the formula:
[0039]
[0040] Among them, L l represents the face image sequence, L represents the number of sampling frames of the face image sequence, and F face and F conv They represent the face extraction operation and the convolution branch operation of the face image, respectively. v_conv It represents the multi-scale face image features after the convolution branch operation.
[0041] Step 105 , aligning the multi-scale video frame image feature sequence with the multi-scale face image feature sequence to obtain a visual alignment information sequence.
[0042] In some embodiments, the execution subject may align the multiscale video frame image feature sequence with the multiscale face image feature sequence to obtain a visual alignment information sequence, wherein the multiscale video frame image features in the multiscale video frame image feature sequence correspond to the multiscale face image features in the multiscale face image feature sequence.
[0043] In practice, the above-mentioned execution entity can perform the following processing steps for each multi-scale video frame image feature in the multi-scale video frame image feature sequence: align and fuse the multi-scale video frame image feature with the corresponding multi-scale face image feature to obtain visual alignment information.
[0044] For example, the multi-scale video frame image feature sequence and the multi-scale face image feature sequence may be aligned using the formula:
[0045] f v =F flatten (F interact (f v_conv , f v_tran ).
[0046] Among them, F flatten F stands for Flatten. interactIt means that the information of the two branches (the multi-scale video frame image feature sequence and the multi-scale face image feature sequence) is fully interactively operated.
[0047] Step 106: extract text features from the transcribed text corresponding to the user video to obtain a transcribed text feature sequence.
[0048] In some embodiments, the execution subject may perform text feature extraction on the transcribed text corresponding to the user video to obtain a transcribed text feature sequence. The transcribed text may be a text transcribed from each frame of audio in the user video. For example, the BERT model may be used to perform text feature extraction on the transcribed text corresponding to the user video to obtain a transcribed text feature sequence.
[0049] Step 107: extract audio features from the audio corresponding to the user video to obtain an audio feature sequence.
[0050] In some embodiments, the above-mentioned execution subject can perform audio feature extraction on the audio corresponding to the user video to obtain an audio feature sequence. For example, the audio analysis tool library can be used to extract the logarithmic Mel spectrum features of the audio signal, and these features can be input into the convolutional neural network VGGish model to extract the audio feature embedding. Through this process, it is intended to fully capture the rich information contained in the audio signal. As a powerful audio feature extraction tool, the VGGish model can effectively convert the original audio signal into an embedded representation with high-level semantic information, thereby providing a solid foundation for subsequent personality prediction audio analysis tasks. This method not only improves the efficiency of personality feature extraction, but also enhances the personality model's ability to understand and process complex audio data.
[0051] In the audio modality, personality emotion information is conveyed through intonation and pitch changes, while the text modality expresses emotions and implicit personality meanings through vocabulary and sentence structure. The combination of the two can provide a more comprehensive personality prediction analysis. The audio modality contains information such as the pitch, prosody, rhythm, and emotion of the speech, which are crucial for understanding the speaker's emotional state, tone changes, and language expression. The text modality mainly contains semantic and syntactic structures, providing clear vocabulary and sentence meanings. Therefore, when processing non-visual modality data, this study explores the alignment of text and audio modalities and adopts an early fusion strategy to map text and audio features into a representation space of the same dimensionality.
[0052] For example, when processing audio data, this study used three different methods to explore which method is most suitable for personality trait prediction. First, each audio segment was uniformly converted to a 16kHz sampling rate, and the librosa tool library was used to extract the audio's Mel-frequency cepstral coefficients and logarithmic Mel-frequency spectrum, which contains the original key information in the audio signal; secondly, the convolutional layer and pooling layer of the VGGish model were used to convert the audio information into a 128-dimensional embedding vector, providing a compressed and informative personality representation for the audio features; finally, the Resnet model structure was used to extract audio features as a comparative experiment. For text features, the study used the advanced BERT model for processing and converted the text into a 768-dimensional feature vector. The powerful contextual understanding ability of the BERT model provides strong support for capturing subtle language differences related to personality traits.
[0053] Step 108: align and fuse the transcribed text feature sequence and the audio feature sequence to obtain a non-visual alignment information sequence.
[0054] In some embodiments, the execution subject may perform alignment and fusion processing on the transcribed text feature sequence and the audio feature sequence to obtain a non-visual alignment information sequence, wherein the transcribed text features in the transcribed text feature sequence correspond to the audio features in the audio feature sequence.
[0055] In practice, the execution subject may perform the following processing steps for each transcribed text feature in the transcribed text feature sequence: align and fuse the transcribed text feature with the corresponding audio feature to obtain non-visual alignment information. That is, the non-visual alignment information includes: the transcribed text feature and the corresponding audio feature.
[0056] Step 109 : training the initial aligned personality recognition model according to the visual alignment information sequence and the non-visual alignment information sequence to obtain a trained aligned personality recognition model.
[0057] In some embodiments, the execution subject may train the initial alignment personality recognition model according to the visual alignment information sequence and the non-visual alignment information sequence to obtain a trained alignment personality recognition model, wherein the visual alignment information in the visual alignment information sequence corresponds to the non-visual alignment information in the non-visual alignment information sequence.
[0058] In practice, the above execution entity can train the initial aligned personality recognition model through the following steps to obtain a trained aligned personality recognition model:
[0059] In the first step, the visual alignment information sequence is input into the initial visual personality recognition model included in the initial alignment personality recognition model to obtain a visual personality recognition result sequence. One visual alignment information corresponds to one visual personality recognition result.
[0060] For example, in the process of visual modality alignment, the extracted visual features are mapped to spatial dimensions through a linear layer:
[0061]
[0062] Among them, MLP() represents a feed-forward artificial neural network model. Represents the visual personality recognition result (the visual Big Five personality prediction result by averaging the two-branch results).
[0063] The second step is to determine the visual personality loss value between the visual personality recognition result sequence and the corresponding visual personality recognition label based on the preset visual personality loss function.
[0064] For example, the visual personality loss function can be:
[0065]
[0066] Among them, y i Indicates the corresponding visual personality recognition tag. Here, the visual personality recognition tag can be pre-set and indicates the corresponding real personality tag.
[0067] The third step is to input the non-visual alignment information sequence into the initial non-visual personality recognition model included in the initial alignment personality recognition model to obtain a non-visual personality recognition result sequence. One non-visual alignment information corresponds to one non-visual personality recognition result.
[0068] For example, the transcribed text features and the corresponding audio features included in the non-visual alignment information may be input into the initial non-visual personality recognition model using the following formula:
[0069]
[0070] in, Represents the non-visual personality recognition result. t Represents the features of the transcribed text. a Represents the corresponding audio features.
[0071] The fourth step is to determine the non-visual personality loss value between the non-visual personality recognition result sequence and the corresponding non-visual personality recognition label based on the preset non-visual personality loss function.
[0072] For example, a non-visual personality loss function can be:
[0073]
[0074] Among them, L n Represents the non-visual personality loss value. N represents the number of non-visual personality recognition results in the non-visual personality recognition result sequence.
[0075] In a fifth step, based on the visual alignment information sequence and the non-visual alignment information sequence, a fusion personality recognition result is generated by using an initial fusion personality recognition model included in the initial alignment personality recognition model.
[0076] The fifth step may include the following sub-steps:
[0077] In the first sub-step, for each visual alignment information in the visual alignment information sequence, the following fusion steps are performed:
[0078] 1. Input the visual alignment information and the corresponding non-visual alignment information into the joint attention mechanism layer included in the initial fusion personality recognition model to obtain the joint attention feature.
[0079] 2. Perform vector mapping on the visual alignment information to obtain a visual alignment information vector.
[0080] 3. Perform vector mapping on the non-visual alignment information corresponding to the visual alignment information to obtain a non-visual alignment information vector.
[0081] 4. Construct a cross-modal interactive visual feature based on the visual alignment information vector and the joint attention feature.
[0082] 5. Construct a cross-modal interactive non-visual feature based on the non-visual alignment information vector and the joint attention feature.
[0083] 6. Construct a fusion feature based on the cross-modal interaction visual feature and the cross-modal interaction non-visual feature.
[0084] As an example, the fusion step can be performed by the following formula:
[0085]
[0086] f′ n =FFN n (softmax(f Attn_joint T )proj v (f υ ))
[0087] f′ v =FFN v (softmax(f Attn_joint)proj n (f n )).
[0088] Among them, f n and f v They represent cross-modal interactive non-visual features and cross-modal interactive visual features respectively. And proj v (f v ) and proj n (f n ) are respectively the mapping operations of the visual alignment information and the non-visual alignment information query vector, which are used to transform the original features into the query vector. Attn_joint The joint attention mechanism of visual and non-visual features improves the personality prediction performance of the model by integrating the information of the two modalities. FFN stands for the feedforward neural network structure, which is used to further process the features after the joint attention. n and f′ v They are the cross-modal interactive non-visual features and cross-modal interactive visual features output after the interactive operation, which are used for subsequent personality prediction tasks.
[0089] In the second sub-step, the fused feature sequence is input into the initial fused personality recognition model to obtain a fused personality recognition result sequence.
[0090] For example, the fusion personality recognition result can be output by the following formula:
[0091]
[0092] in, Indicates the fusion personality recognition result.
[0093] The sixth step is to determine the fusion loss value between the fusion personality recognition result sequence and the corresponding fusion personality recognition label based on the preset fusion personality recognition loss function.
[0094] For example, the fusion loss value between the fusion personality recognition result sequence and the corresponding fusion personality recognition label can be determined by the following fusion personality recognition loss function:
[0095]
[0096] Among them, L c Represents the fusion loss value.
[0097] In the seventh step, weighted fusion is performed on the visual personality loss value, the non-visual personality loss value and the fusion loss value to obtain a model loss value.
[0098] For example, the visual personality loss value, the non-visual personality loss value and the fusion loss value may be weightedly fused by the following formula:
[0099] L=αL v +βL n +γL c .
[0100] Among them, α, β, and γ are all constant values, and L is the final model loss value.
[0101] In the eighth step, in response to determining that the model loss value is less than or equal to the preset loss value, the initial aligned personality recognition model is determined as the trained aligned personality recognition model.
[0102] Optionally, in response to receiving a target user video to be detected, the target user video is input into the aligned personality recognition model to obtain a user personality recognition result.
[0103] In some embodiments, the execution subject may, in response to receiving a target user video to be detected, input the target user video into the aligned personality recognition model to obtain a user personality recognition result. The target user video to be detected may be a video including the user's face. The user personality recognition result may represent five personality recognition results, namely, openness, responsibility, extroversion, agreeableness, and neuroticism.
[0104] like Figure 4The example shows a model training diagram of steps 101 to 109, including: In terms of visual feature extraction, convolution operation is a basic and powerful tool in deep learning. This operation captures local features and learns the spatial hierarchy of data by sliding the convolution kernel (or filter) on the local area of the input data. This mechanism allows the model to recognize basic image elements such as edges and textures, and as the network layer deepens, it can gradually recognize more complex patterns and objects. Therefore, convolutional neural networks (CNNs) have achieved great success in fields such as image recognition, classification, and video analysis. However, although CNN is extremely effective in extracting local features, it faces certain challenges in capturing global representations. Traditional convolutional networks often have difficulty handling long-distance dependency problems due to their inherent structural limitations. For example, when processing high-resolution images, local convolution operations may not be able to fully capture the relationship between different areas of the image, which may lead to performance limitations in tasks that require global understanding, such as scene classification or overall structural analysis of images. To address this problem, the emergence of the Transformer model provides an effective solution. Through its self-attention mechanism, the Transformer can directly calculate the dependency between any two positions in the sequence, no matter how far apart they are. This feature enables the Transformer to perform well when processing data with long-distance dependencies, such as text or time series data. The introduction of the self-attention mechanism not only improves the model's ability to handle complex spatial transformations, but also builds a global representation of the data, which is crucial for understanding complex structures and semantics. However, compared with convolutional neural networks, the Transformer has certain shortcomings in capturing local details. Although it can effectively capture long-distance feature dependencies and form a global representation, it may not be as sensitive as a convolutional network when processing tasks that require fine local information, such as detailed textures in images or subtle semantic differences in text. In this case, the Transformer may ignore some important local information, thereby affecting the performance of the model on specific tasks. For example, in image segmentation tasks, the distinguishability of foreground and background is crucial, but the Transformer may ignore local differences because it focuses too much on global information.
[0105] In order to overcome these limitations and make full use of the respective advantages of CNN and Transformer, this personality prediction study proposes a new visual modality alignment structure. Figure 5As shown in the figure, this study uses a dual-branch structure for personality assessment experimental framework: the processing of face images is one branch, and the processing of panoramic images is another branch, and the dual-branch parallel operation is performed. By integrating convolution operations and self-attention modules, the dual branches not only enhance the ability to capture local features of the image, but also maintain sensitivity to global information. With the effective supplementation and enhancement of global and local features, the model's ability to recognize character features and analyze overall personality traits is improved.
[0106] Specifically, for the facial image processing task branch, special attention is paid to the subtle facial expression changes and body movements of individuals. In order to extract more detailed personality information from facial images, the experiment uses a convolutional neural network (CNN) structure to extract local facial features. By gradually reducing the scale of the feature map during the convolution process, this branch constructs a feature pyramid structure. This structure combines visual features of different scales as the network depth deepens and the number of channels increases. This method takes into account that in a time series segment, the feature changes of the character are mainly concentrated in specific expressions such as the corners of the mouth and eyes, while the position of the facial features and the facial contour information remain basically unchanged. Therefore, based on this observation, the experiment adopts a multi-scale image feature extraction method. Different from the traditional method of extracting multi-scale features for a single static image, this personality prediction study reduces the scale of image features in the time series segment according to the time order and integrates them into multi-scale image features on the time series. A multi-scale image feature extraction method based on time series is formed. This multi-scale feature extraction method focuses more on capturing detailed features in the time series, while reducing the number of model parameters. Its performance is comparable to that of the traditional single-scale time series feature training method for personality prediction.
[0107] like Figure 6 As shown in the figure, in order to effectively align the information extracted from panoramic images and facial expressions in the visual modality and promote the complementarity and flow of local and global information, this study implemented the adjustment and transformation of feature scales in the dual-branch interaction process of image processing. This visual two-way interaction mechanism aims to achieve the fusion of personality information at different levels, thereby enhancing the personality expression ability and recognition accuracy of the model. In this way, local features and global features can be more closely combined, which not only improves the system's performance in dealing with personality prediction tasks in complex vision, enables the model to understand and utilize visual information more comprehensively, but also improves the accuracy and reliability of personality trait prediction.
[0108] After successfully extracting visual and non-visual features, this study proposed a visual and non-visual modality alignment enhancement strategy, which aims to further explore and utilize the complementary personality information between modalities while retaining the uniqueness of each modality. This strategy strengthens the common personality content of each modality through cross-modal feature fusion. The specific method is to input visual and non-visual features into a specially designed personality feature enhancer, which consists of a multi-layer network specifically designed to process and fuse personality information from different modalities, thereby achieving more accurate personality trait prediction.
[0109] Inside the feature enhancer, this study implements a cross-attention mechanism from visual to non-visual features, and from non-visual to visual features. This design aims to promote the flow and integration of personality information between different modalities. Figure 7 As shown, through such a design, each modality can not only respond to changes in its own personality traits, but also receive and utilize complementary personality information from other modalities, thereby achieving a deeper fusion of personality traits.
[0110] Further references Figure 2 As an implementation of the methods shown in the above figures, the present disclosure provides some embodiments of an aligned personality recognition model training device based on a multimodal scenario. These embodiments of the aligned personality recognition model training device based on a multimodal scenario are similar to Figure 1 Corresponding to the method embodiments shown, the aligned personality recognition model training device based on a multimodal scenario can be specifically applied to various electronic devices.
[0111] like Figure 2As shown, the aligned personality recognition model training device 200 based on a multimodal scenario in some embodiments includes: an extraction unit 201, a feature extraction unit 202, a clipping unit 203, a face extraction unit 204, a first alignment unit 205, a text extraction unit 206, an audio extraction unit 207, a second alignment unit 208 and a training unit 209. Among them, the extraction unit 201 is configured to perform video frame extraction on the user video to obtain a video frame image sequence; the feature extraction unit 202 is configured to perform multi-scale feature extraction on each frame image in the video frame image sequence to generate multi-scale video frame image features and obtain a multi-scale video frame image feature sequence; the interception unit 203 is configured to perform face interception on the face in each video frame image in the video frame image sequence to generate a face image and obtain a face image sequence; the face extraction unit 204 is configured to perform multi-scale feature extraction on each face image in the face image sequence to generate multi-scale face image features and obtain a multi-scale face image feature sequence; the first alignment unit 205 is configured to align the multi-scale video frame image feature sequence with the multi-scale face image feature sequence to obtain a visual alignment information sequence, wherein the multi-scale video frame image features in the multi-scale video frame image feature sequence are aligned with each other. The multi-scale facial image features in the multi-scale facial image feature sequence correspond to the multi-scale facial image feature sequence; the text extraction unit 206 is configured to perform text feature extraction on the transcribed text corresponding to the user video to obtain a transcribed text feature sequence; the audio extraction unit 207 is configured to perform audio feature extraction on the audio corresponding to the user video to obtain an audio feature sequence; the second alignment unit 208 is configured to perform alignment and fusion processing on the transcribed text feature sequence and the audio feature sequence to obtain a non-visual alignment information sequence, wherein the transcribed text features in the transcribed text feature sequence correspond to the audio features in the audio feature sequence; the training unit 209 is configured to train the initial aligned personality recognition model according to the visual alignment information sequence and the non-visual alignment information sequence to obtain a trained aligned personality recognition model, wherein the visual alignment information in the visual alignment information sequence corresponds to the non-visual alignment information in the non-visual alignment information sequence.
[0112] It can be understood that the units recorded in the alignment personality recognition model training device 200 based on the multimodal scenario are similar to the reference Figure 1 Therefore, the operations, features and beneficial effects described above for the method are also applicable to the aligned personality recognition model training device 200 based on the multimodal scenario and the units contained therein, and will not be described in detail here.
[0113] Reference below Figure 3, which shows a schematic diagram of the structure of an electronic device 300 (e.g., a computing device) suitable for implementing some embodiments of the present disclosure. The electronic devices in some embodiments of the present disclosure may include, but are not limited to, mobile terminals such as mobile phones, notebook computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 3 The electronic device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present disclosure.
[0114] like Figure 3 As shown, the electronic device 300 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 301, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 302 or a program loaded from a storage device 308 into a random access memory (RAM) 303. In the RAM 303, various programs and data required for the operation of the electronic device 300 are also stored. The processing device 301, the ROM 302, and the RAM 303 are connected to each other via a bus 304. An input / output (I / O) interface 305 is also connected to the bus 304.
[0115] Typically, the following devices may be connected to the I / O interface 305: input devices 306 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; output devices 307 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; storage devices 308 including, for example, a magnetic tape, a hard disk, etc.; and communication devices 309. The communication devices 309 may allow the electronic device 300 to communicate with other devices wirelessly or by wire to exchange data. Although Figure 3 The electronic device 300 is shown with various devices, but it should be understood that it is not required to implement or possess all the devices shown. More or fewer devices may be implemented or possessed instead. Figure 3 Each block shown in the figure may represent one device, or may represent multiple devices as required.
[0116] In particular, according to some embodiments of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, some embodiments of the present disclosure include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In some such embodiments, the computer program can be downloaded and installed from the network through the communication device 309, or installed from the storage device 308, or installed from the ROM 302. When the computer program is executed by the processing device 301, the above-mentioned functions defined in the method of some embodiments of the present disclosure are executed.
[0117] It should be noted that the computer-readable medium recorded in some embodiments of the present disclosure may be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In some embodiments of the present disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in combination with an instruction execution system, device or device. In some embodiments of the present disclosure, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries a computer-readable program code. This propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer readable signal medium may also be any computer readable medium other than a computer readable storage medium, which may send, propagate or transmit a program for use by or in conjunction with an instruction execution system, apparatus or device. The program code contained on the computer readable medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.
[0118] In some embodiments, the client and the server may communicate using any currently known or future developed network protocol such as HTTP (HyperText Transfer Protocol), and may be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), an internet (e.g., the Internet), and a peer-to-peer network (e.g., an ad hoc peer-to-peer network), as well as any currently known or future developed network.
[0119] The above-mentioned computer-readable medium may be included in the above-mentioned electronic device; or it may exist independently without being assembled into the electronic device. The above-mentioned computer-readable medium carries one or more programs. When the above-mentioned one or more programs are executed by the electronic device, the electronic device: performs video frame extraction on the user video to obtain a video frame image sequence; performs multi-scale feature extraction on each frame image in the video frame image sequence to generate multi-scale video frame image features to obtain a multi-scale video frame image feature sequence; performs face capture on the face in each video frame image in the video frame image sequence to generate a face image to obtain a face image sequence; performs multi-scale feature extraction on each face image in the face image sequence to generate a multi-scale face image feature to obtain a multi-scale face image feature sequence; performs alignment processing on the multi-scale video frame image feature sequence with the multi-scale face image feature sequence to obtain a visual alignment information sequence, wherein the multi-scale video frame image The multi-scale video frame image features in the feature sequence correspond to the multi-scale face image features in the multi-scale face image feature sequence; text features are extracted from the transcribed text corresponding to the user video to obtain a transcribed text feature sequence; audio features are extracted from the audio corresponding to the user video to obtain an audio feature sequence; the transcribed text feature sequence and the audio feature sequence are aligned and fused to obtain a non-visual alignment information sequence, wherein the transcribed text features in the transcribed text feature sequence correspond to the audio features in the audio feature sequence; an initial aligned personality recognition model is trained according to the visual alignment information sequence and the non-visual alignment information sequence to obtain a trained aligned personality recognition model, wherein the visual alignment information in the visual alignment information sequence corresponds to the non-visual alignment information in the non-visual alignment information sequence.
[0120] Computer program code for performing the operations of some embodiments of the present disclosure may be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0121] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present disclosure. In this regard, each square box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some implementations as replacements, the functions marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two square boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0122] The units described in some embodiments of the present disclosure may be implemented by software or by hardware. The described units may also be provided in a processor, for example, may be described as: a processor comprising: an extraction unit, a feature extraction unit, a capture unit, a face extraction unit, a first alignment unit, a text extraction unit, an audio extraction unit, a second alignment unit and a training unit. Among them, the names of these units do not constitute a limitation on the units themselves in certain cases, for example, the extraction unit may also be described as "a unit for extracting video frames from user videos to obtain a video frame image sequence".
[0123] The functions described above herein may be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), and the like.
[0124] The above descriptions are only some preferred embodiments of the present disclosure and an explanation of the technical principles used. Those skilled in the art should understand that the scope of the invention involved in the embodiments of the present disclosure is not limited to the technical solutions formed by a specific combination of the above-mentioned technical features, but should also cover other technical solutions formed by any combination of the above-mentioned technical features or their equivalent features without departing from the above-mentioned inventive concept. For example, the above-mentioned features are replaced with the technical features with similar functions disclosed in the embodiments of the present disclosure (but not limited to) and the technical solutions formed.
Claims
1. A method for training an aligned personality recognition model based on a multimodal scenario, comprising: Extract video frames from user videos to obtain video frame image sequences; Performing multi-scale feature extraction on each frame image in the video frame image sequence to generate multi-scale video frame image features and obtain a multi-scale video frame image feature sequence; Performing face capture on a human face in each video frame image in the video frame image sequence to generate a face image and obtain a face image sequence; Performing multi-scale feature extraction on each face image in the face image sequence to generate multi-scale face image features and obtain a multi-scale face image feature sequence; Aligning the multiscale video frame image feature sequence with the multiscale face image feature sequence to obtain a visual alignment information sequence, wherein the multiscale video frame image features in the multiscale video frame image feature sequence correspond to the multiscale face image features in the multiscale face image feature sequence; Performing text feature extraction on the transcribed text corresponding to the user video to obtain a transcribed text feature sequence; Extracting audio features from the audio corresponding to the user video to obtain an audio feature sequence; Performing alignment and fusion processing on the transcribed text feature sequence and the audio feature sequence to obtain a non-visual alignment information sequence, wherein the transcribed text features in the transcribed text feature sequence correspond to the audio features in the audio feature sequence; The initial aligned personality recognition model is trained according to the visual alignment information sequence and the non-visual alignment information sequence to obtain a trained aligned personality recognition model, wherein the visual alignment information in the visual alignment information sequence corresponds to the non-visual alignment information in the non-visual alignment information sequence.
2. The method according to claim 1, wherein: The method further comprises: In response to receiving the target user video to be detected, the target user video is input into the aligned personality recognition model to obtain a user personality recognition result.
3. The method according to claim 1, wherein: The aligning process of the multi-scale video frame image feature sequence and the multi-scale face image feature sequence to obtain visual alignment information includes: For each multi-scale video frame image feature in the multi-scale video frame image feature sequence, the following processing steps are performed: The multi-scale video frame image features and the corresponding multi-scale face image features are aligned and fused to obtain visual alignment information.
4. The method according to claim 1, wherein: The aligning and fusing the transcribed text feature sequence with the audio feature sequence to obtain a non-visual alignment information sequence includes: For each transcribed text feature in the transcribed text feature sequence, the following processing steps are performed: The transcribed text features are aligned and fused with the corresponding audio features to obtain non-visual alignment information.
5. The method according to claim 1, wherein: The step of training the initial aligned personality recognition model according to the visual alignment information sequence and the non-visual alignment information sequence to obtain a trained aligned personality recognition model includes: Inputting the visual alignment information sequence into the initial visual personality recognition model included in the initial alignment personality recognition model to obtain a visual personality recognition result sequence; Based on a preset visual personality loss function, determining a visual personality loss value between a visual personality recognition result sequence and a corresponding visual personality recognition label; Inputting the non-visual alignment information sequence into an initial non-visual personality recognition model included in the initial alignment personality recognition model to obtain a non-visual personality recognition result sequence; Based on a preset non-visual personality loss function, determining a non-visual personality loss value between a non-visual personality recognition result sequence and a corresponding non-visual personality recognition label; Based on the visual alignment information sequence and the non-visual alignment information sequence, generating a fusion personality recognition result sequence through an initial fusion personality recognition model included in the initial alignment personality recognition model; Based on a preset fusion personality recognition loss function, determining a fusion loss value between the fusion personality recognition result sequence and the corresponding fusion personality recognition label; Performing weighted fusion on the visual personality loss value, the non-visual personality loss value and the fusion loss value to obtain a model loss value; In response to determining that the model loss value is less than or equal to the preset loss value, the initial aligned personality recognition model is determined as the trained aligned personality recognition model.
6. The method according to claim 5, wherein: The generating a fusion personality recognition result based on the visual alignment information sequence and the non-visual alignment information sequence by using the initial fusion personality recognition model included in the initial alignment personality recognition model includes: For each visual alignment information in the visual alignment information sequence, the following fusion steps are performed: Inputting the visual alignment information and the corresponding non-visual alignment information into a joint attention mechanism layer included in the initial fusion personality recognition model to obtain a joint attention feature; Performing vector mapping on the visual alignment information to obtain a visual alignment information vector; Performing vector mapping on the non-visual alignment information corresponding to the visual alignment information to obtain a non-visual alignment information vector; constructing a cross-modal interactive visual feature according to the visual alignment information vector and the joint attention feature; constructing a cross-modal interactive non-visual feature according to the non-visual alignment information vector and the joint attention feature; constructing a fusion feature according to the cross-modal interaction visual feature and the cross-modal interaction non-visual feature; The fused feature sequence is input into the initial fused personality recognition model to obtain the fused personality recognition result.
7. A training device for an aligned personality recognition model based on a multimodal scenario, comprising: An extraction unit is configured to extract video frames from a user video to obtain a video frame image sequence; A feature extraction unit is configured to perform multi-scale feature extraction on each frame image in the video frame image sequence to generate multi-scale video frame image features and obtain a multi-scale video frame image feature sequence; A capture unit is configured to capture a face in each video frame image in the video frame image sequence to generate a face image and obtain a face image sequence; A face extraction unit is configured to perform multi-scale feature extraction on each face image in the face image sequence to generate multi-scale face image features and obtain a multi-scale face image feature sequence; A first alignment unit is configured to perform alignment processing on the multiscale video frame image feature sequence and the multiscale face image feature sequence to obtain a visual alignment information sequence, wherein the multiscale video frame image features in the multiscale video frame image feature sequence correspond to the multiscale face image features in the multiscale face image feature sequence; A text extraction unit is configured to extract text features from the transcribed text corresponding to the user video to obtain a transcribed text feature sequence; an audio extraction unit, configured to extract audio features from the audio corresponding to the user video to obtain an audio feature sequence; A second alignment unit is configured to perform alignment and fusion processing on the transcribed text feature sequence and the audio feature sequence to obtain a non-visual alignment information sequence, wherein the transcribed text features in the transcribed text feature sequence correspond to the audio features in the audio feature sequence; The training unit is configured to train the initial aligned personality recognition model according to the visual alignment information sequence and the non-visual alignment information sequence to obtain a trained aligned personality recognition model, wherein the visual alignment information in the visual alignment information sequence corresponds to the non-visual alignment information in the non-visual alignment information sequence.
8. An electronic device comprising: one or more processors; a storage device having one or more programs stored thereon; When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 6.
9. A computer readable medium having a computer program stored thereon, wherein: When the program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Personality detection method based on multi-modal alignment and multi-vector representation
CN111259976A
Face anti-fraud method based on cross-domain feature alignment network
CN114120401A
Multi-modal face emotion recognition method and device
CN114399818A
Abnormal emotion inference system based on attitude feature alignment
CN118135664A
Large five personality prediction method and device of multi-modal and graph structure feature learning network
CN118245965A