Multi-modal personality detection model training method, user information pushing method and device

By using multimodal features of video, audio and text to train the initial multimodal personality detection model, the problems of low detection accuracy and large computing resource utilization in the prior art are solved, and more efficient personality detection is achieved.

CN119964052APending Publication Date: 2025-05-09SHENZHEN JUSI ARTIFICIAL INTELLIGENCE TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510028730.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-08
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

The existing personality detection model has low detection accuracy when using single-modal sample data, and the model structure is complex when using multi-modal sample data, and occupies more computing resources.

Method used

The initial multimodal personality detection model is trained using three modal data: video, audio and text. Feature extraction and fusion are performed through the video audio processing module and the video text processing module to reduce the number of feature extraction branches and avoid the introduction of shared encoder.

Benefits of technology

The detection accuracy of personality detection model is improved, the complexity of model structure is reduced, and the use of computing resources is reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119964052A_ABST
    Figure CN119964052A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a multi-modal personality detection model training method and a user information pushing method and device. One specific embodiment of the method comprises the following steps: acquiring an original audio-visual data set and an initial multi-modal personality detection model; preprocessing each piece of original audio-visual data to obtain a target sample data set; selecting at least one target sample data, and executing the following steps: inputting sample video data and sample audio features included in the target sample data into a video and audio processing module to output a first modal fusion result; inputting sample video data and sample text features included in the target sample data into a video text processing module to output a second modal fusion result; inputting the first modal fusion result and the corresponding second modal fusion result into a prediction module to output a personality detection result; generating a training loss value; and determining the trained initial multi-modal personality detection model as a multi-modal personality detection model. According to the embodiment, the accuracy of personality detection can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present disclosure relate to the field of personality prediction, and specifically to a multimodal personality detection model training method, a user information push method and a device. Background Art

[0002] Personality detection plays an important role in selecting job candidates to better fit the required positions. At present, when conducting personality detection, the usual method is to use single-modal sample data or multi-modal sample data to train the personality detection model for personality detection of job candidates. Among them, when using multi-modal sample data to train the personality detection model, it is usually necessary to design different feature extraction branches for each different modality of data, and introduce a shared encoder to achieve cross-modal information fusion.

[0003] However, it is found in practice that when the above method is adopted, the following technical problems often occur:

[0004] If single-modal sample data is used to train the personality detection model, the detection accuracy of the personality detection model will be low due to the relatively simple sample data. If multi-modal sample data is used to train the personality detection model, the structure of the personality detection model will be more complicated due to the need to design different feature extraction branches for each different modality and introduce a shared encoder for information fusion, thus occupying more computing resources.

[0005] The above information disclosed in this Background section is only for enhancement of understanding of the background of the present disclosure concept and therefore it may contain information that does not form the prior art that is already known to a person of ordinary skill in the art. Summary of the invention

[0006] The content of this disclosure is used to introduce concepts in a brief form, which will be described in detail in the detailed implementation section below. The content of this disclosure is not intended to identify the key features or essential features of the technical solution claimed for protection, nor is it intended to limit the scope of the technical solution claimed for protection.

[0007] Some embodiments of the present disclosure propose a multimodal personality detection model training method, a user information push method and a device to solve one or more of the technical problems mentioned in the above background technology section.

[0008] In a first aspect, some embodiments of the present disclosure provide a multimodal personality detection model training method, the method comprising: obtaining an original audio-visual data set and an initial multimodal personality detection model, wherein each original audio-visual data in the original audio-visual data set comprises original video data, original audio data, speech text data and sample labels, the initial multimodal personality detection model comprises a video and audio processing module, a video and text processing module and a prediction module, the video and audio processing module comprises a first feature extraction submodule, a first spatiotemporal feature enhancement submodule and a first multimodal fusion submodule, the video and text processing module comprises a second feature extraction submodule, a second spatiotemporal feature enhancement submodule and a second multimodal fusion submodule, the first multimodal fusion submodule and the second multimodal fusion submodule are both used to perform cross-modal fusion on features of modalities such as video, audio and text; pre-processing each original audio-visual data in the original audio-visual data set to obtain a target sample data set, wherein each target sample data in the target sample data set comprises sample video data, sample audio features and sample text features; selecting at least one target sample data from the target sample data set, and The personality detection model performs the following training steps: inputting the sample video data and sample audio features included in each target sample data in the at least one target sample data into a video and audio processing module included in the initial multimodal personality detection model to output a first modal fusion result, thereby obtaining at least one first modal fusion result; inputting the sample video data and sample text features included in each target sample data in the at least one target sample data into a video and text processing module included in the initial multimodal personality detection model to output a second modal fusion result, thereby obtaining at least one second modal fusion result; inputting each first modal fusion result in the at least one first modal fusion result and the second modal fusion result corresponding to the first modal fusion result into a prediction module included in the initial multimodal personality detection model to output a personality detection result, thereby obtaining at least one personality detection result; generating a training loss value based on the sample label and the personality detection result corresponding to each target sample data in the at least one target sample data by a preset hybrid loss function; and determining the initial multimodal personality detection model after training as a multimodal personality detection model in response to determining that the training loss value is less than a preset training loss threshold.

[0009] In a second aspect, some embodiments of the present disclosure provide a user information push method, the method comprising: obtaining user audio-visual data corresponding to a target user, wherein the user audio-visual data comprises user video data, user audio data and voice text data, and the target user corresponds to a user identifier; performing standardization processing on the user video data included in the user audio-visual data to obtain target video data; performing feature extraction processing on the user audio data included in the user audio-visual data to obtain target audio features; performing text embedding processing on the voice text data included in the user audio-visual data to obtain target text features; inputting the target video data, the target audio features and the target text features into a pre-trained multimodal personality detection model to obtain a personality detection result corresponding to the target user, wherein the multimodal personality detection model is pre-trained by the method described in any implementation manner of the first aspect; in response to determining that the personality detection result meets a preset job requirement condition, obtaining resume information corresponding to the target user; determining the user identifier and the resume information corresponding to the target user as user information, and sending the user information to a target terminal.

[0010] In a third aspect, some embodiments of the present disclosure provide a multimodal personality detection model training device, the device comprising: an acquisition unit, configured to acquire an original audio-visual data set and an initial multimodal personality detection model, wherein each original audio-visual data in the original audio-visual data set comprises original video data, original audio data, speech text data and sample labels, the initial multimodal personality detection model comprises a video and audio processing module, a video and text processing module and a prediction module, the video and audio processing module comprises a first feature extraction submodule, a first spatiotemporal feature enhancement submodule and a first multimodal fusion submodule, the video and text processing module comprises a first feature extraction submodule, a first spatiotemporal feature enhancement submodule and a first multimodal fusion submodule, The processing module includes a second feature extraction submodule, a second spatiotemporal feature enhancement submodule and a second multimodal fusion submodule, wherein the first multimodal fusion submodule and the second multimodal fusion submodule are both used for cross-modal fusion of features of modes such as video, audio and text; a preprocessing unit is configured to preprocess each original audio-visual data in the original audio-visual data set to obtain a target sample data set, wherein each target sample data in the target sample data set includes sample video data, sample audio features and sample text features; a selection and execution unit is configured to select at least one from the target sample data set The invention relates to target sample data, and based on the initial multimodal personality detection model, performing the following training steps: inputting the sample video data and sample audio features included in each target sample data in the at least one target sample data into the video and audio processing module included in the initial multimodal personality detection model to output a first modal fusion result, thereby obtaining at least one first modal fusion result; inputting the sample video data and sample text features included in each target sample data in the at least one target sample data into the video and text processing module included in the initial multimodal personality detection model to output a second modal fusion result, thereby obtaining at least one second modal fusion result; inputting each first modal fusion result in the at least one first modal fusion result and the second modal fusion result corresponding to the first modal fusion result into the prediction module included in the initial multimodal personality detection model to output a personality detection result, thereby obtaining at least one personality detection result; generating a training loss value based on the sample label and the personality detection result corresponding to each target sample data in the at least one target sample data through a preset mixed loss function; and determining the initial multimodal personality detection model after training as the multimodal personality detection model in response to determining that the training loss value is less than a preset training loss threshold.

[0011] In a fourth aspect, some embodiments of the present disclosure provide a user information push device, the device comprising: a first acquisition unit, configured to acquire user audio-visual data corresponding to a target user, wherein the user audio-visual data comprises user video data, user audio data and voice text data, and the target user corresponds to a user identifier; a standardization processing unit, configured to perform standardization processing on the user video data included in the user audio-visual data to obtain target video data; a feature extraction processing unit, configured to perform feature extraction processing on the user audio data included in the user audio-visual data to obtain target audio features; a text embedding processing unit, configured to perform text embedding processing on the voice text data included in the user audio-visual data The invention relates to a method for detecting a personality of a target user through a plurality of methods for detecting a personality of a target user. The method comprises: a first step of: performing text embedding processing to obtain target text features; an input unit configured to input the target video data, the target audio features and the target text features into a pre-trained multimodal personality detection model to obtain a personality detection result corresponding to the target user, wherein the multimodal personality detection model is pre-trained by the method described in any implementation of the first aspect; a second acquisition unit configured to obtain resume information corresponding to the target user in response to determining that the personality detection result meets the preset job requirement conditions; a determination and sending unit configured to determine the user identifier corresponding to the target user and the resume information as user information, and send the user information to the target terminal.

[0012] In a fifth aspect, some embodiments of the present disclosure provide an electronic device, comprising: one or more processors; a storage device on which one or more programs are stored, and when the one or more programs are executed by the one or more processors, the one or more processors implement the method described in any implementation of the first or second aspect above.

[0013] In a sixth aspect, some embodiments of the present disclosure provide a computer-readable medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the method described in any implementation of the first aspect or the second aspect is implemented.

[0014] The above-mentioned various embodiments of the present disclosure have the following beneficial effects: through the multimodal personality detection model training method of some embodiments of the present disclosure, the detection accuracy of the personality detection model can be improved and the occupation of computing resources can be reduced. Specifically, the reason for the low detection accuracy of the personality detection model or the occupation of more computing resources is that if the personality detection model is trained with single-modal sample data, the detection accuracy of the personality detection model will be low due to the relatively simple sample data. If the personality detection model is trained with multimodal sample data, the structure of the personality detection model will be more complicated due to the need to design different feature extraction branches for each different modality and introduce a shared encoder for information fusion, thereby resulting in the occupation of more computing resources. Based on this, the multimodal personality detection model training method of some embodiments of the present disclosure uses three modal data of video, audio and text to train the initial multimodal personality detection model. Among them, the initial multimodal personality detection model includes a video and audio processing module, a video and text processing module and a prediction module. The above-mentioned video and audio processing module includes a first feature extraction submodule, a first spatiotemporal feature enhancement submodule and a first multimodal fusion submodule. The above-mentioned video and text processing module includes a second feature extraction submodule, a second spatiotemporal feature enhancement submodule and a second multimodal fusion submodule. First, each original audio-visual data in the acquired original audio-visual data set is preprocessed to obtain a target sample data set suitable for inputting the initial multimodal personality detection model. Then, in the specific training process, the video and audio combination modality and the video and text combination modality are first used as two different processing branches, and feature extraction and feature fusion are performed respectively. Thus, the joint features after video and audio alignment and the joint features after video and text alignment can be obtained. And on this basis, through the first multimodal fusion submodule and the second multimodal fusion submodule, the cross-modal feature fusion of the video and audio joint features and the video and text joint features can be realized. Further, the two processing branches can output modal fusion results respectively. Further, the prediction module can be used to determine the final predicted personality detection result according to the two modal fusion results. Finally, a training loss value is determined according to the sample label and the predicted personality detection result, and in response to determining that the training loss value is less than a preset training loss threshold, the trained initial multimodal personality detection model is determined as a multimodal personality detection model. Therefore, the multimodal personality detection model training method of some embodiments of the present disclosure can improve the detection accuracy of the multimodal personality detection model by aligning the multimodal features of vision, audio, and text and fusing cross-modal features.Furthermore, because visual features usually contain more personality information, using vision and audio as combined modal branches, and vision and text as combined modal branches can reduce the number of feature extraction branches, and cross-modal fusion can be achieved without additionally introducing a shared encoder in the first multimodal fusion submodule and the second multimodal fusion submodule. Therefore, the structural complexity of the personality detection model can be reduced and the occupation of computing resources can be reduced. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] The above and other features, advantages and aspects of the embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the accompanying drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and that components and elements are not necessarily drawn to scale.

[0016] Figure 1 is a flowchart of some embodiments of the multimodal personality detection model training method according to the present disclosure;

[0017] Figure 2 is a schematic diagram of the structure of a multimodal personality detection model according to some embodiments of the present disclosure;

[0018] Figure 3 is a schematic structural diagram of a first multimodal fusion submodule of a multimodal personality detection model in some embodiments of the present disclosure;

[0019] Figure 4 is a model performance evaluation data graph of a multimodal personality detection model according to some embodiments of the present disclosure;

[0020] Figure 5 is a flow chart of some embodiments of the user information push method according to the present disclosure;

[0021] Figure 6 is a schematic diagram of the structure of some embodiments of the multimodal personality detection model training device according to the present disclosure;

[0022] Figure 7 is a schematic diagram of the structure of some embodiments of the user information push device according to the present disclosure;

[0023] Figure 8 It is a schematic diagram of the structure of an electronic device suitable for implementing some embodiments of the present disclosure. DETAILED DESCRIPTION

[0024] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as being limited to the embodiments set forth herein. On the contrary, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not intended to limit the scope of protection of the present disclosure.

[0025] It should also be noted that, for ease of description, only the parts related to the invention are shown in the drawings. In the absence of conflict, the embodiments and features in the embodiments of the present disclosure can be combined with each other.

[0026] It should be noted that the concepts such as "first" and "second" mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.

[0027] It should be noted that the modifications of "one" and "plurality" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, it should be understood as "one or more".

[0028] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.

[0029] With regard to the collection, storage, and use of user personal information (such as original audio-visual data sets and user audio-visual data) involved in this disclosure, before performing the corresponding operations, the relevant organizations or individuals shall fulfill their obligations, including conducting personal information security impact assessments, fulfilling the obligation to inform the personal information subject, and obtaining the authorization and consent of the personal information subject in advance.

[0030] The present disclosure will be described in detail below with reference to the accompanying drawings and in conjunction with embodiments.

[0031] Figure 1 The process of some embodiments of the multimodal personality detection model training method according to the present disclosure is shown. The multimodal personality detection model training method 100 includes the following steps:

[0032] Step 101: Obtain an original audio-visual dataset and an initial multimodal personality detection model.

[0033] In some embodiments, the execution subject (e.g., a computing device) of the multimodal personality detection model training method may obtain an original audio-visual data set and an initial multimodal personality detection model from a database through a wired connection or a wireless connection, wherein each original audio-visual data in the original audio-visual data set may include original video data, original audio data, voice text data, and sample labels. The original video data and the original audio data may be a video frame sequence and an audio frame sequence obtained by simultaneously performing image acquisition and voice acquisition on a target object through a recording device. The target object may be an individual for personality detection. The voice text data may be a text obtained by transcribing the original audio data. The sample label may be a personality trait score information sequence. The personality trait score information in the personality trait score information sequence may include a personality trait type identifier and a feature score. The personality trait type identifier may be a unique identifier of a personality trait type. The personality trait type may be a dimension in the Big Five personality model. For example, the personality trait type may be one of the following: openness, conscientiousness, extroversion, agreeableness, and neuroticism. The personality trait score information sequence may be an ordered set of personality trait score information arranged according to preset personality trait sequence information. The preset personality trait sequence information may be {"openness": 1, "conscientiousness": 2, "extraversion": 3, "agreeableness": 4, "neuroticism": 5}.

[0034] In addition, the initial multimodal personality detection model may include a video and audio processing module, a video and text processing module, and a prediction module. The video and audio processing module may include a first feature extraction submodule, a first spatiotemporal feature enhancement submodule, and a first multimodal fusion submodule. The video and text processing module may include a second feature extraction submodule, a second spatiotemporal feature enhancement submodule, and a second multimodal fusion submodule. The first feature extraction submodule may be used to perform dimension transformation and feature extraction on video data and audio features. The first spatiotemporal feature enhancement submodule may be used to extract joint features of vision and audio. The first multimodal fusion submodule may be used to share information with the second multimodal fusion submodule to achieve cross-modal feature fusion of video, audio, and text. The second feature extraction submodule may be used to perform dimension transformation and feature extraction on video data and text features. The second spatiotemporal feature enhancement submodule may be used to extract joint features of vision and text. The second multimodal fusion submodule may be used to share information with the first multimodal fusion submodule to achieve cross-modal feature fusion of video, audio, and text. The above prediction module can be used to output a predicted personality trait score sequence based on the fused multimodal features.

[0035] It should be noted that the above-mentioned wireless connection methods may include but are not limited to 3G / 4G connection, WiFi connection, Bluetooth connection, WiMAX connection, Zigbee connection, UWB (ultra wideband) connection, and other wireless connection methods currently known or to be developed in the future.

[0036] Optionally, the first multimodal fusion submodule and the second multimodal fusion submodule may each include a preset number of stacked layers. The preset number may be the number of pre-set stacked layers. Each stacked layer corresponds to a unique layer sequence number. The layer sequence number of the stacked layer may characterize the execution order of the stacked layer in the corresponding multimodal fusion submodule. For example, when the preset number is 5, the layer sequence number set of each stacked layer of the first multimodal fusion submodule and the layer sequence number set of each stacked layer of the second multimodal fusion submodule may be represented by {1, 2, 3, 4, 5}. Each stacked layer may include a spatiotemporal attention layer, a cross attention layer and a feedforward neural network layer. Among them, the spatiotemporal attention layer may be used to extract dynamic change features between frames and spatial static features within each video frame. The cross attention layer may be used to adjust the interaction strength between the video audio branch and the video text branch and realize information sharing through self-learning. The feedforward neural network layer may be used to extract features of three modal information: video, audio and text.

[0037] As an example, see the attached Figure 2 . Figure 2 An example structure of a multimodal personality detection model in some embodiments of the present disclosure is shown.

[0038] Step 102 : pre-process each original audio-visual data in the original audio-visual data set to obtain a target sample data set.

[0039] In some embodiments, the execution subject may pre-process each original audio-visual data in the original audio-visual data set in various ways to obtain a target sample data set. Each target sample data in the target sample data set may include sample video data, sample audio features, and sample text features. The sample video data may be the normalized original video data. The sample audio features may be the spectral features corresponding to the original audio data. The sample text features may be the speech text data after text embedding.

[0040] In some optional implementations of some embodiments, the execution subject may pre-process the original audio-visual data for each initial sample data in the initial sample data set through the following steps to generate target sample data in the target sample data set:

[0041] In the first step, the original video data included in the initial sample data is downsampled to obtain the downsampled video data. The downsampled video data may be the original video data after downsampling. The original video data included in the initial sample data may be downsampled by a preset downsampling processing method to obtain the downsampled video data. For example, the downsampling processing method may be one of the following: a content-based downsampling method and an adaptive downsampling method.

[0042] In practice, the above-mentioned execution entity reduces the number of video frames from more than 400 frames in each original video data to 100 frames through the above-mentioned downsampling processing method, which can reduce the burden of calculation and storage and also retain the key visual information of the video.

[0043] The second step is to perform frame extraction processing on the downsampled video data to obtain frame extracted video data. The frame extracted video data may be a video frame sequence with a preset number of frames. The preset number of frames may be a number of pre-set video frames. The frame extracted video data may be obtained by extracting a preset number of video frames from the downsampled video data by a random extraction method, and arranging the extracted video frames in chronological order.

[0044] The third step is to resize the above-mentioned video data after frame extraction to obtain target video data. The above-mentioned target video data may be a video frame sequence whose video frame resolution is a preset resolution. The above-mentioned preset resolution may be a preset resolution of a video frame. For example, the above-mentioned preset resolution may be 224 pixels*224 pixels. The above-mentioned video data after frame extraction may be resized by a preset image scaling algorithm to obtain the target video data. For example, the above-mentioned image scaling algorithm may be, but is not limited to, one of the following: nearest neighbor interpolation, bilinear interpolation, and bicubic interpolation.

[0045] In the fourth step, the original audio data included in the initial sample data is resampled to obtain resampled audio data. The resampled audio data may be audio data having a sampling rate of a preset sampling rate. The preset sampling rate may be a preset audio sampling rate. For example, the preset sampling rate may be 16 Hz. The original audio data included in the initial sample data may be resampled by a preset interpolation algorithm to obtain resampled audio data. For example, the interpolation algorithm may be, but is not limited to, one of the following: nearest neighbor interpolation, bilinear interpolation, and bicubic interpolation.

[0046] The fifth step is to perform spectrum feature extraction processing on the above-mentioned resampled audio data to obtain target audio features. Among them, the above-mentioned target audio features can be composed of the logarithmic Mel filter group feature sequences corresponding to the above-mentioned resampled audio data. The logarithmic Mel filter group features in each logarithmic Mel filter group feature sequence can be 128 dimensions. The above-mentioned resampled audio data can be subjected to spectrum feature extraction processing by a preset spectrum feature extraction processing method to obtain the target audio features. The above-mentioned spectrum feature extraction processing method can include the steps of pre-emphasis, framing, windowing, fast Fourier transform, calculating power spectrum, applying Mel filter group and taking logarithm.

[0047] The sixth step is to perform text embedding processing on the speech text data included in the above-mentioned initial sample data to obtain target text features. The above-mentioned target text features may be text embedding representations of preset embedding dimensions that characterize the above-mentioned speech text data. The above-mentioned preset embedding dimensions may be dimensions of pre-set text embedding representations. For example, the above-mentioned preset embedding dimensions may be 768 dimensions. The speech text data included in the above-mentioned initial sample data may be subjected to text embedding processing by a preset text embedding method to obtain target text features. For example, the above-mentioned text embedding method may be a BERT (Bidirectional Encoder Representations from Transformers) language representation model.

[0048] In the seventh step, the target video data, the target audio features and the target text features are determined as target sample data.

[0049] Step 103, selecting at least one target sample data from the target sample data set, and performing the following training steps based on the initial multimodal personality detection model:

[0050] Step 1031: input the sample video data and sample audio features included in each target sample data in at least one target sample data into a video and audio processing module included in an initial multimodal personality detection model to output a first modality fusion result, thereby obtaining at least one first modality fusion result.

[0051] In some embodiments, the execution subject may input the sample video data and sample audio features included in each target sample data in the at least one target sample data into the video and audio processing module included in the initial multimodal personality detection model in various ways to output a first modality fusion result, thereby obtaining at least one first modality fusion result. The first modality fusion result in the at least one first modality fusion result may be a joint feature after the fusion of video, audio and text output by the video and audio processing module.

[0052] In some optional implementations of some embodiments, the execution subject may input the sample video data and sample audio features included in each target sample data in the at least one target sample data into a video and audio processing module included in the initial multimodal personality detection model, and output a first modality fusion result through the following steps to obtain at least one first modality fusion result:

[0053] In the first step, the sample video data and sample audio features included in each target sample data in the at least one target sample data are input into the first feature extraction submodule included in the video and audio processing module to output an initial video and audio splicing feature, thereby obtaining at least one initial video and audio splicing feature. The first feature extraction submodule may include a convolutional neural network. The convolutional neural network included in the first feature extraction submodule is used to extract features of video data and features of audio data. The initial video and audio splicing feature may be a joint feature of the splicing of the video feature extraction result and the audio feature extraction result.

[0054] As an example, the above-mentioned execution entity can, for each target sample data in the above-mentioned at least one target sample data, perform feature extraction on the sample video data and sample audio features included in the above-mentioned target sample through the convolutional neural network included in the above-mentioned first feature extraction submodule, to obtain video feature extraction results and audio feature extraction results, and through a splicing operation, perform splicing processing on the above-mentioned video feature extraction results and audio feature extraction results along the channel dimension to obtain initial video and audio splicing features.

[0055] In the second step, for each of the at least one initial video and audio splicing feature, perform the following steps:

[0056] The first sub-step is to determine the first position coding feature and the first timing coding feature corresponding to the initial video and audio splicing feature. The first position coding feature can characterize the relative position relationship between elements in the initial video and audio splicing feature. The first timing coding feature can characterize the time dependency relationship between elements in the initial video and audio splicing feature.

[0057] As an example, the execution subject may first determine the first position coding feature corresponding to the initial video and audio splicing feature by using a sine and cosine function or an index based on a video frame. Then, the first temporal coding feature corresponding to the initial video and audio splicing feature may be determined by calculating the difference between adjacent frames and using a sequence modeling method such as a recurrent neural network or a long short-term memory network.

[0058] The second sub-step is to perform an addition operation on the above-mentioned initial video and audio splicing features, the above-mentioned first position coding features and the above-mentioned first timing coding features to obtain the initial video and audio splicing features with position and timing information as the target video and audio splicing features.

[0059] As an example, the above-mentioned execution entity can adopt the method of adding elements at corresponding positions to perform an addition operation on the above-mentioned initial video and audio splicing features, the above-mentioned first position coding features and the above-mentioned first timing coding features to obtain the initial video and audio splicing features with position and timing information as the target video and audio splicing features.

[0060] In the third step, each target video and audio splicing feature of the obtained at least one target video and audio splicing feature is input into the first spatiotemporal feature enhancement submodule included in the above-mentioned video and audio processing module to output a first bimodal feature enhancement result, and obtain at least one first bimodal feature enhancement result. Among them, the above-mentioned first spatiotemporal feature enhancement submodule may include a preset number of stacked layers. The above-mentioned preset number of layers may be the number of pre-set stacked layers. Each stacked layer of the above-mentioned first spatiotemporal feature enhancement submodule may include a spatiotemporal attention layer and a feedforward neural network layer. The above-mentioned spatiotemporal attention layer may be composed of a time and space attention module in a TimeSformer (Time-Space transformer) structure. Each of the above-mentioned at least one first bimodal feature enhancement result may be a joint feature of a visual and audio combination modality.

[0061] As an example, for each target video and audio splicing feature, the execution subject may output the target video and audio splicing feature to the first spatiotemporal feature enhancement submodule, and the target video and audio splicing feature is processed by the spatiotemporal attention module and the residual structure in each stacking layer to obtain the visual and audio combined modal joint feature. After that, the visual and audio combined modal joint feature is input into the feedforward neural network layer included in the first spatiotemporal feature enhancement submodule, and the first bimodal feature enhancement result is output.

[0062] The fourth step is to input each of the at least one first bimodal feature enhancement result into the first multimodal fusion submodule included in the video and audio processing module to output a first modal fusion result, thereby obtaining at least one first modal fusion result.

[0063] It should be noted that in order to avoid introducing a shared encoder into the first multimodal fusion submodule and the second multimodal fusion submodule, increasing the structural complexity of the multimodal personality detection model, and combining the technical advantages of the solution development team in the field of deep learning, this disclosure decides to adopt the following solution.

[0064] Optionally, the input data of any cross-attention layer in the first multimodal fusion submodule may include the query vector and target video text attention data output by the spatiotemporal attention layer in the same stacking layer. The target video text attention data may be the key vector and value vector output by the spatiotemporal attention layer in the second multimodal fusion submodule that matches the any cross-attention layer. The stacking layer that matches the any cross-attention layer may be the one whose layer number corresponding to the spatiotemporal attention layer in the second multimodal fusion submodule is the same as the layer number of the stacking layer where the any cross-attention layer is located.

[0065] Optionally, the input data of the feedforward neural network layer in the first multimodal fusion submodule may be determined by the following formula:

[0066]

[0067] Where k is the layer number of the stacked layer. β is a learnable control parameter. CA(·) represents the cross attention layer. Represents the joint feature output of the k-th layer of visual and audio combined modality. Represents the joint feature output of the k-1th layer of visual-text combination modality. Represents the joint feature output of the k-1th layer of visual and audio combined modality. Represents the joint feature output of the k-th layer of visual and audio combination modality after passing through the spatiotemporal attention layer.

[0068] As an example, see the attached Figure 3 .in, Figure 3 An example structure of a first multimodal fusion submodule of a multimodal personality detection model in some embodiments of the present disclosure is shown. Figure 3 The interaction data between the first multimodal fusion submodule and the second multimodal fusion submodule is also shown. a Represents the key vector corresponding to the joint feature of video and audio. a K represents the value vector corresponding to the joint feature of video and audio. t V represents the key vector corresponding to the joint feature of video and text. t Represents the value vector corresponding to the joint features of video and text.

[0069] The above-mentioned first modal fusion result generation step and its related contents are an inventive point of an embodiment of the present disclosure, which solves the above-mentioned technical problem of "increasing the structural complexity of the multimodal personality detection model". The reason for increasing the structural complexity of the multimodal personality detection model is that the shared encoder is introduced into the first multimodal fusion submodule. If this solution solves the above-mentioned problem, the effect of reducing the structural complexity of the personality detection model can be achieved. In order to achieve this effect, in each stacking layer in the first multimodal fusion submodule, the video and audio joint features corresponding to the current branch are extracted through the spatiotemporal attention layer, and then the extracted video and audio joint features are input into the cross attention layer together with the video and text joint features output by the spatiotemporal attention layer with the same layer number in the second multimodal fusion submodule, so as to realize the interaction and sharing of modal information between the two branches, and further realize the cross-modal fusion of the three modalities. Therefore, there is no need to add a shared encoder, so that the structural complexity of the multimodal personality detection model can be reduced. In addition, because the interaction intensity can be adjusted by self-learning according to the control parameters, the information interaction and sharing strategy between the modalities can be adjusted according to actual needs.

[0070] Step 1032: input the sample video data and sample text features included in each target sample data in at least one target sample data into the video text processing module included in the initial multimodal personality detection model to output a second modality fusion result, thereby obtaining at least one second modality fusion result.

[0071] In some embodiments, the execution subject may input the sample video data and sample text features included in each target sample data in the at least one target sample data into the video text processing module included in the initial multimodal personality detection model in various ways to output the second modality fusion result, and obtain at least one second modality fusion result. The second modality fusion result in the at least one second modality fusion result may be a joint feature after the video, audio and text are fused, which is output by the video text processing module.

[0072] In some optional implementations of some embodiments, the execution subject may input the sample video data and sample audio features included in each target sample data in the at least one target sample data into a video and audio processing module included in the initial multimodal personality detection model, and output a second modality fusion result through the following steps to obtain at least one second modality fusion result:

[0073] In the first step, the sample video data and sample text features included in each target sample data in the at least one target sample data are input into the second feature extraction submodule included in the video text processing module to output the initial video text splicing feature, and obtain at least one initial video text splicing feature. The second feature extraction submodule may include a convolutional neural network. The convolutional neural network included in the second feature extraction submodule is used to extract features of video data and features of text data. The initial video text splicing feature can be a joint feature after the video feature extraction result and the text feature extraction result are spliced.

[0074] As an example, the execution entity may perform feature extraction on the sample video data and sample text features included in the target sample for each target sample data in the at least one target sample data through the convolutional neural network included in the second feature extraction submodule to obtain video feature extraction results and text feature extraction results, and perform splicing processing on the video feature extraction results and text feature extraction results along the channel dimension through a splicing operation to obtain initial video text splicing features.

[0075] In the second step, for each of the at least one initial video text splicing feature, the following steps are performed:

[0076] The first sub-step is to determine the second position coding feature and the second timing coding feature corresponding to the initial video text splicing feature. The second position coding feature can characterize the relative position relationship between elements in the initial video text splicing feature. The second timing coding feature can characterize the time dependency relationship between elements in the initial video text splicing feature. The generation steps of the first position coding feature and the first timing coding feature can be referred to and will not be repeated here.

[0077] In the second sub-step, an addition operation is performed on the above-mentioned initial video text splicing feature and the corresponding second position coding feature and the corresponding second timing coding feature to obtain the initial video text splicing feature with position and timing information as the target video text splicing feature.

[0078] As an example, the above-mentioned execution entity can adopt the method of adding elements at corresponding positions to perform an addition operation on the above-mentioned initial video text splicing features and the corresponding second position coding features and the corresponding second timing coding features to obtain the initial video text splicing features with position and timing information as the target video text splicing features.

[0079] In the third step, each of the obtained at least one target video text splicing feature is input into the second spatiotemporal feature enhancement submodule included in the video text processing module to output a second bimodal feature enhancement result, thereby obtaining at least one second bimodal feature enhancement result. The second spatiotemporal feature enhancement submodule may include a spatiotemporal attention layer and a feedforward neural network layer. Each of the at least one second bimodal feature enhancement result may be a joint feature of the visual text combination modality.

[0080] As an example, for each target video text splicing feature, the execution subject may output the target video text splicing feature to the second spatiotemporal feature enhancement submodule, and the target video text splicing feature is processed by the spatiotemporal attention module and the residual structure in each stacking layer to obtain the visual text combined modal joint feature. After that, the visual text combined modal joint feature is input into the feedforward neural network layer included in the second spatiotemporal feature enhancement submodule, and the second bimodal feature enhancement result is output.

[0081] The fourth step is to input each of the at least one second bimodal feature enhancement results into the second multimodal fusion submodule included in the video text processing module to output a second modal fusion result, thereby obtaining at least one second modal fusion result.

[0082] Optionally, the input data of any cross-attention layer in the second multimodal fusion submodule may include the query vector and target video and audio attention data output by the spatiotemporal attention layer in the same stacking layer. The target video and audio attention data may be the key vector and value vector output by the spatiotemporal attention layer in the first multimodal fusion submodule that matches any cross-attention layer in the second multimodal fusion submodule. Matching with any cross-attention layer in the second multimodal fusion submodule may be: the layer number of the stacking layer corresponding to the spatiotemporal attention layer in the first multimodal fusion submodule is the same as the layer number of the stacking layer where any cross-attention layer in the second multimodal fusion submodule is located. The input data of the feedforward neural network layer in the second multimodal fusion submodule may be determined by the following formula:

[0083]

[0084] Here, α represents a learnable control parameter. Represents the final joint feature output of the k-th layer of visual-textual combination modality. Represents the joint feature output of the k-th layer of visual-text combination modality after passing through the spatiotemporal attention layer.

[0085] Step 1033: Input each first modal fusion result in at least one first modal fusion result and the second modal fusion result corresponding to the first modal fusion result into a prediction module included in the initial multimodal personality detection model to output a personality detection result, thereby obtaining at least one personality detection result.

[0086] In some embodiments, the execution subject may input each first modal fusion result in the at least one first modal fusion result and the second modal fusion result corresponding to the first modal fusion result into the prediction module included in the initial multimodal personality detection model in various ways to output the personality detection result, thereby obtaining at least one personality detection result. The personality detection result in the at least one personality detection result may be a predicted personality trait score sequence. Each personality trait score in the personality trait score sequence may be a predicted score on the corresponding personality trait type.

[0087] In some optional implementations of some embodiments, the prediction module may include a multi-layer perceptron submodule. The multi-layer perceptron submodule may be used to predict the score of each personality trait type. The execution subject may input each first modal fusion result of the at least one first modal fusion result and the second modal fusion result corresponding to the first modal fusion result into the prediction module included in the initial multimodal personality detection model, and output the personality detection result through the following steps to obtain at least one personality detection result:

[0088] In the first step, for each of the at least one first modality fusion result, the following steps are performed:

[0089] The first sub-step is to input the first modality fusion result into the multi-layer perceptron sub-module included in the prediction module to obtain a first detection result. The first detection result may be a personality trait score sequence predicted based on the first modality fusion result.

[0090] The second sub-step is to input the second modality fusion result corresponding to the first modality fusion result into the multi-layer perceptron sub-module included in the prediction module to obtain a second detection result. The second detection result may be a personality trait score sequence predicted based on the second modality fusion result.

[0091] The third sub-step is to generate a personality detection result based on the first detection result and the second detection result. The personality detection result can be obtained by averaging the corresponding position elements of the first detection result and the second detection result by arithmetic averaging.

[0092] Step 1034: Generate a training loss value based on a sample label and a personality detection result corresponding to each target sample data in at least one target sample data by using a preset hybrid loss function.

[0093] In some embodiments, the execution entity may generate a training loss value based on a sample label and a personality detection result corresponding to each target sample data in the at least one target sample data by means of a preset hybrid loss function.

[0094] Optionally, the above hybrid loss function can be:

[0095]

[0096] Where L represents the training loss value. N represents the total number of target sample data selected in one iteration of training. i represents the sequence number of the target sample data. λ represents the learnable hyperparameter. y i Represents the sample label corresponding to the i-th target sample data. represents the personality detection result corresponding to the predicted i-th target sample data. Tanh(·) represents the hyperbolic tangent function. ||·||2 represents the second norm.

[0097] Step 1035 , in response to determining that the training loss value is less than a preset training loss threshold, determining the trained initial multimodal personality detection model as the multimodal personality detection model.

[0098] In some embodiments, the execution subject may determine the trained initial multimodal personality detection model as the multimodal personality detection model in response to determining that the training loss value is less than a preset training loss threshold, wherein the preset training loss threshold may be an upper limit value of the pre-set training loss value.

[0099] Optionally, the execution subject may also adjust the network parameters of the initial multimodal personality detection model in response to determining that the training loss value is not less than the preset training loss threshold, and use unused target sample data to form a target sample data set, and use the adjusted initial multimodal personality detection model as the initial multimodal personality detection model to perform the training step again. The network parameters of the initial multimodal personality detection model may be adjusted by a gradient descent optimization method.

[0100] As an example, Figure 4 The model performance evaluation data of the multimodal personality detection model in some embodiments of the present disclosure are shown. Among them, Figure 4 There are two training strategies: cross-modal interaction and non-cross-modal interaction. Cross-modal interaction is the training strategy adopted by the disclosed method. Each training strategy corresponds to three feature fusion methods: averaging, concatenation, and attention mechanism. Figure 4 It can be seen that after the cross-modal interaction of the disclosed method, the dual-branch training model using the feature splicing strategy has the best average personality prediction performance.

[0101] The above-mentioned various embodiments of the present disclosure have the following beneficial effects: through the multimodal personality detection model training method of some embodiments of the present disclosure, the detection accuracy of the personality detection model can be improved and the occupation of computing resources can be reduced. Specifically, the reason for the low detection accuracy of the personality detection model or the occupation of more computing resources is that if the personality detection model is trained with single-modal sample data, the detection accuracy of the personality detection model will be low due to the relatively simple sample data. If the personality detection model is trained with multimodal sample data, the structure of the personality detection model will be more complicated due to the need to design different feature extraction branches for each different modality and introduce a shared encoder for information fusion, thereby resulting in the occupation of more computing resources. Based on this, the multimodal personality detection model training method of some embodiments of the present disclosure uses three modal data of video, audio and text to train the initial multimodal personality detection model. Among them, the initial multimodal personality detection model includes a video and audio processing module, a video and text processing module and a prediction module. The above-mentioned video and audio processing module includes a first feature extraction submodule, a first spatiotemporal feature enhancement submodule and a first multimodal fusion submodule. The above-mentioned video and text processing module includes a second feature extraction submodule, a second spatiotemporal feature enhancement submodule and a second multimodal fusion submodule. First, each original audio-visual data in the acquired original audio-visual data set is preprocessed to obtain a target sample data set suitable for inputting the initial multimodal personality detection model. Then, in the specific training process, the video and audio combination modality and the video and text combination modality are first used as two different processing branches, and feature extraction and feature fusion are performed respectively. Thus, the joint features after video and audio alignment and the joint features after video and text alignment can be obtained. And on this basis, through the first multimodal fusion submodule and the second multimodal fusion submodule, the cross-modal feature fusion of the video and audio joint features and the video and text joint features can be realized. Further, the two processing branches can output modal fusion results respectively. Further, the prediction module can be used to determine the final predicted personality detection result according to the two modal fusion results. Finally, a training loss value is determined according to the sample label and the predicted personality detection result, and in response to determining that the training loss value is less than a preset training loss threshold, the trained initial multimodal personality detection model is determined as a multimodal personality detection model. Therefore, the multimodal personality detection model training method of some embodiments of the present disclosure can improve the detection accuracy of the multimodal personality detection model by aligning the multimodal features of vision, audio, and text and fusing cross-modal features.Furthermore, because visual features usually contain more personality information, using vision and audio as combined modal branches, and vision and text as combined modal branches can reduce the number of feature extraction branches, and cross-modal fusion can be achieved without additionally introducing a shared encoder in the first multimodal fusion submodule and the second multimodal fusion submodule. Therefore, the structural complexity of the personality detection model can be reduced and the occupation of computing resources can be reduced.

[0102] Figure 5 The process of some embodiments of the user information push method according to the present disclosure is shown. The user information push method 500 includes the following steps:

[0103] Step 501: Acquire user audiovisual data corresponding to a target user.

[0104] In some embodiments, the execution subject (such as a computing device) of the user information push method can obtain the user audio-visual data of the corresponding target user from the database through a wired connection or a wireless connection. Among them, the above-mentioned user audio-visual data may include user video data, user audio data and voice text data. The above-mentioned target user may correspond to the user identifier one by one. The above-mentioned target user may be an individual to be subjected to personality detection. The above-mentioned user video data and the above-mentioned user audio data may be a video frame sequence and an audio frame sequence obtained by performing image acquisition and voice acquisition on the above-mentioned target user through a recording device. The above-mentioned voice text data may be a text obtained by transcribing the above-mentioned user audio data. It should be noted that the above-mentioned wireless connection method may include but is not limited to 3G / 4G connection, WiFi connection, Bluetooth connection, WiMAX connection, Zigbee connection, UWB (ultra wideband) connection, and other wireless connection methods currently known or to be developed in the future.

[0105] Step 502: Standardize the user video data included in the user audio-visual data to obtain target video data.

[0106] In some embodiments, the execution subject may perform standardization processing on the user video data included in the user audio-visual data to obtain target video data. The target video data may be user video data after scaling and resizing. The user video data included in the user audio-visual data may be standardized by a preset standardization processing method to obtain target video data. The standardization processing method may include steps such as downsampling processing, frame extraction processing, and image scaling. Specifically, the user video data included in the user audio-visual data may be standardized with reference to step 102 to obtain target video data, which will not be described in detail herein.

[0107] Step 503: Perform feature extraction processing on the user audio data included in the user audio-visual data to obtain target audio features.

[0108] In some embodiments, the execution subject may perform feature extraction processing on the user audio data included in the user audio-visual data to obtain target audio features. The target audio features may be a logarithmic Mel filter bank feature sequence of a preset feature dimension. The user audio data included in the user audio-visual data may be subjected to spectrum feature extraction processing by a preset feature extraction method to obtain target audio features. The feature extraction method may include steps such as resampling processing and spectrum feature extraction processing. Specifically, referring to step 102, feature extraction processing may be performed on the user audio data included in the user audio-visual data to obtain target audio features, which will not be described in detail herein.

[0109] Step 504: Perform text embedding processing on the speech text data included in the user's audio-visual data to obtain target text features.

[0110] In some embodiments, the execution subject may perform text embedding processing on the voice text data included in the user audio-visual data to obtain target text features. The target text features may be embedded representations of the voice text data. The voice text data included in the user audio-visual data may be embedded in the text embedding method to obtain target text features.

[0111] Step 505: input the target video data, target audio features, and target text features into a pre-trained multimodal personality detection model to obtain a personality detection result corresponding to the target user.

[0112] In some embodiments, the execution subject may input the target video data, the target audio features, and the target text features into a pre-trained multimodal personality detection model to obtain a personality detection result corresponding to the target user. The multimodal personality detection model may be pre-trained by the method of steps 103-1035. The personality detection result may be a personality trait score sequence corresponding to the target user. For example, the personality detection result may be {0.912, 0.718, 0.883, 0.794, 0.156}.

[0113] Step 506, in response to determining that the personality detection result meets the preset job requirement conditions, obtaining the resume information corresponding to the target user.

[0114] In some embodiments, the execution subject may obtain the resume information corresponding to the target user in response to determining that the personality test result meets the preset job requirement. The preset job requirement may be that the personality test result matches the preset job personality trait type score threshold information group. The job personality trait type score threshold information in the job personality trait type score threshold information group may include a reference personality trait type identifier, a rule identifier, and a score threshold. The reference personality trait type identifier may be an identifier of a personality trait type for reference. The rule identifier may be one of the following: 0, 1. When the rule identifier is 0, the personality trait score corresponding to the reference personality trait type identifier in the personality test result representing the target user must be less than the score threshold corresponding to the reference personality trait type identifier. When the rule identifier is 1, the personality trait score corresponding to the reference personality trait type identifier in the personality test result representing the target user must be greater than the score threshold corresponding to the reference personality trait type identifier. The resume information may be information about the resume of the target user obtained from a database.

[0115] As an example, the execution subject may firstly perform detection processing on the personality detection result through a preset information detection interface based on the above-mentioned job personality trait type score threshold information group to obtain a detection result. The above-mentioned information detection interface may be encapsulated with an information detection function. The above-mentioned information detection function may detect whether each personality trait score meets the threshold rule requirements, and when each personality trait score meets the threshold rule requirements, a preset qualified mark is determined as the detection result. The above-mentioned preset qualified mark may indicate that the personality trait score matches the job requirements. Then, when the detection result includes the preset qualified mark, it may be determined that the above-mentioned personality detection result meets the preset job requirement conditions.

[0116] Step 507: determine the user identification and resume information corresponding to the target user as user information, and send the user information to the target terminal.

[0117] In some embodiments, the execution subject may determine the user identification corresponding to the target user and the resume information as user information, and send the user information to a target terminal, wherein the target terminal may be a terminal for displaying resume information.

[0118] The above-mentioned embodiments of the present disclosure have the following beneficial effects: through the user information push method of some embodiments of the present disclosure, personnel matching positions can be screened according to personality traits, which can improve the matching accuracy between personnel and positions. Thus, the quality and efficiency of personnel selection can be improved. In addition, because only the personnel information whose personality traits meet the requirements of the position is pushed to the target terminal, the occupation of transmission resources can also be reduced.

[0119] Further references Figure 6 As an implementation of the methods shown in the above figures, the present disclosure provides some embodiments of a multimodal personality detection model training device. These device embodiments are similar to Figure 1 Corresponding to the method embodiments shown, the multimodal personality detection model training device 600 can be specifically applied to various electronic devices.

[0120] like Figure 6As shown, the multimodal personality detection model training device 600 of some embodiments includes: an acquisition unit 601, a preprocessing unit 602 and a selection and execution unit 603. The acquisition unit 601 is configured to acquire an original audio-visual data set and an initial multimodal personality detection model, wherein each original audio-visual data in the original audio-visual data set includes original video data, original audio data, speech text data and sample labels, and the initial multimodal personality detection model includes a video and audio processing module, a video and text processing module and a prediction module, wherein the video and audio processing module includes a first feature extraction submodule, a first spatiotemporal feature enhancement submodule and a first multimodal fusion submodule, and the video and text processing module includes a second feature extraction submodule, a second spatiotemporal feature enhancement submodule and a second multimodal fusion submodule, and the first multimodal fusion submodule and the second multimodal fusion submodule are both used to fuse features of modalities such as video, audio and text; the preprocessing unit 602 is configured to preprocess each original audio-visual data in the original audio-visual data set to obtain a target sample data set, wherein each target sample data in the target sample data set includes sample video data, sample audio features and sample text features; the selection and execution unit 603 is configured to select at least one target sample data from the target sample data set, and to select at least one target sample data based on the initial multimodal The method comprises the steps of: inputting the sample video data and the sample audio features included in each target sample data in the at least one target sample data into the video and audio processing module included in the initial multimodal personality detection model to output a first modal fusion result, thereby obtaining at least one first modal fusion result; inputting the sample video data and the sample text features included in each target sample data in the at least one target sample data into the video and text processing module included in the initial multimodal personality detection model to output a second modal fusion result, thereby obtaining at least one second modal fusion result; inputting each first modal fusion result in the at least one first modal fusion result and the second modal fusion result corresponding to the first modal fusion result into the prediction module included in the initial multimodal personality detection model to output a personality detection result, thereby obtaining at least one personality detection result; generating a training loss value based on the sample label and the personality detection result corresponding to each target sample data in the at least one target sample data through a preset mixed loss function; and determining the initial multimodal personality detection model after training as the multimodal personality detection model in response to determining that the training loss value is less than a preset training loss threshold.

[0121] It can be understood that the units recorded in the multimodal personality detection model training device 600 are similar to the reference Figure 1Therefore, the operations, features and beneficial effects described above for the method are also applicable to the multimodal personality detection model training device 600 and the units contained therein, and will not be described in detail here.

[0122] Further references Figure 7 As an implementation of the methods shown in the above figures, the present disclosure provides some embodiments of a user information push device, and these device embodiments are similar to Figure 5 Corresponding to the method embodiments shown, the user information pushing device 700 can be specifically applied to various electronic devices.

[0123] like Figure 7 As shown, the user information push device 700 of some embodiments includes: a first acquisition unit 701, a standardization processing unit 702, a feature extraction processing unit 703, a text embedding processing unit 704, an input unit 705, a second acquisition unit 706 and a determination and sending unit 707. The first acquisition unit 701 is configured to acquire user audio-visual data corresponding to a target user, wherein the user audio-visual data includes user video data, user audio data and voice text data, and the target user corresponds to a user identifier; the standardization processing unit 702 is configured to perform standardization processing on the user video data included in the user audio-visual data to obtain target video data; the feature extraction processing unit 703 is configured to perform feature extraction processing on the user audio data included in the user audio-visual data to obtain target audio features; the text embedding processing unit 704 is configured to perform text embedding processing on the voice text data included in the user audio-visual data to obtain target text Features; an input unit 705, configured to input the target video data, the target audio features and the target text features into a pre-trained multimodal personality detection model to obtain a personality detection result corresponding to the target user, wherein the multimodal personality detection model is pre-trained by the method described in any implementation of the first aspect; a second acquisition unit 706, configured to obtain resume information corresponding to the target user in response to determining that the personality detection result meets the preset job requirement conditions; a determination and sending unit 707, configured to determine the user identifier corresponding to the target user and the resume information as user information, and send the user information to the target terminal.

[0124] It is understandable that the units recorded in the user information push device 700 are similar to those in the reference Figure 5 Therefore, the operations, features and beneficial effects described above for the method are also applicable to the user information push device 700 and the units included therein, and will not be described in detail here.

[0125] Further references Figure 8 , which shows a structural schematic diagram of an electronic device 800 suitable for implementing some embodiments of the present disclosure. Figure 8 The electronic device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present disclosure.

[0126] like Figure 8 As shown, the electronic device 800 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 801, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 802 or a program loaded from a storage device 808 into a random access memory (RAM) 803. In the RAM 803, various programs and data required for the operation of the electronic device 800 are also stored. The processing device 801, the ROM 802, and the RAM 803 are connected to each other via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.

[0127] Typically, the following devices may be connected to the I / O interface 805: input devices 806 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; output devices 807 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; storage devices 808 including, for example, a magnetic tape, a hard disk, etc.; and communication devices 809. The communication device 809 may allow the electronic device 800 to communicate with other devices wirelessly or by wire to exchange data. Although Figure 8 The electronic device 800 is shown with various devices, but it should be understood that it is not required to implement or possess all the devices shown. More or fewer devices may be implemented or possessed instead. Figure 8 Each block shown in the figure may represent one device, or may represent multiple devices as required.

[0128] In particular, according to some embodiments of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, some embodiments of the present disclosure include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In some such embodiments, the computer program can be downloaded and installed from the network through the communication device 809, or installed from the storage device 808, or installed from the ROM 802. When the computer program is executed by the processing device 801, the above-mentioned functions defined in the method of some embodiments of the present disclosure are executed.

[0129] It should be noted that the computer-readable medium in some embodiments of the present disclosure may be a computer-readable signal medium or a computer-readable storage medium or any combination of the two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In some embodiments of the present disclosure, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in combination with an instruction execution system, device or device. In some embodiments of the present disclosure, the computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, in which a computer-readable program code is carried. This propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer readable signal medium may also be any computer readable medium other than a computer readable storage medium, which may send, propagate or transmit a program for use by or in conjunction with an instruction execution system, apparatus or device. The program code contained on the computer readable medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.

[0130] In some embodiments, the client and the server may communicate using any currently known or future developed network protocol such as HTTP (HyperText Transfer Protocol), and may be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), an internet (e.g., the Internet), and a peer-to-peer network (e.g., an ad hoc peer-to-peer network), as well as any currently known or future developed network.

[0131] The computer-readable medium may be included in the device, or may exist independently without being installed in the electronic device. The computer-readable medium carries one or more programs. When the one or more programs are executed by the electronic device, the electronic device may execute:

[0132] Computer program code for performing the operations of some embodiments of the present disclosure may be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, C++, and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0133] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present disclosure. In this regard, each square box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some implementations as replacements, the functions marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two square boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0134] The units described in some embodiments of the present disclosure may be implemented by software or hardware. The units described may also be set in a processor, for example, it may be described as: a processor includes: an acquisition unit and a selection and execution unit. The names of these units do not constitute a limitation on the units themselves in some cases, for example, the acquisition unit may also be described as a "unit".

[0135] The functions described above herein may be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), and the like.

[0136] The above descriptions are only some preferred embodiments of the present disclosure and an explanation of the technical principles used. Those skilled in the art should understand that the scope of the invention involved in the embodiments of the present disclosure is not limited to the technical solutions formed by a specific combination of the above-mentioned technical features, but should also cover other technical solutions formed by any combination of the above-mentioned technical features or their equivalent features without departing from the above-mentioned inventive concept. For example, the above-mentioned features are replaced with the technical features with similar functions disclosed in the embodiments of the present disclosure (but not limited to) and the technical solutions formed.

Claims

1. A multimodal personality detection model training method, comprising: Acquire an original audio-visual data set and an initial multimodal personality detection model, wherein each original audio-visual data in the original audio-visual data set includes original video data, original audio data, speech text data and sample labels, and the initial multimodal personality detection model includes a video and audio processing module, a video and text processing module and a prediction module, the video and audio processing module includes a first feature extraction submodule, a first spatiotemporal feature enhancement submodule and a first multimodal fusion submodule, the video and text processing module includes a second feature extraction submodule, a second spatiotemporal feature enhancement submodule and a second multimodal fusion submodule, and the first multimodal fusion submodule and the second multimodal fusion submodule are both used to perform cross-modal fusion on features of modalities such as video, audio and text; Preprocessing each original audio-visual data in the original audio-visual data set to obtain a target sample data set, wherein each target sample data in the target sample data set includes sample video data, sample audio features, and sample text features; At least one target sample data is selected from the target sample data set, and the following training steps are performed based on the initial multimodal personality detection model: Inputting the sample video data and the sample audio features included in each target sample data in the at least one target sample data into the video and audio processing module included in the initial multimodal personality detection model to output a first modality fusion result, thereby obtaining at least one first modality fusion result; Inputting the sample video data and sample text features included in each target sample data in the at least one target sample data into a video text processing module included in the initial multimodal personality detection model to output a second modality fusion result, thereby obtaining at least one second modality fusion result; Inputting each first modality fusion result of the at least one first modality fusion result and a second modality fusion result corresponding to the first modality fusion result into a prediction module included in the initial multimodal personality detection model to output a personality detection result, thereby obtaining at least one personality detection result; Generate a training loss value based on a sample label and a personality detection result corresponding to each target sample data in the at least one target sample data by using a preset hybrid loss function; In response to determining that the training loss value is less than a preset training loss threshold, the trained initial multimodal personality detection model is determined as the multimodal personality detection model.

2. The method according to claim 1, wherein: The method further comprises: In response to determining that the training loss value is not less than the preset training loss threshold, adjusting the network parameters of the initial multimodal personality detection model, using unused target sample data to form a target sample data set, using the adjusted initial multimodal personality detection model as the initial multimodal personality detection model, and performing the training step again.

3. The method according to claim 1, wherein: The preprocessing of each original audio-visual data in the original audio-visual data set to obtain a target sample data set includes: For each initial sample data in the initial sample data set, the following steps are performed: Downsampling the original video data included in the initial sample data to obtain downsampled video data; Performing frame extraction processing on the down-sampled video data to obtain frame extracted video data; Resizing the frame-extracted video data to obtain target video data; Resampling the original audio data included in the initial sample data to obtain resampled audio data; Performing spectrum feature extraction processing on the resampled audio data to obtain target audio features; Performing text embedding processing on the speech text data included in the initial sample data to obtain target text features; The target video data, the target audio features and the target text features are determined as target sample data.

4. The method according to claim 1, wherein: The step of inputting the sample video data and the sample audio features included in each target sample data in the at least one target sample data into the video and audio processing module included in the initial multimodal personality detection model to output a first modality fusion result, and obtaining at least one first modality fusion result, comprises: Inputting the sample video data and the sample audio features included in each target sample data of the at least one target sample data into the first feature extraction submodule included in the video and audio processing module to output the initial video and audio splicing features, thereby obtaining at least one initial video and audio splicing feature; For each of the at least one initial video and audio splicing feature, the following steps are performed: Determine a first position coding feature and a first timing coding feature corresponding to the initial video and audio splicing feature; Performing an addition operation on the initial video and audio splicing feature, the first position coding feature and the first timing coding feature to obtain an initial video and audio splicing feature with position and timing information as a target video and audio splicing feature; Inputting each target video and audio splicing feature of the obtained at least one target video and audio splicing feature into a first spatiotemporal feature enhancement submodule included in the video and audio processing module to output a first bimodal feature enhancement result, thereby obtaining at least one first bimodal feature enhancement result; Each of the at least one first bimodal feature enhancement result is input into the first multimodal fusion submodule included in the video and audio processing module to output a first modal fusion result, thereby obtaining at least one first modal fusion result.

5. The method according to any one of claims 1 to 4, wherein: The step of inputting the sample video data and sample text features included in each target sample data in the at least one target sample data into a video text processing module included in the initial multimodal personality detection model to output a second modality fusion result, and obtaining at least one second modality fusion result, comprises: Inputting the sample video data and the sample text features included in each target sample data in the at least one target sample data into the second feature extraction submodule included in the video text processing module to output the initial video text splicing feature, thereby obtaining at least one initial video text splicing feature; For each of the at least one initial video text splicing feature, the following steps are performed: Determine a second position coding feature and a second temporal coding feature corresponding to the initial video text splicing feature; Performing an addition operation on the initial video text splicing feature and the corresponding second position coding feature and the corresponding second time sequence coding feature to obtain the initial video text splicing feature with position and time sequence information as the target video text splicing feature; Inputting each target video text splicing feature of the obtained at least one target video text splicing feature into a second spatiotemporal feature enhancement submodule included in the video text processing module to output a second bimodal feature enhancement result, thereby obtaining at least one second bimodal feature enhancement result; Each of the at least one second bimodal feature enhancement result is input into the second multimodal fusion submodule included in the video text processing module to output a second modal fusion result, thereby obtaining at least one second modal fusion result.

6. The method according to claim 1, wherein: The prediction module includes a multi-layer perceptron submodule; And the step of inputting each first modal fusion result of the at least one first modal fusion result and the second modal fusion result corresponding to the first modal fusion result into a prediction module included in the initial multimodal personality detection model to output a personality detection result, and obtaining at least one personality detection result includes: For each first modality fusion result of the at least one first modality fusion result, the following steps are performed: Inputting the first modality fusion result into the multi-layer perceptron submodule included in the prediction module to obtain a first detection result; Inputting the second modality fusion result corresponding to the first modality fusion result into the multi-layer perceptron submodule included in the prediction module to obtain a second detection result; A personality detection result is generated based on the first detection result and the second detection result.

7. The method according to claim 1, wherein: The first multimodal fusion submodule and the second multimodal fusion submodule each include a preset number of stacked layers, each stacked layer includes a spatiotemporal attention layer, a cross attention layer and a feedforward neural network layer.

8. A method for pushing user information, comprising: Acquire user audiovisual data corresponding to a target user, wherein the user audiovisual data includes user video data, user audio data and voice text data, and the target user corresponds to a user identifier; Performing standardization processing on the user video data included in the user audio-visual data to obtain target video data; Performing feature extraction processing on the user audio data included in the user audio-visual data to obtain target audio features; Performing text embedding processing on the speech text data included in the user audio-visual data to obtain target text features; Inputting the target video data, the target audio features, and the target text features into a pre-trained multimodal personality detection model to obtain a personality detection result corresponding to the target user, wherein the multimodal personality detection model is pre-trained by the method described in any one of claims 1 to 7; In response to determining that the personality detection result meets the preset job requirement conditions, obtaining resume information corresponding to the target user; The user identification corresponding to the target user and the resume information are determined as user information, and the user information is sent to a target terminal.

9. A multimodal personality detection model training device, comprising: An acquisition unit is configured to acquire an original audio-visual data set and an initial multimodal personality detection model, wherein each original audio-visual data in the original audio-visual data set includes original video data, original audio data, voice text data and sample labels, and the initial multimodal personality detection model includes a video and audio processing module, a video and text processing module and a prediction module, the video and audio processing module includes a first feature extraction submodule, a first spatiotemporal feature enhancement submodule and a first multimodal fusion submodule, the video and text processing module includes a second feature extraction submodule, a second spatiotemporal feature enhancement submodule and a second multimodal fusion submodule, and the first multimodal fusion submodule and the second multimodal fusion submodule are both used to perform cross-modal fusion on features of modalities such as video, audio and text; A preprocessing unit is configured to preprocess each original audio-visual data in the original audio-visual data set to obtain a target sample data set, wherein each target sample data in the target sample data set includes sample video data, sample audio features and sample text features; The selection and execution unit is configured to select at least one target sample data from the target sample data set, and perform the following training steps based on the initial multimodal personality detection model: Inputting the sample video data and the sample audio features included in each target sample data in the at least one target sample data into the video and audio processing module included in the initial multimodal personality detection model to output a first modality fusion result, thereby obtaining at least one first modality fusion result; Inputting the sample video data and sample text features included in each target sample data in the at least one target sample data into a video text processing module included in the initial multimodal personality detection model to output a second modality fusion result, thereby obtaining at least one second modality fusion result; Inputting each first modality fusion result of the at least one first modality fusion result and a second modality fusion result corresponding to the first modality fusion result into a prediction module included in the initial multimodal personality detection model to output a personality detection result, thereby obtaining at least one personality detection result; Generate a training loss value based on a sample label and a personality detection result corresponding to each target sample data in the at least one target sample data by using a preset hybrid loss function; In response to determining that the training loss value is less than a preset training loss threshold, the trained initial multimodal personality detection model is determined as the multimodal personality detection model.

10. A user information push device, comprising: A first acquisition unit is configured to acquire user audiovisual data corresponding to a target user, wherein the user audiovisual data includes user video data, user audio data and voice text data, and the target user corresponds to a user identifier; a standardization processing unit configured to perform standardization processing on the user video data included in the user audiovisual data to obtain target video data; a feature extraction processing unit configured to perform feature extraction processing on the user audio data included in the user audio-visual data to obtain a target audio feature; A text embedding processing unit, configured to perform text embedding processing on the speech text data included in the user audio-visual data to obtain target text features; an input unit, configured to input the target video data, the target audio feature, and the target text feature into a pre-trained multimodal personality detection model to obtain a personality detection result corresponding to the target user, wherein the multimodal personality detection model is pre-trained by the method according to any one of claims 1 to 7; A second acquisition unit is configured to acquire resume information corresponding to the target user in response to determining that the personality detection result meets the preset job requirement condition; The determining and sending unit is configured to determine the user identification corresponding to the target user and the resume information as user information, and send the user information to the target terminal.