Multi-modal data-based micro-movie user portrait analysis system and multi-modal data-based micro-movie user portrait analysis method

Through multimodal data acquisition and deep forest model generation of dynamic user portraits, the problem of insufficient single-modal data in traditional user portrait analysis systems is solved, and higher precision and dynamic user portrait analysis is achieved.

CN120354262AActive Publication Date: 2025-07-22SHANGHAI FANGONG CULTURE MEDIA CO LTD

Patent Information

Application Number
CN202510273497.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-10
Publication Date
2025-07-22
Estimated Expiration
2045-03-10

AI Technical Summary

Technical Problem

Traditional user portrait analysis systems rely on single modal data and cannot fully portray the user's complex behavior characteristics in micro-short drama scenes. The static tag system cannot dynamically reflect user interest migration, and lacks efficient cross-modal feature alignment and fusion methods.

Method used

Multimodal data acquisition, hierarchical attention network model and online deep forest model are used to obtain multimodal micro-short drama user data, feature extraction and fusion are performed, dynamic user portraits are generated, and a portrait analysis report is generated using Shapley values.

Benefits of technology

It improves the accuracy and richness of user portraits, can dynamically respond to changes in user preferences, and generate more accurate user portrait analysis reports.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120354262A_ABST
    Figure CN120354262A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data analysis and processing, and discloses a multi-modal data-based micro and short play user portrait analysis system and a multi-modal data-based micro and short play user portrait analysis method, and the system comprises a data acquisition module which is used for acquiring multi-modal micro and short play user data; the feature extraction module is used for performing feature extraction according to the multi-modal micro-movie user data to obtain a plurality of user features; the fusion module is used for inputting the plurality of user features into a pre-trained hierarchical attention network model to obtain a joint feature vector; the dynamic portrait generation module is used for inputting the joint feature vector into a preset online deep forest model to generate a dynamic portrait; and the portrait analysis module is used for calculating the Shapril value of each dimension of the dynamic portrait and generating a portrait analysis report according to the Shapril value. Compared with single-mode data portraits which are commonly adopted in the prior art and are richer in dimensionality, the dynamic portraits are generated by adopting the deep forest model, and the portraits can be modified in time according to user preferences.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data analysis and processing, and in particular, to a system and method for analyzing the user portrait of micro-short dramas based on multi-modal data. Background Art

[0002] Traditional user portrait analysis systems mainly rely on single-modal data (such as click-through rate, viewing duration, etc.), and it is difficult to comprehensively depict the complex behavioral characteristics of users in the micro-short drama scenario. Micro-short dramas have characteristics such as fragmented content, high-frequency interaction, and intensive emotional expression. The existing technologies have the following problems:

[0003] Single-modal data cannot capture multi-dimensional feedback such as users' expressions, voices, and barrage interactions;

[0004] The static label system cannot dynamically reflect the migration of users' interests, resulting in insufficient recommendation accuracy;

[0005] Multi-modal data has strong heterogeneity, and there is a lack of efficient cross-modal feature alignment and fusion methods. Summary of the Invention

[0006] In view of this, embodiments of the present invention provide a system and method for analyzing the user portrait of micro-short dramas based on multi-modal data to solve at least one of the above technical problems.

[0007] To achieve the above object, in a first aspect, a system for analyzing the user portrait of micro-short dramas based on multi-modal data is provided, including:

[0008] A data acquisition module for acquiring multi-modal micro-short drama user data;

[0009] A feature extraction module for extracting features according to the multi-modal micro-short drama user data to obtain a number of user features;

[0010] A fusion module for inputting a number of the user features into a pre-trained hierarchical attention network model to obtain a joint feature vector;

[0011] A dynamic portrait generation module for inputting the joint feature vector into a preset online deep forest model to generate a dynamic portrait;

[0012] A portrait analysis module for calculating the Shapley value of each dimension of the dynamic portrait and generating a portrait analysis report according to the Shapley value.

[0013] In a second aspect, the present invention provides a method for analyzing the user portrait of micro-short dramas based on multi-modal data, including:

[0014] Acquiring multi-modal micro-short drama user data;

[0015] Feature extraction is performed on the multi-modal micro short drama user data to obtain a number of user features;

[0016] A number of the user features are input into a pre-trained hierarchical attention network model to obtain a joint feature vector;

[0017] The joint feature vector is input into a preset online deep forest model to generate a dynamic portrait;

[0018] The Shapley value of each dimension of the dynamic portrait is calculated, and a portrait analysis report is generated according to the Shapley value.

[0019] In a third aspect, an electronic device is provided, which includes:

[0020] One or more processors;

[0021] A storage device for storing one or more programs,

[0022] When the one or more programs are executed by the one or more processors, the one or more processors implement a method for analyzing a micro short drama user portrait based on multi-modal data as described in the second aspect.

[0023] In a fourth aspect, a computer-readable storage medium is provided, on which a computer program is stored, and when the program is executed by a processor, a method for analyzing a micro short drama user portrait based on multi-modal data as described in the second aspect is implemented.

[0024] The above technical solutions have the following beneficial technical effects:

[0025] The present invention generates a user portrait by obtaining multi-modal micro short drama user data, which has higher accuracy and richer dimensions compared with the single-modal data portraits commonly used in the prior art. The present invention uses a hierarchical attention network model to fuse a number of user features to obtain a joint feature vector, which greatly improves the fusion accuracy and lays a solid foundation for generating an accurate portrait subsequently. The present invention uses a deep forest model to generate a dynamic portrait, which is more flexible than the existing user portraits and can be modified in a timely manner according to user preferences. Description of the Drawings

[0026] The drawings are used to better understand the present invention and do not constitute an improper limitation to the present invention. Among them:

[0027] Figure 1 is a structural block diagram of a system for analyzing a micro short drama user portrait based on multi-modal data in an embodiment of the present invention;

[0028] Figure 2 is a structural block diagram of a data acquisition module in an embodiment of the present invention;

[0029] Figure 3It is the structural block diagram of the feature extraction module in the embodiment of the present invention;

[0030] Figure 4 It is the structural block diagram of the fusion module in the embodiment of the present invention;

[0031] Figure 5 It is the structural block diagram of the dynamic portrait generation module in the embodiment of the present invention;

[0032] Figure 6 It is the structural block diagram of the portrait analysis module in the embodiment of the present invention;

[0033] Figure 7 It is the flowchart of a method for analyzing the user portrait of micro short plays based on multi-modal data in the embodiment of the present invention;

[0034] Figure 8 It is the structural schematic diagram of the computer system in the embodiment of the present invention. Detailed implementation manners

[0035] The following describes the exemplary embodiments of the present invention with reference to the accompanying drawings. Various details of the embodiments of the present invention are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present invention. Similarly, for the sake of clarity and conciseness, the description below omits the description of well-known functions and structures.

[0036] Embodiment 1

[0037] As Figure 1 shown, this embodiment provides a system for analyzing the user portrait of micro short plays based on multi-modal data, including:

[0038] A data acquisition module for acquiring multi-modal micro short play user data;

[0039] A feature extraction module for extracting features according to the multi-modal micro short play user data to obtain a number of user features;

[0040] A fusion module for inputting a number of the user features into a pre-trained hierarchical attention network model to obtain a joint feature vector;

[0041] A dynamic portrait generation module for inputting the joint feature vector into a preset online deep forest model to generate a dynamic portrait;

[0042] A portrait analysis module for calculating the Shapley values (Shapley Additive exPlanations, SHAP) of each dimension of the dynamic portrait and generating a portrait analysis report according to the Shapley values.

[0043] Specifically, as Figure 2 shown, the data acquisition module includes: a data acquisition unit, a time alignment unit, and a first preprocessing unit.

[0044] The data acquisition unit is used to acquire initial multimodal data from different data sources;

[0045] The time alignment unit is used to synchronize the timestamps of the initial multimodal data from different sources according to the Network Time Protocol to obtain time-synchronized initial multimodal data;

[0046] The first preprocessing unit is used to process the time-synchronized initial multimodal data by using the cubic spline interpolation method and the noise suppression algorithm to obtain multimodal data.

[0047] Specifically, in the data acquisition unit, the initial multimodal data includes initial behavior data, initial user facial data, initial user voice data, and initial text data. The initial behavior data includes click frequency (single / double click mode), sliding trajectory (direction / speed), pause / rewind operation, and device gyroscope attitude data. The acquisition frequency of the initial behavior data can be 10Hz. The initial user facial data captures the user's facial video stream through the front camera of the mobile phone at a frame rate of 30fps, and uses the MediaPipe framework to construct a three-dimensional facial mesh of 62 key points in real time, and focuses on extracting micro-expression parameters such as eyelid opening and closing degree, glabella muscle displacement, and mouth corner upward angle. The initial user voice data is obtained through a microphone. The initial user voice data is the user's real-time voice comment. After using the noise suppression algorithm to eliminate environmental noise, it is saved as a Pulse Code Modulation (PCM) format audio segment at a sampling rate of 16kHz. The initial text data is obtained in real time through the web socket (websocket) protocol. The initial text data includes the barrage content sent by the user and the sending time point.

[0048] Specifically, the 62 key points are not fixed. The number of key points can be adjusted according to the specific application scenario and algorithm model. The face mesh model of MediaPipe provides different numbers of key points with different precisions, ranging from dozens to hundreds. Selecting the number of key points requires balancing computational complexity, detection accuracy, and real-time performance. More key points can provide more detailed facial features, but also increase the computational burden.

[0049] Specifically, the MediaPipe framework is used to build complex computer vision and machine learning pipelines. This framework provides a rich set of pre-trained models and tools, supporting various computer vision tasks such as face detection, pose estimation, and gesture recognition. Its greatest feature is the ability to quickly deploy machine learning models and achieve real-time performance on mobile devices, web pages, and desktop platforms, enabling efficient multimedia processing through a graph computing model. The MediaPipe framework can be replaced by the deep learning face detection module of the open-source computer vision library (OpenCV), the facial landmark detection library of Dlib, the face recognition software development kit of Face++, the Vision framework of Apple Inc., and machine learning toolkits. These alternative solutions have their own characteristics in terms of real-time performance, key point accuracy, and cross-platform support. OpenCV provides high-performance deep learning models, Dlib is famous for its precise landmark localization, Face++ has advantages in commercial scenarios, Apple's Vision framework has high integration in the iOS ecosystem, and the machine learning toolkit provides a lightweight cross-platform solution similar to MediaPipe. For the multi-modal data acquisition scenario of the present invention, these alternative solutions all support real-time face mesh construction, can extract key points of eyes, mouth, and facial micro-expressions, and have low computational overhead and good mobile device adaptability, capable of meeting the technical requirements of real-time interactive data acquisition.

[0050] Specifically, pulse code modulation converts continuous analog signals into discrete digital signals by performing periodic sampling, quantization, and encoding on the analog signals. There are various alternative ways of pulse code modulation, such as Advanced Audio Coding, MP3, or Opus, etc. These formats provide different balances between compression ratio and sound quality. Compared with pulse code modulation, these alternative formats have higher compression efficiency and smaller storage space occupancy.

[0051] Specifically, in the time alignment unit, the Network Time Protocol (NTP) is a network protocol used to synchronize the time of computer systems. During the data acquisition process, millisecond-level timestamp synchronization is performed through the Network Time Protocol, and a sliding window algorithm is used to slice the data with a 5-second time unit.

[0052] Specifically, in the first preprocessing unit, when there are missing values in the initially time-synchronized multi-modal data, cubic spline interpolation is used to complete the time series. By constructing piecewise cubic polynomials to estimate the missing data points, the multi-modal data is ensured to be aligned in the spatio-temporal dimension. The noise suppression algorithm is a method used in signal processing to reduce background noise, including frequency domain filtering or spectral subtraction, etc.

[0053] Specifically, as Figure 3As shown, in the feature extraction module, it specifically includes:

[0054] The first extraction unit is used to extract features from the behavior data by using a temporal convolutional network to obtain a number of first user features;

[0055] The second extraction unit is used to extract features from the user facial data by using a pre-trained object detection model according to the visual database to obtain a number of second user features;

[0056] The third extraction unit is used to input the user speech data into a pre-trained speech representation model for feature extraction to obtain a number of third user features;

[0057] The fourth extraction unit is used to input the text data into a pre-trained language model for feature extraction to obtain fourth user features.

[0058] Specifically, in the first extraction unit, by inputting the behavior data into a preset temporal convolutional network for feature extraction, a number of the first user features are obtained. The first user features include viewing concentration features (4 dimensions), content preference intensity features (4 dimensions), and interaction behavior features (4 dimensions). The first user features are obtained through weighted calculation by the viewing concentration features, the content preference intensity features, and the interaction behavior features. The viewing concentration features include operation interval variance, single viewing continuous duration, interrupted viewing frequency, and repeated playback segment ratio. These indicators together reflect the user's concentration and interest intensity in the content. The content preference intensity features focus on the number of repeated plays, viewing completeness, content stay duration, and fast forward / rewind operation frequency of the user for specific types of content, thus revealing the user's content preference pattern. The interaction behavior features include click / slide frequency, device gyroscope attitude change rate, night / day viewing ratio, and social sharing willingness index, comprehensively expressing the user's viewing behavior characteristics. The temporal convolutional network is a deep learning model specifically used to process sequence data, directly processing time series information through convolutional operations. Compared with traditional recurrent neural networks, the temporal convolutional network has advantages such as parallel computing, long-term dependence modeling, and better gradient flow. The temporal convolutional network extracts features from time series data by using one-dimensional convolutional kernels, and can capture dependencies at different time scales through the dilated convolution mechanism while maintaining a low computational complexity. In the fields of audio processing, video analysis, user behavior prediction, etc., the temporal convolutional network can effectively capture time patterns and long-term dependencies in sequence data.

[0059] Specifically, the viewing concentration feature is determined based on the variance of operation intervals, the continuous viewing duration per session, the frequency of viewing interruptions, and the proportion of repeated play segments. The variance of operation intervals is calculated by recording the time intervals between every two consecutive operations (such as pausing, fast-forwarding, and rewinding) during the user's viewing process, and then calculating the standard deviation of these intervals. The calculation formula for the variance of operation intervals is as follows:

[0060]

[0061] In the formula, σ represents the variance of operation intervals, x i is each operation interval, μ is the average interval, and n is the number of operations. The smaller σ is, the more regular the user's operations are and the higher the concentration.

[0062] Specifically, the continuous viewing duration per session is obtained by counting the user's continuous viewing time without interruption and calculating the duration of the continuous viewing interval. The continuous viewing duration per session is defined as max(t_end - t_start), where t_end and t_start are the end timestamp and start timestamp of the continuous viewing interval respectively. The longer the duration, the higher the concentration.

[0063] Specifically, the frequency of viewing interruptions represents the number of interruption behaviors such as pausing, switching, and exiting during the user's viewing process. The calculation method for the frequency of viewing interruptions is count(interruption_events), that is, counting the total number of interruption events. The fewer the interruption times, the higher the concentration.

[0064] Specifically, the proportion of repeated play segments is determined by calculating the proportion of the user's repeated play of specific video segments. The calculation formula for the repeated play segments is [(duration of repeated play segments / total viewing duration) * 100%]. By analyzing whether the user repeatedly watches certain segments, it reflects the degree of the user's interest in the content.

[0065] Specifically, the content preference intensity feature is obtained by weighted calculation based on the number of repeated plays, the viewing completeness, the content staying duration, and the fast-forward / rewind operation frequency. The number of repeated plays is obtained by counting the total number of repeated plays of specific types or specific content by the user. The calculation method is sum(replay_count), that is, accumulating the number of replays of a certain type of content. The more the number of replays, the stronger the user's preference for this content.

[0066] Specifically, the viewing completeness is used to measure whether the user watches the content completely. The calculation formula for the viewing completeness is [(actual viewing duration / total video duration) * 100%]. The higher the viewing completeness, the stronger the user's interest in the content.

[0067] Specifically, the content stay duration is used to record the total time that a user stays on a specific content type or a specific video. The calculation method of the content stay duration is sum(stay_duration), that is, to accumulate the stay time of the user on a certain type of content. The longer the stay duration, the more interested the user is.

[0068] Specifically, the fast forward / rewind operation frequency is used to count the number of fast forward or rewind operations of the user during the viewing process. The calculation formula of the fast forward / rewind operation frequency is [(number of fast forward operations + number of rewind operations) / total viewing duration]. The lower the operation frequency, the more focused the user is on the content.

[0069] Specifically, the interactive behavior characteristics are calculated by weighting the click / slide frequency, the device gyroscope attitude change rate, the night / day viewing ratio, and the social sharing willingness index. The click / slide frequency is used to record the frequency of click or slide operations of the user during the viewing process. The calculation formula of the click / slide frequency is [(number of click operations + number of slide operations) / total viewing duration], and the click / slide frequency reflects the interactive activity degree of the user.

[0070] Specifically, the device gyroscope attitude change rate analyzes the device gyroscope data to calculate the frequency and amplitude of the device attitude change during the viewing process. The calculation formula of the device gyroscope attitude change rate is as follows:

[0071]

[0072] In the formula, Δθ i is the attitude angle change value of the i-th time, n is the total number of attitude changes, T is the total viewing duration (in seconds or minutes), and the device gyroscope attitude change rate reflects the concentration and comfort of the user during viewing.

[0073] Specifically, the night / day viewing ratio: counts the proportion of the viewing duration of the user in different time periods (night / day). The calculation formula is (night viewing duration / total viewing duration)*100%, indicating the viewing habits and time preferences of the user.

[0074] Specifically, the social sharing willingness index analyzes the sharing behavior of the user to calculate its recognition degree and dissemination willingness for the content. The calculation formula of the social sharing willingness index: [(number of sharing operations / number of viewing operations)*100%], reflecting the social value evaluation of the user for the content.

[0075] Specifically, in the second extraction unit, the visual database includes a simplified computer vision library or an open-source library focusing on human pose recognition, etc. The optical flow matrix between consecutive frames is calculated through the visual database, and then the optical flow matrix between consecutive frames is input into the pre-trained object detection model (such as Residual Neural Network-18 model, lightweight convolutional neural network model, Transformer vision model or other pre-trained object detection models) to extract the second user features such as the pupil diameter change rate, blink frequency in the eye region (ROI_eye), and the muscle tremor amplitude in the mouth region (ROI_mouth). The second user features include eye region features (4D), mouth region features (3D), and comprehensive facial features (2D).

[0076] Specifically, the eye region features are obtained by weighted calculation based on the pupil diameter change rate, blink frequency, eye muscle tension, and fixation duration change. The mouth region features are obtained by weighted calculation based on the angle of upward curvature of the mouth corners, muscle tremor amplitude, and mouth micro-action frequency. The comprehensive facial features are obtained by weighted calculation through the facial muscle tension index and emotional expression intensity. The comprehensive facial features provide an overall assessment of the user's emotional state and provide important non-verbal interaction information for the multi-modal user portrait.

[0077] Specifically, the pupil diameter change rate extracts the change in the pupil region in consecutive frames by using edge detection and contour recognition algorithms in the visual database. By tracking the tiny changes in the pupil size, it reflects the user's emotional and attention fluctuations. The calculation formula for the pupil diameter change rate is as follows:

[0078]

[0079] Specifically, the blink frequency identifies blink events in consecutive video frames using an object detection model. The calculation formula for the blink frequency is: Blink frequency = Total number of blinks / Observation time length. The unit of the observation time length is preferably seconds. The higher the blink frequency, the more nervous, fatigued, or distracted.

[0080] Specifically, the eye muscle tension is calculated by analyzing the tiny deformation of the periorbital muscles and using the optical flow method to calculate the standard deviation of the pixel point displacement in the eye region. The calculation formula for the eye muscle tension is as follows:

[0081]

[0082] In the formula, MT represents the eye muscle tension, Δp i is the displacement difference of the i-th pixel point, and N is the total number of pixel points. The larger the value of the eye muscle tension, the higher the muscle tension, reflecting stress or emotional fluctuations.

[0083] Specifically, the change in fixation duration is obtained by tracking the stable fixation time of the eyeball in consecutive frames and calculating the variance of the fixation duration. The calculation formula for the change in fixation duration is as follows:

[0084]

[0085] In the formula, T i is the fixation duration of the i-th frame, T μ is the average fixation duration, and n is the total number of frames. The change in fixation duration can reflect the changes in cognitive load and concentration.

[0086] Specifically, the angle of upward curvature of the mouth can extract the coordinates of the corners of the mouth using a key-point detection algorithm and calculate the change in the angle between the corners of the mouth and the horizontal line. The calculation formula for the angle of upward curvature of the mouth is as follows:

[0087]

[0088] In the formula, θ up is the angle of upward curvature of the mouth, (x1, y1) and (x2, y2) are the coordinates of the two key points of the corners of the mouth, and arctan is the arctangent function used to calculate the angle. The angle of upward curvature of the mouth can reflect the emotional state, such as smiling or being nervous.

[0089] Specifically, the amplitude of muscle tremors analyzes the displacement intensity of pixels in the mouth region in consecutive frames using the optical flow method. The calculation formula for the amplitude of muscle tremors: Amplitude of muscle tremors = max(length of pixel displacement vector). The larger the amplitude of muscle tremors, the more intense the muscle activity, indicating emotional fluctuations.

[0090] Specifically, the frequency of micro-movements of the mouth is obtained by identifying the minute movements between consecutive frames in the mouth region and counting the number of movements per unit time. The calculation method for the frequency of micro-movements of the mouth: Frequency of micro-movements of the mouth = total number of micro-movements / observation time length. The unit of the observation time length is preferably seconds, and the frequency of micro-movements of the mouth can reflect the degree of tension and emotional expression.

[0091] Specifically, the facial muscle tension index calculates the comprehensive standard deviation of the displacement of the entire facial muscles by integrating the muscle activity data in the eye and mouth regions. The calculation formula for the facial muscle tension index is as follows:

[0092]

[0093] In the formula, MTI is the facial muscle tension index, Δq i is the displacement difference of the i-th facial pixel point, and M is the total number of facial pixel points. The larger the value of the facial muscle tension index, the higher the overall muscle tension.

[0094] Specifically, the emotional expression intensity is obtained by constructing a comprehensive emotional intensity scoring model by combining

[0095] the eye region features and the mouth region features. A machine learning model (such as a support vector machine) can be used for training to output an emotional intensity value between 0 and 1. The higher the emotional intensity value, the more obvious the emotional expression. First, the eye region features and the mouth region features are preprocessed to unify the dimensions. Specifically, the min-max normalization method can be adopted to map each feature to the [0,1] interval. The formula is X_normalized = (X - X_min) / (X_max - X_min). Through this standardization, the scale differences between features can be eliminated, enabling the subsequent machine learning model to more fairly evaluate the importance of each feature and improving the learning efficiency and generalization ability of the model. Secondly, a support vector machine is selected as the algorithm for the emotional intensity scoring model, which can effectively process multi-dimensional non-linear features. The emotional intensity scoring model learns from the labeled training data set to construct an optimal classification hyperplane. Then, the emotional intensity output by the emotional intensity scoring model is within [0,1], where 0 indicates extremely weak emotional expression, close to a state of no perception; 1 indicates extremely strong emotional expression, almost out of control. The intermediate values reflect the gradual change of emotional intensity. This quantification method provides an objective measure of emotional intensity.

[0096] Specifically, in the third extraction unit, the user voice data is first segmented into effective voice segments, and then the effective voice segments are converted into 128-dimensional Mel spectrograms and input into a pre-trained voice representation model (such as the VGG-based Feature Extraction for Audio (VGGish)) to obtain voice emotion features (6-dimensional). The voice emotion features are the third user features. The voice emotion features extract the multi-dimensional emotional states of the user through in-depth analysis of the audio signal. The excitement level and disappointment intensity reflect the user's immediate emotional reaction to the content, and the surprise index captures the emotional fluctuations brought by unexpected or breakthrough content. The anger level and fear degree reveal the user's negative emotional feedback to the content, while the neutral emotion ratio provides an index of the balance of emotional expression. These six-dimensional features together constitute a comprehensive voice emotion portrait, which can deeply analyze the subtle changes in the user's emotions.

[0097] Specifically, the VGGish network can be replaced by the wav2vec 2.0 model or the Hidden-Unit Bidirectional Encoder Representations from Transformers (HuBERT). The wav2vec 2.0 model has excellent performance in speech feature extraction. By performing self-supervised pre-training on the original audio signal, it can capture richer and deeper speech features and maintain high accuracy even with a small amount of labeled data. Compared with VGGish, wav2vec 2.0 has more advantages in feature expression and generalization ability. The HuBERT model can learn powerful speech features from unlabeled audio data through a clustering and BERT-like masked prediction strategy. HuBERT performs well in speech emotion recognition tasks, and its model structure is more flexible, which can well adapt to different speech emotion feature extraction requirements.

[0098] Specifically, the speech emotion features are obtained by weighting the excitement degree feature, disappointment intensity feature, surprise index feature, anger level feature, fear degree feature, and neutral emotion ratio feature. The calculation formula for the excitement degree feature is as follows:

[0099] Excitement degree feature = α * speech energy + β * fundamental frequency change rate + γ * speech rate;

[0100] In the formula, α, β, and γ are weight coefficients, which are dynamically adjusted by machine learning methods such as gradient descent. Speech energy reflects the sound intensity, the fundamental frequency change rate reflects the pitch fluctuation, and the speech rate characterizes the speaking rhythm. The higher the excitement degree, the larger the values of these three indicators. The three indicators of the excitement degree feature are obtained through complex acoustic signal processing techniques. Speech energy is obtained by performing a short-time Fourier transform (STFT) on the audio signal, calculating the sum of the squares of the signal amplitudes in each time window and taking the logarithm, which reflects the physical intensity and amplitude change of the sound. The fundamental frequency change rate uses a pitch detection algorithm (such as the autocorrelation method or frequency domain analysis) to extract the fundamental frequency (F0) of each speech frame, and quantifies the dynamic fluctuation degree of the pitch by calculating the change rate and amplitude of the fundamental frequency between adjacent frames. Speech rate measurement involves speech segmentation and phoneme recognition. First, an endpoint detection algorithm (such as the energy threshold method or zero crossing rate) is used to segment continuous speech into independent speech segments, and then the number of speech segments or syllables per unit time is counted and standardized in combination with the speech duration. These three indicators dynamically adjust the weight coefficients through machine learning methods (such as gradient descent) to construct a multi-dimensional and dynamically adaptable excitement degree evaluation model, which can accurately capture the subtle emotional change features in speech.

[0101] Specifically, the calculation formula for the disappointment intensity feature is as follows:

[0102]

[0103] In the formula, DI represents the disappointment intensity feature, σ represents the Pitch Variance. D represents the Amplitude Decay Coefficient. e is the base of the natural logarithm. δ and ε are adjustment parameters, and the Sigmoid function is used to map the feature to the [0,1] interval. The smaller the pitch variance of the voice and the faster the amplitude decay, the higher the disappointment intensity. The acquisition of the pitch variance and amplitude decay coefficient of the voice involves complex acoustic signal processing techniques. The pitch variance is obtained through the statistical analysis of the fundamental frequency (F0). First, the fundamental frequency extraction algorithm (such as the autocorrelation method or the Hilbert-Huang transform) is used to accurately locate the fundamental frequency in each speech frame, and then the statistical variance of these fundamental frequencies is calculated. Specifically, the original speech signal needs to be preprocessed, including denoising, normalization, and framing. Each frame is 20-40 milliseconds, with 50% overlap to ensure continuity. The fundamental frequency extraction algorithm is applied to each frame to obtain the fundamental frequency sequence, and its standard deviation is calculated. The smaller the variance, the more single the pitch change, reflecting the monotony of the speech and the lack of emotional fluctuations. The amplitude decay coefficient of the voice is obtained by analyzing the exponential decay characteristics of the speech signal envelope. First, the Hilbert envelope transform is performed on the speech signal to extract the instantaneous amplitude of the signal, and then the exponential decay model is used to fit the change trend of the amplitude with time. The specific method is to take the logarithm of the amplitude and then perform linear regression. The slope is the amplitude decay coefficient, which reflects the degree of rapid decline of the speech intensity with time. Combining these two indicators with the Sigmoid non-linear mapping constructs a disappointment intensity evaluation model that can sensitively capture subtle emotional changes in speech.

[0104] Specifically, the calculation formula for the surprise index feature is: Surprise Index = max(|Fundamental Frequency Mutation Rate|, |Speech Energy Mutation Rate|) * Mutation Duration. By capturing the sudden changes in the speech signal, it quantifies the intensity of unexpected or unanticipated emotions. The greater the mutation rate and the longer the duration, the higher the surprise index. Obtaining the fundamental frequency mutation rate and the speech energy mutation rate requires the use of advanced signal processing and feature extraction techniques. The calculation of the fundamental frequency mutation rate first involves the precise extraction of the fundamental frequency (F0). A multi-algorithm fusion method is adopted, such as the autocorrelation function, cepstrum analysis, and harmonic matching method, to ensure the accurate localization of the fundamental tone in a complex speech background. The fundamental frequency of consecutive speech frames (each frame is 20 - 30 milliseconds) is detected, and the instantaneous change rate of the fundamental frequency between adjacent frames is calculated. Specifically, it is obtained by dividing the absolute value of the difference in the fundamental frequency between two consecutive frames by the inter-frame time interval. The speech energy mutation rate is based on the short-time Fourier transform (STFT) and energy envelope analysis. First, the speech signal is decomposed into a time-frequency representation, and the signal energy of each time window is calculated. The relative change rate of the energy between adjacent time windows is used to quantify it. The mutation detection algorithm combines wavelet transform and adaptive threshold methods to identify significant energy jump points and exclude the interference of background noise. The calculation of these two mutation rates not only focuses on instantaneous changes but also considers the duration and amplitude of the changes. Through multi-scale analysis, it captures the subtle but obvious unexpected features in the speech signal, thereby constructing a computational model that can sensitively reflect the intensity of surprise emotions.

[0105] Specifically, the calculation formula for the anger level feature is as follows:

[0106] Anger Level = ζ * Speech Roughness + η * Speech Rate Volatility + θ * Fundamental Frequency Instability;

[0107] In the formula, ζ, η, and θ are adjustment weights. The voice roughness reflects the degree of hoarseness, the speech rate volatility reflects the emotional instability, and the fundamental frequency instability characterizes the drastic change in pitch. The acquisition of voice roughness, speech rate volatility, and fundamental frequency instability requires the comprehensive application of various acoustic signal processing techniques. The voice roughness is quantified by the irregularity of glottal closure and the harmonic-to-noise ratio (HNR). First, the linear predictive cepstral coefficients (LPCC) and mel-frequency cepstral coefficients (MFCC) are used to extract the vocal tract features, and then combined with the glottal pulse analysis algorithm to accurately evaluate the proportion of the aperiodic components in the speech signal. Specifically, by calculating the modulation spectrum and harmonic structure of each speech frame, the degree of harmonic distortion of the sound is quantified. The higher the roughness, the more hoarse and harsh the sound is. The speech rate volatility is based on the statistical analysis of the time intervals of speech segmentation and the syllable durations. The dynamic time warping (DTW) algorithm is used to accurately extract the time features of the speech segment, and the coefficient of variation of adjacent syllable durations is calculated to reflect the irregularity of the speech rate. The fundamental frequency instability is obtained through multi-dimensional pitch trajectory analysis. First, the improved YIN algorithm (the YIN algorithm is an algorithm for pitch frequency detection, which can accurately extract the pitch frequency or pitch from the audio signal) and the harmonic matching technique are used to accurately extract the fundamental frequency sequence, and then the short-term variance and peak difference of the fundamental frequency are calculated to quantify the degree of drastic fluctuation of the pitch. These three features are combined through weighted linear combination to construct a computational model that can sensitively capture the subtle changes in the angry emotion in speech, not only paying attention to the instantaneous features, but also considering the overall dynamic characteristics of the speech signal.

[0108] Specifically, the calculation formula for the fear degree feature is as follows:

[0109]

[0110] In the formula, F represents the voice tremor frequency. A represents the voice volume weakening coefficient. e is the base of the natural logarithm. κ and λ are adjustment parameters, and are normalized using the Sigmoid function. The higher the voice tremor frequency and the weaker the voice volume, the greater the degree of fear. Obtaining the voice tremor frequency and the voice volume weakening coefficient requires the use of complex acoustic signal processing and psycholinguistic analysis techniques. The calculation of the voice tremor frequency is first based on time-frequency analysis methods such as the Short-Time Fourier Transform (STFT) and the Hilbert-Huang Transform (HHT), which decompose the voice signal into multi-scale oscillation components. By extracting the envelope and instantaneous frequency of the voice signal, the involuntary tremor characteristics in the voice are identified. The specific methods include: calculating the change rate of the envelope of the voice signal, extracting the tremor frequency band (in the range of 4 - 8 Hz), and analyzing the frequency distribution and duration of the tremor. At the same time, combining wavelet transform and Empirical Mode Decomposition (EMD) techniques, the weak tremor components in the voice signal are accurately extracted, excluding the interference of background noise. The voice volume weakening coefficient is comprehensively evaluated through multi-dimensional acoustic features. First, logarithmic energy normalization and perceived sound level calculation methods are used to extract the short-time energy and sound pressure level of the voice signal, and then combined with the psychoacoustic model to quantify the relative weakness of the voice volume. The specific calculations include: calculating the normalized energy envelope of the voice segment and analyzing its deviation from the standard voice energy distribution; using the perceived weight function to consider the sensitivity of the human ear to sounds of different frequencies; introducing a dynamic range compression algorithm to highlight the characteristics of weak sounds. Through fine signal processing and psychoacoustic modeling, these two features can sensitively capture the subtle acoustic features in the voice that reflect fear emotions, providing a detailed calculation basis for the quantification of emotional intensity.

[0111] Specifically, the calculation formula for the neutral emotion proportion feature is as follows:

[0112] Neutral emotion proportion feature = 1 - (excitement degree feature + disappointment intensity feature + surprise index feature + anger level feature + fear degree feature);

[0113] The neutral emotion proportion feature is calculated by subtraction, reflecting the proportion of non-extreme emotional states. The closer the proportion is to 1, the more stable and neutral the emotional expression is.

[0114] Specifically, in the fourth extraction unit, the text data line is input into the Bidirectional Encoder Representations from Transformers-Whole Word Masking (BERT-wwm) model after being segmented by Jieba. A 768-dimensional semantic vector is obtained through the Classification Token (CLS), and then reduced to the bullet screen emotion feature (8 dimensions) by Principal Component Analysis (PCA). Each feature extractor is connected to an intra-modal attention layer at the backend, and learnable parameters are used to calculate the feature importance weights. For example, in facial features, the indication of the blink frequency on the interest attenuation is automatically strengthened. The fourth user feature is the bullet screen emotion feature.

[0115] Through in-depth analysis of the user group interaction, the bullet screen emotion feature extracts collective emotion features. The positive emotion intensity feature and the negative emotion intensity feature directly reflect the overall evaluation of the content by users. The network hot word usage density feature reflects the popularity of the content and the language features of the user group. The emotion fluctuation frequency feature and the interaction enthusiasm feature reveal the participation degree of the user group, while the emotion consistency feature and the group emotion synchronization feature reflect the resonance degree of the collective emotion. There is also an emotion depth feature. These eight-dimensional features not only capture individual emotions but also present the collective emotion features of the user group. In other embodiments, BERT-wwm can be replaced by the Robustly Optimized BERT Pretraining Approach (RoBERTa), the Knowledge Enhanced Representation Model ERNIE, or other pre-trained language models to extract and optimize the semantic features of text data.

[0116] Specifically, Jieba segmentation can be replaced by the HanLP Chinese natural language processing toolkit or the THULAC Chinese word segmentation tool. HanLP can provide a more refined word segmentation mode, support custom dictionaries, and has stronger ambiguity elimination capabilities. Especially in network texts and colloquial corpora, HanLP's word segmentation performance is more outstanding. HanLP has a high-precision word segmentation algorithm and low computational overhead. It supports multiple word segmentation modes and has a dictionary trained with a large-scale corpus, performing excellently in processing complex Chinese texts. Compared with Jieba, THULAC has more advantages in word segmentation for professional fields and academic texts.

[0117] Specifically, the calculation formula for the positive emotion intensity feature is as follows:

[0118]

[0119] In the formula, PS is the positive emotion intensity, Wi + is the weight of the i-th positive sentiment word. These weights are determined by a pre-trained sentiment dictionary and represent the contribution degree of the sentiment word to the emotion. F i is the word frequency of the i-th sentiment word, that is, the number of times the sentiment word appears in the text. N is the total number of words, representing the number of all words in the text, which is used to normalize the sentiment intensity.

[0120] Specifically, the calculation formula of the negative sentiment intensity feature is as follows:

[0121]

[0122] In the formula, NS is the negative sentiment intensity feature, W i - is the weight of the i-th negative sentiment word, which is also determined by the sentiment dictionary and represents the influence of the negative sentiment word on the emotion.

[0123] Specifically, the calculation formula of the network buzzword usage density feature is as follows:

[0124] Network buzzword density feature = (Number of times network buzzword appears / Total number of words) * log(Spread range of network buzzword);

[0125] By dividing the number of times the network buzzword appears by the total number of words, the appearance frequency of the network buzzword is obtained. Then, combined with the spread range of the network buzzword, it reflects the popularity of the content and the language characteristics of the user group. Using the logarithmic function weakens the influence of extreme values and makes the feature more representative. The acquisition of the spread range of the network buzzword is achieved by comprehensively analyzing multi-dimensional indicators of online social media, mainly including the number of forwards, comments, interaction times, and cross-platform mentions. The specific method is to use web crawlers, big data analysis platforms, and social network application programming

[0126] interfaces to count the spread breadth of the network buzzword in different platforms, regions, and user groups. The calculation of the spread range of the network buzzword integrates factors such as the number of forwards, the number of platforms, user group diversity, and time span, and can be expressed as a continuous floating-point number or a standardized discrete value to accurately reflect the spread intensity and influence of the network buzzword. This quantification method can not only objectively measure the popularity of the buzzword but also reveal the language characteristics of the user group and the information dissemination dynamics.

[0127] Specifically, the calculation formula of the emotion fluctuation frequency feature is as follows:

[0128] Emotion fluctuation frequency feature = Number of consecutive emotional polarity changes / Total number of bullet comments;

[0129] The emotional fluctuation frequency feature quantifies the instability of the user group's emotions by tracking the rapid switching of the emotional polarity (positive / negative / neutral) between consecutive bullet comments. The more frequent the fluctuations, the higher the feature value.

[0130] Specifically, the calculation formula for the interactive enthusiasm feature is as follows:

[0131] Interactive enthusiasm feature = (number of replies * weight of interactive words) / total number of bullet comments;

[0132] The interactive enthusiasm feature comprehensively considers the number of bullet comment replies and the usage frequency of interactive words (such as "666", "amazing", etc.). By quantifying the degree of users' active participation and positive interaction, it reflects the enthusiasm of the group's participation.

[0133] Specifically, the calculation formula for the emotional consistency feature is as follows:

[0134] Emotional consistency feature = 1 - (standard deviation of emotional polarity / mean of emotional polarity);

[0135] The emotional consistency feature calculates the degree of dispersion of the emotional polarity of the bullet comment group. The smaller the standard deviation, the more consistent the emotions of the user group. Subtracting the normalized coefficient of dispersion from 1 makes the larger the feature value represent higher consistency.

[0136] Specifically, the calculation formula for the group emotional synchronization feature is as follows:

[0137] Group emotional synchronization feature = Pearson correlation coefficient (emotional polarity sequence 1, emotional polarity sequence 2);

[0138] The group emotional synchronization feature quantifies the synchronization degree of the emotions of the user group by calculating the Pearson correlation coefficient of the emotional polarity sequences within different time windows. The closer the correlation coefficient is to 1, the more synchronized the group emotions are. The Pearson correlation coefficient is used to measure the linear correlation degree between two variables, with a value range between -1 and 1. When the coefficient is close to 1, it indicates a positive correlation, that is, the two variables change in the same direction; when close to -1, it indicates a negative correlation, that is, the variables change in the opposite direction; when close to 0, it means there is almost no linear correlation between the variables. In the analysis of group emotional synchronization, the emotional polarity sequence is obtained through natural language processing (NLP) technology. The specific method is to perform sentiment analysis on the bullet comment text, and use sentiment classification algorithms (such as Naive Bayes, Support Vector Machine, or deep learning models) to classify each bullet comment into positive, negative, or neutral emotions, and assign corresponding numerical values (such as positive +1, negative -1, neutral 0), forming an emotional polarity sequence arranged in chronological order. By calculating the Pearson correlation coefficient of these two sequences, the consistency of the emotional changes of different user groups within a specific time window can be quantified, thereby revealing the dynamic characteristics and synchronization mechanism of collective emotions.

[0139] Specifically, the calculation formula for the emotional depth feature is as follows:

[0140] Emotional depth feature = (number of emotional words / total number of words) * log(average weight of emotional words);

[0141] Specifically, the emotional depth feature reflects the depth and complexity of emotional expression in the bullet comments by combining the proportion of emotional words and the average weight of emotional words. The logarithmic function is used to balance extreme values and quantify the richness of the emotional expression of the user group. The whole-word masking BERT model is an important pre-trained model in the field of natural language processing, which improves the language understanding ability of the model through an improved masking strategy. The classification token is located at the beginning of the input sequence, and aggregates the semantic information of the entire sequence through the attention mechanism to generate a global semantic representation. During the text feature extraction process, the classification token can capture the overall semantic features of the input text, providing a semantic basis for subsequent emotional analysis and user portrait construction. During the pre-training and fine-tuning of the whole-word masking BERT model, the classification token is placed at the beginning of the input sequence, and its role is to aggregate the semantic information of the entire input sequence through the self-attention mechanism to generate a fixed-length global semantic representation vector. This token was originally designed for processing classification tasks, and can capture the overall semantic features of the input text, integrating the context information of each token in the sequence through the multi-head attention mechanism, so as to provide a semantically rich and compact representation for subsequent text classification, emotional analysis and other tasks. In the multi-modal user portrait system, the classification token can not only extract the global semantic features of the text, but also serve as a bridge for cross-modal feature fusion, helping the system understand the deep semantic intentions behind user behaviors.

[0142] Specifically, the first user feature, the second user feature, the third user feature, the fourth user feature, and the features subordinate to each user feature dynamically adjust the contribution degrees of different features through learning parameters. The weight calculation uses the softmax function, which non-linearly transforms the original features through a learnable weight matrix W and a bias term b to obtain the importance distribution of each feature. This adaptive weight learning mechanism enables the model to automatically adjust the weights of each modal feature according to different scenarios and user features, achieving a more accurate representation of the user portrait.

[0143] Specifically, as Figure 4 shown, in the fusion module, it specifically includes:

[0144] A second preprocessing unit for preprocessing a plurality of the user features to obtain a plurality of preprocessed features;

[0145] A intra-modal attention unit for dividing the plurality of preprocessed features into a plurality of modalities, calculating the attention scores of each of the preprocessed features within each modality, and selecting key sub-features from each modality according to the attention scores;

[0146] A cross-modal attention unit for calculating the correlation degree between pairwise modalities using a cross-modal attention algorithm;

[0147] A fusion unit for generating a joint feature vector based on the key sub-features and the correlation degree.

[0148] Specifically, in the second preprocessing unit, the preprocessing includes standardization and normalization. Since the user features include data of multiple modalities (such as text, speech, facial expressions, etc.), for each modality m, its feature vector By using batch normalization or layer normalization, the scale differences in the feature spaces of different modalities are eliminated, ensuring the fairness and stability of subsequent attention mechanism calculations. The normalization process helps to reduce the deviation in the feature distributions between modalities.

[0149] Specifically, in the intra-modal attention unit, when dividing several preprocessed features into several modalities, the division is performed according to the data sources of the preprocessed features. For example, the preprocessed features of the speech category are divided into the unified speech modality. For each modality m, through a learnable weight matrix W m and V m , the attention score α m is calculated. The calculation formula for the attention score is as follows:

[0150]

[0151] In the formula, a m is the attention score of modality m, representing the importance weights of each dimension of the features of this modality. W m is the weight matrix of modality m, used for calculating the weighting of the preprocessed features within modality m. V m is another learnable matrix of modality m, used for performing a non-linear transformation on the input preprocessed feature h m . h m is the preprocessed feature within modality m, representing the input features of the modality. Tanh is a non-linear activation function used to introduce non-linear characteristics. Softmax is used to normalize the obtained attention scores to ensure that the sum of the attention scores of all modalities is 1. is the transpose matrix of W m .

[0152] Specifically, according to the attention scores, the preprocessed feature with the largest weight is selected from each modality as the key sub-feature.

[0153] Specifically, in the cross-modal attention unit, the semantic correlation degree between different modalities is calculated through learnable query matrix Q (Query), key matrix K (Key), and value matrix V (Value). Specifically, for modality i and modality j, their correlation degree β ij is calculated through the dot product attention mechanism, and the calculation formula for the correlation degree is as follows:

[0154]

[0155] In the formula, β ij is the correlation degree between modality i and modality j. q i is the query vector (Query) of modality i. k j is the key vector (Key) of modality j. k k represents the key vector (Key Vector) of the k-th modality. d is the dimension of the query vector and the key vector, which is used to scale the dot product to prevent the numerical value from being too large. exp represents the exponential function, which is used to calculate the exponential value after the dot product. ∑k is the weighted sum of all modalities to ensure the normalization of the correlation degree value. K is the total number of modalities.

[0156] Specifically, in the fusion unit, according to the correlation degree between modalities and several key sub-features, the key sub-features of different modalities are integrated through weighted summation or more complex fusion strategies. The joint feature vector comprehensively reflects the semantic information of each modality and the interaction relationship between modalities, providing rich multi-modal representations for subsequent sentiment classification tasks.

[0157] Specifically, it further includes a dynamic feature enhancement unit, which is used to detect several of the key sub-features. When a special field appears in the several key sub-features, the weight of the key sub-feature with the special field when generating the joint feature vector is adjusted. The special field, for example, is a barrage containing "reverse wonderful". When there is a special field, the dynamic feature enhancement mechanism is activated, such as automatically adjusting the weight coefficient of the key sub-feature related to barrage feedback, so as to achieve semantic enhancement of the text-voice modality. This mechanism dynamically adjusts the weight allocation of cross-modal feature fusion based on the real-time feedback of user interaction, improving

[0158] the adaptability of the model to complex interaction scenarios.

[0159] Specifically, it further includes a post-processing unit, which is used to post-process the joint feature vector (128-dimensional). The post-processing includes but is not limited to: dropout regularization, layer normalization, or residual connection. These techniques are beneficial to reducing overfitting, stabilizing model training, and preparing high-quality feature representations for the subsequent sentiment classification layer.

[0160] Specifically, such asFigure 5 As shown in Figure 5 , in the dynamic portrait generation module, it specifically includes:

[0161] A subset establishment unit, configured to select several feature subsets with different granularities from the combined feature vectors;

[0162] An attenuation unit, configured to adjust the influence weights of several of the feature subsets according to a preset time attenuation factor;

[0163] A portrait generation unit, configured to input several of the feature subsets and the influence weights into the pre-trained online deep forest model to obtain an initial user portrait;

[0164] An update unit, configured to perform incremental learning at preset time intervals to update the user portrait and obtain a dynamic portrait.

[0165] Specifically, in the subset establishment unit, the random subspace method, feature importance evaluation method, or principal component analysis method is used for the combined feature vectors to select feature subsets with different granularities. This method can effectively reduce the dimension, prevent overfitting, and at the same time enhance the generalization ability of the model, enabling each base learner to capture the subtle changes in user interests from different perspectives.

[0166] Specifically, in the attenuation unit, the expression of the attenuation factor is: λ = 0.98 t , where t is the data time difference in hours. By setting the attenuation factor, the weights of historical samples are dynamically adjusted, enabling the online deep forest model (ODF-MT) to pay more attention to recent behavior patterns; it can exponentially decay the influence of historical data. For example, the weight of data 1 hour ago is approximately 0.98, the weight of data 24 hours ago drops to 0.376, and the weight of data 72 hours ago is only 0.014. Through this refined weight adjustment strategy, the model can quickly respond to the real-time changes in user interests, while retaining the reference value of historical behaviors, improving the sensitivity and adaptability to recent behavior patterns, and thus achieving dynamic and accurate updates of the user portrait.

[0167] Specifically, in the portrait generation unit, the online deep forest model includes 3 incremental random forest models and 2 dynamic gradient boosting tree models, and each base learner receives feature subsets with different granularities; a base learner is a single classification or regression model in ensemble learning. By inputting the feature subsets and the corresponding influence weights of each feature subset into different models, an initial user portrait is obtained.

[0168] Specifically, in the update unit, the preset time interval is preferably 30 minutes. By adopting a misaligned binning strategy to update the leaf node distribution (please explain the misaligned binning strategy), when the KL divergence of the new data distribution exceeds the threshold θ = 0.15, the node is automatically split. At the same time, the system maintains the user status matrix U ∈ R 128 , and fuses long-term and short-term interest features through the LSTM update gate mechanism. A leaf node is the bottommost node in a decision tree / forest model and will not be split anymore, representing the final classification or regression result. The misaligned binning strategy is an advanced data distribution adjustment method. By dividing the data space into multiple overlapping sub-intervals, when the new data distribution drifts significantly, it automatically triggers the splitting of leaf nodes. This strategy enables the model to dynamically adjust its internal structure and capture the subtle changes and potential trends of user interests in real time. The user status matrix U ∈ R 128 is obtained according to the joint feature vector and is continuously updated and optimized by combining the user's historical interaction data. The system uses the LSTM (Long Short-Term Memory) update gate mechanism to fuse the long-term and short-term interest features of the user. The update gate controls the flow of information into the long-term memory unit, balancing the volatility of short-term interests and the stability of long-term interests. The LSTM update gate mechanism realizes the dynamic fusion of long-term and short-term interest features through a complex gating structure.

[0169] Specifically, as Figure 6 shown, in the portrait analysis module, it specifically includes:

[0170] A Shapley value calculation unit for calculating the Shapley values of each dimension in the dynamic portrait;

[0171] A label generation unit for generating a number of multi-dimensional labels according to each dimension in the dynamic portrait and the corresponding Shapley values of each dimension;

[0172] An analysis unit for generating a portrait analysis report according to a number of the multi-dimensional labels.

[0173] Specifically, in the Shapley value calculation unit, when extracting features from the dynamic portrait, each dimension of the dynamic portrait corresponds to a portrait feature, obtaining a number of portrait features, and inputting the number of portrait features into the Shapley library to calculate the Shapley values of each portrait feature.

[0174] Specifically, the working process of the Shapley value calculation unit includes the following steps:

[0175] Step 1: Feature extraction. First, in the Shapley value calculation unit, each dimension of the dynamic portrait corresponds to a portrait feature. These features are extracted based on the user's behavior data or sentiment data. For example, instantaneous sentiment value, sentiment fluctuation index, content preference degree, consumption decision, etc. Each dimension represents the user's behavior or sentiment state in a specific field. The process of feature extraction includes identifying and extracting specific data related to each dimension from the original user data and converting it into a feature vector convenient for Shapley value calculation. The feature values extracted from each dimension are used as model inputs, reflecting the user's specific performance in that dimension.

[0176] Step 2: After feature extraction, the next step is to define the feature set. Assume that the dynamic portrait has n features, and each feature represents a dimension in the dynamic portrait. For example: f1 represents the instantaneous sentiment value, f2 represents the sentiment fluctuation index, f3 represents the content preference degree, f4 represents the consumption decision, etc. The feature set N is the set of all these features: N = {f1, f2, …, fn}; At this time, the feature vectors of each dimension have been extracted and formed a feature set. Next, the system will use these features for Shapley value calculation.

[0177] Step 3: Calculate the value of feature combinations. The calculation of the Shapley value is through evaluating the marginal contribution of each feature in different feature combinations. Specifically, the system needs to consider all feature combinations S and calculate the marginal contribution of each feature in these combinations. Assume that S is a subset of the feature set N, which does not contain the feature f for which the Shapley value is to be calculated i . For each subset S, the system will evaluate the value difference between the combination S and the combination S ∪ {f i}}. This difference represents the marginal contribution of the feature f i .

[0178] When calculating the value of the feature combination, the system inputs the feature set S into the pre-trained model to obtain the output value of the combination: v(S) = model output(S); where v(S) is the value of the feature combination S, representing the user portrait result predicted by the model when the feature set is S.

[0179] Step 4: Calculate the marginal contribution. Once the value v(S) of the feature combination S is obtained, the system needs to calculate the marginal contribution of the current feature f i . The marginal contribution of the feature f i is the change in the model output value when the feature f i is added to the feature combination S. Specifically, the calculation formula for the marginal contribution is:

[0180] Marginal contribution i (S) = v(S ∪ {f i}) - v(S);

[0181] Among them, v(S ∪ {f i ) is the model output value after adding feature f i to combination S, representing the result predicted by the model when feature f i is included.

[0182] Step 5: Calculate the Shapley value. Once the marginal contributions of each feature in different combinations are calculated, the next step is to calculate the Shapley value. The idea of the Shapley value is to determine the contribution of each feature to the final user profile by taking a weighted average of the marginal contributions of all feature combinations. The purpose of weighting is to consider the frequency of feature f i appearing in different feature combinations. In the calculation of the Shapley value, the marginal contribution of each combination is weighted according to the size of the combination. The calculation formula is:

[0183]

[0184] Among them, φ i (v) is the Shapley value of feature f i , representing the contribution degree of feature f i to the final profile. S is a subset of the feature set N and S does not contain element f i . v(S) is the value of feature combination S (i.e., the model output result). v(S ∪ {f i}) is the value of the feature combination S ∪ {f i} that contains feature f i . is the weight coefficient, representing the frequency of feature f i appearing in combination S. In the Shapley value calculation formula, lowercase n represents the total number of features. That is to say, n is the number of all features in the dynamic profile. For example, if the dynamic profile contains 5 features (instantaneous emotion value, emotion fluctuation index, content preference degree, consumption decision, etc.), then n = 5. The exclamation mark in the formula represents factorial. Factorial is a mathematical operation representing the product of a positive integer and all positive integers below it. In the Shapley value calculation formula, factorial is used to calculate the number of combinations, that is, the different permutations and combinations of features. For example, ∣S∣! represents the factorial of the number of elements in feature set S, representing the number of permutation ways of all features in this set. (n - ∣S∣ - 1)! is the factorial of the remaining features, and n! is the total factorial of all features.

[0185] Step 6: Output the Shapley value. By taking a weighted average of the marginal contributions of each feature, the Shapley value of each feature will finally be calculated. The higher the Shapley value of a feature

[0186] The greater the contribution to the final portrait. Through this calculation process, the system can quantify the impact of each feature on the dynamic portrait and generate a detailed analysis of the user portrait. The final Shapley value can be used to analyze the importance of different features in the user portrait, supporting further personalized recommendations and the generation of analysis reports.

[0187] Specifically, in the label generation unit, a number of multi-dimensional labels are generated according to the Shapley value and each dimension in the dynamic portrait. The multi-dimensional labels include instantaneous emotion value, emotion fluctuation index, content preference degree, and consumption decision-making mode.

[0188] Specifically, the instantaneous emotion value adopts a continuous real-time scoring mechanism, ranging from -1 to 1. A negative value represents a negative emotion, and a positive value represents a positive emotion.

[0189] Specifically, the emotion fluctuation index is based on the variance statistics of emotion values within a certain time window, reflecting the stability and fluctuation degree of the user's emotions. The calculation formula of the emotion fluctuation index is as follows:

[0190]

[0191] In the formula, EI represents the emotion fluctuation index, x i is the emotion value at each time point, μ is the average emotion value, n is the number of samples, and the emotion value is determined according to the Shapley value related to the emotion dimension.

[0192] Specifically, the content preference degree is realized through a non-linear dimensionality reduction method. By compressing the behavioral features involved in the dynamic portrait (such as likes, views, collections, etc.), the multi-dimensional features are compressed into a two-dimensional plane to achieve content label clustering. Then, a non-linear dimensionality reduction method is used to map the high-dimensional features to a low-dimensional space to identify clustering labels such as "suspense preference" and "urban emotion". The preference intensity coefficient is calculated based on the clustering density and behavioral frequency, and is normalized to the 0-1 interval, where 0 represents no interest at all, and 1 represents extreme preference. The calculation formula of the content preference degree is as follows:

[0193]

[0194] In the formula, PI represents the content preference degree, w i is the weight of the i-th cluster, f i is the behavioral frequency of the i-th cluster, max(w i ·f i ) is the maximum value of the product of the weights and behavioral frequencies of all clusters.

[0195] Specifically, the consumption decision-making mode is determined by the Shapley values of the dimensions involving micro-expressions, eye movement tracking, and historical recharge records in the dynamic portrait. For "impulsive" users, the system captures the characteristics of a decision-making time less than 3 seconds and an upward angle of the mouth corner greater than 15°; for "rational" users, the stability of multiple re-look behaviors and pupil changes is analyzed. Decision type recognition combines multi-modal features through a machine learning classifier to construct a probability model of decision-making behavior. Key features include decision-making time, facial micro-expression angle, eye movement trajectory stability, etc., and the user's decision type is comprehensively judged through an ensemble learning algorithm.

[0196] Specifically, the D3.js visualization library can be used to dynamically display the evolution path of the user's content preferences based on the timeline. Taking "the preference for ancient costume dramas changes from 0.72 to 0.35, and the preference for workplace dramas changes from 0.18 to 0.61" as an example, the dynamic migration process of the user's interests is intuitively presented through color gradients and curve trajectories. In the visualization design, the color maps the change in emotional tendency from cold tones to warm tones, and the smoothness of the curve reflects the smoothness of the preference change. The interactive chart allows users to deeply explore the evolution process of their personal portraits.

[0197] Specifically, in the analysis unit, the portrait analysis report includes several multi-dimensional tags, the values of each multi-dimensional tag, and the meanings of each multi-dimensional tag. The portrait analysis report also includes the evolution path of content preferences, which is used to reflect the process of the user's preference change.

[0198] Specifically, it also includes a dual-channel feedback mechanism. The dual-channel feedback mechanism includes an explicit feedback channel and an implicit feedback channel. The explicit feedback channel updates the tag weights through the user's active scoring (e.g., 1 star to 5 stars); the implicit feedback channel monitors indicators such as the viewing completion rate (>95% is counted as a positive sample) and the next-day retention rate. When a new user undergoes cold start, a collaborative filtering algorithm based on content similarity is used to generate an initial portrait, which is gradually replaced by the multi-modal analysis result within 24 hours. The model performance monitoring module continuously tracks indicators such as the Area Under the Curve (AUC) and the Normalized Discounted Cumulative Gain (NDCG). When the NDCG@10 drops by more than 5%, a full-scale model retraining is automatically triggered to ensure the continuous optimization of the portrait system. The cold start is for the generation of new user portraits, and the system uses a collaborative filtering algorithm based on content similarity. In the initial stage, a sparse feature vector is constructed through the user's registration information, initial interest tags, etc. The similarity calculation uses the cosine similarity algorithm, which will match the most similar historical user portrait for the new user and generate a preliminary recommendation. Within 24 hours, the system gradually replaces the initial portrait with the multi-modal depth analysis result to achieve rapid personalization.

[0199] Specifically, the dual-channel feedback mechanism further includes a performance monitoring mechanism, which is used to track key evaluation metrics, including AUC (Area Under the Curve) and NDCG (Normalized Discounted Cumulative Gain). The formula for NDCG@10 is: NDCG@10 = DCG@10 / IDCG@10, where DCG (Discounted Cumulative Gain) measures the relevance of the recommendation list, and IDCG is the best gain in the ideal case. When the NDCG@10 metric drops by more than 5%, the system automatically triggers the full-scale model retraining process. The retraining process includes incremental data collection, feature engineering reconstruction, model parameter reset, multiple rounds of cross-validation, and gray release verification to ensure the continuous stability and improvement of model performance. The model iteration is based on a closed-loop feedback system, constructing a multi-dimensional and adaptive optimization framework. The strategies include dynamically adjusting feature weights, incremental learning mechanisms, multi-modal feature fusion, and model integration. By real-time monitoring user interaction signals and model performance metrics, the system can intelligently adjust the learning algorithm parameters. For example, when the contribution of certain features to user portrait prediction decreases, their weights can be dynamically reduced through an automatic feature selection algorithm (such as mutual information-based feature importance evaluation). Model integration improves the robustness and generalization ability of the overall system by combining the prediction results of multiple different algorithms (such as deep neural networks, tree models, and probabilistic graphical models). During the user portrait iteration process, the system strictly follows the principles of data desensitization and anonymization. Differential privacy technology is used to introduce random noise during model training to protect personal privacy. Specifically, Gaussian noise is added during gradient update to ensure that individual user data does not have a significant impact on the model output. The formula can be expressed as: Noisy Gradient = Original Gradient + N(0,σ2), where σ is the privacy budget control parameter.

[0200] Embodiment 2

[0201] As Figure 7 shown, this embodiment provides a method for analyzing the user portrait of micro-short dramas based on multi-modal data, including:

[0202] S10: Obtain multi-modal micro-short drama user data;

[0203] S20: Extract features from the multi-modal micro-short drama user data to obtain several user features;

[0204] S30: Input several of the user features into a pre-trained hierarchical attention network model to obtain a joint feature vector;

[0205] S40: Input the joint feature vector into a preset online deep forest model to generate a dynamic portrait;

[0206] S50: Calculate the Shapley value of each dimension of the dynamic portrait, and generate a portrait analysis report according to the Shapley value.

[0207] Specifically, in step S10, the following steps are included:

[0208] S11: Obtain initial multimodal data from different data sources;

[0209] S12: Synchronize the timestamps of the initial multimodal data from different sources according to the Network Time Protocol to obtain time-synchronized initial multimodal data;

[0210] S13: Process the time-synchronized initial multimodal data using cubic spline interpolation and noise suppression algorithms to obtain multimodal data.

[0211] Specifically, in step S10, the initial multimodal data includes initial behavior data, initial user facial data, initial user voice data, and initial text data. The initial behavior data includes click frequency (single / double click mode), sliding trajectory (direction / speed), pause / rewind operations, and device gyroscope attitude data. The acquisition frequency of the initial behavior data can be 10 Hz. The initial user facial data captures the user's facial video stream through the front camera of the mobile phone at a frame rate of 30 fps, and uses the MediaPipe framework to construct a three-dimensional facial mesh of 62 key points in real time, and focuses on extracting micro-expression parameters such as eyelid opening degree, glabella muscle displacement, and mouth corner upward angle. The initial user voice data is obtained through a microphone. The initial user voice data is the user's real-time voice comment. After using a noise suppression algorithm to eliminate ambient noise, it is saved as a Pulse Code Modulation (PCM) format audio segment at a sampling rate of 16 kHz. The initial text data is obtained in real time through the WebSocket protocol. The initial text data includes the bullet screen content sent by the user and the sending time point.

[0212] Specifically, in step S20, the following steps are included:

[0213] S21: Use a temporal convolutional network to extract features from the behavior data to obtain a number of first user features;

[0214] S22: Use a pre-trained object detection model to extract features from the user facial data according to the visual database to obtain a number of second user features;

[0215] S23: Input the user voice data into a pre-trained voice representation model for feature extraction to obtain a number of third user features;

[0216] S24: Input the text data into a pre-trained language model for feature extraction to obtain fourth user features.

[0217] Specifically, in steps S21 to S24, the first user feature is obtained by weighting the viewing concentration feature, the content preference intensity feature, and the interaction behavior feature;

[0218] The second user feature is obtained by weighting the eye region feature, the mouth region feature, and the comprehensive facial feature;

[0219] The third user feature is obtained by weighting the excitement level feature, the disappointment intensity feature, the surprise index feature, the anger level feature, the fear level feature, and the neutral emotion ratio feature;

[0220] The fourth user feature is obtained by weighting the positive emotion intensity feature, the negative emotion intensity feature, the network buzzword usage density feature, the emotion fluctuation frequency feature, the interaction enthusiasm feature, the emotion consistency feature, the group emotion synchronization feature, and the emotion depth feature.

[0221] Specifically, the first user feature, the second user feature, the third user feature, the fourth user feature, and the features subordinate to each user feature dynamically adjust the contribution degrees of different features through learning parameters. The weight calculation uses the softmax function, and the original features are non-linearly transformed through the learnable weight matrix W and the bias term b to obtain the importance distribution of each feature. This adaptive weight learning mechanism enables the model to automatically adjust the weights of each modal feature according to different scenarios and user features, achieving a more accurate user portrait representation.

[0222] Specifically, in step S30, the following steps are included:

[0223] S31: Preprocess several of the user features to obtain several preprocessed features;

[0224] S32: Divide several of the preprocessed features into several modalities, calculate the attention scores of each preprocessed feature within each modality, and select key sub-features from each modality according to the attention scores;

[0225] S33: Use the cross-modal attention algorithm to calculate the correlation degrees between pairwise modalities;

[0226] S34: Generate a joint feature vector according to the key sub-features and the correlation degrees.

[0227] Specifically, in step S31, the preprocessing includes standardization and normalization. Since the user features include data of multiple modalities (such as text, speech, facial expressions, etc.), for each modality m, its feature vector By using batch normalization or layer normalization, eliminate the scale differences in the feature spaces of different modalities to ensure the fairness and stability of subsequent attention mechanism calculations. Normalization processing helps reduce the bias in the feature distributions between modalities.

[0228] Specifically, in step S32, when dividing the several preprocessed features into several modalities, divide them according to the data sources of the preprocessed features. For example, the preprocessed features of the voice category are divided into the unified voice modality.

[0229] Specifically, in step S33, calculate the semantic correlation degrees between different modalities through the learnable query matrix Q (Query), key matrix K (Key), and value matrix V (Value).

[0230] Specifically, in step S34, according to the correlation degrees between modalities and several key sub-features, integrate the key sub-features of different modalities through weighted summation or more complex fusion strategies. The joint feature vector comprehensively reflects the semantic information of each modality and the interaction relationship between modalities, providing rich multi-modal representations for subsequent sentiment classification tasks.

[0231] Specifically, in step S40, it includes the following steps:

[0232] S41: Select several feature subsets with different granularities from the joint feature vector;

[0233] S42: Adjust the influence weights of the several feature subsets according to a preset time decay factor;

[0234] S43: Input the several feature subsets and the influence weights into the pre-trained online deep forest model to obtain an initial user profile;

[0235] S44: Perform incremental learning at preset time intervals to update the user profile and obtain a dynamic profile.

[0236] Specifically, in step S41, adopt the random subspace method, feature importance evaluation method, or principal component analysis method for the joint feature vector to select feature subsets with different granularities. This method can effectively reduce dimensions, prevent overfitting, and at the same time enhance the generalization ability of the model, enabling each base learner to capture the subtle changes in user interests from different perspectives.

[0237] Specifically, in step S42, the expression of the decay factor is: λ = 0.98 t, where \(t\) is the data time difference in hours. By setting the decay factor, the weights of historical samples are dynamically adjusted, enabling the Online Deep Forest Model (ODF-MT) to pay more attention to recent behavior patterns. It can exponentially decay the influence of historical data. For example, the weight of data 1 hour ago is approximately 0.98, the weight of data 24 hours ago drops to 0.376, and the weight of data 72 hours ago is only 0.014. Through this refined weight adjustment strategy, the model can quickly respond to real-time changes in user interests, while retaining the reference value of historical behavior, improving the sensitivity and adaptability to recent behavior patterns, and thus achieving dynamic and accurate updates of user portraits.

[0238] Specifically, in step S43, the online deep forest model includes 3 incremental random forest models and 2 dynamic gradient boosting tree models. Each base learner receives feature subsets with different granularities. A base learner is a single classification or regression model in ensemble learning. By inputting the feature subsets and the corresponding influence weights of each feature subset into different models, an initial user portrait is obtained.

[0239] Specifically, in step S44, the preset time interval is preferably 30 minutes. By adopting a misaligned binning strategy to update the leaf node distribution (please explain the misaligned binning strategy), the node is automatically split when the KL divergence of the new data distribution exceeds the threshold \(\theta = 0.15\). At the same time, the system maintains the user status matrix \(U\in R\) 128 , and fuses long-term and short-term interest features through the LSTM update gate mechanism.

[0240] Specifically, in step S50, it includes the following steps:

[0241] S51: Calculate the Shapley values of each dimension in the dynamic portrait;

[0242] S52: Generate a number of multi-dimensional labels according to each dimension in the dynamic portrait and the corresponding Shapley values;

[0243] S53: Generate a portrait analysis report according to a number of the multi-dimensional labels.

[0244] Specifically, in step S51, when extracting features from the dynamic portrait, each dimension of the dynamic portrait corresponds to a portrait feature, obtaining a number of portrait features. The Shapley values of each portrait feature are calculated by inputting the number of portrait features into the Shapley library.

[0245] Specifically, in the step S52, a number of multi-dimensional tags are generated according to the Shapley value and each dimension in the dynamic portrait. The multi-dimensional tags include instantaneous emotion value, emotion fluctuation index, content preference degree, and consumption decision. The instantaneous emotion value adopts a continuous real-time scoring mechanism, with a range between -1 and 1. A negative value represents a negative emotion, and a positive value represents a positive emotion. The emotion fluctuation index is based on the variance statistics of emotion values within a certain time window, reflecting the stability and fluctuation degree of the user's emotion. The content preference degree is achieved through a non-linear dimensionality reduction method. By compressing the behavioral characteristics (such as likes, views, collections, etc.) involved in the dynamic portrait, the multi-dimensional features are compressed into a two-dimensional plane to achieve content label clustering. Then, the non-linear dimensionality reduction method is used to map the high-dimensional features to a low-dimensional space to identify clustering labels such as "suspense preference" and "urban emotion". The preference intensity coefficient is calculated based on clustering density and behavioral frequency, and is normalized to the range of 0-1, where 0 means completely uninterested and 1 means extremely preferred. The consumption decision is determined by the Shapley value of the dimensions involving micro-expressions, eye movement tracking, and historical recharge records in the dynamic portrait. For "impulsive" users, the system captures the characteristics of a decision-making time less than 3 seconds and an upward angle of the mouth corner greater than 15°. For "rational" users, the stability of multiple re-watching behaviors and pupil changes is analyzed. The decision type recognition combines multi-modal features through a machine learning classifier to construct a probability model of decision-making behavior. The key features include decision-making time, facial micro-expression angle, eye movement trajectory stability, etc., and the user's decision type is comprehensively judged through an ensemble learning algorithm.

[0246] Specifically, in the step S53, the portrait analysis report includes a number of multi-dimensional tags, the values of each multi-dimensional tag, and the meanings of each multi-dimensional tag. The portrait analysis report also includes the evolution path of content preference, which is used to reflect the process of the user's preference change.

[0247] Those skilled in the art can clearly understand that, for the convenience and simplicity of description, only the above-mentioned division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiment can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of the present application. The specific working processes of the units and modules in the above system can refer to the corresponding processes in the foregoing method embodiments and will not be repeated here.

[0248] An embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements any one of the above methods.

[0249] The present invention also provides an electronic device. The electronic device according to an embodiment of the present invention includes: one or more processors; a storage device for storing one or more programs, and when the one or more programs are executed by the one or more processors, the one or more processors implement a method for analyzing the user portrait of micro short plays based on multi-modal data provided by the present invention.

[0250] Reference is made below Figure 8 , which shows a schematic structural diagram of a computer system 800 of an electronic device suitable for implementing an embodiment of the present invention. Figure 8 The shown electronic device is only an example and should not bring any limitation to the functions and usage scope of the embodiments of the present invention.

[0251] As Figure 8 shown, the computer system 800 includes a central processing unit (CPU) 801, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 802 or the program loaded from the storage section 808 into the random access memory (RAM) 803. In the RAM 803, various programs and data required for the operation of the computer system 800 are also stored. The CPU 801, ROM 802, and RAM 803 are connected to each other via a bus 804. The input / output (I / O) interface 805 is also connected to the bus 804.

[0252] The following components are connected to the I / O interface 805: an input section 806 including a keyboard, a mouse, etc.; an output section 807 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage section 808 including a hard disk, etc.; and a communication section 809 including a network interface card such as a LAN card, a modem, etc. The communication section 809 performs communication processing via a network such as the Internet. A drive 810 is also connected to the I / O interface 805 as required. A removable medium 811, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 810 as required, so that the computer program read from it is installed into the storage section 808 as required.

[0253] The above specific embodiments do not constitute a limitation to the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can occur depending on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A micro short drama user portrait analysis system based on multimodal data, characterized in that, Including: A data acquisition module for acquiring multi-modal micro short drama user data; A feature extraction module for extracting features based on the multi-modal micro short drama user data to obtain a number of user features; A fusion module for inputting a number of the user features into a pre-trained hierarchical attention network model to obtain a joint feature vector; A dynamic portrait generation module for inputting the joint feature vector into a preset online deep forest model to generate a dynamic portrait; A portrait analysis module for calculating the Shapley value of each dimension of the dynamic portrait and generating a portrait analysis report based on the Shapley value.

2. The multimodal data-based micro short drama user portrait analysis system according to claim 1, wherein In the data acquisition module, it specifically includes: A data acquisition unit for acquiring initial multi-modal data from different data sources; A time alignment unit for synchronizing the timestamps of the initial multi-modal data from different sources according to the network time protocol to obtain time-synchronized initial multi-modal data; A first preprocessing unit for processing the time-synchronized initial multi-modal data using cubic spline interpolation and a noise suppression algorithm to obtain multi-modal micro short drama user data.

3. The micro short drama user portrait analysis system based on multimodal data according to claim 1, wherein The multi-modal micro short drama user data includes behavior data, user facial data, user voice data, and text data. The number of the user features includes a number of first user features, a number of second user features, a number of third user features, and a number of fourth user features. In the feature extraction module, it specifically includes: A first extraction unit for extracting features from the behavior data using a temporal convolutional network to obtain a number of first user features; A second extraction unit for extracting features from the user facial data using a pre-trained object detection model according to a visual database to obtain a number of second user features; A third extraction unit for inputting the user voice data into a pre-trained speech representation model for feature extraction to obtain a number of third user features; A fourth extraction unit for inputting the text data into a pre-trained language model for feature extraction to obtain fourth user features.

4. The multi-modal data-based micro short drama user portrait analysis system according to claim 1, wherein, In the fusion module, it specifically includes: A second preprocessing unit for preprocessing a number of the user features to obtain a number of preprocessed features; An intra-modal attention unit for dividing a number of the preprocessed features into a number of modalities, calculating the attention scores of each of the preprocessed features within each modality, and selecting key sub-features from each modality according to the attention scores; A cross-modal attention unit for calculating the correlation between pairwise modalities using a cross-modal attention algorithm; A fusion unit for generating a joint feature vector based on the key sub-features and the correlation.

5. The micro-short drama user portrait analysis system based on multi-modal data according to claim 1, characterized in that, In the dynamic portrait generation module, it specifically includes: A subset establishment unit for selecting a number of feature subsets with different granularities from the joint feature vector; An attenuation unit for adjusting the influence weights of a number of the feature subsets according to a preset time decay factor; A portrait generation unit for inputting a number of the feature subsets and the influence weights into the pre-trained online deep forest model to obtain an initial user portrait; An update unit for performing incremental learning at a preset time interval to update the user portrait to obtain a dynamic portrait.

6. The multimodal data-based micro short drama user portrait analysis system according to claim 1, wherein, In the portrait analysis module, it specifically includes: The Shapley value calculation unit is used to calculate the Shapley values of each dimension in the dynamic portrait; The label generation unit is used to generate a number of multi-dimensional labels according to each dimension in the dynamic portrait and the Shapley value corresponding to each dimension; The analysis unit is used to generate a portrait analysis report according to a number of the multi-dimensional labels.

7. The system for analyzing the user portrait of micro-short plays based on multi-modal data according to claim 3, wherein, The first user feature is obtained by weighting the viewing concentration feature, the content preference intensity feature, and the interaction behavior feature; The second user feature is obtained by weighting the eye region feature, the mouth region feature, and the comprehensive facial feature; The third user feature is obtained by weighting the excitement degree feature, the disappointment intensity feature, the surprise index feature, the anger level feature, the fear degree feature, and the neutral emotion ratio feature; The fourth user feature is obtained by weighting the positive emotion intensity feature, the negative emotion intensity feature, the network hot word usage density feature, the emotion fluctuation frequency feature, the interaction enthusiasm feature, the emotion consistency feature, the group emotion synchronization feature, and the emotion depth feature.

8. The multimodal data-based micro short drama user portrait analysis system according to claim 4, characterized in that, In the dynamic portrait generation module, there is also a dynamic feature enhancement unit, and the dynamic feature enhancement unit is used to detect a number of the key sub-features, and when a special field appears in the number of the key sub-features, adjust the weight of the key sub-feature with the special field when generating the joint feature vector.

9. The system for analyzing the user portrait of micro short plays based on multi-modal data according to claim 5, wherein The online deep forest model includes three incremental random forest models and two dynamic gradient boosting tree models.

10. A method for analyzing the user portrait of micro short plays based on multi-modal data, characterized in that, Including: S10: Obtain multi-modal micro short drama user data; S20: Extract features according to the multi-modal micro short drama user data to obtain a number of user features; S30: Input the number of the user features into a pre-trained hierarchical attention network model to obtain a joint feature vector; S40: Input the joint feature vector into a preset online deep forest model to generate a dynamic portrait; S50: Calculate the Shapley values of each dimension of the dynamic portrait, and generate a portrait analysis report according to the Shapley values.

Citation Information

Patent Citations

  • Multi-mode jeamosity detection method and device, computer equipment and storage medium

    CN117633516A

  • Travel user portrait construction method and system

    CN117972204A

  • Security propaganda and education recommendation method and system based on demand portrait and content label

    CN118797173A

  • Big data-based live broadcast room crowd portrait generation method and system

    CN118916731A

  • Big data-based network live broadcast e-commerce marketing management system and method

    CN119151591A

Cited By

  • Emotion analysis method and system based on multi-modal enhanced retrieval generation technology

    CN120524444A

  • New media product user portrait analysis system and method based on media big data

    CN120894063A

  • Fitness equipment control method and system

    CN121239721A