Short video traffic prediction method based on multi-modal data
By constructing a multimodal short video traffic prediction model, the problems of insufficient dynamic temporal evolution of multimodal features and cross-modal correlation in short video traffic prediction are solved, achieving more accurate and stable traffic prediction and supporting content recommendation and advertising placement decisions.
Patent Information
- Application Number
- CN202511093352.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-06
- Publication Date
- 2025-10-28
AI Technical Summary
Existing short video traffic prediction methods struggle to effectively capture the dynamic temporal evolution of multimodal features, fail to fully explore the potential correlations between video, audio, and text modalities, and the text information cannot be correlated with frame-level video/audio, leading to prediction bias and noise.
A multimodal short video traffic prediction model is constructed, including a multimodal feature extraction module, a temporal modeling module, a modal attention module, a global modulation module, and a prediction head module. Through multimodal feature extraction, temporal modeling, modal attention fusion, and adaptive loss optimization, cross-modal interaction modeling and information fusion are achieved.
It significantly improves the accuracy and stability of short video traffic prediction, reduces reliance on manual feature engineering, and provides a reliable basis for content recommendation and advertising.
Smart Images

Figure CN120856909A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of artificial intelligence and big data analysis technology, specifically to a method for predicting short video traffic based on multimodal data. Background Art
[0002] Short videos are a new content format characterized by their short duration and integration of audio, visual, and text elements. With their fragmented, entertaining, and easily shareable nature, they have become an important medium for information dissemination and mass entertainment. Short video traffic refers to the reach and attention a short video receives on online platforms, typically measured by metrics such as views, likes, comments, shares, and saves. Higher traffic indicates greater video exposure, a wider audience, and greater influence.
[0003] Short video traffic prediction refers to estimating future traffic metrics such as views, likes, and comments for short videos by analyzing various relevant data and factors. Short video traffic prediction is one of the core technologies for short video platform operation and content recommendation. Existing methods mainly rely on single-modal data (such as time-series statistics of video views or text tag analysis) for modeling, which has the following problems:
[0004] 1) Short video traffic is affected by multiple factors such as video content, user interaction, and platform strategy. Traditional methods are unable to capture the dynamic temporal evolution of multimodal features.
[0005] 2) Potential correlations between video, audio, and text modalities (such as background music and user emotions, text descriptions and content dissemination) were not effectively explored, leading to prediction bias.
[0006] 3) Text information is often a whole paragraph description, which cannot be matched one-to-one with frame-level video / audio, causing the text semantics to fail to play a role at the correct time and increasing prediction noise. Summary of the Invention
[0007] To address the shortcomings of existing short video traffic prediction methods, such as insufficient utilization of multimodal data, weak cross-modal interaction modeling capabilities, and inaccurate capture of temporal dynamic dependencies, this invention provides a short video traffic prediction method based on multimodal data. The method involves constructing and training a multimodal short video traffic prediction model, and then using this trained model for traffic prediction. The multimodal short video traffic prediction model includes a multimodal feature extraction module, a temporal modeling module, a modal attention module, a global modulation module, and a prediction head module.
[0008] The training process of the multimodal short video traffic prediction model includes the following steps:
[0009] S1. Collect a short video dataset consisting of multiple sets of samples; each set of samples includes structured video modality data, structured audio modality data, and structured text modality data of the same short video;
[0010] S2. For each set of samples, extract video modal feature representation, audio modal feature representation, and text modal feature representation through the multimodal feature extraction module;
[0011] S3. The temporal modeling module is used to process the video modal feature representation and the audio modal feature representation to obtain the temporal dependency representation of the video modality and the temporal dependency representation of the audio modality;
[0012] S4. A modal attention module is used to process the video modal temporal dependency representation and the audio modal temporal dependency representation to obtain a dual-modal fused frame vector sequence;
[0013] S5. Input the text modal feature representation and the dual-modal fusion frame vector sequence into the global modulation module to obtain the frame-level fusion feature sequence; send the frame-level fusion feature sequence into the prediction head module to obtain the prediction result;
[0014] S6. Calculate the adaptive multimodal collaborative optimization loss based on the prediction results, and iteratively train the model through backpropagation until the model converges.
[0015] The beneficial effects of this invention are:
[0016] This invention introduces different pre-trained networks for different modal data and maps the features of each modality to a shared space through a unified projection, which greatly enriches the available information sources.
[0017] This invention introduces a nonlinear temporal context aggregator to fully capture cross-frame evolution patterns, significantly improving the ability to characterize the temporal dynamics of short videos.
[0018] This invention explicitly measures frame-level intramodal dependencies and heteromodal complementarities through intramodal self-attention and cross-modal mutual attention, fully exploring potential connections such as the rhythm of background music and the mood of the scene, and the semantics of the text and the content of the shot, thus solving the problem of insufficient cross-modal interaction in existing methods.
[0019] This invention uses text modal feature representation to perform gated modulation on the dual-modal fusion frame vector, so that the static text semantics have a consistent impact on the entire video, avoiding information noise caused by timeline misalignment.
[0020] This invention combines the outputs of the master prediction head and three single-modal prediction heads to construct an adaptive multimodal collaborative optimization loss. During the training phase, it dynamically allocates the weights of each modality and constrains the fusion consistency, thereby improving the model convergence speed and suppressing overfitting.
[0021] The method proposed in this invention significantly improves the accuracy and stability of short video traffic prediction through multimodal deep fusion, cross-modal interactive modeling, and adaptive loss optimization, while reducing the reliance on manual feature engineering. It can provide a more reliable decision-making basis for scenarios such as content recommendation, advertising, and bandwidth resource allocation. Attached Figure Description
[0022] Figure 1 This is a flowchart of the model processing framework shown in some embodiments of the present invention;
[0023] Figure 2 This is a schematic diagram illustrating the specific processing flow of the method shown in some embodiments of the present invention. Detailed Implementation
[0024] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0025] Figure 1 This is a flowchart of a model processing framework according to some embodiments of the present invention. Figure 2 This is a schematic diagram of the specific processing flow of the method according to some embodiments of the present invention.
[0026] Some embodiments of the present invention provide a method for predicting short video traffic based on multimodal data, including constructing and training a multimodal short video traffic prediction model, and using the trained multimodal short video traffic prediction model to predict traffic; the multimodal short video traffic prediction model includes a multimodal feature extraction module, a temporal modeling module, a modal attention module, a global modulation module, and a prediction head module.
[0027] like Figure 1 , Figure 2 As shown, the training process of the multimodal short video traffic prediction model includes the following steps:
[0028] S1. Collect a short video dataset consisting of multiple sets of samples; each set of samples includes structured video modality data, structured audio modality data, and structured text modality data of the same short video.
[0029] In some embodiments, step S1 obtains a set of samples, including:
[0030] S11. Obtain the video modal data, audio modal data, and text modal data corresponding to a short video after it has been published for a period of time. The text modal data includes the title, description, subtitles, user interaction information, and the number of short video views.
[0031] S12. Perform length standardization processing on the video modal data to obtain video modal structured data, including: determining whether the number of frames in the video modal data is less than a preset number L; if so, fill in predefined frames to make the number of frames reach L to form video modal structured data; otherwise, randomly select L frames to form video modal structured data.
[0032] S13. The audio modal data is length normalized, converted to mono, and then resampled to 16kHz. The resampled data is subjected to a short-time Fourier transform with a 25ms Hanning window and a 10ms step size to obtain the amplitude spectrum. The amplitude spectrum is mapped to a 64-dimensional Mel filter bank to obtain the Mel spectrum. The Mel spectrum is logarithmically compressed and dynamically clipped (Top-dB=80dB) to enhance weak signals, thus obtaining structured audio modal data.
[0033] The length normalization process for audio modal data is similar to that for video modal data. This operation enhances sample diversity and model generalization ability while ensuring the consistency of input length.
[0034] Preferably, in the inference stage, to more comprehensively cover the temporal information of the input data, a sliding window approach is used to segment the video modal data and audio modal data respectively. That is, for any one modal data, a fixed window length L and a sliding step size S are used to continuously truncate and generate multiple local subsequences. Each local subsequence of the two modal data is input into the model for prediction. Each local subsequence is independently sent to the subsequent S2~S5 pipelines. Finally, the results of the same video are averaged, thereby improving the perception ability and prediction accuracy of temporal changes in long sequences.
[0035] S14. Normalize the text modal data to Unicode, remove emojis, desensitize URLs and @user tags, unify full-width and half-width characters, and unify uppercase and lowercase. Use a Chinese word segmentation tool in conjunction with a user dictionary to perform maximum positive matching on the processed text to obtain the text modal structured data.
[0036] S2. For each set of samples, extract video modal feature representation, audio modal feature representation, and text modal feature representation through the multimodal feature extraction module.
[0037] In some embodiments, in step S2, to better extract features from the structured data of each modality, different deep learning networks are selected for different modalities, and the selected deep learning networks are pre-trained. During the pre-training process of each deep learning network, to prevent overfitting and accelerate convergence, the task-related layers after the penultimate layer of each deep learning network are removed, and the remaining parameters are... Freeze, only the mapping layer Participate in fine-tuning.
[0038] For example, for video modalities, ResNet-50Inflated-3D (I3D-ResNet50) can be used to obtain frame-level features that combine static appearance and motion dynamics; for audio modalities, VGGish can be used, with the addition of Temporal-ConvModule to distinguish between short and long audio events; for text modalities, BGE-Large-zh (based on the RoBERTa architecture) can be used to obtain sentence-level and paragraph-level dual-scale representations.
[0039] For example, the overall process of step S2 includes:
[0040] S21. Use pre-trained I3D-ResNet50 to extract video features from video modal structured data. The pre-trained VGGish was used to extract audio features from the audio modal structured data. ;d V d represents the original video feature dimension (the number of channels in the global pooling output at the end of I3D-ResNet50). A This represents the original audio feature dimension (the length of the VGGish (plus TCM) semantic embedding). To represent the field of real numbers, i.e., to denote... Represents the dimension The real vector space.
[0041] S22. The word-level matrix and paragraph-level representation vector of the text modal structured data are extracted using pre-trained BGE-Large-zh. To enhance multi-granularity information, residual mean pooling is performed to obtain a static global text representation, expressed as:
[0042]
[0043] Among them, g T Represents a static global text representation. Represents the word-level matrix E T The element in the i-th row, n represents the row number, s T The layer represents the paragraph-level representation vector, and LayerNorm() represents layer normalization; d T This represents the original text feature dimension (BGE-Large-zh Hidden size).
[0044] S23. Project the video features, audio features, and static global text representation onto the same dimension to obtain video modal feature representation, audio modal feature representation, and text modal feature representation.
[0045] For example, to ensure consistency across the three modal dimensions and seamless integration with subsequent modules, a unified projection matrix is introduced. , , d C This represents the preset common feature dimension hyperparameter; after processing, we obtain:
[0046]
[0047]
[0048] For video modal feature representation, For audio modal feature representation, , All are frame-level sequences of length L; This represents the text modal features.
[0049] For example, d C The value can be 256, 512, etc.
[0050] S3. The temporal modeling module is used to process the video modal feature representation and the audio modal feature representation to obtain the temporal dependency representation of the video modality and the temporal dependency representation of the audio modality.
[0051] In some embodiments, step S3 includes:
[0052] S31. Perform preliminary temporal modeling on the video feature representation and audio feature representation to obtain the video modal comprehensive evolution and audio modal comprehensive evolution at different time steps, expressed as:
[0053]
[0054] in, This represents the D-mode synthesis evolution at time step t. The D-modal feature represents the high-dimensional feature vector at time step t. It is obtained by forward reasoning of D-modal structured data by a pre-trained deep learning network, and is used to represent the instantaneous information of time step t; The D-modal feature represents the high-dimensional feature vector at time step t-1, which is related to... We define a common temporal difference to characterize the evolution of features over time; D={A,V}, where A is the audio modality and V is the video modality; f() represents a learnable multi-input nonlinear combination function used to jointly process historical features, current features, temporal derivatives, and integral terms.
[0055] For example, f() can be implemented using a multilayer perceptron (MLP), a gated recursive unit (GRU), or other neural network structures with general approximation capabilities.
[0056] S32. A nonlinear temporal context aggregator is used to update the video modal synthesis evolution and audio modal synthesis evolution at each time step to obtain the video modal temporal dependency representation and audio modal temporal dependency representation at the corresponding time step.
[0057] For each time t, the update of the modal features depends not only on the features at the current time, but also on the features of past and future times within the window. The update formula is expressed as:
[0058]
[0059] Where k represents a hyperparameter, which represents the size of the half-window and is used to control the length of the timing context; This represents the D-mode temporal dependency representation at time step t. This represents the D-mode synthesis evolution of the i-th frame within the sliding window. express First-order difference, σ 2 This represents the Gaussian kernel standard deviation of the hyperparameter, used to control how quickly the weights decay with distance. For follow Increase the weight decay of frames (i.e., frames farther from the center frame of the window) to suppress the excessive influence of frames far from the center frame of the window on the aggregation result and improve temporal smoothness. It uses a Gaussian kernel weight to highlight frames that are closer in time to the center frame of the window.
[0060] This invention takes into account the interaction between history and the future and nonlinear transformations, and introduces a nonlinear temporal context aggregator. By weighted summing of the comprehensive evolution and its temporal derivative of each frame within the sliding window, it achieves effective fusion of historical and future time information, thereby obtaining a more temporally informative feature representation.
[0061] S4. The modal attention module is used to process the video modal temporal dependency representation and the audio modal temporal dependency representation to obtain the dual-modal fused frame vector sequence.
[0062] In some embodiments, step S4 includes:
[0063] S41. For video modality temporal dependency representation and audio modality temporal dependency representation respectively, a self-attention mechanism is used to calculate the correlation between each feature in the temporal dependency representation, and each feature is weighted to ensure that important features dominate the prediction.
[0064] Specifically, the video modal self-attention weights and audio modal self-attention weights are calculated as follows:
[0065] All time steps Stacking yields the feature matrix H of the D modes D :
[0066]
[0067] Through linear projection, we obtain:
[0068]
[0069]
[0070]
[0071]
[0072]
[0073] in, This represents the temporal dependency of the D modalities at time step t, where D = {A, V}, A is the audio modality, and V is the video modality; L represents the sequence length (the number of time steps per sample after standardization), and Q... D K represents the query matrix of mode D. D V represents the bond matrix of the D modes. D The value matrix representing the D modes; , , For trainable weight matrix, Indicates self-attention bias, used to encode relative position or mask information; d k S represents the dimension of the key matrix. D This represents the self-attention scoring matrix (an unnormalized similarity matrix, with weights obtained after Softmax). ), Softmax() represents the activation function, α D This represents the D-mode self-attention weight, used to characterize the inter-frame correlation within the same modality;
[0074] S42. Calculate the video modal update features and audio modal update features, represented as follows:
[0075]
[0076] in, Denotes the D-modal update features, LayerNorm() denotes layer normalization, ε D ρ represents the regularization term for stabilization.D Indicates the learnable bias term; The result of a frame-by-frame weighted summation of the Value vector.
[0077] S43. To ensure It can input subsequent cross-modal mutual attention, and the dimension is the same as the common feature dimension d. C To achieve consistency, the video modality update features and audio modality update features are projected onto each other to obtain the video modality projection feature matrix and the audio modality projection feature matrix, which are represented as follows:
[0078]
[0079] in, The projected weight matrix is a learnable matrix. This represents the feature matrix of the D-mode projection.
[0080] S44. To capture the interaction relationships between different modalities, calculate the cross-modal mutual attention coefficients in the video-to-audio and audio-to-video directions, expressed as:
[0081]
[0082]
[0083]
[0084] in, The query matrix represents the D1 modality. The key matrix representing the D2 mode. This represents the characteristic matrix of the D1 modal projection. D1 represents the characteristic matrix of the D2 modal projection, where D1 = {A, V}; , This represents the trainable directional projection matrix from mode D1 to mode D2. This represents the bias term, used to supplement the calculation of cross-modal attention; Represents the cross-modal mutual attention coefficients from mode D1 to mode D2, used to quantify the degree of alignment between the two modes at the frame level; (D1,D2)∈{(A,V), (V,A)}; The modal direction indicates that the modal direction is as follows: For Query, read information.
[0085] S45. Calculate the bimodal fused frame vector at each time step to form a bimodal fused frame vector sequence, represented as:
[0086]
[0087]
[0088] in, Represents the cross-modal weighting matrix from mode D1 to mode D2 ( from (The context representation after reading) The characteristic matrix representing the D2 mode, This represents a trainable weight matrix. This represents the row vector of the D1 modal projection eigenma at time step t. express In the row vector at time step t, μ represents the hyperparameter used to suppress the large influence of a single mode; Φ fusion This indicates a learnable bias. This represents the dual-modal fused frame vector at time step t; This is an element-wise product.
[0089] S5. Input the text modal feature representation and the dual-modal fusion frame vector sequence into the global modulation module to obtain the frame-level fusion feature sequence; send the frame-level fusion feature sequence into the prediction head module to obtain the prediction result.
[0090] In some embodiments, in the global modulation module, the bimodal fusion frame vector at each time step is gated using text modal feature representation to obtain frame-level fusion features, represented as:
[0091]
[0092] in, U represents the frame-level fusion feature at time step t. b z represents the learnable parameter. T Let σ represent the text modal feature representation, and let σ() represent the Sigmoid activation function.
[0093] Through global gating factors This avoids frame-level timeline misalignment between text modal and audio / video modal, while also ensuring consistent modulation of text semantics across all frames.
[0094] In some embodiments, the prediction head module includes a main prediction head, a video prediction head, an audio prediction head, and a text prediction head, wherein:
[0095] S51. The frame-level fused feature sequence is averaged and then fed into the main prediction head to obtain the fused prediction result;
[0096] S52. After performing average pooling along the time dimension on the video modality projection feature matrix, output the video modality prediction result through the video prediction head;
[0097] S53. After performing average pooling along the time dimension on the audio modality projection feature matrix, output the audio modality prediction result through the audio prediction head;
[0098] S54. Represent the text modal features through the text prediction head and output the text modal prediction results.
[0099] S6. Calculate the adaptive multimodal collaborative optimization loss based on the prediction results, and iteratively train the model through backpropagation until the model converges.
[0100] In some embodiments, the process of calculating the Weighted Multi-modal Fusion Loss is as follows:
[0101] S61. Calculate the error of each modality (video, audio, text) on the sample, and apply it using learnable weights. Adjust its contribution:
[0102]
[0103]
[0104]
[0105] In the formula, y i This represents the true label of the i-th sample. T' represents the D' mode prediction result of the i-th sample, N represents the total number of training samples; T' is a preset hyperparameter (usually selected between 0.1 and 10) that controls the smoothness of Softmax and is used to adjust the weight difference of different modes in the loss function, balancing the sensitivity of the network to mode preference during training. Representing modes The confidence scalar logit on the i-th sample; and For each modality D', the learnable parameters are defined as follows: for each modality D'∈{A,V,T}, where T represents the text modality, the backbone features are used to... Output an unbiased scalar In order to achieve The purpose is to adaptively change D''∈{A,V,T} for different samples and different training stages.
[0106] For video / audio modalities, the video modal feature representation / Audio Modal Feature Representation Average pooling is performed to obtain / For text modalities, text modal features are used directly. As the input vector for weight calculation .
[0107] S62. Calculate the multimodal cooperative consistency loss:
[0108] Define cross-modal consistency error as the mean squared difference between the three-modal prediction results and the fused prediction result:
[0109]
[0110] Construct a two-part collaborative loss and control the weights using adjustable coefficients λ1 and λ2:
[0111]
[0112] In the formula, This represents the fusion prediction result for the i-th sample. The mean squared difference between the three-modal prediction results and the fusion prediction result for the i-th sample is represented by γ; γ represents a hyperparameter used to control... Regarding the growth rate of the penalty term, when γ=0, the term is a linear penalty term; when γ>0, the term is a sublinear penalty term, and its growth rate varies with... Increases and decreases, with the supremum being λ1 / γ; τ represents a nonnegative scalar, and λ1 and λ2 represent adjustable coefficients used to control logarithmic smoothing; the logarithmic term... It is used to smooth out large errors, making the model less sensitive to outliers in the samples.
[0113] S63. Add the single-modal loss to the multimodal fusion loss to obtain the adaptive multimodal collaborative optimization loss L. total :
[0114]
[0115] In the formula, λ modal λ represents the weight of the hyperparameter single-mode loss. fusion L represents the weights of the hyperparameter multimodal loss. modal L represents the single-mode loss. fusion This represents the multimodal loss.
[0116] In this invention, unless otherwise explicitly specified and limited, the terms "installation," "setting," "connection," "fixing," "rotation," etc., should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral part; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; they can refer to the internal communication of two components or the interaction between two components. Unless otherwise explicitly limited, those skilled in the art can understand the specific meaning of the above terms in this invention according to the specific circumstances.
[0117] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.
Claims
1. A method for predicting short video traffic based on multimodal data, characterized in that, A multimodal short video traffic prediction model is constructed and trained, and the trained multimodal short video traffic prediction model is used for traffic prediction; the multimodal short video traffic prediction model includes a multimodal feature extraction module, a temporal modeling module, a modal attention module, a global modulation module, and a prediction head module; The training process of the multimodal short video traffic prediction model includes the following steps: S1. Collect a short video dataset consisting of multiple sets of samples; each set of samples includes structured video modality data, structured audio modality data, and structured text modality data of the same short video; S2. For each set of samples, extract video modal feature representation, audio modal feature representation, and text modal feature representation through the multimodal feature extraction module; S3. The temporal modeling module is used to process the video modal feature representation and the audio modal feature representation to obtain the temporal dependency representation of the video modality and the temporal dependency representation of the audio modality; S4. A modal attention module is used to process the video modal temporal dependency representation and the audio modal temporal dependency representation to obtain a dual-modal fused frame vector sequence; S5. Input the text modal feature representation and the dual-modal fusion frame vector sequence into the global modulation module to obtain the frame-level fusion feature sequence; send the frame-level fusion feature sequence into the prediction head module to obtain the prediction result; S6. Calculate the adaptive multimodal collaborative optimization loss based on the prediction results, and iteratively train the model through backpropagation until the model converges.
2. The short video traffic prediction method based on multimodal data according to claim 1, characterized in that, Step S1 obtains a set of samples, including: S11. Obtain the video modal data, audio modal data, and text modal data corresponding to a short video; S12. Perform length standardization processing on the video modal data to obtain video modal structured data, including: determining whether the number of frames in the video modal data is less than a preset number L; if so, fill in predefined frames to make the number of frames reach L to form video modal structured data; otherwise, randomly select L frames to form video modal structured data. S13. The audio modal data is length normalized, converted into mono, and then resampled to 16kHz. The resampled data is subjected to short-time Fourier transform with a 25ms Hanning window and a 10ms step size to obtain the amplitude spectrum. The amplitude spectrum is mapped to a 64-dimensional Mel filter bank to obtain the Mel spectrum. The Mel spectrum is logarithmically compressed and dynamically truncated to obtain the structured audio modal data. S14. Normalize the text modal data to Unicode, remove emojis, desensitize URLs and @user tags, unify full-width and half-width characters, and unify uppercase and lowercase. Use a Chinese word segmentation tool in conjunction with a user dictionary to perform maximum positive matching on the processed text to obtain the text modal structured data.
3. The short video traffic prediction method based on multimodal data according to claim 1, characterized in that, Step S2 includes: S21. Use pre-trained I3D-ResNet50 to extract video features from video modal structured data, and use pre-trained VGGish to extract audio features from audio modal structured data; S22. Using pre-trained BGE-Large-zh, word-level matrices and paragraph-level representation vectors are extracted from the text modal structured data. Residual mean pooling is then performed to obtain a static global text representation, represented as follows: , Among them, g T Represents a static global text representation. Represents the word-level matrix E T The element in the i-th row, n represents the row number, s T This represents the paragraph-level representation vector, and LayerNorm() represents layer normalization; S23. Project the video features, audio features, and static global text representation onto the same dimension to obtain video modal feature representation, audio modal feature representation, and text modal feature representation.
4. The short video traffic prediction method based on multimodal data according to claim 1, characterized in that, Step S3 includes: S31. Perform preliminary temporal modeling on the video feature representation and audio feature representation to obtain the video modal comprehensive evolution and audio modal comprehensive evolution at different time steps, expressed as: , in, This represents the D-mode synthesis evolution at time step t. Let D represent the high-dimensional feature vector corresponding to time step t, and f() represent the learnable multi-input nonlinear combination function; D={A,V}, where A is the audio modality and V is the video modality; S32. A nonlinear temporal context aggregator is used to update the video modal synthesis evolution and audio modal synthesis evolution at each time step, resulting in the corresponding time step's video modal temporal dependency representation and audio modal temporal dependency representation, expressed as: , in, This represents the D-mode temporal dependency representation at time step t, where k represents the hyperparameter. This represents the D-mode synthesis evolution of the i-th frame within the sliding window. express First-order difference, σ 2 This represents the standard deviation of the Gaussian kernel hyperparameter.
5. The short video traffic prediction method based on multimodal data according to claim 1, characterized in that, Step S4 includes: S41. Calculate the video modal self-attention weights and audio modal self-attention weights, expressed as follows: , , , , , , in, This represents the D-modal temporal dependency representation at time step t, where D = {A, V}, A is the audio modality, V is the video modality, L represents the sequence length, and H represents the sequence length. D The characteristic matrix Q represents the D-mode. D K represents the query matrix of mode D. D V represents the bond matrix of the D modes. D The value matrix representing the D modes; , , For trainable weight matrix, Indicates self-attention bias, d k S represents the dimension of the key matrix. D Let α represent the self-attention scoring matrix, Softmax() represent the activation function, and α represent the self-attention scoring matrix. D Represents the self-attention weights of the D-mode; S42. Calculate the video modal update features and audio modal update features, represented as follows: , in, Denotes the D-modal update features, LayerNorm() denotes layer normalization, ε D ρ represents the regularization term. D Indicates the learnable bias term; S43. Project the video modality update features and audio modality update features to obtain the video modality projection feature matrix and the audio modality projection feature matrix; S44. Calculate the cross-modal mutual attention coefficients for the video-to-audio and audio-to-video modal directions, expressed as: , , , in, The query matrix representing the D1 mode. The bond matrix representing the D2 mode. This represents the characteristic matrix of the D1 modal projection. D1 represents the characteristic matrix of the D2 modal projection, where D1 = {A, V}; , This represents the trainable directional projection matrix from mode D1 to mode D2. Indicates the bias term. Let (D1,D2) denote the cross-modal mutual attention coefficients from mode D1 to mode D2, where (D1,D2)∈{(A,V), (V,A)}; S45. Calculate the bimodal fused frame vector at each time step to form a bimodal fused frame vector sequence, represented as: , , in, This represents the cross-modal weighting matrix from mode D1 to mode D2. The characteristic matrix representing the D2 mode, This represents a trainable weight matrix. This represents the row vector of the D1 modal projection eigenma at time step t. express In the row vector at time step t, μ represents the hyperparameter, Φ fusion This indicates a learnable bias. This represents the dual-modal fused frame vector at time step t; This is an element-wise product.
6. The short video traffic prediction method based on multimodal data according to claim 1, characterized in that, In the global modulation module, the text modal feature representation is used to perform gated modulation on the dual-modal fused frame vector at each time step to obtain the frame-level fusion feature, which is represented as: , in, Let z represent the frame-level fusion features at time step t, where U and b represent learnable parameters. T Let σ represent the text modal feature representation, and let σ() represent the Sigmoid activation function.
7. The short video traffic prediction method based on multimodal data according to claim 1, characterized in that, The prediction head module includes a main prediction head, a video prediction head, an audio prediction head, and a text prediction head, among which: S51. The frame-level fused feature sequence is averaged and then fed into the main prediction head to obtain the fused prediction result; S52. After performing average pooling along the time dimension on the video modality projection feature matrix, output the video modality prediction result through the video prediction head; S53. After performing average pooling along the time dimension on the audio modality projection feature matrix, output the audio modality prediction result through the audio prediction head; S54. Represent the text modal features through the text prediction head and output the text modal prediction results.
8. The short video traffic prediction method based on multimodal data according to claim 1, characterized in that, The adaptive multimodal collaborative optimization loss L is calculated based on the prediction results. total , represented as , In the formula, λ modal λ represents the weight of the hyperparameter single-mode loss. fusion L represents the weights of the hyperparameter multimodal loss. modal L represents the single-mode loss. fusion Denotes the multimodal loss, where: , , In the formula, y i This represents the true label of the i-th sample. Let N represent the D' mode prediction result of the i-th sample, and N represent the total number of training samples; Represents the learnable weights of the D' mode; This represents the fusion prediction result for the i-th sample. Let represent the mean squared difference between the three-modal prediction results and the fusion prediction results for the i-th sample, γ represent the hyperparameter, τ represent the nonnegative scalar, and λ1 and λ2 represent adjustable coefficients; D'∈{A,V,T}, where T represents the text modality; where: , In the formula, Representing modes The confidence scalar on the i-th sample, T' represents the preset hyperparameter, and D''∈{A,V,T}.
Citation Information
Cited By
Multi-modal data multi-model combined training method and system
CN121144858A
A method and system for training multiple models combined with multimodal data
CN121144858B
Video analysis report generation method and system
CN121665033A