Consumer participation prediction method and system fused with goods-carrying short video multi-modal information
Through a fully connected network and a cross-modal cross-attention mechanism, the text, audio and visual characteristics of short videos are integrated, and combined with the Transformer encoder, the problem of underutilization of multimodal information in short videos is solved, and the prediction accuracy and robustness of consumer participation behavior is improved.
Patent Information
- Application Number
- CN202510579869.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-07
- Publication Date
- 2025-08-01
AI Technical Summary
The prior art fails to fully utilize multimodal information features and interactions between modals in short video sales scenarios, resulting in inaccurate prediction of consumer participation behavior and insufficient application of machine learning methods.
The text, audio and visual modal feature matrix are mapped to the shared space through a fully connected network, and the modal cross-attention mechanism is used to calculate the modal information fusion vector, and feature representation and prediction are combined with a four-layer Transformer encoder.
Improve the prediction accuracy of consumer participation behaviors (such as likes, comments, and retweets), and enhance the robustness of predictions and the consideration of modal interactions.
Smart Images

Figure CN120407852A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of consumer behavior prediction, and in particular to a method and system for predicting consumer participation by integrating multimodal information of short videos promoting products. Background Art
[0002] With the rapid development of technologies such as mobile internet and 5G, short video live streaming, as an emerging marketing method, has rapidly emerged and garnered widespread attention due to its unique appeal. Through short videos lasting 15 seconds to one minute, using a trimodal approach of text, audio, and video, short video live streaming quickly delivers product information to potential consumers, thereby quickly capturing their attention. While watching these live streaming videos, consumers can interact with the creators and other consumers through various means, such as liking, commenting, and forwarding. This engagement not only increases video exposure but also strengthens social recognition and brand loyalty, promotes the widespread dissemination of user-generated content, and ultimately increases revenue for the creators. Therefore, a deep understanding and accurate prediction of consumer engagement behavior is crucial for enhancing the marketing effectiveness of short video live streaming.
[0003] However, in the short video sales scenario, although the multimodal information of the short video content, such as text, auditory and visual, has an impact on the viewers at the same time, the existing inventions still have significant shortcomings in fully utilizing the internal characteristics of these multimodal information and the interaction between modalities to predict consumer participation behavior.
[0004] Secondly, while extensive research has explored how single-modal information independently influences and predicts consumer engagement (such as likes, comments, and reposts), the fact that the fusion of multimodal information features holds even greater potential in stimulating active consumer interaction is crucial. In the innovative context of short video e-commerce, the close integration and simultaneous presentation of text, voice, and visual modalities not only significantly enriches the dimensionality of information but also, through their potential interactions, creates a highly engaging information environment. This multimodal interaction is not simply a superposition, but rather a complex process of mutual reinforcement and complementation, making information more vivid and multifaceted, thereby more effectively reaching and influencing consumer engagement decisions. However, current literature remains insufficient in exploring the impact of multimodal information fusion on consumer engagement.
[0005] Furthermore, while a few studies have examined the impact of multimodal information combinations on consumer engagement, most have primarily employed empirical empirical testing using quantitative analysis and experimental designs, with relatively few predictive studies based on machine learning and deep learning methods. This, to a certain extent, limits the potential of multimodal information in predicting consumer engagement behavior. Summary of the Invention
[0006] In view of the above situation, the main object of the present invention is to propose a consumer participation prediction method and system that integrates multi-modal information of short live-streaming videos for goods sales, so as to solve the above technical problems.
[0007] The present invention proposes a consumer participation prediction method that integrates multi-modal information of short live-streaming videos for goods sales, and the method includes the following steps: Step 1: Obtain a text modality feature matrix, an audio modality feature matrix, and a visual modality feature matrix by using short live-streaming videos for goods sales; Step 2: Map the text modality feature matrix, the audio modality feature matrix, and the visual modality feature matrix to a shared space of the same target dimension through a fully connected network, and respectively obtain a text feature representation after linear transformation, an audio feature representation after linear transformation, and a visual feature representation after linear transformation; Based on the text feature representation after linear transformation, the audio feature representation after linear transformation, and the visual feature representation after linear transformation, calculate through a cross-modal cross-attention mechanism to obtain a modality information fusion vector representation; Concatenate the modality information fusion vector representations to obtain a concatenated vector representation; Step 3: Process the concatenated vector representation through a four-layer Transformer encoder to obtain a feature vector representation output by the encoder; Predict the feature vector representation output by the encoder to obtain a prediction result.
[0008] The present invention also proposes a consumer participation prediction system that integrates multi-modal information of short live-streaming videos for goods sales, and the system includes: A modality information extraction module, used for: Obtaining a text modality feature matrix, an audio modality feature matrix, and a visual modality feature matrix by using short live-streaming videos for goods sales; A cross-modal cross-attention module, used for: Mapping the text modality feature matrix, the audio modality feature matrix, and the visual modality feature matrix to a shared space of the same target dimension through a fully connected network, and respectively obtaining a text feature representation after linear transformation, an audio feature representation after linear transformation, and a visual feature representation after linear transformation; Based on the text feature representation after linear transformation, the audio feature representation after linear transformation, and the visual feature representation after linear transformation, calculate through a cross-modal cross-attention mechanism to obtain a modality information fusion vector representation; Concatenating the modality information fusion vector representations to obtain a concatenated vector representation; A prediction result module, used for: Predicting the feature vector representation output by the encoder to obtain a prediction result.
[0009] Compared with the prior art, the beneficial effects of the present invention are as follows: 1. The present invention makes full use of the text, audio, and visual information in short videos through a deep learning model to improve the accuracy of predicting consumer engagement behaviors (such as the number of likes, comments, and forwards); 2. The present invention not only focuses on the local and global features of audio modal information but also considers the interactions within and between modalities in order to achieve the best prediction performance and robustness.
[0010] Additional aspects and advantages of the present invention will be given in part in the following description, will become apparent in part from the following description, or will be understood through the embodiments of the present invention. Brief Description of the Drawings
[0011] Figure 1 is a flowchart of the consumer engagement prediction method that fuses multi-modal information of short videos for goods promotion proposed by the present invention; Figure 2 is a flowchart of audio feature vector extraction of the consumer engagement prediction method that fuses multi-modal information of short videos for goods promotion proposed by the present invention; Figure 3 is a line graph comparing the prediction performances of different prediction methods under different modalities; Figure 4 is a schematic diagram of the overall framework of the consumer engagement prediction system that fuses multi-modal information of short videos for goods promotion proposed by the present invention. Detailed Embodiments
[0012] The embodiments of the present invention will be described in detail below. Examples of the embodiments are shown in the drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below by referring to the drawings are exemplary only for explaining the present invention and should not be construed as limiting the present invention.
[0013] Referring to the following description and drawings, these and other aspects of the embodiments of the present invention will be clear. In these descriptions and drawings, some specific embodiments of the embodiments of the present invention are specifically disclosed to represent some ways of implementing the principles of the embodiments of the present invention, but it should be understood that the scope of the embodiments of the present invention is not limited thereto.
[0014] Please refer to Figure 1 , the embodiments of the present invention propose a consumer engagement prediction method that fuses multi-modal information of short videos for goods promotion. The method includes the following steps: Step 1. Obtain a text modal feature matrix, an audio modal feature matrix, and a visual modal feature matrix by using short videos for goods promotion; In step 1, a text modality feature matrix, an audio modality feature matrix, and a visual modality feature matrix are obtained from the short video with goods promotion. The specific steps are as follows: The ELECTRA_CHINESE model is used to process the short video with goods promotion to obtain text feature vectors, and a text modality feature matrix is constructed based on the text feature vectors; An open-source tool for audio signal processing is used to process the short video with goods promotion to obtain audio feature vectors, and an audio modality feature matrix is constructed based on the audio feature vectors; A residual network is used to process the short video with goods promotion to obtain visual feature vectors, and a visual modality feature matrix is constructed based on the visual feature vectors.
[0015] Specifically, in this step, a Douyin short video URL is generated according to the unique identification code of the short video, and the source file of the short video with goods promotion is obtained through a crawler program. After data cleaning to remove incomplete information and invalid or unplayable short video files, a dataset of short videos with goods promotion with complete information is finally constructed. Specifically, the text modality information includes the title text and product description text of the short video with goods promotion, the audio modality information is the wav audio of the entire video, and the visual modality information includes the video cover image and key frame images of the video.
[0016] Further, please refer to Figure 2 , in this step, in the process of obtaining the text modality feature matrix, ELECTRA_CHINESE is selected to embed the text feature vectors, and the information of the last hidden layer is extracted as the representation of the text feature vectors. The ELECTRA_CHINESE model is a pre-trained model jointly developed by Harbin Institute of Technology and iFlytek based on the official ELECTRA training code and a large amount of Chinese data. Compared with the common Bert model, this model can capture richer context semantic information.
[0017] In the process of obtaining the audio modality feature matrix, the basic acoustic features of the audio, such as fundamental frequency, energy, pitch, etc., 64-dimensional basic acoustic features are extracted using the ComParE_2016 feature set of openSMILE. This method takes frames as the extraction unit, that is, frame-level acoustic features are extracted. The frame-level acoustic features are calculated within a short time window, which can provide local information of the audio and reflect the short-term changes of the audio over time.
[0018] In the process of obtaining the visual modality feature matrix, ten image frames are evenly extracted as video key frames according to the total duration of the video. Next, the video cover image and the ten key frame images are resized to the same size (224*224) and normalized. Then, ResNet-18 is used to extract image features, and the last fully connected layer is selected as the feature output. Finally, an aggregation operation is performed on the extracted image features to form a visual feature representation of the video, and a visual feature vector is finally output.
[0019] Step 2: Map the text modality feature matrix, the audio modality feature matrix, and the visual modality feature matrix to a shared space of the same target dimension through a fully connected network, and respectively obtain the text feature representation after linear transformation, the audio feature representation after linear transformation, and the visual feature representation after linear transformation; Based on the text feature representation after linear transformation, the audio feature representation after linear transformation, and the visual feature representation after linear transformation, through cross-modal cross-attention mechanism calculation, obtain the modal information fusion vector representation; Concatenate the modal information fusion vector representations to obtain a concatenated vector representation; In Step 2, map the text modality feature matrix, the audio modality feature matrix, and the visual modality feature matrix to a shared space of the same target dimension through a fully connected network, and respectively obtain the text feature representation after linear transformation, the audio feature representation after linear transformation, and the visual feature representation after linear transformation. The relational expressions existing in the corresponding process are: ; Among them, represents the modal feature representation after linear transformation, represents the modal feature matrix, represents the weight matrix, represents the bias term, represents the modal type, represents the text modality, represents the audio modality, represents the visual modality; Based on the text feature representation after linear transformation, the audio feature representation after linear transformation, and the visual feature representation after linear transformation, through cross-modal cross-attention mechanism calculation, obtain the modal information fusion vector representation. The relational expressions existing in the corresponding process are: ; Among them, represents the cross-attention representation of the modal , represents the cross-attention representation of the modal , represents the modal type, Indicates being processed by the multi-head self-attention mechanism, Indicates the modality of the query vector, Indicates the modality of the key vector, Indicates the modality of the value vector, Indicates the modality of the query vector, Indicates the modality of the key vector, Indicates the modality of the value vector, Indicates the modality information fusion vector representation, Indicates being calculated and processed by the cross-modal cross-attention mechanism, Indicates the modality after linear transformation Indicates being processed by the Transformer layer, Indicates being processed by concatenation; Concatenate the modality information fusion vector representations to obtain the concatenated vector representation. The relational expression for the corresponding process is: ; Among them, Indicates the concatenation matrix, Indicates the text modality-audio modality concatenated vector representation, Indicates the text modality-audio modality concatenated vector representation, Indicates the text modality-audio modality concatenated vector representation.
[0020] Furthermore, in this step, learn the interaction of multi-modal information in the live-streaming short video through the cross-modal cross-attention mechanism.
[0021] Step 3: Process the concatenated vector representation through four layers of Transformer encoders to obtain the feature vector representation output by the encoder; Predict the feature vector representation output by the encoder to obtain the prediction result; In Step 3, process the concatenated vector representation through four layers of Transformer encoders to obtain the feature vector representation output by the encoder. The relational expression for the corresponding process is: ; Among them, Indicates the feature vector representation output by the encoder, Indicates being processed by four layers of Transformer encoders, Indicates being processed by the first layer of Transformer encoder, Indicates processing by the second-layer Transformer encoder, Indicates processing by the third-layer Transformer encoder, Indicates processing by the fourth-layer Transformer encoder; Predict the feature vector representation output by the encoder to obtain a prediction result. The relational expression existing in the corresponding process is: ; Among them, represents the weight of the output layer, represents the bias term of the output layer.
[0022] Furthermore, in this step, the Transformer encoder includes a self-attention mechanism (Self-Attention, SA), a multi-head self-attention mechanism (Multi-Head Self-Attention, MHSA), and a feed-forward network (Feed-Forward Network, FNN). Each mechanism and network has a residual connection and a normalization layer (Add&Norm). Each layer of the Transformer encoder is an 8-head attention mechanism. Map the concatenated vector representation containing multimodal features to a shared space through a fully connected layer to form a unified feature representation.
[0023] Please refer to Figure 4 , this embodiment of the present invention also provides a consumer participation prediction system that fuses multimodal information of live-streaming short videos. The system includes: A modal information extraction module for: Obtain a text modal feature matrix, an audio modal feature matrix, and a visual modal feature matrix using the live-streaming short video; A cross-modal cross-attention module for: Map the text modal feature matrix, the audio modal feature matrix, and the visual modal feature matrix to a shared space of the same target dimension through a fully connected network, and respectively obtain a linearly transformed text feature representation, a linearly transformed audio feature representation, and a linearly transformed visual feature representation; Based on the linearly transformed text feature representation, the linearly transformed audio feature representation, and the linearly transformed visual feature representation, calculate through a cross-modal cross-attention mechanism to obtain a modal information fusion vector representation; Concatenate the modal information fusion vector representations to obtain a concatenated vector representation; A prediction result module for: Predict the feature vector representation output by the encoder to obtain a prediction result.
[0024] To verify the effectiveness of short-video multimodal information features in predicting consumer participation, the performance of the present invention is compared with that of a baseline model with the same order of parameter quantity, as shown in Table 1 and Figure 3 It can be seen that the present invention has certain advantages in capturing real data sets from short videos.
[0025] Table 1 Comparison of Results of Different Prediction Methods
[0026] At the same time, the cross-modal cross-attention connection mechanism used in the present invention is compared with the feature-level fusion and decision-level fusion of general methods in terms of prediction performance. It can be seen from Table 2 that the cross-modal cross-attention connection mechanism used in the present invention is more suitable for the downstream task of predicting consumer participation in short videos for goods promotion.
[0027] Table 2 Comparison of Prediction Results of Different Fusion Methods
[0028] It should be understood that each part of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following well-known technologies in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application-specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.
[0029] In the description of this specification, the description with reference to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.
[0030] The above-described embodiments only represent several implementation manners of the present invention, and their descriptions are relatively specific and detailed, but should not be construed as limiting the scope of the patent of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the patent of the present invention should be subject to the appended claims.
Claims
1. A method for predicting consumer participation by integrating multi-modal information of short live-streaming videos for product promotion, characterized in that The method includes the following steps: Step 1: Obtain a text modality feature matrix, an audio modality feature matrix, and a visual modality feature matrix using a short video for promoting goods; Step 2: Map the text modality feature matrix, the audio modality feature matrix, and the visual modality feature matrix to a shared space of the same target dimension through a fully connected network, and respectively obtain a text feature representation after linear transformation, an audio feature representation after linear transformation, and a visual feature representation after linear transformation; Based on the text feature representation after linear transformation, the audio feature representation after linear transformation, and the visual feature representation after linear transformation, calculate through a cross-modal cross-attention mechanism to obtain a modality information fusion vector representation; Concatenate the modality information fusion vector representations to obtain a concatenated vector representation; Step 3: Process the concatenated vector representation through four layers of Transformer encoders to obtain a feature vector representation output by the encoder; Predict the feature vector representation output by the encoder to obtain a prediction result.
2. The consumer participation prediction method for integrating multi-modal information of short videos with goods promotion according to claim 1, wherein In the said Step 1, the specific steps for obtaining a text modality feature matrix, an audio modality feature matrix, and a visual modality feature matrix using a short video for promoting goods are as follows: Process the short video for promoting goods using the ELECTRA_CHINESE model to obtain text feature vectors, and construct a text modality feature matrix based on the text feature vectors; Process the short video for promoting goods using an open-source tool for audio signal processing to obtain audio feature vectors, and construct an audio modality feature matrix based on the audio feature vectors; Process the short video for promoting goods using a residual network to obtain visual feature vectors, and construct a visual modality feature matrix based on the visual feature vectors.
3. The consumer participation prediction method for integrating short video multi-modal information with goods promotion according to claim 2, wherein, In the said Step 2, when mapping the text modality feature matrix, the audio modality feature matrix, and the visual modality feature matrix to a shared space of the same target dimension through a fully connected network, and respectively obtaining a text feature representation after linear transformation, an audio feature representation after linear transformation, and a visual feature representation after linear transformation, the relational expressions existing in the corresponding process are: ; Among them, represents the modality after linear transformation feature representation, represents the modality feature matrix, represents the weight matrix, represents the bias term, represents the modality type, represents the text modality, represents the audio modality, represents the visual modality.
4. The consumer participation prediction method for fusing multi-modal information of short live-streaming videos with goods sales according to claim 3, wherein, In the said Step 2, when calculating through a cross-modal cross-attention mechanism based on the text feature representation after linear transformation, the audio feature representation after linear transformation, and the visual feature representation after linear transformation to obtain a modality information fusion vector representation, the relational expressions existing in the corresponding process are: ; Among them, represents the cross-attention representation of modality ; represents the cross-attention representation of modality ; represents the modality type, represents being processed by the multi-head self-attention mechanism, represents the query vector of modality ; represents the key vector of modality ; represents the value vector of modality ; represents the query vector of modality ; represents the key vector of modality ; represents the value vector of modality ; represents the modality information fusion vector representation, represents being calculated and processed by the cross-modal cross-attention mechanism, represents the modality feature representation after linear transformation, represents being processed by the Transformer layer, represents being processed by concatenation.
5. The consumer participation prediction method for fusing multi-modal information of short live-streaming videos with goods sales according to claim 4, wherein, In the said Step 2, when concatenating the modality information fusion vector representations to obtain a concatenated vector representation, the relational expressions existing in the corresponding process are: ; Among them, represents a splicing matrix, represents the text modality-audio modality splicing vector representation, represents the text modality-audio modality splicing vector representation, represents the text modality-audio modality splicing vector representation.
6. The consumer participation prediction method for integrating multi-modal information of short videos with goods promotion according to claim 5, characterized in that In the said Step 3, when processing the concatenated vector representation through four layers of Transformer encoders to obtain a feature vector representation output by the encoder, the relational expressions existing in the corresponding process are: ; Among them, represents the feature vector representation output by the encoder, represents being processed by a four-layer Transformer encoder, represents being processed by the first-layer Transformer encoder, represents being processed by the second-layer Transformer encoder, represents being processed by the third-layer Transformer encoder, represents being processed by the fourth-layer Transformer encoder.
7. The consumer participation prediction method for fusing multi-modal information of short live-streaming videos with goods sales according to claim 6, wherein In the said Step 3, when predicting the feature vector representation output by the encoder to obtain a prediction result, the relational expressions existing in the corresponding process are: ; Among them, represents the weight of the output layer, represents the bias term of the output layer.
8. A consumer participation prediction system that integrates multi-modal information of short live-streaming videos for sales promotion, characterized in that, The system applies the consumer participation prediction method for fusing multi-modal information of short videos for promoting goods according to any one of claims 1 to 7. The system includes: A modality information extraction module, used for: Obtaining a text modality feature matrix, an audio modality feature matrix, and a visual modality feature matrix using a short video for promoting goods; A cross-modal cross-attention module, used for: The text modality feature matrix, audio modality feature matrix, and visual modality feature matrix are mapped to a shared space of the same target dimension through a fully connected network to obtain the linearly transformed text feature representation, linearly transformed audio feature representation, and linearly transformed visual feature representation respectively; Based on the linearly transformed text feature representation, linearly transformed audio feature representation, and linearly transformed visual feature representation, through cross-modal cross-attention mechanism calculation, a modal information fusion vector representation is obtained; The modal information fusion vector representations are concatenated to obtain a concatenated vector representation; A prediction result module, which is used for: Predicting the feature vector representation output by the encoder to obtain a prediction result.