Marketing activity index prediction model establishment method
Through multimodal feature extraction and cross-modal information fusion, combined with static and dynamic prediction modules, the problem of insufficient multimodal data fusion in marketing activity indicator prediction is solved, accurate prediction of marketing activities is achieved, and the adaptability and robustness of the model are enhanced.
Patent Information
- Application Number
- CN202511178621.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-22
- Publication Date
- 2025-09-19
- Estimated Expiration
- 2045-08-22
AI Technical Summary
Existing marketing activity indicator prediction methods lack an effective cross-modal semantic alignment mechanism when processing multimodal data, and are unable to establish deep correlations between different modalities, resulting in the inability to fully utilize the complementary information between text, images and structured features. At the same time, they lack multi-scale temporal modeling capabilities, affecting the accuracy of the prediction model.
A multimodal feature extraction module is used for encoding through the BERT and ResNet-50 models, a dynamic cross-modal information fusion mechanism is designed, and a four-channel bidirectional attention mechanism is constructed to realize the interactive fusion of feature vectors of different modalities. Static and dynamic prediction modules are combined, and a multi-scale sliding window is used to extract time series features. Adaptive fusion is performed through a dynamic weight matrix, and the model is optimized using historical data.
It achieves deep integration of multimodal data, improves the accuracy and adaptability of marketing activity indicator predictions, can process static and dynamic features at the same time, enhances the adaptability to complex marketing environments, and improves the comprehensiveness and robustness of predictions.
Smart Images

Figure CN120672381A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of large models, and in particular relates to a method for establishing a marketing activity indicator prediction model. Background Art
[0002] In the field of intelligent marketing forecasting, traditional marketing activity indicator prediction methods primarily rely on single-modal data analysis techniques, such as natural language processing methods based solely on text descriptions, computer vision techniques based solely on image content, or statistical analysis models based solely on structured data. When dealing with comprehensive marketing scenarios that include text descriptions, promotional images, and feature parameters, these methods often adopt a strategy of simple splicing after independent processing. However, existing multimodal fusion technologies have significant drawbacks, primarily manifested in the lack of effective cross-modal semantic alignment mechanisms, an inability to establish deep correlations between different modalities, and an inability to fully utilize the complementary information between text, image, and structured features. Furthermore, existing methods lack multi-scale temporal modeling capabilities when dealing with temporal dynamic features. In current marketing activity forecasting applications, due to the multimodal nature of marketing data, which includes rich text descriptions, diverse visual content, and complex temporal features, existing technologies struggle to achieve deep fusion and collaborative modeling of information from each modality. This results in the inability of prediction models to accurately capture the comprehensive feature representation of marketing activities, which in turn affects the accuracy of indicator predictions. Summary of the Invention
[0003] In view of this, the present invention provides a method for establishing a marketing activity indicator prediction model, which can solve the technical problem in the prior art that insufficient fusion of multimodal data features of marketing activities leads to low prediction accuracy.
[0004] The present invention is implemented as follows: the present invention provides a method for establishing a marketing activity indicator prediction model to collect multimodal data of marketing activities, including marketing activity text descriptions, marketing activity promotional images and marketing activity feature parameters; construct a multimodal feature extraction module, encode the marketing activity text description through the BERT model to obtain a text feature vector, encode the marketing activity promotional image through the ResNet-50 model to obtain a visual feature vector, and encode the marketing activity feature parameters through a structured encoder to obtain an activity feature vector; design a cross-modal information dynamic fusion mechanism, construct a four-channel bidirectional attention mechanism to achieve interactive fusion between different modal feature vectors, and use the maximum weight bipartite graph matching algorithm to combine different modal feature vectors. Vectors are constructed as bipartite graph nodes and the optimal cross-modal feature pairing weights are solved to generate a multimodal fusion representation; a static prediction module is established, and the multimodal fusion representation is input into the stacked residual multi-layer perceptron network to realize the regression prediction of the total marketing activity indicator and obtain the static prediction result; a dynamic prediction module is constructed, and a multi-scale sliding window is used to extract time series features. The dynamic prediction of the marketing activity indicator sequence is realized through the dual-branch collaborative mechanism of the time series evolution branch and the cross-modal interaction branch; a dynamic fusion output layer is designed, and a dynamic weight matrix is generated through cross-modal attention. The output of the time series evolution branch and the output of the cross-modal interaction branch are adaptively fused to obtain the dynamic prediction result; historical marketing activity data are used to train and optimize the marketing activity indicator prediction model.
[0005] Among them, the marketing activity characteristic parameters specifically include time characteristics and activity strategy characteristics. The time characteristics use discrete coding to capture the natural time sequence patterns of weeks, months and holidays. The activity strategy characteristics use structured representation to characterize the marketing strategy attributes of activity budget intensity, activity type and duration cycle.
[0006] Among them, the dynamic cross-modal information fusion mechanism specifically achieves fine-grained cross-modal semantic alignment and information enhancement through differentiated query key-value pair configurations. The four independent attention channels in the four-channel bidirectional attention mechanism respectively model the asymmetric dependency relationship between different modalities. The text feature vector is used as the query vector to actively align the visual content to ensure the semantic consistency between the activity description and the image theme.
[0007] Among them, the four-channel bidirectional attention mechanism specifically realizes the interactive fusion between text feature vectors and visual feature vectors, activity feature vectors and visual feature vectors, activity feature vectors and text feature vectors, and text feature vectors and activity feature vectors. First, query key-value pairs are generated through independent projection matrices, and then the attention distribution of feature vectors to other modal contents is calculated through scaled dot product attention.
[0008] Among them, the static prediction module specifically realizes the regression prediction of the total index of the marketing activity through a processing chain composed of a batch normalization layer, a self-attention layer and a residual connection. The stacked residual multi-layer perceptron network includes a feature transformation path and a residual direct connection path. The feature transformation path captures nonlinear relationships through a progressive mapping from 512 dimensions to 384 dimensions to 256 dimensions, and the residual direct connection path retains the original information to prevent the gradient from disappearing.
[0009] Among them, the multi-scale sliding window specifically uses a parallel multi-scale convolution kernel group to extract time series patterns of different granularities, using a 3-day short kernel to capture sudden promotion fluctuations, a 7-day medium kernel to analyze periodic consumption patterns, and a 30-day long kernel to identify trend billing patterns.
[0010] Among them, the temporal evolution branch is specifically based on the sequential pattern capture of the causal Transformer, adopts the causal mask mechanism to constrain the attention range, ensures that the prediction time step only accesses historical information, and generates temporal representation by aggregating sequence features and learnable position encoding.
[0011] Among them, the cross-modal interaction branch specifically realizes feature enhancement through the interaction between the temporal global context and the multimodal fusion representation, generates a query vector through pooling operation, generates a key-value pair through the multimodal fusion representation, and outputs the interaction feature for subsequent fusion processing.
[0012] Among them, the dynamic weight matrix specifically reflects the intensity of attention to the time series pattern at different prediction steps, realizes the dynamic adjustment of the modal importance in the prediction scenario, and generates dynamic weights through the interactive calculation of the time series branch projection features and the cross-modal interaction features.
[0013] Among them, the visual feature vector encoding specifically uses a deep residual network to extract fine-grained spatial features, extracts abstract features step by step through the convolutional layer group of ResNet-50, and obtains the visual feature vector through spatial attention pooling and linear projection.
[0014] The text feature vector encoding is specifically a dynamic semantic extractor based on a pre-trained language model, which uses the BERT model to generate context-aware semantic vectors and performs encoding processing through the text projection matrix and the marketing activity text description input.
[0015] Among them, the environmental perception gating mechanism specifically splices the output features of the four independent attention channels in the four-channel bidirectional attention mechanism with the original activity feature vector, generates the adaptive fusion weights of the four channels through a learnable parameter matrix, and the final multimodal fusion representation is generated through weighted combination.
[0016] Among them, the multi-scale feature extraction formula is Where i represents the index of convolution kernels of different scales, To input time series data, the time series representation generation formula is in Aggregate sequence features A learnable positional encoding.
[0017] Among them, the query vector generation formula is: The key-value pair generation formula is and , the static prediction result calculation formula is: , the dynamic weight matrix generation formula is: .
[0018] Among them, the visual feature vector encoding formula is: in is the trainable projection matrix is the bias term, and the text feature vector encoding formula is in is the text projection matrix Enter a text description for the campaign.
[0019] The game model includes an upper model with the goal of maximizing marketing effects and a lower model with the goal of optimizing resource allocation. The objective function of the upper model is , the objective function of the lower model is .
[0020] The present invention realizes the deep interactive fusion of multimodal features by constructing a four-channel bidirectional attention mechanism, and designs a cross-modal information dynamic fusion mechanism to establish the semantic alignment relationship between text, visual and activity features, which effectively solves the technical defect of insufficient multimodal data fusion in traditional methods. The present invention adopts a multi-scale sliding window to extract time series features of different granularities, and realizes the organic combination of static features and dynamic features through the dual-branch collaborative mechanism of time series evolution branch and cross-modal interaction branch, overcomes the shortcomings of the existing technology in time series modeling, and significantly improves the ability to capture complex spatiotemporal patterns of marketing activities. In summary, the present invention fundamentally solves the technical problem of low prediction accuracy caused by insufficient fusion of multimodal data features of marketing activities in the existing technology by establishing a complete multimodal fusion prediction framework. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 is a flow chart of the method of the present invention.
[0022] Figure 2 Schematic diagram of the structure of the model established for the present invention.
[0023] Figure 3 Schematic diagram of the residual multi-layer perceptron block.
[0024] Figure 4 This is the structure diagram of the static prediction module.
[0025] Figure 5 This is the structural diagram of the dynamic prediction module. DETAILED DESCRIPTION
[0026] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention.
[0027] like Figure 1 FIG. 1 is a flow chart of a method for establishing a marketing activity indicator prediction model provided by the present invention. The method includes the following steps: S01. Collect multimodal data of a marketing activity, wherein the multimodal data includes a text description of the marketing activity, a promotional image of the marketing activity, and characteristic parameters of the marketing activity, wherein the characteristic parameters of the marketing activity include time characteristics and activity strategy characteristics; S02. Construct a multimodal feature extraction module, encode the marketing campaign text description using the BERT model to obtain a text feature vector, encode the marketing campaign promotional image using the ResNet-50 model to obtain a visual feature vector, and encode the marketing campaign feature parameters using a structured encoder to obtain an activity feature vector; S03. Design a dynamic cross-modal information fusion mechanism and construct a four-way bidirectional attention mechanism to achieve interactive fusion between text feature vectors and visual feature vectors, activity feature vectors and visual feature vectors, activity feature vectors and text feature vectors, and text feature vectors and activity feature vectors. Use the maximum weighted bipartite graph matching algorithm to construct different modal feature vectors into bipartite graph nodes and solve the optimal cross-modal feature pairing weights to generate a multimodal fusion representation. S04. Establish a static prediction module, input the multimodal fusion representation into a stacked residual multilayer perceptron network, and implement regression prediction of the total marketing activity indicator through a processing chain consisting of a batch normalization layer, a self-attention layer, and a residual connection to obtain a static prediction result; S05. Build a dynamic prediction module, use a multi-scale sliding window to extract time series features, and achieve dynamic prediction of marketing activity indicator sequences through a dual-branch collaborative mechanism of time series evolution branch and cross-modal interaction branch; S06. Design a dynamic fusion output layer, generate a dynamic weight matrix through cross-modal attention, adaptively fuse the temporal evolution branch output and the cross-modal interaction branch output, and obtain dynamic prediction results; S07. Use historical marketing activity data to train and optimize the marketing activity indicator prediction model, balance the dual goals of maximizing marketing effects and optimizing resource allocation through a game model, use a backpropagation algorithm to update model parameters, and use the static prediction results and the dynamic prediction results to achieve accurate prediction of key indicators of marketing activities.
[0028] The structure of the model constructed by the present invention is as follows Figures 2 to 5 shown.
[0029] The cross-modal information dynamic fusion mechanism achieves fine-grained cross-modal semantic alignment and information enhancement through differentiated query key-value pair configurations. The four independent attention channels in the four-channel bidirectional attention mechanism respectively model the asymmetric dependency relationships between different modalities. The text feature vector is used as a query vector to actively align visual content to ensure semantic consistency between the activity description and the image theme. The activity feature vector is used as a query to drive attention to key areas of the image so that parameter settings can locate the corresponding expressions in the image. The activity feature vector actively verifies the compatibility of the description text, and the text feature vector actively queries related parameter information. The multi-scale sliding window uses a parallel multi-scale convolution kernel group to extract time series patterns of different granularities. The 3-day short kernel is used to capture sudden promotion fluctuations, the 7-day medium kernel is used to analyze periodic consumption patterns, and the 30-day long kernel is used to identify trend billing patterns. The multi-scale feature extraction formula is: Where i represents the index of convolution kernels of different scales, The temporal evolution branch is based on the sequential pattern capture of the causal transformer and uses the causal mask mechanism to constrain the attention range to ensure that the prediction time step only accesses historical information. The temporal representation generation formula is in Aggregate sequence features is a learnable position code. The cross-modal interaction branch achieves feature enhancement through the interaction of temporal global context and multimodal fusion representation. The query vector generation formula is: The key-value pair generation formula is and ,in The stacked residual multilayer perceptron network contains a feature transformation path and a residual direct connection path. The feature transformation path captures nonlinear relationships through a progressive mapping from 512 dimensions to 384 dimensions to 256 dimensions. The residual direct connection path retains the original information to prevent gradient disappearance. The static prediction result calculation formula is: in represents the Lth residual multilayer perceptron block, The dynamic weight matrix elements reflect the attention intensity of different prediction steps on the temporal pattern, and realize the dynamic adjustment of the importance of the modality in the prediction scenario. The dynamic weight matrix generation formula is: in is the temporal branch projection feature It is a cross-modal interaction feature.
[0030] The visual feature vector encoding uses a deep residual network to extract fine-grained spatial features, extracts abstract features step by step through the convolutional layer group of ResNet-50, and obtains the visual feature vector through spatial attention pooling and linear projection. The encoding formula is: in is the trainable projection matrix is the bias term The text feature vector encoding is based on the dynamic semantic extractor of the pre-trained language model, and the BERT model is used to generate context-aware semantic vectors. The encoding formula is: in is the text projection matrix Input is the text description of the marketing activity. The activity feature vector encoding decomposes the key factors affecting user behavior into two dimensions: time feature and activity strategy feature. The time feature uses discrete encoding to capture the natural temporal patterns of weeks, months, and holidays. The activity strategy feature uses structured representation to depict the marketing strategy attributes of activity budget intensity, activity type, and duration cycle. The encoding formula is: in is the trainable parameter matrix Time characteristics It is the activity strategy feature.
[0031] The calculation process of the text feature vector to visual feature vector path in the four-path bidirectional attention mechanism is to first generate the query key-value pair through an independent projection matrix 、 、 , and then calculate the attention distribution of the text feature vector to the visual content by scaling the dot product attention The other three paths calculate the interactive representation of the activity feature vector to the visual feature vector, the activity feature vector to the text feature vector, and the text feature vector to the activity feature vector. The environment-aware gating mechanism concatenates the output features of the four independent attention paths in the four-path bidirectional attention mechanism with the original activity feature vector. , generating four-channel adaptive fusion weights through a learnable parameter matrix , the final multimodal fusion representation is generated by weighted combination .
[0032] The multi-scale feature extraction is based on the coupling characteristics of short-term promotional pulses and long-term cyclical laws in user consumption behavior. A multi-scale causal convolution architecture is used to extract hierarchical temporal features. Given an input sequence, a parallel multi-scale convolution kernel group is used to extract temporal patterns of different granularities. Dynamic feature aggregation introduces a learnable parameter matrix based on multi-granularity features to achieve cross-scale dynamic calibration. The weight matrix is broadcast along the time dimension and then element-by-element multiplied with the feature map. The dual-branch collaborative mechanism decouples static and dynamic features. The temporal evolution branch focuses on the periodic and trending changes in user behavior, while the cross-modal interaction branch models the immediate impact of marketing materials. Both are dynamically coupled through a dynamic fusion output layer.
[0033] The dynamic fusion output layer realizes the adaptive fusion of temporal features and cross-modal information through three steps. First, the output of the temporal evolution branch is projected into the feature space. , and then generate a dynamic weight matrix through cross-modal attention , the final output layer fuses information and maps the output dimension to obtain dynamic prediction results in Strengthening the timing-dominant mode Preserving cross-modal baseline information.
[0034] The game model includes an upper model with the goal of maximizing marketing effects and a lower model with the goal of optimizing resource allocation. The objective function of the upper model is , the constraints are 、 、 The objective function of the lower model is , the constraints are 、 The marketing effect maximization function is used to optimize the overall revenue performance of the marketing campaign, and the input includes user conversion rate Derived from historical marketing activity data statistics and activity coverage Derived from marketing activity characteristic parameters and user engagement Derived from user behavior data analysis and market competition intensity The resource allocation optimization function is used to minimize the resource input cost of marketing activities, and the input includes advertising cost. Derived from marketing budget allocation data and human resource costs Derived from staffing plans and technical maintenance costs Derived from system operating expenses and risk control costs Derived from risk assessment model and user conversion rate Derived from the output of the upper model, the output is the total cost control indicator used for resource allocation decision-making. Represents the mutual influence between user conversion rate and advertising cost.
[0035] Optionally, an adaptive learning rate adjustment function is also included for adjusting the learning rate parameters of the neural network. The function calculates an adjustment factor value based on the training loss change rate derived from model training process monitoring, the gradient norm derived from back propagation calculation, the model convergence speed derived from validation set performance tracking, and the historical performance indicators derived from model evaluation records. When the adjustment factor value is in the range of 0 to 0.3, it is used to adjust the learning rate parameters of the neural network using an exponential decay adjustment strategy. When the adjustment factor value is in the range of 0.3 to 0.7, it is used to adjust the learning rate parameters of the neural network using a linear decrease adjustment strategy. When the adjustment factor value is in the range of 0.7 to 1.0, it is used to adjust the learning rate parameters of the neural network using a cosine annealing adjustment strategy. The adjusted learning rate parameters are used to optimize the model training effect.
[0036] Optionally, a gating weight function is also included for adjusting the weight distribution of the gating mechanism of the neural network. The function calculates a balance value based on the modal feature importance score derived from attention weight analysis, the historical prediction accuracy derived from model performance records, the current training stage derived from training progress monitoring, and the model parameter complexity index derived from network structure analysis. When the balance value is in the range of 0 to 0.4, a conservative weight adjustment function is used to adjust the gating parameters. When the balance value is in the range of 0.4 to 0.6, a balanced weight adjustment function is used to adjust the gating parameters. When the balance value is in the range of 0.6 to 1.0, an aggressive weight adjustment function is used to adjust the gating parameters. The adjusted gating parameters are used to optimize the multimodal information fusion effect.
[0037] The specific implementation of the above steps is described in detail below.
[0038] The specific implementation of step S01 involves building a multi-dimensional marketing campaign data collection system. This system, based on the principle of data source diversity, enables comprehensive information acquisition. First, a text data collection module is established. Using natural language processing technology, it performs structured extraction of textual information such as marketing campaign titles, product descriptions, and promotional copy. Word segmentation algorithms and semantic analysis techniques are used to identify key marketing elements, converting unstructured text into a processable digital format. Next, an image data collection module is established. Computer vision technology is used to automatically collect and preprocess visual content such as marketing posters, product display images, and advertising creative materials. Image enhancement algorithms are used to improve image quality and standardize formatting. Next, a feature parameter collection module is established to systematically collect temporal characteristics of marketing campaigns, such as seasonal factors, holiday identifiers, and campaign cycles. Key parameters related to campaign strategies, such as budget size, campaign type, target audience, and channel distribution, are also collected. Finally, a data quality control mechanism is implemented to ensure the reliability of collected data through data integrity verification, outlier detection, and data consistency verification. A data standardization process is established to standardize data from different sources to the same dimensionality and format.
[0039] The specific implementation of step S02 involves establishing a three-way parallel multimodal feature encoding architecture. This architecture, based on deep learning and representation learning principles, achieves unified feature space mapping for heterogeneous data. The text feature extraction path utilizes pre-trained bidirectional encoder representation technology. The multi-head self-attention mechanism of the BERT model captures long-range dependencies between words in a text sequence. The encoder parameters are initialized using pre-trained weights from positional encoding and a masked language model, mapping variable-length text sequences into a dense vector representation of fixed dimension. The image feature extraction path utilizes a deep residual network architecture. Using a cascade of convolutional layers from the ResNet-50 model, the network extracts hierarchical visual features from edge texture to semantic concepts layer by layer. Skip connections are used to address the vanishing gradient problem of deep networks. Global average pooling is used to compress two-dimensional feature maps into one-dimensional feature vectors. The structured feature extraction path designs a hybrid encoding strategy for both discrete and continuous parameters of marketing campaigns. For temporal features, a combination of one-hot encoding and periodic encoding is used to capture periodic patterns in time series. A densely connected network is used to perform nonlinear transformations on campaign strategy features. Batch normalization is used to stabilize the training process and accelerate convergence.
[0040] The specific implementation of step S03 involves constructing a cross-modal information fusion architecture based on an attention mechanism. This architecture, based on the principle of multisensory information integration in cognitive psychology, achieves semantic alignment and information enhancement between features of different modalities. First, a four-pathway bidirectional attention network is established. Each path utilizes a query-key-value triplet attention computation model. Input features are projected into the query space, key space, and value space via a learnable linear transformation matrix. A scaled dot-product attention mechanism is then used to calculate the correlation weight distribution between features of different modalities. The text-to-image path uses text features as query vectors and image features as key-value pairs to actively retrieve and semantically match visual content from text descriptions, ensuring thematic consistency between marketing copy and visual materials. The image-to-text path employs the opposite query direction, using visual features to drive attention allocation to text keywords, strengthening the semantic correspondence between image content and text descriptions. The two paths, from activity features to text and image, respectively, model the constraints and guidance of structured parameters on unstructured content. Parameterized attention weights enable targeted optimization of creative content through marketing strategies. Then, an environmental perception gating mechanism is constructed to concatenate the output features of the four attention pathways with the original activity features. A softmax normalization function is used to generate adaptive fusion weights, achieving dynamic balancing and optimal combination of information from different pathways. Finally, a maximum-weighted bipartite graph matching algorithm is used to optimize cross-modal feature pairing. Feature vectors from different modalities are constructed as nodes of a bipartite graph. The Hungarian algorithm is used to find the optimal matching solution, ensuring maximum information preservation and optimal configuration during the feature fusion process.
[0041] The specific implementation of step S04 is to establish a static prediction architecture based on a stacked residual network. This architecture realizes regression prediction of marketing indicators based on the residual learning principle and ensemble learning ideas in deep learning. First, the backbone structure of the multi-layer perceptron network is constructed, and a progressive dimension mapping strategy from 512 dimensions to 384 dimensions and then to 256 dimensions is adopted. The expressive power of the model is introduced through nonlinear activation functions. Dropout regularization technology is added after each hidden layer to prevent overfitting. The dropout probability is set between 0.2 and 0.3 to balance the model capacity and generalization ability. Then, a residual connection mechanism is introduced to establish a parallel structure of feature transformation path and identity mapping path in each multi-layer perceptron block. The feature transformation path is responsible for learning the complex nonlinear mapping relationship from input to output, and the identity mapping path retains the original feature information through direct connection. The outputs of the two paths are fused through element-by-element addition operation, effectively alleviating the gradient vanishing problem in deep network training. Next, we add a processing chain consisting of batch normalization and self-attention layers. The batch normalization layer stabilizes the distribution of activation values during training through standardization, while the self-attention layer captures correlation patterns between elements within the feature vector through inner product operations, enhancing the model's ability to identify important feature dimensions. Finally, we establish a regression output layer, which uses a linear transformation to map the high-dimensional feature vector to the numerical space of marketing metrics. The mean squared error loss function is used to optimize prediction accuracy, and the static prediction results are output as a benchmark for evaluating the overall effectiveness of the marketing campaign.
[0042] The specific implementation of step S05 involves constructing a dual-branch collaborative dynamic prediction architecture. This architecture, based on time series analysis theory and multi-scale signal processing principles, enables serialized prediction of marketing indicators. The multi-scale sliding window module utilizes a parallel convolution kernel architecture. Three one-dimensional convolution kernels of different scales, 3 days, 7 days, and 30 days, are designed to capture temporal patterns representing short-term fluctuations, medium-term cycles, and long-term trends, respectively. Each convolution kernel employs a causal convolution structure to ensure that only historical information is used at the time of prediction. Nonlinear transformation capabilities are introduced through the ReLU activation function. Multi-scale features are fused through a weighted aggregation mechanism. The weight matrix is broadcast along the time dimension and then element-wise multiplied with the feature map. The temporal evolution branch models long-range dependencies of sequential patterns based on a causal Transformer architecture. A causal mask matrix is used to restrict the scope of attention computation, ensuring that each time step only accesses historical information from the current moment and before. A multi-head attention mechanism is used to parallelize the attention distribution across different representation subspaces. The position encoding module assigns a learnable position identifier to each time step in the sequence, enhancing the model's perception of temporal information. The cross-modal interaction branch establishes an interactive mechanism between the temporal global context and the multimodal fusion representation. Through pooling, temporal features are compressed into a global context vector as query information. The multimodal fusion representation undergoes a linear transformation to generate key-value pairs. The attention calculation results represent the contribution of different modal information to temporal prediction. The dual-branch collaborative mechanism integrates information from the temporal evolution and cross-modal interaction branches through a gated fusion strategy. The weight distribution of the two branches is dynamically adjusted based on the different prediction scenarios, increasing the weight of the temporal branch in trend prediction scenarios and enhancing the importance of the cross-modal branch in emergency prediction scenarios.
[0043] The specific implementation of step S06 involves designing a dynamic fusion output architecture with adaptive weight allocation. This architecture, based on the attention mechanism and information fusion theory, achieves the optimal combination of multi-source prediction information. First, the output features of the temporal evolution branch undergo a spatial projection transformation. A causal one-dimensional convolutional network is used to map the temporal features into a representation space that matches the cross-modal features. The convolution kernel size is set between 3 and 5 to balance receptive field coverage and computational complexity, and a zero-padding strategy is used to maintain sequence length consistency. A cross-modal attention mechanism is then constructed to generate a dynamic weight matrix. The projected temporal features serve as query vectors, and the cross-modal interaction features serve as key vectors. An inner product operation is used to calculate the similarity scores between features. A softmax function is used for normalization to obtain the attention weight distribution. Each element of the weight matrix reflects the importance of the corresponding time step and modality dimension. Next, a weighted fusion operation is performed: a matrix product operation is performed on the dynamic weight matrix and the temporal projection features. This highlights the key information in the temporal pattern while retaining the cross-modal baseline features as supplementary information. A linear combination of the two types of information is achieved through element-by-element addition. Finally, an output mapping layer is established, and a multi-layer perceptron network is used to map the fused high-dimensional feature vector to the target prediction dimension. The fitting ability of the model is enhanced by a nonlinear activation function. The output layer does not use an activation function to support numerical predictions in any range, and the final dynamic prediction results are obtained.
[0044] The specific implementation of step S07 is to establish a multi-objective optimization training mechanism within a game theory framework. This mechanism, based on two-level optimization theory and reinforcement learning principles, achieves the coordinated optimization of marketing effectiveness and resource allocation. The upper-level game model optimizes marketing effectiveness. The objective function uses a combination of polynomials to comprehensively consider four key factors: user conversion rate, campaign coverage, user engagement, and market competition intensity. The logarithmic and square root terms are used to model the law of diminishing marginal returns, the quadratic term is used to characterize the negative impact of competition intensity, and the interaction term is used to capture the synergistic effect of conversion rate and coverage. The coefficients are determined by fitting historical data. Constraints include practical business constraints such as budget caps, coverage caps, and competition intensity caps. The lower-level game model optimizes resource allocation. The objective function uses a combination of quadratic, linear, exponential, and interaction terms to model the nonlinear relationship between advertising costs, human resource costs, technical maintenance costs, and risk control costs. The quadratic term is used to characterize scale effects, the exponential term is used to model the rapid growth of technical costs, and the interaction term establishes the coupling relationship between the upper and lower-level models. Constraints include resource constraints such as the total cost cap and the risk control cap. The game is solved using an iterative optimization algorithm. The upper-level model uses gradient ascent to find the optimal solution for marketing effectiveness, while the lower-level model uses gradient descent to find the cost-minimizing solution. The two-level models exchange information through coupling terms, and the decision variables are dynamically adjusted during the iteration process until a Nash equilibrium is reached. A backpropagation algorithm is used to update neural network parameters. An adaptive learning rate adjustment strategy is used to calculate the adjustment factor based on the rate of change of training loss, gradient norm, model convergence speed, and historical performance indicators. When the adjustment factor is in the range of 0 to 0.3, an exponential decay strategy is used; when it is in the range of 0.3 to 0.7, a linear decrease strategy is used; and when it is in the range of 0.7 to 1.0, a cosine annealing strategy is used to ensure the stability and convergence of model training.
[0045] The key technical ideas of the present invention include the following aspects. Multimodal feature extraction and cross-modal information fusion technology uses deep neural networks to perform feature encoding on text, images and structured parameters respectively, and adopts a four-channel bidirectional attention mechanism to achieve semantic alignment and information enhancement between different modalities. Compared with the traditional single-modal prediction method, it can make full use of the multi-dimensional information of marketing activities, overcome the limitations of incomplete information of a single data source, and improve the richness of feature representation and prediction accuracy. The dynamic and static combined dual-branch prediction architecture captures the overall effect benchmark of the marketing activities through the static prediction module, and the dynamic prediction module models the time series change pattern. The two complement each other through an adaptive fusion mechanism. Compared with the traditional single prediction model, it can simultaneously process the steady-state characteristics and time-varying characteristics of the marketing activities, and enhance the model's adaptability to complex marketing environments. Multi-scale time series feature extraction technology uses parallel multi-scale convolution kernel groups to capture the time series patterns of short-term fluctuations, medium-term cycles and long-term trends respectively. Combined with the long-distance dependency modeling capability of the causal Transformer, it can more comprehensively characterize the multi-level time laws of user consumption behavior compared to the traditional single-scale time series analysis method, thereby improving the accuracy and robustness of time series prediction. A game-theory-driven multi-objective optimization framework optimizes the dual objectives of maximizing marketing effectiveness and optimizing resource allocation through upper and lower-level game models, achieving a balanced decision-making process between effectiveness and cost in marketing strategy formulation. Compared to traditional single-objective optimization methods, it better reflects the complexity and multi-objective nature of actual marketing decisions. The synergy of these technical approaches forms a complete marketing forecasting and optimization system. Multimodal information fusion provides a rich feature foundation for forecasting, a dynamic and static forecasting architecture ensures the comprehensiveness of forecast results, multi-scale time series modeling enhances sensitivity to temporal changes, and a game optimization framework ensures the practicality of forecast results. Compared to existing technologies, this system achieves technological innovation throughout the entire process, from data processing and feature extraction to model forecasting and decision optimization.
[0046] The detailed structure of the marketing campaign indicator prediction model consists of five main components: a data input layer, a multimodal feature extraction layer, a cross-modal information fusion layer, a dual-branch prediction layer, and an output fusion layer. The data input layer receives three types of heterogeneous data: text descriptions of marketing campaigns, promotional images, and feature parameters. The data preprocessing module standardizes the format and performs quality control. The multimodal feature extraction layer uses a BERT encoder to process text data, a ResNet-50 network to process image data, and a structured encoder to process parameter data, mapping the heterogeneous inputs into a unified high-dimensional feature space. The cross-modal information fusion layer uses a four-way bidirectional attention network to establish interactive relationships between features of different modalities and optimizes feature pairing using a maximum weighted bipartite graph matching algorithm to generate a fused multimodal representation. The dual-branch prediction layer consists of a static prediction branch and a dynamic prediction branch. The static branch uses a stacked residual multilayer perceptron to achieve overall indicator prediction, while the dynamic branch combines multi-scale temporal feature extraction and a causal transformer to achieve sequential prediction. The output fusion layer adaptively combines the prediction results of the two branches using a dynamic weight matrix to generate the final marketing indicator prediction value.
[0047] The detailed steps for model training begin with data preparation and preprocessing. Historical marketing campaign data is collected and divided into training, validation, and test sets. Text data is segmented and a vocabulary is constructed. Image data is size-normalized and augmented. Feature engineering and missing value handling are performed on structural parameters. Model parameters are then initialized. The BERT encoder is initialized using pretrained weights, the ResNet-50 network uses ImageNet pretrained weights, and other network layers use Xavier or He initialization strategies. Multi-stage training is then performed. In the first stage, the pretrained layer parameters are frozen and only the newly added network layers are trained. In the second stage, all parameters are unfrozen for end-to-end fine-tuning. In the third stage, a game optimization framework is introduced for multi-objective joint training. During training, adaptive learning rate adjustment and gated weight adjustment strategies are used. Hyperparameters are dynamically adjusted based on validation set performance. Early stopping is used to prevent overfitting. Finally, model performance is evaluated on the test set.
[0048] The marketing activity indicator prediction model established by this invention is capable of handling multi-source heterogeneous data and complex temporal dependencies in marketing environments. Traditional marketing forecasting methods, primarily based on statistical regression models or simple time series analysis, can only process a single type of structured data, struggle to integrate unstructured information such as text and images, and are unable to effectively model the interactions between multimodal data. Compared to traditional methods, the advantages of the model presented in this invention lie in its multimodal information fusion capabilities, its combined dynamic and static prediction mechanism, and its game-theoretic optimization framework. Multimodal information fusion uses deep learning technology to map heterogeneous data into a unified representation space, overcoming the inadequate information utilization of traditional methods and improving the comprehensiveness and accuracy of predictions. This combined dynamic and static prediction mechanism simultaneously considers both the steady-state and time-varying characteristics of marketing activities, enabling a more comprehensive characterization of complex marketing patterns than single static or dynamic prediction methods, enhancing the model's adaptability and robustness. The game-theoretic optimization framework jointly optimizes marketing effectiveness maximization and resource allocation optimization as dual objectives. Compared to traditional single-objective optimization methods, this framework better reflects the multi-objective nature of actual marketing decisions, improving the practical value and decision-making guidance of the prediction results.
[0049] It should be noted that the present invention also solves the following technical problems: First, the present invention solves the technical problem that the separation of static and dynamic feature modeling in marketing activity prediction leads to inaccurate capture of temporal dependencies. Traditional prediction methods usually process the static attributes and dynamic temporal features of marketing activities independently, and cannot effectively model the interaction between static and dynamic features. The present invention designs a dual-branch collaborative mechanism, in which the temporal evolution branch captures the periodic and trend changes in user behavior, and the cross-modal interaction branch models the immediate impact of marketing materials. The two branches are adaptively coupled through a dynamic fusion output layer, so that the static marketing activity attributes can be organically combined with the dynamic temporal evolution pattern, thereby accurately capturing the complex dependencies between static and dynamic features in marketing activities. Secondly, the present invention solves the technical problem that the fixed weight distribution in the multimodal feature fusion process leads to poor adaptability to different prediction scenarios. Existing multimodal fusion methods usually adopt a fixed weight distribution strategy, which cannot dynamically adjust the importance of each modality according to the specific prediction scenario and data characteristics. By introducing an environmental perception gating mechanism and a dynamic weight matrix, the present invention can adaptively adjust the fusion weights of text, visual, and activity features according to the current prediction scenario. The elements of the dynamic weight matrix reflect the intensity of attention paid to temporal patterns at different prediction steps, thereby realizing dynamic adjustment of modal importance in the prediction scenario and significantly improving the adaptability and prediction accuracy of the model in different marketing environments.
[0050] Specifically, the present invention addresses the technical issue of insufficient feature fusion in multimodal marketing campaign data. This is primarily based on the following technical principles: First, the textual descriptions, promotional images, and feature parameters of marketing campaigns are inherently correlated in semantic space. A four-way bidirectional attention mechanism establishes a bidirectional interaction channel between different modalities, enabling textual features to actively align with visual content to ensure semantic consistency between the campaign description and the image theme. Activity features can also drive attention to key image regions, thereby achieving fine-grained cross-modal semantic alignment. Second, marketing campaign data exhibits multi-scale temporal regularities, and user behavior follows natural time cycles and cognitive psychological laws. A multi-scale sliding window can simultaneously capture short-term promotional pulses, medium-term cyclical patterns, and long-term trend changes. Multi-scale feature extraction conforms to the hierarchical temporal structure of user consumption behavior. Third, the static prediction module implements nonlinear feature transformation through a stacked residual multi-layer perceptron network, with residual connections ensuring efficient gradient propagation. The dynamic prediction module employs a causal Transformer to ensure causal constraints in time series modeling. A dual-branch collaborative mechanism achieves decoupled modeling and adaptive fusion of static and dynamic features. Finally, the maximum weight bipartite graph matching algorithm constructs different modal features into a graph structure, and achieves global optimization fusion of multimodal information by solving the optimal cross-modal feature pairing weights. The dynamic fusion output layer further realizes the adaptive adjustment of modal importance in the prediction scenario through the dynamic weight matrix generated by cross-modal attention, ensuring the effectiveness of multimodal fusion representation and the accuracy of prediction results.
[0051] A specific embodiment 1 of the present invention is provided below. The specific implementation of each step in this embodiment 1 is described in detail as follows.
[0052] In this embodiment, the specific implementation of step S01 is the same as above and will not be described in detail here.
[0053] The specific implementation of step S02 is to build a multimodal feature extraction module and encode different types of data through a deep neural network. The visual feature vector encoding uses a deep residual network to extract fine-grained spatial features, which is specifically expressed as follows: ; Where, is the visual feature vector; is a trainable projection matrix with a dimension of 512×256; It is the feature extraction function of the ResNet-50 network for the input image; Input for marketing campaign promotional images, size 224×224×3; is the bias vector with a dimension of 256 × 1. The text feature vector encoding is based on the dynamic semantic extractor of the pre-trained language model, which is specifically expressed as follows: ; Where, is the text feature vector; is the text projection matrix, with a dimension of 768×256; The encoding function of the BERT model for the input text; Enter a text description for the campaign; is a text bias term with a dimension of 256 × 1. Activity feature vector encoding decomposes the key factors affecting user behavior into two dimensions: time features and activity strategy features, which are specifically expressed as follows: ; Where, is the activity feature vector; is a trainable parameter matrix with a dimension of 128×256; Time features include discrete codes such as week, month, and holidays, with a dimension of 64; Activity strategy characteristics, including budget intensity, activity type, duration and other parameters, with a dimension of 64; Represents feature concatenation operation; is the activity feature bias term with a dimension of 256×1.
[0054] The specific implementation of step S03 is to design a dynamic cross-modal information fusion mechanism, which realizes the interactive fusion of different modal features through a four-path bidirectional attention mechanism. The calculation process of the text feature vector to visual feature vector path first generates query key-value pairs through an independent projection matrix, which is specifically expressed as follows: , , ; Where, is the query matrix for text to image; is the image key matrix; is the image value matrix; 、 、 The projection weight matrices of query, key, and value are 256×64 respectively. The attention distribution of text feature vector to visual content is then calculated by scaling dot product attention, which is specifically expressed as follows: ; Where, Output features for text-to-image attention; is the model dimension parameter, the value is 64; is a normalization function. The context-aware gating mechanism concatenates the output features of the four independent attention channels in the four-channel bidirectional attention mechanism with the original activity feature vector, which is specifically expressed as follows: ; Where, is the comprehensive feature matrix after splicing; The attention output of activity features to images; Attention output from activity features to text; Attention output for text to activity features; is the environment perception projection matrix with a dimension of 256 × 64. The adaptive fusion weights of the four channels are generated through the learnable parameter matrix, which is specifically expressed as follows: ; Where, is the adaptive fusion weight vector with a dimension of 4×1; is the weight generation matrix with a dimension of 320 × 4. The final multimodal fusion representation is generated through weighted combination, which is specifically expressed as follows: ; Where, It is a multimodal fusion representation; is the fusion weight of the kth path; is the output feature of the kth attention path, where 、 、 、 They correspond to the outputs of the four pathways: text to image, activity feature to image, activity feature to text, and text to activity feature.
[0055] The specific implementation of step S04 is to establish a static prediction module, input the multimodal fusion representation into the stacked residual multilayer perceptron network to achieve regression prediction of the total index of the marketing activity, which is specifically expressed as follows: ; Where, is the static prediction result; is the multi-layer perceptron function; Represents the Lth residual multilayer perceptron block, where L is 3; Represents a function composite operation; Represent input for multimodal fusion.
[0056] The specific implementation of step S05 is to build a dynamic prediction module and use a multi-scale sliding window to extract time series features. Multi-scale feature extraction uses a parallel multi-scale convolution kernel group to extract time series patterns of different granularities, which is specifically expressed as follows: ; Where, is the temporal feature of the i-th scale; Represents the index of convolution kernels of different scales, with values of 1, 2, and 3 corresponding to 3-day, 7-day, and 30-day windows respectively; is the rectified linear unit activation function; is the i-th one-dimensional convolution operation; is the input time series data sequence, where is the index of the current time step. Dynamic feature aggregation introduces a learnable parameter matrix based on multi-granularity features to achieve cross-scale dynamic calibration, which is specifically expressed as follows: ; Where, is the aggregation sequence feature; is the number of scales, the value is 3; is the weight matrix of the i-th scale, with the same dimension as same; Represents an element-by-element product operation. The temporal evolution branch is based on the sequential pattern capture of the causal Transformer, and the temporal representation generation is specifically expressed as follows: ; Where, It is the temporal evolution feature; is the causal transformer function; is the learnable position encoding matrix, with the same dimension as The cross-modal interaction branch achieves feature enhancement through the interaction of temporal global context and multimodal fusion representation. The query vector generation is specifically expressed as follows: ; Where, Query vector for cross-modal interaction; It is a global pooling operation; To query the projection matrix, the dimension is 256×64. The key-value pair generation is specifically expressed as follows: , ; Where, and are the key matrix and value matrix of cross-modal interactions, respectively; and are the projection weight matrices of keys and values, respectively, both with dimensions of 256 × 64. The cross-modal interaction features are calculated through the attention mechanism, and are specifically expressed as follows: ; Where, It is a cross-modal interaction feature; It is the interactive feature dimension and its value is 64.
[0057] The specific implementation of step S06 is to design a dynamic fusion output layer and generate a dynamic weight matrix through cross-modal attention. First, the output of the temporal evolution branch is projected into the feature space, which is specifically expressed as follows: ; Where, is the temporal branch projection feature; is a causal one-dimensional convolution function with a kernel size of 3. The dynamic weight matrix is generated by cross-modal attention, which is specifically expressed as follows: ; Where, is the dynamic weight matrix; It is a cross-modal interaction feature; is the feature dimension, with a value of 256. The dynamic fusion output layer fuses information and maps the output dimension to obtain the dynamic prediction result, which is specifically expressed as follows: ; Where, For dynamic prediction results; is the output layer multilayer perceptron; Strengthen the time-dominant model; Preserving cross-modal baseline information.
[0058] The specific implementation of step S07 is to use historical marketing activity data to train and optimize the marketing activity indicator prediction model, and balance the dual goals of maximizing marketing effectiveness and optimizing resource allocation through a game model. The upper-level model aims to maximize marketing effectiveness, and the objective function is specifically expressed as follows: ; Where, Function for maximizing marketing effect; is the user conversion rate, ranging from 0.01 to 0.5; is the activity coverage, ranging from 0.1 to 1.0; is user engagement, ranging from 0.05 to 0.8; is the market competition intensity, ranging from 0.1 to 1.0; 、 、 、 are undetermined coefficients, which are 2.5, 1.8, 0.3, and 1.2 respectively. The constraints include 、 、 ,in is the budget constraint parameter, which takes the value of 1.2; is the lower limit of participation, with a value of 0.1; The upper limit of competition intensity is 0.9. The lower model aims to optimize resource allocation, and the objective function is specifically expressed as follows: ; Where, Optimize functions for resource allocation; The advertising cost ranges from 1000 to 50000. is the human resource cost, ranging from 500 to 20,000; Technical maintenance cost, ranging from 200 to 10,000; is the risk control cost, ranging from 100 to 5000; 、 、 、 is the cost weight coefficient, with values of 0.001, 1.0, 0.5, and 2.0 respectively; is the coupling coefficient, which takes a value of 10000. The constraints include 、 ,in is the upper limit of the total cost, which is 60000; It is the lower limit of risk control and its value is 200.
[0059] The principles and effects of each formula are explained as follows. Visual feature vector encoding formula Based on the feature extraction principle of deep convolutional neural networks, the spatial semantic information of images is extracted through the hierarchical feature learning mechanism of the residual network. The linear projection operation maps high-dimensional visual features to a unified feature space. The bias term is used to adjust the mean shift of the feature distribution. Compared with traditional manual feature extraction methods, it can automatically learn richer visual representations, improve the utilization efficiency of image information and the semantic understanding ability of marketing materials, and provide a high-quality visual feature foundation for subsequent cross-modal fusion. Based on the context-aware encoding principle of the pre-trained language model, the bidirectional attention mechanism of the BERT model is used to capture the complex semantic relationships between words in the text sequence. The projection matrix realizes the transformation from the pre-trained feature space to the task-adapted feature space. Compared with the traditional bag-of-words model or TF-IDF method, it can better understand the semantic content and sentiment tendency of marketing texts, enhance the expressiveness and semantic consistency of text information, and provide deep semantic support for text understanding of marketing activities. Based on the encoding principle of multi-dimensional feature fusion, the heterogeneous information of the time dimension and the strategy dimension is uniformly encoded through feature splicing operations, and the linear transformation matrix learns the interaction relationship between different feature dimensions. Compared with the single-dimensional feature encoding method, it can more comprehensively characterize the multifaceted attributes of marketing activities, improve the information density and prediction relevance of structured parameters, and provide the model with rich marketing environment context information.
[0060] Query key-value pair generation formula in four-channel bidirectional attention mechanism 、 、 Based on the information retrieval principle of the attention mechanism, different modal features are mapped to mutually matching query search spaces through independent projection transformations, achieving accurate alignment and correlation calculation of cross-modal information. Compared with simple feature splicing methods, it can establish more refined semantic correspondences between modalities, improving the accuracy and effectiveness of multimodal information fusion. Based on the weighted aggregation principle of soft attention, through inner product operation Calculate the similarity score between the query and the key, scaling factor Prevent the vanishing gradient problem, Normalization ensures the probabilistic nature of weight distribution, which can achieve smoother information integration and more stable training process compared to the hard attention mechanism, and enhances the robustness and generalization ability of cross-modal attention. Based on the fusion principle of adaptive weight aggregation, the learned weight coefficients Dynamically balancing the contributions of different attention pathways achieves the optimal combination of multimodal information. Compared with the fixed-weight fusion method, it can adaptively adjust the fusion strategy according to the characteristics of the input content, thereby improving the expressive power and task adaptability of the multimodal representation.
[0061] Multi-scale feature extraction formula Based on the principle of multi-resolution time series analysis, the time series patterns of short-term fluctuations, medium-term cycles and long-term trends are extracted in parallel through convolution kernels of different window sizes. The activation function introduces nonlinear transformation capabilities, which can more comprehensively capture the multi-level temporal patterns of user consumption behavior compared to single-scale time series analysis methods, improving the accuracy of time series prediction and the ability to model complex time patterns. Dynamic feature aggregation formula Based on the principle of multi-scale information fusion, adaptive weighted aggregation of different scale features is performed through a learnable weight matrix, and element-by-element product operation is performed. This achieves fine-grained feature adjustment, which can better balance the importance of different time scales compared to simple average or maximum aggregation methods, and enhances the expressive richness and predictive relevance of temporal features. Based on the information retrieval principle of the attention mechanism, the query vector With the key matrix The inner product operation of is used to calculate the correlation score. Normalization ensures the probabilistic characteristics of weight distribution. Compared with simple feature splicing methods, it can achieve more accurate cross-modal information alignment and dynamic weight distribution, thereby improving the effectiveness and prediction accuracy of multimodal information fusion.
[0062] Dynamic weight matrix generation formula Based on the dynamic weight allocation principle of the self-attention mechanism, the inner product operation of temporal features and cross-modal features is performed Calculate the correlation matrix, Normalization ensures the probability distribution characteristics of weights, scaling factors Maintaining numerical stability, compared to the fixed weight fusion method, it can dynamically adjust the importance of different information sources according to the prediction scenario, improving the adaptive ability of the prediction model and its responsiveness to complex marketing environments. Dynamic prediction result formula Based on the prediction principle of weighted information fusion, the dynamic weight matrix Strengthening the timing-dominant mode , while preserving cross-modal baseline information As a supplement, The final nonlinear mapping is achieved, which can better integrate the complementary information of temporal evolution and cross-modal interaction compared to the prediction method of a single information source, thereby improving the accuracy and stability of dynamic prediction.
[0063] Upper-level game model objective function Based on the economic principle of multi-objective utility maximization, The marginal effect of modeling conversion rate and coverage is diminishing. The term characterizes the concave function characteristics of participation, The term represents the negative impact of competitive intensity, The interaction term captures the synergistic effect of conversion rate and coverage. Compared with the single-objective linear optimization method, it can more accurately reflect the complex nonlinear relationship and multi-factor interaction of marketing effects, thus improving the scientificity and practicality of marketing strategy optimization. Based on the principle of resource allocation based on cost minimization, Modeling the diseconomies of scale in advertising, The term represents the proportional relationship of labor costs, The rapid growth of technology maintenance costs is characterized by The item reflects the cost-saving effect of risk control. The coupling term establishes the mutual influence relationship between the upper and lower layer models. Compared with the traditional single-layer optimization method, it can better balance the dual constraints of marketing effect and resource input, and achieve the overall optimization of marketing decisions and the efficiency of resource allocation.
[0064] To better understand and implement the present invention, Example 2 of a specific application scenario is provided below: A technical team established an indicator prediction model for a mobile communications service-based phone recharge marketing campaign. This campaign, titled "Spring Festival Phone Recharge Discount Campaign," featured a promotional text describing the campaign as "Recharge 100 yuan during the Spring Festival and receive a 20 yuan bonus, limited to 7 days." The promotional image was a poster featuring Spring Festival elements and discount information. The campaign's characteristic parameters included a time feature (during the Spring Festival, lasting 7 days) and a strategy feature (recharge 100 yuan, receive a 20 yuan bonus, and offer an 80% discount). The technical team used a method for establishing a marketing campaign indicator prediction model to perform both static and dynamic predictions of user participation and recharge amounts.
[0065] The technical team first collected multimodal data from historical marketing campaigns as training samples. The text description data included 3,000 promotional copy from historical campaigns, such as "Recharge 50 yuan on weekends and get 10 yuan free" and "Recharge 200 yuan on National Day and get 50 yuan free." After encoding with the BERT model, a 768-dimensional text feature vector was generated. The image data consisted of 2,800 corresponding promotional posters, sized 1080 × 1920 pixels. A 2048-dimensional visual feature vector was extracted using the ResNet-50 model. The campaign feature parameters included time features (day of the week, month, holiday) and campaign strategy features (recharge amount, gift ratio, campaign type, and duration). After processing with a structured encoder, a 256-dimensional campaign feature vector was generated.
[0066] In the multimodal feature extraction stage, the technical team maps the feature vectors of the three modalities to a 512-dimensional shared latent space through a trainable projection matrix. The visual feature vector is obtained by encoding "Recharge 100 yuan during the Spring Festival and get 20 yuan of phone credit, limited to 7 days" through the BERT model. The activity feature vector is obtained by encoding the Spring Festival theme poster through the ResNet-50 model. The time features (Spring Festival, 7 days) and activity strategy features (100 yuan recharge, 20 yuan gift, 20% discount rate) are structuredly coded.
[0067] The dynamic cross-modal information fusion mechanism uses a four-pathway bidirectional attention mechanism to achieve feature interaction. The text-to-visual feature vector pathway calculates an attention weight matrix. The "Spring Festival" keyword in the text and the Spring Festival decorations in the poster have a high attention weight of 0.85, while the phrase "recharge 100 yuan" and the amount display area in the poster have an attention weight of 0.78. In the activity-to-visual feature vector pathway, the temporal feature "Spring Festival" and the festive atmosphere of the poster have an association weight of 0.82, while the activity strategy feature "20% discount" and the discount logo in the poster have an attention weight of 0.76. The activity-to-text feature vector pathway verifies the compatibility of the activity parameters with the text description. The recharge amount parameter matches the "100 yuan" in the text with a matching degree of 0.91. The text-to-activity feature vector pathway extracts the temporal parameter information corresponding to the phrase "limited to 7 days" in the text, with a matching weight of 0.88.
[0068] The environment-aware gating mechanism adaptively fuses the outputs of the four pathways, as shown in Table 1: Table 1 Four-channel attention weight distribution table
[0069] Multimodal fusion representation generated by weighted combination The dimension is 512, which contains comprehensive information of text semantics, visual content and activity parameters.
[0070] The static prediction module uses multimodal fusion representation to predict the total activity indicators. The technical team built a stacked network containing 4 residual multi-layer perceptron blocks. Each block uses a 512→384→256-dimensional progressive feature transformation path and a residual direct connection path. The input multimodal fusion features are sequentially constrained to a zero-mean unit variance through the batch normalization layer. The self-attention layer dynamically filters key information and assigns an attention weight of 0.73 to the recharge amount-related features, a weight of 0.68 to the time node features, and a weight of 0.52 to the visual attractiveness features. After processing by the stacked residual network, the static prediction module outputs that the expected number of participating users for the Spring Festival recharge activity is 28,500, and the expected total recharge amount is Yuan.
[0071] The dynamic prediction module uses a multi-scale time series modeling approach to process historical activity data sequences. The technical team collected historical activity metrics data from the past 180 days, including daily participating users and top-up amounts. A multi-scale sliding window uses convolutional kernels of three, seven, and fourteen, scales to extract time series features of varying granularity. The three-day short kernel captured fluctuations in user behavior within the 72-hour period of the weekend promotion. Extracted features show that top-up volume increased 65% on Friday evening compared to weekdays, peaked on Saturday, and then declined on Sunday. The seven-day medium kernel analyzed cyclical consumption patterns, identifying that weekly top-up peaks are concentrated on Fridays and Saturdays, accounting for 42% of the total weekly top-up volume. The 14-day long kernel identified the impact of the bi-weekly payroll cycle on top-up behavior, finding that top-up activity at the beginning and middle of the month was 23% and 18% higher, respectively, than during other periods.
[0072] As shown in Table 2, the weight distribution of multi-scale feature extraction: Table 2 Multi-scale convolution kernel weight distribution table
[0073] The dual-branch collaborative mechanism includes a temporal evolution branch and a cross-modal interaction branch. The temporal evolution branch is based on the causal Transformer architecture, which inputs aggregated multi-scale features and learnable position encoding to generate temporal representations. This branch identifies the three-stage characteristics of user recharge behavior during the Spring Festival: the recharge volume gradually increases in the week before the Spring Festival, the recharge activity is relatively slow from New Year's Eve to the third day of the New Year, and the recharge volume quickly rebounds after the fourth day of the New Year as work resumes. The cross-modal interaction branch generates query vectors, key vectors, and value vectors through the interaction of temporal global context and multimodal fusion representation, and calculates the interaction features. ,This feature reflects the impact of activity materials on immediate user behavior.
[0074] The dynamic fusion output layer realizes the adaptive fusion of temporal features and cross-modal information through three steps. The output of the temporal evolution branch is projected into the feature space through causal convolution to obtain the projected features. . Cross-modal attention mechanism generates dynamic weight matrix The weight elements of this matrix reflect the intensity of attention paid to temporal patterns at different forecasting steps. During the 7-day forecast period for the Spring Festival recharge activity, the weights for the first three days favored short-term promotional effects, with weights of 0.82, 0.79, and 0.74, respectively. The weights for the last four days were more focused on long-term trend patterns, with weights of 0.65, 0.58, 0.52, and 0.47, respectively.
[0075] The technical team optimizes model parameters through historical data training and uses a game model to balance the dual goals of maximizing marketing effects and optimizing resource allocation. The upper model aims to maximize marketing effects, and the input parameters include user conversion rate. 0.156, activity coverage 0.734, user engagement 0.682, market competition intensity The objective function is calculated to get a marketing effect score of 4.23. The lower model aims to optimize resource allocation, and the input parameters include advertising cost. for Yuan, human resource cost for Yuan, technical maintenance cost for Yuan, risk control cost for Yuan, the total cost control index is calculated as Yuan.
[0076] During model training, an adaptive learning rate adjustment function was used. A correction factor of 0.54 was calculated based on the training loss change rate of 0.023, the gradient norm of 1.45, the model convergence speed of 0.67, and the historical performance index of 0.84. Since the correction factor value falls within the range of 0.3 to 0.7, the system adopted a linear decreasing adjustment strategy, gradually adjusting the learning rate from an initial value of 0.001 to 0.0003. The gating weight function calculated an equilibrium value of 0.65 based on a modal feature importance score of 0.76, a historical prediction accuracy of 0.82, a current training stage of 0.65, and a model parameter complexity index of 0.58. An aggressive weight adjustment function was used to optimize the gating parameters and improve the multimodal information fusion effect.
[0077] After 200 rounds of iterative training, the performance of the model on the validation set is shown in Table 3: Table 3 Model validation set performance indicators
[0078] The technical team applied the trained model to the actual prediction of the Spring Festival recharge event. The static prediction results showed that the expected number of participating users during the event was 28,500, and the expected total recharge amount was Yuan, the average number of daily recharge users is 4071, and the average daily recharge amount is The dynamic prediction results show the changing trend of indicators during the 7-day activity: on the first day, affected by the activity launch effect, the number of participating users is predicted to be 5,200, and the recharge amount is Yuan; it maintained a high level on the 2nd and 3rd days, with a predicted daily average of 4,800 participating users; it fell slightly on the 4th and 5th days, with the average daily number of participating users dropping to 4,200; on the 6th and 7th days, as the event neared its end, there was a final sprint effect, and the predicted number of participating users rebounded to 4,600.
[0079] The model also predicts user behavior patterns at different times of day. It predicts that top-up users will account for 26% of the morning 10:00-12:00 PM, 22% of the afternoon 2:00-16:00 PM, 31% of the evening 7:00-21:00 PM, and 21% of other times of day. Age distribution forecasts show that users aged 25-35 are the most engaged, accounting for 38% of all participating users. Users aged 36-45 account for 32%, users aged 18-25 account for 21%, and other age groups account for 9%.
[0080] Compared with traditional marketing campaign forecasting methods, this model building method represents a significant technological advancement. Traditional forecasting methods are typically based on single-modal data, such as using only historical numerical data for time series analysis, or performing rule matching based solely on campaign parameters. These methods struggle to capture the complex interactions between text semantics, visual content, and campaign features. The present invention achieves deep semantic alignment of heterogeneous data through multimodal feature extraction and a dynamic cross-modal information fusion mechanism. This allows for a more comprehensive understanding of marketing scenarios by simultaneously leveraging the expressiveness of promotional copy, the visual appeal of promotional posters, and the structured information of campaign parameters. Traditional time series modeling methods often employ fixed-window or single-scale analysis, failing to simultaneously capture the coupling effects of short-term promotional pulses and long-term consumption patterns. The present invention's multi-scale sliding window and dual-branch collaborative mechanism, through differentiated time granularity extraction and decoupling of static and dynamic features, can accurately model multi-level time series features, from hourly promotional responses to monthly consumption cycles. Traditional forecasting models often separate static planning from dynamic regulation, resulting in significant deviations between expectations during the campaign planning phase and actual performance during the execution phase. The static prediction and dynamic prediction collaborative architecture of the present invention realizes full-cycle integrated prediction from activity planning to execution monitoring through shared multimodal fusion representation and consistent feature coding system, thereby improving the continuity and credibility of the prediction results.
[0081] It should be noted that the detailed explanation of the variables involved in the present invention is shown in Table 4.
[0082] Table 4 Variable Explanation Table
[0083] The above description is only a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any technician familiar with this technical field can easily think of changes or replacements within the technical scope disclosed by the present invention, which should be covered by the scope of protection of the present invention.
Claims
1. A method for establishing a marketing activity indicator prediction model, characterized in that: Collect multimodal data of marketing activities, including marketing activity text descriptions, marketing activity promotional images, and marketing activity feature parameters; construct a multimodal feature extraction module, encode the marketing activity text descriptions through the BERT model to obtain text feature vectors, encode the marketing activity promotional images through the ResNet-50 model to obtain visual feature vectors, and encode the marketing activity feature parameters through the structured encoder to obtain activity feature vectors; design a cross-modal information dynamic fusion mechanism, construct a four-channel bidirectional attention mechanism to achieve interactive fusion between different modal feature vectors, construct different modal feature vectors into bipartite graph nodes through the maximum weight bipartite graph matching algorithm and solve the optimal cross-modal feature pairing weights to generate a multimodal fusion representation; establish a static prediction module, input the multimodal fusion representation into the stacked residual multi-layer perceptron network to achieve regression prediction of the total marketing activity indicator and obtain static prediction results; construct a dynamic prediction module, use a multi-scale sliding window to extract time series features, and realize dynamic prediction of the marketing activity indicator sequence through a dual-branch collaborative mechanism of time series evolution branch and cross-modal interaction branch; A dynamic fusion output layer is designed to generate a dynamic weight matrix through cross-modal attention, adaptively fuse the temporal evolution branch output and the cross-modal interaction branch output to obtain dynamic prediction results; historical marketing activity data is used to train and optimize the marketing activity indicator prediction model.
2. The method for establishing a marketing activity indicator prediction model according to claim 1, characterized in that: The marketing activity characteristic parameters specifically include time characteristics and activity strategy characteristics. The time characteristics use discrete coding to capture the natural time sequence patterns of weeks, months and holidays. The activity strategy characteristics use structured representation to characterize the marketing strategy attributes of activity budget intensity, activity type and duration cycle.
3. The method for establishing a marketing activity indicator prediction model according to claim 2, characterized in that: The dynamic cross-modal information fusion mechanism achieves fine-grained cross-modal semantic alignment and information enhancement through differentiated query key-value pair configurations. The four independent attention channels in the four-channel bidirectional attention mechanism respectively model the asymmetric dependencies between different modalities. The text feature vector is used as the query vector to actively align the visual content to ensure semantic consistency between the activity description and the image theme.
4. The method for establishing a marketing activity indicator prediction model according to claim 3, characterized in that: The four-path bidirectional attention mechanism specifically realizes the interactive fusion between text feature vectors and visual feature vectors, activity feature vectors and visual feature vectors, activity feature vectors and text feature vectors, and text feature vectors and activity feature vectors. First, query key-value pairs are generated through independent projection matrices, and then the attention distribution of feature vectors to other modal contents is calculated through scaled dot product attention.
5. The method for establishing a marketing activity indicator prediction model according to claim 4, characterized in that: The static prediction module specifically realizes the regression prediction of the total index of the marketing activity through a processing chain composed of a batch normalization layer, a self-attention layer and a residual connection. The stacked residual multi-layer perceptron network includes a feature transformation path and a residual direct connection path. The feature transformation path captures nonlinear relationships through a progressive mapping from 512 dimensions to 384 dimensions to 256 dimensions, and the residual direct connection path retains the original information to prevent the gradient from disappearing.
6. The method for establishing a marketing activity indicator prediction model according to claim 5, characterized in that: The multi-scale sliding window specifically uses a parallel multi-scale convolution kernel group to extract time series patterns of different granularities, using a 3-day short kernel to capture sudden promotion fluctuations, a 7-day medium kernel to analyze periodic consumption patterns, and a 30-day long kernel to identify trend billing patterns.
7. The method for establishing a marketing activity indicator prediction model according to claim 6, characterized in that: The temporal evolution branch is specifically based on the sequential pattern capture of the causal Transformer, which uses a causal mask mechanism to constrain the attention scope, ensuring that the prediction time step only accesses historical information, and generates temporal representation by aggregating sequence features and learnable position encoding.
8. The method for establishing a marketing activity indicator prediction model according to claim 7, characterized in that: The cross-modal interaction branch specifically realizes feature enhancement through the interaction between the temporal global context and the multimodal fusion representation, generates a query vector through a pooling operation, generates a key-value pair through the multimodal fusion representation, and outputs the interaction features for subsequent fusion processing.
9. The method for establishing a marketing activity indicator prediction model according to claim 8, characterized in that: The dynamic weight matrix specifically includes elements that reflect the intensity of attention paid to the temporal pattern at different prediction steps, thereby achieving dynamic adjustment of modal importance in the prediction scenario, and generating dynamic weights through interactive calculation of temporal branch projection features and cross-modal interaction features.
10. The method for establishing a marketing activity indicator prediction model according to claim 9, characterized in that: The visual feature vector encoding specifically uses a deep residual network to extract fine-grained spatial features, extracts abstract features step by step through the convolutional layer group of ResNet-50, and obtains the visual feature vector through spatial attention pooling and linear projection.
Citation Information
Patent Citations
Dual-angle marketing activity effect prediction method, medium and system
CN117593044A
Sales volume estimation method and device, electronic equipment, storage medium and program product
CN118396658A
Text mining data query method and system based on cross-modal similarity
CN119311854A
Commercial vitality prediction and business district evaluation method based on multi-modal feature fusion
CN119918981A
Digital marketing system based on big data
CN120494905A
Cited By
Identification method for matching scene behaviors by using multi-modal features
CN121213967A
A recognition method for matching scene behavior by using multi-modal features
CN121213967B
Software test expected result prediction method and system based on deep learning
CN121434100A
Fashion preference prediction method and device
CN121640486A
Multi-mode collaborative cognitive improvement oscillation-induced nerve regulation method and device
CN121730843A