Intelligent marketing decision optimization method based on multi-dimensional data fusion

Through multi-dimensional data fusion and multi-modal dynamic timing embedding technology, combined with lightweight reinforcement learning and multi-objective Bayesian optimization, dynamically adjusting weights and strategies, the problem of difficulty in efficiently synergistic optimization of multi-objectives in the dynamic environment in the existing technology is solved, and efficient and accurate decision-making results are achieved.

CN120012984AInactive Publication Date: 2025-05-16SHENZHEN YOUQIAN INFORMATION TECH CO LTD
View PDF 0 Cites 7 Cited by

Patent Information

Application Number
CN202510066772.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-16
Publication Date
2025-05-16
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing technology is difficult to efficiently coordinate the optimization of multiple goals in dynamic environments, resulting in a decrease in decision-making effect, ineffective computing efficiency and insufficient model accuracy.

Method used

Through the intelligent marketing decision optimization method of multi-dimensional data fusion, multi-modal dynamic timing embedding and lightweight reinforcement learning strategy optimization are adopted, and combined with multi-objective Bayesian optimization and closed-loop optimization processes, the weights and strategies are dynamically adjusted to achieve efficient decision-making.

Benefits of technology

It achieves efficient collaborative optimization of multiple goals in a dynamic environment, improves the real-time and accuracy of decisions, and reduces computing costs and computing power requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120012984A_ABST
    Figure CN120012984A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of big data marketing decision, and discloses an intelligent marketing decision optimization method based on multi-dimensional data fusion. By introducing an attention mechanism and a multi-target Bayesian optimization model, dynamic perception of real-time environment data and accurate adjustment of multi-dimensional weight parameters are realized; a soft maximization function adaptive weight distribution technology is combined, so that key nodes can be efficiently identified in a complex data scene; meanwhile, through the dynamic fusion of the multi-dimensional score function, the precision of the target decision and the robustness of the system are improved. In the system design, a reinforcement learning feedback mechanism and real-time updating logic are further combined, and full-link closed-loop optimization from data input to decision output is realized. Compared with the prior art, the technical scheme of the invention breaks through the limitation of a traditional single model, and through the synergistic effect of multiple links, the decision-making efficiency and the adaptive capacity in a dynamic environment are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of big data marketing decision-making, and in particular to an intelligent marketing decision-making optimization method for multi-dimensional data fusion. Background Art

[0002] In the application scenarios of modern industrial intelligence and dynamic optimization of complex systems, real-time analysis and weight allocation of multi-dimensional data are usually required to support accurate decision-making. Typical applications of dynamic optimization technology include equipment health management, supply chain optimization, and dynamic resource allocation. In these scenarios, the system often faces trade-offs between multiple objectives, such as how to find the optimal solution between performance, cost, and time. However, existing technologies still have many shortcomings in dynamic weight allocation and real-time decision optimization.

[0003] In the prior art, in order to solve the multi-objective optimization problem, the following methods are usually adopted: 1. Static weight allocation method: The industry generally adopts a fixed weight method, and the priority of each goal is set in advance during the system design stage. For example, in supply chain management, the transportation cost is set as a fixed weight. However, this method cannot adapt to changes in the dynamic environment. When external conditions (such as market demand or logistics conditions) change, it is difficult for the system to adjust the optimization plan in real time, resulting in a significant decrease in decision-making effect.

[0004] 2. Single model optimization: Another common approach is to use a single optimization model, such as linear programming or traditional machine learning algorithms, to handle multi-objective problems. Although this approach works well in simple scenarios, as the data dimension and complexity increase, the processing capacity of a single model is limited, and it is prone to low computational efficiency or insufficient model accuracy. For example, in equipment health management, it is difficult for a single model to simultaneously balance aging trend prediction and maintenance cost control.

[0005] 3. Phased optimization strategy: In order to overcome the above problems, some existing technologies adopt a phased processing approach, such as first roughly allocating weights through preset rules and then optimizing using models. However, this approach increases the complexity of system design. At the same time, since the weight adjustment and optimization processes are independent of each other and cannot form a closed loop, the optimization effect is limited and it is difficult to achieve high-precision dynamic decision-making.

[0006] These methods in the prior art have problems such as lack of dynamic response capability to real-time environmental changes, resulting in optimization results lagging behind actual needs; insufficient coordination of system multi-objective optimization, and inability to simultaneously take into account the dynamic adjustment of different objectives; and insufficient computational efficiency and robustness during the optimization process, making it difficult to meet real-time requirements in high-dimensional complex scenarios.

[0007] Therefore, how to efficiently and collaboratively optimize multiple objectives in a dynamic environment and balance the relationship between performance, cost and time has become a technical problem to be solved by the present invention. Summary of the invention

[0008] The present invention provides an intelligent marketing decision optimization method based on multi-dimensional data fusion, the main purpose of which is to solve the problem of efficient collaborative optimization of multiple objectives in a dynamic environment and balance the flexibility of performance, cost and weight distribution.

[0009] To achieve the above object, the present invention provides a multi-dimensional data fusion intelligent marketing decision optimization method, comprising the following steps: Step 1, data collection and processing: Collect marketing-related data from multiple heterogeneous data sources in real time, including: a. Structured data, including user attributes and transaction records; b. Unstructured data, including social media comments and product images; c. Semi-structured data, including web logs; Normalize the collected data and generate an initial feature matrix through principal component analysis and feature dimension reduction method, and the initial feature matrix is ​​used for subsequent data modeling; Step 2: Multimodal dynamic temporal embedding: A dynamic multimodal graph model is constructed based on multimodal dynamic graph embedding technology, in which: a. Data dimensions are represented as graph nodes; b. The associations between different data dimensions are represented as weighted edges of the graph; Extract time series features and perform multi-scale modeling of feature sequences through a temporal convolutional network to capture the dynamic changes of time series; Use the attention mechanism to perform weighted optimization on key nodes and edges in the graph model; Dynamically adjust the weights of nodes and edges. The weights are calculated based on the interaction strength between nodes and their changing trends over time. The specific calculation formula is:

[0010] in: Representation Node and nodes In time The edge weight of Represents the similarity function between node features, calculated using cosine similarity or Euclidean distance; The adjustment coefficient representing the interaction intensity ranges from 0.1 to 1.0 and is set according to the amount of data and the frequency of interaction; Represents the weight normalization base, ranging from 1 to 10, set according to the distribution of connection strength between nodes; Indicates the total number of nodes in the current graph model; Step 3: Lightweight reinforcement learning strategy optimization: Pre-train the reinforcement learning model through meta-learning methods to generate initial strategies suitable for different marketing scenarios; Construct a reward function based on environmental feedback. The reward function is calculated by the weighted sum of the three objectives of conversion rate, user satisfaction, and cost control. The specific weighted formula is:

[0011] in: is the comprehensive reward value; Indicates conversion rate; Indicates user satisfaction; It means cost control; , , is the target weight, provided by the dynamic adjustment module; The updated gradient of the strategy is calculated using the following formula:

[0012] in: represents the objective function of the strategy; Indicates the current status The policy function under represents the immediate reward value, and Consistency; Based on the state The baseline value of is used to reduce the variance of the policy gradient; Step 4: Dynamic adjustment of multi-objective weights: Use multi-objective Bayesian optimization to dynamically adjust the weights of conversion rate, user satisfaction, and cost control; The priority of each target is calculated through environmental feedback and neural network, and the target score function Calculated by linear regression fitting based on current status and historical performance; The target weight is dynamically updated according to the target score. The specific calculation formula is:

[0013] in: Indicates the target In time The weight of Indicates the target In time The scores were fitted by a linear regression model; Indicates the total number of targets; Step 5, closed-loop optimization: The closed-loop optimization process is formed by combining data collection, dynamic modeling, reinforcement learning strategy optimization and dynamic adjustment of target weights, including: The data acquisition module periodically updates the initial feature matrix; The dynamic graph features generated by the dynamic temporal embedding module according to the feature matrix are directly passed to the reinforcement learning module; After the reinforcement learning module outputs the optimization strategy, the target priority is updated in real time through the target weight adjustment module to generate the next step of strategy optimization input.

[0014] As a further solution of the present invention, the weight parameter and The specific range of is [0, 1], and its setting basis is the normalized value of the target node interaction frequency and the node attribute similarity. The node interaction frequency is obtained through the real-time data acquisition module, and the node attribute similarity is obtained by calculating the cosine similarity between the node feature vectors.

[0015] As a further solution of the present invention, the score function The calculation method adopts a linear regression model, the input features of which include the historical number of interactions of the node, the average stay time and the user feedback score, and the output is the current time node The target node score.

[0016] As a further solution of the present invention, the similarity function The calculation method of is cosine similarity calculation, which is defined as:

[0017] in, and Node and No. eigenvalues, is the feature dimension of the node.

[0018] As a further solution of the present invention, the reward function In the calculation of , and The dynamic adjustment is based on the following relationship:

[0019] in, Indicates the target , , Normalized weight value in the most recent feedback cycle.

[0020] As a further solution of the present invention, the dynamic adjustment module uses an attention mechanism to weight node features, and the calculation formula of the attention weight is:

[0021] in, and Respectively represent nodes and The query vector and key vector of Indicates the dimension of the key vector.

[0022] As a further solution of the present invention, the edge weights in the graph model are The dynamic update period is based on the feedback frequency and system processing delay. The calculation formula of the update period is:

[0023] in, is the minimum update cycle, is the number of feedbacks per unit time.

[0024] As a further embodiment of the present invention, the target value , , The method of obtaining is based on the weighted sliding average calculation of real-time user interaction data, where the sliding window size is , the weighting method is:

[0025] As a further solution of the present invention, the reward value of the reinforcement learning module is calculated based on the target achievement rate in the environmental feedback, and the target achievement rate is defined as:

[0026] When the target achievement rate is lower than the preset threshold, the reward value is adjusted according to the proportional penalty.

[0027] As a further solution of the present invention, in the multi-objective Bayesian optimization module, the objective function is constructed based on a weighted summation method, and the expression of the objective function is:

[0028] in, Indicates the target , , The weight of Indicates the score of the corresponding target.

[0029] Compared with the problems described in the background technology, the beneficial effects of the present invention are: 1. By introducing dynamic optimization of multiple parameters into multiple links of weight calculation, scoring function and similarity calculation, a full-link closed-loop logic from input data to output results is formed. The dynamic adjustment of weight parameters combined with real-time environmental feedback realizes multi-dimensional data collaborative optimization, enabling the system to accurately adapt to real-time scene changes based on the interaction frequency, attribute similarity and historical feedback data of different nodes. This technical feature not only enhances the system's ability to adapt to dynamic changes in the environment, but also avoids the performance limitations caused by the static setting of a single parameter in the existing technology.

[0030] 2. Through the multimodal dynamic time series embedding model, combined with the time series convolutional network (TCN) and the dynamic multimodal graph neural network, it can effectively capture the dynamic interaction relationship and time series changes between heterogeneous data; real-time dynamic update: the dynamic adjustment mechanism of weights between time nodes is used to enhance the adaptability of the model to real-time data and improve the accuracy and timeliness of data fusion; improve data processing efficiency: dynamically adjust weights and graph structures to reduce the interference of invalid interactions on calculations, while maintaining high precision, optimizing processing speed, and at the same time, by introducing meta-learning technology and using the "strategy pre-training + fast adaptation" mechanism, the training overhead of reinforcement learning in dynamic scenarios is reduced, significantly reducing computing power costs.

[0031] 3. The attention mechanism is used to perform weighted calculations on node features and interaction frequencies, and the weight distribution between nodes is dynamically adjusted through the soft maximization function, effectively highlighting the key nodes with high correlation. This mechanism optimizes the traditional static weight distribution into a dynamic adaptive distribution, enabling the system to automatically identify key nodes in complex data scenarios, improving the robustness of the overall performance and the ability to adapt to scenarios; and the multi-objective Bayesian optimization module realizes the fusion of multiple scoring dimensions through weighted summation. The weight of each target score is dynamically adjusted according to real-time environmental data, effectively balancing the conflicts between different optimization objectives.

[0032] 4. The introduction of dynamic update cycle automatically calculates the minimum update cycle and the number of feedbacks to ensure the synchronization of the update frequency of edge weights with environmental changes. Combined with the feedback mechanism of reinforcement learning, the system can correct the path optimization decision in real time during the edge weight adjustment, further improving the adaptability of the solution to complex dynamic environments; the design of the reward function strengthens the coupling between the target achievement rate and the system performance by combining multiple weight factors with environmental feedback. The weighted sliding average calculation method realizes the comprehensive optimization of historical data and real-time data by introducing time weights, and can extract the optimal decision path from short-term and long-term data; and the node similarity function based on cosine similarity realizes the accurate measurement of node associations in high-dimensional feature space by combining attribute similarity with interaction frequency, and significantly reduces the redundant data noise in complex networks through multi-dimensional similarity calculation, and improves the specificity and stability of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] Figure 1 This is a flow chart of the intelligent marketing decision optimization method based on multi-dimensional data fusion of the present invention.

[0034] Figure 2 This is a diagram of the multimodal dynamic graph modeling architecture of the present invention.

[0035] The purpose, features and advantages of the present invention will be further described with reference to the accompanying drawings in conjunction with the embodiments. DETAILED DESCRIPTION

[0036] It should be understood that the specific embodiments described herein are only used to explain the present invention, and are not used to limit the present invention.

[0037] The present application embodiment provides a multi-dimensional data fusion intelligent marketing decision optimization method, comprising the following steps: Step 1, data collection and processing: Collect marketing-related data in real time from multiple heterogeneous data sources, including: a. Structured data, including user attributes and transaction records; b. Unstructured data, including social media comments and product images; c. Semi-structured data, including web logs; The collected data are normalized, and the initial feature matrix is ​​generated through principal component analysis and feature dimension reduction method. The initial feature matrix is ​​used for subsequent data modeling.

[0038] Step 2: Multimodal dynamic temporal embedding: A dynamic multimodal graph model is constructed based on multimodal dynamic graph embedding technology. In the graph model: a. Data dimensions are represented as graph nodes; b. The associations between different data dimensions are represented as weighted edges of the graph; Extract time series features and perform multi-scale modeling of feature sequences through a temporal convolutional network to capture the dynamic changes of time series; Use the attention mechanism to perform weighted optimization on key nodes and edges in the graph model; Dynamically adjust the weights of nodes and edges. The weights are calculated based on the interaction strength between nodes and their changing trends over time. The specific calculation formula is:

[0039] in: Representation Node and nodes In time The edge weight of Represents the similarity function between node features, calculated using cosine similarity or Euclidean distance; The adjustment coefficient representing the interaction intensity ranges from 0.1 to 1.0 and is set according to the amount of data and the frequency of interaction; Represents the weight normalization base, ranging from 1 to 10, set according to the distribution of connection strength between nodes; Indicates the total number of nodes in the current graph model.

[0040] Step 3: Lightweight reinforcement learning strategy optimization: Pre-train the reinforcement learning model through meta-learning methods to generate initial strategies suitable for different marketing scenarios; Construct a reward function based on environmental feedback. The reward function is calculated by the weighted sum of the three objectives of conversion rate, user satisfaction, and cost control. The specific weighted formula is:

[0041] in: is the comprehensive reward value; Indicates conversion rate; Indicates user satisfaction; It means cost control; , , is the target weight, provided by the dynamic adjustment module; The updated gradient of the strategy is calculated using the following formula:

[0042] in: represents the objective function of the strategy; Indicates the current status The policy function under represents the immediate reward value, and Consistency; Based on the state The baseline value of is used to reduce the variance of the policy gradient.

[0043] Step 4: Dynamic adjustment of multi-objective weights: Use multi-objective Bayesian optimization to dynamically adjust the weights of conversion rate, user satisfaction, and cost control; The priority of each target is calculated through environmental feedback and neural network, and the target score function Calculated by linear regression fitting based on current status and historical performance; The target weight is dynamically updated according to the target score. The specific calculation formula is:

[0044] in: Indicates the target In time The weight of Indicates the target In time The scores were fitted by a linear regression model; Indicates the total number of targets.

[0045] Step 5, closed-loop optimization: The closed-loop optimization process is formed by combining data collection, dynamic modeling, reinforcement learning strategy optimization and dynamic adjustment of target weights, including: The data acquisition module periodically updates the initial feature matrix; The dynamic graph features generated by the dynamic temporal embedding module according to the feature matrix are directly passed to the reinforcement learning module; After the reinforcement learning module outputs the optimization strategy, the target priority is updated in real time through the target weight adjustment module to generate the next step of strategy optimization input.

[0046] Weight Parameters and The specific range of is [0, 1], and its setting basis is the normalized value of the target node interaction frequency and the node attribute similarity. The node interaction frequency is obtained through the real-time data acquisition module, and the node attribute similarity is obtained by calculating the cosine similarity between the node feature vectors.

[0047] Score function The calculation method uses a linear regression model. The input features of the linear regression model include the historical number of interactions of the node, the average stay time and the user feedback score. The output is the current time node The target node score of . Similarity function The calculation method of is cosine similarity calculation, which is defined as:

[0048] in, and Node and No. eigenvalues, is the feature dimension of the node.

[0049] Reward Function In the calculation of , and The dynamic adjustment is based on the following relationship:

[0050] in, Indicates the target , , Normalized weight value in the most recent feedback cycle.

[0051] The dynamic adjustment module uses the attention mechanism to weight node features. The calculation formula of attention weight is:

[0052] in, and Respectively represent nodes and The query vector and key vector of Indicates the dimension of the key vector.

[0053] Edge weights in graph models The dynamic update cycle is based on the feedback frequency and system processing delay. The calculation formula for the update cycle is:

[0054] in, is the minimum update cycle, is the number of feedbacks per unit time.

[0055] Target value , , The method of obtaining is based on the weighted sliding average calculation of real-time user interaction data, where the sliding window size is , the weighting method is:

[0056] The reward value of the reinforcement learning module is calculated based on the target achievement rate in the environmental feedback. The target achievement rate is defined as:

[0057] When the target achievement rate is lower than the preset threshold, the reward value is adjusted according to the proportional penalty.

[0058] In the multi-objective Bayesian optimization module, the objective function is constructed based on the weighted summation method, and the expression of the objective function is:

[0059] in, Indicates the target , , The weight of Indicates the score of the corresponding target.

[0060] Embodiment 1: This embodiment more clearly demonstrates the specific implementation path of the technical solution of the present invention.

[0061] 1. Data collection and feature processing.

[0062] In this path, the data collection module is implemented in such a way that the system collects relevant marketing data in real time through a variety of heterogeneous data sources. Specifically, the collected data includes: a. Structured data: such as the user's age, gender, consumption records, etc. b. Unstructured data: such as the text content of social media comments, product images, etc. c. Semi-structured data: such as log files of users visiting web pages. Various types of data are transmitted to the central database through standardized interfaces (such as RESTful API or Kafka message queues) and stored in distributed storage systems that support large-scale parallel processing (such as HDFS or Amazon S3).

[0063] Then feature extraction and preprocessing are performed, that is, after data collection is completed, the dimension differences between different data dimensions are reduced through normalization. The specific method is:

[0064] in, is the original eigenvalue, and Represent the minimum and maximum values ​​of the feature respectively; use principal component analysis (PCA) to reduce the dimension of the feature and retain 95% of the cumulative variance contribution rate to form the initial feature matrix.

[0065] This is followed by data cleaning, where the system uses automated cleaning tools (such as OpenRefine or the Pandas library) to remove redundant and noisy data while ensuring data integrity and consistency.

[0066] 2. Multimodal dynamic graph modeling.

[0067] In this step, we first construct a multimodal graph, that is, map the collected multidimensional data into a dynamic multimodal graph model. The graph is constructed in the following ways: a. Each data dimension is represented as a graph node; b. The relationship between different dimensions is represented by weighted edges, and the weight of the edge is calculated by the following formula:

[0068] in: :node With Node In time The edge weight of : The similarity function between node features can be calculated using cosine similarity:

[0069] : The adjustment coefficient of interaction strength, ranging from 0.1 to 1.0; : Normalized base number, ranging from 1 to 10.

[0070] Then, time series feature modeling is performed, that is, the time series features of the graph nodes are extracted using the temporal convolutional network (TCN). The core number of the model is set to 3 layers, and the convolution kernel size of each layer is 5 to capture the dynamic change pattern of the time series.

[0071] Then the attention mechanism is optimized. That is, the system uses the attention mechanism to dynamically optimize the key nodes and edge weights in the graph model. The specific formula is:

[0072] in: and :respectively nodes and nodes The query vector and key vector of ; : The dimension of the key vector, used to normalize the weight distribution.

[0073] 3. Reinforcement learning strategy optimization.

[0074] In this step, the initial strategy is first generated, that is, the reinforcement learning model is pre-trained through the meta-learning method, and the initial strategy suitable for different marketing scenarios is generated by combining historical marketing data.

[0075] Next, the definition of the reward function constructs a multi-objective reward function with conversion rate, user satisfaction and cost control as the goals:

[0076] in: : Conversion rate, which indicates the proportion of final purchasing users to the total visiting users; : User satisfaction, calculated through real-time user feedback ratings; : Cost control, considering the advertising cost per click; , , : are the weight parameters of each objective, which are optimized through the dynamic adjustment module.

[0077] As well as updating the strategy, the reinforcement learning module dynamically updates the strategy gradient through the following formula:

[0078] in: : Current strategy; : Instant reward, with Consistency; : Based on status The baseline value of is used to reduce the variance of the policy gradient.

[0079] 4. Dynamic weight adjustment.

[0080] In this step, the dynamic adjustment logic of weights, such as the adjustment of target weights, is implemented through multi-objective Bayesian optimization, and the optimization formula is:

[0081] in: :Target In time The scores were calculated by linear regression fitting; : Total number of targets.

[0082] Then, closed-loop optimization is carried out. The system forms a closed loop between data collection, dynamic graph modeling, strategy optimization and weight adjustment, and optimizes the entire process through real-time environmental feedback.

[0083] 5. Verification of implementation effect.

[0084] In actual test scenarios, such as the marketing campaign on an e-commerce platform, this solution was implemented in a test environment with 1 million visits per day and a variety of heterogeneous data (such as click logs, user comments, product images, etc.); compared with traditional methods (fixed weight optimization, single model optimization), the target conversion rate increased by 18.5%, the cost decreased by 12.3%, and the user satisfaction score increased by 15%.

[0085] Embodiment 2: This embodiment is based on the existing multi-dimensional data fusion intelligent marketing decision optimization method, through the specific details of data collection, feature processing, dynamic modeling, weight adjustment, and reward mechanism.

[0086] Step 1: Data collection and processing. First, the system collects marketing-related data in real time from multiple heterogeneous data sources (such as user attributes, transaction records, social media comments, product images, etc.). The data types include structured data, unstructured data, and semi-structured data.

[0087] Data preprocessing: All collected data are first normalized. The specific method is:

[0088] in, is the original eigenvalue, and Represent the minimum and maximum values ​​of the feature respectively. is the standardized data.

[0089] Feature dimensionality reduction: Use principal component analysis (PCA) to reduce the dimension of features, retain 95% of the cumulative variance contribution rate, and generate the initial feature matrix.

[0090] Step 2: Multimodal dynamic graph modeling.

[0091] In this step, the system uses multimodal dynamic graph embedding technology to build a dynamic multimodal graph model. The construction of the graph model includes the following: Node representation: Each data dimension (such as user behavior, product features, etc.) is represented as a node in the graph.

[0092] Edge representation: The relationship between different data dimensions is represented by weighted edges. The edge weight calculation formula is:

[0093] in, Indicates at time node and nodes The edge weights between is the adjustment coefficient of interaction strength, is the normalized base, Represents the similarity function between node features, calculated using cosine similarity or Euclidean distance.

[0094] Similarity function: The similarity between node features is calculated by cosine similarity, which is defined as:

[0095] in, and Respectively represent nodes and nodes In the The value of a feature, is the dimension of the feature.

[0096] Step 3: Reinforcement learning strategy optimization. In the reinforcement learning module, the reinforcement learning model is pre-trained through the meta-learning method to generate initial strategies suitable for different marketing scenarios. The reward function is calculated by the weighted sum of the following objectives:

[0097] in, represents the conversion rate, Indicates user satisfaction, Indicates cost control, , , are the weights of the targets, respectively, provided by the dynamic adjustment module.

[0098] Policy update: Use the policy gradient descent method to update the optimization strategy. The formula is:

[0099] in, represents the current policy function, For instant rewards, is the baseline value.

[0100] Step 4: Dynamic adjustment of multi-objective weights.

[0101] Use multi-objective Bayesian optimization method to dynamically adjust the weights of conversion rate, user satisfaction and cost control. The specific formula is:

[0102] in, For time Target The weight of is the target score, calculated by linear regression model fitting, is the total number of targets.

[0103] Target score function: Target score function Calculated using a linear regression model, based on current status and historical performance.

[0104] Step 5: Closed-loop optimization. The closed-loop optimization process is formed by combining data collection, dynamic modeling, reinforcement learning strategy optimization, and dynamic adjustment of target weights in the following ways: Data acquisition module: Regularly update the initial feature matrix to ensure the timeliness of the data.

[0105] Dynamic modeling module: Generates dynamic graph features based on the feature matrix and passes them to the reinforcement learning module.

[0106] Reinforcement learning module: After outputting the optimization strategy, the target priority is updated in real time through the target weight adjustment module to generate the next step of strategy optimization input.

[0107] Step 6: Dynamic update and optimization.

[0108] Based on real-time feedback, the weights and strategies are adjusted dynamically to ensure that the entire system can respond to market changes and user needs in a timely manner. During the feedback cycle, the dynamic update cycle of the weights is based on the minimum update cycle. and feedback frequency Calculate the update cycle formula:

[0109] in, is the minimum update cycle, is the number of feedbacks per unit time.

[0110] Embodiment 3: This embodiment collects marketing-related data from multiple heterogeneous data sources in real time, including structured data (such as user attributes and transaction records), unstructured data (such as social media comments and product images), and semi-structured data (such as web logs). The data collection module uses a standardized interface (such as RESTful API) to send the data to a central database and store it in a distributed storage system (such as HDFS).

[0111] The data preprocessing steps include normalization and principal component analysis (PCA) dimensionality reduction. Normalization is achieved through the following formula:

[0112] in, is the original eigenvalue, and Represent the minimum and maximum values ​​of the feature respectively. is the normalized eigenvalue.

[0113] Next, PCA was used to reduce the dimension of the data, retaining 95% of the cumulative variance contribution rate and generating the initial feature matrix, which was used for subsequent modeling and analysis.

[0114] Furthermore, this embodiment constructs a dynamic multimodal graph model through multimodal dynamic graph embedding technology. Each data dimension is represented as a node in the graph, and the relationship between different data dimensions is represented by weighted edges. The weight of the edge is calculated based on the feature similarity between nodes, and the specific calculation formula is:

[0115] in, For Node With Node In time The edge weight of Representation Node With Node The similarity between features is calculated using cosine similarity or Euclidean distance; is the interaction intensity adjustment coefficient, is the weight normalization base, is the total number of nodes in the graph model.

[0116] The similarity function between nodes is calculated by cosine similarity as follows:

[0117] in, and Node and nodes In the The value of the feature dimension, is the dimension of the feature.

[0118] In addition, the meta-learning method can be used to pre-train the reinforcement learning model to generate initial strategies suitable for different marketing scenarios. The reward function of reinforcement learning is calculated by weighting the three objectives of conversion rate (CR), customer satisfaction (CS) and cost control (CC). The specific formula is:

[0119] in, is the comprehensive reward value, , and is the target weight, provided by the dynamic adjustment module. The following policy gradient update formula is used to optimize the strategy:

[0120] in, represents the policy objective function, Indicates the current status The policy function is is the immediate reward value, is the baseline value.

[0121] Multi-objective Bayesian optimization is used to dynamically adjust the weights of conversion rate, user satisfaction, and cost control. The priority of each goal is calculated through environmental feedback and neural networks, and the goal score function Through linear regression model fitting calculation, the target weight update formula is:

[0122] in, For the goal In time The weight of Indicates the target In time The score, is the total number of targets.

[0123] The entire system continuously optimizes the decision-making process through a closed-loop optimization process. Specifically, the data acquisition module regularly updates the initial feature matrix, and the dynamic time series embedding module generates dynamic graph features based on the matrix and passes them to the reinforcement learning module. After the reinforcement learning module outputs the optimization strategy, the target weight adjustment module updates the target priority in real time and generates the next step of strategy optimization input.

[0124] On the basis of real-time feedback, the dynamic update cycle of edge weight is calculated based on the feedback frequency and system processing delay. The formula of the update cycle is:

[0125] in, is the minimum update cycle, is the number of feedbacks per unit time.

[0126] In actual application, the intelligent marketing system implementing this solution was verified in the marketing delivery of e-commerce platforms. The test environment includes 1 million visits per day, covering structured, unstructured and semi-structured data. Compared with traditional methods, the system has significantly improved indicators such as conversion rate, user satisfaction and cost control. The specific data are as follows: conversion rate increased by 18.5%; user satisfaction score increased by 15%; advertising delivery cost decreased by 12.3%.

[0127] Embodiment 4: This embodiment is further optimized on the basis of Embodiment 3, aiming to further refine the variable definition, optimize the dynamic adjustment mechanism and clarify the source of the formula.

[0128] In data collection and feature processing, the system first collects marketing-related data from multiple heterogeneous data sources in real time, including the following three types of data: structured data: such as user attributes and transaction records. Unstructured data: such as social media comments and product images. Semi-structured data: such as web logs. After data collection, all data is normalized using the following formula:

[0129] in, is the original eigenvalue, and are the minimum and maximum values ​​of the feature, respectively. is the normalized eigenvalue. Normalization ensures that different data dimensions are consistent in the subsequent modeling process. Subsequently, principal component analysis (PCA) is used to reduce the dimension of the features, retaining 95% of the cumulative variance contribution rate to form the initial feature matrix for dynamic model construction.

[0130] In multimodal dynamic graph modeling, the system builds a dynamic multimodal graph model to capture the complex relationships and dynamic changes between data. Its node representation: data dimensions (such as user behavior, product features, etc.) are represented as nodes in the graph; edge weight calculation: the following formula is used to dynamically calculate the association strength between nodes:

[0131] in, For time node and nodes The edge weights between is the interaction strength adjustment coefficient (ranging from 0.1 to 1.0), is the normalized base (ranging from 1 to 10), is the node feature similarity function, calculated using cosine similarity:

[0132] in, and Node and nodes In the The value of the dimension feature, is the feature dimension.

[0133] In the process of optimizing the reinforcement learning strategy, the meta-learning module generates the initial strategy through the meta-learning method and dynamically optimizes the following objectives, such as conversion rate ( ): The proportion of final purchasing users to total visiting users; User satisfaction ( ): Calculated through real-time user feedback rating; cost control ( ): The advertising cost per click; its reward function is defined as:

[0134] in, , and For dynamic weights, use the following formula to update:

[0135] in, For the goal In time The priority score of is fitted by a linear regression model, is the total number of targets.

[0136] The reinforcement learning module dynamically updates the strategy through the following policy gradient formula:

[0137] in, is the current policy function, For immediate rewards, Consistent, is the baseline value used to reduce the variance of the policy gradient.

[0138] In the closed-loop optimization stage, a closed-loop optimization process is formed by data collection, dynamic modeling, reinforcement learning strategy optimization and dynamic adjustment of target weights. Specifically, the data collection module regularly updates the initial feature matrix; the dynamic modeling module generates dynamic graph features and passes them to the reinforcement learning module; after the reinforcement learning module outputs the optimization strategy, the target weight adjustment module updates the priority in real time to form the next step of strategy optimization input.

[0139] Embodiment 5: This example uses the e-commerce marketing scenario as the background to demonstrate how to build a dynamic optimization process based on multimodal data. Its technical implementation path includes data collection and preprocessing, that is, real-time collection of marketing-related data from multiple heterogeneous data sources, including: Structured data: user attributes (such as age, gender) and transaction records (such as purchase frequency and amount); unstructured data: social media comment text, product pictures uploaded by users; semi-structured data: user access behavior logs (such as page dwell time); data is transmitted to the central database through standardized interfaces (such as Kafka message queues) and stored using distributed storage systems (such as HDFS). The following processing is performed on the collected data: Normalization: in, is the original data value, , are the minimum and maximum values ​​of this feature respectively.

[0140] Feature dimensionality reduction: Use principal component analysis (PCA) to retain 95% of the cumulative variance contribution and generate the initial feature matrix.

[0141] A dynamic multimodal graph model is constructed through multimodal dynamic graph embedding technology, including node and edge representation: each data type is mapped to a graph node, and the relationship between nodes is represented as a weighted edge. The edge weight calculation formula is as follows:

[0142] :node and In time The edge weight of : Cosine similarity function, defined as:

[0143] : Interaction intensity adjustment coefficient, ranging from 0.1 to 1.0, set according to the interaction frequency; : Weight normalization base, ranging from 1 to 10, is set according to the distribution of connection strength between nodes.

[0144] The reinforcement learning module uses meta-learning method for pre-training to generate initial strategies suitable for multiple scenarios. The specific implementation steps are as follows, and its reward function is defined as:

[0145] : Conversion rate, which indicates the proportion of users who complete a purchase to the total number of visiting users; : User satisfaction, calculated based on feedback ratings; : Advertising costs; , , : Dynamically adjusted target weight; its strategy update formula:

[0146] in, represents the policy function, For immediate rewards, Consistent, is the baseline value, used to reduce the gradient variance.

[0147] In addition, in the dynamic adjustment of multi-objective weights, the multi-objective Bayesian optimization can be combined to dynamically adjust the weights. The formula is as follows:

[0148] :Target The priority score is fitted by linear regression, and the input features include the number of historical interactions and user feedback.

[0149] This embodiment implements the following closed-loop process: the data acquisition module regularly updates the initial feature matrix; the dynamic modeling module generates dynamic graph features and passes them to the reinforcement learning module; the reinforcement learning module outputs the optimization strategy, updates the target weight in real time, and generates the next step strategy input.

[0150] It is obvious to those skilled in the art that the present invention is not limited to the details of the above exemplary embodiments, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention.

[0151] Finally, it should be noted that the above embodiments are only used to illustrate the technical solution of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solution of the present invention can be modified or replaced by equivalents without departing from the spirit and scope of the technical solution of the present invention.

Claims

1. An intelligent marketing decision optimization method based on multi-dimensional data fusion, characterized in that: The following steps are involved: Step 1, data collection and processing: Collect marketing-related data from multiple heterogeneous data sources in real time, including: Structured data, including user attributes and transaction records; Unstructured data, including social media comments and product images; Semi-structured data, including web logs; Normalize the collected data and generate an initial feature matrix through principal component analysis and feature dimension reduction method, and the initial feature matrix is ​​used for subsequent data modeling; Step 2: Multimodal dynamic temporal embedding: A dynamic multimodal graph model is constructed based on multimodal dynamic graph embedding technology, in which: Data dimensions are represented as graph nodes; The associations between different data dimensions are represented as weighted edges of the graph; Extract time series features and perform multi-scale modeling of feature sequences through a temporal convolutional network to capture the dynamic changes of time series; Use the attention mechanism to perform weighted optimization on key nodes and edges in the graph model; Dynamically adjust the weights of nodes and edges. The weights are calculated based on the interaction strength between nodes and their changing trends over time. The specific calculation formula is: in: Representation Node and nodes In time The edge weight of Represents the similarity function between node features, calculated using cosine similarity or Euclidean distance; The adjustment coefficient representing the interaction intensity ranges from 0.1 to 1.0 and is set according to the amount of data and the frequency of interaction; Represents the weight normalization base, ranging from 1 to 10, set according to the distribution of connection strength between nodes; Indicates the total number of nodes in the current graph model; Step 3: Lightweight reinforcement learning strategy optimization: Pre-train the reinforcement learning model through meta-learning methods to generate initial strategies suitable for different marketing scenarios; Construct a reward function based on environmental feedback. The reward function is calculated by the weighted sum of the three objectives of conversion rate, user satisfaction, and cost control. The specific weighted formula is: in: is the comprehensive reward value; Indicates conversion rate; Indicates user satisfaction; It means cost control; , , is the target weight, provided by the dynamic adjustment module; The updated gradient of the strategy is calculated using the following formula: in: represents the objective function of the strategy; Indicates the current status The policy function under represents the immediate reward value, and Consistency; Based on the state The baseline value of is used to reduce the variance of the policy gradient; Step 4: Dynamic adjustment of multi-objective weights: Use multi-objective Bayesian optimization to dynamically adjust the weights of conversion rate, user satisfaction, and cost control; The priority of each target is calculated through environmental feedback and neural network, and the target score function Calculated by linear regression fitting based on current status and historical performance; The target weight is dynamically updated according to the target score. The specific calculation formula is: in: Indicates the target In time The weight of Indicates the target In time The scores were fitted by a linear regression model; Indicates the total number of targets; Step 5, closed-loop optimization: The closed-loop optimization process is formed by combining data collection, dynamic modeling, reinforcement learning strategy optimization and dynamic adjustment of target weights, including: The data acquisition module periodically updates the initial feature matrix; The dynamic graph features generated by the dynamic temporal embedding module according to the feature matrix are directly passed to the reinforcement learning module; After the reinforcement learning module outputs the optimization strategy, the target priority is updated in real time through the target weight adjustment module to generate the next step of strategy optimization input.

2. The intelligent marketing decision optimization method based on multi-dimensional data fusion according to claim 1, characterized in that: The weight parameter and The specific range of is [0, 1], and its setting basis is the normalized value of the target node interaction frequency and the node attribute similarity. The node interaction frequency is obtained through the real-time data acquisition module, and the node attribute similarity is obtained by calculating the cosine similarity between the node feature vectors.

3. The intelligent marketing decision optimization method based on multi-dimensional data fusion according to claim 1, characterized in that: The score function The calculation method adopts a linear regression model, the input features of which include the historical number of interactions of the node, the average stay time and the user feedback score, and the output is the current time node The target node score.

4. The intelligent marketing decision optimization method based on multi-dimensional data fusion according to claim 1, characterized in that: The similarity function The calculation method of is cosine similarity calculation, which is defined as: in, and Node and No. eigenvalues, is the feature dimension of the node.

5. The intelligent marketing decision optimization method based on multi-dimensional data fusion according to claim 1, characterized in that: The reward function In the calculation of , and The dynamic adjustment is based on the following relationship: in, Indicates the target , , Normalized weight value in the most recent feedback cycle.

6. The intelligent marketing decision optimization method based on multi-dimensional data fusion according to claim 1, characterized in that: The dynamic adjustment module uses the attention mechanism to weight the node features, and the calculation formula of the attention weight is: in, and Respectively represent nodes and The query vector and key vector of Indicates the dimension of the key vector.

7. The intelligent marketing decision optimization method based on multi-dimensional data fusion according to claim 1, characterized in that: The edge weights in the graph model The dynamic update period is based on the feedback frequency and system processing delay. The calculation formula of the update period is: in, is the minimum update cycle, is the number of feedbacks per unit time.

8. The intelligent marketing decision optimization method based on multi-dimensional data fusion according to claim 1, characterized in that: The target value , , The method of obtaining is based on the weighted sliding average calculation of real-time user interaction data, where the sliding window size is , the weighting method is: 。 9. The intelligent marketing decision optimization method based on multi-dimensional data fusion according to claim 1, characterized in that: The reward value of the reinforcement learning module is calculated based on the target achievement rate in the environmental feedback. The target achievement rate is defined as: When the target achievement rate is lower than the preset threshold, the reward value is adjusted according to the proportional penalty.

10. The intelligent marketing decision optimization method based on multi-dimensional data fusion according to claim 1, characterized in that: In the multi-objective Bayesian optimization module, the objective function is constructed based on the weighted summation method, and the expression of the objective function is: in, Indicates the target , , The weight of Indicates the score of the corresponding target.

Citation Information

Cited By

  • Multi-contact marketing effect attribution method based on AI big data fusion

    CN120338856A

  • Digitized marketing probability-initiated split marketing method and system

    CN120612127A

  • Logistics scheduling planning method and system based on graph neural network and reinforcement learning

    CN120806801A

  • Logistics scheduling planning method and system based on graph neural network and reinforcement learning

    CN120806801B

  • Marketing intelligent decision-making system based on multi-modal learning

    CN120975862A