A comprehensive intelligent marketing method based on multimodal information fusion
By combining multimodal deep learning and reinforcement learning models with game theory, multimodal data fusion of the full-domain intelligent marketing system was achieved, solving the problems of user intent recognition bias and dynamic bidding, and improving advertising efficiency and business negotiation effectiveness.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- JUDAOLAIKE (SHANDONG) BIG DATA SERVICE CO LTD
- Filing Date
- 2025-12-22
- Publication Date
- 2026-08-04
AI Technical Summary
Existing marketing systems are unable to effectively establish semantic mapping between heterogeneous data, leading to biases in user intent recognition. They also lack dynamic bidding mechanisms based on long-term returns and automated game negotiation strategies, resulting in low return on investment for advertising and low business conversion efficiency.
By employing multimodal deep learning, reinforcement learning, and game theory models, and through multimodal information fusion, a comprehensive intelligent marketing method is constructed to achieve user intent recognition, precise ad delivery, and automated business negotiation.
It improved the accuracy of user intent recognition, optimized advertising strategies, increased business conversion rates and inquiry conversion efficiency, and reduced manual operation costs.
Smart Images

Figure CN121746014B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of internet data processing and artificial intelligence technology, specifically to a comprehensive intelligent marketing method based on multimodal information fusion. Background Technology
[0002] With the rapid development of cross-border e-commerce and digital marketing technologies, businesses are increasingly diversifying their customer acquisition channels, generating massive amounts of user interaction data. This data exhibits significant multimodal characteristics, encompassing traditional text logs, visual information from product images, and temporal behavioral data related to user browsing patterns. Existing marketing systems typically process this data using single-modal analysis or simple feature stitching. This approach often overlooks the deep semantic relationships between different modalities, such as failing to effectively establish a logical mapping between the visual content of video ads and user text comments. This leads to an inaccurate understanding of users' potential needs, consequently affecting the accuracy of user profiling and the reliability of intent recognition.
[0003] In the ad placement decision-making process, most existing real-time bidding (RTB) systems rely on rule-based static strategies or linear bidding models based solely on estimated click-through rates (CTR). However, the actual bidding environment is highly dynamic and uncertain, with market traffic prices and competition intensity fluctuating over time. Traditional bidding methods lack a mechanism to balance long-term returns with immediate costs, making it difficult to adjust bidding strategies in real time based on remaining budget and campaign progress. This can easily lead to problems such as premature budget exhaustion or insufficient campaign effort, resulting in an inability to achieve optimal return on investment (ROI).
[0004] Furthermore, in the lead conversion phase, current customer service systems primarily rely on human agents or automated responses based on keyword matching. When faced with complex business scenarios involving price negotiations and trade terms, existing systems lack dynamic interactive capabilities based on game theory strategies. They cannot flexibly adjust pricing strategies according to the urgency of the customer's decision and profit targets, nor can they automatically find alternative benefit-sharing solutions when prices are deadlocked. This not only increases the company's human resource operating costs but also limits the potential for improving inquiry conversion rates and average order value. Therefore, there is an urgent need for a comprehensive marketing solution that can integrate multimodal data, possess adaptive targeting capabilities, and execute intelligent business negotiations. Summary of the Invention
[0005] To address the shortcomings of existing technologies, this invention provides a comprehensive intelligent marketing method based on multimodal information fusion. This method solves the problems of existing marketing systems, such as biased user intent recognition due to the inability to effectively establish semantic mapping between heterogeneous data, imbalance between advertising budget control and conversion effect due to the lack of a dynamic bidding mechanism based on long-term returns, and inefficiency in the business conversion process due to the lack of automated game negotiation strategies.
[0006] To achieve the above objectives, the present invention provides the following technical solution: The first aspect of this invention provides a comprehensive intelligent marketing method based on multimodal information fusion. This method integrates multimodal deep learning, reinforcement learning, and game theory models to execute a processing flow from user intent recognition and precise ad targeting to automated business negotiations. The method specifically includes the following implementation steps:
[0007] During the data access and preprocessing phase, the system establishes API connection channels between the external interface module and the target social media and e-commerce platforms, and initializes real-time data stream transmission services. The collected raw data, including user interaction behavior, platform content, and environmental parameters, is cleaned and standardized. Separate preprocessing pipelines are constructed for text, visual, and behavioral time-series data to ensure that the input data meets the model training distribution requirements.
[0008] In the multimodal feature extraction and fusion stage, the core computing engine module performs deep feature extraction on text, visual, and behavioral data respectively. For text data, the Transformer architecture is used to compute the global semantic representation of the classification identifier position; for visual data, a deep residual network is used to extract convolutional feature maps and perform global average pooling; for behavioral sequences, a long short-term memory network is used to extract temporal dependency features. The system constructs a multimodal feature joint matrix and introduces a modality-level attention module to learn the global context vector. Calculate the first Attention score for each modality : ;
[0009] in, For the first Feature vectors of each modality and These are the parameters. The normalized weight coefficients are then calculated using the Softmax function. Based on this weight, a fusion feature vector is generated, and it is projected onto the common semantic space through a multilayer perceptron to establish a unified geometric representation.
[0010] In the holographic user profile construction and intent recognition stage, a high-dimensional sparse user vector containing inherent attributes, long-term interests, and short-term behaviors is generated based on fused features. A time-decay accumulation algorithm is used to dynamically update the user profile vector at the current time step. The calculation is as follows: ; in, For memory retention coefficient, For the new feature vector, The time interval is defined as [time interval]. Based on this dynamic profile, an attention classification network is used to calculate the conditional probability of a user belonging to different intent stages, thereby identifying the user's intent state.
[0011] In the intelligent traffic generation and execution phase, a dual-tower deep neural network is used to predict click-through rate, constructing a real-time bidding model based on reinforcement learning. A Markov decision process is established, defining the state space as including remaining budget, campaign duration, and estimated conversion rate, and the action space as continuous bid adjustment coefficients. Under the constraint of maximizing total conversion value and controlling costs, a reward function is defined. : ; in, For conversion identifiers, For order value, To incur costs, The balancing coefficient is used. A deep Q-network is used to fit the Q-value function, and the strategy parameters are updated by minimizing the time difference error to output the optimal bid order. A PID controller is then used to adjust the budget consumption rate.
[0012] During the intelligent interaction and business negotiation phase, quotations are generated or automated negotiations are conducted based on customer inquiries. For quotation generation, a base price is calculated by considering production costs, logistics costs, and exchange rate risks, and a dynamic pricing coefficient is determined based on the strength of the user's purchasing intent. For business negotiations, a game theory model based on time-dependent strategies is introduced. The game resilience coefficient is dynamically adjusted according to the urgency of the customer's decision. and in the Calculating the proposed offer during the round of negotiations : ; in, For the expected quote, As the bottom price, This is the maximum number of rounds. When price negotiations fail to reach an agreement, a multi-dimensional game mechanism is activated to search for Pareto improvement solutions and generate negotiation responses by utilizing the exchange of benefits in non-price clauses.
[0013] During the feedback and model optimization phase, an online incremental learning method was employed, using streaming data to update model parameters in real time. The Thompson sampling algorithm was introduced for traffic allocation, and conversion rates of advertising creatives were sampled based on a Beta distribution. Furthermore, model distillation techniques using teacher-student networks were employed to transfer the inference capabilities of the high-precision model to a lightweight model.
[0014] A second aspect of the present invention provides a comprehensive intelligent marketing system based on multimodal information fusion, the system comprising a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the comprehensive intelligent marketing method based on multimodal information fusion described in the first aspect.
[0015] This invention addresses the problem of heterogeneous data representation through a multimodal attention mechanism, improving the accuracy of user intent recognition; it achieves dynamic bidding optimization in complex bidding environments through reinforcement learning algorithms, controlling customer acquisition costs; and it automates and strategizes business negotiations through game theory models, increasing inquiry conversion rates. The system possesses adaptive evolution capabilities, enabling real-time strategy adjustments based on market feedback.
[0016] This invention provides a comprehensive intelligent marketing method based on multimodal information fusion. It offers the following advantages: 1. This invention achieves effective fusion of heterogeneous data such as text, visual data, and behavioral temporal data by constructing a modal-level attention mechanism and a common semantic space mapping model. This method dynamically calculates the weight coefficients of each modality using global context vectors, avoiding the information fragmentation and feature interference problems caused by direct splicing or single-modal analysis. This enables the system to accurately measure the semantic similarity between advertising materials and user profiles within a unified geometric space, thereby improving the accuracy of user purchase intent recognition. 2. This invention utilizes a reinforcement learning model based on a deep Q-network and a PID controller to achieve dynamic strategy generation in complex bidding environments. By incorporating the remaining budget, campaign duration, and estimated conversion rate into the state space, the model can automatically adjust the bidding coefficient based on real-time feedback and find the optimal solution to maximize the total conversion value under budget constraints. Simultaneously, the PID control mechanism smooths the budget consumption rate, preventing premature budget exhaustion due to early aggressive bidding, effectively controlling customer acquisition costs while ensuring full-time coverage. 3. This invention introduces a game theory model based on time-dependent strategies and a multi-dimensional terms exchange mechanism, realizing the automation and strategization of the business negotiation process. The system can dynamically adjust the negotiation resilience coefficient based on the decision urgency in the customer profile, use a non-linear concession strategy to simulate real business games, and seek Pareto improvement solutions by adjusting non-price factors such as payment terms or minimum order quantity when prices are deadlocked. This mechanism reduces manual operating costs, standardizes the price approval process, and improves inquiry conversion efficiency. Attached Figure Description
[0017] Figure 1 This is a block diagram of the overall system architecture of the present invention; Figure 2 This is a detailed flowchart of the multimodal data fusion and intent recognition module according to an embodiment of the present invention; Figure 3 This is a flowchart illustrating the overall process of the intelligent advertising traffic decision-making method according to an embodiment of the present invention. Detailed Implementation
[0018] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0019] Example: Please refer to the appendix Figure 1 - Appendix Figure 3 This invention provides a comprehensive intelligent marketing method based on multimodal information fusion, comprising: S100, System initialization and multi-source data access; Receives multi-platform account authorization information from users through the client terminal, verifies the validity of authorization credentials through the backend service module, establishes API connection channels between the external interface module and the target social media platform and e-commerce platform, and initializes real-time data stream transmission service; S200, Multimodal Data Acquisition and Preprocessing: Raw data is collected from the access platform through an external interface module. The raw data includes user interaction behavior data, platform content data, and environmental parameter data. The backend service module cleans and standardizes the raw data and stores the structured data in the data storage module. S300, Multimodal Feature Extraction and Fusion: It utilizes the deep learning model in the core computing engine module to extract semantic feature vectors from text data, visual feature vectors from visual data, and time-series feature vectors from user behavior. The feature vectors of different modalities are then mapped to a unified semantic space for splicing and fusion processing to obtain multimodal fused features. S400, Holographic User Profile Construction and Intent Recognition; Based on the multimodal fusion features, a dynamic user profile vector containing long-term interest tags, real-time purchase intent, and user behavior prediction values is generated, and a classification model is used to identify the user's current decision-making stage; S500, intelligent traffic delivery strategy generation and execution; calculates the similarity between the feature vector of the advertising material to be promoted and the user dynamic profile vector, outputs the delivery strategy instruction based on the similarity score and bidding environment parameters according to the reinforcement learning model, and converts the delivery strategy instruction into an API request to execute the advertising delivery; S600: Intelligent Interaction and Business Negotiation; Receives customer inquiry messages and converts them into standard intermediate language, extracts key business parameters, combines preset cost calculation models and profit strategies to generate quotation schemes or negotiation response data, and pushes the response content through real-time communication units. S700, Effect Feedback and Model Optimization: Monitor the execution results of the delivery task and the conversion status of the interactive session to collect feedback data, and use the feedback data to update and iterate the model parameters in the core computing engine module.
[0020] The multimodal data acquisition and preprocessing step S200 in this embodiment of the invention specifically includes the following sub-steps: In sub-step S210, distributed collection and classification of the entire dataset are performed. The external interface module works in parallel with the distributed web crawler cluster through a pre-built API connector to obtain the raw data stream from the target platform. The collected dataset is denoted as... The set consists of four heterogeneous subsets: .in, This represents a subset of text data, including user comments, private message records, search queries, and webpage text content. This represents a subset of visual data, including images from user-generated content, keyframes from short videos, and product display images. This represents a subset of behavioral data, including click event streams, page dwell time, browsing paths, and interaction timestamps. This represents a subset of environmental parameter data, including exchange rate fluctuation data of the target market, time zone information, and trade compliance documents. Countermeasures against anti-scraping strategies and IP proxy pool scheduling during web crawling are well-known techniques and will not be elaborated upon here.
[0021] In sub-step S220, text data cleaning and standardization are performed. The backend service module... Each text record in Noise filtering is performed. Regular expressions are used to remove HTML tags, special characters, and non-target language characters, converting the text into a clean text sequence. Subsequently, a natural language processing toolkit is used for word segmentation and stop word filtering. Let the original text sequence be... The cleaned text sequence is represented as ,in The remaining valid word units are used for preprocessing. For multilingual scenarios, the system calls the language detection interface to identify the language of the text and marks non-baseline languages (such as Chinese or English) as untranslatable. However, in this preprocessing stage, the original semantic structure is mainly preserved for subsequent feature extraction.
[0022] In sub-step S230, normalization processing of the visual data is performed. Image or video frame data with varying resolution, aspect ratios, and encoding formats The system first adjusts them to a fixed pixel size. To eliminate illumination differences and accelerate the convergence of subsequent neural network models, Z-score normalization is applied to the image pixel values. Let the original image pixel matrix be... Standardized pixel matrix The calculation formula is as follows: ; in, To train the dataset, the pixel mean vector across the RGB channels, This is the corresponding pixel standard deviation vector. For video data, the system extracts keyframes at preset time intervals (e.g., extracting 1 frame per second) and performs the above normalization process on each keyframe as independent image data.
[0023] In sub-step S240, temporal structuring processing of the behavioral data is performed. (Targeting...) In the discrete interaction logs, the system clusters them according to User ID and Session ID to construct a behavioral sequence representation with time attributes. Let the behavioral sequence of user uu within the time window TT be... It is represented as: ; in, Indicates the first The type of action in the interaction (such as clicking, swiping, or typing). This represents the normalized relative timestamp. This indicates the characteristics of the operation object corresponding to the action (such as ad slot ID or product category ID). The system will interpolate and complete the missing timestamps, and remove invalid interaction records with a dwell time lower than a preset threshold (e.g., 0.5 seconds).
[0024] In sub-step S250, privacy desensitization processing is performed on sensitive data. Before storing the processed data in the database, for fields containing user privacy information (such as device identification codes and geographic coordinates), the system uses a differential privacy mechanism to add perturbation noise, or uses a hash salting algorithm for irreversible encryption. For numerical statistical data... The Laplace algorithm is used for processing, and the desensitized values are output. : ; in, Denotes the Laplace distribution sampling function. For the sensitivity of the query function, This is a privacy budget parameter. This processing ensures that sensitive attributes of a specific user cannot be reverse-engineered during subsequent data mining, thus meeting data compliance requirements in international trade.
[0025] In the multimodal feature extraction and fusion step S300 of this embodiment of the invention, the multimodal feature extraction mechanism based on deep neural networks specifically includes the following sub-steps: In sub-step S310, deep vectorization extraction of text semantic features is performed. The multimodal feature extraction unit in the core computing engine module is configured with a pre-trained language model (such as the BERT model) based on the Transformer architecture. For the preprocessed text sequence data, the model first transforms it into a composite embedding representation containing character vectors, position vectors, and paragraph vectors. Subsequently, a multi-head self-attention mechanism is used to capture long-distance contextual dependencies. The system selects the hidden state vector corresponding to the classification identifier ([CLS] token) in the model output layer as the global semantic representation of the text paragraph. Let the input text embedding sequence be... ,in For sequence length, For the embedding dimension, the mathematical expression of the text feature extraction process is as follows: ; ; in, This represents the context hidden layer matrix output by the encoder. This represents the feature vector corresponding to the [CLS] position. and These are the linear transformation weight matrix and bias term used for dimension mapping, respectively. This is the final output fixed-length text semantic feature vector.
[0026] In sub-step S320, convolutional feature extraction of visual content is performed. For image or video keyframe data, a deep residual network (ResNet-50) is used as the visual backbone network. The network structure contains multiple cascaded residual blocks, each containing convolutional layers, batch normalization layers, and rectified linear unit (ReLU) activation functions. For the input normalized image matrix... Deep visual features are extracted through forward propagation. To obtain a spatially independent global visual representation, the system extracts the feature map output from the last convolutional layer in the network and performs global average pooling on it, instead of directly using the classification results from the fully connected layers. This process is calculated as follows: ; ; in, This represents the mapping function of a convolutional neural network excluding fully connected layers. For network parameters, For the feature map tensor, ( These represent the height, width, and number of channels of the feature map, respectively. This is the output visual feature vector. The specific implementation details of the ResNet-50 network architecture and convolution operations are well-known to those skilled in the art and will not be elaborated upon here.
[0027] In sub-step S330, temporal feature extraction of user interaction behavior is performed. For discrete behavior sequence data containing timestamp information, a Long Short-Term Memory (LSTM) network is used for modeling. The LSTM network solves the gradient vanishing problem in long sequence training through forget gates, input gates, and output gates, thereby effectively capturing the long-term dependencies of user behavior. The user behavior input vector at time t is (Including the concatenation of action type embedding and object feature embedding), the hidden state at the previous time step was... The cell state is The update calculation at the current moment is as follows: ; ; ; ; ; ; in, This represents the Sigmoid activation function. This represents element-wise multiplication. and For the corresponding weight matrix and bias terms , , These represent the activation values of the forget gate, input gate, and output gate, respectively. The system will use the last time step of the sequence. Hidden state As a behavioral feature vector representing users' recent behavioral patterns .
[0028] In substep S340, dimensional alignment of the multimodal features is performed. Because... , and Generated by different network structures, the vectors have different dimensions. The system uses independent fully connected layers (DenseLayer) as projection heads to map the three feature vectors to a common semantic space of the same dimension, resulting in aligned feature vectors. , and This provides a unified data foundation for subsequent cross-modal fusion computing.
[0029] In the multimodal feature extraction and fusion step S300 of this embodiment of the invention, the specific implementation method for cross-modal feature fusion and semantic space mapping includes the following sub-steps: In sub-step S350, a multimodal feature joint matrix is constructed. The core computing engine module receives the dimension-aligned feature vectors output from the previous stage, i.e., the text feature vectors. Visual feature vectors and behavioral feature vectors The system stacks the aforementioned feature vectors along the feature channel dimension to construct a multimodal joint feature matrix. Let the aligned feature dimensions be... The number of modes is (In this embodiment) ),but This matrix represents the original state combinations of a single user or object under different information sources, but it does not yet include the interaction weight information between modalities.
[0030] In sub-step S360, modality weights based on an attention mechanism are calculated. To address the issue of varying contributions from different modalities in specific business scenarios, the system introduces a Modality-Level Attention Module. This module learns a global context vector. Calculate the relevance score of each modality feature relative to the global context. For the , Feature vectors of each modality (in ), their attention score The calculation formula is as follows: ; in, Transform the weight matrix for attention. For bias vectors, The context query vector learned during training. The hyperbolic tangent activation function is used. Then, the Softmax function is applied to the score. After normalization, the final weight coefficients of each mode are obtained. : ; This weighting coefficient This reflects the importance of the corresponding modality in constructing the final profile and satisfies... The constraints.
[0031] In sub-step S370, a multimodal weighted fusion vector is generated. This is based on the calculated weight coefficients. The system performs a weighted summation of the original feature vectors to generate a fused intermediate feature vector. The calculation formula is as follows:
[0032] ; The weighted summation method used here instead of direct concatenation can effectively control the dimensionality of the final feature vector, avoid the dimensionality curse caused by the increase in the number of modalities, and retain the key high-weight feature information.
[0033] In substep S380, a nonlinear mapping of the common semantic space is performed. To give the fused feature vectors measurable geometric properties, the system maps them to the final common semantic space through a multilayer perceptron (MLP). This MLP contains two fully connected layers and an intermediate ReLU activation layer. The mapping process is as follows: ; in, , and , These are the weight matrices and bias terms for each layer. The output vector... This represents the holographic feature representation of the user or object. In this common semantic space, the geometric distance between different objects (such as user profile vectors and advertising creative vectors) directly characterizes their semantic similarity. The system uses cosine similarity as the metric within this space. ; The forward propagation computation logic of the multilayer perceptron is well-known to those skilled in the art and will not be elaborated upon here. This mapping step ensures that the system can perform efficient retrieval and matching based on a unified vector standard in subsequent steps.
[0034] The holographic user profile construction and intent recognition step S400 in this embodiment of the invention specifically includes the following sub-steps: In sub-step S410, a structured representation of the holographic user feature vector is established. The user profile construction unit is based on the multimodal fusion feature vector output from the previous step. Construct high-dimensional sparse user vectors The vector is logically divided into three feature subspaces: the intrinsic attribute subspace, the long-term interest subspace, and the short-term behavior subspace. The intrinsic attribute subspace stores the user's static tags, including enterprise type (manufacturer / trader / retailer), geographical region, and estimated purchase size; the long-term interest subspace stores the user's preference weights for specific product categories over historical periods; and the short-term behavior subspace stores the user's real-time interaction features within the current session window. The system utilizes hash mapping technology to map discrete tag data to corresponding vector dimensions, ensuring both the density of the vector space and computational efficiency.
[0035] In sub-step S420, the user profile is dynamically updated over time. To reflect real-time changes in user needs, the system uses an exponential decay accumulation algorithm to continuously update the user vector. When the system captures new user interactions and generates new feature vectors... At that time, the time decay factor was used to modify the historical portrait vector. Perform weighted processing to calculate the user profile vector at the current time. : ; in, This is the memory retention coefficient, used to balance the influence weights of historical preferences and current behavior; The time interval between the last interaction and the current interaction; The time decay function is typically used. In form, This is the decay rate parameter. This update mechanism ensures that the profile can quickly respond to changes in the user's immediate intent, while retaining long-term accumulated behavioral patterns.
[0036] In sub-step S430, an attribute prediction model based on multi-task learning is constructed. For cases where some user attributes are missing, the core computing engine module deploys a multilayer perceptron (MLP) classifier to infer unknown attributes using known behavioral features. This model employs a multi-task learning architecture, sharing underlying feature extraction parameters, and setting multiple independent fully connected branches in the output layer to correspond to different attribute tasks (such as "whether it is a wholesaler" or "budget range prediction"). Let the input vector be... , No. Predicted output for each task The calculation is as follows: ;
[0037] in, and For the weights and biases of the shared layer, and For parameters of a specific task branch. The system selects attribute values with a predicted probability greater than a preset confidence threshold (e.g., 0.85) and automatically populates them into the inherent attribute subspace of the user profile.
[0038] In sub-step S440, the classification and identification of the user's purchase intent stage is performed. The system divides the user's current decision-making process into a discrete set of intent stages. These correspond to four stages: unconscious browsing, information gathering, option comparison, and purchase decision-making. The intent recognition model employs an attention-based classification network, with user profile vectors as input. The output is the conditional probability distribution belonging to each intent stage. The calculation formula is as follows: ; ; in, and For the corresponding category The regression parameters. The system selects the category with the highest probability as the user's current intent state. If the recognition result is... (During the purchase decision phase), the system will generate a high-priority marketing trigger signal and transmit it to the subsequent intelligent traffic decision unit for adjusting the bidding strategy. The construction of the multi-class cross-entropy loss function and the backpropagation training process are well-known techniques to those skilled in the art and will not be elaborated upon here.
[0039] The intelligent flow allocation decision generation and execution step S500 in this embodiment of the invention specifically includes the following sub-steps:
[0040] In sub-step S510, precise matching and click-through rate prediction of ad creatives are performed. The ad delivery decision unit uses a dual-tower deep neural network (DNN) architecture for relevance calculation. This network contains two independent branches: the user tower and the creative tower, which receive user profile vectors respectively. With the feature vector of candidate ad creative As input, the two side vectors are mapped to the same-dimensional interaction space through a nonlinear mapping of multiple fully connected layers. The system calculates the inner product of the two output vectors and outputs the predicted click-through rate (PredictedCTR, pCTR) through the Sigmoid activation function. The calculation formula is as follows: ; in, For the Sigmoid function, and These are the mapping functions for the user tower and the resource tower, respectively. The system is based on... The candidate creatives are sorted by value, and the top-N highly relevant ad creatives are selected to enter the bidding queue.
[0041] In sub-step S520, a reinforcement learning model based on the real-time bidding environment is constructed. The system models the ad delivery process as a Markov decision process (MDP) and defines the state space. Action space and reward function .
[0042] state space By quadruplets Composition, in which This represents the remaining budget at the current moment. Indicates the remaining delivery time. The estimated conversion rate (pCVR) for the current traffic requests. This represents the estimated cost per thousand impressions in the current market.
[0043] Action space Defined as a continuous bid adjustment factor The system controls the aggressiveness of the final bid by adjusting this coefficient.
[0044] reward function The aim is to maximize total conversion value (GMV) while controlling costs within budget constraints. Instant rewards for steps Defined as: ; in, Indicates whether a transformation has occurred. The estimated value of the order. This refers to the actual advertising costs incurred. This is a hyperparameter used to balance the return on investment (ROI).
[0045] In sub-step S530, policy optimization based on a deep Q-network (DQN) is performed. The system utilizes a deep neural network to fit the data. Value function Used to evaluate in state Take action below The long-term expected return. Network parameters. Iterative updates are performed by minimizing the Time Difference Error (TDError). Loss function The calculation is as follows: ; in, As a discount factor, For the parameters of the target network, These are the parameters for the current evaluation network. To address the sample correlation problem, the system maintains an ExperienceReplayBuffer to store historical transition quadruples. Random sampling is performed during training.
[0046] In sub-step S540, dynamic bidding instructions and budget smoothing control are generated. Based on the trained Q-network, the bidding decision unit performs the following at each time step: Based on the current state Select the optimal action : ; Final bid amount Determined by both the fundamental value and the adjustment factor: ; in, The system sets a target conversion cost for advertisers. Simultaneously, a PID controller is introduced to smooth the rate of remaining budget depletion. If the current depletion rate exceeds the preset timeline, the PID controller will output a negative feedback signal to reduce it. The value is adjusted to prevent the budget from being exhausted prematurely before the end of the day, ensuring that advertising coverage reaches the target market during its most active periods. The specific parameter tuning of the PID control algorithm can be adjusted by those skilled in the art based on actual traffic fluctuations, and will not be elaborated upon here.
[0047] In the intelligent interaction and business negotiation step S600 of this embodiment of the invention, the specific implementation method of the AI agent's autonomous quotation and cost calculation model includes the following sub-steps: In sub-step S610, inquiry intent parsing and parameter structure extraction are performed. After receiving the translated standard intermediate language text, the intelligent interaction unit uses a Named Entity Recognition (NER) model to extract the set of key parameters needed to construct the quotation. This parameter set is represented as ,in Code the inventory quantity unit for specific products. For the quantity requested by the customer. Trade terms (such as FOB, CIF, DDP). This refers to the port of destination or delivery address. If any of the above necessary parameters are missing from the inquiry information, the system will trigger follow-up queries; if the parameters are complete, the cost calculation engine will be invoked for calculation.
[0048] In sub-step S620, the basic cost of goods sold is calculated based on real-time production data. The system retrieves the corresponding cost through the internal ERP interface. Bill of Materials (BOM) and current production schedule data. Basic ex-factory cost. The calculation covers raw material costs, dynamic labor costs, and allocated manufacturing overhead. The calculation formula is as follows: ; in, For the first The consumption of various raw materials, This is the real-time purchase price of the raw material. Standard manufacturing time, The rate is based on the unit hourly rate. Fixed costs such as depreciation and energy consumption for this batch of equipment. This refers to the production batch quantity. For the interface calls to ERP data and the basic accounting logic, those skilled in the art can implement them using existing enterprise resource management systems, and will not be elaborated upon here.
[0049] In sub-step S630, cross-border logistics and compliance costs are calculated. The system calculates based on... Calculate the total volume based on product packaging parameters. With total weight Subsequently, data from the port of shipment to... was obtained via an external logistics API. Real-time shipping rates and insurance premium rates Based on the extracted trade terms Determine logistics costs The boundary. Taking CIF (Cost, Insurance and Freight) terms as an example, logistics costs are calculated as follows: ; in, For the value of goods, For export customs declaration and port charges. If If it is FOB, then and Set to 0. The system has a built-in port rate database and supports automatic matching of different shipping routes and container types (20GP / 40HC).
[0050] In sub-step S640, a final quote based on dynamic profit margin and exchange rate hedging is generated. The system introduces dynamic pricing coefficients. This coefficient is determined by customer ratings and historical transaction probabilities in the user profile. Simultaneously, the system obtains real-time exchange rates. (e.g., the buying price of USD against CNY), and introduce a reserve factor for exchange rate fluctuation risk. Final price for a single item. The calculation model is as follows: ; in, The preset value range is [0.05, 0.4], and the specific value is determined by the positive correlation between the strength of the user's purchase intent output in the previous step S400; The value is usually set to the maximum fluctuation range of the exchange rate within a preset period (e.g., 0.02).
[0051] In sub-step S650, a tiered quantity discount check is performed. The system has a pre-stored tiered pricing rule function. The system will calculate With the number of requests Substitute the values into the function for correction. If Once the preset wholesale threshold range is reached, the system automatically applies the corresponding discount factor and generates a final quotation data package to be sent to the customer. This data package includes the unit price, total price, validity period, and a detailed parameter specification table, and is rendered into PDF or HTML format by the backend service module before being sent to the client.
[0052] In the intelligent interaction and business negotiation step S600 of this embodiment of the invention, the specific implementation method of the dynamic negotiation and objection handling strategy based on game theory includes the following sub-steps: In sub-step S660, negotiation intent identification and state space initialization are performed. The intelligent interaction unit performs semantic parsing on the feedback text sent back by the client to identify the core objections. The system defines a set of objections. These correspond to objections regarding price, payment method, delivery date, and quality terms, respectively. Once an objection type is identified, the system immediately initializes the negotiation game model and sets a negotiation round counter. And determine the maximum number of rounds of negotiation allowed. At the same time, the system retrieves the reserved floor price for this transaction. (i.e., the lowest acceptable transaction price for the company) and the initial expected price. Construct the Zone of Possible Agreement (ZOPA) for negotiation.
[0053] In sub-step S670, a concession function based on a time-dependent tactic is constructed. To simulate the psychological game in real business negotiations, the system does not employ a linear price reduction strategy, but rather a nonlinear concession model based on a polynomial function. During the round of negotiations, the system calculates the currently recommended offer. The calculation formula is as follows: ; in, represents the game resilience coefficient. When... At this point, the model exhibits a "Boulware strategy," meaning it makes minimal concessions in the early stages of negotiations, maintaining a tough stance until the close of negotiations. Only then did they compromise; when In this scenario, the model exhibits a "Conceder strategy," meaning it makes rapid concessions initially to facilitate a quick transaction. This coefficient... The urgency of the decision is dynamically determined by the "decision urgency" feature in the customer profile generated in the previous steps. If a customer is identified as urgently needing to make a purchase, the system automatically increases the urgency level. The value is adjusted accordingly; conversely, it is adjusted to a smaller value.
[0054] In sub-step S680, a decision based on the utility function is performed. This occurs when a counter-quote is received from the customer. At that time, the system calculates the utility value of the counter-offer to the company. and the threshold utility of the current round. Comparisons are made. The utility function is defined as: ; like The system generates an "Accept" instruction, locking in the transaction price; if and The system generates a "Reject and Counter-offer" instruction, which will... Send it to the customer as a new quote; if If no agreement is reached, a manual takeover process will be triggered.
[0055] In sub-step S690, the trade-off process for multi-dimensional terms is executed. This occurs when negotiations on a single price dimension reach a stalemate (i.e., the price changes for both parties in two consecutive rounds are less than a preset threshold). When this occurs, the system activates a multi-dimensional game mechanism. The system retrieves a pre-defined terms-of-contract matrix to find Pareto improvement solutions. For example, the system generates a statement containing concessions: "If you accept the 30% prepayment ratio being adjusted to 50%, we can accept your proposed unit price." This logic introduces non-price variables (such as payment terms). Minimum order quantity This expands the game space. The new combinatorial utility is calculated as follows: ;
[0056] in, , , These are the weighting coefficients for price, payment terms, and quantity, respectively. The system iterates through all feasible combinations and selects... Reply with the maximum possible solution that exceeds the current minimum requirement.
[0057] In sub-step S700, the negotiation script is dynamically generated. The system has a built-in pragmatics-based negotiation corpus, which categorizes the scripts according to sentiment tags such as "tough," "moderate," "regretful," and "inducing." Based on the current concession level and the stage of the negotiation, the system selects the corresponding sentiment template and uses the Natural Language Generation (NLG) module to generate the numerical offer. Alternatively, the terms of exchange can be transformed into natural language text that conforms to business etiquette. Template filling and smoothing techniques in the natural language generation process are well-known to those skilled in the art and will not be elaborated upon here.
[0058] The customer lifecycle prediction and intelligent wake-up mechanism in this embodiment of the invention specifically includes the following sub-steps: In sub-step S710, a customer churn early warning model based on survival analysis is constructed. The core computing engine module uses the Cox Proportional Hazards Model to model the customer's historical interaction data to assess the probability risk of customer churn or cessation of interaction at a specific point in time. Let the customer's feature vector be... (Including activity level, historical transaction volume, and the interval between the most recent interactions extracted in the preceding step S400), the time variable is... Then the customer churn risk function at time tt Defined as: ; in, The baseline risk function represents the trend of basic churn risk over time under the influence of no covariates. For the first Features The regression coefficients are used to quantify the weight of this feature's impact on churn risk. The system uses maximum likelihood estimation to evaluate the parameters. Perform training and solve the problem. If the calculated cumulative survival probability... If the user's ID falls below a preset warning threshold (e.g., 0.4), the system marks the user as a "high-risk churned user" and pushes the user's ID to the wake-up task queue.
[0059] In sub-step S720, the optimal wake-up time window is calculated. To avoid unnecessary disturbance to customers, the system analyzes the time interval distribution of customers' historical orders to predict the expected time when they will next generate a purchasing demand. Assume that the customers' historical purchase intervals follow a normal distribution. The system is set to wake up based on the current time. satisfy: ; in, The time of the last interaction or transaction. For average procurement cycle, This is the tolerance window width. For new customers with no prior transaction history... The value is taken as the average conversion cycle of all customers in the same industry category.
[0060] In sub-step S730, differentiated wake-up strategy content is generated. The system generates targeted marketing materials based on the "churn reason attribution" in the customer profile. If the attribution result is "price-sensitive churn," the intelligent interaction unit automatically generates email content containing limited-time discount coupons or a stock clearance list; if the attribution result is "demand mismatch churn," the system calls the recommendation algorithm to filter new products (NewArrivals) with the highest similarity to the customer's historical browsing history and generates a product recommendation catalog. The system uses slot-filling technology to embed specific product information into a pre-set EDM (email marketing) template to generate a wake-up message package to be sent.
[0061] In sub-step S740, multi-path outreach based on channel preferences is performed. The system maintains a probability table containing the success rate of multi-channel outreach. , indicating user In channels ( The response probability on the given information. The system selects the channel that maximizes the expected response. Send push notifications: ; in, The expected conversion value of the message. This refers to the sending cost for the corresponding channel. If no response is received from the preferred channel within a preset time (e.g., 24 hours), the system automatically switches to the secondary channel for a second outreach, but the total number of outreaches within the same wake-up cycle does not exceed a preset limit (e.g., 3 times) to prevent excessive marketing that may cause user resentment. The API call logic for the SMTP email sending protocol and instant messaging tools is well-known to those skilled in the art and will not be elaborated upon here.
[0062] The data feedback closed loop and model adaptive evolution step S700 in this embodiment of the invention specifically includes the following sub-steps: In sub-step S750, attribution and sample construction of delayed feedback data are performed. Due to the time lag between ad placement and actual transactions, the backend service module uses a time window-based delayed labeling mechanism to construct the training sample set. The system sets the observation window. (For example, 7 days), for conversion events occurring within this window, the system backtracks and retrieves the corresponding exposure logs and click logs, and then extracts the original feature vectors. With the final transformation result ( ) to construct positive samples by association; for more than Clicks that do not result in a conversion are marked as negative samples. The system assigns a reliability weight to each sample. This weight decays over time to reduce the impact of outdated data on the model.
[0063] In sub-step S760, online incremental learning and parameter updates are performed. To adapt to the rapidly changing market environment, the core computing engine module does not use full-data retraining, but instead uses a stochastic gradient descent (SGD) algorithm based on streaming data to fine-tune the model parameters. Let the model parameters at the current time be... The newly arrived data batch (Mini-batch) contains Sample The system calculates the gradient of the average loss function for this batch. And update the parameters: ; in, For learning rate, This is the regularization coefficient, used to prevent overfitting; For click-through rate prediction tasks, a log-loss function is used as the loss function. This step allows the model to absorb the latest user behavior patterns in real time without interrupting the service.
[0064] In sub-step S770, a traffic exploration strategy based on Thompson sampling is implemented. To strike a balance between utilizing the existing optimal model and exploring potentially better strategies, the traffic allocation decision unit introduces the Multi-Armed Bandit (MAB) algorithm for traffic allocation. The system assumes that the conversion rate of each ad creative follows a certain pattern. distributed At each decision, a predicted conversion rate is generated by sampling from the posterior distribution. : ; in, , These are prior parameters. The cumulative number of conversions for this material. This represents the number of times the sample was not converted. The system selects the sample value. The system will deploy the most effective creative or strategy. As the amount of data accumulates and the variance decreases, the system will automatically converge to the best-performing strategy, achieving adaptive optimization of traffic allocation.
[0065] In sub-step S780, knowledge distillation and lightweight deployment are performed. To improve the response speed of the online inference service, the system uses a teacher-student network architecture for model compression. The core computing engine module maintains a high-precision teacher model with a large number of parameters. And train a simplified student model. During training, the student model learns not only the true labels YY, but also the soft labels (SoftTargets) output by the teacher model. The total loss function... Defined as: ; in, For cross-entropy loss, The Kullback-Leibler divergence measures the difference between the teacher's output distribution and the student's output distribution. , These are the Logits outputs for the teacher and student networks, respectively. For temperature parameters, This is the balancing coefficient. This step transfers the inference capabilities of complex models to lightweight models, thereby reducing reliance on hardware computing resources. For the mathematical properties of the stochastic gradient descent algorithm and KL divergence, those skilled in the art can refer to relevant machine learning theories, which will not be elaborated upon here.
Claims
1. A comprehensive intelligent marketing method based on multimodal information fusion, characterized in that, Includes the following steps: S100, System Initialization and Multi-Source Data Access: Establish API connection channels between external interface modules and target social media platforms and e-commerce platforms, and initialize real-time data stream transmission services; S200, Multimodal data acquisition and preprocessing: Raw data including user interaction behavior data, platform content data and environmental parameter data are acquired through the external interface module. The raw data is cleaned and standardized by the backend service module, and the structured data is stored in the data storage module. S300, Multimodal Feature Extraction and Fusion: The deep learning model in the core computing engine module is used to extract the semantic feature vector of text data, the visual feature vector of visual data, and the time series feature vector of user behavior, respectively. The feature vectors of different modalities are mapped to a unified semantic space for splicing and fusion processing to obtain multimodal fusion features. S400, Holographic User Profile Construction and Intent Recognition: Based on the multimodal fusion features, a dynamic user profile vector containing long-term interest tags, real-time purchase intent, and user behavior prediction values is generated, and a classification model is used to identify the user's current decision-making stage; S500, Intelligent Traffic Flow Strategy Generation and Execution: Calculate the similarity between the feature vector of the advertising material to be promoted and the user dynamic profile vector, output the delivery strategy instruction based on the similarity score and bidding environment parameters according to the reinforcement learning model, and convert the delivery strategy instruction into an API request to execute the advertising delivery; The intelligent traffic delivery strategy generation and execution specifically includes: A dual-tower deep neural network is used to process the user profile vector and the candidate ad creative feature vector respectively, calculate the inner product of the two output vectors and output the estimated click-through rate; Construct a reinforcement learning model based on a real-time bidding environment, defining a state space that includes remaining budget, remaining campaign duration, estimated conversion rate, and estimated cost per thousand impressions, as well as an action space defined as the continuous bid adjustment coefficient; By fitting the Q-value function using a deep Q-network, the long-term expected return of taking a specific bid adjustment coefficient in the current state is evaluated, and the optimal action is selected to generate a bid order based on the principle of maximizing the Q-value. A PID controller is introduced to monitor the consumption rate of the remaining budget. When the consumption rate is higher than the preset time progress curve, a negative feedback signal is output to reduce the bid adjustment coefficient. S600, Intelligent Interaction and Business Negotiation: Receives customer inquiry messages and extracts key business parameters, generates a quotation plan by combining a preset cost calculation model and profit strategy, or generates negotiation response data based on a game theory model, and pushes the response content through a real-time communication unit; S700, Effect Feedback and Model Optimization: Monitor the execution results of the delivery task and the conversion status of the interactive session to collect feedback data, and use the feedback data to update and iterate the model parameters in the core computing engine module.
2. The omni-channel intelligent marketing method based on multimodal information fusion according to claim 1, characterized in that, In step S200, the multimodal data acquisition and preprocessing specifically includes: The entire dataset is divided into subsets of text data, visual data, behavioral data, and environmental parameter data. For a subset of text data, HTML tags and special characters are removed using regular expressions, while preserving the original semantic structure of the non-baseline language; For the visual data subset, the keyframes of the images or videos are adjusted to a fixed pixel size and Z-Score normalization is performed based on the pixel mean vector and pixel standard deviation vector in the RGB channels of the training dataset. For the subset of behavioral data, cluster according to user ID and session ID to construct a behavioral sequence representation with normalized relative timestamps; For fields containing user privacy information, a differential privacy mechanism is used to add Laplace noise for desensitization.
3. The omni-channel intelligent marketing method based on multimodal information fusion according to claim 1, characterized in that, In step S300, the multimodal feature extraction specifically includes: By using a pre-trained language model based on the Transformer architecture, the text sequence is transformed into a composite embedding representation of character vectors, position vectors, and paragraph vectors. The hidden state vector corresponding to the classification identifier is selected as the text semantic feature vector. Visual data is processed using a deep residual network as the visual backbone network. The feature map output by the last convolutional layer is extracted, and a global average pooling operation is performed on the feature map to obtain a visual feature vector. Long Short-Term Memory (LSTM) networks are used to model behavioral sequences containing timestamp information. The unit states are updated through forget gate, input gate, and output gate mechanisms, and the hidden state of the last time step of the sequence is used as the behavioral feature vector. An independent fully connected layer is set up as a projection head to map the text semantic feature vector, visual feature vector and behavioral feature vector to a common semantic space of the same dimension, so as to obtain dimension-aligned feature vectors.
4. The omni-channel intelligent marketing method based on multimodal information fusion according to claim 3, characterized in that, In step S300, the fusion process specifically includes: The feature vectors of each modality after dimension alignment are stacked along the feature channel dimension to construct a multimodal joint feature matrix; By learning a global context vector, the relevance score of each modality feature relative to the global context is calculated, and the weight coefficient of each modality is obtained by normalizing the relevance score using the Softmax function. The original feature vectors are weighted and summed based on the weighting coefficients to generate a fused intermediate feature vector. By using a multilayer perceptron containing fully connected layers and activation layers, the intermediate feature vectors are mapped to a common semantic space to obtain a holographic feature representation, where the geometric distance between different objects in the common semantic space represents their semantic similarity.
5. The omni-channel intelligent marketing method based on multimodal information fusion according to claim 1, characterized in that, In step S400, the holographic user profile construction and intent recognition specifically include: Construct a high-dimensional sparse user vector that includes an intrinsic attribute subspace, a long-term interest subspace, and a short-term behavior subspace; When a new interaction is captured, the historical profile vector is weighted using a time decay function, and the updated user profile vector is calculated by combining it with the new feature vector at the current moment. Using an attention-based classification network, with user profile vectors as input, we calculate the conditional probability distribution of the user's intent in four stages: unconscious browsing, information gathering, solution comparison, and purchase decision. The category with the highest probability is selected as the current intent state.
6. The omni-channel intelligent marketing method based on multimodal information fusion according to claim 1, characterized in that, In step S600, the generation of the quotation scheme specifically includes: Extract the inventory unit code, purchase quantity, trade terms, and destination port parameters from the inquiry message; The system calls upon the enterprise resource management system to obtain the bill of materials and production schedule, and calculates the basic cost of the factory, which includes raw material costs, dynamic labor costs, and allocated manufacturing overhead. Calculate the total volume and weight of the goods based on the purchase quantity, obtain real-time freight rates and insurance rates by combining the parameters of the destination port, and calculate the cross-border logistics costs. Obtain real-time exchange rates and exchange rate fluctuation risk reserve ratios, determine dynamic pricing coefficients based on the intensity of user purchase intent, and generate a final price quote for a single item by combining the aforementioned basic factory cost and cross-border logistics cost. The final price is adjusted for quantity discounts using a pre-stored tiered pricing rule function.
7. The omni-channel intelligent marketing method based on multimodal information fusion according to claim 1, characterized in that, In step S600, the generation of negotiation response corpus based on the game theory model specifically includes: Identify the types of objections in customer feedback texts, initialize the negotiation game model, and set the maximum allowed negotiation rounds, reserve bottom price, and initial expected offer to construct the negotiation feasibility domain; Construct a concession function based on a time-dependent strategy, dynamically adjust the game resilience coefficient according to the decision urgency characteristics in the customer profile, and use a polynomial function to calculate the suggested offer for the current round; Calculate the utility value of the customer's counter-offer to the enterprise. If the utility value is lower than the threshold utility of the current round and the maximum negotiation round has not been reached, generate a rejection and counter-offer instruction. When negotiations on a single price dimension reach a stalemate, a multi-dimensional game mechanism is activated to search for Pareto improvement solutions by retrieving the terms exchange matrix and calculating the combined utility by introducing payment terms or minimum order quantity variables. Based on the current concession level and the stage of the game, select an emotional template and use the natural language generation module to convert numerical offers or exchange terms into negotiation response text.
8. The omni-channel intelligent marketing method based on multimodal information fusion according to claim 1, characterized in that, In step S700, the effect feedback and model optimization specifically include: A time-window-based delayed labeling mechanism is adopted to backtrack and retrieve exposure and click logs corresponding to conversion events to construct a training sample set, and confidence weights are configured for the sample scores. An online incremental learning method is adopted, and the gradient of the average loss function is calculated based on the stochastic gradient descent algorithm of streaming data to update the model parameters in the core computing engine module in real time. Traffic allocation is performed using a multi-armed slot machine algorithm. Based on the Beta distribution, Thompson sampling is used to evaluate the conversion rate of ad creatives, and the strategy with the largest sample value is selected for deployment. A teacher-student network architecture is used for model distillation. The student model is trained to learn both the real labels and the soft labels output by the teacher model, thus transferring the reasoning ability of the teacher model to the lightweight student model.
9. A system based on the marketing method of claim 1, characterized in that, include: Memory, used to store computer programs; A processor, configured to implement the steps of the multimodal information fusion-based omni-channel intelligent marketing method as described in any one of claims 1 to 8 when executing the computer program.