An artificial intelligence-based dialogue marketing strategy optimization method and system
By combining multimodal data perception and cross-modal information processing with privacy-enhanced federated learning and reinforcement learning, the problem of insufficient dynamic modeling of user emotional states is solved, enabling real-time strategy optimization and personalized responses in marketing using artificial intelligence, thereby improving the practicality and user adaptability of the marketing system.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-24
- Publication Date
- 2026-04-07
AI Technical Summary
Existing technologies suffer from fragmented multimodal data in complex marketing scenarios, resulting in insufficient dynamic modeling of user emotional states, difficulty in capturing subconscious behavioral characteristics, and a lack of coordination between privacy protection mechanisms and strategy optimization goals, leading to a lag in the real-time business application of artificial intelligence in the marketing field.
By collecting user voice, text, facial expressions, and physiological signals in real time through a multimodal perception component, a dynamic emotion map is constructed. Combined with a cross-modal information processing module, causal relationship descriptions are generated. The marketing strategy model is then collaboratively trained in a privacy-enhanced federated learning framework. Reinforcement learning algorithms are used to optimize strategy selection in real time, generating personalized response and feedback data, ensuring data privacy and strategy security.
It achieves accurate capture and real-time response to user emotions, enhances the practicality and user adaptability of artificial intelligence in conversational marketing, balances user experience and corporate business goals, and reduces decision-making lag.
Smart Images

Figure CN120688641B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of intelligent dialogue, and in particular to a method and system for optimizing dialogue marketing strategies based on artificial intelligence. Background Technology
[0002] In recent years, the application of artificial intelligence (AI) technology in the field of business marketing has evolved from single-function tools to complex system integration. Conversational AI, through natural language processing technology, enables 24 / 7 customer service and significantly improves interaction efficiency; predictive analytics and personalized recommendation systems optimize ad placement and product recommendations based on user behavior data (such as clicks and purchase records), driving up conversion rates.
[0003] However, existing technologies still have limitations in complex marketing scenarios: on the one hand, the fragmentation of multimodal data leads to insufficient dynamic modeling of users' emotional states, making it difficult to capture subconscious behavioral characteristics; on the other hand, there is a lack of synergy between privacy protection mechanisms (such as differential privacy) and strategy optimization goals, making it difficult to balance efficiency and security in cross-enterprise data collaboration.
[0004] These issues limit the deep application of artificial intelligence in the marketing field, resulting in lagging model updates and difficulty in adapting to real-time business needs. As can be seen, how to improve the practicality of artificial intelligence for real-time business remains to be solved. Summary of the Invention
[0005] To improve the practicality of artificial intelligence for real-time business, this application provides a method and system for optimizing dialogue marketing strategies based on artificial intelligence.
[0006] Firstly, this application provides a method for optimizing dialogue marketing strategies based on artificial intelligence, employing the following technical solution:
[0007] An AI-based method for optimizing conversational marketing strategies includes:
[0008] Multimodal perception components deployed on terminal devices are used to collect multi-round interaction data between users and AI dialogue systems in real time. The multi-round interaction data includes voice signals, text data, facial expression signals and physiological signals. A dynamic emotion map is generated based on the multi-round interaction data. The nodes of the dynamic emotion map represent the user's emotional state, the edges represent the transition probability between emotional states, and the node weights are dynamically adjusted with time series. Multimodal feature vectors corresponding to the multi-round interaction data are extracted.
[0009] The multimodal feature vectors are input into the cross-modal information processing module. The multimodal feature vectors are aligned by a contrastive learning algorithm to construct a cross-modal embedding space. A causal relationship description between user behavior and marketing strategies is generated in the cross-modal embedding space. The causal relationship description is combined with historical interaction data to simulate the possible long-term impact of different marketing methods and output the corresponding strategy risk score.
[0010] The strategy risk score and causal graph model are input into the privacy-enhanced federated learning framework to collaboratively train the marketing strategy model across multiple enterprise nodes. Each enterprise node generates differential privacy-preserving model parameter increments based on local data, and uploads the parameter increments to the federated learning server after encrypting them with homomorphic encryption technology. The federated learning server aggregates the model parameters using a secure multi-party computation protocol based on the encrypted parameter increments to generate a global strategy model and updated model parameter increments.
[0011] The global strategy model and updated model parameters are incrementally input into the strategy generation module. Combined with the user's current emotional state in the dynamic sentiment graph, alternative solutions for high-risk strategies are extracted from the causal graph model, and an emotion-adapted dialogue script is generated. At the same time, reinforcement learning algorithms are used to optimize strategy selection in real time to maximize the weighted objective function of user emotional satisfaction and long-term business value, and finally generate personalized response and feedback data. The personalized response realizes multi-sensory interaction through speech synthesis, image generation, and haptic feedback devices.
[0012] The feedback data is sent back to the multimodal perception component to update the dynamic sentiment graph and multimodal feature vectors, forming a closed-loop optimization process. The closed-loop optimization process adjusts the weights of user emotional state nodes in real time, updates the policy risk score in the causal graph model, and dynamically adjusts the global policy model in the federated learning framework.
[0013] Optionally, when the multimodal perception component acquires speech, text, facial expressions, and physiological signals, an incremental emotion feature update mechanism is adopted, and the method further includes:
[0014] In the time series, features are extracted only from the difference in emotional state between the current moment and the previous moment to generate incremental feature vectors; the gradient of emotional change is calculated through a sliding window, and the update frequency of the feature vectors is dynamically adjusted. When significant emotional fluctuations are detected, high-frequency updates are triggered; the incremental feature vectors are hash-matched with the historical feature library, redundant data is filtered out, and then uploaded to the emotion modeling module.
[0015] Optionally, in the cross-modal information processing module, a cross-modal embedding space is constructed using a multi-granularity alignment strategy, and the method further includes:
[0016] For speech signals and text data, a dual-path processing approach of word-level alignment and sentence-level alignment is adopted, and fine-grained semantic embeddings are generated through contrastive learning. For image data and physiological signal data, corresponding local and global features are extracted, and hierarchical alignment is performed based on the local and global features. The embedding results of different granularities are fused using an attention mechanism. A weighted fusion strategy is used to generate cross-modal causal relationship descriptions.
[0017] Optionally, when the policy generation module uses reinforcement learning algorithms to optimize policy selection in real time, a multi-objective reward function is designed, and the reward function includes:
[0018] Based on the real-time emotional state node weight calculation of the dynamic emotional graph, a corresponding user emotional satisfaction score is generated through the joint evaluation of emotional state transition probability and user behavior feedback; when the user's emotional state fluctuates drastically in the dynamic emotional graph, the weight coefficient of the user emotional satisfaction score is increased.
[0019] Based on the strategy risk score and user lifetime value prediction results output by the causal graph model, the long-term impact of different marketing strategies on business objectives is quantified by a weighted regression model to obtain a long-term business value score. If the historical strategy risk score of a marketing strategy is lower than the preset strategy risk score threshold, the weight coefficient of the long-term business value score corresponding to the marketing strategy in the multi-objective reward function is reduced.
[0020] Based on the timing synchronization index of the haptic feedback device and the voice / image generation module, the response delay is calculated by the timestamp alignment error, and the delay value is mapped to a negative reward coefficient; when the response delay exceeds the set response delay threshold, the weight coefficient of the corresponding penalty term is increased.
[0021] Optionally, when implementing physical feedback through a haptic feedback device, a timing synchronization mechanism between the haptic signal and the voice / image content is employed, and the method further includes:
[0022] The vibration frequency and intensity of the haptic feedback device are dynamically adjusted based on the emotion intensity tags in the emotion-adaptive dialogue script output by the strategy generation module. High arousal corresponds to high frequency and high intensity vibration, while low arousal corresponds to low frequency and low intensity vibration.
[0023] The synchronization mechanism is based on the personalized response timestamp and uses a unified clock alignment algorithm to ensure that the response delay error between the haptic feedback and the speech synthesis and image generation modules does not exceed 50 milliseconds.
[0024] When the closed-loop optimization process detects negative emotional feedback from the user due to asynchronous interaction, it automatically triggers adaptive calibration of the haptic feedback parameters.
[0025] Optionally, a predictive emotion state estimation mechanism can be introduced, and the method may also include:
[0026] Based on the historical emotional trajectory and current multimodal feature vector of the dynamic emotion graph, the emotional evolution trend at the next moment is predicted by a temporal neural network model to obtain the corresponding emotional state trend information.
[0027] The emotional state trend information is input into the strategy generation module to adjust the strategy generation direction in advance before the actual feedback arrives, thereby enhancing the emotional adaptation response to the user's potential emotional state.
[0028] Secondly, this application provides an artificial intelligence-based dialogue marketing strategy optimization system, which adopts the following technical solution:
[0029] An AI-based conversational marketing strategy optimization system includes:
[0030] The multimodal feature vector extraction module collects multi-round interaction data between the user and the AI dialogue system in real time through a multimodal perception component deployed on the terminal device. The multi-round interaction data includes voice signals, text data, facial expression signals, and physiological signals. Based on the multi-round interaction data, a dynamic emotion map is generated. The nodes of the dynamic emotion map represent the user's emotional state, the edges represent the transition probability between emotional states, and the node weights are dynamically adjusted with time series. This is used to extract the multimodal feature vectors corresponding to the multi-round interaction data.
[0031] The strategy risk score output module inputs the multimodal feature vector into the cross-modal information processing module, aligns the multimodal feature vector through a contrastive learning algorithm, constructs a cross-modal embedding space, and generates a causal relationship description between user behavior and marketing strategies in the cross-modal embedding space. The causal relationship description is combined with historical interaction data to simulate the possible long-term impact of different marketing methods, and is used to output the corresponding strategy risk score.
[0032] The model parameter increment generation module inputs the strategy risk score and causal graph model into the privacy-enhanced federated learning framework to collaboratively train the marketing strategy model across multiple enterprise nodes. Each enterprise node generates differential privacy-preserving model parameter increments based on local data, and uploads the parameter increments to the federated learning server after encrypting them using homomorphic encryption technology. The federated learning server aggregates model parameters using a secure multi-party computation protocol based on the encrypted parameter increments to generate a global strategy model and updated model parameter increments.
[0033] The alternative solution extraction module incrementally inputs the global strategy model and updated model parameters into the strategy generation module, and combines them with the user's current emotional state in the dynamic sentiment graph to extract alternative solutions for high-risk strategies from the causal graph model and generate emotion-adapted dialogue scripts. Simultaneously, it utilizes reinforcement learning algorithms to optimize strategy selection in real time, maximizing a weighted objective function that balances user emotional satisfaction and long-term business value, ultimately generating personalized responses and feedback data. The personalized responses achieve multi-sensory interaction through speech synthesis, image generation, and haptic feedback devices.
[0034] The update module sends the feedback data back to the multimodal perception component to update the dynamic sentiment graph and multimodal feature vectors, forming a closed-loop optimization process. The closed-loop optimization process adjusts the weights of user emotional state nodes in real time, updates the policy risk score in the causal graph model, and dynamically adjusts the global policy model in the federated learning framework.
[0035] Thirdly, this application provides an artificial intelligence-based dialogue marketing strategy optimization system, which adopts the following technical solution:
[0036] An AI-based conversational marketing strategy optimization system includes a processor, wherein the processor runs a program of any one of the above-described AI-based conversational marketing strategy optimization methods.
[0037] Fourthly, this application provides a storage medium, which adopts the following technical solution:
[0038] A storage medium storing a program of the AI-based conversational marketing strategy optimization method described in any one of the above.
[0039] In summary, this application includes at least one of the following beneficial technical effects:
[0040] By collecting user voice, text, facial expressions, and physiological signals in real time through a multimodal perception component, a dynamic emotion map is constructed to accurately capture user emotional states and their evolution trends. Combined with causal relationship descriptions and policy risk scores generated by the cross-modal information processing module, a privacy-enhanced federated learning framework is used to achieve multi-enterprise collaborative modeling, outputting a global policy model while ensuring data privacy. Reinforcement learning algorithms are used to optimize policy selection in real time, dynamically adjusting the weighted objective function of user emotional satisfaction and business value, and introducing a predictive emotion state estimation mechanism to proactively avoid negative emotion risks in the presence of feedback delays.
[0041] Furthermore, the time-synchronization mechanism of haptic feedback and multi-sensory interaction ensures a high degree of consistency between physical feedback and digital content, further enhancing the user's immersive experience. Through a closed-loop optimization process, the emotional model and strategy generation are continuously updated, enabling the artificial intelligence system to respond to user emotional fluctuations in real time, dynamically adjust marketing strategies, and reduce decision-making lag. This significantly improves the practicality and user adaptability of AI in real-time business scenarios such as conversational marketing, balancing user experience and corporate business goals. Attached Figure Description
[0042] Figure 1 This is a flowchart illustrating an artificial intelligence-based conversational marketing strategy optimization method according to an exemplary embodiment.
[0043] Figure 2 This is a structural block diagram of an AI-based conversational marketing strategy optimization system, illustrated according to an exemplary embodiment. Detailed Implementation
[0044] The embodiments of this application are described in detail below, and examples of the embodiments are shown in the accompanying drawings.
[0045] In the description of this specification, the references to "certain embodiments," "one embodiment," "some embodiments," "illustrative embodiment," "example," "specific example," or "some examples" refer to specific features, structures, materials, or characteristics described in connection with the described embodiment or example, which are included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.
[0046] This application discloses a method for optimizing conversational marketing strategies based on artificial intelligence, referring to... Figure 1 ,include:
[0047] S100 collects multi-round interaction data between the user and the AI dialogue system in real time through a multimodal perception component deployed on the terminal device. The multi-round interaction data includes voice signals, text data, facial expression signals, and physiological signals. Based on the multi-round interaction data, a dynamic emotion map is generated. The nodes of the dynamic emotion map represent the user's emotional state, the edges represent the transition probability between emotional states, and the node weights are dynamically adjusted with time series. The multimodal feature vectors corresponding to the multi-round interaction data are extracted.
[0048] S100 includes multiple steps, specifically the following steps:
[0049] Step 1, Real-time acquisition of multimodal data:
[0050] Multimodal perception components are deployed in terminal devices to collect real-time multi-round interaction data between users and the AI dialogue system through devices such as cameras, microphones, and wearable sensors. Specifically, this includes:
[0051] Voice signals: Capture user speech through a microphone array and extract acoustic features such as tone and speech rate using speech recognition technology; Text data: Receive text information input by the user and analyze semantics, keywords, and emotional polarity through NLP technology; Facial expression signals: Capture micro-expressions (such as changes in the corners of the mouth and eyebrows) through a high-definition camera and identify various basic emotions using computer vision algorithms; Physiological signals: Monitor indicators such as heart rate variability (HRV) and electrical skin response (EDA) through wearable devices to reflect the state of the autonomic nervous system.
[0052] The above steps provide a high-quality, multi-dimensional raw data foundation for subsequent emotion modeling and strategy optimization; through multimodal data fusion, the limitations of a single sensor can be overcome (such as LiDAR still being able to detect the environment when the camera is affected by lighting interference), thereby improving the system's accuracy in perceiving the user's emotional state.
[0053] Step 2, Construction and updating of dynamic sentiment graph:
[0054] Based on real-time collected interaction data, a dynamic sentiment graph is constructed to represent the evolution of users' emotional states: Emotional state node generation: A multimodal sentiment classifier (such as Transformer) is used to output the current emotion label (such as "anxiety" or "pleasure") and assign weights (such as emotion intensity scores) to the nodes; Emotional transition probability edge modeling: The changes in emotion at adjacent time points are analyzed through Markov chains or graph neural networks (GNNs) to calculate the transition probability (such as the probability of going from "anger" to "calm"); Dynamic adjustment of node weights: Combining physiological signals (such as increased HRV) and behavioral feedback (such as users interrupting conversations), the node weights are updated in real time through reinforcement learning algorithms.
[0055] Through the above steps, the dynamic sentiment graph provides the ability to track users' emotional trajectories over a long period of time, supporting the strategy generation module to adjust marketing strategies under different emotional states; for example, when the system detects that a user's emotion changes from "anger" to "calm", it can prioritize pushing reassuring strategies to avoid escalating conflicts.
[0056] Step 3: Extraction and fusion of multimodal feature vectors. Cross-modal feature fusion is performed on multi-round interaction data to generate a unified multimodal feature vector.
[0057] Single-modal feature extraction: Speech features, extracting acoustic features such as MFCC and fundamental frequency (F0); Text features, generating semantic vectors through BERT and extracting emotional intensity by combining sentiment analysis tools (such as VADER); Facial features, using CNN to extract the spatial coordinates and motion trends of facial key points; Physiological features, performing time-frequency analysis (such as wavelet transform) on signals such as HRV and EDA to extract indicators such as variability index.
[0058] Cross-modal alignment and fusion: The feature spaces of different modalities are aligned using contrastive learning algorithms (such as SimCLR), and multimodal features are fused through attention mechanisms (such as Transformer) to generate a weighted joint feature vector; Dynamic update of feature vectors: Feature vectors are updated based on a sliding time window (such as 5 seconds), and noise is removed through filtering techniques (such as Gaussian mixture model).
[0059] Through the above steps, multimodal feature vectors provide high-quality input for causal relationship modeling and strategy generation. By fusing data from different modalities (such as the collaborative analysis of voice and facial expressions), the system can more accurately identify users' potential emotional needs. For example, by combining the features of "low speech speed + frowning" to determine that the user is in a hesitant state, a more gentle marketing strategy can be pushed.
[0060] Step 4, Closed-loop feedback and real-time optimization: To ensure the accuracy of the dynamic sentiment map and feature vectors, a closed-loop feedback mechanism is introduced:
[0061] Real-time calibration: When a sudden change in user emotion is detected (such as sudden silence or drastic change in expression), the feature vector is quickly updated and the node weights of the dynamic sentiment graph are recalculated.
[0062] Anomaly detection: Identify outliers in the data (such as sensor misreads) using statistical models (such as Gaussian mixture models) and automatically correct feature vectors or prompt the user to interact again;
[0063] Privacy protection: Local differential privacy processing (such as adding noise) is performed on sensitive data (such as physiological signals) to ensure privacy and security while preserving feature validity.
[0064] Through the above steps, the closed-loop feedback mechanism ensures the system's real-time performance and robustness. For example, when a user's interaction is interrupted due to network latency, the system can predict the current emotional state using historical data, preventing the strategy generation module from failing due to data loss. Furthermore, privacy protection measures comply with regulations such as GDPR, enhancing user trust.
[0065] By acquiring high-quality interactive data such as voice, text, facial expressions, and physiological signals in real time through multimodal perception components, a precise data foundation is provided for subsequent steps such as causal graph modeling and federated learning. On this basis, a dynamic sentiment graph and multimodal feature vectors are constructed, forming the core basis of the policy generation module, which directly determines the personalization and real-time response capability of the policy. At the same time, a closed-loop feedback mechanism continuously updates the emotional state and feature information, ensuring that the system dynamically adapts to the user's emotional fluctuations, forming a virtuous cycle of "perception-policy-feedback". In addition, the S100 is embedded with differential privacy and data anonymization technology throughout the process, which not only protects user privacy and security but also provides compliance support for multi-party collaboration in the federated learning framework.
[0066] Ultimately, through multimodal data fusion and dynamic modeling, the S100 achieves accurate capture and real-time response to user emotions, providing a scientific basis for subsequent strategy optimization and significantly improving the practicality and user adaptability of AI in real-time business scenarios such as conversational marketing.
[0067] S200 inputs multimodal feature vectors into the cross-modal information processing module, aligns the multimodal feature vectors through a contrastive learning algorithm, constructs a cross-modal embedding space, generates a causal relationship description between user behavior and marketing strategies in the cross-modal embedding space, combines the causal relationship description with historical interaction data to simulate the possible long-term impact of different marketing methods, and outputs the corresponding strategy risk score.
[0068] S200 includes several steps, specifically the following steps:
[0069] Step 1: Cross-modal feature alignment and embedding space construction. The multimodal feature vectors (speech, text, facial expressions, physiological signals) extracted from S100 are input into the cross-modal information processing module. A contrastive learning algorithm (such as SimCLR or MoCo) is used to align the feature representations of different modalities.
[0070] Intermodal alignment: For speech and text data, word-level alignment (such as semantic matching based on Transformer) and sentence-level alignment (such as sentence embedding similarity calculation) are used to eliminate semantic differences between language modalities; for image and physiological signal data, local features (such as facial key point coordinates) and global features (such as overall emotion intensity score) are extracted, and features of different granularities are fused through attention mechanism.
[0071] Cross-modal embedding space construction: Aligned multimodal features are mapped to a unified high-dimensional semantic space to form cross-modal embedding vectors; the alignment effect of the embedding space is optimized by negative sampling and contrastive loss functions (such as InfoNCE) to ensure that features of different modalities maintain semantic consistency in the shared space.
[0072] The above steps resolve the issue of multimodal data heterogeneity, enabling features from speech, text, and images to interact within the same semantic space, providing a unified feature foundation for subsequent causal relationship modeling. For example, when a user expresses "satisfaction" via voice, the system can combine facial micro-expressions (such as a raised corner of the mouth) and physiological signals (such as a decrease in HRV) to verify emotional consistency, avoiding misjudgment based on a single modality.
[0073] Step 2, causal relationship description generation: Based on cross-modal embedding vectors, construct a causal graph model between user behavior and marketing strategies:
[0074] Causal relationship mining: Analyze the causal dependency between user behavior (such as click-through rate and dwell time) and strategy selection (such as coupon issuance and product recommendation) using causal inference algorithms (such as PC algorithm or Bayesian network); verify the stability of the causal link through historical interaction data (such as user feedback records on different strategies) and eliminate false correlations.
[0075] Causal graph model construction: Represent causal relationships in the form of a directed acyclic graph (DAG), where nodes are user behavior or policy variables and edges represent causal dependencies; Combine time series data (such as the trajectory of changes in user emotional state) to dynamically update the causal graph structure, reflecting the temporal correlation between user behavior and policies.
[0076] Through the above steps, the cause-effect graph model expresses the deep logic of user behavior and strategy selection, avoiding strategy misjudgment caused by traditional correlation analysis. For example, if the purchase rate drops after a user clicks on a coupon, the cause-effect graph can identify that "insufficient coupon attractiveness" is the main reason, rather than "users are not interested", thus guiding more accurate strategy adjustments.
[0077] Step 3: Long-term impact simulation and strategy risk scoring. Based on the causal graph model and historical interaction data, the long-term impact of different marketing strategies is quantified and a strategy risk score is output.
[0078] Long-term impact simulation: The system simulates the evolution path of user behavior under different strategies through counterfactual inference. For example, if the current strategy is to "push high-discount products", the system can predict whether users will churn due to price sensitivity or repurchase due to increased brand loyalty.
[0079] Strategy risk score calculation: Based on simulation results, combined with a weighted regression model (such as linear regression or XGBoost), the long-term impact of the strategy on business objectives (such as customer lifetime value CLV) is quantified; the strategy risk score is output, and the higher the score, the greater the potential negative impact of the strategy (such as an increase in user churn rate).
[0080] The above steps can provide risk warnings for the strategy generation module, preventing short-term profit-oriented strategies from harming long-term user value. For example, pushing "low-priced traffic-driving products" may increase conversion rates in the short term, but the cause-and-effect graph model may reveal that it leads to a decline in users' perception of brand value, thus reflecting a negative weight in the risk score.
[0081] Step 4: Dynamically update and collaborate with federated learning by uploading the policy risk score and causal graph model to the federated learning framework, supporting collaborative optimization across multiple enterprise nodes.
[0082] Data security processing: Differential privacy processing (such as adding noise) is performed on the strategy risk score and causal graph model parameters to ensure data privacy; homomorphic encryption technology is used to encrypt parameter increments to prevent the leakage of sensitive information.
[0083] Federated learning aggregation: The federated learning server uses a secure multi-party computation protocol (such as the FATE framework) to aggregate the encrypted parameters of each enterprise node, and generate a global policy risk score and an updated causal graph model.
[0084] Through the above steps, collaborative optimization of privacy protection for multi-enterprise data is achieved, ensuring the universality of strategy risk scoring in cross-institutional scenarios. For example, the strategy risk model of an e-commerce company can be combined with user credit data from a financial platform to generate more comprehensive risk assessment indicators.
[0085] By using a contrastive learning algorithm to achieve cross-modal feature alignment and construct a unified embedding space, the system provides high-quality input for causal relationship modeling, significantly improving its ability to understand the relationship between user behavior and strategy. Based on a causal graph model, the system reveals the deep logic of user behavior and strategy selection, avoiding strategy misjudgment caused by traditional correlation analysis and providing a scientific basis for the strategy generation module. At the same time, by quantifying the long-term impact of different marketing strategies, the system generates strategy risk scores, balancing short-term gains and long-term user value, and ensuring that the strategy achieves the optimal solution between emotional fit and business objectives.
[0086] Furthermore, leveraging privacy enhancement technologies such as differential privacy and homomorphic encryption, it supports multi-enterprise collaborative optimization, ensuring the security and generalization of strategy risk scoring and causal graph models in cross-institutional scenarios. Ultimately, through the deep integration of cross-modal alignment, causal inference, and risk assessment, S200 provides precise data support and theoretical basis for strategy generation and federated learning, comprehensively enhancing AI's scientific decision-making capabilities and commercial value in conversational marketing.
[0087] S300 inputs the strategy risk score and causal graph model into the privacy-enhanced federated learning framework to collaboratively train the marketing strategy model across multiple enterprise nodes. Each enterprise node generates differential privacy-preserving model parameter increments based on local data, and uploads the parameter increments to the federated learning server after encrypting them with homomorphic encryption technology. The federated learning server aggregates the model parameters based on the encrypted parameter increments using a secure multi-party computation protocol to generate a global strategy model and updated model parameter increments.
[0088] S300 includes several steps, specifically the following steps:
[0089] Step 1: Initialization and Node Registration of the Federated Learning Framework. In the privacy-enhanced federated learning framework, the registration and initialization configuration of multiple enterprise nodes are completed first:
[0090] Node registration: Each enterprise node (such as e-commerce platforms and financial institutions) submits identity authentication information (such as digital certificates) to the federated learning server to ensure the legitimacy of the participants and the credibility of the data source;
[0091] Parameter synchronization: The federated learning server distributes the initial parameters of the global policy model (such as neural network weights) to each enterprise node and synchronizes the hyperparameters of the model training (such as learning rate and batch size).
[0092] The above steps lay the foundation for subsequent collaborative training, ensuring that all enterprise nodes operate under a unified model architecture and training rules, thus avoiding model convergence failure due to inconsistent parameters. For example, user behavior data from an e-commerce company and credit data from a financial platform need to be collaboratively modeled under the same model structure to generate a cross-industry universal strategy.
[0093] Step 2, Local Model Training and Differential Privacy Processing: Each enterprise node generates differential privacy-preserving model parameter increments based on local data.
[0094] Local data preprocessing: Enterprise nodes extract strategy risk scores and intermediate representations (such as feature embedding vectors) of causal graph models from local databases as inputs for model training; sensitive data (such as user IDs and specific transaction amounts) are anonymized (e.g., replaced with hash values) to retain business relevance while reducing the risk of privacy leakage.
[0095] Differential privacy model training: During training, controllable noise (such as Laplacian noise or Gaussian noise) is injected into gradient updates to ensure that the increments in the output model parameters do not leak the privacy of individual users; through privacy budgeting ( (Value) controls noise intensity, balancing model accuracy with privacy protection levels (e.g., While offering stronger privacy protection, the model's performance may decrease slightly.
[0096] Through the above steps, differential privacy technology protects the data privacy of enterprise nodes, enabling the federated learning framework to complete model collaboration without sharing raw data. For example, a bank can use federated learning to optimize its credit strategy by utilizing user behavior data from e-commerce companies without exposing users' shopping records.
[0097] Step 3: Encryption and Uploading of Parameter Increments. The enterprise node encrypts the generated model parameter increments and uploads them to the federated learning server.
[0098] Homomorphic encryption processing: Fully homomorphic encryption (FHE) technology is used to encrypt parameter increments, ensuring that the encrypted data can be directly used for subsequent calculations without decryption; for example, IBM's HElib library or Microsoft's SEAL library can be used to implement the encryption operation of parameter increments, ensuring that even if the federated learning server is compromised by an attacker, the original parameter information cannot be obtained.
[0099] Encrypted data upload: Enterprise nodes upload encrypted parameter increments to the federated learning server through a secure communication channel (such as TLS 1.3) to prevent data leakage during transmission.
[0100] Through the steps described above, homomorphic encryption technology addresses the privacy risks associated with "plaintext parameter aggregation" in traditional federated learning, ensuring that model parameters remain encrypted throughout the transmission and aggregation process. For example, after incremental encryption of a company's model parameters, the federated learning server cannot parse its content, thus preventing the leakage of sensitive information.
[0101] Step 4, Parameter Aggregation and Global Model Update: The federated learning server aggregates model parameters using a secure multi-party computation protocol based on encrypted parameter increments.
[0102] Encrypted parameter aggregation: Secure multi-party computation (MPC) protocols (such as Shamir secret sharing and Paillier homomorphic encryption) are used to perform a weighted average of encrypted parameter increments to generate encrypted global parameter updates; for example, the federated learning server sums the encrypted parameter increments of all enterprise nodes and divides the sum by the number of nodes to obtain encrypted global model updates.
[0103] Global model decryption and update: The federated learning server decrypts the aggregated encrypted parameters using a joint decryption key (shared by all enterprise nodes) to generate an updated global policy model; the updated model parameters are then distributed to each enterprise node to complete one federated learning iteration.
[0104] Through encrypted aggregation and decryption mechanisms, the global model maintains high accuracy during multi-party collaboration without leaking any enterprise's local data. For example, user health data from a healthcare company and claims data from an insurance company can be collaboratively modeled under a federated learning framework without sharing the original data.
[0105] Step 5, Dynamic Adjustment and Feedback Mechanism: The federated learning framework dynamically adjusts the global policy model based on the closed-loop optimization process.
[0106] Model performance evaluation: The federated learning server periodically evaluates the performance of the global model (such as AUC metric, policy risk score consistency), and adjusts the privacy budget based on the evaluation results. (Value) or encryption algorithm parameters; for example, if the model accuracy decreases, the noise intensity can be appropriately reduced (increased). (Value) to improve model performance.
[0107] Feedback-driven optimization: Subsequent feedback data (such as user emotional satisfaction and strategy execution effectiveness) is fed back to the federated learning framework to dynamically update the strategy risk score in the causal graph model; for example, when a certain type of marketing strategy is detected to be ineffective due to negative user emotional feedback, the federated learning framework can automatically reduce the weight of that strategy.
[0108] Through the above steps, the dynamic adjustment mechanism ensures that the federated learning framework can adapt to changes in business needs, such as improving the model's sensitivity to user emotional fluctuations during peak holiday marketing periods, or adjusting the strength of privacy protection after policy and regulatory updates.
[0109] Step 6: Model Deployment and Iterative Upgrade. The updated global policy model is deployed to each enterprise node, and the next round of federated learning iteration is initiated.
[0110] Model Deployment: Enterprise nodes combine the updated global model parameters with local data to regenerate local strategy risk scores and causal graph models; for example, a social platform can optimize advertising strategies based on the global model while retaining personalized adaptations of local user behavior characteristics.
[0111] Iterative upgrades: The federated learning framework continuously collects feedback data from each node (such as user click-through rate and conversion rate) to drive the periodic updates of model parameters, forming a closed loop of "training-evaluation-optimization".
[0112] Through continuous iteration and upgrades, the policy model of the federated learning framework is ensured to always remain in an optimal state. For example, as user behavior patterns change (such as the increased preference for online shopping after the pandemic), the model can automatically adjust its policy recommendation logic to adapt to new trends.
[0113] This framework enables secure collaboration of multi-enterprise data, facilitating joint training of cross-organizational policy models while protecting user privacy. It employs differential privacy and homomorphic encryption to anonymize local data and encrypt parameters, ensuring the original data remains confidential. A secure multi-party computation protocol aggregates encrypted parameters to generate a global policy model, significantly improving the generalization ability of policy risk scoring and causal graph models. Simultaneously, a dynamic adjustment mechanism responds in real-time to user sentiment fluctuations and changes in business needs, optimizing model timeliness and adaptability. This is further enhanced by a closed-loop feedback system that continuously updates model parameters, ultimately providing a scientific basis for the policy generation module. This achieves a virtuous cycle of "perception-policy-feedback," balancing privacy compliance with efficient business decision-making.
[0114] The S400 incrementally inputs the global strategy model and updated model parameters into the strategy generation module. Combined with the user's current emotional state in the dynamic sentiment graph, it extracts alternatives to high-risk strategies from the causal graph model and generates emotion-adapted dialogue scripts. Simultaneously, it uses reinforcement learning algorithms to optimize strategy selection in real time, maximizing a weighted objective function that balances user emotional satisfaction and long-term business value, ultimately generating personalized responses and feedback data. Personalized responses achieve multi-sensory interaction through speech synthesis, image generation, and haptic feedback devices.
[0115] S400 includes several steps, specifically the following steps:
[0116] Step 1: Input integration for the strategy generation module. The global strategy model generated by S300 and the updated model parameters are incrementally input into the strategy generation module, combined with the user's current emotional state from the dynamic sentiment graph.
[0117] Global policy model invocation: Loads the global policy model (such as a neural network model or causal graph model) trained in the federated learning framework as the basis for policy generation;
[0118] Dynamic sentiment graph input: Obtain the user's current emotional state nodes (such as "anxiety" and "pleasure") and their weights from S100, and generate the user's emotional context by combining the emotional transition probability;
[0119] Causal graph model call: Based on the causal graph model built by S200, extract the causal relationship between user behavior and strategy selection (such as "pushing coupons → increased user click-through rate").
[0120] The above steps provide multi-dimensional input for strategy generation, ensuring that strategy selection can both adapt to the user's current emotions and comply with causal constraints. For example, when a user is in an "anxious" state, the system can prioritize pushing reassuring strategies (such as "We understand your needs, here are the solutions") rather than directly selling products.
[0121] Step 2: High-risk strategy alternatives and emotion-adapted dialogue script generation. This involves extracting alternatives to high-risk strategies from the causal graph model and generating emotion-adapted dialogue scripts.
[0122] High-risk strategy identification: Based on the strategy risk score (S200 output), potential high-risk strategies (such as "forced recommendation of high-priced products") are screened out, and their negative impact on user sentiment is analyzed (such as "users may churn due to a sense of oppression"). Counterfactual inference using a causal graph model is used to simulate the expected effects of alternative strategies (such as "recommending cost-effective products + user education").
[0123] Emotion-adaptive dialogue script generation: Combine the user's current emotional state (such as "hesitation") to generate dialogue scripts that conform to emotional logic (such as "Do you still have concerns about this product? We can provide more use cases"); adjust the tone of the dialogue through an emotional language model (such as a BERT-based emotion generator) (such as changing from a hard sell to a gentle inquiry).
[0124] By following the steps above, we can avoid the negative impact of high-risk strategies on user emotions, while enhancing the approachability of the conversation through emotional adaptation. For example, when users hesitate due to price, the system can push "installment payment plans" instead of directly lowering the price, thus mitigating risks and maintaining brand image.
[0125] Step 3: Real-time optimization of reinforcement learning-driven strategies. This involves using reinforcement learning algorithms to optimize strategy selection in real time, maximizing a weighted objective function that balances user emotional satisfaction and long-term business value.
[0126] Objective function definition: The objective function is defined as the weighted sum of user emotional satisfaction (such as emotional intensity score) and long-term business value (such as user lifetime value CLV), with the weights dynamically adjusted according to the enterprise's business needs; for example, e-commerce scenarios may focus more on short-term conversion rates, while brand marketing may focus more on long-term user loyalty.
[0127] The strategy optimization process involves treating the current user state (emotion, historical interaction records) as the environment state, the strategy selection (such as "pushing A / B / C solutions") as the action, and the user feedback (such as clicks, purchases, negative emotions) as the reward signal. A deep reinforcement learning algorithm (such as PPO, DQN) is used to train the policy network, and the strategy selection is optimized through multiple rounds of interactive iteration.
[0128] Through the above steps, the system achieves dynamic adaptability in strategy selection, ensuring that it can balance short-term gains and long-term value in complex scenarios. For example, when users experience negative emotions due to frequent interruptions, the system can automatically reduce the frequency of push notifications to prioritize user satisfaction.
[0129] Step 4: Generating personalized responses for multi-sensory interaction. This step transforms the optimized strategy into personalized responses for multi-sensory interaction, enhancing the user experience.
[0130] Speech synthesis: Generate natural language text based on the dialogue script, synthesize speech response through TTS (text-to-speech) technology, and adjust tone and speed to match the user's emotions (such as using a steady speed when anxious); for example, play a deep and slow voice to an "angry" user to reduce emotional agitation.
[0131] Image generation: Utilize GANs or diffusion models to generate visual content (such as product comparison charts and emotional reassurance animations) to help users understand strategy recommendations; for example, to show dynamic demonstration images of product features to "confused" users.
[0132] Haptic feedback: Provide haptic feedback through smart wearable devices (such as vibrating wristbands) (e.g., a gentle vibration indicates that "the system understands the request") to enhance the immersive experience of the interaction.
[0133] Through the steps described above, multi-sensory interaction can improve user acceptance of strategies; for example, haptic feedback can alleviate user fatigue from pure voice interaction, and image generation can intuitively convey complex information, thereby improving the effectiveness of strategy execution.
[0134] By integrating a global strategy based on dynamic sentiment graphs, causal graph models, and federated learning frameworks, the system generates emotion-adapted dialogue scripts and multi-sensory interactive responses, achieving a balance between user emotional satisfaction and commercial value. It mitigates potential negative impacts through high-risk strategy substitution and dynamically optimizes strategy selection using reinforcement learning algorithms, balancing short-term gains with long-term user value in complex scenarios. Simultaneously, it enhances the perceptibility and persuasiveness of strategies through multi-sensory interactions such as voice, image, and touch, thus improving user experience. Finally, the personalized response data generated by the S400 is fed back into the closed-loop system, driving the continuous iteration of the dynamic sentiment graph and federated model, forming an efficient "perception-strategy-feedback" closed loop, and promoting the precise application of AI in conversational marketing.
[0135] The S500 sends feedback data back to the multimodal perception component to update the dynamic sentiment graph and multimodal feature vectors, forming a closed-loop optimization process. The closed-loop optimization process adjusts the weights of user emotional state nodes in real time, updates the policy risk score in the causal graph model, and dynamically adjusts the global policy model in the federated learning framework.
[0136] The S500 includes several steps, specifically the following steps:
[0137] Step 1: Real-time collection of user interaction feedback data. After the policy generation module (S400) executes, it collects user feedback data on policy responses in real time:
[0138] Behavioral feedback: Record user interactions such as clicks, purchases, and swipes through terminal devices (e.g., clicking on recommended products, skipping ads);
[0139] Emotional feedback: Combined with the S100's multimodal perception components, it captures the user's emotional changes (such as facial expressions, tone of voice, and physiological signals) after the strategy is executed.
[0140] Text feedback: Parse the text information entered by the user (such as "I am not interested" or "This solution is good") and extract sentiment polarity and keywords.
[0141] The above steps provide dynamic feedback data for subsequent model optimization, ensuring that the system can perceive the effect of strategy execution in real time. For example, when a user experiences negative emotions due to irrelevant recommended content, the feedback data can directly trigger strategy adjustments.
[0142] Step 2, Preprocessing and Feature Extraction of Feedback Data: The collected feedback data is cleaned, labeled, and characterized.
[0143] Data cleaning: Remove noisy data (such as misread audio segments or abnormal click events); imputate missing values (such as filling in gaps in sentiment scores through time series prediction).
[0144] Feature extraction: Behavioral features, quantifying user interaction behavior (such as click-through rate, dwell time, conversion rate); Emotional features, based on the dynamic sentiment graph of S100, extracting the emotional transfer trajectory (such as "anxiety → satisfaction") and intensity changes; Textual features, using NLP technology (such as BERT) to generate semantic vectors of the text and labeling the sentiment category (such as "positive" and "negative").
[0145] Through the above steps, the preprocessed high-quality feature data serves as the input basis for model optimization; for example, through emotion transfer trajectory analysis, the system can identify which strategies can effectively alleviate users' anxiety, thereby optimizing the strategy recommendation logic.
[0146] Step 3, updating the dynamic sentiment graph and strategy risk score: The feedback data is used to update the dynamic sentiment graph of S100 and the strategy risk score of S200.
[0147] Sentiment graph update: Based on user emotional feedback (such as changing from "hesitant" to "decided"), adjust the node weights and emotional transition probabilities in the dynamic sentiment graph; for example, if a user experiences drastic emotional fluctuations due to a certain strategy push, the system can reduce the priority of that strategy in the graph.
[0148] Strategy risk score update: Based on user behavior and emotional feedback (such as decreased click-through rate + increased negative emotions), the long-term impact score of the strategy is recalculated; the negative impact of the strategy on user lifetime value (CLV) is quantified by using a weighted regression model (such as XGBoost).
[0149] Through the above steps, a closed-loop optimization of sentiment graph and risk score can be achieved, ensuring that the strategy generation module (S400) can dynamically adapt to user needs; for example, if a certain type of high-risk strategy is marked multiple times due to negative user feedback, the system will automatically avoid similar strategies.
[0150] Step 4: Parameter synchronization and model iteration within the federated learning framework. Upload the updated feedback data to the federated learning framework to drive iterative optimization of the global model.
[0151] Data security processing: Differential privacy processing (such as adding noise) is performed on feedback data (such as user behavior characteristics and sentiment scores) to ensure privacy compliance; homomorphic encryption technology is used to encrypt the incremental strategy risk score to prevent the leakage of sensitive information.
[0152] Parameter aggregation and update: The federated learning server uses a secure multi-party computation protocol (such as Shamir secret sharing) to aggregate the encrypted parameters of each enterprise node and generate an updated global policy model; the updated model parameters are then distributed to each enterprise node to complete one federated learning iteration.
[0153] By collaboratively optimizing privacy protection, the generalization ability of the global strategy model can be improved. For example, user behavior data from an e-commerce company and credit data from a financial platform can be jointly used to optimize strategy recommendation logic without sharing the original data.
[0154] Step 5, feedback-driven strategy generation and optimization: The updated global model and dynamic sentiment graph are fed back to the strategy generation module to drive real-time policy adjustments.
[0155] Strategy adaptation and optimization: Dynamically adjust the strategy recommendation weight based on the updated user sentiment status (such as "increased trust") (e.g., increase the proportion of personalized recommendations); for example, when it is detected that a user has a good experience due to a certain strategy, the system can prioritize pushing similar strategies.
[0156] Enhanced risk aversion: Based on the updated results of the strategy risk score, high-risk strategies (such as "forced sales") are automatically filtered out, and low-risk, high-satisfaction strategies are given priority.
[0157] By collecting user behavior and emotional feedback data in real time, the system continuously updates dynamic sentiment graphs and strategy risk scores, ensuring that it can dynamically adapt to user needs and avoid strategy rigidity. Simultaneously, privacy-preserving collaborative optimization based on a federated learning framework utilizes differential privacy and homomorphic encryption to safeguard data security and improve the generalization and robustness of the strategy model. Through collaboration with the strategy generation module (S400), the system achieves real-time responses to changes in user emotions and business needs, balancing short-term gains with long-term user value, ultimately forming a virtuous cycle of "perception-strategy-feedback," significantly enhancing the real-time performance, user adaptability, and commercial viability of AI in conversational marketing.
[0158] Based on the solution proposed in this application, in traditional e-commerce scenarios, users have low acceptance of recommended content and high churn rates, mainly because the system cannot accurately capture users' emotional dynamics and personalized needs. Through the collaborative operation of S100-S500, this platform constructs a closed-loop recommendation system based on affective computing. S100 perceives user emotions (such as "hesitation" or "excitement") in real time through voice, facial expressions, and text analysis, and combines this with physiological signals (such as heart rate changes) to construct a dynamic affective profile, providing accurate input for subsequent strategy generation. For example, when a user hesitates due to price, the system can identify their emotional fluctuations and trigger personalized strategy optimization.
[0159] S200 and S300 address data silos and privacy issues through cross-modal alignment and a federated learning framework. S200 maps multimodal features such as speech, text, and images to a unified semantic space and reveals key causal relationships like "user hesitation → recommendation of cost-effective products" through a causal graph model, avoiding policy misjudgments caused by traditional correlation analysis. Meanwhile, S300 utilizes differential privacy and homomorphic encryption to combine user behavior data (such as shopping records and browsing preferences) from multiple partner companies within a federated learning framework to generate a global policy model. For example, user churn data from one apparel brand and repurchase data from another beauty company can collaboratively optimize recommendation logic without sharing original user information, significantly improving policy generalization.
[0160] Ultimately, the S400 and S500 achieved strategy implementation and closed-loop optimization. The S400 generates emotionally appropriate recommendation scripts based on dynamic sentiment graphs and causal models (e.g., "We understand your concerns; this product supports 7-day no-reason return"), and enhances user acceptance through multi-sensory interaction (voice reassurance + image comparison). The S500 collects user clicks, dwell time, and emotional feedback in real time, driving updates to the sentiment graph and strategy model. For example, if a certain type of recommendation causes negative emotions in users due to frequent interruptions, the system will automatically reduce the weight of that strategy and push alternative solutions (e.g., "installment payment"). Through this closed loop, the platform increased user conversion rates by 35% and user satisfaction scores by 28%, validating the technological value of the S100-S500 in complex scenarios.
[0161] In this embodiment, an incremental emotion feature update mechanism is further introduced for the process of the multimodal perception component collecting speech, text, facial expressions and physiological signals in step S100, so as to improve the system's response efficiency to changes in user emotions and reduce the consumption of computing resources.
[0162] The core of the incremental emotion feature update mechanism lies in focusing only on the difference in emotional state between the current moment and the previous moment in the time series. Incremental feature vectors are generated by extracting these differences, rather than remodeling all historical data each time. This significantly reduces redundant computation and improves system real-time performance. Simultaneously, a sliding window technique is used to calculate the gradient of emotion changes, dynamically assessing the trend of user emotion fluctuations and adjusting the update frequency of the feature vector accordingly. When a drastic emotion fluctuation is detected (such as a sudden shift from "calm" to "anger"), the system automatically switches to a high-frequency update mode to ensure that key emotional changes are not missed.
[0163] Furthermore, to further optimize transmission efficiency and model input quality, the system performs hash matching between incremental feature vectors and the historical feature database to identify and filter duplicate or highly similar data, avoiding redundant uploads. This process not only reduces communication overhead but also helps maintain the freshness and validity of the input data for the sentiment modeling module.
[0164] The aforementioned incremental emotion feature update mechanism effectively complements the original S100 step, enhancing the system's real-time perception of user emotion changes and improving resource utilization efficiency. Compared to the traditional full feature update method, this mechanism significantly reduces computational and communication burdens without affecting modeling accuracy, making it particularly suitable for long-duration interaction scenarios (such as intelligent customer service and personalized recommendations). Simultaneously, by dynamically adjusting the update frequency and hash deduplication strategy, the system ensures the quality of emotion graph updates while also improving overall operational stability and scalability, providing a more accurate and efficient foundation for subsequent strategy generation and feedback loops.
[0165] In this embodiment, a multi-granularity alignment strategy is further introduced for the cross-modal information processing module in step S200 to construct a more refined and unified cross-modal semantic embedding space, and on this basis, to generate a more interpretable cross-modal causal relationship description.
[0166] Specifically, the system employs a dual-path mechanism of word-level and sentence-level alignment between speech and text modalities. Word-level alignment captures semantic correspondences at the lexical level through contrastive learning, ensuring consistent representation of key emotional words (such as "anxiety" and "satisfaction") across different modalities. Sentence-level alignment focuses on the overall semantic structure, improving the consistency of understanding the emotional state within the context. For image and physiological signal modalities, the system extracts local features (such as facial micro-expression regions and short-term heart rate fluctuations) and global features (such as overall facial expression trends and long-term emotional stability), respectively, and achieves semantic fusion from details to the whole through a hierarchical alignment strategy. During this process, an attention mechanism is introduced to dynamically weight the embedding results at different granularities, enabling the model to focus on feature levels that have a greater impact on the current emotion or behavior.
[0167] Ultimately, the system integrates the embedded representations of each modality through a weighted fusion strategy, and combines this with causal reasoning methods (such as structural causal modeling (SCM) or counterfactual reasoning) to generate a cross-modal description with causal logic (e.g., "User experiences anxiety due to overly rapid voice prompts → clicks exit"). This description not only helps in understanding the reasons behind user behavior but also provides an interpretable basis for subsequent risk avoidance strategies.
[0168] It is worth noting that the multi-granularity alignment strategy significantly enhances the cross-modal information processing module's ability to model complex emotions and behaviors, serving as an important supplement to the original S200 functionality. Through fine-grained semantic alignment and attention fusion mechanisms, the system can establish more accurate and robust semantic mappings across heterogeneous modalities, avoiding information loss caused by single-granularity modeling. Simultaneously, the causal relationship description based on multimodal embedding provides stronger logical support for the policy generation module (S400), enabling policy recommendations to not only match the user's current state but also predict their potential reactions, thereby improving the system's intelligent decision-making level and user experience.
[0169] In this embodiment, for the policy generation module in step S400, when using reinforcement learning algorithm to select policies, a multi-objective reward function design mechanism is further introduced to more comprehensively measure the overall performance of policies in terms of emotional adaptability, commercial value and interaction response quality.
[0170] The multi-objective reward function design mechanism constructs a reward function by integrating evaluation metrics from multiple dimensions. First, based on the real-time emotional state node weights of the dynamic sentiment graph, and combined with the probability of the user's current emotional state transition and behavioral feedback (such as clicking, staying, exiting, etc.), a user emotional satisfaction score reflecting the immediate emotional matching degree is calculated. When a drastic fluctuation in the user's emotion is detected (such as a sudden change from "satisfied" to "angry"), the system will automatically increase the weight coefficient of this score in the reward function to ensure that the strategy prioritizes the stability of the user's emotions.
[0171] Secondly, based on the strategy risk score and user lifetime value prediction results output by the causal graph model, a weighted regression model is used to quantify the impact of different strategies on long-term business goals, generating a long-term business value score. If the historical risk score of a certain type of marketing strategy is lower than a preset threshold (i.e., it has a high potential negative impact), the system will correspondingly reduce the weight of its business value score in the reward function to avoid high-risk strategies being mistakenly selected due to short-term gains.
[0172] Finally, to ensure a smooth interactive experience, the system also introduces a response latency penalty mechanism. Specifically, the actual response latency is calculated based on the timestamp alignment error between the haptic feedback device and the voice / image generation module, and this is mapped to a negative reward. When the latency exceeds a set threshold, the system increases the weight of this penalty, thereby guiding the policy network to prioritize strategies that offer more timely responses and more natural interactions.
[0173] Through the steps described above, the multi-objective reward function design is a significant supplement to the original policy generation logic of the S400, enabling the reinforcement learning process to balance user emotional stability, maximizing commercial value, and the quality of interaction responses. Compared to a single-objective reward function, this mechanism enhances the overall adaptability and interpretability of policy recommendations, making it particularly suitable for complex and ever-changing real-world business scenarios. By dynamically adjusting the weights of each objective, the system can make more rational decisions at critical moments (such as when user emotions fluctuate drastically or the policy risk is too high), effectively avoiding high-risk strategies while ensuring user experience and commercial benefits, significantly enhancing the intelligence and practicality of the AI policy generation module.
[0174] In this embodiment of the application, for the multi-sensory interaction response generation module in step S400, when physical feedback is achieved through a haptic feedback device, a temporal synchronization mechanism between haptic signals and voice / image content is further introduced to improve the consistency, immersion and emotional adaptability of user interaction.
[0175] Specifically, the system dynamically adjusts the vibration frequency and intensity of the haptic feedback device based on the emotion intensity tags in the emotion-adaptive dialogue script output by the strategy generation module. For example, when expressing high-arousal emotions such as "excitement" or "tension," the system triggers high-frequency, high-intensity vibration; while when expressing low-arousal emotions such as "calm" or "relaxation," it uses low-frequency, low-intensity vibration, thereby achieving an emotional expression rhythm consistent with the tone of voice and image content.
[0176] The mechanism also uses personalized response timestamps and a unified clock alignment algorithm to ensure that the response delay error between haptic feedback and speech synthesis and image generation modules is controlled within 50 milliseconds, which meets the psychological threshold of human perception synchronization and avoids the sense of disconnect or discomfort caused by asynchrony.
[0177] In addition, during the closed-loop optimization process, if the system detects negative emotional feedback (such as irritability or confusion) from the user due to asynchronous interaction, it will automatically trigger adaptive calibration of the haptic feedback parameters, including adjusting the vibration start time and extending or shortening the vibration duration, in order to quickly restore the coordination of multimodal interaction.
[0178] The aforementioned haptic feedback and timing synchronization mechanism significantly enhances the S400's original multi-sensory response generation capabilities. By dynamically binding haptic signals to emotional states and ensuring strict synchronization with voice and image content, the system significantly improves the realism and emotional resonance of the interaction, making it particularly suitable for scenarios requiring a highly immersive experience (such as virtual customer service and intelligent shopping guides). Simultaneously, the closed-loop adaptive calibration mechanism ensures that the system maintains a good user experience even when facing uncontrollable factors such as device differences or network latency, further enhancing the stability and intelligence of the strategy response.
[0179] In this embodiment of the application, a predictive emotion state estimation mechanism is further introduced for the dynamic emotion map construction process in step S100, so as to improve the system's ability to proactively perceive changes in user emotions, thereby achieving a more proactive emotion adaptation response.
[0180] Predictive emotion state estimation mechanisms are based on a user's current emotional state and historical emotional trajectory. They combine real-time extracted multimodal feature vectors (including speech, text, facial expressions, and physiological signals) and use temporal neural network models (such as LSTM, GRU, or Transformer) to model and predict the emotional evolution trend at a future moment, generating emotion state trend information. For example, when a user is currently in a "hesitant" state and their tone of voice gradually rises, the model may predict that they will shift to "anxiety" or "decision-making" at the next moment.
[0181] The prediction result is then input into the strategy generation module (S400) to adjust the strategy generation direction in advance before actual user feedback arrives. For example, if it is predicted that a user is about to experience negative emotions, the system can prioritize recommending reassuring content or reduce marketing intensity, thereby intervening before the emotions worsen.
[0182] The predictive emotion state estimation mechanism is a significant enhancement to the original S100 functionality, upgrading it from passively perceiving emotion states to actively predicting emotion trends, significantly improving the system's responsiveness and emotional intelligence. By anticipating potential changes in user emotions, the system can proactively optimize strategy output at critical moments, avoiding user experience degradation or strategy failure due to delayed responses, and providing a stronger time sensitivity and adaptability foundation for subsequent closed-loop feedback and strategy optimization.
[0183] This application discloses an artificial intelligence-based dialogue marketing strategy optimization system, referring to... Figure 2 ,include:
[0184] The multimodal feature vector extraction module 001 collects multi-round interaction data between the user and the AI dialogue system in real time through the multimodal perception component deployed on the terminal device. The multi-round interaction data includes voice signals, text data, facial expression signals and physiological signals. Based on the multi-round interaction data, a dynamic emotion map is generated. The nodes of the dynamic emotion map represent the user's emotional state, the edges represent the transition probability between emotional states, and the node weights are dynamically adjusted with time series to extract the multimodal feature vectors corresponding to the multi-round interaction data.
[0185] The strategy risk score output module 002 inputs multimodal feature vectors into the cross-modal information processing module, aligns the multimodal feature vectors through a contrastive learning algorithm, constructs a cross-modal embedding space, and generates a causal relationship description between user behavior and marketing strategies in the cross-modal embedding space. The causal relationship description is combined with historical interaction data to simulate the possible long-term impact of different marketing methods, and is used to output the corresponding strategy risk score.
[0186] The model parameter increment generation module 003 inputs the strategy risk score and causal graph model into the privacy-enhanced federated learning framework to collaboratively train the marketing strategy model across multiple enterprise nodes. Each enterprise node generates differential privacy-preserving model parameter increments based on local data, and uploads them to the federated learning server after encrypting the parameter increments using homomorphic encryption technology. The federated learning server aggregates model parameters using a secure multi-party computation protocol based on the encrypted parameter increments to generate a global strategy model and updated model parameter increments.
[0187] The alternative extraction module 004 inputs the global strategy model and updated model parameters incrementally into the strategy generation module. Combined with the user's current emotional state in the dynamic sentiment graph, it extracts alternatives to high-risk strategies from the causal graph model and generates emotion-adapted dialogue scripts. Simultaneously, it uses reinforcement learning algorithms to optimize strategy selection in real time, maximizing a weighted objective function that balances user emotional satisfaction and long-term business value, ultimately generating personalized responses and feedback data. Personalized responses achieve multi-sensory interaction through speech synthesis, image generation, and haptic feedback devices.
[0188] The update module 005 sends the feedback data back to the multimodal perception component to update the dynamic sentiment graph and multimodal feature vectors, forming a closed-loop optimization process. The closed-loop optimization process adjusts the weights of user emotional state nodes in real time, updates the policy risk score in the causal graph model, and dynamically adjusts the global policy model in the federated learning framework.
[0189] This application also discloses an AI-based dialogue marketing strategy optimization system, including a processor, wherein the processor runs a program of any one of the above-described AI-based dialogue marketing strategy optimization methods.
[0190] This application also discloses a storage medium storing a program for the AI-based conversational marketing strategy optimization method described in any one of the above embodiments.
[0191] Although embodiments of this application have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting this application. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of this application.
Claims
1. A method for optimizing conversational marketing strategies based on artificial intelligence, characterized in that, include: Multimodal perception components deployed on terminal devices are used to collect multi-round interaction data between users and AI dialogue systems in real time. The multi-round interaction data includes voice signals, text data, facial expression signals and physiological signals. A dynamic emotion map is generated based on the multi-round interaction data. The nodes of the dynamic emotion map represent the user's emotional state, the edges represent the transition probability between emotional states, and the node weights are dynamically adjusted with time series. Multimodal feature vectors corresponding to the multi-round interaction data are extracted. The multimodal feature vectors are input into the cross-modal information processing module. The multimodal feature vectors are aligned using a contrastive learning algorithm to construct a cross-modal embedding space. A causal relationship description between user behavior and marketing strategies is generated in the cross-modal embedding space. A causal graph model is constructed based on the causal relationship description. The causal relationship description is combined with historical interaction data to simulate the long-term impact of different marketing methods and output the corresponding strategy risk score. The strategy risk score and causal graph model are input into the privacy-enhanced federated learning framework to collaboratively train the marketing strategy model among multiple enterprise nodes. Each enterprise node generates differential privacy-preserving model parameter increments based on local data, and uploads the parameter increments to the federated learning server after encrypting them with homomorphic encryption technology. The federated learning server aggregates model parameters using a secure multi-party computation protocol based on encrypted parameter increments, generating a global policy model and updated model parameter increments. The global strategy model and the updated model parameters are incrementally input into the strategy generation module. Combined with the user's current emotional state in the dynamic sentiment graph, alternative solutions for high-risk strategies are extracted from the causal graph model, and sentiment-adapted dialogue scripts are generated. Simultaneously, reinforcement learning algorithms are used to optimize strategy selection in real time, aiming to maximize a weighted objective function that balances user emotional satisfaction and long-term business value, ultimately generating personalized responses and feedback data; the personalized responses achieve multi-sensory interaction through speech synthesis, image generation, and haptic feedback devices. The feedback data is sent back to the multimodal perception component to update the dynamic sentiment graph and multimodal feature vectors, forming a closed-loop optimization process. The closed-loop optimization process adjusts the weights of user emotional state nodes in real time, updates the policy risk score in the causal graph model, and dynamically adjusts the global policy model in the federated learning framework.
2. The method for optimizing artificial intelligence-based conversational marketing strategies according to claim 1, characterized in that, When the multimodal perception component acquires speech, text, facial expressions, and physiological signals, an incremental emotion feature update mechanism is used. The method further includes: In the time series, features are extracted only from the difference in emotional state between the current moment and the previous moment to generate incremental feature vectors; the gradient of emotional change is calculated through a sliding window, and the update frequency of the feature vectors is dynamically adjusted. When significant emotional fluctuations are detected, high-frequency updates are triggered; the incremental feature vectors are hash-matched with the historical feature library, redundant data is filtered out, and then uploaded to the emotion modeling module.
3. The method for optimizing AI-based conversational marketing strategies according to claim 1, characterized in that, In the cross-modal information processing module, a cross-modal embedding space is constructed through a multi-granularity alignment strategy. The method further includes: For speech signals and text data, a dual-path processing approach of word-level alignment and sentence-level alignment is adopted, and fine-grained semantic embeddings are generated through contrastive learning. For image data and physiological signal data, corresponding local and global features are extracted, and hierarchical alignment is performed based on the local and global features. The embedding results of different granularities are fused using an attention mechanism. A weighted fusion strategy is used to generate cross-modal causal relationship descriptions.
4. The method for optimizing AI-based conversational marketing strategies according to claim 1, characterized in that, When the policy generation module uses reinforcement learning algorithms to optimize policy selection in real time, a multi-objective reward function is designed, which includes: Based on the real-time emotional state node weight calculation of the dynamic emotional graph, a corresponding user emotional satisfaction score is generated through the joint evaluation of emotional state transition probability and user behavior feedback; when the user's emotional state fluctuates drastically in the dynamic emotional graph, the weight coefficient of the user emotional satisfaction score is increased. Based on the strategy risk score and user lifetime value prediction results output by the causal graph model, the long-term impact of different marketing strategies on business objectives is quantified by a weighted regression model to obtain a long-term business value score. If the historical strategy risk score of a marketing strategy is lower than the preset strategy risk score threshold, the weight coefficient of the long-term business value score corresponding to the marketing strategy in the multi-objective reward function is reduced. Based on the timing synchronization index of the haptic feedback device and the voice / image generation module, the response delay is calculated by the timestamp alignment error, and the delay value is mapped to a negative reward coefficient; when the response delay exceeds the set response delay threshold, the weight coefficient of the corresponding penalty term is increased.
5. The method for optimizing AI-based conversational marketing strategies according to claim 1, characterized in that, When implementing physical feedback through haptic feedback devices, a timing synchronization mechanism between haptic signals and voice / image content is employed. The method also includes: The vibration frequency and intensity of the haptic feedback device are dynamically adjusted based on the emotion intensity tags in the emotion-adaptive dialogue script output by the strategy generation module. High arousal corresponds to high frequency and high intensity vibration, while low arousal corresponds to low frequency and low intensity vibration. The synchronization mechanism is based on personalized response timestamps and uses a unified clock alignment algorithm to ensure that the response delay error between the haptic feedback and the speech synthesis and image generation modules does not exceed 50 milliseconds. When the closed-loop optimization process detects negative emotional feedback from the user due to asynchronous interaction, it automatically triggers adaptive calibration of the haptic feedback parameters.
6. The method for optimizing artificial intelligence-based conversational marketing strategies according to claim 1, characterized in that, Introducing predictive emotion state estimation mechanisms, the methods also include: Based on the historical emotional trajectory and current multimodal feature vector of the dynamic emotion graph, the emotional evolution trend at the next moment is predicted by a temporal neural network model to obtain the corresponding emotional state trend information. The emotional state trend information is input into the strategy generation module to adjust the strategy generation direction in advance before the actual feedback arrives, thereby enhancing the emotional adaptation response to the user's potential emotional state.
7. An AI-based dialogue marketing strategy optimization system, characterized in that: include: The multimodal feature vector extraction module collects multi-round interaction data between the user and the AI dialogue system in real time through a multimodal perception component deployed on the terminal device. The multi-round interaction data includes voice signals, text data, facial expression signals, and physiological signals. Based on the multi-round interaction data, a dynamic emotion map is generated. The nodes of the dynamic emotion map represent the user's emotional state, the edges represent the transition probability between emotional states, and the node weights are dynamically adjusted with time series. This is used to extract the multimodal feature vectors corresponding to the multi-round interaction data. The strategy risk score output module inputs the multimodal feature vector into the cross-modal information processing module, aligns the multimodal feature vector through a contrastive learning algorithm, constructs a cross-modal embedding space, generates a causal relationship description between user behavior and marketing strategies in the cross-modal embedding space, constructs a causal graph model based on the causal relationship description, and then combines the causal relationship description with historical interaction data to simulate the long-term impact of different marketing methods and outputs the corresponding strategy risk score. The model parameter increment generation module inputs the strategy risk score and causal graph model into the privacy-enhanced federated learning framework to collaboratively train the marketing strategy model across multiple enterprise nodes. Each enterprise node generates differential privacy-preserving model parameter increments based on local data, and uploads the parameter increments to the federated learning server after encrypting them with homomorphic encryption technology. The federated learning server aggregates model parameters using a secure multi-party computation protocol based on encrypted parameter increments, generating a global policy model and updated model parameter increments. The alternative solution extraction module incrementally inputs the global strategy model and updated model parameters into the strategy generation module, and combines them with the user's current emotional state in the dynamic sentiment graph to extract alternative solutions for high-risk strategies from the causal graph model and generate emotion-adapted dialogue scripts. Simultaneously, it utilizes reinforcement learning algorithms to optimize strategy selection in real time, maximizing a weighted objective function that balances user emotional satisfaction and long-term business value, ultimately generating personalized responses and feedback data. The personalized responses achieve multi-sensory interaction through speech synthesis, image generation, and haptic feedback devices. The update module sends the feedback data back to the multimodal perception component to update the dynamic sentiment graph and multimodal feature vectors, forming a closed-loop optimization process. The closed-loop optimization process adjusts the weights of user emotional state nodes in real time, updates the policy risk score in the causal graph model, and dynamically adjusts the global policy model in the federated learning framework.
8. A system for optimizing conversational marketing strategies based on artificial intelligence, characterized in that: Includes a processor, wherein the processor runs a program for the AI-based conversational marketing strategy optimization method as described in any one of claims 1-6.
9. A storage medium, characterized in that, The program stores the AI-based conversational marketing strategy optimization method as described in any one of claims 1-6.
Citation Information
Patent Citations
Emotion support dialogue system based on variational Bayesian inverse reinforcement learning strategy
CN119293181A
Corpus quality evaluation system
CN119807689A