Automatic marketing dialogue method and device based on reinforcement learning

By combining user historical data and real-time interactive data, and using reinforcement learning to optimize marketing strategies, the problem that marketing strategies in the existing technology cannot respond to changes in user needs in a timely manner, achieving higher strategy accuracy and marketing effects.

CN120046729APending Publication Date: 2025-05-27GUANGZHOU YUNDI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510102313.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-22
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The existing automated marketing technology relies on historical data to generate marketing strategies, which leads to the inability to respond to changes in user needs in a timely manner, with low accuracy and reduced results.

Method used

By obtaining the user's historical marketing data and real-time interactive text data, an initial marketing strategy is generated, and based on real-time interactive text data and dialogue emotional data, combined with reinforcement learning algorithm optimization strategies, dynamically generate and instantly adjust marketing strategies.

Benefits of technology

It improves the accuracy and effectiveness of marketing strategies, enhances the effectiveness and accuracy of automated marketing conversations, and can quickly respond to changes in user needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120046729A_ABST
    Figure CN120046729A_ABST
Patent Text Reader

Abstract

The invention discloses an automatic marketing dialogue method and device based on reinforcement learning, and relates to the technical field of language text processing. The dialogue emotion data is obtained according to the real-time interaction text data of the user, so that the emotion behavior track of the user is generated, and the optimal marketing strategy is generated according to the real-time interaction text data and the emotion behavior track, so that the optimal marketing strategy not only can fully fuse real-time preference and demand change of the user, but also can generate the optimal marketing strategy. In this way, the historical marketing situation of the user can be considered, dynamic generation and instant adjustment of the marketing strategy are achieved, the marketing strategy can quickly respond to user demand changes, user demand deviation caused by the fact that in the prior art, the marketing strategy is generated only depending on historical marketing data is avoided, and the user experience is improved. Therefore, the accuracy and the marketing effect of the marketing strategy are improved, and the effectiveness and the accuracy of the automatic marketing dialogue are further improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of language text processing, and particularly relates to an automated marketing dialogue method and device based on reinforcement learning. Background Art

[0002] With the continuous development of artificial intelligence and big data technologies, currently more and more industries rely on data-driven and intelligent means to improve marketing efficiency and user experience. Especially in the field of customer relationship management (CRM), through artificial intelligence and language processing technologies, automated marketing dialogues have been realized in the form of generating automated marketing strategies.

[0003] In current automated marketing technologies, usually big data resources and language models are used to generate automated marketing strategies by means of artificial intelligence technology. Although this method can quickly and automatically generate marketing strategies to achieve automated marketing dialogues, the marketing strategies obtained by this method are trained based on historical data. Once the user's preferences and behavior patterns change, the marketing strategies generated in this way will deviate from the user's needs, resulting in low accuracy of automated marketing dialogues and a decline in the effectiveness of automated marketing strategies. In addition, since current automated marketing strategies mostly rely on the statistical laws of historical data and ignore the instant interaction with users during the implementation of marketing strategies, this also greatly reduces the effectiveness of marketing strategies. Therefore, there is an urgent need for an automated marketing dialogue method and device based on reinforcement learning to solve the defects of the existing technology. Summary of the Invention

[0004] The present invention aims to provide an automated marketing dialogue method and device based on reinforcement learning to solve the above technical problems, and improve the low accuracy of automated marketing dialogues and marketing effectiveness through the user's real-time interaction text data and dialogue emotion data.

[0005] To solve the above technical problems, an embodiment of the present invention provides an automated marketing dialogue method based on reinforcement learning, including:

[0006] Obtain the user's historical marketing data, and perform tagging processing on the historical marketing data to determine user characteristic data;

[0007] Obtain the user's real-time interaction text data, and generate an initial marketing strategy according to the real-time interaction text data and user characteristic data;

[0008] Obtain the user's dialogue emotion data according to the real-time interaction text data, and generate an emotional behavior trajectory of the user based on the dialogue emotion data;

[0009] Optimize the initial marketing strategy based on the real-time interactive text data and the emotional behavior trajectory, and combine with a preset reinforcement learning algorithm to obtain an optimal marketing strategy;

[0010] Generate a marketing response text according to the optimal marketing strategy and the real-time interactive text data, and conduct an automated marketing conversation with the user based on the marketing response text.

[0011] It can be understood that, compared with the prior art, the present invention generates an initial marketing strategy through user feature data and real-time interactive text data, then obtains the user's dialogue emotion data through the real-time interactive text data to generate the user's emotional behavior trajectory, and then optimizes the initial marketing strategy based on the real-time interactive text data and the emotional behavior trajectory, combined with a reinforcement learning algorithm, to obtain an optimal marketing strategy, realizing an automated marketing conversation with the user. The present invention obtains dialogue emotion data from the user's real-time interactive text data, thereby generating the user's emotional behavior trajectory, and generates an optimal marketing strategy based on the real-time interactive text data and the emotional behavior trajectory, so that the optimal marketing strategy can not only fully integrate the user's real-time preferences and demand changes, but also consider the user's historical marketing situation, realizing the dynamic generation and immediate adjustment of the marketing strategy, enabling the marketing strategy to quickly respond to the user's demand changes, avoiding the deviation from the user's needs caused by generating marketing strategies solely relying on historical marketing data in the prior art, thereby improving the accuracy of the marketing strategy and the marketing effect, and further improving the effectiveness and accuracy of the automated marketing conversation.

[0012] As a preferred solution, the obtaining of the user's real-time interactive text data and generating an initial marketing strategy according to the real-time interactive text data and the user feature data specifically includes:

[0013] Obtain the user's historical interactive text data and historical marketing feedback data;

[0014] Train a preset marketing strategy recommendation model according to the historical interactive text data, historical marketing feedback data and user feature data to determine an optimal marketing strategy recommendation model;

[0015] Input the real-time interactive text data and the user feature data into the optimal marketing strategy recommendation model to generate an initial marketing strategy.

[0016] This preferred solution trains a preset marketing strategy recommendation model through historical interactive text data, historical marketing feedback data and user feature data, enabling the optimal marketing strategy recommendation model to fully consider the user's real-time preferences and demand changes, so that the initial marketing strategy can fully integrate the user's real-time preferences and demand changes, thereby improving the accuracy of the marketing strategy and the marketing effect, and further improving the effectiveness and accuracy of the automated marketing conversation.

[0017] As a preferred solution, training the preset marketing strategy recommendation model based on the historical interaction text data, historical marketing feedback data, and user characteristic data to determine the optimal marketing strategy recommendation model specifically includes:

[0018] Obtain the strategy data in the preset strategy database, and perform label encoding processing on the strategy data to determine the first strategy label data;

[0019] Input the historical marketing feedback data, historical interaction text data, user characteristic data, and the first strategy label data into the preset marketing strategy recommendation model, train the preset marketing strategy recommendation model in combination with a preset activation function and a preset loss function, and evaluate the training result of the preset marketing strategy recommendation model in combination with a preset model evaluation index;

[0020] Determine that the evaluation result of the training result of the preset marketing strategy recommendation model meets the preset requirements, complete the training of the preset marketing strategy recommendation model, and obtain the optimal marketing strategy recommendation model.

[0021] This preferred solution trains the preset marketing strategy recommendation model through historical interaction text data, user characteristic data, and the first strategy label data, enabling the optimal marketing strategy recommendation model to fully consider the real-time preferences and demand changes of users, thereby improving the accuracy of the initial marketing strategy output by the optimal marketing strategy recommendation model and the marketing effect, and further improving the effectiveness and accuracy of automated marketing conversations.

[0022] As a preferred solution, obtaining the dialogue emotion data of the user according to the real-time interaction text data, and generating the emotional behavior trajectory of the user based on the dialogue emotion data specifically includes:

[0023] The real-time interaction text data includes: user output text data and text data for replying to the user;

[0024] Analyze the user output text data according to a preset prompt word extraction algorithm to determine the dialogue emotion data of the user;

[0025] Convert the text data for replying to the user to determine the user behavior response data;

[0026] Pair the dialogue emotion data, user behavior response data, and user characteristic data to generate the emotional behavior trajectory of the user.

[0027] This preferred solution analyzes the user's output text data and the text data for replying to the user, generates the user's emotional behavior trajectory, and thus can fully represent the user's real-time preferences and demand changes through the emotional behavior trajectory. Furthermore, it improves the accuracy and marketing effect of the subsequent generated optimal marketing strategy, realizes the dynamic generation and immediate adjustment of the marketing strategy, and further improves the effectiveness and accuracy of the automated marketing dialogue.

[0028] As a preferred solution, analyzing the user's output text data according to the preset prompt word extraction algorithm to determine the user's dialogue emotional data specifically includes:

[0029] Inputting the user's output text data into a preset prompt word model to analyze the user's output text data, and obtaining the user's emotional category and user's emotional intensity;

[0030] Assigning a continuous value to the user's emotional intensity to obtain the user's emotional intensity value;

[0031] Obtaining a prompt word template, and modifying the prompt word template according to the real-time interaction text data, the user's emotional category and the user's emotional intensity value;

[0032] Determining the user's dialogue emotional data based on the modified prompt word template.

[0033] This preferred solution can obtain more accurate user emotional categories and user emotional intensities through the prompt word model. The form of continuous value assignment can avoid information loss caused by traditional discrete classification, improve the accuracy of user emotion evaluation, and improve the accuracy of the user's dialogue emotional data; the prompt word template can simplify the process of obtaining the user's dialogue emotional data and improve the efficiency of obtaining the user's dialogue emotional data.

[0034] As a preferred solution, converting the text data for replying to the user to determine the user's behavior response data specifically includes:

[0035] Performing data mapping and encoding on the text data for replying to the user based on the user feature data and the real-time interaction text data to determine the behavior selection data corresponding to the text data for replying to the user;

[0036] Querying a preset user behavior response database according to the behavior selection data corresponding to the text data for replying to the user to determine the user's behavior response data.

[0037] In this preferred solution, the response to the user's text data is mapped and encoded using user feature data and real-time interaction text data, making the resulting user behavior response data more in line with the user's real-time preferences and changing needs. This improves the accuracy of subsequent optimal marketing strategies and marketing effects, and further enhances the effectiveness and accuracy of automated marketing conversations.

[0038] As a preferred solution, based on the real-time interaction text data and emotional behavior trajectory, the initial marketing strategy is optimized by combining a preset reinforcement learning algorithm to obtain an optimal marketing strategy, which specifically includes:

[0039] Convert the real-time interaction text data into a dialogue history background information vector based on a preset text conversion algorithm;

[0040] Generate an emotional state vector, a user behavior response vector, and a user individual feature vector based on the emotional behavior trajectory;

[0041] Construct a marketing strategy state space according to the dialogue history background information vector, emotional state vector, user behavior response vector, and user individual feature vector;

[0042] Optimize the initial marketing strategy according to a preset reinforcement learning algorithm and the marketing strategy state space to obtain an optimal marketing strategy.

[0043] In this preferred solution, a marketing strategy state space is constructed through a dialogue history background information vector, an emotional state vector, a user behavior response vector, and a user individual feature vector, which can represent the user's real-time preferences and changing needs in the form of a state space. This improves the accuracy and marketing effect of the optimal marketing strategy obtained based on the state space and reinforcement learning algorithm, and further enhances the effectiveness and accuracy of automated marketing conversations.

[0044] As a preferred solution, generating a marketing reply text according to the optimal marketing strategy and real-time interaction text data, and conducting an automated marketing conversation with the user based on the marketing reply text, specifically includes:

[0045] Input the optimal marketing strategy and real-time interaction text data into a preset large language model to generate a marketing reply text;

[0046] Send the marketing reply text to the user to complete the automated marketing conversation with the user.

[0047] In this preferred solution, by inputting the optimal marketing strategy and real-time interaction text data into a preset large language model, the required marketing reply text can be quickly generated, thereby improving the effectiveness and efficiency of automated marketing conversations.

[0048] As a preferred solution, the acquisition of the user's historical marketing data and labeling of the historical marketing data to determine the user's characteristic data specifically includes:

[0049] Acquire historical marketing data of the user, remove duplicate records in the historical marketing data, and fill in missing values ​​in the historical marketing data to obtain first historical marketing data;

[0050] Filtering the first historical marketing data based on a preset data field to obtain second historical marketing data;

[0051] The second historical marketing data is iteratively clustered according to a preset clustering algorithm to obtain an optimal clustering result, and labels are assigned to users according to the optimal clustering result to determine user feature data.

[0052] This preferred solution determines user characteristic data by labeling the historical marketing data, which can avoid the impact of erroneous data in the historical marketing data on subsequent marketing strategies, thereby improving the accuracy and marketing effectiveness of the marketing strategy, and further improving the effectiveness and accuracy of automated marketing conversations.

[0053] Accordingly, an embodiment of the present invention provides an automated marketing dialogue device based on reinforcement learning, comprising: a user feature data acquisition module, an initial marketing strategy generation module, an emotional behavior trajectory acquisition module, an optimal marketing strategy acquisition module, and an automated marketing dialogue module;

[0054] The user characteristic data acquisition module is used to acquire the user's historical marketing data, and label the historical marketing data to determine the user characteristic data;

[0055] The initial marketing strategy generation module is used to obtain the user's real-time interactive text data and generate an initial marketing strategy based on the real-time interactive text data and user feature data;

[0056] The emotion behavior trajectory acquisition module is used to acquire the user's conversation emotion data according to the real-time interactive text data, and generate the user's emotion behavior trajectory based on the conversation emotion data;

[0057] The optimal marketing strategy acquisition module is used to optimize the initial marketing strategy based on the real-time interactive text data and emotional behavior trajectory in combination with a preset reinforcement learning algorithm to obtain the optimal marketing strategy; the automated marketing dialogue module is used to generate a marketing response text based on the optimal marketing strategy and real-time interactive text data, and to conduct an automated marketing dialogue with the user based on the marketing response text.

[0058] It can be understood that, compared with the prior art, the present device generates an initial marketing strategy through user characteristic data and real-time interaction text data, and then obtains the dialogue emotion data of the user through the real-time interaction text data to generate the emotional behavior trajectory of the user. Then, based on the real-time interaction text data and the emotional behavior trajectory, the initial marketing strategy is optimized by combining a reinforcement learning algorithm to obtain the optimal marketing strategy, realizing an automated marketing dialogue with the user. The present device obtains dialogue emotion data from the real-time interaction text data of the user, thereby generating the emotional behavior trajectory of the user, and generates the optimal marketing strategy based on the real-time interaction text data and the emotional behavior trajectory, enabling the optimal marketing strategy to fully integrate the real-time preferences and demand changes of the user, realizing the dynamic generation and immediate adjustment of the marketing strategy, enabling the marketing strategy to quickly respond to the changes in user needs, avoiding the deviation from user needs caused by generating marketing strategies solely relying on historical marketing data in the prior art, thereby improving the accuracy of the marketing strategy and the marketing effect, and further improving the effectiveness and accuracy of the automated marketing dialogue. BRIEF DESCRIPTION OF THE DRAWINGS

[0059] Figure 1 : A flowchart of the steps of a method for automated marketing dialogue based on reinforcement learning provided by an embodiment of the present invention;

[0060] Figure 2 : A schematic structural diagram of an automated marketing dialogue device based on reinforcement learning provided by an embodiment of the present invention;

[0061] Among them, 201: User characteristic data acquisition module; 202: Initial marketing strategy generation module; 203: Emotional behavior trajectory acquisition module; 204: Optimal marketing strategy acquisition module; 205: Automated marketing dialogue module. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0062] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0063] Embodiment 1

[0064] Please refer to Figure 1 , a flowchart of the steps of a method for automated marketing dialogue based on reinforcement learning provided by an embodiment of the present invention, including steps S101 to S105.

[0065] Step S101: Obtain the historical marketing data of the user, and perform tagging processing on the historical marketing data to determine the user characteristic data.

[0066] In this embodiment, the acquisition of the user's historical marketing data and labeling of the historical marketing data to determine the user's characteristic data specifically includes:

[0067] Acquire historical marketing data of the user, remove duplicate records in the historical marketing data, and fill in missing values ​​in the historical marketing data to obtain first historical marketing data;

[0068] Filtering the first historical marketing data based on a preset data field to obtain second historical marketing data;

[0069] The second historical marketing data is iteratively clustered according to a preset clustering algorithm to obtain an optimal clustering result, and labels are assigned to users according to the optimal clustering result to determine user feature data.

[0070] In an optional embodiment, this embodiment designs automation tasks through Process Easy RPA, including: logging into the enterprise micro-backend, reading interface data, etc., and then transmitting the data through the API interface. It also creates a Python script through the Python language, uses its requests library to call the CRM API interface, and obtains the user's telecommunications background system data. Finally, all the data is integrated to obtain the user's historical marketing data.

[0071] It should be noted that FlowEasy RPA is a solution based on Robotic Process Automation (RPA) technology. API interface refers to Application Programming Interface, which is used for mutual communication between different software applications to achieve data exchange and function calls. CRM refers to Customer Relationship Management, which is a strategy and technology for managing the relationship between an enterprise and existing and potential customers, involving sales, marketing, customer service and support. The requests library is a commonly used library in Python language.

[0072] In particular, the above-mentioned method of obtaining users' historical marketing data through ProcessEasy RPA and creating Python scripts is a relatively mature technology at present, which will not be elaborated in detail here.

[0073] In an optional embodiment, the drop_duplicates() function of Pandas is used to remove duplicate records in the historical marketing data, and then the fillna() method in the Python language is used to fill in the missing values to obtain the first historical marketing data. Then, the format of different data types in the first historical marketing data is converted to ensure that all data types are accurate. Then, the first historical marketing data is filtered based on preset data fields, where the preset data fields refer to data fields related to marketing (such as marketing time, product type, etc.) to obtain the second historical marketing data.

[0074] It should be noted that the Pandas library is an open-source data analysis and operation library in the Python language, which provides high-performance and easy-to-use data structures and data analysis tools; the fillna() method is a commonly used function in the Pandas library for filling in missing values. In addition, in this embodiment, missing values can also be filled in by means of the mode, median, etc.

[0075] In this embodiment, iteratively clustering the second historical marketing data according to a preset clustering algorithm, obtaining an optimal clustering result, and assigning labels to users according to the optimal clustering result to determine user feature data specifically includes:

[0076] Filter the initial feature data of the second historical marketing data, and perform numerical standardization processing on the initial feature data to determine the first feature data;

[0077] Cluster the first feature data according to a preset clustering algorithm to obtain an initial clustering result; and evaluate the initial clustering result according to a preset clustering evaluation algorithm to obtain an evaluation result;

[0078] When the evaluation result does not meet the preset clustering requirements, iteratively cluster the first feature data according to the preset clustering algorithm and update the evaluation result until the evaluation result meets the preset clustering requirements, completing the clustering of the first feature data to obtain an optimal clustering result;

[0079] Assign labels to users according to the optimal clustering result, and determine user feature data based on user labels.

[0080] In an optional embodiment, filter the initial feature data of the second historical marketing data. The initial feature data includes: user transaction history, enterprise WeChat interaction times, package type, average consumption amount, etc.; then use sklearn.StandardScaler to perform numerical standardization processing on the initial feature data to determine the first feature data; then use the K-means clustering algorithm (i.e., the preset clustering algorithm) for clustering; where the core equation of the K-means clustering algorithm is:

[0081]

[0082] In the core equation, u ij represents the probability that the sample x i belongs to the j-th class, μ j is the center of the j-th class, n is the number of samples, and k is the number of classes;

[0083] After that, the silhouette coefficient or the Calinski-Harabasz index (i.e., the preset clustering evaluation algorithm) is used to evaluate the clustering result of the K-means clustering algorithm, and the K-means clustering algorithm is adjusted and reclustered according to the clustering result, so that the initial evaluation result meets the preset clustering requirements, completing the clustering of the first feature data and obtaining the optimal clustering result. Then, labels are assigned to each user according to the optimal clustering result, thereby determining the user feature data.

[0084] It should be noted that sklearn.StandardScaler is a class in Scikit-learn (a Python machine learning library) used to standardize the numerical features in a dataset. The preset clustering algorithm described in this embodiment can also be algorithms such as Hierarchical Clustering and DBSCAN clustering (Density-Based Spatial Clustering of Applications with Noise). The silhouette coefficient described in this embodiment is specifically the Silhouette Coefficient algorithm (SC), and the Calinski-Harabasz index is specifically the variance ratio criterion, which measures the compactness and separation of the clustering result by comparing the variance ratio between the sample dispersion within the cluster and the sample dispersion between the clusters. The preset clustering requirement in this embodiment is set such that when using the silhouette coefficient to evaluate the clustering result, the silhouette coefficient threshold is set to 0.9. When the silhouette coefficient corresponding to the initial evaluation result exceeds 0.9, it is considered that the initial evaluation result meets the preset clustering requirements.

[0085] In this embodiment, by performing labeling processing on the historical marketing data to determine the user feature data, it is possible to avoid the impact of incorrect data in the historical marketing data on subsequent marketing strategies, thereby improving the accuracy of the marketing strategy and the marketing effect, and further improving the effectiveness and accuracy of the automated marketing dialogue.

[0086] Step S102: Obtain the real-time interaction text data of the user, and generate an initial marketing strategy according to the real-time interaction text data and the user feature data.

[0087] In this embodiment, obtaining the real-time interactive text data of the user and generating an initial marketing strategy according to the real-time interactive text data and the user characteristic data specifically includes:

[0088] Obtaining the historical interactive text data and historical marketing feedback data of the user;

[0089] Training a preset marketing strategy recommendation model according to the historical interactive text data, historical marketing feedback data and user characteristic data to determine an optimal marketing strategy recommendation model;

[0090] Inputting the real-time interactive text data and user characteristic data into the optimal marketing strategy recommendation model to generate an initial marketing strategy.

[0091] In this embodiment, a preset marketing strategy recommendation model is trained by using historical interactive text data, historical marketing feedback data and user characteristic data, so that the optimal marketing strategy recommendation model can fully consider the real-time preferences and demand changes of the user, so that the initial marketing strategy can fully integrate the real-time preferences and demand changes of the user, thereby improving the accuracy of the marketing strategy and the marketing effect, and further improving the effectiveness and accuracy of the automated marketing dialogue.

[0092] In this embodiment, training a preset marketing strategy recommendation model according to the historical interactive text data, historical marketing feedback data and user characteristic data to determine an optimal marketing strategy recommendation model specifically includes:

[0093] Obtaining the strategy data in the preset strategy database and performing label encoding processing on the strategy data to determine the first strategy label data;

[0094] Inputting the historical marketing feedback data, historical interactive text data, user characteristic data and the first strategy label data into the preset marketing strategy recommendation model, training the preset marketing strategy recommendation model by combining a preset activation function and a preset loss function, and evaluating the training result of the preset marketing strategy recommendation model by combining a preset model evaluation index;

[0095] Determining that the evaluation result of the training result of the preset marketing strategy recommendation model meets the preset requirements, completing the training of the preset marketing strategy recommendation model, and obtaining an optimal marketing strategy recommendation model.

[0096] In an alternative embodiment, obtaining the strategy data in the preset strategy database and performing one-hot encoding or label encoding on the strategy data to obtain the first strategy label data; then constructing a preset marketing strategy recommendation model through a multi-layer perceptron (that is, the preset marketing strategy recommendation model uses stacked fully connected layers, and its last layer is the output node of the marketing strategy);

[0097] After that, the Sigmoid function is selected as the preset activation function, and the cross-entropy loss function is selected as the preset loss function;

[0098] The expression of the Sigmoid function is:

[0099] The expression of the cross-entropy loss function is

[0100] In the cross-entropy loss function, y o,c is a binary indicator (0 or 1), which is 1 if class c is the correct classification of sample o, and 0 otherwise. p o,c is the probability that the model predicts that sample o belongs to class c. C is the total number of classes;

[0101] After that, the historical marketing feedback data, historical interaction text data, user feature data, and first strategy label data are divided into a training set and a validation set, and are batch-loaded using DataLoader. Then, a for loop is used to iterate through the data batches, performing forward propagation, calculating the loss, backpropagation, and parameter updates, and adjusting parameters such as the learning rate and hidden layer size through grid search or random search. Finally, metrics such as accuracy, recall, and F1 score are used to evaluate the training results of the preset marketing strategy recommendation model to obtain the optimal marketing strategy recommendation model.

[0102] It should be noted that the Multilayer Perceptron (MLP) is a feedforward artificial neural network model that is stacked by multiple neuron layers and can handle complex non-linear problems. The Sigmoid function, also known as the S-shaped growth curve or logistic function, is often used as the activation function of neural networks, which maps the input of neurons to the range (0, 1), thus achieving non-linear transformation. The cross-entropy loss function is a commonly used loss function in machine learning and deep learning, especially suitable for classification problems, and is used to evaluate the difference between the probability distribution output by the model and the true probability distribution. DataLoader is an iterator object that receives a dataset as input and loads data according to parameters such as the specified batch size and whether to shuffle the data. By iterating through DataLoader, a batch of data can be extracted each time for training or testing.

[0103] In this embodiment, the preset marketing strategy recommendation model is trained with historical interaction text data, user feature data, and first policy label data, enabling the optimal marketing strategy recommendation model to fully consider the real-time preferences and demand changes of users. As a result, the accuracy of the initial marketing strategy output by the optimal marketing strategy recommendation model and the marketing effect are improved, thereby enhancing the effectiveness and accuracy of automated marketing conversations.

[0104] Step S103: Obtain the dialogue sentiment data of the user according to the real-time interaction text data, and generate an emotional behavior trajectory of the user based on the dialogue sentiment data.

[0105] In this embodiment, the obtaining of the dialogue sentiment data of the user according to the real-time interaction text data and generating an emotional behavior trajectory of the user based on the dialogue sentiment data specifically includes:

[0106] The real-time interaction text data includes: user output text data and replied user text data;

[0107] Analyze the user output text data according to a preset prompt word extraction algorithm to determine the dialogue sentiment data of the user;

[0108] Convert the replied user text data to determine the user behavior response data;

[0109] Pair the dialogue sentiment data, user behavior response data, and user feature data to generate an emotional behavior trajectory of the user.

[0110] In this embodiment, by analyzing the user output text data and replied user text data, an emotional behavior trajectory of the user is generated, enabling the full characterization of the real-time preferences and demand changes of the user through the emotional behavior trajectory. As a result, the accuracy of the subsequent generated optimal marketing strategy and the marketing effect are improved, realizing the dynamic generation and instant adjustment of the marketing strategy, and further enhancing the effectiveness and accuracy of automated marketing conversations.

[0111] In this embodiment, the analyzing of the user output text data according to a preset prompt word extraction algorithm to determine the dialogue sentiment data of the user specifically includes:

[0112] Input the user output text data into a preset prompt word model to analyze the user output text data and obtain the user emotion category and user emotion intensity;

[0113] Assign a continuous value to the user emotion intensity to obtain a user emotion intensity value;

[0114] Obtain a prompt word template, and modify the prompt word template according to the real-time interaction text data, user emotion category, and user emotion intensity value;

[0115] Determine the user's dialogue sentiment data based on the modified prompt template.

[0116] It should be noted that the preset prompt model refers to using Prompt Engineering to call large language models, which is a key concept in the field of artificial intelligence, especially playing an important role in natural language processing (NLP) tasks. Its working principle is based on the parsing and generation capabilities of pre-trained models (such as BERT, GPT, etc., that is, large language models) for prompts. When a user inputs a prompt, the large language model will parse this prompt and generate corresponding outputs according to its content and structure.

[0117] In an optional embodiment, before inputting the user output text data into the preset prompt model, it is also necessary to train and adjust the prompt template for the preset prompt engineering; first, prepare a multi-round dialogue sample set with labels (used to verify the effectiveness and accuracy of the prompt template). Each sample should contain a complete multi-round dialogue record, its corresponding true sentiment label and intensity score. The intensity score uses a continuous value, a floating point number between 0 and 1, which is manually annotated. At the same time, it is necessary to annotate the sentiment state of each round of dialogue; then design the prompt template, where the prompt template includes: background, the user's current reply, historical dialogue, etc. Use JSON format as the output format of the prompt template.

[0118] Then use the prompt template to test the multi-round dialogue sample set, compare the output results of the large language model with the manually annotated results, and thus adjust the prompt content to make the output results of the large model close to the manually annotated results, so as to determine the prompt model and the prompt template.

[0119] Then input the user output text data into the large language model, and through prompt engineering (that is, inputting the preset prompt model), analyze the user output text data to obtain the user sentiment category and user sentiment intensity. Among them, the user sentiment category includes: positive, negative, neutral, doubtful, etc.; then assign a continuous value to the user sentiment intensity to obtain the user sentiment intensity value; among them, the value range of the user sentiment intensity value is from 0 to 1, 0 is used to represent the user's lack of emotion, and 1 is used to represent the user's extremely strong emotion.

[0120] Then modify the variables of the prompt template according to the real-time interaction text data, user sentiment category and user sentiment intensity value, and determine the user's dialogue sentiment data based on the modified prompt template.

[0121] In an optional embodiment, when a user expresses interest in a certain product but is unsure whether it meets their needs, it is marked as having a user emotion category of doubt and a user emotion intensity value of 0.7.

[0122] Through the prompt word model in this embodiment, more accurate user emotion categories and user emotion intensities can be obtained. The form of continuous value assignment can avoid information loss caused by traditional discrete classification, improve the accuracy of user emotion evaluation, and enhance the accuracy of the user's dialogue emotion data. The acquisition process of the user's dialogue emotion data can be simplified through the prompt word template, improving the acquisition efficiency of the user's dialogue emotion data.

[0123] In this embodiment, the conversion of the replied user text data to determine the user behavior response data specifically includes:

[0124] Performing data mapping and encoding on the replied user text data based on the user feature data and the real-time interaction text data to determine the behavior selection data corresponding to the replied user text data;

[0125] Querying a preset user behavior response database according to the behavior selection data corresponding to the replied user text data to determine the user behavior response data.

[0126] In an optional embodiment, the preset user behavior response database includes user behavior response operations such as recommending featured products or services, providing detailed product explanations, performing soothing words, and providing more information. Based on the user feature data and the real-time interaction text data, data mapping and encoding are performed on the replied user text data to determine the behavior selection data corresponding to the replied user text data, and then the preset user behavior response database is queried to select the corresponding user behavior response operation to determine the user behavior response data.

[0127] In an optional embodiment, when the user emotion category is doubt and the user emotion intensity value is 0.7, the user behavior response operation of providing more information is selected, encoded as vector 5, to determine the user behavior response data.

[0128] In this embodiment, data mapping and encoding are performed on the replied user text data through the user feature data and the real-time interaction text data, making the obtained user behavior response data more in line with the user's real-time preferences and demand changes, thereby improving the accuracy of the subsequent optimal marketing strategy and the marketing effect, and further enhancing the effectiveness and accuracy of the automated marketing dialogue.

[0129] In an optional embodiment, the pairing of the dialogue emotion data, the user behavior response data, and the user feature data to generate the user's emotion behavior trajectory specifically includes:

[0130] Encode the user feature data to obtain the user individual feature data;

[0131] Concatenate the dialogue sentiment data, user behavior response data, and user individual feature data in chronological order to form a complete emotion - behavior trajectory (EAT), that is, generate the emotion - behavior trajectory of the user. Each node in the emotion - behavior trajectory includes: dialogue sentiment data, user behavior response data, and user feature data.

[0132] In an optional embodiment, when the user feature data shows that the user has frequently consulted international roaming services in the past three months, it is encoded as vector 5, indicating that the user is interested in international roaming services.

[0133] Please refer to Table 1, which is a schematic table of the emotion - behavior trajectory provided by the embodiment of the present invention; as shown in Table 1, at different time points, such as time point t 0 when the dialogue sentiment data of the user is confusion / 0.7, the user behavior response data is to provide more information (vector 2), and the user individual feature data is high interest in international roaming (vector 5); at time point t 1 when the dialogue sentiment data of the user is positive / 0.9, the user behavior response data is to recommend products (vector 1), and the user individual feature data is preference for long - term contracts (vector 4).

[0134] Table 1 Schematic table of emotion - behavior trajectory

[0135]

[0136]

[0137] Step S104: Optimize the initial marketing strategy based on the real - time interaction text data and the emotion - behavior trajectory, combined with a preset reinforcement learning algorithm, to obtain the optimal marketing strategy.

[0138] In this embodiment, the optimizing the initial marketing strategy based on the real - time interaction text data and the emotion - behavior trajectory, combined with a preset reinforcement learning algorithm, to obtain the optimal marketing strategy specifically includes:

[0139] Convert the real - time interaction text data into a dialogue history background information vector based on a preset text conversion algorithm;

[0140] Generate an emotion state vector, a user behavior response vector, and a user individual feature vector based on the emotion - behavior trajectory;

[0141] Construct a marketing strategy state space according to the dialogue history background information vector, emotion state vector, user behavior response vector, and user individual feature vector;

[0142] Optimize the initial marketing strategy according to the preset reinforcement learning algorithm and the state space of the marketing strategy to obtain the optimal marketing strategy.

[0143] In an alternative embodiment, convert the real-time interactive text data into a fixed-length dialogue history background information vector H based on TextEmbedding. For example, H = [embedding(h)], where embedding(h) is the embedding representation of the dialogue history in the real-time interactive text data.

[0144] It should be noted that TextEmbedding, that is, text embedding, is a crucial technology in natural language processing (NLP), which refers to the technology of representing words, sentences, or entire texts using multi-dimensional vectors.

[0145] In an alternative embodiment, generate an emotional state vector, a user behavior response vector, and a user individual feature vector based on the emotional behavior trajectory, specifically including:

[0146] Define the emotional state vector as E t , representing the user's emotional category and user emotional intensity value at time point t. Among them, the user's emotional category includes: positive, negative, neutral, doubtful, etc., and the value range of the user's emotional intensity value is from 0 to 1. Therefore, the emotional state vector can be expressed as E t = [one-hot(e), e intensity , where one-hot(e) is the one-hot encoding of the user's emotional category, and e intensity ∈ [0, 1] is the user's emotional intensity value;

[0147] Define the user behavior response vector as A t , representing the user's behavior response at time point t, which can be encoded using integer encoding or one-hot encoding; when using one-hot encoding, A t = [one-hot(e)], and when using integer encoding, A t = a;

[0148] Define the user individual feature vector as U, representing the user's features, expressed as U = [u 1 , u 2 , … u n , where u i represents the numerical value of the i-th feature;

[0149] After that, construct the state space of the marketing strategy according to the dialogue history background information vector, emotional state vector, user behavior response vector, and user individual feature vector. The state space of the marketing strategy is expressed as S t = (E t , A t,U,H).

[0150] It should be noted that One-Hot Encoding, also known as one-bit effective encoding, is a common method used in the field of machine learning to process discrete variables (categorical data).

[0151] In an optional embodiment, a Double Deep Q-Network (DDQN) algorithm is selected as a preset reinforcement learning algorithm to optimize the initial marketing strategy to obtain an optimal marketing strategy;

[0152] Specifically, a reward mechanism is set up, with whether a transaction is completed and the user's final emotion as the final reward design, and the action is determined based on the state space; then the Q table is initialized, the ε-greedy strategy is adopted, the selected action is executed, the new state and immediate reward are observed, and then the Q table is updated. The update formula of the Q table is: Q(s,a)←Q(s,a)+α[r+γQ′(s′,argmax a Q(s′,a))-Q(s,a)];

[0153] Among them, Q and Q′ are two independent Q networks, α is the learning rate, γ is the discount factor, s is the current state, a is the action taken, r is the reward, s' is the new state, and a' is the possible action under the new state.

[0154] It should be noted that Double Deep Q-Network (DDQN) is a deep reinforcement learning framework. It is an improved version of Q-learning and aims to solve the problem of decision makers making decisions in uncertain environments. It combines deep neural networks and double Q-learning technology to improve the stability and performance of the Q-learning algorithm. In reinforcement learning, Q-learning is an algorithm for learning how to choose the best action in a given state. DDQN successfully solves the problem of overestimation that may be caused by the target Q value estimation process in traditional Q-learning or DQN (Deep Q-Network) by introducing two independent neural networks (one online network, also called the main network or policy network; the other target network).

[0155] It should be noted that the above optional embodiment uses the Double Deep Q-Network (DDQN) algorithm as the preset reinforcement learning algorithm to optimize the initial marketing strategy to obtain the process description of the optimal marketing strategy. This is only an adaptive description. Due to the current relatively mature technical applications, no further elaboration is given here.

[0156] In this embodiment, a marketing strategy state space is constructed through the dialogue historical background information vector, the emotional state vector, the user behavior response vector, and the user individual characteristic vector, which can represent the real-time preferences and demand changes of users in the form of a state space, so as to improve the accuracy of the optimal marketing strategy obtained based on the state space and the reinforcement learning algorithm and the marketing effect, and further improve the effectiveness and accuracy of the automated marketing dialogue.

[0157] Step S105: Generate a marketing response text according to the optimal marketing strategy and the real-time interaction text data, and conduct an automated marketing dialogue with the user according to the marketing response text.

[0158] In this embodiment, the generating a marketing response text according to the optimal marketing strategy and the real-time interaction text data, and conducting an automated marketing dialogue with the user according to the marketing response text specifically includes:

[0159] Input the optimal marketing strategy and the real-time interaction text data into a preset large language model to generate a marketing response text;

[0160] Send the marketing response text to the user to complete the automated marketing dialogue with the user.

[0161] It should be noted that for the large language model described in this embodiment, models such as Tongyi Qianwen and ChatGLM can be selected.

[0162] In this embodiment, by inputting the optimal marketing strategy and the real-time interaction text data into a preset large language model, the required marketing response text can be quickly generated, thus improving the effectiveness and efficiency of the automated marketing dialogue.

[0163] In an optional embodiment, after conducting an automated marketing dialogue with the user according to the marketing response text, the Double Deep Q-Network (DDQN) algorithm and the marketing state space can also be optimized through the user's real-time response text, thereby improving the immediacy and accuracy of the marketing strategy.

[0164] This embodiment generates an initial marketing strategy through user feature data and real-time interactive text data, then obtains the user's conversation emotion data through real-time interactive text data, generates the user's emotional behavior trajectory, and then optimizes the initial marketing strategy based on the real-time interactive text data and the emotional behavior trajectory in combination with the reinforcement learning algorithm to obtain the optimal marketing strategy, thereby realizing automated marketing dialogue for the user. This embodiment obtains conversation emotion data from the user's real-time interactive text data, thereby generating the user's emotional behavior trajectory, and generates the optimal marketing strategy with real-time interactive text data and emotional behavior trajectory, so that the optimal marketing strategy can not only fully integrate the user's real-time preferences and demand changes, but also consider the user's historical marketing situation, realizing the dynamic generation and immediate adjustment of the marketing strategy, so that the marketing strategy can quickly respond to changes in user demand, avoiding the deviation from user demand caused by relying solely on historical marketing data to generate marketing strategies in the prior art, thereby improving the accuracy and marketing effect of the marketing strategy, and further improving the effectiveness and accuracy of the automated marketing dialogue.

[0165] Embodiment 2

[0166] Please refer to Figure 2 , which is a structural diagram of an automated marketing dialogue device based on reinforcement learning provided by an embodiment of the present invention, comprising: a user feature data acquisition module 201, an initial marketing strategy generation module 202, an emotional behavior trajectory acquisition module 203, an optimal marketing strategy acquisition module 204 and an automated marketing dialogue module 205;

[0167] The user characteristic data acquisition module 201 is used to acquire the user's historical marketing data, and perform labeling on the historical marketing data to determine the user characteristic data.

[0168] In this embodiment, the user characteristic data acquisition module 201 includes: a user characteristic data acquisition unit;

[0169] The user feature data acquisition unit is used to acquire the user's historical marketing data, remove duplicate records in the historical marketing data, and fill in the missing values ​​in the historical marketing data to obtain the first historical marketing data;

[0170] Filtering the first historical marketing data based on a preset data field to obtain second historical marketing data;

[0171] The second historical marketing data is iteratively clustered according to a preset clustering algorithm to obtain an optimal clustering result, and labels are assigned to users according to the optimal clustering result to determine user feature data.

[0172] The initial marketing strategy generation module 202 is used to obtain the real-time interactive text data of the user, and generate an initial marketing strategy according to the real-time interactive text data and the user characteristic data.

[0173] In this embodiment, the initial marketing strategy generation module 202 includes: an initial marketing strategy generation unit;

[0174] The initial marketing strategy generation unit is used to obtain the historical interactive text data and historical marketing feedback data of the user;

[0175] Train a preset marketing strategy recommendation model according to the historical interactive text data, historical marketing feedback data and user characteristic data, and determine an optimal marketing strategy recommendation model;

[0176] Input the real-time interactive text data and user characteristic data into the optimal marketing strategy recommendation model to generate an initial marketing strategy.

[0177] In this embodiment, the initial marketing strategy generation unit includes: a marketing strategy recommendation model training subunit;

[0178] The marketing strategy recommendation model training subunit is used to obtain the strategy data in the preset strategy database, and perform label encoding processing on the strategy data to determine the first strategy label data;

[0179] Input the historical marketing feedback data, historical interactive text data, user characteristic data and the first strategy label data into the preset marketing strategy recommendation model, train the preset marketing strategy recommendation model in combination with a preset activation function and a preset loss function, and evaluate the training result of the preset marketing strategy recommendation model in combination with a preset model evaluation index;

[0180] Determine that the evaluation result of the training result of the preset marketing strategy recommendation model meets the preset requirements, complete the training of the preset marketing strategy recommendation model, and obtain an optimal marketing strategy recommendation model.

[0181] The emotion behavior trajectory acquisition module 203 is used to obtain the dialogue emotion data of the user according to the real-time interactive text data, and generate an emotion behavior trajectory of the user based on the dialogue emotion data.

[0182] In this embodiment, the emotion behavior trajectory acquisition module 203 includes: an emotion behavior trajectory acquisition unit;

[0183] In the emotion behavior trajectory acquisition unit, the real-time interactive text data includes: user output text data and reply user text data;

[0184] The emotional behavior trajectory acquisition unit is used to analyze the user output text data according to a preset prompt word extraction algorithm to determine the user's dialogue emotional data;

[0185] Convert the replied user text data to determine the user behavior response data;

[0186] Pair the dialogue emotional data, user behavior response data, and user feature data to generate the user's emotional behavior trajectory.

[0187] In this embodiment, the emotional behavior trajectory acquisition unit includes: a dialogue emotional data acquisition subunit;

[0188] The dialogue emotional data acquisition subunit is used to input the user output text data into a preset prompt word model, analyze the user output text data, and obtain the user emotional category and user emotional intensity;

[0189] Assign a continuous value to the user emotional intensity to obtain the user emotional intensity value;

[0190] Obtain a prompt word template, and modify the prompt word template according to the real-time interaction text data, user emotional category, and user emotional intensity value;

[0191] Determine the user's dialogue emotional data based on the modified prompt word template.

[0192] In this embodiment, the emotional behavior trajectory acquisition unit includes: a user behavior response data acquisition subunit;

[0193] The user behavior response data acquisition subunit is used to perform data mapping and encoding on the replied user text data based on the user feature data and real-time interaction text data to determine the behavior selection data corresponding to the replied user text data;

[0194] Query a preset user behavior response database according to the behavior selection data corresponding to the replied user text data to determine the user behavior response data.

[0195] The optimal marketing strategy acquisition module 204 is used to optimize the initial marketing strategy based on the real-time interaction text data and emotional behavior trajectory, combined with a preset reinforcement learning algorithm, to obtain the optimal marketing strategy.

[0196] In this embodiment, the optimal marketing strategy acquisition module 204 includes: an optimal marketing strategy acquisition unit;

[0197] The optimal marketing strategy acquisition unit is used to convert the real-time interaction text data into a dialogue historical background information vector based on a preset text conversion algorithm;

[0198] Generate an emotional state vector, a user behavior response vector, and a user individual characteristic vector based on the emotional behavior trajectory;

[0199] Construct a marketing strategy state space according to the dialogue historical background information vector, the emotional state vector, the user behavior response vector, and the user individual characteristic vector;

[0200] Optimize the initial marketing strategy according to a preset reinforcement learning algorithm and the marketing strategy state space to obtain an optimal marketing strategy.

[0201] The automated marketing dialogue module 205 is used to generate a marketing reply text according to the optimal marketing strategy and real-time interaction text data, and conduct an automated marketing dialogue with the user according to the marketing reply text.

[0202] In this embodiment, the automated marketing dialogue module 205 includes: an automated marketing dialogue unit;

[0203] The automated marketing dialogue unit is used to input the optimal marketing strategy and real-time interaction text data into a preset large language model to generate a marketing reply text;

[0204] Send the marketing reply text to the user to complete the automated marketing dialogue with the user.

[0205] In this embodiment, an initial marketing strategy is generated through user characteristic data and real-time interaction text data. Then, the user's dialogue emotion data is obtained through the real-time interaction text data, and the emotional behavior trajectory of the user is generated. Then, based on the real-time interaction text data and the emotional behavior trajectory, combined with the reinforcement learning algorithm, the initial marketing strategy is optimized to obtain an optimal marketing strategy, realizing the automated marketing dialogue with the user. In this embodiment, the dialogue emotion data is obtained from the user's real-time interaction text data, thereby generating the emotional behavior trajectory of the user. The optimal marketing strategy is generated based on the real-time interaction text data and the emotional behavior trajectory, so that the optimal marketing strategy can not only fully integrate the user's real-time preferences and demand changes, but also consider the user's historical marketing situation, realizing the dynamic generation and instant adjustment of the marketing strategy, enabling the marketing strategy to quickly respond to the user's demand changes, avoiding the deviation from the user's needs caused by only relying on historical marketing data to generate the marketing strategy in the prior art, thereby improving the accuracy of the marketing strategy and the marketing effect, and further improving the effectiveness and accuracy of the automated marketing dialogue.

[0206] In summary, in the embodiment of the present invention, an initial marketing strategy is generated based on user feature data and real-time interaction text data. Then, the dialogue emotion data of the user is obtained through the real-time interaction text data, and the emotional behavior trajectory of the user is generated. Then, based on the real-time interaction text data and the emotional behavior trajectory, the initial marketing strategy is optimized by combining a reinforcement learning algorithm to obtain an optimal marketing strategy, realizing an automated marketing dialogue with the user. In the embodiment of the present invention, the dialogue emotion data is obtained from the real-time interaction text data of the user, thereby generating the emotional behavior trajectory of the user. The optimal marketing strategy is generated based on the real-time interaction text data and the emotional behavior trajectory, so that the optimal marketing strategy can not only fully integrate the real-time preferences and demand changes of the user, but also consider the historical marketing situation of the user, realizing the dynamic generation and instant adjustment of the marketing strategy, enabling the marketing strategy to quickly respond to the user demand changes, avoiding the deviation from the user demand caused by generating the marketing strategy only relying on historical marketing data in the prior art, thereby improving the accuracy of the marketing strategy and the marketing effect, and further improving the effectiveness and accuracy of the automated marketing dialogue.

[0207] In the specific embodiments described above, the purpose, technical solutions, and beneficial effects of the present invention have been further described in detail. It should be understood that the above are only specific embodiments of the present invention and are not used to limit the protection scope of the present invention. In particular, it is pointed out that for those skilled in the art, any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. An automated marketing dialogue method based on reinforcement learning, characterized in that: include: Obtaining historical marketing data of the user, and labeling the historical marketing data to determine user feature data; Acquire real-time interactive text data of users, and generate an initial marketing strategy based on the real-time interactive text data and user feature data; Acquire the user's conversation emotion data according to the real-time interactive text data, and generate the user's emotion behavior trajectory based on the conversation emotion data; Based on the real-time interactive text data and emotional behavior trajectory, the initial marketing strategy is optimized in combination with a preset reinforcement learning algorithm to obtain an optimal marketing strategy; A marketing response text is generated according to the optimal marketing strategy and the real-time interactive text data, and an automated marketing dialogue is conducted with the user according to the marketing response text.

2. The automated marketing dialogue method based on reinforcement learning according to claim 1, characterized in that: The step of obtaining the user's real-time interactive text data and generating an initial marketing strategy based on the real-time interactive text data and user feature data specifically includes: Obtain users' historical interaction text data and historical marketing feedback data; Training a preset marketing strategy recommendation model based on the historical interactive text data, historical marketing feedback data, and user feature data to determine an optimal marketing strategy recommendation model; The real-time interactive text data and user feature data are input into the optimal marketing strategy recommendation model to generate an initial marketing strategy.

3. The automated marketing dialogue method based on reinforcement learning as claimed in claim 2, characterized in that: The step of training a preset marketing strategy recommendation model according to the historical interactive text data, historical marketing feedback data and user feature data to determine the optimal marketing strategy recommendation model specifically includes: Obtaining policy data from a preset policy database, and performing label encoding processing on the policy data to determine first policy label data; Inputting the historical marketing feedback data, the historical interaction text data, the user feature data and the first strategy label data into a preset marketing strategy recommendation model, training the preset marketing strategy recommendation model in combination with a preset activation function and a preset loss function, and evaluating the training result of the preset marketing strategy recommendation model in combination with a preset model evaluation indicator; Determine whether the evaluation result of the training result of the preset marketing strategy recommendation model meets the preset requirements, complete the training of the preset marketing strategy recommendation model, and obtain the optimal marketing strategy recommendation model.

4. The automated marketing dialogue method based on reinforcement learning according to claim 1, characterized in that: The acquiring the user's conversation emotion data according to the real-time interactive text data, and generating the user's emotion behavior trajectory based on the conversation emotion data, specifically includes: The real-time interactive text data includes: user output text data and reply user text data; Analyze the user output text data according to a preset prompt word extraction algorithm to determine the user's conversation emotion data; Converting the replying user text data to determine user behavior response data; The conversation emotion data, user behavior response data and user feature data are paired to generate the user's emotion behavior trajectory.

5. The automated marketing dialogue method based on reinforcement learning as claimed in claim 4, characterized in that: The analyzing the user output text data according to the preset prompt word extraction algorithm to determine the user's dialogue emotion data specifically includes: Inputting the user output text data into a preset prompt word model, analyzing the user output text data to obtain the user emotion category and the user emotion intensity; Assigning a continuous value to the user emotion intensity to obtain a user emotion intensity value; Acquire a prompt word template, and modify the prompt word template according to the real-time interactive text data, the user emotion category, and the user emotion intensity value; The user's conversation emotion data is determined based on the modified prompt word template.

6. The automated marketing dialogue method based on reinforcement learning as claimed in claim 4, characterized in that: The converting of the reply user text data to determine the user behavior response data specifically includes: Based on the user feature data and the real-time interactive text data, data mapping and encoding are performed on the reply user text data to determine the behavior selection data corresponding to the reply user text data; A preset user behavior response database is queried according to the behavior selection data corresponding to the reply user text data to determine the user behavior response data.

7. The automated marketing dialogue method based on reinforcement learning as claimed in claim 4, characterized in that: The optimization of the initial marketing strategy based on the real-time interactive text data and the emotional behavior trajectory in combination with a preset reinforcement learning algorithm to obtain the optimal marketing strategy specifically includes: Converting the real-time interactive text data into a conversation history background information vector based on a preset text conversion algorithm; Generate an emotional state vector, a user behavior response vector and a user individual feature vector based on the emotional behavior trajectory; Constructing a marketing strategy state space based on the conversation history background information vector, the emotional state vector, the user behavior response vector, and the user individual feature vector; The initial marketing strategy is optimized according to a preset reinforcement learning algorithm and the marketing strategy state space to obtain an optimal marketing strategy.

8. The automated marketing dialogue method based on reinforcement learning as claimed in claim 1, characterized in that: Generating a marketing response text according to the optimal marketing strategy and the real-time interactive text data, and conducting an automated marketing dialogue with the user according to the marketing response text, specifically includes: Inputting the optimal marketing strategy and real-time interactive text data into a preset large language model to generate a marketing response text; The marketing reply text is sent to the user to complete the automated marketing dialogue with the user.

9. The automated marketing dialogue method based on reinforcement learning as claimed in claim 1, characterized in that: The obtaining of the user's historical marketing data and labeling of the historical marketing data to determine the user's characteristic data specifically includes: Acquire historical marketing data of the user, remove duplicate records in the historical marketing data, and fill in missing values ​​in the historical marketing data to obtain first historical marketing data; Filtering the first historical marketing data based on a preset data field to obtain second historical marketing data; The second historical marketing data is iteratively clustered according to a preset clustering algorithm to obtain an optimal clustering result, and labels are assigned to users according to the optimal clustering result to determine user feature data.

10. An automated marketing dialogue device based on reinforcement learning, characterized in that: include: User feature data acquisition module, initial marketing strategy generation module, emotional behavior trajectory acquisition module, optimal marketing strategy acquisition module and automated marketing dialogue module; The user characteristic data acquisition module is used to acquire the user's historical marketing data, and label the historical marketing data to determine the user characteristic data; The initial marketing strategy generation module is used to obtain the user's real-time interactive text data and generate an initial marketing strategy based on the real-time interactive text data and user feature data; The emotion behavior trajectory acquisition module is used to acquire the user's conversation emotion data according to the real-time interactive text data, and generate the user's emotion behavior trajectory based on the conversation emotion data; The optimal marketing strategy acquisition module is used to optimize the initial marketing strategy based on the real-time interactive text data and emotional behavior trajectory in combination with a preset reinforcement learning algorithm to obtain the optimal marketing strategy; The automated marketing dialogue module is used to generate a marketing response text according to the optimal marketing strategy and the real-time interactive text data, and to conduct an automated marketing dialogue with the user according to the marketing response text.