Scientific and technological product popularization method based on AI
By deploying data acquisition systems and building deep learning behavior models in the insurance company's internal network, combining reinforcement learning and adaptive adjustment mechanisms, the problem of difficulty in personalizing and real-time adjustment of promotion strategies in the existing technology is solved, and efficient and flexible promotion strategy optimization is achieved.
Patent Information
- Application Number
- CN202411960903.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-30
- Publication Date
- 2025-05-23
AI Technical Summary
The existing technology product promotion methods are difficult to meet personalized needs, and lack real-time dynamic adjustment mechanisms, making it difficult to quickly optimize promotion strategies based on employee actual feedback.
Using an automated promotion method based on AI, we deploy a data acquisition system in the insurance company's internal network, collect employee usage behavior data, and build a deep learning behavior model to generate personalized promotion content. Use reinforcement learning algorithm to dynamically adjust the push strategy based on real-time feedback, and optimize the promotion strategy through an adaptive adjustment mechanism.
It realizes the personalization and flexibility of promotion strategies, can quickly adapt to changes in employee needs, and improves the effectiveness of promotion and employee acceptance.
Smart Images

Figure CN120030224A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of product promotion, and in particular to a method for promoting technological products based on AI. Background Art
[0002] With the continuous development of artificial intelligence technology, all walks of life have begun to widely apply AI technology to improve their work efficiency and innovation capabilities. In the field of technology product promotion, traditional promotion methods usually rely on manual intervention and targeted push, which is difficult to meet personalized needs and difficult to adjust promotion strategies in real time to cope with changes in employee or user behavior. Therefore, AI-based automated promotion methods are increasingly becoming the key for major companies to optimize product promotion and increase employee engagement.
[0003] As an information-intensive industry, the insurance industry's product promotion not only involves complex internal processes, but also needs to consider employees' usage habits and needs. Therefore, how to accurately push and evaluate promotional content through intelligent technology to improve employees' acceptance and usage frequency has become a key issue in improving overall work efficiency and business conversion rate. Existing promotion methods mainly rely on static content push and regular evaluation, lacking a real-time dynamic adjustment mechanism, making it difficult to quickly optimize promotion strategies based on employees' actual feedback.
[0004] Therefore, there is an urgent need for an AI-based technology product promotion method to solve the above problems. Summary of the invention
[0005] In order to overcome the above technical problems existing in the prior art, an embodiment of the present invention provides an AI-based technology product promotion method, the method comprising:
[0006] S1: Deploy a data collection system in the insurance company's internal network to collect internal employees' usage behavior data, including application usage frequency, dwell time, and function usage records. The data collection system is connected to GaussDB and OceanBase databases. After data collection is completed, the collected data is processed through preprocessing methods.
[0007] S2: Based on the preprocessed data, a behavior model of internal employees is constructed. The behavior model uses a deep learning algorithm, including a convolutional neural network (CNN) and a long short-term memory (LSTM) network. The input of the behavior model is the usage behavior data of internal employees, and the output is the potential acceptance of internal employees for new applications.
[0008] S3: Generate personalized promotional content based on the output of the behavior model. The promotional content is generated using natural language processing (NLP) technology, including a text generation model. The generated promotional content includes but is not limited to user guides, operation videos, and FAQs.
[0009] S4: Push personalized promotional content to internal employees, using reinforcement learning algorithms, including Q-learning and DQN, to dynamically adjust the push strategy based on real-time feedback from internal employees;
[0010] S5: Collect internal employees’ feedback on the use of promotional content, including click-through rate, usage frequency, and satisfaction evaluation. The feedback data is transmitted to GaussDB and OceanBase databases in real time through the data collection system. The collected feedback data is used to evaluate the promotion effect.
[0011] S6: Based on the feedback evaluation results, optimize the promotion strategy through an adaptive adjustment mechanism.
[0012] Preferably, in S1, a multi-level data preprocessing method is used, including data cleaning, data standardization, and data dimension reduction. The data cleaning method uses Z-score standardization, and the data dimension reduction uses principal component analysis PCA method. The mathematical expression of Z-score standardization is as follows:
[0013]
[0014] Among them, z is the standardized data, x is the original data, μ is the mean, and σ is the standard deviation.
[0015] Preferably, in S2, CNN is used to extract features from employee behavior data, and LSTM is used to capture time series features in employee behavior data. The specific expression of the model is:
[0016] f CNN (X) = ReLU(W CNN ·X+b CNN );
[0017] f LSTM (x) = LSTM(W LSTM ·X+b LSTM );
[0018] f output (X) = softmax(W output ·f LSTM (X)+b output );
[0019] Among them, X is the input data, W and b are weight and bias parameters respectively, ReLU and LSTM are activation functions respectively, and softmax is used to output the probability distribution of potential acceptance.
[0020] Preferably, in the CNN model, an attention mechanism is added to enhance the model's ability to capture key features. The specific expression of the attention mechanism is:
[0021]
[0022] Where Q, K, V are query, key and value vectors respectively, QK is the dimension of the key vector;
[0023] In the LSTM model, the gated recurrent unit GRU is added to improve the model's time series processing capability. The specific expression of GRU is:
[0024] z t =σ(W z ·[h t-1 , x t ]+b z );
[0025] Among them, z t 、r t are the update gate and reset gate respectively, and σ is the sigmoid activation function.
[0026] Preferably, in S3, the specific expression of the text generation model is:
[0027] f NLP (Y) = Transformer (W NLP ·Y+b NLP );
[0028] Among them, Y is the input internal employee potential acceptance data, W NLP and b NLP are weight and bias parameters respectively, and Transformer is a model for generating text.
[0029] Preferably, in S4, the Q-learning algorithm is used to update the Q value according to the feedback of internal employees, and the DQN algorithm is used to optimize the recommendation strategy. The specific expression of Q-learning is:
[0030]
[0031] Among them, s is the current state, a is the current action, r is the reward value, α is the learning rate, and γ is the discount factor. The DQN algorithm approximates the Q value function through a neural network. The specific expression is:
[0032] Q(s, a; θ) = f DQN (s, a; θ);
[0033] Among them, θ is the network parameter, f DQN is the output function of the DQN model.
[0034] Preferably, S4 also includes a multi-agent reinforcement learning MARL mechanism, which optimizes the overall push strategy through collaborative learning of multiple agents. The specific expression of the multi-agent is:
[0035] Q MARL (s,a 1 , a 2 , ..., a n ) = Q(s, a 1 )+Q(s,a 2 )+...+Q(s,a n );
[0036] Among them, a 1 , a 2 , …, a n are the actions of multiple agents,
[0037] Add an adaptive reward mechanism to dynamically adjust the reward value based on the feedback of internal employees to improve learning efficiency; the specific expression of adaptive rewards is:
[0038] r adapt = r + β·Δ;
[0039] Among them, r is the original reward value, β is the adjustment coefficient, and Δ is the feedback change.
[0040] Preferably, in S5, the feedback and evaluation specifically adopt A / B testing and multi-arm bandit algorithm, and the mathematical expression of feedback evaluation is as follows:
[0041] F = l(C, E; β);
[0042] Among them, F is the feedback evaluation result, C is the promotion content, E is the usage feedback of internal employees, and β is the parameter of the evaluation model.
[0043] Preferably, in S6, the step of the adaptive adjustment mechanism includes:
[0044] S6.1: Based on the feedback evaluation results, evaluate the effectiveness of the current promotion strategy and generate a status evaluation value;
[0045] S6.2: Generate new promotion strategies based on state evaluation values through adaptive control theory;
[0046] S6.3: Verify the generated new strategy in a small range and collect the verification results. During the verification process, apply the new strategy to some internal employees, collect their usage feedback and behavior data, and evaluate the effectiveness of the new strategy;
[0047] S6.4: Based on the verification results, the promotion strategy is optimized through adaptive control theory to generate the final promotion strategy;
[0048] S6.5: Apply the optimized promotion strategy to the content pushed to internal employees to achieve dynamic adjustment and optimization. The optimized strategy is applied in real time through the data collection system, and the pushed content is dynamically adjusted based on the real-time feedback of internal employees.
[0049] Preferably, in S6.2, the adaptive control theory is implemented by the following steps:
[0050] S6.2.1: Define the relationship model between feedback evaluation results and current promotion strategies;
[0051] S6.2.2: Adjust the parameters of the promotion strategy based on the feedback evaluation results and model;
[0052] The specific expression is:
[0053] Among them, θ t is the strategy parameter, α t is the learning rate, is the performance indicator, the gradient of J with respect to the parameter θ.
[0054] Through the technical solution provided by the present invention, the present invention has at least the following technical effects:
[0055] 1. The reinforcement learning and adaptive adjustment mechanism enables the promotion strategy to be dynamically adjusted according to the real-time feedback of employees, thereby improving the flexibility and pertinence of the promotion. This means that even if the needs or interests of employees change, the promotion content can be quickly adapted. Other features and advantages of the embodiments of the present invention will be described in detail in the subsequent specific implementation section.
[0056] 2. Through standardization and dimensionality reduction, not only the speed of model training is improved, but also the expressiveness of the model is enhanced, avoiding the negative impact of noise and redundant information in the data on model training.
[0057] 3. The combination of CNN and LSTM can conduct in-depth analysis of employee behavior from different angles (local features and time series features), thereby generating more accurate personalized promotion strategies.
[0058] 4. GRU can maintain long-term memory and effectively suppress the gradient vanishing problem through an effective gating mechanism, and is suitable for processing long-term trends in employee behavior. BRIEF DESCRIPTION OF THE DRAWINGS
[0059] The accompanying drawings are used to provide a further understanding of the embodiments of the present invention and constitute a part of the specification. Together with the following specific implementations, they are used to explain the embodiments of the present invention, but do not constitute a limitation on the embodiments of the present invention. In the accompanying drawings:
[0060] Figure 1 is a flow chart of the method of the present invention;
[0061] Figure 2 is a flow chart of S6 of the present invention;
[0062] Figure 3 This is a flow chart of S6.2 of the present invention; DETAILED DESCRIPTION
[0063] The specific implementation of the embodiment of the present invention is described in detail below in conjunction with the accompanying drawings. It should be understood that the specific implementation described here is only used to illustrate and explain the embodiment of the present invention, and is not used to limit the embodiment of the present invention.
[0064] The terms "system" and "network" in the embodiments of the present invention can be used interchangeably. "Multiple" means two or more than two. In view of this, "multiple" can also be understood as "at least two" in the embodiments of the present invention. "And / or" describes the association relationship of associated objects, indicating that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. In addition, the character " / ", unless otherwise specified, generally indicates that the previous and subsequent associated objects are in an "or" relationship. In addition, it should be understood that in the description of the embodiments of the present invention, the words "first", "second", etc. are only used to distinguish the purpose of description, and cannot be understood as indicating or implying relative importance, nor can they be understood as indicating or implying order.
[0065] See also Figure 1 , a method for promoting scientific and technological products based on AI, the method comprising:
[0066] S1: Deploy a data collection system in the insurance company's internal network to collect internal employees' usage behavior data, including application usage frequency, dwell time, and function usage records. The data collection system is connected to GaussDB and OceanBase databases. After data collection is completed, the collected data is processed through preprocessing methods.
[0067] S2: Based on the preprocessed data, a behavior model of internal employees is constructed. The behavior model uses a deep learning algorithm, including a convolutional neural network (CNN) and a long short-term memory (LSTM) network. The input of the behavior model is the usage behavior data of internal employees, and the output is the potential acceptance of internal employees for new applications.
[0068] S3: Generate personalized promotional content based on the output of the behavior model. The promotional content is generated using natural language processing (NLP) technology, including a text generation model. The generated promotional content includes but is not limited to user guides, operation videos, and FAQs.
[0069] S4: Push personalized promotional content to internal employees, using reinforcement learning algorithms, including Q-learning and DQN, to dynamically adjust the push strategy based on real-time feedback from internal employees;
[0070] S5: Collect internal employees’ feedback on the use of promotional content, including but not limited to click-through rate, usage frequency, and satisfaction evaluation. The feedback data is transmitted to GaussDB and OceanBase databases in real time through the data collection system. The collected feedback data is used to evaluate the promotion effect.
[0071] S6: Based on the feedback evaluation results, optimize the promotion strategy through an adaptive adjustment mechanism.
[0072] In one possible implementation, the data collection system deployed on the internal network of the insurance company is first used to obtain the employee's usage behavior data (such as application usage frequency, dwell time, etc.). These data can reflect the employee's interest in and behavioral characteristics of the current application. The collected raw data will then be processed through preprocessing methods, which may include noise removal, data standardization, etc., to provide a high-quality data foundation for subsequent modeling. Data storage and management are completed through Gaussian GaussDB and OceanBase databases, which provide efficient storage and query capabilities, ensuring the reliability and real-time nature of the data.
[0073] Furthermore, based on the preprocessed data, deep learning algorithms (CNN and LSTM) are used to build employee behavior models. Convolutional neural networks (CNN) are good at extracting useful features from complex time series data, while long short-term memory networks (LSTM) can process long-term dependent sequence data and are suitable for analyzing the trend of employee behavior over time. By inputting employee usage behavior data, the model outputs the employee's potential acceptance of new applications, providing a personalized reference for promotion strategies.
[0074] Furthermore, based on the output of the employee behavior model, the system can generate promotional content that meets the needs of each employee. Natural language processing (NLP) technology, especially the application of text generation models, makes the generated promotional content more personalized. For example, the generated content can be a user guide, operation video, or FAQ, etc., the purpose is to improve employees' understanding and acceptance of new applications. Personalized content can better fit employees' usage habits and learning needs, and improve promotion effects.
[0075] Furthermore, after the generated personalized promotion content is pushed to employees, the system uses reinforcement learning algorithms (such as Q-learning and DQN) for dynamic adjustment. Based on employees' real-time feedback (such as click-through rate, usage frequency, etc.), the system continuously optimizes the push strategy and selects the most appropriate promotion content and push timing. This adjustment mechanism based on real-time feedback makes the promotion strategy more flexible and can be customized among different employees.
[0076] Furthermore, collecting employee feedback on promotional content (click-through rate, satisfaction, etc.) is an important part of evaluating promotion effectiveness. Feedback data is transmitted to GaussDB and OceanBase databases in real time through the data collection system to ensure the real-time and accuracy of the data. With this data, the system can evaluate the promotion effect and provide data support for subsequent optimization of promotion strategies.
[0077] Furthermore, after collecting enough feedback data, the system optimizes the promotion strategy through an adaptive adjustment mechanism. This mechanism continuously adjusts the generation and push strategy of promotion content based on the collected feedback data, thereby maximizing employee acceptance and satisfaction. This process ensures that the promotion strategy can be continuously optimized and gradually adapt to the needs and changes of employees.
[0078] In the embodiment of the present invention, in S1, a multi-level data preprocessing method is adopted, including data cleaning, data standardization, and data dimension reduction. The data cleaning method adopts Z-score standardization, and the data dimension reduction adopts the principal component analysis PCA method. The mathematical expression of Z-score standardization is as follows:
[0079]
[0080] Among them, z is the standardized data, x is the original data, μ is the mean, and σ is the standard deviation.
[0081] In one possible implementation, data cleaning is the first step of preprocessing, which aims to remove noise and erroneous information in the data and ensure that the data input into the model is accurate and valid. This step mainly includes:
[0082] Missing value handling: Identify and fill in missing values in the data to avoid adverse effects on model training.
[0083] Outlier detection: Detect and process outliers through statistical methods or model algorithms to ensure that the data set truly reflects the actual situation.
[0084] Deduplication: Remove duplicate records to avoid repeated calculations and errors during model training.
[0085] Furthermore, after data cleaning, the Z-score standardization method was used to standardize the data. This method converts all features into a standard normal distribution with a mean of 0 and a standard deviation of 1, ensuring that data of different dimensions and ranges are compared on the same scale.
[0086] Furthermore, the data dimension reduction adopts the principal component analysis (PCA) method, which maps the feature space in the original data to a low-dimensional space through linear transformation, while maximally retaining the variance of the data. The core idea of PCA is to select the principal component (that is, the direction that best expresses the change of the data) and reduce redundant information through dimensionality reduction, so that the data in the low-dimensional space can still fully represent the variability of the original data.
[0087] In the embodiment of the present invention, in S2, CNN is used to extract features in employee behavior data, and LSTM is used to capture time series features in employee behavior data. The specific expression of the model is:
[0088] f CNN (X) = ReLU(W CNN ·X+b CNN );
[0089] f LSTM (X) = LSTM(W LSTM ·X+b LSTM );
[0090] f output (X) = softmax(W output ·f LSTM (X)+b output );
[0091] Among them, X is the input data, W and b are weight and bias parameters respectively, ReLU and LSTM are activation functions respectively, and softmax is used to output the probability distribution of potential acceptance.
[0092] In one possible implementation, convolutional neural networks (CNNs) are usually used to process data with spatial structures, such as image data, but here we apply them to feature extraction of employee behavior data. Behavior data may contain information of multiple dimensions, such as employee operation records, activity frequency, click streams, etc. CNN can extract local features from these data through convolutional layers. The specific expression of CNN is: CNN (X) = ReLU(W CNN ·X+b CNN ); The input data X is the employee's behavior data, which may be a continuous numerical feature or an embedded representation after preprocessing. Convolution operation W CNN ·X+b CNNIt is to perform convolution transformation on the input data to extract local features. The activation function ReLU is used to introduce nonlinearity to help the model better fit complex behavior patterns. The features extracted by CNN can capture the behavior patterns of employees in a specific time period or scenario, such as the frequent occurrence of a certain operation or the tendency of a specific behavior. These local features are the basis for further analysis by LSTM.
[0093] Furthermore, the long short-term memory network (LSTM) is a powerful tool for processing time series data, which can capture long-term dependencies in the data. Employee behavior data is usually time series data, such as the user's click records and operation time series within a certain period of time. LSTM can effectively capture this time dependency. The specific expression of LSTM is:
[0094] f LSTM (X) = LSTM (W LSTM ·X+b LSTM );
[0095] The input data X is the data after CNN extracts features, which contain key local information in employee behavior. The LSTM operation is used to learn the time series features in the data, retaining and forgetting information through cell states and gating mechanisms, thereby capturing the time dependency of employee behavior (for example, whether an employee will continue to perform a certain operation within a certain period of time). LSTM output can extract long-term patterns of employee behavior, which is crucial for predicting employees' future behavior and acceptance. LSTM can capture dynamic information such as employee interest evolution and behavior changes through historical behavior data, thereby helping to accurately predict employees' reactions to future promotional content.
[0096] Furthermore, the data processed by CNN and LSTM is input into the output layer, which uses the softmax activation function to generate the probability distribution of employee acceptance. Its expression is:
[0097] f output (X) = softmax(W output ·f LSTM (X)+b output );
[0098] The output layer generates the final prediction based on the results of LSTM.
[0099] W output ·f LSTM (X)+b outputis a linear transformation operation that maps the output of the LSTM to the prediction space of potential acceptance. The softmax activation function converts the output into a probability distribution, which represents the probability of an employee accepting or rejecting a promotional content. Through softmax, we can get the probability of each promotion strategy and make decisions accordingly.
[0100] In an embodiment of the present invention, an attention mechanism is added to the CNN model to enhance the model's ability to capture key features. The specific expression of the attention mechanism is:
[0101]
[0102] Where Q, K, V are query, key and value vectors respectively, QK is the dimension of the key vector;
[0103] In the LSTM model, the gated recurrent unit GRU is added to improve the model's time series processing capability. The specific expression of GRU is:
[0104] z t =σ(W z ·[h t-1 , x t ]+b z );
[0105] Among them, z t 、r t are the update gate and reset gate respectively, and σ is the sigmoid activation function.
[0106] In one possible implementation, the purpose of adding an attention mechanism to the CNN model is to enhance the model's ability to focus on key features, so that the model can automatically assign different attention weights in multi-dimensional employee behavior data, thereby better capturing and processing key features related to promotion effects. The attention mechanism calculates the similarity between the query vector and the key vector, determines which value vectors are more important in the current task, and weights and sums these important information. This weighted sum will be used as the input of the CNN model to more effectively capture key information related to promotion in employee behavior data.
[0107] Specifically, first, the employee's behavior data is mapped to the query (Q), key (K), and value (V) space. Furthermore, the similarity between the query vector and the key vector is calculated to obtain the attention weight. Furthermore, the attention weight is applied to the value vector to obtain the weighted feature representation. Finally, these weighted features are used as the input of the CNN model to further extract useful local features through the convolutional layer.
[0108] Furthermore, in order to enhance the processing capability of time series data, the gated recurrent unit (GRU) is introduced, which can better capture the temporal dependencies in employee behavior data. Compared with the standard LSTM, the GRU has a simpler structure and higher computational efficiency, but it can also effectively capture information in long time series.
[0109] Specifically, the update gate and reset gate are calculated using the input of the current time step and the hidden state of the previous time step. Further, the hidden state of the previous time step is updated according to the reset gate, and the candidate hidden state is calculated together with the current input. Further, the update gate is used to weight the combination of the current hidden state and the candidate hidden state to obtain the hidden state of the current time step. Further, the final hidden state is input to the next time step, or used as the output of the LSTM model.
[0110] In the entire model, the CNN+attention mechanism part first extracts local features from employee behavior data and automatically focuses on key features, which enhances the model's ability to capture important behavior patterns. Next, the GRU+LSTM part processes time series data and uses the GRU's gating mechanism to strengthen the learning of long-term time dependencies. Finally, combining the outputs of these two parts can generate employee acceptance predictions for future promotion strategies. Through this combination, when processing large-scale and complex employee behavior data, the model can not only identify key features more accurately, but also flexibly adjust in the time dimension to provide more personalized and dynamic promotion strategies, thereby effectively improving the success rate of promotion and user engagement.
[0111] In the embodiment of the present invention, in S3, the specific expression of the text generation model is:
[0112] f NLP (Y) = Transformer (W NLP ·Y+b NLP );
[0113] Among them, Y is the input internal employee potential acceptance data, W NLP and b NLP are weight and bias parameters respectively, and Transformer is a model for generating text.
[0114] In a possible implementation, first, the input employee potential acceptance data (obtained through the previous steps (such as employee behavior analysis, market feedback analysis, etc.) represents each employee's potential reaction or acceptance to the product promotion strategy, and is usually a high-dimensional vector.
[0115] Furthermore, in order to adapt this potential acceptance data to the input space of the Transformer model, the input data is first linearly transformed. This step converts the input data into a form suitable for Transformer processing, ensuring that the model can fully utilize the employee's potential acceptance information to generate appropriate text content.
[0116] Furthermore, the linearly transformed data is passed to the Transformer model. The core of the Transformer model is the self-attention mechanism, which can effectively capture the relationship between different parts of the input data. Through multiple encoding layers, the Transformer deeply represents the information in the input data and then generates text content related to the input data (employee acceptance).
[0117] Finally, the Transformer decoder generates the corresponding text output. These text contents will be personalized promotion strategies, copywriting or suggestions related to employee acceptance data, so that the promotion strategy can be more in line with employee preferences and behavior patterns.
[0118] In the embodiment of the present invention, in S4, the Q-learning algorithm is used to update the Q value according to the feedback of internal employees, and the DQN algorithm is used to optimize the recommendation strategy. The specific expression of Q-learning is:
[0119]
[0120] Among them, s is the current state, a is the current action, r is the reward value, α is the learning rate, and γ is the discount factor. The DQN algorithm approximates the Q value function through a neural network. The specific expression is:
[0121] Q(s, a; θ) = f DQN (s, a; θ);
[0122] Among them, θ is the network parameter, f DQN is the output function of the DQN model.
[0123] In one possible implementation, Q-learning is a reinforcement learning algorithm used to solve model-free decision-making problems. Its core is to learn an optimal decision-making strategy by updating the Q value. In step S4, the specific update rules of the Q-learning algorithm are as follows:
[0124] Among them, s is the current state, which represents the execution environment of the current promotion strategy among employees, such as employee feedback, current employee emotions, participation, etc.
[0125] a represents the current action, which is the promotion strategy adopted in the current state, such as recommending a specific product or promotion content.
[0126] r is the reward value, which reflects the employee's feedback on the current strategy (such as satisfaction, participation, etc.), and is usually a positive or negative numerical value.
[0127] α is the learning rate, which determines the influence degree of new information on the Q-value update.
[0128] γ is the discount factor, which represents the degree of value decay of future rewards and is usually set to a value less than 1.
[0129] By continuously updating the Q-value, the Q-learning algorithm can learn the optimal promotion strategy in different states, enabling the system to adjust the promotion method according to the employee's feedback, thereby optimizing the promotion effect.
[0130] Furthermore, in traditional Q-learning, the Q-value is usually updated by looking up a table. However, when the state and action spaces are very large, the method of looking up a table becomes infeasible. To solve this problem, the DQN (Deep Q-Network) algorithm is introduced, which approximates the Q-value function through a deep neural network. The Q-value approximation expression of DQN is:
[0131] Q(s, a; θ) = f DQN (s, a; θ);
[0132] where θ is the network parameter, which is continuously optimized by training the network.
[0133] f DQN is the output function of the DQN model, which is used to predict the Q-value under the given state and action.
[0134] DQN approximates the Q-value function through a deep neural network and can efficiently handle high-dimensional state and action spaces. In practical applications, DQN can optimize the recommendation strategy through the following steps:
[0135] The neural network is trained based on employee feedback (i.e., reward value) and the current state and action. Through the back-propagation algorithm, the parameters of the neural network are continuously optimized so that the network can accurately predict the Q value under a given state. Furthermore, the Q value generated by DQN is used to guide the update of the recommendation strategy. By maximizing the Q value, DQN can select the optimal promotion action, thereby maximizing employee participation and promotion effect. Furthermore, the DQN algorithm usually adopts the target network strategy, that is, two networks are used: one for generating Q values (current network) and the other for calculating the target Q value (target network). This strategy helps to reduce instability during training and improve training results.
[0136] In the embodiment of the present invention, S4 also includes a multi-agent reinforcement learning MARL mechanism, which optimizes the overall push strategy through collaborative learning of multiple agents. The specific expression of the multi-agent is:
[0137] Q MARL (s,a 1 , a 2 , ..., a n ) = Q(s, a 1 )+Q(s,a 2 )+...+Q(s,a n );
[0138] Among them, a 1 , a 2 , …, a n are the actions of multiple agents,
[0139] Add an adaptive reward mechanism to dynamically adjust the reward value based on the feedback of internal employees to improve learning efficiency; the specific expression of adaptive rewards is:
[0140] r adapt = r + β·Δ;
[0141] Among them, r is the original reward value, β is the adjustment coefficient, and Δ is the feedback change.
[0142] In one possible implementation, in traditional reinforcement learning (RL), there is usually only one agent that interacts with the environment and learns based on rewards. In multi-agent reinforcement learning (MARL), the system uses multiple agents to learn and collaborate with each other simultaneously to optimize the overall promotion strategy. The specific multi-agent Q value expression is:
[0143] Q MARL (s,a 1 , a 2 , ..., a n ) = Q(s, a 1)+Q(s,a 2 )+...+Q(s,a n );
[0144] Among them, a 1 , a 2 , …, a n are the actions of multiple agents, s is the current state of the environment. Each agent learns independently through interaction with the environment and adjusts its strategy by updating its own Q value.
[0145] Specifically, state representation: each agent observes the same environmental state (i.e., the same promotion task or product), but their decisions (actions) may be different.
[0146] Action selection: Each agent selects an action based on its state and current policy. The actions of different agents can be the same or different, depending on the system design.
[0147] Q-value update: Each agent updates its corresponding Q-value according to its own reward signal, and the overall Q-value represents the overall feedback of the system by summing up the Q-values of all agents. This mechanism encourages each agent to not only optimize its own strategy, but also consider the impact of other agents' behaviors on the overall reward.
[0148] Furthermore, the adaptive reward mechanism improves learning efficiency by dynamically adjusting the reward value. Especially in multi-agent systems, the feedback of each agent may change continuously, so the reward value needs to be adjusted in real time according to these changes to ensure the effectiveness of the learning process.
[0149] The expression of adaptive reward is: adapt = r + β·Δ;
[0150] Among them, r is the original reward value, which means the initial reward obtained by the agent after taking an action in the current state;
[0151] β is the adjustment coefficient, which is used to control the magnitude of reward adjustment;
[0152] Δ is the change in feedback. It indicates the degree of change in feedback compared to the previous state. Usually, Δ can be calculated based on changes in employee engagement, satisfaction, or behavior.
[0153] Specifically, reward adjustment: After the agent obtains the initial reward, the reward value r is adjusted by calculating the feedback change. adapt This process ensures that rewards not only reflect immediate results, but also take into account the changing trend of feedback, thereby helping the agent better adapt to the environment.
[0154] Reward Update: Adjusted Reward r adaptIt is passed to the agent to update its Q value. This adaptive reward mechanism can prompt the agent to adjust its learning strategy to make it more efficient when facing different environmental feedback.
[0155] In the embodiment of the present invention, in S5, the feedback and evaluation specifically adopt A / B test and multi-arm bandit algorithm, and the mathematical expression of feedback evaluation is as follows:
[0156] F = l(C, E; β);
[0157] Among them, F is the feedback evaluation result, C is the promotion content, E is the usage feedback of internal employees, and β is the parameter of the evaluation model.
[0158] In a possible implementation, in S5, the feedback and evaluation process is optimized by combining A / B testing and multi-arm bandit algorithm. The mathematical expression of feedback evaluation is:
[0159] F = l(C, E; β);
[0160] Among them, F is the feedback evaluation result, C is the promotion content, E is the usage feedback of internal employees, and β is the parameter of the evaluation model. This process combines A / B testing and multi-arm bandit algorithm to dynamically evaluate and optimize the promotion strategy.
[0161] Specifically, A / B testing is an experimental method used to compare the effectiveness of different promotional content or strategies. In A / B testing, promotional content is randomly assigned to different groups of employees (i.e., Group A and Group B), and then the feedback data of the two groups is compared to evaluate the effectiveness of the promotional content. A / B testing provides a basic framework for quickly verifying the effectiveness of different strategies in the early stages and helping to identify which promotional content or programs are most likely to attract employees' interest.
[0162] The multi-armed bandit algorithm maximizes expected rewards by dynamically allocating resources. During the promotion process, the system can display multiple versions of promotional content (for example, multiple ads or push notifications) at the same time and monitor employee feedback in real time. Based on these real-time feedback, the system will gradually adjust the distribution ratio of promotional content and allocate more resources (such as display frequency) to content with better performance, thereby maximizing the overall effect.
[0163] Specifically, the A / B testing phase:
[0164] First, employees are randomly assigned to different groups (e.g., Group A and Group B), and each group receives different promotional content. Further, feedback from each group is collected and analyzed, such as employee engagement, click-through rate, or product usage rate. Finally, the feedback from Group A and Group B is compared using statistical analysis methods to evaluate the initial effects of each promotional content.
[0165] Multi-arm bandit algorithm stage:
[0166] First, the promotional content that initially performed well in the A / B test is further optimized as multiple "arms". The system dynamically adjusts its allocation strategy based on the performance of each "arm". Furthermore, the system continuously obtains employee feedback from each promotional content, and updates the reward value of each content in real time, optimizing the probability of content display based on feedback signals. Finally, through the exploration and development mechanism of the multi-arm bandit algorithm, the efficiency of the promotion strategy is gradually improved, and ultimately the focus is on the most effective promotional content.
[0167] See also Figure 2 In this embodiment of the present invention, in S6, the steps of the adaptive adjustment mechanism include:
[0168] S6.1: Based on the feedback evaluation results, evaluate the effectiveness of the current promotion strategy and generate a status evaluation value;
[0169] S6.2: Generate new promotion strategies based on state evaluation values through adaptive control theory;
[0170] S6.3: Verify the generated new strategy in a small range and collect the verification results. During the verification process, apply the new strategy to some internal employees, collect their usage feedback and behavior data, and evaluate the effectiveness of the new strategy;
[0171] S6.4: Based on the verification results, the promotion strategy is optimized through adaptive control theory to generate the final promotion strategy;
[0172] S6.5: Apply the optimized promotion strategy to the content pushed to internal employees to achieve dynamic adjustment and optimization. The optimized strategy is applied in real time through the data collection system, and the pushed content is dynamically adjusted based on the real-time feedback of internal employees.
[0173] In one possible implementation, first, a state evaluation value is calculated by collecting feedback data (such as employee usage behavior data, participation, feedback, etc.) and combining the evaluation results of the aforementioned A / B test and multi-arm bandit algorithm. This evaluation value reflects the execution effect of the current promotion strategy. The generation of the state evaluation value is a quantitative process, which evaluates whether the current strategy has achieved the expected goals based on certain key performance indicators (KPIs) such as conversion rate and user satisfaction.
[0174] Furthermore, after obtaining the state evaluation value, the current promotion strategy is adjusted using adaptive control theory. Adaptive control theory optimizes the strategy through control algorithms and generates new promotion plans. The core of this process is to identify the state of the system based on the feedback results and adjust the control strategy based on this state so that the system maintains optimal operation in a dynamically changing environment. In this way, the promotion strategy can respond to the feedback and demand changes of internal employees in real time, adapt to different environmental changes, and optimize the promotion effect.
[0175] Furthermore, the generated new promotion strategy will be verified in a small range in this step. This small range can select some internal employee groups, and collect their usage feedback and behavior data by limiting the experimental scenario. This verification stage is crucial because it helps to judge the actual effect of the new strategy and discover potential strategy flaws. Employee feedback and behavior data will directly affect the further optimization of the strategy at this stage.
[0176] Furthermore, based on the results of small-scale verification, the new strategy is further adjusted and optimized through adaptive control theory. This process ensures that the strategy can be continuously improved and ultimately generates the optimal promotion strategy. The optimized strategy takes into account factors such as employee feedback, strategy implementation effect, and employee participation, and is deeply fine-tuned to adapt it to the needs of large-scale promotion.
[0177] Furthermore, after finally generating the optimal promotion strategy, the system applies the strategy to actual promotion activities. This process includes collecting employee feedback in real time through the data collection system and dynamically adjusting the push content based on this feedback. Through continuous dynamic adjustment, the promotion content can be consistent with the needs and interests of employees, thereby improving the overall effect of the promotion.
[0178] See also Figure 3 In the embodiment of the present invention, in S6.2, the adaptive control theory is implemented by the following steps:
[0179] S6.2.1: Define the relationship model between feedback evaluation results and current promotion strategies;
[0180] S6.2.2: Adjust the parameters of the promotion strategy based on the feedback evaluation results and model;
[0181] The specific expression is:
[0182] Among them, θ t is the strategy parameter, α t is the learning rate, is the performance indicator, the gradient of J with respect to the parameter θ.
[0183] In a possible implementation, first, step S6.2.1 aims to define a relationship model between feedback evaluation results and the current promotion strategy. The model describes the relationship between promotion strategy parameters (such as selection of promotion content, push frequency, personalized adjustment, etc.) and feedback evaluation results (such as employee engagement, conversion rate, satisfaction, etc.) through mathematical or machine learning methods. Usually, this relationship model can be a regression model, a neural network, a support vector machine, etc., to represent the mapping relationship between strategy parameters and feedback indicators.
[0184] For example, suppose that during the promotion process, the system pushes different content multiple times and collects employee participation feedback data. In this process, by analyzing the employee behavior data, a model can be built to predict how the employee's response or participation changes under different strategy parameters. This model can not only help us understand which strategy parameters have the greatest impact on the promotion effect, but also help us discover potential optimization directions.
[0185] Furthermore, in step S6.2.2, the relationship model established in S6.2.1 is used to adjust the parameters of the promotion strategy according to the current feedback evaluation results. Specifically, the model can calculate the performance indicators (such as conversion rate, satisfaction, etc.) of the promotion strategy under the current parameter settings, and further adjust the parameters through the gradient update rule so that these performance indicators can achieve better performance in the next round of strategy push. Through the adjustment method formula, the strategy parameters will be adjusted in the direction of optimal performance in each iteration. In other words, through gradient updates, the promotion strategy gradually converges to the parameter combination that can maximize the effect.
[0186] The two steps in S6.2 achieve dynamic adjustment and optimization of promotion strategies by establishing a relationship model between feedback and promotion strategies and using gradient optimization to update strategy parameters. This process ensures continuous improvement of promotion activities and enables optimal decisions to be made based on real-time data, thereby significantly improving promotion results.
[0187] The above describes in detail the optional implementation modes of the embodiments of the present invention in combination with the accompanying drawings. However, the embodiments of the present invention are not limited to the specific details in the above implementation modes. Within the technical concept of the embodiments of the present invention, various simple modifications can be made to the technical scheme of the embodiments of the present invention, and these simple modifications all belong to the protection scope of the embodiments of the present invention.
[0188] It should also be noted that the various specific technical features described in the above specific embodiments can be combined in any suitable manner without contradiction. To avoid unnecessary repetition, the embodiments of the present invention will not further describe various possible combinations.
[0189] Those skilled in the art can understand that all or part of the steps in implementing the methods of the above embodiments can be completed by instructing relevant hardware through a program. The program is stored in a storage medium, including several instructions to enable a single-chip microcomputer, a chip, or a processor to execute all or part of the steps of the methods described in various embodiments of the present application. The foregoing storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk, or an optical disc.
[0190] In addition, any combination can be made among various different embodiments of the embodiments of the present invention, as long as it does not violate the idea of the embodiments of the present invention, and it should also be regarded as the content disclosed in the embodiments of the present invention.
Claims
1. A method for promoting scientific and technological products based on AI, characterized in that: The method comprises: S1: Deploy a data collection system in the insurance company's internal network to collect internal employees' usage behavior data, including application usage frequency, dwell time, and function usage records. The data collection system is connected to GaussDB and OceanBase databases. After data collection is completed, the collected data is processed through preprocessing methods. S2: Build a behavioral model of internal employees based on the preprocessed data; S3: Generate personalized promotional content based on the output of the behavior model. The promotional content is generated using natural language processing (NLP) technology, including a text generation model. The generated promotional content includes but is not limited to user guides, operation videos, and FAQs. S4: Push personalized promotional content to internal employees, using reinforcement learning algorithms, including Q-learning and DQN, to dynamically adjust the push strategy based on real-time feedback from internal employees; S5: Collect internal employees’ feedback on the use of promotional content, including click-through rate, usage frequency, and satisfaction evaluation. The feedback data is transmitted to GaussDB and OceanBase databases in real time through the data collection system. The collected feedback data is used to evaluate the promotion effect. S6: Based on the feedback evaluation results, optimize the promotion strategy through an adaptive adjustment mechanism.
2. The AI-based technology product promotion method according to claim 1, characterized in that: In S1, a multi-level data preprocessing method is used, including data cleaning, data standardization, and data dimensionality reduction. The data cleaning method uses Z-score standardization, and the data dimensionality reduction uses principal component analysis (PCA) method.
3. The AI-based technology product promotion method according to claim 1, characterized in that: In S2, the behavior model uses deep learning algorithms, including convolutional neural networks (CNN) and long short-term memory networks (LSTM). The input of the behavior model is the usage behavior data of internal employees, and the output is the potential acceptance of new applications by internal employees.
4. The AI-based technology product promotion method according to claim 3, characterized in that: In the CNN model, an attention mechanism is added to enhance the model's ability to capture key features. The specific expression of the attention mechanism is: Where Q, K, V are query, key and value vectors respectively, QK is the dimension of the key vector; In the LSTM model, the gated recurrent unit GRU is added to improve the model's time series processing capability. The specific expression of GRU is: z t =σ(W z ·[h t-1 ,x t ]+b z ); Among them, z t 、r t are the update gate and reset gate respectively, and σ is the sigmoid activation function.
5. The AI-based technology product promotion method according to claim 1, characterized in that: In S3, the specific expression of the text generation model is: f NLP (Y)=Transformer(W NLP ·Y+b NLP ); Among them, Y is the input internal employee potential acceptance data, W NLP and b NLP are weight and bias parameters respectively, and Transformer is a model for generating text.
6. The AI-based technology product promotion method according to claim 1, characterized in that: In S4, the Q-learning algorithm is used to update the Q value based on the feedback from internal employees, and the DQN algorithm is used to optimize the recommendation strategy.
7. The AI-based technology product promotion method according to claim 6, characterized in that: In S4, it also includes the multi-agent reinforcement learning MARL mechanism, which optimizes the overall push strategy through the collaborative learning of multiple agents; It also includes adding an adaptive reward mechanism to dynamically adjust reward values based on internal employee feedback to improve learning efficiency.
8. The AI-based technology product promotion method according to claim 1, characterized in that: In S5, feedback and evaluation specifically use A / B testing and multi-arm bandit algorithm.
9. The AI-based technology product promotion method according to claim 1, characterized in that: In S6, the steps of the adaptive adjustment mechanism include: S6.1: Based on the feedback evaluation results, evaluate the effectiveness of the current promotion strategy and generate a status evaluation value; S6.2: Generate new promotion strategies based on state evaluation values through adaptive control theory; S6.3: Verify the generated new strategy in a small range and collect the verification results. During the verification process, apply the new strategy to some internal employees, collect their usage feedback and behavior data, and evaluate the effectiveness of the new strategy; S6.4: Based on the verification results, the promotion strategy is optimized through adaptive control theory to generate the final promotion strategy; S6.5: Apply the optimized promotion strategy to the content pushed to internal employees to achieve dynamic adjustment and optimization. The optimized strategy is applied in real time through the data collection system, and the pushed content is dynamically adjusted based on the real-time feedback of internal employees.
10. The AI-based technology product promotion method according to claim 1, characterized in that: In S6.2, the adaptive control theory is implemented by the following steps: S6.2.1: Define the relationship model between feedback evaluation results and current promotion strategies; S6.2.2: Adjust the parameters of the promotion strategy based on the feedback evaluation results and model.