Strategy generation method and device based on artificial intelligence, computer equipment and medium
By using a strategy generation method based on deep reinforcement learning networks, the problem of the inability to dynamically adjust sales scripts in traditional insurance training is solved. This method enables intelligent sales script recommendations based on changes in customer status, thereby improving training effectiveness.
Patent Information
- Application Number
- CN202610051077.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-15
- Publication Date
- 2026-04-28
AI Technical Summary
Traditional insurance industry training models rely on pre-set script libraries and rule engines, which cannot provide optimized script strategies based on trainees' actual performance and real-time changes in customer status, making it difficult for trainees to effectively deal with real and complex customers.
A strategy generation method based on deep reinforcement learning networks is adopted. By listening to the dialogue between the trainee and the virtual customer, the semantic vector is transformed and combined with the customer state space to select the speech action strategy, update the customer state, and extract features to match the historical dialogue scenario to obtain the speech strategy under the termination condition.
It enables the automatic and intelligent provision of real-time sales script strategies to trainees based on the reactions and status changes of virtual customers, thereby improving the intelligence and accuracy of sales script strategy recommendations.
Smart Images

Figure CN121935349A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology and can be applied to the financial technology field, particularly to artificial intelligence-based strategy generation methods, devices, computer equipment, and storage media. Background Technology
[0002] In traditional insurance industry training models, insurance sales training primarily relies on pre-set script libraries and rule engines. Specifically, by pre-setting scripts for various customer types and fixed interaction rules, trainees practice simulated sales based on these pre-set elements. While this training method improves upon the shortcomings of purely theoretical instruction to some extent, exposing trainees to a range of simulated customer scenarios, its limitations are also significant. Because the interaction process is often linear and static, the system can only respond according to predetermined rules and script libraries, lacking flexibility and dynamism. It struggles to provide effective decision support based on trainees' actual performance and real-time changes in customer status, failing to offer optimized sales strategies and hindering the overall improvement of trainees' sales capabilities.
[0003] For example, in the auto insurance sales scenario within the financial insurance sector, traditional training systems based on pre-set script libraries and rule engines can only provide simple responses to complex and flexible questions from customers regarding the correlation between different car models, driving habits, and premium discounts. They cannot delve into the underlying needs of the customer's questions, nor can they dynamically adjust their communication strategies according to the specific customer situation. This makes it difficult for trainees to effectively handle real, complex customer situations. Therefore, there is an urgent need for a new type of intelligent training system to overcome existing limitations, improve training effectiveness, and enhance the professional capabilities of insurance sales personnel. Summary of the Invention
[0004] The purpose of this application is to propose a strategy generation method, apparatus, computer device, and storage medium based on artificial intelligence, so as to solve the technical problem that existing training methods that rely on preset script libraries and rule engines cannot provide trainees with optimized script strategies.
[0005] Firstly, an artificial intelligence-based strategy generation method is provided, including: After listening to the dialogue request triggered by the student corresponding to the virtual customer, the system receives the first speech data input by the student. The first discourse data is converted into a semantic vector, and the semantic vector is combined with the customer state space of the virtual customer to obtain the corresponding initial state representation; Based on a pre-trained deep reinforcement learning network, the initial state representation is processed to select the speech action strategy, and the corresponding target speech action is obtained. Receive the second speech data input by the student corresponding to the target speech action, and update the customer state space based on the second speech data and preset dimension information to obtain the corresponding target customer state space; Determine whether the student's current conversation meets the termination condition; If the termination condition is met, feature extraction is performed on the target customer's state space to obtain state features, and it is determined whether there is a target dialogue scene that matches the state features in the preset historical dialogue scene feature library. If so, obtain the dialogue strategy corresponding to the target dialogue scenario and push the dialogue strategy to the student.
[0006] Secondly, an artificial intelligence-based strategy generation device is provided, comprising: The receiving module is used to receive the first speech data input by the student after listening to the dialogue request triggered by the student corresponding to the virtual customer; The first processing module is used to convert the first discourse data into a semantic vector and combine the semantic vector with the customer state space of the virtual customer to obtain the corresponding initial state representation. The selection module is used to select the speech action strategy based on the initial state representation using a pre-trained deep reinforcement learning network to obtain the corresponding target speech action. The update module is used to receive the second speech data corresponding to the target speech action input by the trainee, and update the customer state space based on the second speech data and preset dimension information to obtain the corresponding target customer state space. The judgment module is used to determine whether the student's current dialogue meets the termination condition; The second processing module is used to extract features from the target customer's state space to obtain state features if the termination condition is met, and to determine whether there is a target dialogue scene that matches the state features in the preset historical dialogue scene feature library. The third processing module is used to, if so, obtain the dialogue strategy corresponding to the target dialogue scenario and push the dialogue strategy to the student.
[0007] Thirdly, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the above-described artificial intelligence-based strategy generation method.
[0008] Fourthly, a computer-readable storage medium is provided, which stores a computer program that, when executed by a processor, implements the steps of the above-described artificial intelligence-based strategy generation method.
[0009] In the above-mentioned scheme implemented by the AI-based strategy generation method, apparatus, computer equipment, and storage medium, after listening to the dialogue request triggered by the student corresponding to the virtual customer, the system first receives the first utterance data input by the student; then, the first utterance data is converted into a semantic vector, and the semantic vector is combined with the customer state space of the virtual customer to obtain the corresponding initial state representation; then, based on a pre-trained deep reinforcement learning network, the initial state representation is processed to select a speech action strategy to obtain the corresponding target speech action; subsequently, the system receives the second utterance data input by the student corresponding to the target speech action, and updates the customer state space based on the second utterance data and preset dimension information to obtain the corresponding target customer state space; further, it determines whether the student's current dialogue meets the termination condition; if it meets the termination condition, the target customer state space is feature extracted to obtain state features, and it is determined whether there is a target dialogue scenario matching the state features in a preset historical dialogue scenario feature library; if so, the speech strategy corresponding to the target dialogue scenario is obtained, and the speech strategy is pushed to the student. Based on the above automated processing flow, this application, through the use of deep reinforcement learning networks, can automatically, intelligently, and accurately provide trainees with real-time script strategies according to the reactions and state changes of virtual customers, thereby improving the intelligence and accuracy of script strategy recommendations. Attached Figure Description
[0010] To more clearly illustrate the solutions in this application, the accompanying drawings used in the description of the embodiments of this application will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0011] Figure 1 This is an exemplary system architecture diagram to which this application can be applied; Figure 2 This is a flowchart of an embodiment of the AI-based strategy generation method according to this application; Figure 3 This is a schematic diagram of the structure of an embodiment of the AI-based strategy generation apparatus according to this application; Figure 4 This is a schematic diagram of the structure of one embodiment of the computer device according to this application. Detailed Implementation
[0012] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains; the terminology used herein in the specification of the application is for the purpose of describing particular embodiments only and is not intended to be limiting of the application; the terms "comprising" and "having," and any variations thereof, in the specification, claims, and foregoing drawings of this application, are intended to cover non-exclusive inclusion. The terms "first," "second," etc., in the specification, claims, or foregoing drawings of this application are used to distinguish different objects, not to describe a particular order.
[0013] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.
[0014] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings.
[0015] like Figure 1 As shown, system architecture 100 may include terminal device 101, network 102, and server 103. Terminal device 101 may be a laptop 1011, tablet 1012, or mobile phone 1013. Network 102 is used as a medium to provide a communication link between terminal device 101 and server 103. Network 102 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.
[0016] Users can use terminal device 101 to interact with server 103 via network 102 to receive or send messages, etc. Various communication client applications can be installed on terminal device 101, such as web browser applications, shopping applications, search applications, instant messaging tools, email clients, social media platform software, etc.
[0017] Terminal device 101 can be various electronic devices with a display screen and support web browsing. In addition to laptops 1011, tablets 1012, or mobile phones 1013, terminal device 101 can also be an e-book reader, an MP3 player (Moving Picture Experts Group Audio Layer III), an MP4 player (Moving Picture Experts Group Audio Layer IV), a laptop computer, and a desktop computer, etc.
[0018] Server 103 can be a server that provides various services, such as a backend server that supports the pages displayed on terminal device 101.
[0019] It should be noted that the AI-based strategy generation method provided in this application is generally executed by a server / terminal device, and correspondingly, the AI-based strategy generation device is generally located in the server / terminal device.
[0020] It should be understood that Figure 1 The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included.
[0021] Continue to refer to Figure 2 The flowchart illustrates an embodiment of the AI-based strategy generation method according to this application. The order of steps in the flowchart can be changed, and some steps can be omitted, depending on different needs. The AI-based strategy generation method provided in this application can be applied to any scenario requiring strategy generation, and therefore can be applied to products in these scenarios, such as strategy generation products in the financial insurance field. The AI-based strategy generation method includes the following steps: Step S201: After listening to the dialogue request triggered by the student corresponding to the virtual customer, the first speech data input by the student is received.
[0022] In this embodiment, the AI-based strategy generation method runs on an electronic device (e.g., Figure 1The server / terminal device shown can acquire the first speech data input by the trainee through wired or wireless connections. It should be noted that the aforementioned wireless connection methods may include, but are not limited to, 3G / 4G / 5G connections, WiFi connections, Bluetooth connections, WiMAX connections, Zigbee connections, UWB (ultra-wideband) connections, and other currently known or future-developed wireless connection methods. This application can be applied to speech training scenarios related to insurance personnel in the financial and insurance field. The implementing entity of this application is specifically a strategy generation system, also known as a training system, which can be simply referred to as the system. In this system, trainees actively initiate dialogues with virtual customers, which can be accomplished by clicking corresponding buttons on the interface or inputting specific instructions. For example, in a webpage interface for insurance sales training, when a trainee clicks the "Start Practice" button, the system knows that the trainee wants to start a communication with a virtual customer and inputs corresponding speech data, which can be speech-to-text or direct text input.
[0023] The system initializes customer states based on pre-set virtual customer profiles. These profiles include basic attributes such as age, occupation, and spending power. For example, a virtual customer might be a 35-year-old office worker with a monthly income of 10,000 yuan and a certain willingness to purchase insurance. The system assigns initial values to various relevant dimensions within the corresponding customer state space based on these attributes. For instance, the initial emotional state is set to "neutral," the insurance demand intensity is set to "moderate," and the historical conversation memory is initially empty. Furthermore, by engaging in dialogue practice with virtual customers, trainees can hone their sales skills, communication abilities, and adaptability in an environment closely resembling a real sales scenario. The system provides real-time sales strategy suggestions based on the virtual customer's reactions and state changes, helping trainees continuously improve their sales methods and increase their sales success rate. Simultaneously, trainees can accumulate sales experience by repeatedly practicing dialogues with virtual customers, becoming familiar with the characteristics and needs of different customer types.
[0024] Step S202: Convert the first discourse data into a semantic vector, and combine the semantic vector with the virtual customer's customer state space to obtain the corresponding initial state representation.
[0025] In this embodiment, the system utilizes a pre-trained language model (such as BERT, DeepSeek, etc.) to process the input first utterance data. The pre-trained language model acts like a language expert trained on a large amount of text data, capable of transforming the input text into a high-dimensional semantic vector. This vector contains semantic information about the utterance, such as the emotional tone and key themes expressed. For example, if a student inputs "I want to learn about insurance products," the pre-trained language model will convert it into a high-dimensional vector containing semantic information such as "inquiry" and "insurance products."
[0026] The system then combines the transformed semantic vector with the initialized customer state space to form a complete initial state representation. This initial state representation is structured data that integrates basic customer information and the semantic information of the learner's first round of utterances, providing a comprehensive foundation for subsequent decision-making. For example, the initial state representation includes information such as the customer's age and occupation, as well as the intent expressed in the learner's utterances.
[0027] By initializing the customer state space, the system can simulate a virtual customer with specific characteristics, making the dialogue more realistic. Converting student utterances into semantic vectors and representing them as initial states transforms complex natural language dialogues into structured data that computers can understand and process, providing the necessary data input for subsequent decision-making using a trained deep reinforcement learning network.
[0028] Step S203: Based on the pre-trained deep reinforcement learning network, the initial state representation is processed to select the speech action strategy to obtain the corresponding target speech action.
[0029] In this embodiment, the specific implementation process of selecting the speech action strategy based on the pre-trained deep reinforcement learning network to obtain the corresponding target speech action will be further described in detail in subsequent specific embodiments of this application, and will not be elaborated on here.
[0030] Step S204: Receive the second speech data input by the student corresponding to the target speech action, and update the customer state space based on the second speech data and preset dimension information to obtain the corresponding target customer state space.
[0031] In this embodiment, the system prompts the learner to respond based on the selected target dialogue action. The learner executes the target dialogue action as prompted. For example, if the system selects the dialogue action as "proving claims processing speed with data," the learner will say something like, "Our company's claims processing speed is very fast; according to statistics, 90% of cases are processed within 3 business days." The system then updates the state based on the learner's dialogue and the virtual customer's own state simulation rules. By utilizing the customer state simulation module, the system comprehensively considers the learner's dialogue content (secondary dialogue data), combined with the virtual customer's basic attributes, real-time emotional state, historical dialogue memory, and insurance demand intensity to update the customer state space, thus obtaining the updated target customer state space. For example, if the learner uses the dialogue "proving claims processing speed with data," and this dialogue is effective, the customer may change from a "skeptical" emotional state to "interested," and the intensity of their insurance demand may increase accordingly.
[0032] This step embodies the dynamism and interactivity of the dialogue. After the trainee performs the spoken actions, they will react and update their status accordingly to the virtual customer based on the content of the speech and their own state. This updating allows the customer's state to continuously change as the conversation progresses, more closely resembling the psychological and behavioral changes of a real customer during communication with a salesperson. Simultaneously, it provides new state information for subsequent decision-making, enabling the system to adjust its speaking strategies based on the customer's evolving state.
[0033] Step S205: Determine whether the student's current dialogue meets the termination condition.
[0034] In this embodiment, the system presets a series of dialogue termination conditions, such as the customer explicitly stating their intention to buy, refusing to buy, or reaching the maximum number of rounds in the dialogue. These conditions are set according to the actual sales scenario and training needs, and are used to determine whether the dialogue should end. The system monitors the progress of the dialogue in real time and determines whether the current dialogue meets the preset termination conditions. For example, if the customer says, "I have decided to buy this insurance," the system determines that the dialogue meets the termination condition of "the customer explicitly stating their intention to buy." If the dialogue has gone through 20 rounds, exceeding the preset maximum of 15 rounds, the system also determines that the dialogue has ended. If the dialogue has not ended, the system returns to step S202, using the new target customer state space as input, and continues with the next round of script selection; if the dialogue has ended, the system proceeds to the step "providing script strategies to trainees."
[0035] Step S206: If the termination condition is met, feature extraction is performed on the target customer state space to obtain state features, and it is determined whether there is a target dialogue scene that matches the state features in the preset historical dialogue scene feature library.
[0036] In this embodiment, the system extracts features from the target customer's state space to obtain the virtual customer's state features, forming a current state feature vector. These state features may include basic attributes (such as age, gender, occupation, income level, etc.), emotional states (such as happy, neutral, depressed, doubtful, etc.), insurance demand intensity (such as strong, average, weak, etc.), and historical dialogue memories (such as previously mentioned insurance products, points of interest, etc.). Furthermore, the collected feature data can be further processed using feature engineering. For example, age can be segmented into different age ranges; emotional states can be quantified into numerical values, such as 3 points for happy, 2 points for neutral, 1 point for depressed, and 1.5 points for doubtful; and for historical dialogue memories, key themes and keywords can be extracted as features.
[0037] Next, the similarity between the current state feature vector and each feature vector in the high-success-rate scenario feature library (i.e., the historical dialogue scenario feature library) is calculated. Similarity calculation methods can include cosine similarity, Euclidean distance, etc. For example, cosine similarity can be used to calculate the cosine of the angle between the current state feature vector and a feature vector in the feature library; the closer the value is to 1, the higher the similarity. A similarity threshold is set; when the similarity between the current state feature vector and a feature vector in the historical dialogue scenario feature library exceeds this threshold, it is determined that the current virtual customer's state is similar to that high-success-rate scenario (target dialogue scenario), thereby triggering the dialogue strategy suggestion mechanism.
[0038] In addition, the construction process of the aforementioned historical dialogue scenario feature library includes: Training data screening: Screening successful insurance purchase dialogue records from historical training data. These records contain complete customer status information at the time and the evolution of the customer's status during the dialogue. Feature extraction and storage: For each successful insurance purchase dialogue record, extracting the customer's status features at each stage of the dialogue, and combining and storing these features to form a high-success-rate scenario feature library, i.e., the historical dialogue scenario feature library. For example, for a successful case, the following features are recorded: the customer's initial age is 30 years old, occupation is white-collar worker, emotional state is neutral, insurance demand intensity is moderate, and during the dialogue, the emotion gradually changes to happy, and the insurance demand intensity becomes strong.
[0039] Step S207: If yes, obtain the dialogue strategy corresponding to the target dialogue scenario and push the dialogue strategy to the student.
[0040] In this embodiment, the specific implementation process of obtaining the dialogue strategy corresponding to the target dialogue scenario and pushing the dialogue strategy to the student will be described in more detail in subsequent specific embodiments of this application, and will not be elaborated on here.
[0041] This application first receives first utterance data input by the student after listening to a dialogue request triggered by a virtual customer. Then, it converts the first utterance data into a semantic vector and combines the semantic vector with the virtual customer's state space to obtain a corresponding initial state representation. Next, it uses a pre-trained deep reinforcement learning network to select a dialogue action strategy for the initial state representation, obtaining a corresponding target dialogue action. Subsequently, it receives second utterance data input by the student corresponding to the target dialogue action and updates the customer state space based on the second utterance data and preset dimension information to obtain a corresponding target customer state space. It further determines whether the student's current dialogue meets the termination condition. If it does, it extracts features from the target customer state space to obtain state features and determines whether a target dialogue scenario matching the state features exists in a preset historical dialogue scenario feature library. If so, it obtains the dialogue strategy corresponding to the target dialogue scenario and pushes the dialogue strategy to the student. Based on this automated processing flow, this application, through the use of a deep reinforcement learning network, can automatically, intelligently, and accurately provide students with real-time dialogue strategies based on the virtual customer's reactions and state changes, improving the intelligence and accuracy of dialogue strategy recommendations.
[0042] In some alternative implementations, step S203 includes the following steps: Invoke a pre-trained deep reinforcement learning network.
[0043] In this embodiment, the training process of the aforementioned deep reinforcement learning network aims to enable the Deep Q-Network (DQN) to learn the optimal verbal strategy. The following steps are executed: 1. Start the dialogue and initial state representation: Initiate the training dialogue, initialize the customer state space, and convert student utterances into semantic vectors to form an initial state representation, providing basic data for training. 2. Input the state representation into the DQN network for action selection: Input the initial state representation into the online network and calculate the Q-value of each action in the action space, providing a basis for strategy selection. 3. Select verbal actions using an ε-greedy strategy: Use an ε-greedy strategy to balance exploration and exploitation, selecting verbal actions to discover better strategies. 4. Execute the verbal action and update the customer state: Simulate the execution of the verbal action, updating the customer state space according to the simulation rules based on the verbal action and the customer's own state, providing a basis for subsequent reward calculation and strategy optimization. 5. Calculate the immediate reward: Calculate the immediate reward for this verbal action based on a multi-dimensional reward function, quantitatively evaluating the verbal effect. 6. Store Experience Tuples: Store the experience tuples from this interaction into the experience replay buffer, breaking data correlation and making network learning more stable. 7. Determine if the Dialogue Has Ended: Determine if the dialogue has ended based on preset conditions. If not, continue the loop; if ended, proceed to the next step. 8. Iterative Optimization of Network Parameters: After the dialogue ends, sample from the experience replay buffer, calculate the target Q-value using the target network, and update the online network parameters by minimizing the mean squared error loss, making Q-value prediction more accurate and gradually converging to learn the optimal strategy.
[0044] The specific network training process includes: Step 1: Speech Sequence Encoding and Multidimensional Representation of Customer State. First, the system needs to convert the trainee's dialogue content into machine-understandable feature vectors. This application uses pre-trained language models (such as BERT, DeepSeek, etc.) to convert each round of trainee input (whether speech-to-text or direct text input) into a high-dimensional semantic vector. This vector not only captures keywords but also contains semantic sentiment and intent information. Simultaneously, the system constructs a multidimensional state space (S) for the virtual customer, which includes at least the following dimensions: Basic attributes: such as age, occupation, and spending power (e.g., 35 years old, technical background, middle-to-high income). Real-time emotional state: such as interest, hesitation, aversion, and skepticism (e.g., price sensitivity, emotional state is "skepticism"). Historical dialogue memory: The system records summaries of previous dialogues to ensure that the customer can respond based on context. For example, if the trainee has already explained the claims process, the customer's subsequent questions will be more in-depth, rather than repeating basic questions. Insurance demand intensity: a hidden quantitative value that simulates the customer's purchase intention; this value dynamically changes with the trainee's persuasive skills. This step transforms complex sales dialogue scenarios into a structured, quantifiable state representation S_t, laying the foundation for subsequent decision optimization.
[0045] Step Two: DQN-Based Script Action Strategy Selection and Value Assessment The core of this solution is using the DQN algorithm to model the optimization process of script strategies. Each type of script response available to the trainee (e.g., "proving claims processing speed with data," "narrating successful claims cases") is defined as an action (A), and the set of all possible actions constitutes the action space (A). The core task of the DQN network is to learn a value function Q(S, A), which evaluates the cumulative expected return of taking a certain script action A_t under a specific customer state S_t. Forward Propagation: The current customer state representation S_t is input into the DQN online network, which calculates a Q-value for each possible action A_i in the action space A. This Q-value represents the long-term value of choosing that script. Strategy Selection: The system uses an ε-greedy strategy to select actions. In the early stages of training, actions are randomly selected with a high probability ε to fully explore the effects of various scripts; as training progresses, the probability of selecting the action with the highest current Q-value is gradually increased, i.e., utilizing the learned excellent strategies. For example, when faced with a customer who questions the speed of claims processing, the system might explore the action of "providing third-party data," even if it currently believes that "promising service" has the highest Q-value. Experience replay: The system stores the experience tuple (S_t, A_t, R_t, S_{t+1}) for each interaction in an experience replay buffer. Here, R_t is the immediate reward, and S_{t+1} is the customer's new state after taking the action. By randomly sampling experiences from the buffer for training, the correlation between data can be broken, making the network learning more stable.
[0046] Step 3: Multi-dimensional Reward Function Design and Real-time Feedback Reward Function. R serves as the guiding principle for the DQN agent to learn "excellent sales pitches." The system's reward function is meticulously designed, incorporating multiple quantitative indicators to comprehensively evaluate the immediate and long-term effects of the sales pitch: Customer Satisfaction Reward (R1): Based on emotional changes output by the customer state simulation module. For example, a positive reward is given when a customer changes from "hesitant" to "interested." Process Progression Reward (R2): Encourages the conversation to move towards a sale. For example, a higher reward is given if the sales pitch successfully guides the customer to the "quote" or "confirm information" stage. Sales Pitch Quality Reward (R3): Scored based on preset elements of excellent sales pitches (e.g., moderate speaking speed, no taboo words, complete coverage of key information points). For example, a positive reward is given for mentioning "specific disease coverage," while a negative reward is given for using prohibited promises such as "guaranteed compensation." Efficiency Reward (R4): Encourages the quick and effective resolution of customer objections. For example, resolving customer doubts about price within fewer conversation rounds earns an efficiency bonus. The total reward R_t = w1*R1 + w2*R2 + w3*R3 + w4*R4, where w1 to w4 are adjustable weighting coefficients. After each round of dialogue, the system calculates and displays the reward composition in real time, allowing trainees to clearly understand the strengths and weaknesses of their communication skills. For example: "Your explanation using a specific case effectively increased customer interest (R1++), but failed to guide them to the next step (R2-)."
[0047] Step Four: Iterative Optimization of Network Parameters and Evolution of Script Strategy. This is a crucial step for the system's self-learning and improvement. DQN uses a target network to calculate the target Q-value: Target Q = R_t + γ *max_{A} Q(S_{t+1}, A; θ-), where γ is the discount factor and θ- is the parameter of the target network. Then, the parameters θ of the online network are updated by minimizing the mean squared error loss: Loss = [Target Q - Q(S_t, A_t; θ)]^2. This optimization process uses the gradient descent algorithm to make the prediction of the Q-value increasingly accurate. Finally, after training with a large amount of dialogue data, the Q-network will gradually converge and learn an optimal policy π, that is, for any customer state S_t, it can give the script action A_t that will obtain the maximum long-term reward, thus obtaining a well-trained deep reinforcement learning network. This means that the system can not only provide practice, but also internalize a set of "golden script" strategies and can dynamically recommend them to learners. For example, when the system recognizes that the current virtual customer's status is similar to a high-success-rate scenario in the training data, it will directly prompt the trainee: "For this type of customer, it is recommended to use the 'case-first, data-supported' sales pitch strategy. Historical data shows that the success rate has increased by 28%."
[0048] Based on the deep reinforcement learning network, the corresponding value data is calculated for each action in the preset action space according to the initial state representation.
[0049] In this embodiment, the initial state representation is used as input data and prepared to be fed into the online network of the trained deep reinforcement learning network. The online network is the part of the DQN network used for real-time computation and decision-making; it has learned the relationship between different states and actions through a large amount of training data. Then, the online network calculates the Q-value (value data) for each possible action in the action space. The Q-value represents the long-term value that choosing that particular dialogue action can bring in the current customer state. For example, if the action space includes dialogue actions such as "introducing the coverage of insurance products," "emphasizing the cost-effectiveness of insurance products," and "sharing successful claims cases," the online network will calculate the Q-value of each of these actions in the current state.
[0050] Get a preset random value.
[0051] In this embodiment, the system generates a random number between 0 and 1.
[0052] Based on the random value and the value data, a target greedy strategy is used to filter out the corresponding specified action from the action space.
[0053] In this embodiment, when applying the trained network to actual practice, the system employs an ε-greedy strategy (target-greedy strategy) to select dialogue actions, and sets a relatively small fixed ε value, such as 0.1. This ε value is used to balance exploration and exploitation. A smaller ε value means that the system tends to utilize the learned best strategies more, while a smaller exploration probability retains a certain degree of flexibility to cope with various unexpected situations that may arise in actual dialogue. Furthermore, the system randomly selects an action from the action space with probability ε. For example, if there are 5 possible dialogue actions in the action space, the system generates a random number between 0 and 1. If this random number is less than ε (0.1), one of the 5 actions is randomly selected as the selected dialogue action (designated action). Alternatively, the system selects the action with the highest current Q value with probability 1-ε. For example, after online network calculation, the action "using data to prove the speed of claims settlement" has the highest Q value, and the system has a probability of (0.9) of selecting this action as the designated action.
[0054] The specified action is taken as the target speech action.
[0055] This application utilizes a pre-trained deep reinforcement learning network. Based on this network, it calculates the corresponding value data for each action in a preset action space according to the initial state representation. Then, it obtains a preset random value. Subsequently, based on the random value and the value data, it uses a target-greedy strategy to select the corresponding designated action from the action space. Finally, it uses the designated action as the target dialogue action. Based on this process, this application, by using a deep reinforcement learning network to calculate the corresponding value data for each action in a preset action space according to the initial state representation, and then using a target-greedy strategy to select the corresponding designated action as the required target dialogue action based on the preset random value and value data, can achieve the selection of the optimal dialogue action in most cases, while trying other actions with a certain probability, potentially discovering new and better strategies. This continuously improves the system's dialogue generation capability and enhances the intelligence and accuracy of the generated target dialogue actions.
[0056] In some optional implementations of this embodiment, step S207 includes the following steps: Retrieve the dialogue strategy that matches the target dialogue scenario from the preset strategy library.
[0057] In this embodiment, the aforementioned strategy library is an optimal strategy library built based on a deep reinforcement learning network. The construction process of the optimal strategy library includes: strategy collection and organization: During training, the system records high-performing sales strategies and corresponding success stories under various customer states. These strategies include the way the sales pitch is expressed, the logical order, and the key points emphasized. For example, for young customers who have some knowledge of insurance but not a strong need for it, an excellent sales strategy might start with a lighthearted topic and then stimulate their interest by sharing insurance claim cases from peers. Strategy classification and storage: The collected sales strategies are classified according to customer characteristics, insurance product type, and dialogue stage, and stored in the optimal strategy library. For example, strategies are categorized by customer age (young customer strategy, middle-aged customer strategy, elderly customer strategy, etc.) and by insurance product type (health insurance strategy, accident insurance strategy, life insurance strategy, etc.).
[0058] Specifically, when the system identifies that the current virtual customer's state characteristics are similar to a high-success-rate scenario (target dialogue scenario), it searches the optimal strategy library for a matching sales pitch strategy. For example, if the current virtual customer's state is similar to a scenario where health insurance was successfully sold to a middle-aged customer, the system will search for an excellent sales pitch strategy targeting middle-aged customers for health insurance.
[0059] Determine whether the number of the stated communication strategies is multiple.
[0060] In this embodiment, the number of dialogue strategies can be counted to determine whether there are multiple dialogue strategies. If there is only one dialogue strategy, then that strategy is directly pushed to the learner.
[0061] If so, each of the stated dialogue strategies is evaluated based on the preset strategy evaluation rules to obtain the strategy score of each stated dialogue strategy.
[0062] In this embodiment, the specific implementation process of evaluating each of the above-mentioned dialogue strategies based on preset strategy evaluation rules to obtain the strategy score of each of the dialogue strategies will be further described in detail in subsequent specific embodiments of this application, and will not be elaborated on here.
[0063] Select the target verbal strategy with the highest strategy score from all the stated verbal strategies.
[0064] In this embodiment, the target verbal strategy with the highest strategy score can be selected by comparing the strategy scores of all verbal strategies.
[0065] The target sales script strategy is pushed to the trainee.
[0066] In this embodiment, the specific implementation process of pushing the target speech strategy to the trainee will be described in more detail in subsequent specific embodiments of this application, and will not be elaborated on here.
[0067] Based on the above processing flow, this application uses a strategy library to query for dialogue strategies matching the target dialogue scenario. When multiple dialogue strategies are detected, it automatically and intelligently evaluates each strategy based on strategy evaluation rules, obtaining a strategy score for each strategy. Then, it selects the target dialogue strategy with the highest score from all strategies and pushes it to the learner. This ensures that when multiple dialogue strategies are matched, the highest-scoring target dialogue strategy that best fits the current dialogue context is intelligently and accurately pushed to the learner as a suggestion, thus providing targeted dialogue strategy recommendations and improving the accuracy and intelligence of dialogue strategy recommendations.
[0068] In some optional implementations, the evaluation of each of the dialogue strategies based on preset strategy evaluation rules to obtain a strategy score for each dialogue strategy includes the following steps: The success rate of a specified script strategy is evaluated to obtain the corresponding success rate value; wherein, the specified script strategy is any one of all the script strategies.
[0069] In this embodiment, for multiple matched sales pitch strategies, the number of times each strategy (such as a specified sales pitch strategy) was used and the number of successful transactions in historical successful cases are counted, and its success rate is calculated to obtain the corresponding success rate value. For example, if a certain sales pitch strategy was used 5 times in 10 successful cases, and all 5 times successfully facilitated a transaction, then the success rate of the strategy is 100%.
[0070] A satisfaction feedback evaluation is performed on the specified dialogue strategy to obtain the corresponding feedback value.
[0071] In this embodiment, in addition to success rate, the impact of customer feedback on the dialogue strategy is also considered. Data such as customer satisfaction ratings, objections raised, and resolution status during the dialogue process are collected. If a given dialogue strategy receives high customer satisfaction ratings in most successful cases and effectively resolves customer objections, then the given dialogue strategy performs well in terms of customer feedback. Specifically, the first number of satisfaction ratings corresponding to the given dialogue strategy and the second number of all ratings can be statistically analyzed, and the quotient between the first and second numbers can be calculated and used as the feedback value for the given dialogue strategy.
[0072] Obtain a first weight corresponding to the success rate value, and obtain a second weight corresponding to the feedback value.
[0073] In this embodiment, the values of the first weight corresponding to the success rate indicator and the second weight corresponding to the satisfaction indicator are not specifically limited, and can be determined according to the actual business evaluation needs. For example, the first weight corresponding to the success rate is 0.6, and the second weight corresponding to the satisfaction indicator is 0.4.
[0074] Based on the preset scoring formula, the success rate value, the feedback value, the first weight and the second weight are calculated and processed to obtain the corresponding calculation results.
[0075] In this embodiment, the above-mentioned score calculation formula is specifically a comprehensive evaluation formula based on weighted summation. It can be calculated by substituting the above-mentioned success rate value, feedback value, first weight and second weight into the corresponding positions in the above-mentioned score calculation formula, and the obtained calculation result is used as the specified strategy score of the specified speech strategy.
[0076] The calculation result is used as the designated strategy score for the designated speech strategy.
[0077] This application obtains a success rate value by evaluating the success rate of a specified dialogue strategy; wherein the specified dialogue strategy is any one of all dialogue strategies; then, a satisfaction feedback evaluation is performed on the specified dialogue strategy to obtain a corresponding feedback value; and a first weight corresponding to the success rate value and a second weight corresponding to the feedback value are obtained; then, based on a preset scoring formula, the success rate value, the feedback value, the first weight, and the second weight are calculated to obtain a corresponding calculation result; subsequently, the calculation result is used as the specified strategy score of the specified dialogue strategy. Based on the above processing flow, this application obtains the corresponding success rate and feedback value by evaluating the success rate and satisfaction feedback of a specified dialogue strategy, and then calculates the success rate, feedback value, and the obtained first and second weights based on the use of a scoring formula, and uses the obtained calculation result as the specified strategy score of the specified dialogue strategy, thereby achieving intelligent and accurate strategy evaluation processing of the specified dialogue strategy and ensuring the accuracy of the generated specified strategy score data.
[0078] In some optional implementations, pushing the target script strategy to the learner includes the following steps: Obtain auxiliary information corresponding to the target speech strategy.
[0079] In this embodiment, the aforementioned auxiliary information may include the background of the use of the target communication strategy in historical successful cases, the results achieved, etc., so that trainees can better understand and apply it.
[0080] The auxiliary information and the target speech strategy are integrated and processed to obtain the corresponding strategy prompt information.
[0081] In this embodiment, the above-mentioned auxiliary information and target speech strategy can be integrated and processed, and the resulting integrated information can be used as the corresponding strategy prompt information.
[0082] Get the preset push notification method.
[0083] In this embodiment, the selection of the above-mentioned push method is not specifically limited and can be determined according to the actual business needs. For example, any one of the various methods such as interface pop-up, voice prompts, and text annotations can be used to display the above-mentioned strategy prompt information to the students.
[0084] Based on the aforementioned push method, the strategy prompt information is sent to the student.
[0085] In this embodiment, the generated strategy prompts can be presented to trainees in a clear and concise manner, depending on the selected push method. The prompts may include the specific wording of the sales pitch, key points emphasized, and the appropriate timing for its use. For example, trainees may be prompted that "for this type of customer, we recommend using a 'case-based, data-supported' sales pitch strategy, which historical data shows increases the success rate by 28%."
[0086] This application obtains auxiliary information corresponding to the target dialogue strategy; then integrates the auxiliary information with the target dialogue strategy to obtain corresponding strategy prompt information; subsequently, it obtains a preset push method; and then sends the strategy prompt information to the trainee based on the push method. Based on the above processing flow, this application improves the data richness of the generated strategy prompt information by obtaining auxiliary information corresponding to the target dialogue strategy and integrating the auxiliary information with the target dialogue strategy. Furthermore, by using a push method to send the strategy prompt information to the trainee, it ensures that the trainee can obtain the strategy prompt information in a timely and accurate manner, which is beneficial to improving the trainee's user experience and training effectiveness.
[0087] In some optional implementations of this embodiment, after step S207, the electronic device may further perform the following steps: After processing the dialogue request, relevant data corresponding to the target reward function is collected; wherein, the target reward function is a multi-dimensional reward function corresponding to the deep reinforcement learning network.
[0088] In this embodiment, the construction of the aforementioned target reward function includes: Customer Satisfaction Reward (R1): Calculated based on indicators such as the virtual customer's emotional changes during the dialogue and their feedback after the dialogue. For example, the system can set up an emotion rating table, assigning corresponding scores based on the customer's emotional state changes during the dialogue; a change from frustration to happiness earns a higher score. Additionally, if the customer expresses satisfaction or willingness to learn more after the dialogue, the customer satisfaction reward score will also increase. Process Progression Reward (R2): Measures whether the dialogue proceeds according to the expected process and effectively advances the sales process. For example, the system sets key nodes in the sales process, such as product introduction, needs confirmation, objection handling, and closing the deal. A corresponding score reward is given for each successful advancement of a key node. Script Quality Reward (R3): Evaluated from aspects such as the logic, professionalism, and attractiveness of the script. The system can pre-set some evaluation criteria, such as whether the script is clear, whether appropriate cases and data are used, and whether it effectively addresses customer concerns. The script is scored based on these criteria to derive the script quality reward. Efficiency Reward (R4): Considers the duration and number of rounds of the dialogue. Higher efficiency rewards are given if the conversation can achieve the expected goal, such as completing a product introduction or closing a deal, in a shorter duration and fewer rounds.
[0089] The relevant data is processed by the target reward function to obtain the corresponding target reward score data.
[0090] In this embodiment, the specific implementation process of performing reward calculation on the relevant data based on the target reward function to obtain the corresponding target reward score data will be further described in detail in subsequent specific embodiments of this application, and will not be elaborated on here.
[0091] Feedback analysis is performed on the target reward score data to obtain corresponding feedback analysis data.
[0092] In this embodiment, the system can analyze the trainee's performance based on the obtained target reward score data, indicating which aspects the trainee performed well in and which aspects need improvement, thus obtaining corresponding feedback analysis data. For example, if a trainee's speech quality reward is low, the system can prompt the trainee to improve the logic and professionalism of their speech.
[0093] The target reward score data and the feedback analysis data are displayed and processed based on a preset display strategy.
[0094] In this embodiment, the calculated reward scores for each dimension and the target reward score data (total reward) can be displayed to trainees in an intuitive way. For example, charts and numerical displays on the interface can be used to allow trainees to clearly understand their performance in each dimension. For example, it can display "Customer satisfaction reward: 8 points, process progress reward: 6 points, script quality reward: 7 points, efficiency reward: 5 points, total reward: 26 points".
[0095] After processing the dialogue request, this application collects relevant data corresponding to the target reward function, wherein the target reward function is a multi-dimensional reward function corresponding to the deep reinforcement learning network. Then, based on the target reward function, reward calculation is performed on the relevant data to obtain the corresponding target reward score data. Afterwards, feedback analysis is performed on the target reward score data to obtain corresponding feedback analysis data. Subsequently, the target reward score data and the feedback analysis data are displayed based on a preset display strategy. Based on the above processing flow, this application, through the use of the multi-dimensional reward function corresponding to the deep reinforcement learning network, can automatically and intelligently calculate the student's target reward score data, perform feedback analysis on the target reward score data to obtain feedback analysis data, and then display the target reward score data and feedback analysis data based on the use of a display strategy. This allows students to clearly understand the strengths and weaknesses of their dialogue skills, thereby enabling targeted improvements. It also provides a reference for the system to further optimize dialogue strategy suggestions.
[0096] In some optional implementations of this embodiment, the step of performing reward calculation processing on the relevant data based on the target reward function to obtain the corresponding target reward score data includes the following steps: The relevant data are calculated and processed based on the target reward function to obtain the corresponding customer satisfaction reward score, process progress reward score, communication quality reward score, and efficiency reward score.
[0097] In this embodiment, after a student completes a round of dialogue, the system collects various data related to the target reward function, including records of the virtual customer's emotional changes, the completion status of key nodes in the dialogue process, the content of the dialogue used by the student, the dialogue duration, and the number of rounds. The reward calculation includes: calculating scores for customer satisfaction reward (R1), process progress reward (R2), dialogue quality reward (R3), and efficiency reward (R4) based on the pre-set target reward function and the collected data. For example, the R1 score is calculated according to the customer's emotional changes using an emotion rating scale; the R2 score is calculated based on the completion status of key nodes in the dialogue process; the R3 score is obtained by scoring the dialogue according to the dialogue evaluation criteria; and the R4 score is calculated based on the dialogue duration and the number of rounds.
[0098] Obtain the preset weighting coefficients.
[0099] In this embodiment, the aforementioned weighting coefficients are weighting coefficients corresponding to the reward scores of each dimension. There are no specific limitations on the selection of the weighting coefficients; they can be set according to actual business needs and adjusted based on the training focus and objectives. For example, if customer satisfaction is a greater focus, a higher weight can be assigned to R1.
[0100] The customer satisfaction reward score, the process advancement reward score, the communication quality reward score, and the efficiency reward score are calculated based on the weighting coefficients to obtain the corresponding total score data.
[0101] In this embodiment, the total reward for this round of dialogue, i.e., the total score data, can be obtained by using the weighting coefficients mentioned above to match and weight the reward scores of each dimension, namely customer satisfaction reward score, process progress reward score, dialogue quality reward score and efficiency reward score.
[0102] The total score data is used as the target reward score data.
[0103] This application calculates and processes the relevant data based on the target reward function to obtain corresponding customer satisfaction reward scores, process advancement reward scores, script quality reward scores, and efficiency reward scores; then, it obtains preset weighting coefficients; subsequently, it calculates and processes the customer satisfaction reward scores, process advancement reward scores, script quality reward scores, and efficiency reward scores based on the weighting coefficients to obtain corresponding total score data; finally, it uses the total score data as the target reward score data. Based on the above processing flow, this application recalculates the customer satisfaction reward scores, process advancement reward scores, script quality reward scores, and efficiency reward scores obtained by calculating and processing the relevant data based on the target reward function using the obtained weighting coefficients, and uses the obtained total score data as the corresponding target reward score data. This enables automatic and accurate completion of reward calculation processing for relevant data, improves the processing efficiency of reward calculation, and ensures the accuracy of the obtained target reward score data.
[0104] In some alternative implementations, the user information obtained is subject to user consent and complies with relevant laws and policies.
[0105] Furthermore, any software tools or components not belonging to our company that appear in the embodiments of this application are merely illustrative examples and do not represent actual use.
[0106] Furthermore, this application has the following advantages over the prior art: 1. Adaptive and Personalized Training Paths: This system completely revolutionizes the "one-size-fits-all" training model. Through meticulous quantitative evaluation of each trainee's performance in every round of dialogue, the DQN model dynamically perceives each trainee's unique strengths and weaknesses (e.g., someone excels at product presentations but struggles with handling objections). Consequently, the system automatically adjusts the type of virtual customer, the difficulty of the dialogue scenarios, and the focus of the challenges in subsequent practice sessions, tailoring a progressive growth path from easy to difficult for each trainee. This highly personalized training model ensures that every exercise directly addresses the trainee's "pain points," thereby significantly shortening the training cycle and helping newcomers reach or even surpass the practical skills of experienced sales personnel more quickly.
[0107] 2. Enhance the ability to handle complex real-world scenarios: By introducing customer state sequences and multi-round game modeling, this solution creates a highly realistic training environment, far surpassing older systems based on fixed scripts. Trainees face "virtual customers" whose emotions and needs change in real time, and who can remember their words and actions. This effectively trains their ability to adapt to dynamic environments, control the pace of the game, and make continuous decisions. This training enables trainees to remain calm and flexible when facing unexpected reactions from real customers, seamlessly transferring training results to real-world work scenarios and significantly improving the success rate of closing deals in real-world situations.
[0108] 3. From "Error Correction" to "Empowerment": Providing Quantitative Strategy Support: The biggest leap forward for this system is that it is not merely a "judge," but a "strategic coach." Based on the optimal strategy library learned by DQN, the system can provide trainees with data-driven, positive suggestions for optimizing their sales pitches. For example, it will explicitly point out: "When facing budget-conscious customers, using the 'comparing long-term benefits' strategy in the third interaction is expected to increase the probability of closing the deal by 20% compared to continuing to emphasize 'brand advantages.'" This insight shifts the training focus from avoiding mistakes to learning and mastering high-win-rate strategies, accelerating the growth of novices and providing a scientific and quantifiable reference for experienced sales personnel to improve their skills.
[0109] This application creates a continuously evolving and highly intelligent training environment by deeply integrating the DQN reinforcement learning algorithm with the needs of sales script training. It not only solves the core pain points of existing training technologies, but also brings revolutionary changes to the talent training model of the insurance industry, effectively helping practitioners to make the leap from "qualified" to "excellent".
[0110] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0111] It should be emphasized that, to further ensure the privacy and security of the above-mentioned communication strategies, these strategies can also be stored in a blockchain node.
[0112] The blockchain referred to in this application is a novel application model of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanisms, and encryption algorithms. Essentially, a blockchain is a decentralized database, a chain of data blocks linked together using cryptographic methods. Each data block contains information about a batch of network transactions, used to verify the validity of the information (anti-counterfeiting) and generate the next block. A blockchain can include an underlying blockchain platform, a platform product service layer, and an application service layer.
[0113] The embodiments of this application can acquire and process relevant data based on artificial intelligence technology. Artificial intelligence (AI) refers to the theories, methods, technologies, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results. Foundational technologies in artificial intelligence generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies mainly encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning.
[0114] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by instructing related hardware through computer-readable instructions. These computer-readable instructions can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. The aforementioned storage medium can be a non-volatile storage medium such as a magnetic disk, optical disk, or read-only memory (ROM), or random access memory (RAM).
[0115] It should be understood that although the steps in the flowcharts of the accompanying figures are shown sequentially as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the accompanying figures may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the sub-steps or stages of other steps.
[0116] Further reference Figure 3 As a response to the above Figure 2 To implement the method shown, this application provides an embodiment of an artificial intelligence-based strategy generation device, which is similar to... Figure 2 Corresponding to the method embodiments shown, this device can be specifically applied to various electronic devices.
[0117] like Figure 3As shown, the AI-based strategy generation device 300 described in this embodiment includes: a receiving module 301, a first processing module 302, a selection module 303, an update module 304, a judgment module 305, a second processing module 306, and a third processing module 307. Wherein: The receiving module 301 is used to receive the first speech data input by the student after listening to the dialogue request triggered by the student corresponding to the virtual customer; The first processing module 302 is used to convert the first discourse data into a semantic vector and combine the semantic vector with the customer state space of the virtual customer to obtain the corresponding initial state representation. Selection module 303 is used to select the speech action strategy based on the initial state representation based on a pre-trained deep reinforcement learning network to obtain the corresponding target speech action. The update module 304 is used to receive the second speech data input by the trainee corresponding to the target speech action, and update the customer state space based on the second speech data and preset dimension information to obtain the corresponding target customer state space. The judgment module 305 is used to determine whether the current dialogue of the student meets the termination condition; The second processing module 306 is used to extract features from the target customer's state space to obtain state features if the termination condition is met, and to determine whether there is a target dialogue scene that matches the state features in the preset historical dialogue scene feature library. The third processing module 307 is used to, if so, obtain the dialogue strategy corresponding to the target dialogue scenario and push the dialogue strategy to the student.
[0118] In some optional implementations of this embodiment, the selection module 303 includes: Calling submodules is used to invoke pre-trained deep reinforcement learning networks; The first calculation submodule is used to calculate the corresponding value data for each action in the preset action space based on the deep reinforcement learning network and the initial state representation. The first acquisition submodule is used to acquire a preset random value; The first filtering submodule is used to filter out the corresponding specified action from the action space based on the random value and the value data using a target greedy strategy. The first determining submodule is used to take the specified action as the target speech action.
[0119] In some optional implementations of this embodiment, the third processing module 307 includes: The query submodule is used to retrieve the dialogue strategy that matches the target dialogue scenario from the preset strategy library; The judgment submodule is used to determine whether the number of the speech strategies is multiple; The evaluation submodule is used to evaluate each of the stated dialogue strategies based on preset strategy evaluation rules, and obtain the strategy score of each stated dialogue strategy. The second filtering submodule is used to filter out the target verbal strategy with the highest strategy score from all the verbal strategies. The push submodule is used to push the target speech strategy to the trainees.
[0120] In some optional implementations of this embodiment, the evaluation submodule includes: The first evaluation unit is used to evaluate the success rate of a specified script strategy and obtain the corresponding success rate value; wherein, the specified script strategy is any one of all the script strategies. The second evaluation unit is used to evaluate the satisfaction of the specified script strategy and obtain the corresponding feedback value. The first acquisition unit is used to acquire a first weight corresponding to the success rate value and to acquire a second weight corresponding to the feedback value. The calculation unit is used to calculate the success rate value, the feedback value, the first weight and the second weight based on a preset score calculation formula to obtain the corresponding calculation result. A determining unit is used to use the calculation result as the specified strategy score of the specified speech strategy.
[0121] In some optional implementations of this embodiment, the push submodule includes: The second acquisition unit is used to acquire auxiliary information corresponding to the target speech strategy; An integration unit is used to integrate the auxiliary information and the target speech strategy to obtain corresponding strategy prompt information; The third acquisition unit is used to acquire the preset push method; The sending unit is used to send the strategy prompt information to the student based on the push method.
[0122] In some optional implementations of this embodiment, the AI-based strategy generation device further includes: The collection module is used to collect relevant data corresponding to the target reward function after processing the dialogue request; wherein the target reward function is a multi-dimensional reward function corresponding to the deep reinforcement learning network. The calculation module is used to perform reward calculation processing on the relevant data based on the target reward function to obtain the corresponding target reward score data; The analysis module is used to perform feedback analysis on the target reward score data to obtain corresponding feedback analysis data; The display module is used to display the target reward score data and the feedback analysis data based on a preset display strategy.
[0123] In some optional implementations of this embodiment, the calculation module includes: The second calculation submodule is used to calculate and process the relevant data based on the target reward function to obtain the corresponding customer satisfaction reward score, process progress reward score, script quality reward score and efficiency reward score; The second acquisition submodule is used to acquire preset weighting coefficients; The third calculation submodule is used to calculate the customer satisfaction reward score, the process advancement reward score, the speech quality reward score and the efficiency reward score based on the weighting coefficients to obtain the corresponding total score data; The second determining submodule is used to use the total score data as the target reward score data.
[0124] To address the aforementioned technical problems, embodiments of this application also provide a computer device. Please refer to [link / reference needed]. Figure 4 , Figure 4 This is a basic structural block diagram of the computer device in this embodiment.
[0125] The computer device 4 includes a memory 41, a processor 42, and a network interface 43 that are interconnected via a system bus. It should be noted that only the computer device 4 with components 41-43 is shown in the figure; however, it should be understood that it is not required to implement all the shown components, and more or fewer components can be implemented alternatively. Those skilled in the art will understand that the computer device described here is a device capable of automatically performing numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes, but is not limited to, microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.
[0126] The computer device can be a desktop computer, laptop, handheld computer, or cloud server, etc. The computer device can interact with the user via a keyboard, mouse, remote control, touchpad, or voice control.
[0127] The memory 41 includes at least one type of readable storage medium, including flash memory, hard disk, multimedia card, card-type memory (e.g., SD or DX memory), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the memory 41 may be an internal storage unit of the computer device 4, such as the hard disk or memory of the computer device 4. In other embodiments, the memory 41 may also be an external storage device of the computer device 4, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc., equipped on the computer device 4. Of course, the memory 41 may also include both the internal storage unit and its external storage device of the computer device 4. In this embodiment, the memory 41 is typically used to store the operating system and various application software installed on the computer device 4, such as computer-readable instructions for strategy generation methods based on artificial intelligence. In addition, the memory 41 can also be used to temporarily store various types of data that have been output or will be output.
[0128] In some embodiments, the processor 42 may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chip. The processor 42 is typically used to control the overall operation of the computer device 4. In this embodiment, the processor 42 is used to execute computer-readable instructions stored in the memory 41 or to process data, for example, to execute computer-readable instructions of the artificial intelligence-based strategy generation method.
[0129] The network interface 43 may include a wireless network interface or a wired network interface, which is typically used to establish communication connections between the computer device 4 and other electronic devices.
[0130] This application also provides another embodiment, namely, providing a computer-readable storage medium storing computer-readable instructions that can be executed by at least one processor to cause the at least one processor to perform the steps of the artificial intelligence-based policy generation method described above.
[0131] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of this application.
[0132] Obviously, the embodiments described above are only some embodiments of this application, not all embodiments. The accompanying drawings show preferred embodiments of this application, but do not limit the patent scope of this application. This application can be implemented in many different forms; rather, the purpose of providing these embodiments is to provide a more thorough and comprehensive understanding of the disclosure of this application. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing specific embodiments, or make equivalent substitutions for some of the technical features. Any equivalent structures made using the content of this application's specification and drawings, directly or indirectly applied to other related technical fields, are similarly within the scope of patent protection of this application.
Claims
1. A strategy generation method based on artificial intelligence, characterized in that, Includes the following steps: After listening to the dialogue request triggered by the student corresponding to the virtual customer, the system receives the first speech data input by the student. The first discourse data is converted into a semantic vector, and the semantic vector is combined with the customer state space of the virtual customer to obtain the corresponding initial state representation; Based on a pre-trained deep reinforcement learning network, the initial state representation is processed to select the speech action strategy, and the corresponding target speech action is obtained. Receive the second speech data input by the student corresponding to the target speech action, and update the customer state space based on the second speech data and preset dimension information to obtain the corresponding target customer state space; Determine whether the student's current conversation meets the termination condition; If the termination condition is met, feature extraction is performed on the target customer's state space to obtain state features, and it is determined whether there is a target dialogue scene that matches the state features in the preset historical dialogue scene feature library. If so, obtain the dialogue strategy corresponding to the target dialogue scenario and push the dialogue strategy to the student.
2. The strategy generation method based on artificial intelligence according to claim 1, characterized in that, The step of selecting a speech action strategy based on the initial state representation using a pre-trained deep reinforcement learning network to obtain the corresponding target speech action specifically includes: Call upon a pre-trained deep reinforcement learning network; Based on the deep reinforcement learning network, the corresponding value data is calculated for each action in the preset action space according to the initial state representation. Get a preset random value; Based on the random value and the value data, a target greedy strategy is used to filter out the corresponding specified action from the action space; The specified action is taken as the target speech action.
3. The strategy generation method based on artificial intelligence according to claim 1, characterized in that, The step of acquiring the dialogue strategy corresponding to the target dialogue scenario and pushing the dialogue strategy to the trainee specifically includes: Retrieve the dialogue strategy that matches the target dialogue scenario from the preset strategy library; Determine whether the number of the stated communication strategies is multiple; If so, each of the aforementioned dialogue strategies is evaluated based on preset strategy evaluation rules to obtain a strategy score for each of the aforementioned dialogue strategies; Select the target conversation strategy with the highest strategy score from all the stated conversation strategies; The target sales script strategy is pushed to the trainee.
4. The strategy generation method based on artificial intelligence according to claim 3, characterized in that, The step of evaluating each of the dialogue strategies based on preset strategy evaluation rules to obtain a strategy score for each dialogue strategy specifically includes: The success rate of a specified script strategy is evaluated to obtain the corresponding success rate value; wherein, the specified script strategy is any one of all the script strategies. A satisfaction feedback evaluation is performed on the specified script strategy to obtain the corresponding feedback value; Obtain a first weight corresponding to the success rate value, and obtain a second weight corresponding to the feedback value; Based on the preset scoring formula, the success rate value, the feedback value, the first weight and the second weight are calculated and processed to obtain the corresponding calculation results; The calculation result is used as the designated strategy score for the designated speech strategy.
5. The strategy generation method based on artificial intelligence according to claim 3, characterized in that, The step of pushing the target script strategy to the trainee specifically includes: Obtain auxiliary information corresponding to the target speech strategy; The auxiliary information and the target speech strategy are integrated and processed to obtain the corresponding strategy prompt information; Get the preset push notification method; Based on the aforementioned push method, the strategy prompt information is sent to the student.
6. The strategy generation method based on artificial intelligence according to claim 1, characterized in that, After the step of obtaining the dialogue strategy corresponding to the target dialogue scenario and pushing the dialogue strategy to the student, the method further includes: After processing the dialogue request, relevant data corresponding to the target reward function is collected; wherein, the target reward function is a multi-dimensional reward function corresponding to the deep reinforcement learning network; Based on the target reward function, the relevant data is processed to obtain the corresponding target reward score data; The target reward score data is subjected to feedback analysis to obtain corresponding feedback analysis data; The target reward score data and the feedback analysis data are displayed and processed based on a preset display strategy.
7. The strategy generation method based on artificial intelligence according to claim 6, characterized in that, The step of performing reward calculation processing on the relevant data based on the target reward function to obtain the corresponding target reward score data specifically includes: Based on the target reward function, the relevant data are calculated and processed to obtain the corresponding customer satisfaction reward score, process progress reward score, script quality reward score, and efficiency reward score. Obtain the preset weighting coefficients; Based on the weighting coefficients, the customer satisfaction reward score, the process advancement reward score, the script quality reward score, and the efficiency reward score are calculated and processed to obtain the corresponding total score data. The total score data is used as the target reward score data.
8. A strategy generation device based on artificial intelligence, characterized in that, include: The receiving module is used to receive the first speech data input by the student after listening to the dialogue request triggered by the student corresponding to the virtual customer; The first processing module is used to convert the first discourse data into a semantic vector and combine the semantic vector with the customer state space of the virtual customer to obtain the corresponding initial state representation. The selection module is used to select the speech action strategy based on the initial state representation using a pre-trained deep reinforcement learning network to obtain the corresponding target speech action. The update module is used to receive the second speech data corresponding to the target speech action input by the trainee, and update the customer state space based on the second speech data and preset dimension information to obtain the corresponding target customer state space. The judgment module is used to determine whether the student's current dialogue meets the termination condition; The second processing module is used to extract features from the target customer's state space to obtain state features if the termination condition is met, and to determine whether there is a target dialogue scene that matches the state features in the preset historical dialogue scene feature library. The third processing module is used to, if so, obtain the dialogue strategy corresponding to the target dialogue scenario and push the dialogue strategy to the student.
9. A computer device, characterized in that, The system includes a memory and a processor, wherein the memory stores computer-readable instructions, and the processor executes the computer-readable instructions to implement the steps of the artificial intelligence-based strategy generation method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-readable instructions, which, when executed by a processor, implement the steps of the artificial intelligence-based strategy generation method as described in any one of claims 1 to 7.