Sale partner training method and system based on GPT and multi-modal large model
Through the sales training method based on GPT and multimodal large models, the problem of insufficient response to customer emotional changes in traditional sales training has been solved, personalized sales scenario simulation and real-time evaluation have been achieved, and the sales staff's response ability and training effect have been improved.
Patent Information
- Application Number
- CN202510858960.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-25
- Publication Date
- 2025-09-19
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Traditional sales coaching methods are unable to effectively respond to changes in customer emotions and personality differences, resulting in a decrease in the sales staff's interaction success rate in real sales scenarios. The training results are difficult to reflect the actual response level and transaction skills.
A sales coaching method based on GPT and multimodal large models is adopted to generate personalized sales dialogue scenarios by obtaining basic customer attributes, financial product characteristics and market hot topics. The multimodal Transformer encoder is used to screen interaction features, build virtual sales roles, monitor and evaluate the response capabilities of sales personnel in real time, and provide quantitative training scores.
It has significantly improved the coverage and flexibility of sales training, enhanced the realism of scenario simulation and the accuracy of predicting customer response behavior, enhanced the immediacy of training results and the accuracy of quantitative evaluation, and improved sales staff's emotional sensitivity and topic handling ability.
Smart Images

Figure CN120672533A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of sales training technology, and in particular to a sales training method and system based on GPT and a multimodal large model. Background Art
[0002] Sales coaching involves simulating real-life sales situations and using systematic, standardized processes to provide targeted training for salespeople in communication skills, demand generation, objection handling, proposal presentation, and deal closure. Specific methods include role-playing simulations, the use of standardized conversation scripts, scenario-based case studies, close-selling speech drills, and behavioral assessment scales. The goal is to systematically enhance salespeople's ability to respond and close effectively in the real sales process through quantifiable training metrics and continuous practice.
[0003] Sales coaching refers to a specific program designed and implemented to train salespeople, enhancing their response skills and closing capabilities in diverse customer scenarios. Its uses include strengthening sales pitch proficiency through multiple rounds of role-playing and scenario simulations, improving the accuracy of customer needs identification through standardized process training, and continuously optimizing sales performance through immediate feedback and skill scoring, ultimately achieving the goals of shortening sales cycles, increasing closing rates, and steadily improving the overall performance of the sales team.
[0004] Traditional coaching methods use fixed scripts and single scenarios for role-playing and case analysis. The training process is prone to modeling and homogenization, resulting in sales personnel showing obvious inadaptability and insufficient adaptability when facing changes in actual customer situations. This is manifested in the single expression of demand exploration questions, mechanical way of handling objections, and lack of flexibility in closing sales scripts, resulting in a decrease in the success rate of customer interaction in real sales scenarios. Traditional methods of standardized dialogue scripts and scenario simulations are usually based on static scenario presets, and lack detailed responses to customer emotional changes and personality differences. As a result, sales training cannot fully reflect the real reactions of customers and subtle emotional fluctuations during the interaction process, affecting the sales staff's ability to recognize customer emotions and adjust scripts, making it difficult for training results to accurately reflect the sales staff's real response level and closing skills, restricting the improvement of sales training quality and long-term effect guarantee. Summary of the Invention
[0005] The purpose of the present invention is to solve the shortcomings of the prior art and propose a sales training method and system based on GPT and multimodal large models.
[0006] To achieve the above objectives, the present invention adopts the following technical solution: a sales training method based on GPT and a multimodal large model, comprising the following steps:
[0007] S1: Obtain sales training needs of bank sales staff, call customer basic attribute parameters, financial product function feature parameters, and market hot topic parameters, and use the GPT-4 language generation model to expand the complete conversation opening statement, demand discovery questions, and product introduction sentences based on the content to generate sales conversation scenario description content;
[0008] S2: Based on the description of the sales conversation scenario, call the multimodal Transformer structure encoder to encode the features into feature vectors, screen the interactive feature combinations according to the customer personality and emotion correlation standard, and establish a multimodal training input set;
[0009] S3: Call the multimodal training input set, input it into the GPT-4 language model according to the sales training requirements, specify the role attributes, analyze the emotion intensity change parameters and the frequency of conversation topic switching, mark the emotion turning points and topic transfer nodes in the multi-round conversation generation process, and construct the virtual sales training role data;
[0010] S4: Based on the virtual sales sparring role data, call the speech flow generated by GPT-4 to train sales personnel to input sales speech, monitor the input voice stream and text stream in real time, record the customer's emotional state and the accuracy of the financial advisor's response to the speech in each round of interaction, and generate a sales sparring interaction record set.
[0011] The present invention has improvements in that the sales dialogue scenario description content includes customer attribute description, product feature description, sales topic design, standardized opening statement, demand mining question set and product introduction statement set; the multimodal sparring input set includes a voice feature coding set, a text feature vector set, an expression feature coding set and an interaction feature matching result set; the virtual sales sparring role data includes customer basic information settings, financial management interest preference settings, risk tolerance level, emotion change label set and topic transfer node set; the sales sparring interaction record set includes voice input content record, text input content record, emotion classification tag record, speech matching score record and interaction round process record.
[0012] The present invention is improved in that the steps of obtaining the sales dialogue scene description content are specifically as follows:
[0013] S111: Obtain customer types, product categories, and sales topic items from the sales training needs of bank sales personnel, call customer basic information parameters, financial product function feature parameters, and market hot topic parameters, cross-screen the sales topic items and market hot topic parameters, extract corresponding macroeconomic indicators and policy change items, and establish customer product topic combination items;
[0014] S112: Based on the customer-product topic combination items, the customer portrait feature items, the product function feature items, and the market hot topic items are arranged in order, the contents are sequentially connected, and a unified format is arranged according to the text paragraphing rules to generate a standardized description segment, thereby obtaining the customer portrait and product element topic description segment;
[0015] S113: Based on the customer portrait and product element topic description segments, they are input into the GPT-4 language generation model, with sales opening remarks, demand mining questions and product introduction sentences as prompt items, and the generated text output content is called to extract text paragraphs that completely cover customer feature questions, product element introductions and product introduction content, and arrange them in the order of the standard sales dialogue process to generate sales dialogue scene description content.
[0016] The present invention is improved in that the step of obtaining the multimodal training input set is specifically as follows:
[0017] S211: Obtain the sales dialogue scenario description content, collect the fundamental frequency, energy, and speech rate data from the voice feature parameters, the rising and falling tone turning point data from the intonation feature parameters, the word frequency TF-IDF value and sentiment polarity score from the text feature parameters, and the FACS encoding data from the image expression feature parameters, and perform feature normalization processing uniformly according to the data structure format to generate multimodal feature normalized data;
[0018] S212: Based on the multimodal feature normalized data, calling a multimodal Transformer structure encoder to perform feature vector encoding on the voice, intonation, text, and image features in the normalized data, obtaining feature vector groups after encoding, and calculating the interactive behavior matching degree between the customer personality feature vector and the emotion feature vector to obtain the interactive behavior matching coefficient;
[0019] S213: Based on the interaction behavior matching coefficient, screen the feature combinations whose matching coefficients are greater than the set matching threshold, combine the corresponding speech feature subsets, intonation feature subsets, text feature subsets and image expression feature subsets, integrate them into a unified input sample format, and obtain a multimodal training input set.
[0020] The present invention is improved in that the step of acquiring the virtual sales sparring role data is specifically as follows:
[0021] S311: Calling the multimodal training input set, setting a role parameter set according to the sales training requirements, collecting name setting items, total assets setting items, investment knowledge level setting items, and risk tolerance level setting items, combining the setting items into a role feature vector based on customer group standards, unifying the value of each parameter, and standardizing the feature scale to generate a role feature normalized vector;
[0022] S312: Based on the normalized character feature vector, the vector is input into the GPT-4 language model, and the normalized parameter content is used to generate the character setting text, extract the emotion intensity change parameter and the conversation topic switching frequency, call the emotion intensity change curve and the topic change node marking content, sort the nodes, mark the emotion turning points and topic transfer nodes, and obtain the character emotion and topic marking data;
[0023] S313: Based on the character emotion and topic annotation data, filter the emotion fluctuation nodes and topic connection nodes that meet the natural language interaction process standards, combine the emotion corpus and topic corpus corresponding to the sales training scenario, integrate the text, emotion labels and topic jump relationships according to the character feature vector requirements, and establish virtual sales training character data.
[0024] The present invention is improved in that the steps of obtaining the sales practice interaction record set are specifically as follows:
[0025] S411: Call the virtual sales training role data, configure the sales training platform, input the speech flow generated by GPT-4, set the customer role attributes according to the sales training requirements, set the sales staff training goals and response range, establish the basic sales training scenario, and generate sales training scenario configuration data;
[0026] S412: Based on the sales practice scenario configuration data, the voice and text streams input by the salesperson are monitored in real time, the fundamental frequency changes and energy fluctuations in the voice stream are monitored, key semantic features of the input speech in the text stream are extracted, the speech emotion category and the text speech content are detected, an interaction deviation value is calculated, and valid interaction records with interaction deviation values less than an interaction stability threshold are screened to obtain sales practice behavior evaluation data;
[0027] S413: Based on the sales sparring behavior evaluation data, record each round of data that meets the interaction deviation less than the interaction stability threshold, extract the salesperson's input text, voice emotion category, emotion deviation and text matching score, mark the sparring customer feedback actions and emotional state, combine and archive them into a standard training record format, and establish a sales sparring interaction record set.
[0028] The present invention is improved in that the method further comprises the following steps:
[0029] S5: Calling the sales sparring interaction record set, extracting the frequency of product introduction speech, the correct response rate for objection handling, and the proportion of emotionally positive responses, normalizing each indicator, assigning a corresponding weight to each indicator, calculating the training score, and generating sales sparring training score information;
[0030] The sales coaching training score information includes a product introduction frequency score, an objection handling accuracy score, an emotional positive response ratio score, and a comprehensive weighted training score.
[0031] The present invention is improved in that the step of obtaining the sales sparring training score information is specifically as follows:
[0032] S511: Calling the sales sparring interaction record set, extracting product introduction speech frequency data, objection handling correct response ratio data, and emotionally positive response ratio data, and standardizing each data item to establish a sales sparring indicator standardized data set;
[0033] S512: Based on the standardized sales training indicator dataset, weights for product introduction speech frequency, objection handling response ratio, and emotionally positive response ratio are set according to the indicator importance, and corresponding weight values are assigned to each. Calculation is performed to obtain a sales training score, and training samples with scores greater than a training threshold are screened to obtain a sales training score calculation result.
[0034] S513: Based on the sales coaching score calculation results, filter the score samples that meet the training standards, extract the training score and indicator details of each sample, integrate the frequency of product introduction scripts, the correct response rate for objection handling, and the proportion of positive emotional responses, and establish sales coaching training score information.
[0035] A sales training system based on GPT and a large multimodal model, wherein the sales training system based on GPT and a large multimodal model is used to implement the above-mentioned sales training method based on GPT and a large multimodal model, and the system comprises:
[0036] The conversation scenario analysis module obtains the sales training needs of bank sales staff, calls on basic customer attribute parameters, financial product function feature parameters, and market hot topic parameters, and uses the GPT-4 language generation model to expand the complete conversation opening statement, demand discovery questions, and product introduction sentences based on the content to generate sales conversation scenario description content;
[0037] The training input analysis module uses a multimodal Transformer structure encoder to encode the features into feature vectors based on the description of the sales conversation scenario, and selects the interaction feature combination according to the customer personality and emotion correlation standard to establish a multimodal training input set;
[0038] The virtual character construction module calls the multimodal training input set, inputs it into the GPT-4 language model according to the sales training requirements, specifies the character attributes, analyzes the emotion intensity change parameters and the frequency of conversation topic switching, marks the emotion turning points and topic transfer nodes in the multi-round conversation generation process, and constructs the virtual sales training character data;
[0039] The training recording module uses the data of the virtual sales sparring role to call the speech flow generated by GPT-4 to train the salesperson to input sales speech, monitor the input voice stream and text stream in real time, record the customer's emotional state and the accuracy of the financial advisor's response in each round of interaction, and generate a sales training interaction record set;
[0040] The training effect analysis module calls the sales sparring interaction record set, extracts the frequency of product introduction scripts, the correct response rate for objection handling, and the proportion of positive emotional responses, standardizes each indicator, assigns a corresponding weight to each indicator, calculates the training score, and generates sales sparring training score information.
[0041] Compared with the prior art, the advantages and positive effects of the present invention are:
[0042] In the present invention, by calling customer attribute parameters, financial product function parameters and market hot topic parameters, customer portraits, product elements and sales topics are automatically combined in a scene splicing manner, and the language generation model is used to realize the automatic generation of sales dialogue scenes, which significantly improves the degree of content personalization and scene matching accuracy, allowing sales personnel to be exposed to more diverse situations that are close to real customer needs, significantly improving the coverage and flexibility of sales training, and based on customer emotions and personality parameters, through multimodal feature coding and interactive behavior matching analysis, automatically screening out highly correlated interactive feature combinations to accurately evaluate the emotional fluctuations of customer feedback and sales The correlation between human interaction actions improves the realism of scenario simulation and the accuracy of predicting customer response behavior. During the sales talk rehearsal, it dynamically monitors voice emotions and text words, automatically identifies and marks customer emotional transitions and topic transfer nodes, accurately tracks the real-time performance of sales personnel in responding to customer emotional changes, and significantly improves the effectiveness of emotional sensitivity and topic handling ability training. Through standardized indicator evaluation of product introduction frequency, objection response accuracy and positive emotion ratio, it dynamically provides training score results, making sales skills evaluation results more objective and accurate, and enhancing the immediacy of training result feedback and the accuracy of quantitative evaluation. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] Figure 1 is a flow chart of the method of the present invention;
[0044] Figure 2 A flowchart for obtaining a sales conversation scenario description for the present invention;
[0045] Figure 3 A flowchart of obtaining a multimodal training input set for the present invention;
[0046] Figure 4 A flowchart of the present invention for obtaining data of a virtual sales sparring role;
[0047] Figure 5A flowchart for obtaining a sales sparring interaction record set for the present invention;
[0048] Figure 6 This is a flow chart of the present invention for obtaining sales sparring training score information. DETAILED DESCRIPTION
[0049] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0050] In the description of the present invention, it should be understood that the terms "length", "width", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside", etc., indicating directions or positional relationships, are based on the directions or positional relationships shown in the accompanying drawings and are only for the convenience of describing the present invention and simplifying the description. They do not indicate or imply that the devices or elements referred to must have a specific direction, be constructed and operate in a specific direction, and therefore should not be understood as limiting the present invention. In addition, in the description of the present invention, the meaning of "plurality" is two or more, unless otherwise clearly and specifically defined.
[0051] See also Figure 1 The present invention provides a technical solution: a sales training method based on GPT and a multimodal large model, comprising the following steps:
[0052] S1: Obtain sales training requirements for bank sales staff, which include customer type, product category, and sales topic items. The training requires basic customer attribute parameters, financial product function feature parameters, and market hot topic parameters. The customer portrait description, product element description, and sales topic description are arranged in sequence using scene splicing actions. The GPT-4 language generation model is used to expand the complete conversation opening statement, demand discovery questions, and product introduction sentences based on the content to generate a sales conversation scenario description.
[0053] Basic customer attribute parameters include name, gender, age, occupation, and asset size; financial product functional characteristic parameters include return type, risk level, and investment period; market hot topic parameters include changes in macroeconomic indicators and policy release events;
[0054] S2: Based on the description of the sales conversation scenario, obtain voice feature parameters, intonation feature parameters, text feature parameters, and image expression feature parameters. Call the multimodal Transformer structure encoder to encode the features into feature vectors. Based on the correlation between customer personality and emotion, use the Pearson correlation coefficient to calculate the interaction behavior matching coefficient. Filter the interaction feature combinations with matching degrees above the matching threshold to establish a multimodal training input set.
[0055] The speech feature parameter is speaking rate, the intonation feature parameter is the fundamental frequency change rate, the text feature parameter is the TF-IDF keyword weight distribution, and the image expression feature parameter is specifically the expression recognition FACS code, such as AU12 smile, AU4 frown, and AU45 blink. Speaking rate refers to the number of speech words uttered per unit time and can be extracted through automatic speech recognition technology. The fundamental frequency change rate refers to the rate of change of the basic frequency of the tone per unit time, reflecting emotional fluctuations. The TF-IDF keyword weight is a standard parameter in the text field, used to measure the importance of a word to a specific document in the corpus;
[0056] S3: Call the multimodal training input set and input it into the GPT-4 language model based on the sales training requirements. Specify the role attributes, including name, total assets, investment knowledge level, and risk tolerance level. Comprehensively analyze the emotion intensity change parameters and the frequency of conversation topic switching. During the multi-round conversation generation process, mark the emotional turning points and topic transfer nodes, and construct virtual sales training role data with natural language interaction and emotional feedback capabilities.
[0057] S4: Based on the data of the virtual sales training role, a sales training platform is configured. The sales staff is trained to input sales scripts by calling the script flow generated by GPT-4. The platform monitors the input voice and text streams in real time, detects the voice emotion category, determines the voice emotion deviation, analyzes the match between the input content and the standard script, records the customer's emotional state and the accuracy of the financial advisor's response in each round of interaction, and generates a sales training interaction record set.
[0058] S5: Call the sales sparring interaction record set to extract the frequency of product introduction dialogue, the correct response rate for objection handling, and the proportion of emotionally positive responses. Each indicator is normalized, assigned a corresponding weight, and the training score is calculated to generate sales sparring training score information.
[0059] The correct response ratio for objection handling is the number of correct responses / total number of objections; the positive emotion response ratio is the number of positive emotions / total number of interactions;
[0060] The description of the sales dialogue scenario includes customer attribute description, product feature description, sales topic design, standardized opening statements, demand mining question set and product introduction statement set. The multimodal training input set includes voice feature coding set, text feature vector set, expression feature coding set and interaction feature matching result set. The virtual sales training role data includes customer basic information settings, financial management interest preference settings, risk tolerance level, emotion change label set and topic transfer node set. The sales training interaction record set includes voice input content records, text input content records, emotion classification label records, speech matching score records and interaction round process records. The sales training training score information includes product introduction frequency score, objection handling accuracy score, emotion positive response ratio score and comprehensive weighted training score.
[0061] See also Figure 2 The specific steps for obtaining the sales dialogue scenario description content are as follows:
[0062] S111: Obtain customer types, product categories, and sales topic items from the sales training needs of bank sales personnel, call customer basic information parameters, financial product functional characteristic parameters, and market hot topic parameters, match customer types with customer basic information parameters, select age, occupation, and asset size characteristic items that meet customer profile standards, compare product categories with financial product functional characteristic parameters, select functional characteristic items that meet investment period, return method, and risk level standards, cross-screen sales topic items with market hot topic parameters, extract corresponding macroeconomic indicators and policy change items, and establish customer product topic combination items;
[0063] Obtain the customer type, product category and sales topic items in the sales training needs of bank sales personnel. First, make a preliminary classification based on the customer type field and divide the customers into three levels: ordinary customers, high-quality customers and high-net-worth customers. When screening, call the age, occupation and asset size of the customer basic information parameter set. By setting the screening criteria, the age of ordinary customers is set to be between 25 and 40 years old, the age of high-quality customers is between 40 and 55 years old, and the age of high-net-worth customers is set to be over 55 years old. The asset size is set to be less than 500,000 yuan, between 500,000 yuan and 5 million yuan, and more than 5 million yuan respectively. The occupation is divided into ordinary employees, middle-level managers and senior managers according to the job level. Collect actual sample data, see Table 1, for the customers in the table The basic information of each customer is collected and the customer type that meets the classification criteria is obtained through screening. The age range is set according to the consumption capacity classification standard of the population life cycle issued by the National Bureau of Statistics. Among them, the 25-40-year-old group is in the early stage of career development, with relatively active consumption capacity and initial financial management awareness. The 40-55-year-old group enters the peak of wealth accumulation, and the group over 55 enters the retirement stage with a relatively concentrated wealth scale. Therefore, this interval division is set. The asset scale interval is set according to the customer stratification data disclosed in the bank's annual report. The median asset scale of ordinary customers is 300,000 yuan, the median asset scale of high-quality customers is 1.5 million yuan, and the median asset scale of high-net-worth customers is 6 million yuan. Therefore, the upper and lower intervals are divided into 500,000 yuan and 5 million yuan nodes, as shown in Table 1. In the data shown, customer A is 35 years old, has an asset size of 300,000 yuan, and is an ordinary employee, which means he is an ordinary customer. Then, for the product category field, the financial product function characteristic parameters are called to extract the three characteristic items of investment period, income method and risk level. The short-term investment period is set to 0-1 years, the medium-term is 1-3 years, and the long-term is more than 3 years. The investment period interval is divided according to the financial product term management regulations issued by the State Financial Supervision and Administration Bureau. Short-term products are controlled within 1 year, the medium-term is the standard financial management cycle, and the long-term corresponds to the large asset allocation needs. The income method is divided into fixed income and floating income. The risk level is divided into low, medium and high according to the bank's internal rating. The risk level standard is set according to the financial industry standard. The low risk standard The annualized return of the quasi-volatility range is less than 2%, the annualized return of the medium-risk range is between 2% and 5%, and the high-risk range is greater than 5%. For example, a financial product B with a 2-year investment period, fixed income, and medium risk is matched as a medium-term fixed-income medium-risk product. Then, the market hot topic parameters are called for the sales topic item to filter topic nodes closely related to current macroeconomic indicators and policy changes. The filtering conditions are set as changes in GDP growth rate, benchmark interest rate adjustments, and real estate policy changes in the past three months. The screening period is set to three months. Based on the frequency of economic data released by the National Bureau of Statistics, the data is updated once a quarter, so a three-month window period is determined to collect data on changes in economic indicators. For example, the recent GDP growth rate was lowered by 0.3 percentage points, the benchmark interest rate was lowered by 5 basis points, and the related topic item extracted was "Increased demand for asset preservation." Finally, based on the matching content extracted from the above customer profile, product elements, and market topics, a customer-product topic combination item was established.
[0064] S112: Based on the customer-product topic combination items, the customer profile feature items, product function feature items, and market hot topic items are arranged in order, and the name, occupation, and asset size parameters in the customer basic information, the income method and risk level parameters in the financial product function feature items, and the economic indicator changes in the market hot topic parameters are called. The contents are sequentially connected and organized in a unified format according to the text paragraphing rules to generate a standardized description segment, thereby obtaining the customer profile and product element topic description segment;
[0065] Based on the customer product topic combination item, the customer portrait feature item, product function feature item and market hot topic item are arranged in order, and the name, occupation and asset size parameters in the customer basic information, the income method and risk level parameters in the financial product function feature are called. Combined with the changes in economic indicators in the market hot topic parameters, the content is connected in sequence. First, based on the customer portrait feature item, the customer name "Li", occupation "middle-level manager", and asset size of 1.2 million yuan are extracted. Then the product function feature item is extracted, and the income method is set to "fixed income" and the risk level is set to "medium risk". The market hot topic parameters are collected. Click on the topic content "Increasing demand for asset preservation" and connect and organize the above information in order to standardize it as: "Customer Li, currently a middle-level manager with an asset size of 1.2 million yuan, is interested in fixed-income medium-risk financial products. Affected by the recent increase in demand for asset preservation, he tends to choose stable investment products." In the process of concatenating the content, in order to avoid field omissions, the content concatenation completeness detection standard is adopted to ensure that the six elements of name, occupation, asset size, income method, risk level, and topic item are complete. The concatenation completeness detection standard is set to require that all six elements appear in the content and the coverage rate reaches 100%. According to the ISO 9001 document integrity inspection requirements, the lack of any element is considered an unqualified record. If any item is missing, the system prompts to re-collect parameters to complete the gap. Verification is performed on the test sample, and the concatenation content completeness rate reaches 100%. Finally, a unified text paragraph is generated to obtain the customer portrait and product element topic description segment.
[0066] S113: Based on the customer profile and product element topic description segments, the GPT-4 language generation model is input. Using the sales opening statement, demand exploration questions, and product introduction sentences as prompts, the model generates text output content. The model extracts text segments that fully cover customer characteristic questions, product element introductions, and product introductions. These segments are arranged in the order of the standard sales dialogue process to generate sales dialogue scenario description content.
[0067] Based on the customer portrait and product element topic description segments, they are input into the GPT-4 language generation model, and the sales opening remarks, demand mining questions and product introduction sentences are set as prompt items. The specific setting examples are: the opening remarks prompt item is set to "Do you have any financial management needs in the near future?", the demand mining question is set to "Are you interested in products with stable returns?", and the product introduction sentence is set to "I can recommend a stable return product portfolio based on your asset size." By inputting the prompt items, the output GPT-4 generates standardized speech output, extracts the paragraphs in the generated text that completely cover customer feature questions, product element introductions and product introduction content, and detects whether the generated text covers the above three types of content. If the coverage rate is less than 100%, it is readjusted. After adjusting the prompt items, the text is regenerated. In the actual operation process, five groups of different customer portrait input examples are selected to test the generation coverage. Four groups are fully covered, and one group is insufficiently covered. After readjusting the demand mining questions, the coverage rate is generated again to reach 100%. The text content that meets the standards is sorted out and arranged according to the standard process sequence of sales dialogue, that is, the opening remarks-demand mining-product introduction process. Finally, standardized sales dialogue paragraphs are formed, and the sales dialogue scene description content is generated. The prompt item lengths of the opening remarks, demand mining, and product introduction stages are set to 15-25 words respectively. According to the user attention maintenance theory, 15-25 words are the optimal output information length range, ensuring that the language prompts cover key information while avoiding redundancy.
[0068] See also Figure 3 ,The specific steps for obtaining the multimodal training input set are:
[0069] S211: Obtain the sales conversation scenario description content, collect the fundamental frequency, energy, and speech rate data from the voice feature parameters, the rising and falling tone turning point data from the intonation feature parameters, the word frequency TF-IDF value and sentiment polarity score from the text feature parameters, and the FACS encoding data from the image expression feature parameters, and perform feature normalization processing on them in a unified manner according to the data structure format to generate multimodal feature normalized data;
[0070] The description of the sales conversation scene is obtained, and the fundamental frequency, energy, and speaking rate data of the voice feature parameters are collected. The fundamental frequency is extracted by analyzing the fundamental frequency change curve of the audio signal per second, and the fundamental frequency extreme value is extracted and averaged to obtain the fundamental frequency parameter. The sampling frequency is set to 16kHz. In a 5-second segment of the voice signal, 16,000 points are sampled per second. The maximum and minimum fundamental frequencies are extracted and averaged to obtain a fundamental frequency average of 230Hz. The energy is normalized by summing the square of the amplitude of each frame of the audio signal. The average energy of 5 seconds of sampling is 0.75. The speaking rate is measured by the number of phonemes recognized per second. In the standard scenario, an average of 13.5 phonemes per second are recognized. The rising and falling turning points in the intonation feature parameters are obtained by judging the fundamental frequency change trend of consecutive frames. When the fundamental frequency rises for 5 consecutive frames and the increase exceeds 30Hz, it is recorded as a rising turning point. When it falls for 5 consecutive frames and the decrease exceeds 30Hz, it is recorded as a falling turning point. In the collection of text feature parameters, the word frequency uses the TF-IDF algorithm to count the frequency of keyword occurrences, and the total vocabulary of the sample is set for the sample text. The number of words is 500, the frequency of the keyword "investment" is 0.008, the sentiment polarity score is scored using the sentiment dictionary matching rule, and the corresponding text segment score is 0.65. In the image expression feature parameter collection, FACS encoding extracts the strength of the three key action units AU12, AU6, and AU1 through the facial expression action unit detection module, which are standardized to 0.7, 0.5, and 0.4 respectively. In order to ensure that the feature data can be processed uniformly, all features are uniformly normalized, and the normalization method adopts the maximum and minimum The parameters were normalized to the range of 0 to 1 using the minor normalization method. The fundamental frequency normalization value was 0.58, the energy normalization value was 0.75, the speech rate normalization value was 0.67, the rising tone normalization value was 0.6, the falling tone normalization value was 0.62, the TF-IDF normalization value was 0.66, the emotion polarity normalization value was 0.65, and the FACS normalization values remained unchanged, as shown in Table 1. The table lists the specific values of each feature after normalization, see Table 1, and finally generated multimodal feature normalized data.
[0071] Table 1 Multimodal feature normalization values
[0072] Feature Type Feature Item Normalized values Voice Features fundamental frequency 0.58 Voice Features energy 0.75 Voice Features speaking speed 0.67 Intonation features rising turning point 0.60 Intonation features Downward turning point 0.62 Text features TF-IDF value 0.66 Text features Sentiment polarity score 0.65 Image features FACS code AU12 0.70 Image features FACS code AU6 0.50 Image features FACS code AU1 0.40
[0073] As shown in Table 1, all multimodal features are normalized.
[0074] S212: Based on the multimodal feature normalized data, the multimodal Transformer structure encoder is called to perform feature vector encoding on the speech, intonation, text, and image features in the normalized data. After encoding, feature vector groups are obtained respectively using the formula:
[0075]
[0076] Calculate the interactive behavior matching degree between the customer's personality feature vector and the emotional feature vector to obtain the interactive behavior matching coefficient;
[0077] Among them, S m represents the interaction behavior matching coefficient, p i represents the i-th normalized speech fundamental frequency eigenvalue, z i represents the i-th normalized speech energy eigenvalue, v i represents the sentiment polarity score of the i-th normalized text, t i represents the i-th normalized image expression FACS feature encoding value, n represents the total number of feature samples, and i is the feature index;
[0078] Based on the multimodal feature normalized data, the multimodal Transformer structure encoder is called to encode the voice, intonation, text, and image features in the normalized data into feature vectors. After encoding, feature vector groups are obtained respectively. In the process of calculating the interactive behavior matching degree, the formula is used:
[0079]
[0080] Among them, S m represents the interaction behavior matching coefficient, p i represents the i-th normalized speech fundamental frequency eigenvalue, z i represents the i-th normalized speech energy eigenvalue, v i represents the sentiment polarity score of the i-th normalized text, t i Represents the i-th normalized image expression FACS feature encoding value, n represents the total number of feature samples, i is the feature index number, and the summation symbol Σ represents the accumulation of all sample features. The formula parameter assignment is as follows. Suppose n = 4, that is, four groups of data are collected, the fundamental frequency features p are 0.58, 0.62, 0.60, and 0.59, the energy features z are 0.75, 0.73, 0.78, and 0.76, the text sentiment polarity scores v are 0.65, 0.67, 0.66, and 0.68, and the image expression features t are 0.70, 0.72, 0.69, and 0.71, respectively. Substituting into the formula, the calculation process is as follows:
[0081] First calculate:
[0082]
[0083] Then calculate the sum of squares of the denominator:
[0084]
[0085] Substituting the result into the overall formula, we get:
[0086]
[0087] Therefore, the final calculation results in the interaction behavior matching coefficient S m It is 1.347, which indicates the overall matching performance under the current feature combination;
[0088] S213: Based on the interaction behavior matching coefficient, screen feature combinations whose matching coefficients are greater than a set matching threshold, combine corresponding speech feature subsets, intonation feature subsets, text feature subsets, and image expression feature subsets, and integrate them into a unified input sample format to obtain a multimodal training input set;
[0089] According to the interaction behavior matching coefficient, the feature combination with a matching coefficient greater than the set matching threshold is screened. The matching threshold is set to 1.200. The threshold is based on the mean of 1.150 of the previous 50 groups of actual interaction sample data, combined with the standard deviation of 0.10, and set up by 0.5σ to ensure the screening of effective interaction samples. When the interaction behavior matching coefficient S m When it is greater than 1.200, the corresponding feature subset is extracted. In the actual case, the calculation result of 1.347 is higher than the threshold of 1.200, and the screening is effective. The corresponding speech feature subset (fundamental frequency, energy normalized vector), intonation feature subset (rising tone, falling tone turning point normalized vector), text feature subset (word frequency TF-IDF, emotion polarity normalized vector) and image expression feature subset (FACS encoding normalized vector) are combined and arranged according to a unified input sample format. The format is set to vector splicing and labeling, where the label is set according to the customer ID + interaction number rule, and finally the multimodal training input set is obtained.
[0090] See also Figure 4 The specific steps for obtaining the virtual sales sparring role data are as follows:
[0091] S311: Calling the multimodal training input set, setting the role parameter set according to the sales training requirements, collecting the name setting item, the total assets setting item, the investment knowledge level setting item, and the risk tolerance level setting item, combining the setting items into a role feature vector based on the customer group standard, unifying the value of each parameter, and standardizing the feature scale to generate a role feature normalized vector;
[0092] Call the multimodal training input set, set the role parameter set according to the sales training needs, collect name setting items, total asset setting items, investment knowledge level setting items and risk tolerance level setting items. The name setting item is collected based on user input and stored uniformly in string encoding. The total asset setting item uses the amount interval division standard and is set to [0-500,000 yuan], [500,000-3 million yuan], [3 million-10 million yuan], and [10 million yuan or more]. The asset size is expressed in median, which is set to 250,000 yuan, 1.75 million yuan, 6.5 million yuan and 15 million yuan respectively. The total asset setting interval is set according to the customer stratification standards for different wealth levels in the survey data of the "China Family Wealth Report 2023" to ensure that most potential customer portraits are covered. The investment knowledge level setting item is standardized to 1-5 points through the self-assessment questionnaire, representing no knowledge, understanding of basic concepts, etc. , have some experience, relatively rich experience and professional knowledge. The risk tolerance level setting items are set according to international risk assessment standards and are divided into levels 1 to 5. Level 1 is extremely low risk tolerance and level 5 is extremely high risk tolerance. The setting standard is that level 1 corresponds to only accepting money fund products, and level 5 accepts stock and derivative investments. Through the above settings, they are combined into a role feature vector, and each parameter is unified as numerical data. Among them, the name is converted into a coding index, the total asset setting item is represented by the asset size value, and the investment knowledge level and risk tolerance both maintain the score value. For each value of the role feature vector, standardized feature scale processing is performed, and the maximum and minimum normalization method is used to uniformly compress it to the [0,1] interval. The standardized benchmark is taken from the maximum and minimum values in the training set. For example, if the minimum value of the asset size sample is 250,000 yuan and the maximum value is 15 million yuan, and a customer has 6.5 million yuan in assets, then the normalized value calculation process is:
[0093] (650-25) / (1500-25)=0.421, and the character feature normalized vector is finally generated.
[0094] S312: Based on the normalized character feature vector, input it into the GPT-4 language model, use the normalized parameter content to generate character setting text, extract the emotion intensity change parameter and the conversation topic switching frequency, call the emotion intensity change curve and topic change node marking content, sort the nodes, mark the emotion turning points and topic transfer nodes, and obtain the character emotion and topic annotation data;
[0095] Based on the normalized vector of role features, it is input into the GPT-4 language model, and the normalized parameter content is used to generate the role setting text. The text structure is generated in the order of name-total assets-investment knowledge level-risk tolerance. For example, the name code 0021 corresponds to the name "Wang Mou", the normalized value of total assets 0.421 corresponds to the original assets of 6.5 million yuan, the investment knowledge level is 4 points (relatively rich experience), and the risk tolerance is 3 points (medium risk acceptance). The generated role setting text is: "Wang Mou, assets of 6.5 million yuan, relatively rich investment knowledge level, medium risk tolerance". After generation, the emotion intensity change curve and the dialogue topic change node mark content are called. The emotion intensity change curve is set by setting Normalize the emotional intensity value, extract the emotional fluctuation value after each round of dialogue, set the emotional intensity change node threshold to 0.15, the threshold is obtained based on the standard deviation analysis in the natural language interaction emotional response experiment, and is set by 50% on the basis of the emotional standard deviation of 0.10. If the amplitude of the emotional change in a certain round exceeds 0.15, it is marked as an emotional turning point. The frequency of dialogue topic switching is calculated by counting the number of topic changes in each round. It is set that if the topic ID appears different in more than two consecutive rounds, it is switched once. The topic switching frequency is greater than 1 time / 5 rounds, it is considered a high-frequency switching node. The nodes are sorted and numbered in the order of occurrence time. Finally, the emotional turning points and topic transfer nodes are marked in sequence according to the node sequence number to obtain the character emotion and topic labeling data.
[0096] S313: Based on the character emotion and topic labeling data, select emotion fluctuation nodes and topic connection nodes that meet the natural language interaction process standards, combine the emotion corpus and topic corpus corresponding to the sales training scenario, integrate the text, emotion labels and topic jump relationships according to the character feature vector requirements, and establish virtual sales training character data;
[0097] According to the character emotion and topic annotation data, the emotion fluctuation nodes and topic connection nodes that meet the natural language interaction process standards are screened. The natural language interaction process standard is set as the ratio of emotion fluctuation nodes to the total number of nodes is between 10%-30%, and the ratio of topic connection nodes to the total number of nodes is between 20%-40%. The emotion fluctuation node screening standard is determined according to the interaction stability requirements in the interaction experiment. If it is too low, it is not enough to reflect the naturalness of the interaction, and if it is too high, it will lead to incoherent interaction. Therefore, the ratio interval is set to 10%-30%. In the screening step, the number of all nodes is first counted, and the ratio of nodes that meet the emotion fluctuation conditions and the topic connection conditions is calculated. If the ratio is within the interval, it passes the screening Otherwise, the node extraction process is retraced to adjust the parameters. The nodes selected are combined to correspond to the speech sentiment corpus and topic text corpus in the sales training scenario. The speech sentiment corpus is organized according to the sentiment classification standard and divided into three categories: positive motivation, interrogative language, and explanatory sentences. The topic corpus is divided into three categories: product introduction, customer care, and investment advice based on the content theme classification. According to the role feature vector requirements, emotion labels and topic jump relationship markers are added to the text paragraphs. The label setting format is a combination of emotion category-topic category-round number, such as "positive motivation-customer care-round 3". Finally, the text content, emotion labels and topic jump relationships are unified and integrated to establish the virtual sales training role data.
[0098] See also Figure 5 The specific steps for obtaining the sales practice interaction record set are as follows:
[0099] S411: Call the virtual sales training role data, configure the sales training platform, input the GPT-4 generated speech flow, set the customer role attributes according to the sales training requirements, set the sales staff training goals and response range, establish the basic sales training scenario, and generate the sales training scenario configuration data;
[0100] Call the data of the virtual sales training role, configure the sales training platform, and input the speech flow generated by GPT-4. First, set the customer role attributes according to the sales training needs. The customer role attribute setting involves name, asset size, investment experience level and risk preference level. The asset size is divided into four levels according to the "China Wealth Management Annual Report", which are 250,000 yuan, 1.75 million yuan, 6.5 million yuan and 15 million yuan as the typical range median. The investment experience level is set by a five-level classification standard, 1 point for no experience, 5 points for professional investment knowledge, and the risk preference level is also divided into five levels, 1 for extremely conservative, and 5 for aggressive investors. In actual configuration, taking the sample customer Wang as an example, the asset size is 6.5 million yuan, the investment experience level is 4 points, and the risk preference is 3 points. The name is uniformly encoded, and the asset size, investment experience and risk level are normalized. The normalization method adopts maximum and minimum standardization. The normalized benchmark values are set as the maximum asset of 15 million yuan and the minimum asset of 250,000 yuan in the training sample. For example, the asset size of 6.5 million yuan is normalized to (650-25) / (1500-25)=0.421, the investment experience is normalized to (4-1) / (5-1)=0.75, and the risk preference is normalized to (3-1) / (5-1)=0.5. The normalized results are combined into a role feature vector [0.421,0.75,0.5]. Then, sales staff practice goals are set. The goals involve response speed, speech completeness, and emotional coping level. The response speed goal is set to an average response time of no more than 3 seconds. The speech completeness requirement covers the four modules of opening remarks, demand exploration, product introduction, and product introduction. The emotional coping level is based on a standard emotion recognition accuracy rate of more than 85%. Basic sales practice scenarios are established. By mapping the role feature vectors with the sales practice goals, sales practice scenario configuration data is generated.
[0101] S412: Based on the sales practice scenario configuration data, the salesperson's voice and text streams are monitored in real time. The fundamental frequency changes and energy fluctuations in the voice stream are monitored. The key semantic features of the input speech in the text stream are extracted. The speech emotion category and text speech content are detected using the formula:
[0102]
[0103] Calculate the interaction deviation value, filter out valid interaction records where the interaction deviation value is less than the interaction stability threshold, and obtain sales sparring behavior evaluation data;
[0104] Among them, C s represents the interaction deviation value, E j Represents the normalized value of the emotional intensity of the j-th round of input speech, E′ j represents the normalized value of the emotional baseline of the standard speech in round j, S jRepresents the normalized matching score of the j-th round input text, S′ j represents the normalized expected score of the standard speech in round j, and N represents the number of interaction rounds;
[0105] Based on the configuration data of the sales practice scenario, the voice and text streams input by the salesperson are monitored in real time. The voice stream sampling frequency is set to 16kHz. The fundamental frequency change value and energy change value are collected per second during the monitoring process. The maximum and minimum fundamental frequency differences and energy fluctuation amplitudes are extracted by dividing each second into 100 frames. The key semantic features of the speech are extracted synchronously from the text stream. Keywords are annotated by combining word frequency statistics and sentiment dictionary recognition. The speech emotion category is detected based on the sentiment fundamental frequency model classification. The text speech content is classified based on the TF-IDF and topic classification standards. The calling formula is:
[0106]
[0107] Among them, C s is the interaction deviation value, E j The normalized value of the emotional intensity of the input speech in round j, E′ j is the standard emotion benchmark normalized value, S j S′ is the normalized value of the j-th round text speech score, j is the normalized value of the standard text score, N is the number of interaction rounds, and for example, assume N = 3 rounds, the input sentiment intensities are 0.65, 0.72, and 0.68, the standard sentiment benchmarks are 0.70, 0.75, and 0.67, the input text scores are 0.82, 0.78, and 0.85, and the standard text scores are 0.80, 0.80, and 0.83. Substituting into the formula:
[0108]
[0109] Interaction deviation value C s =0.1734, set the interaction stability threshold to 0.20, the threshold is set based on the mean deviation of the stable interaction samples in the previous interaction experiment of 0.18 and 10% higher, so C s If the value is less than the threshold, it is filtered as a valid interaction record and the sales practice behavior evaluation data is obtained.
[0110] S413: Based on the sales sparring behavior evaluation data, record each round of data that meets the interaction deviation less than the interaction stability threshold, extract the salesperson's input speech text, voice emotion category, emotion deviation, and text matching score, annotate the sparring customer's feedback actions and emotional state, combine and archive them in a standard training record format, and establish a sales sparring interaction record set;
[0111] Based on the sales practice behavior evaluation data, record each round of data that meets the interaction deviation less than the interaction stability threshold, extract the salesperson's input text, voice emotion category, emotion deviation and text matching score. The emotion deviation is calculated as the absolute value of the difference between the input emotion intensity and the standard emotion intensity. For example, in the first round, the input emotion is 0.65, the standard emotion is 0.70, and the deviation is 0.05. The text matching score is calculated as the overlap rate of the input and standard keywords. If the standard keywords are ["financial management", "stable", "income"], the input text covers ["financial management" , "revenue"], the matching score is 2 / 3 = 0.666. Emotional categories are divided into three categories: calm, positive, and negative based on the fundamental frequency change characteristics. The customer feedback actions of the training partner are marked, and the action classification is set as smile, nod, and frown. Emotional state annotation is divided according to the normalized emotion intensity. Above 0.6 is classified as positive, 0.4-0.6 is neutral, and below 0.4 is negative. The extracted text, emotion labels, and topic jump relationships are combined and archived according to the interaction round number. Finally, the unified format is output as a standard training record format to establish a sales training interaction record set.
[0112] See also Figure 6 The specific steps for obtaining sales training score information are as follows:
[0113] S511: Call the sales training interaction record set, extract the product introduction speech frequency data, objection handling correct response ratio data, and emotionally positive response ratio data, standardize each data item, and establish a standardized sales training indicator data set;
[0114] The sales practice interaction record set was called to extract the frequency data of product introduction words, the correct response rate data of objection handling and the positive emotion response rate data. The frequency of product introduction words was obtained by counting the number of product introduction words in every 100 rounds of dialogue. The correct response rate of objection handling was calculated as the ratio of the number of correct answers to objection scenarios to the total number of objection scenarios. The positive emotion response rate was calculated by the proportion of positive emotion samples in the marked emotion classification. Each data was standardized separately. The standardization method used the maximum and minimum normalization method to process it to the [0,1] interval. The maximum and minimum values were set according to the extreme values of each data in the actual training samples. Taking the frequency of product introduction scripts as an example, if the maximum value of the sample is 12 times / 100 rounds, the minimum value is 4 times / 100 rounds, and a certain sample is 8 times / 100 rounds, then the normalized value is (8-4) / (12-4)=0.5. If the maximum correct response rate for objection handling is 95% and the minimum value is 70%, a certain sample is 85%, and the normalized value is (85-70) / (95-70)=0.6. If the maximum positive emotional response rate is 90% and the minimum is 50%, a certain sample is 70%, and the normalized value is (70-50) / (90-50)=0.5. After normalization, the three items are merged into standardized sample records, and finally a standardized data set of sales training indicators is established.
[0115] S512: Based on the standardized data set of sales training indicators, set the weight of product introduction speech frequency, objection handling response ratio, and emotional positive response ratio according to the importance of the indicators, and assign corresponding weight values to each. The formula is:
[0116]
[0117] Calculate and obtain the sales training score, filter the training samples whose scores are greater than the training threshold, and obtain the sales training score calculation result;
[0118] Among them, T s Represents the sales training score, F a Represents the frequency of occurrence of the product introduction words in item a, represents the average frequency of product introduction words, R a Represents the correct response ratio value of the objection handling of Article a, represents the mean correct response ratio of objection handling, P a represents the proportion of positive emotional responses in the a-th article, w1 represents the frequency weight of product introduction speech, w2 represents the weight of objection handling response ratio, w3 represents the weight of positive emotional responses, and A represents the total number of extracted samples;
[0119] Based on the standardized data set of sales training indicators, weights are set according to the importance of the indicators. The weight of product introduction speech frequency is set to w1 = 0.4, the weight of objection handling response ratio is set to w2 = 0.35, and the weight of positive emotional response ratio is set to w3 = 0.25. The weight setting is based on the average score of the importance of financial advisors' feedback in the transaction stage in the questionnaire. That is, the importance score of product introduction is 8 points, the importance score of objection handling is 7 points, and the importance score of positive emotional feedback is 5 points. After normalization, the weights are set to 0.4, 0.35, and 0.25 respectively. The formula is used:
[0120]
[0121] Assume that the sample size A=3, and the sample data are: F=[0.5, 0.7, 0.6], R=
[0122] 0.6,0.8,0.7], P=[0.5,0.6,0.55], substitute into the formula and expand:
[0123] The first step is to calculate the brackets within each sample:
[0124] Sample 1:
[0125] Sample 2:
[0126] Sample 3:
[0127] Calculated separately:
[0128] Sample 1: 0.4 × 0.1 + 0.35 × 0.01 + 0.25 × 0.7071 = 0.04 + 0.0035 +
[0129] 0.1768=0.2203;
[0130] Sample 2: 0.4 × 0.1 + 0.35 × 0.01 + 0.25 × 0.7746 = 0.04 + 0.0035 +
[0131] 0.1936=0.2371;
[0132] Sample 3: 0.4 × 0 + 0.35 × 0 + 0.25 × 0.7416 = 0 + 0 + 0.1854 = 0.1854;
[0133] The second step is to square and sum:
[0134] 0.2203 2 +0.2371 2 +0.1854 2=0.0485+0.0562+0.0344=0.1391;
[0135] The third step is to substitute the formula:
[0136]
[0137] Therefore, the sales training score is T s =0.2153, set the training threshold to 0.20, which is set 5% higher than the mean score of the standard training sample of 0.19. Therefore, T s If the value is greater than the qualified threshold, it is filtered as a qualified training sample and the sales training score calculation result is obtained.
[0138] S513: Based on the sales training score calculation results, select the scoring samples that meet the training standards, extract the training score and indicator details of each sample, integrate the frequency of product introduction speech, the correct response rate for objection handling, and the proportion of positive emotional responses, and establish the sales training score information;
[0139] According to the calculation results of the sales training score, the scoring samples that meet the training standards are screened. The screening rule is that the score value is higher than the training standard threshold of 0.20. The training score and indicator details of each sample are extracted. The indicator details include the normalized value of the frequency of occurrence of product introduction words, the normalized value of the correct response ratio of objection handling, and the normalized value of the proportion of positive emotional responses. After combination, they are organized into standardized sample records. For example, the training score of sample 1 is 0.2153, which corresponds to the normalized value of 0.5 for product introduction, 0.6 for objection handling, and 0.5 for positive emotional responses. They are summarized into the sales training score information set, as shown in Table 2. The data is displayed in a standardized manner to facilitate subsequent evaluation and comparison, and finally the sales training score information is established.
[0140] Table 2 Sales sparring training score data table
[0141]
[0142] As shown in Table 2 , the sales coaching training scores and the corresponding standardized indicator data are listed to facilitate the comprehensive evaluation of the sample training level.
[0143] A sales training system based on GPT and a large multimodal model is used to implement the above-mentioned sales training method based on GPT and a large multimodal model. The system includes:
[0144] The conversation scenario analysis module obtains the sales training needs of bank sales staff, calls on basic customer attribute parameters, financial product function feature parameters, and market hot topic parameters, and uses the GPT-4 language generation model to expand the complete conversation opening statement, demand discovery questions, and product introduction sentences based on the content to generate sales conversation scenario description content;
[0145] The training input analysis module uses a multimodal Transformer structure encoder to encode features into feature vectors based on the sales conversation scenario description. It then selects interaction feature combinations based on the correlation between customer personality and emotions to establish a multimodal training input set.
[0146] The virtual character construction module calls the multimodal training input set and inputs it into the GPT-4 language model based on the sales training requirements. It specifies the character attributes, analyzes the emotion intensity change parameters and the frequency of conversation topic switching, and annotates the emotion turning points and topic transfer nodes during the multi-round conversation generation process to construct the virtual sales training character data.
[0147] The practice recording module uses data from a virtual sales sparring partner to train salespeople on sales pitches using the speech flow generated by GPT-4. It monitors the input voice and text streams in real time, records the customer's emotional state and the accuracy of the financial advisor's response in each round of interaction, and generates a sales practice interaction record set.
[0148] The training effect analysis module calls the sales sparring interaction record set to extract the frequency of product introduction scripts, the correct response rate for objection handling, and the proportion of emotionally positive responses. By standardizing each indicator, each indicator is assigned a corresponding weight, the training score is calculated, and the sales sparring training score information is generated.
[0149] The above are merely preferred embodiments of the present invention and do not limit the present invention in any other form. Any technician familiar with the profession may use the technical content disclosed above to change or modify it into an equivalent embodiment with equivalent changes and apply it to other fields. However, any simple modification, equivalent change and modification made to the above embodiment based on the technical essence of the present invention without departing from the content of the technical solution of the present invention shall still fall within the scope of protection of the technical solution of the present invention.
Claims
1. A sales training method based on GPT and multimodal large model, characterized by: The following steps are involved: S1: Obtain sales training needs of bank sales staff, call customer basic attribute parameters, financial product function feature parameters, and market hot topic parameters, and use the GPT-4 language generation model to expand the complete conversation opening statement, demand discovery questions, and product introduction sentences based on the content to generate sales conversation scenario description content; S2: Based on the description of the sales conversation scenario, call the multimodal Transformer structure encoder to encode the features into feature vectors, screen the interactive feature combinations according to the customer personality and emotion correlation standard, and establish a multimodal training input set; S3: Call the multimodal training input set, input it into the GPT-4 language model according to the sales training requirements, specify the role attributes, analyze the emotion intensity change parameters and the frequency of conversation topic switching, mark the emotion turning points and topic transfer nodes in the multi-round conversation generation process, and construct the virtual sales training role data; S4: Based on the virtual sales sparring role data, call the speech flow generated by GPT-4 to train sales personnel to input sales speech, monitor the input voice stream and text stream in real time, record the customer's emotional state and the accuracy of the financial advisor's response to the speech in each round of interaction, and generate a sales sparring interaction record set.
2. The sales training method based on GPT and multimodal large model according to claim 1 is characterized in that: The sales dialogue scenario description content includes customer attribute description, product feature description, sales topic design, standardized opening statement, demand mining question set and product introduction statement set; the multimodal sparring input set includes voice feature coding set, text feature vector set, expression feature coding set and interaction feature matching result set; the virtual sales sparring role data includes customer basic information setting, financial management interest preference setting, risk tolerance level, emotion change label set and topic transfer node set; the sales sparring interaction record set includes voice input content record, text input content record, emotion classification tag record, speech matching score record and interaction round process record.
3. The sales training method based on GPT and multimodal large model according to claim 2 is characterized in that: The steps for obtaining the sales dialogue scenario description content are specifically as follows: S111: Obtain customer types, product categories, and sales topic items from the sales training needs of bank sales personnel, call customer basic information parameters, financial product function feature parameters, and market hot topic parameters, cross-screen the sales topic items and market hot topic parameters, extract corresponding macroeconomic indicators and policy change items, and establish customer product topic combination items; S112: Based on the customer-product topic combination items, the customer portrait feature items, the product function feature items, and the market hot topic items are arranged in order, the contents are sequentially connected, and a unified format is arranged according to the text paragraphing rules to generate a standardized description segment, thereby obtaining the customer portrait and product element topic description segment; S113: Based on the customer portrait and product element topic description segments, they are input into the GPT-4 language generation model, with sales opening remarks, demand mining questions and product introduction sentences as prompt items, and the generated text output content is called to extract text paragraphs that completely cover customer feature questions, product element introductions and product introduction content, and arrange them in the order of the standard sales dialogue process to generate sales dialogue scene description content.
4. The sales training method based on GPT and multimodal large model according to claim 3 is characterized in that: The steps for obtaining the multimodal training input set are specifically as follows: S211: Obtain the sales dialogue scenario description content, collect the fundamental frequency, energy, and speech rate data from the voice feature parameters, the rising and falling tone turning point data from the intonation feature parameters, the word frequency TF-IDF value and sentiment polarity score from the text feature parameters, and the FACS encoding data from the image expression feature parameters, and perform feature normalization processing uniformly according to the data structure format to generate multimodal feature normalized data; S212: Based on the multimodal feature normalized data, calling a multimodal Transformer structure encoder to perform feature vector encoding on the voice, intonation, text, and image features in the normalized data, obtaining feature vector groups after encoding, and calculating the interactive behavior matching degree between the customer personality feature vector and the emotion feature vector to obtain the interactive behavior matching coefficient; S213: Based on the interaction behavior matching coefficient, screen the feature combinations whose matching coefficients are greater than the set matching threshold, combine the corresponding speech feature subsets, intonation feature subsets, text feature subsets and image expression feature subsets, integrate them into a unified input sample format, and obtain a multimodal training input set.
5. The sales training method based on GPT and multimodal large model according to claim 4 is characterized in that: The steps for obtaining the virtual sales training role data are as follows: S311: Calling the multimodal training input set, setting a role parameter set according to the sales training requirements, collecting name setting items, total assets setting items, investment knowledge level setting items, and risk tolerance level setting items, combining the setting items into a role feature vector based on customer group standards, unifying the value of each parameter, and standardizing the feature scale to generate a role feature normalized vector; S312: Based on the normalized character feature vector, the vector is input into the GPT-4 language model, and the normalized parameter content is used to generate the character setting text, extract the emotion intensity change parameter and the conversation topic switching frequency, call the emotion intensity change curve and the topic change node marking content, sort the nodes, mark the emotion turning points and topic transfer nodes, and obtain the character emotion and topic marking data; S313: Based on the character emotion and topic annotation data, filter the emotion fluctuation nodes and topic connection nodes that meet the natural language interaction process standards, combine the emotion corpus and topic corpus corresponding to the sales training scenario, integrate the text, emotion labels and topic jump relationships according to the character feature vector requirements, and establish virtual sales training character data.
6. The sales training method based on GPT and multimodal large model according to claim 5 is characterized in that: The steps for obtaining the sales practice interaction record set are as follows: S411: Call the virtual sales training role data, configure the sales training platform, input the speech flow generated by GPT-4, set the customer role attributes according to the sales training requirements, set the sales staff training goals and response range, establish the basic sales training scenario, and generate sales training scenario configuration data; S412: Based on the sales practice scenario configuration data, the voice and text streams input by the salesperson are monitored in real time, the fundamental frequency changes and energy fluctuations in the voice stream are monitored, key semantic features of the input speech in the text stream are extracted, the speech emotion category and the text speech content are detected, an interaction deviation value is calculated, and valid interaction records with interaction deviation values less than an interaction stability threshold are screened to obtain sales practice behavior evaluation data; S413: Based on the sales sparring behavior evaluation data, record each round of data that meets the interaction deviation less than the interaction stability threshold, extract the salesperson's input text, voice emotion category, emotion deviation and text matching score, mark the sparring customer feedback actions and emotional state, combine and archive them into a standard training record format, and establish a sales sparring interaction record set.
7. The sales training method based on GPT and multimodal large model according to claim 6 is characterized in that: The method further comprises the following steps: S5: Calling the sales sparring interaction record set, extracting the frequency of product introduction speech, the correct response rate for objection handling, and the proportion of emotionally positive responses, normalizing each indicator, assigning a corresponding weight to each indicator, calculating the training score, and generating sales sparring training score information; The sales coaching training score information includes a product introduction frequency score, an objection handling accuracy score, an emotional positive response ratio score, and a comprehensive weighted training score.
8. The sales training method based on GPT and multimodal large model according to claim 7 is characterized in that: The steps for obtaining the sales training score information are as follows: S511: Calling the sales sparring interaction record set, extracting product introduction speech frequency data, objection handling correct response ratio data, and emotionally positive response ratio data, and standardizing each data item to establish a sales sparring indicator standardized data set; S512: Based on the standardized sales training indicator dataset, weights for product introduction speech frequency, objection handling response ratio, and emotionally positive response ratio are set according to the indicator importance, and corresponding weight values are assigned to each. Calculation is performed to obtain a sales training score, and training samples with scores greater than a training threshold are screened to obtain a sales training score calculation result. S513: Based on the sales coaching score calculation results, filter the score samples that meet the training standards, extract the training score and indicator details of each sample, integrate the frequency of product introduction scripts, the correct response rate for objection handling, and the proportion of positive emotional responses, and establish sales coaching training score information.
9. A sales training system based on GPT and multimodal large models, characterized by: The system is used to implement the sales training method based on GPT and a multimodal large model according to any one of claims 1 to 7, and the system includes: The conversation scenario analysis module obtains the sales training needs of bank sales staff, calls on basic customer attribute parameters, financial product function feature parameters, and market hot topic parameters, and uses the GPT-4 language generation model to expand the complete conversation opening statement, demand discovery questions, and product introduction sentences based on the content to generate sales conversation scenario description content; The training input analysis module uses a multimodal Transformer structure encoder to encode the features into feature vectors based on the description of the sales conversation scenario, and selects the interaction feature combination according to the customer personality and emotion correlation standard to establish a multimodal training input set; The virtual character construction module calls the multimodal training input set, inputs it into the GPT-4 language model according to the sales training requirements, specifies the character attributes, analyzes the emotion intensity change parameters and the frequency of conversation topic switching, marks the emotion turning points and topic transfer nodes in the multi-round conversation generation process, and constructs the virtual sales training character data; The training recording module uses the data of the virtual sales sparring role to call the speech flow generated by GPT-4 to train the salesperson to input sales speech, monitor the input voice stream and text stream in real time, record the customer's emotional state and the accuracy of the financial advisor's response in each round of interaction, and generate a sales training interaction record set; The training effect analysis module calls the sales sparring interaction record set, extracts the frequency of product introduction scripts, the correct response rate for objection handling, and the proportion of positive emotional responses, standardizes each indicator, assigns a corresponding weight to each indicator, calculates the training score, and generates sales sparring training score information.
Citation Information
Cited By
User emotion recognition method and device, electronic equipment and storage medium
CN120951293A
Financial investment adviser training system and method
CN121032746A
Method and system for dynamically grading transaction prediction probability based on user behavior data
CN121504532A
Customer service staff training method and system in post-loan management scene
CN121563735A