Model training method and question and answer method
By introducing multimodal data processing and modal parameter adjustment methods into the question-and-answer model, the problem of single output of the question-and-answer model and lack of personalization in salesperson training is solved, multimodal data processing and personalized business training are realized, and user experience and training effects are improved.
Patent Information
- Application Number
- CN202510341694.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-05-23
AI Technical Summary
In the prior art, the result output method of the question and answer model is relatively simple, and users can only obtain answers by viewing text, which lacks personalization, resulting in poor user experience. At the same time, traditional salesperson training methods lack personalization and interactivity, making it difficult to meet the needs of salesperson personalized training.
Through a model training method, the modal parameters and input preset questions are used to determine the modality of the output result, generate the target answer, and adjust the modal parameters of the question-and-answer model until the iterative training termination condition is met, and the training completed question-and-answer model is obtained. The model is able to blend a variety of modal data such as text, video and images to generate diverse, intuitive and accurate answers.
The multimodal data processing capability of the question-and-answer model is realized, and more rich, intuitive and accurate answers are generated, which improves the user experience, and through personalized business training, the efficiency, accuracy and personalization of the training are improved.
Smart Images

Figure CN120030353A_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the field of computer technology, and in particular, relates to a model training method and a question-answering method. Background Art
[0002] At present, the big language model analyzes the user's questions through natural language understanding and machine learning models, and generates answers corresponding to the questions based on the question-answering knowledge base. However, in the related technologies, the result output method of the question-answering model is relatively simple, and the user can only get the answer they want by viewing the text. This lacks personalization and makes the user experience poor. At the same time, traditional sales training methods mostly rely on offline courses and paper materials. These methods often lack personalization and interactivity. Therefore, how to achieve personalized training for salesmen based on the big language model has become an urgent problem that needs to be solved. Summary of the invention
[0003] The embodiments of the present application provide a model training method and a question-answering method, which can realize personalized training of sales staff based on a large language model.
[0004] In a first aspect, an embodiment of the present application provides a model training method, the method comprising: determining a second modality of an output result based on modal parameters and a preset question of a first modality of input, wherein the first modality is used to characterize the form of expression of the preset question, the second modality is used to characterize the form of expression of the output result, and the modal parameters are used to balance the contribution of different modalities; generating a target answer corresponding to the preset question according to the second modality; adjusting the modal parameters of the question-answering model according to the target answer and the preset answer corresponding to the preset question until the iterative training termination condition is met, thereby obtaining the trained question-answering model.
[0005] In a second aspect, an embodiment of the present application provides a question-answering method, the method comprising: obtaining a target question; inputting the target question into a question-answering model to generate answer information corresponding to the target question; wherein the question-answering model is trained based on the training method described in the first aspect; and outputting the answer information.
[0006] In the third aspect, an embodiment of the present application provides a model training device, which includes: a determination module, used to determine a second modality of an output result based on modal parameters and a preset question of an input first modality, wherein the first modality is used to characterize the form of expression of the preset question, the second modality is used to characterize the form of expression of the output result, and the modal parameters are used to balance the contribution of different modalities; a generation module, used to generate a target answer corresponding to the preset question according to the second modality; an adjustment module, used to adjust the modal parameters of the question and answer model according to the target answer and the preset answer corresponding to the preset question, until the iterative training termination condition is met to obtain the trained question and answer model.
[0007] In a fourth aspect, an embodiment of the present application provides a question-and-answer device, comprising: an acquisition module for acquiring a target question; a generation module for inputting the target question into a question-and-answer model to generate answer information corresponding to the target question; wherein the question-and-answer model is an output module trained based on the training method described in the first aspect, and is used to output the answer information.
[0008] In a fifth aspect, an embodiment of the present application provides an electronic device, comprising a processor, a memory, and a program or instruction stored in the memory and executable on the processor, wherein the program or instruction, when executed by the processor, implements the steps of the method described in the first aspect, or implements the steps of the method described in the second aspect.
[0009] In a sixth aspect, an embodiment of the present application provides a computer-readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, the steps of the method described in the first aspect are implemented, or the steps of the method described in the second aspect are implemented.
[0010] In the seventh aspect, an embodiment of the present application provides a chip, which includes a processor and a communication interface, wherein the communication interface is coupled to the processor, and the processor is used to run a program or instructions to implement the steps of the method described in the first aspect, or to implement the steps of the method described in the second aspect.
[0011] In an eighth aspect, an embodiment of the present application provides a computer program product, comprising a computer program stored on a non-transitory computer-readable storage medium, wherein the computer program comprises program instructions, and when the program instructions are executed by a computer, the computer implements the steps of the method described in the first aspect, or implements the steps of the method described in the second aspect.
[0012] In an embodiment of the present application, a second modality of the output result is determined according to modal parameters and a preset question of the first modality input, wherein the first modality is used to characterize the expression form of the preset question, the second modality is used to characterize the expression form of the output result, and the modal parameters are used to balance the contribution of different modalities. Then, according to the second modality, a target answer corresponding to the preset question is generated, and then the modal parameters of the question-answering model are adjusted according to the target answer and the preset answer corresponding to the preset question until the termination condition of the iterative training is met, and a trained question-answering model is obtained, so that the trained question-answering model can fully understand the input preset question by integrating multiple modal data such as text, video and image, and then determine the target modality, and finally generate an accurate target answer. At the same time, by identifying the input preset question, the question-answering model can accurately determine whether the user's intention is to obtain text information or picture content or other modal content, and then return the corresponding modal answer according to the user's needs, which also improves the diversity, intuitiveness and accuracy of the expression form of the target answer, and can be used for business training, and realize the efficiency, precision and personalization of business training. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] Figure 1 It is a flow chart of a model training method provided in an embodiment of the present application; Figure 2 It is a flowchart of a question-and-answer method provided in an embodiment of the present application; Figure 3 This is a flow chart of a salesperson training method provided by an embodiment of the present application; Figure 4 It is a structural schematic diagram of a model training device provided in an embodiment of the present application; Figure 5 is a structural schematic diagram of a question-and-answer device provided in an embodiment of the present application; Figure 6 It is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0014] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0015] In the following, in combination with the accompanying drawings, a model training method and a question-answering method provided in an embodiment of the present application are described in detail through specific embodiments and their application scenarios.
[0016] Figure 1 A model training method provided by an embodiment of the present application is shown, and the method can be executed by an electronic device, and the electronic device may include: a server and / or a terminal device. In other words, the method can be executed by software or hardware installed in the electronic device, and the method includes the following steps: S110: Determine a second mode of the output result according to the modal parameters and the preset problem of the input first mode.
[0017] Among them, the first mode is used to characterize the expression form of the preset problem, the second mode is used to characterize the expression form of the output result, and the modal parameters are used to balance the contribution of different modes.
[0018] The first modality may include but is not limited to at least one of the following: text, image, audio, video, table, chart. The second modality may include but is not limited to at least one of the following: text, image, audio, video, table, chart. The preset question may include but is not limited to: business consultation, business operation, product information query, etc.
[0019] It is understandable that by analyzing the preset question, it is determined in which mode or modes the answer corresponding to the preset question should be output. In this way, after determining the second mode of the output result, the answer corresponding to the preset question will be generated according to the second mode in the subsequent answer generation process. At the same time, in the embodiment of the present application, the specific mode of the input preset question is not limited. The preset question can be multimodal data, for example, it can be a combination of text and image, that is, the question-answering model can not only recognize text data, but also recognize other data such as images and videos, so that the information of different modes is associated and integrated, so as to achieve a richer and more comprehensive interpretation of the preset question.
[0020] S120: Generate a target answer corresponding to the preset question according to the second mode.
[0021] It can be understood that after determining the second modality, the question-answering model will generate the target answers corresponding to the preset questions according to the second modality, which can realize the generation of precisely customized training content to meet the learning needs of different salesmen.
[0022] S130: According to the target answer and the preset answer corresponding to the preset question, the modal parameters of the question-answering model are adjusted until the iterative training termination condition is met, thereby obtaining the trained question-answering model.
[0023] It can be understood that by adjusting the modal parameters, the model can accurately output the target answer that matches the preset answer based on the input of the preset question until the iterative training termination condition is met, and finally a trained model is obtained.
[0024] Among them, in an exemplary embodiment, the input of the model can be expressed as: F=αFt+βFv+γFi, where Ft, Fv and Fi are feature vectors of text, video and image respectively, α, β and γ can be the weights of each modality included in the modal parameters, and the fused feature vector F is used as input to train the question-answering model. Among them, the question-answering model adopts an encoder-decoder architecture, the encoder encodes the input question, and the decoder generates the answer. In an embodiment of the present application, training is achieved by combining static modal parameters with dynamic modal weights, that is, the static modal parameters are optimized for a long time through back propagation (based on output errors), and the weights are flexibly allocated according to the input questions, so that it can quickly adapt to new business scenarios and customer needs, and ensure the timeliness and accuracy of the answers.
[0025] It should be noted that the embodiments of the present application are only exemplified using three modes: text, video and image, and do not specifically limit the type of mode.
[0026] In an exemplary embodiment, S130 may include: calculating the difference between the target answer and the preset answer; and then adjusting the modal parameters using an optimization algorithm until the difference between the target answer and the preset answer reaches a preset termination condition after multiple iterations, thereby completing the model training. Optionally, the termination condition may be a cross entropy loss: ; Among them, yi represents the predicted probability corresponding to the target answer, and y^i is the label probability corresponding to the preset answer.
[0027] In addition, during the training process, the validation set can be used to evaluate the model and adjust hyperparameters such as learning rate and weight coefficient. Finally, the model with the best performance is selected for saving. For example, 1,000 conversation records between salesmen and customers, 500 product demonstration videos, and 300 product display images are collected. After training the question-answering model using the above method, if the accuracy of the question-answering model on the validation set reaches more than 90%, the question-answering model can be saved.
[0028] In an embodiment of the present application, a second modality of the output result is determined according to a modality parameter and a preset question of the first modality input, wherein the first modality is used to characterize the expression form of the preset question, the second modality is used to characterize the expression form of the output result, and the modality parameter is used to balance the contribution of different modalities. Then, according to the second modality, a target answer corresponding to the preset question is generated, and then the modality parameter of the question-answering model is adjusted according to the target answer and the preset answer corresponding to the preset question until the termination condition of the iterative training is met, and a trained question-answering model is obtained, so that the trained question-answering model can fully understand the input preset question by integrating multiple modality data such as text, video and image, and then determine the target modality, and finally generate an accurate target answer. At the same time, by identifying the input preset question, the question-answering model can accurately determine whether the user's intention is to obtain text information or picture content or other modal content, and then return the corresponding modal answer according to the user's needs, which also improves the diversity, intuitiveness and accuracy of the expression form of the target answer, and can be used for business training to achieve efficient, precise and personalized business training.
[0029] In an exemplary embodiment, the above S120 may include: S122: Determine the second weight of each third modality by making a first adjustment to the first weight corresponding to each third modality according to the question type of the preset question.
[0030] Among them, the question type is used to indicate the intention of asking the question, the first weight corresponding to each of the third modalities is determined based on the target analysis method and the modal parameters, and the third modality includes the second modality.
[0031] S124: Determine the second mode according to the second weight of each of the third modes.
[0032] It is understandable that in order to improve the accuracy of the determined second modality, the second modality can be determined from the third modality according to the type of the preset question. The question type may include but is not limited to: concept, operation, details, etc. For example, if the preset question indicates obtaining product details or operation steps, the weight of the video modality or image modality is increased; if the question indicates about concepts or explanations, the weight of the text modality is increased.
[0033] In an exemplary embodiment, the above S122 may include the following steps: Step 1: Obtain intent keywords from the preset question.
[0034] Step 2: Obtain the modal attributes corresponding to the intent keyword.
[0035] Step 3: Determine the second weight of each of the third modalities by making a first adjustment to the first weight of each of the third modalities according to the modal attribute corresponding to the intention keyword.
[0036] It is understandable that in the process of determining the second weight of each third modality by making a first adjustment to the first weight corresponding to each third modality according to the question type of the preset question, different intent keywords correspond to different modal attributes. Therefore, after extracting each intent keyword from the preset question, further, according to the modal attribute corresponding to the intent keyword, a first adjustment is made to the first weight of each third modality to determine the second weight of each third modality. Exemplarily, assuming that the intent keyword is "operation steps", the modal attributes corresponding to the intent keyword are "pictures" and "videos", then the first weights of the picture modality and the video modality are increased, thereby obtaining the second weight of each third modality.
[0037] In this embodiment, for the preset questions input by the user, the intent keywords used to express the information type and needs are identified, and then the corresponding modal attributes are obtained based on the intent keywords. Finally, the return content is generated according to the user's needs. If it is a text modality, the processed text information is directly returned to the user; if it is a picture modality, the relevant image data is retrieved or generated, and the format is automatically converted or optimized to ensure that it is presented to the user in the most appropriate form, thereby improving the user experience and satisfaction.
[0038] In another exemplary embodiment, the above S124 may include: determining the third weight of each of the third modalities by making a second adjustment to the second weight of each of the third modalities according to the scenario characteristics of the preset question, wherein the scenario characteristics are used to characterize the personalized characteristics of the user; and determining the second modality according to the third weight of each of the third modalities.
[0039] It is understandable that after adjusting the first weight to obtain the second weight, a secondary adjustment can be made to the first weight according to the scene characteristics of the preset question. The scene characteristics can be determined according to the device currently used by the user, or according to the characteristics indicated in the preset question. For example, if the user is using a mobile device, he or she may prefer text answers, so the weight of the text modality can be increased; if the user is in a demonstration scenario or using a large-screen device, he or she may prefer videos or images, so the weights of the video modality and the image modality can be increased.
[0040] In an exemplary embodiment, before determining the second weight of each third modality by first adjusting the first weight corresponding to each third modality according to the question type of the preset question, the method further includes: Obtaining initial weights of each third mode included in the modal parameters; Determine a primary weight corresponding to each of the third modes according to the first analysis method and the initial weight of each of the third modes; Determine the secondary weight corresponding to each of the third modes according to the second analysis method and the initial weight of each of the third modes; Determine the first weight according to the primary weight and the secondary weight corresponding to each of the third modes; Among them, the first analysis method is one of the hierarchical analysis method, the entropy analysis method, and the principal component analysis method, and the second analysis method is any analysis method among the hierarchical analysis method, the entropy analysis method, and the principal component analysis method other than the first analysis method.
[0041] It is understandable that by adjusting the static modal parameters through the target analysis method, the first weight determined according to the adjusted primary weight and secondary weight can quickly adapt to customer needs and ensure the accuracy of the answer. At the same time, the embodiment of the present application first determines the primary weight by one analysis method, and then uses another analysis method to determine the secondary weight, and then multiplies the combined weights and performs normalization to obtain the final weight coefficient, i.e., the first weight. Among them, the analytic hierarchy process (AHP) is a subjective weighting method based on expert scoring, that is, by constructing a judgment matrix, then comparing the importance of different modes pairwise through expert scoring, then calculating the maximum eigenvalue and corresponding eigenvector of the judgment matrix, and finally normalizing the eigenvector to obtain the weight of each mode; principal component analysis (PCA) is an objective weighting method based on data variance explanation rate, that is, by standardizing multimodal data, then calculating the variance explanation rate of each mode, and then normalizing according to the variance explanation rate to obtain the weight of each mode; entropy method is an objective weighting method based on information entropy, that is, by calculating the entropy value ej of each modal data, then calculating the difference coefficient gj=1−ej, and then normalizing the difference coefficient to obtain the weight of each mode.
[0042] Exemplarily, assuming that the first analysis method is a hierarchical analysis method, the second analysis method is a principal component analysis method, and the third modality may include a text modality, a video modality, and an image modality, wherein, according to the first analysis method, determining the primary weight corresponding to each of the third modalities may include the following steps: Step 1: Construct a judgment matrix. The importance of the three third modes is compared in pairs through expert scoring or salesperson scoring, as shown in Table 1: Table 1
[0043] According to the above scores, the judgment matrix A is constructed as: ; The rows of the matrix represent the “relative importance” and the columns represent the “modalities being compared”. For example, the importance of text relative to video is 2, which means that text is twice as important as video.
[0044] Step 2: Calculate the weight vector: Calculate the maximum eigenvalue (λmax) of the judgment matrix and the corresponding eigenvector (weight vector). The eigenvector can be determined by: First, normalize each column of the judgment matrix: ; Then find the average value of the normalized matrix row by row to get the eigenvector: Eigenvector = ; The calculation result is: eigenvector = (0.56, 0.28, 0.16). After multiplying the eigenvector of each third mode by the corresponding initial weight, normalization is performed to obtain the first-level weight.
[0045] In addition, in order to ensure the consistency of the judgment matrix, a consistency test is required to calculate the consistency index (CI) and consistency ratio (CR).
[0046] Among them, the consistency index (CI) can be expressed as: ; Among them, λmax is the maximum eigenvalue of the judgment matrix, and n is the dimension of the matrix (n=3 in this example).
[0047] Assuming that the maximum eigenvalue λmax=3.05 through calculation, then: ; Among them, the consistency ratio (CR) can be expressed as: ; where RI is the random consistency index (for n=3, RI=0.58).
[0048] ; If CR is less than the target threshold, the consistency of the judgment matrix is considered acceptable. Assuming that the target threshold is 0.1, in this example, CR=0.043<0.1, so the consistency check passes.
[0049] Among them, according to the second analysis method, determining the secondary weight corresponding to each of the third modes may include: calculating the variance explanation rate of each mode, performing normalization processing according to the variance explanation rate, and obtaining the weight of each mode. Exemplarily, the variance explanation rates of text, video and image obtained by principal component analysis are 45%, 30% and 25% respectively, then the characteristic vectors of each third mode determined by the second analysis method can be calculated as: ; ; .
[0050] After multiplying the eigenvector of each third mode by the corresponding initial weight, normalization is performed to obtain the secondary weight.
[0051] After determining the primary weight and secondary weight of each third mode, the tertiary weight of each third mode can be obtained by combining and multiplying, for example: α=0.56*0.45=0.252, β=0.28*0.30=0.084, γ=0.16*0.25=0.04, and then the three-level weights are normalized to obtain the first weight of each third mode.
[0052] In the above example, the first-level weights of the three third modes are determined through the analytic hierarchy process (AHP). This method combines the needs of experts or salesmen, can scientifically balance the contributions of different modes, and improve the accuracy and practicality of determining the target mode. The second-level weights of the three third modes are then determined through principal component analysis. The principal component analysis method can maximize the explanation of the variation of the original modal data, thereby selecting the features that have the greatest impact on data changes as new feature representations. During the processing process, it can be decided which components to retain based on the explanatory contribution of the principal components, making the decision more accurate and reliable.
[0053] In an exemplary embodiment, generating a target answer corresponding to the preset question according to the second modality includes: retrieving candidate answers corresponding to the preset question from target salesperson service data and training data corresponding to the target product according to the second modality; and obtaining the target answer by arranging the candidate answers.
[0054] Among them, the target salesperson’s service data may include daily work service videos, customer feedback audios, and customer service feedback record documents; training data may include training demonstration videos, work attitude service description documents, business processing description documents, and communication skills description documents.
[0055] It can be understood that by understanding the user's needs, that is, expecting to obtain information in text mode, image mode or other mode, the question-and-answer model accurately retrieves answers that match the questions from the target salesperson's service data and training data, and presents them according to the mode specified by the user, so that the user can receive the required information clearly and intuitively.
[0056] In another exemplary embodiment, before retrieving candidate answers corresponding to the preset question from the target salesperson service data and training data corresponding to the target product in accordance with the second mode, the method may further include the following steps: Step 1: Acquire multiple groups of initial salesperson service data corresponding to the target product.
[0057] Among them, multiple groups of initial salesman service data can be obtained from historical salesman service data, and the initial salesman service data can include multi-modal data.
[0058] In an exemplary embodiment, after obtaining multiple groups of initial salesperson service data, each group of initial salesperson service data can be preprocessed, wherein the preprocessing includes: preliminary preprocessing of the collected multimodal data, such as text cleaning, image recognition, video summarization, audio transcription, etc., and extracting key features.
[0059] Step 2: Determine the service quality weight of each group of the initial salesperson service data.
[0060] In an exemplary embodiment, determining the service quality weight of each group of the initial salesperson service data includes: for each group of the initial salesperson service data, determining that the service quality weight is composed of at least one of the following: (1) A fourth weight determined according to the customer sentiment tendency represented by the initial salesperson service data.
[0061] Wherein, determining the fourth weight may include but is not limited to one of the following: If the customer sentiment tendency identified at the initial stage of each round of dialogue is negative, and the customer sentiment tendency identified at the end stage of each round of dialogue is neutral, the fourth weight is assigned the first value; If the customer sentiment tendency identified at the initial stage of each round of dialogue is negative, and the customer sentiment tendency identified at the end stage of each round of dialogue is positive, the fourth weight is assigned the second value; If the customer sentiment tendency identified at the initial stage of each round of dialogue is neutral, and the customer sentiment tendency identified at the end stage of each round of dialogue is positive, the fourth weight is assigned the third value; If the customer sentiment tendency identified at the initial stage of each round of dialogue is positive, and if the customer sentiment tendency identified at the end stage of each round of dialogue is positive, a fourth weight is assigned as a fourth value; Wherein, the first value < the second value < the third value < the fourth value. For example, the first value is 1, the second value is 2, the third value is 1.5, and the fourth value is 2.
[0062] (2) A fifth weight determined according to the customer feedback data included in the initial salesperson service data.
[0063] Wherein, determining the fifth weight may include: identifying emotional keywords in the customer feedback data, and determining the fifth weight by the emotional keywords. Exemplarily, when the emotional keywords in the audio are identified as praise types, such as "Your service is attentive" or "Your service is good", etc., the fifth weight is assigned as the fifth value; when the customer service feedback record document is identified as a praise type in which the emotional keywords actively proposed by the user to the salesperson are recorded, the fifth weight is assigned as the fifth value.
[0064] (3) A sixth weight determined based on the question and answer records included in the initial salesperson service data.
[0065] Determining the sixth weight may include: identifying emotional keywords in the question and answer record, and determining the sixth weight based on the emotional keywords. After determining the emotional keywords, the sixth weight may be determined by referring to the method in (2) above, which will not be described in detail here.
[0066] (4) A seventh weight determined according to the service behavior accuracy score corresponding to the initial salesperson service data.
[0067] Among them, in an exemplary embodiment, before determining the seventh weight according to the service behavior accuracy score, the method also includes: determining the service behavior accuracy score according to the behavior recognition accuracy score, customer feedback satisfaction and document record consistency score, wherein the behavior recognition accuracy score is determined by identifying the actions performed by the salesperson in the process of serving the customer, the customer feedback satisfaction is determined based on the emotional inclination of the customer during the service process, and the document record consistency score is determined by comparing the interactive behavior recorded in the document with the actual interactive behavior during the service process.
[0068] Among them, the determination method of the behavior accuracy score can be expressed by the following formula: Behavior accuracy score = α*behavior recognition accuracy score + β*customer feedback satisfaction + γ*documentation consistency score; Among them, α, β, and γ are weight coefficients, which respectively represent the importance of behavior recognition accuracy score, customer feedback satisfaction score, and document record consistency score. Specifically: Through video analysis, identify the actions of the salesperson during the service process, such as "nodding" to indicate approval, "spreading hands" to indicate helplessness, etc. If the salesperson frequently "spreads his hands" when the customer asks a question, it may indicate that he is unfamiliar with the question. In an exemplary embodiment, the accuracy of action recognition can be evaluated by precision (Precision) and recall rate (Recall). Exemplarily, precision indicates the proportion of actions correctly recognized by the model to the total number of actions recognized, and recall rate indicates the proportion of actions correctly recognized by the model to the total number of actual actions. The accuracy of action recognition can be calculated by the following formula: Action recognition accuracy = Precision × Recall; ; ; TP (True Positive): the number of actions correctly identified; FP (False Positive): the number of actions that are incorrectly identified; FN (False Negative): The number of actual actions that were not recognized.
[0069] Among them, customer feedback satisfaction is determined based on the emotional tendency of the initial stage, the emotional tendency of the middle stage, and the emotional tendency of the final stage of each round of dialogue. When determining the emotional tendency of each stage, the conversation content between the salesperson and the customer is extracted, and the conversation content is input into the sentiment analysis model for deep semantic analysis to identify the emotional tendency in the service, wherein the emotional tendency includes positive, neutral, and negative. In an exemplary embodiment, the emotional tendency score S can be calculated by the following formula: ; Among them, Ppositive and Pnegative are the probabilities of positive and negative emotions respectively.
[0070] The emotional tendency is judged according to the score S: if S>0: the emotional tendency is positive; S=0: the emotional tendency is neutral; S<0: negative.
[0071] Assume that, in the initial stage, "Customer: I want to consult about the package change. Salesperson: OK, what do you want to change specifically?" is input into the sentiment analysis model. If the probability output by the sentiment analysis model is: positive: 0.2; neutral: 0.6; negative: 0.2; then the sentiment tendency score is: ; Therefore, the emotional tendency of customers is neutral in the initial stage.
[0072] In the middle stage, "Customer: I hope to change to a cheaper package. Salesperson: Sure, I will help you check the packages that meet the conditions." is input into the sentiment analysis model. If the probability output by the sentiment analysis model is: positive: 0.1; neutral: 0.7; negative: 0.2; then the sentiment tendency score is:
[0073] Therefore, the emotional tendency of customers in the middle stage is neutral to negative.
[0074] At the end stage, "Customer: Thank you for your help, I am very satisfied! Salesperson: You are welcome, it is a pleasure to serve you!" is input into the sentiment analysis model. If the probability output by the sentiment analysis model is: positive: 0.9; neutral: 0.1; negative: 0.0; then the sentiment tendency score is: ; Therefore, the emotional tendency of customers in the closing stage is positive.
[0075] The document record consistency score is determined by comparing the interaction behavior in the document record with the actual interaction behavior in the service process. The steps for determining the document record consistency score may include: Step 1: Content extraction, which uses natural language processing (NLP) technology to extract key information from documents (such as customer issues, solutions, service results, etc.).
[0076] Step 2: Comparative analysis, that is, comparing the interactive behaviors recorded in the document with the actual interactive behaviors during the service process (such as the conversation content in the video, service actions, etc.) to check whether there are inconsistencies.
[0077] Among them, the document record consistency score can be quantitatively evaluated by the following formula: ; Among them, the number of inconsistent items is the number of items that are inconsistent with the actual service content in the document record. The total number of inspection items is the total number of key information involved in the document record. For example: During the service process, a salesperson reported 3 problems to the customer, and the salesperson recorded the following in the document: Question 1: The records are consistent with the actual service content.
[0078] Issue 2: The records were inconsistent with the actual service content (a key issue raised by the customer was not mentioned in the document).
[0079] Question 3: The records are consistent with the actual service content.
[0080] According to the above formula, the document record consistency score is:
[0081] The document record consistency score shows that the consistency between the salesperson's document record and the actual service content is 67%, and there are certain record inaccuracies. In the salesperson's behavior accuracy score, the document record consistency score can be used as an important indicator. For example: if the consistency score is lower than a certain threshold (such as 70%), it may indicate that there is a problem with the salesperson's work attitude. If the consistency score is high (such as above 90%), it means that the salesperson's record is more accurate and the behavior is good. The document record consistency score can objectively and quantitatively evaluate the service behavior of the salesperson, help screen the service data of high-quality salespeople, and automatically identify and analyze the key behaviors of the salesperson, improve the efficiency and objectivity of behavior analysis, and help optimize salesperson training and improve service quality.
[0082] Step 3: Determine the target salesperson service data from multiple groups of initial salesperson service data according to the service quality weight.
[0083] It can be understood that according to the size of the service quality weight, high-quality salesperson service data, namely target salesperson service data, are screened out from multiple groups of initial salesperson service data. Among them, if the service quality weight is greater than a threshold, the corresponding initial salesperson service data is the target salesperson service data.
[0084] like Figure 2 As shown, the embodiment of the present application also provides a question-answering method, which includes the following steps: S210: Obtain the target question.
[0085] Among them, the target problem may include multimodal data.
[0086] S220: Input the target question into the question-answering model to generate answer information corresponding to the target question.
[0087] The question-answering model is based on Figure 1 The model is obtained by training using the model training method in the illustrated embodiment.
[0088] S230: Output the answer information.
[0089] For example, suppose a training salesperson inputs a question about a product function: a text modal question "How to use function Y of product X?" The question-answering model will determine the target modality based on the question. If the determined target modality is text, video, and image, the relevant text, video, and image will be retrieved, and finally the target answer corresponding to the preset question will be generated: a text answer "Function Y of product X can be used by following the steps below..." will be generated, and relevant video demonstrations and product pictures will be displayed.
[0090] In the embodiment of the present application, the target question is first obtained, and then the target question is input into the question-answering model to generate answer information corresponding to the target question, and finally the answer information is output. Figure 1 The model trained by the model training method provided in the illustrated embodiment can efficiently process multimodal data and provide more accurate and rich answers, thereby avoiding the training based solely on mechanical training data in related technologies, achieving personalized business training, and effectively improving training results.
[0091] Optionally, an embodiment of the present application further provides a scenario simulation system, which is used to automatically generate a service analysis report based on the training staff's behavior recognition and question-and-answer interaction, and point out the training staff's service advantages and room for improvement.
[0092] In this embodiment, the learning experience and knowledge application ability of the salesperson are enhanced through interactive learning means, and it is also convenient for the salesperson to experience the real situation in the simulated scene in an immersive way, so that the user's emotional tendency can be felt, which is more motivating, conducive to work service experience, and improves the training effect. In addition, the evaluation score in the service analysis report can be calculated using the output layer activation function of the neural network: S=σ(W⋅Fcombined+b) where S is the evaluation score, W is the weight matrix, b is the bias term, and σ is the activation function, such as sigmoid or softmax function.
[0093] Figure 3 A training method based on a large language model provided in an embodiment of the present application is shown, which may include the following steps: S310: Obtain multimodal training data and perform preprocessing.
[0094] S320: Filter high-quality salesperson service data based on the weight values of emotional tendency, customer feedback, question and answer pairs, and service behavior accuracy.
[0095] S330: Build a real-time question-and-answer system to obtain text, videos, and images input by training salespeople, combine high-quality salesperson service data and training data, and output accurate answers and explanations.
[0096] S340: Build a scenario simulation system. Based on the behavior recognition and question-answer interaction of the training staff, the system automatically generates a service analysis report, pointing out the service advantages and room for improvement of the training staff.
[0097] In this embodiment, not only text data can be recognized, but also image and video data can be recognized, which enables the model to play a role in a wider range of scenarios, associate and fuse information of different modalities, and thus provide richer and more comprehensive information interpretation. When recognizing image and video data, the model can extract key information and understand the meaning and contextual relationship behind it. This capability enables the model to more accurately identify targets in images and videos and understand scenes and situations. Based on the large language model and computer vision technology, the accuracy of service behavior, sentiment analysis, and question-answer pair data and result analysis are extracted and analyzed. By assigning weights and screening, historical high-quality work service data is obtained, and a real-time question-answering system and a scenario simulation system are constructed, so that users can quickly and accurately understand and remember the training content, without having to read and understand the text content for a long time, and it is easy to quickly receive knowledge, saving sales staff learning time and reducing human resource training costs.
[0098] Figure 4 A schematic diagram of the structure of a model training device provided by an embodiment of the present application is shown as follows: Figure 4 As shown, the model training device 400 may include: a determination module 410, a generation module 420 and an adjustment module 430.
[0099] In this embodiment, a determination module 410 is used to determine a second modality of an output result based on modal parameters and a preset question of an input first modality, wherein the first modality is used to characterize the form of expression of the preset question, the second modality is used to characterize the form of expression of the output result, and the modal parameters are used to balance the contribution of different modalities; a generation module 420 is used to generate a target answer corresponding to the preset question according to the second modality; an adjustment module 430 is used to adjust the modal parameters of the question-answering model according to the target answer and the preset answer corresponding to the preset question, until the iterative training termination condition is met, thereby obtaining the trained question-answering model.
[0100] In one implementation, the determination module 410 is specifically used to: determine the second weight of each of the third modalities by making a first adjustment to the first weight corresponding to each of the third modalities according to the question type of the preset question, wherein the question type is used to indicate the intention of the question, the first weight corresponding to each of the third modalities is determined based on the target analysis method and the modal parameters, and the third modality includes the second modality; determine the second modality according to the second weight of each of the third modalities.
[0101] In one implementation, the determination module 410 is specifically used to: obtain the intent keyword from the preset question; obtain the modal attribute corresponding to the intent keyword; and determine the second weight of each of the third modalities by making a first adjustment to the first weight of each of the third modalities according to the modal attribute corresponding to the intent keyword.
[0102] In one implementation, the determination module 410 is specifically used to: determine the third weight of each of the third modalities by making a second adjustment to the second weight of each of the third modalities according to the scenario characteristics of the preset question, wherein the scenario characteristics are used to characterize the personalized characteristics of the user; and determine the second modality according to the third weight of each of the third modalities.
[0103] In one implementation, the model training device also includes a first processing module, which is used to: obtain the initial weights of each third mode included in the modal parameters; determine the first-level weight corresponding to each third mode according to the first analysis method and the initial weights of each third mode; determine the second-level weight corresponding to each third mode according to the second analysis method and the initial weights of each third mode; determine the first weight according to the first-level weight and the second-level weight corresponding to each third mode; wherein the first analysis method is one of the hierarchical analysis method, the entropy analysis method, and the principal component analysis method, and the second analysis method is any one of the hierarchical analysis method, the entropy analysis method, and the principal component analysis method that is not the first analysis method.
[0104] In one implementation, the generation module 420 is specifically used to: retrieve candidate answers corresponding to the preset question from the target salesperson service data and training data corresponding to the target product according to the second mode; and obtain the target answer by sorting the candidate answers.
[0105] In one implementation, the model training device also includes a second processing module, which is used to: obtain multiple groups of initial salesperson service data corresponding to the target product; determine the service quality weight of each group of the initial salesperson service data; and determine the target salesperson service data from the multiple groups of the initial salesperson service data based on the service quality weight.
[0106] In one implementation, the second processing module is specifically used to: for each group of the initial salesperson service data, determine that the service quality weight is composed of at least one of the following: a fourth weight determined based on the customer emotional tendency represented by the initial salesperson service data; a fifth weight determined based on the customer feedback data included in the initial salesperson service data; a sixth weight determined based on the question and answer records included in the initial salesperson service data; a seventh weight determined based on the service behavior accuracy score corresponding to the initial salesperson service data.
[0107] In one implementation, the model training device also includes a third processing module, which is used to determine the service behavior accuracy score based on the behavior recognition accuracy score, customer feedback satisfaction score and document record consistency score, wherein the behavior recognition accuracy score is determined by identifying the actions performed by the salesperson in the process of serving the customer, the customer feedback satisfaction score is determined based on the emotional inclination of the customer during the service process, and the document record consistency score is determined by comparing the interactive behavior recorded in the document with the actual interactive behavior during the service process.
[0108] The training device provided in the embodiment of the present application can achieve Figure 1 To avoid repetition, the various processes implemented in the illustrated method embodiment will not be described again here.
[0109] Figure 5 A schematic diagram showing the structure of a question-answering device provided by an embodiment of the present application is shown in FIG. Figure 5 As shown, the question-answering device 500 may include: an acquisition module 510 , a generation module 520 and an output module 530 .
[0110] In this embodiment, the acquisition module 510 is used to acquire the target question; the generation module 520 is used to input the target question into the question-answering model to generate answer information corresponding to the target question; wherein the question-answering model is based on Figure 1 The method embodiment shown is obtained by training; the output module 530 is used to output the answer information.
[0111] The question-answering device provided in the embodiment of the present application can realize Figure 2 To avoid repetition, the various processes implemented in the illustrated method embodiment will not be described again here.
[0112] A model training device and a question-answering device in the embodiment of the present application can be a device, or a component, an integrated circuit, or a chip in an electronic device. The embodiment of the present application is not specifically limited.
[0113] A model training device and a question-answering device in the embodiment of the present application may be a device having an operating system. The operating system may be an Android operating system, an iOS operating system, or other possible operating systems, which are not specifically limited in the embodiment of the present application.
[0114] Optional, such as Figure 6 As shown, an embodiment of the present application also provides an electronic device 600, including a processor 610, a memory 620, and a program or instruction stored in the memory 620 and executable on the processor 610. When the program or instruction is executed by the processor 610, the various processes of the above-mentioned model training method embodiment are implemented, or the various processes of the above-mentioned question-and-answer method embodiment are implemented, and the same technical effect can be achieved. To avoid repetition, it will not be repeated here.
[0115] An embodiment of the present application also provides a computer-readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, the various processes of the above-mentioned model training method embodiment are implemented, or the various processes of the above-mentioned question-and-answer method embodiment are implemented, and the same technical effect can be achieved. To avoid repetition, it will not be repeated here.
[0116] The processor is the processor in the electronic device described in the above embodiment. The readable storage medium includes a computer readable storage medium, such as a computer read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0117] An embodiment of the present application further provides a chip, which includes a processor and a communication interface, wherein the communication interface is coupled to the processor, and the processor is used to run programs or instructions to implement the various processes of the above-mentioned model training method embodiment, or to implement the various processes of the above-mentioned question-and-answer method embodiment, and can achieve the same technical effect. To avoid repetition, it will not be repeated here.
[0118] It should be understood that the chip mentioned in the embodiments of the present application can also be called a system-level chip, a system chip, a chip system or a system-on-chip chip, etc.
[0119] An embodiment of the present application also provides a computer program product, which includes a computer program stored on a non-transitory computer-readable storage medium, and the computer program includes program instructions. When the program instructions are executed by a computer, the computer implements the various processes of the above-mentioned model training method embodiment, or implements the various processes of the above-mentioned question-and-answer method embodiment, and can achieve the same technical effect. To avoid repetition, they are not repeated here.
[0120] It should be noted that, in this article, the terms "comprise", "include" or any other variant thereof are intended to cover non-exclusive inclusion, so that the process, method, article or device including a series of elements includes not only those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise one..." do not exclude the presence of other identical elements in the process, method, article or device including the element. In addition, it should be pointed out that the scope of the methods and devices in the embodiments of the present application is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in reverse order according to the functions involved, for example, the described method may be performed in an order different from that described, and various steps may also be added, omitted, or combined. In addition, the features described with reference to certain examples may be combined in other examples.
[0121] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, a magnetic disk, or an optical disk), and includes a number of instructions for a terminal (which can be a mobile phone, a computer, a server, an air conditioner, or a network device, etc.) to execute the methods described in each embodiment of the present application.
[0122] The embodiments of the present application are described above in conjunction with the accompanying drawings, but the present application is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of the present application, ordinary technicians in this field can also make many forms without departing from the purpose of the present application and the scope of protection of the claims, all of which are within the protection of the present application.
Claims
1. A model training method, characterized in that: The method comprises: Determine a second mode of the output result according to the modal parameters and the preset problem of the first mode of input, wherein the first mode is used to characterize the expression form of the preset problem, the second mode is used to characterize the expression form of the output result, and the modal parameters are used to balance the contribution of different modes; According to the second mode, generating a target answer corresponding to the preset question; According to the target answer and the preset answer corresponding to the preset question, the modal parameters of the question-answering model are adjusted until the iterative training termination condition is met, thereby obtaining the trained question-answering model.
2. The method according to claim 1, characterized in that The step of determining the second mode of the output result according to the modal parameters and the preset problem of the input first mode includes: Determining a second weight of each third modality by first adjusting the first weight corresponding to each third modality according to the question type of the preset question, wherein the question type is used to indicate the intention of asking a question, the first weight corresponding to each third modality is determined based on the target analysis method and the modal parameters, and the third modality includes the second modality; The second mode is determined according to the second weight of each of the third modes.
3. The method according to claim 2, characterized in that The determining the second weight of each third modality by first adjusting the first weight corresponding to each third modality according to the question type of the preset question includes: Obtaining intent keywords from the preset questions; Obtaining the modal attribute corresponding to the intent keyword; The second weight of each of the third modalities is determined by performing a first adjustment on the first weight of each of the third modalities according to the modal attribute corresponding to the intention keyword.
4. The method according to claim 2, characterized in that: The determining the second mode according to the second weight of each of the third modes includes: Determine the third weight of each of the third modalities by performing a second adjustment on the second weight of each of the third modalities according to the scenario characteristics of the preset question, wherein the scenario characteristics are used to characterize the personalized characteristics of the user; The second mode is determined according to the third weight of each of the third modes.
5. The method according to claim 2, characterized in that: Before determining the second weight of each third modality by first adjusting the first weight corresponding to each third modality according to the question type of the preset question, the method further includes: Obtaining initial weights of each third mode included in the modal parameters; Determine a primary weight corresponding to each of the third modes according to the first analysis method and the initial weight of each of the third modes; Determine the secondary weight corresponding to each of the third modes according to the second analysis method and the initial weight of each of the third modes; Determine the first weight according to the primary weight and the secondary weight corresponding to each of the third modes; Among them, the first analysis method is one of the hierarchical analysis method, the entropy analysis method, and the principal component analysis method, and the second analysis method is any analysis method among the hierarchical analysis method, the entropy analysis method, and the principal component analysis method other than the first analysis method.
6. The method according to claim 1, characterized in that Generating a target answer corresponding to the preset question according to the second mode includes: According to the second mode, a candidate answer corresponding to the preset question is retrieved from the target salesperson service data and training data corresponding to the target product; The target answer is obtained by sorting the candidate answers.
7. The method according to claim 6, characterized in that Before retrieving candidate answers corresponding to the preset question from the target salesperson service data and training data corresponding to the target product in accordance with the second mode, the method further includes: Acquire multiple groups of initial salesperson service data corresponding to the target product; Determining a service quality weight for each group of the initial salesperson service data; The target salesperson service data is determined from multiple groups of initial salesperson service data according to the service quality weight.
8. The method according to claim 7, characterized in that The determining of the service quality weight of each group of the initial salesperson service data includes: For each set of the initial salesperson service data, determining the service quality weight is composed of at least one of the following: A fourth weight determined according to the customer sentiment tendency represented by the initial salesperson service data; a fifth weight determined according to customer feedback data included in the initial salesperson service data; a sixth weight determined according to the question and answer records included in the initial salesperson service data; A seventh weight is determined according to the service behavior accuracy score corresponding to the initial salesperson service data.
9. The method according to claim 8, characterized in that Before determining the seventh weight according to the service behavior accuracy score, the method further includes: The service behavior accuracy score is determined based on the behavior recognition accuracy score, customer feedback satisfaction score and document record consistency score, wherein the behavior recognition accuracy score is determined by identifying the actions performed by the salesperson in the process of serving the customer, the customer feedback satisfaction score is determined based on the customer's emotional inclination during the service process, and the document record consistency score is determined by comparing the interactive behavior recorded in the document with the actual interactive behavior during the service process.
10. A question-answering method, characterized in that: The method comprises: Get the target question; Inputting the target question into a question-answering model to generate answer information corresponding to the target question; wherein the question-answering model is trained based on the training method described in any one of claims 1 to 9; The answer information is output.
11. A model training device, characterized in that: include: A determination module, configured to determine a second mode of an output result according to a modal parameter and a preset problem of an input first mode, wherein the first mode is used to characterize a representation of the preset problem, the second mode is used to characterize a representation of the output result, and the modal parameter is used to balance the contribution of different modes; A generating module, configured to generate a target answer corresponding to the preset question according to the second mode; An adjustment module is used to adjust the modal parameters of the question-answering model according to the target answer and the preset answer corresponding to the preset question until the iterative training termination condition is met to obtain the trained question-answering model.
12. A question-answering device, characterized in that: include: An acquisition module is used to obtain the target problem; A generation module, used to input the target question into a question-answering model to generate answer information corresponding to the target question; wherein the question-answering model is trained based on the training method according to any one of claims 1 to 9; An output module is used to output the answer information.
13. An electronic device, comprising: processor; A memory for storing instructions executable by the processor; wherein the processor is configured to execute the instructions to implement the method according to any one of claims 1 to 10.
14. A computer-readable storage medium having a program or instruction stored thereon, wherein when the instructions in the storage medium are executed by a processor of an electronic device, the electronic device is enabled to execute the method according to any one of claims 1 to 10.
15. A computer program product, comprising a computer program stored on a non-transitory computer-readable storage medium, wherein the computer program comprises program instructions, and when the program instructions are executed by a computer, the computer implements the method according to any one of claims 1 to 10.