Multi-modal training method and device based on intelligent quality inspection, equipment and medium
Through intelligent quality inspection methods, multi-dimensional scoring of the conversation recordings and clue conversion rates of the invitees is scored, quality inspection information data is generated and personalized training is provided, which solves the problems of inaccurate quality inspection and high labor costs in the existing technology, and achieves more efficient and accurate quality inspection and training results.
Patent Information
- Application Number
- CN202510268962.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-07
- Publication Date
- 2025-06-24
AI Technical Summary
In the prior art, the quality inspection of the invitee is inaccurate and incomplete, and the labor cost is high, making it difficult to guarantee the accuracy and consistency of the quality inspection results.
A multimodal training method based on intelligent quality inspection is adopted. By obtaining the original dialogue recording and clue conversion rate between employees and users, text conversion and cleaning is performed, recording text data is intelligently scored according to preset scoring rules, quality inspection information data is generated, and text or audio-visual training tasks are output.
A comprehensive assessment of the comprehensive capabilities of invitees has been achieved, the accuracy and coverage of quality inspection have been improved, labor costs have been reduced, and training results and invitation conversion rates have been enhanced.
Smart Images

Figure CN120198012A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular, to a multi-modal training method, device, equipment, and medium based on intelligent quality inspection. Background Art
[0002] An invitation clerk is a type of business-oriented salesperson who is mainly responsible for inviting customers to meetings, events, product exhibitions, etc. through means such as phone calls, emails, and text messages. His job nature involves multiple aspects such as customer relationship management, sales, and marketing, and he needs to have good communication skills, customer service awareness, and certain sales promotion experience. Taking the invitation of truck drivers as an example, the invitation clerk needs to contact registered users by phone and invite them to become registered drivers to improve the platform's transportation capacity.
[0003] The work content of the invitation clerk is extremely important. Experienced invitation clerks have a high lead conversion rate, while inexperienced invitation clerks have a low lead conversion rate and may even leave a bad impression on users. Therefore, it is often necessary to conduct quality inspections on invitation clerks to help managers better formulate policies.
[0004] Currently, quality inspection mainly relies on manpower and simple regular expression matching, which results in low accuracy and coverage rates (<1%). This not only increases labor costs but also makes it difficult to ensure the accuracy and consistency of quality inspection results. The existing methods for evaluating the capabilities of invitation clerks are relatively single, mainly relying on a few indicators such as invitation conversion rates, and cannot comprehensively reflect the comprehensive capabilities of invitation clerks. In addition, the training of invitation clerks mainly relies on offline manpower training, which requires a large amount of manpower and time, and the training effect is also difficult to be personalized and targeted. Summary of the Invention
[0005] In view of the above defects of the prior art, the present invention provides a multi-modal training method, device, equipment, and medium based on intelligent quality inspection to solve the technical problems of inaccurate, incomplete quality inspection of invitation clerks and high labor costs.
[0006] To achieve the above and other related objectives, the present invention provides a multi-modal training method based on intelligent quality inspection, including: obtaining the original conversation recording between the employee to be quality inspected and the user, as well as the lead conversion rate of the employee to be quality inspected; performing text conversion and cleaning on the original conversation recording to obtain recorded text data; performing quality inspection scoring on the recorded text data according to a preset scoring rule to obtain multi-dimensional scores, where the multi-dimensional scores are used to characterize the invitation capabilities of the employee to be quality inspected; generating quality inspection information data corresponding to the employee to be quality inspected when the lead conversion rate and the multi-dimensional scores meet the quality inspection conditions; and outputting text or audio-visual training tasks according to the quality inspection information data corresponding to the employee to be quality inspected.
[0007] In an embodiment of the present invention, the original conversation recording is subjected to text conversion and cleaning to obtain recorded text data, including: performing text conversion on the original conversation recording using a speech recognition model to obtain original recorded text data; and performing hot-word error correction on the original recorded text data to obtain the recorded text data.
[0008] In an embodiment of the present invention, quality inspection scoring is performed on the recorded text data according to a preset scoring rule to obtain multi-dimensional scores, including: judging whether there is business clue text or business clue sentence segments in the recorded text data according to the recorded text data and a preset business knowledge base: if so, obtaining the business accuracy score in the multi-dimensional scores according to the business clues included in the recorded text data and the reply of the employee to be quality inspected for the business clues; if not, the business accuracy score in the multi-dimensional scores is the full score.
[0009] In an embodiment of the present invention, obtaining the business accuracy score in the multi-dimensional scores according to the business clues included in the recorded text data and the reply of the employee to be quality inspected for the business clues includes: using multiple different first large models to evaluate the business clues included in the recorded text data and the reply of the employee to be quality inspected for the business clues to obtain the score corresponding to each first large model; judging whether the number of scores greater than a preset score value in the scores corresponding to each first large model exceeds a preset number: if so, the business accuracy score is the full score; if not, the business accuracy score is deducted according to the preset score value.
[0010] In an embodiment of the present invention, quality inspection scoring is performed on the recorded text data according to a preset scoring rule to obtain multi-dimensional scores, including: using a second large model to analyze the emotion, semantics, and context information of the words of the employee to be quality inspected in the recorded text data, and judging whether the words of the employee to be quality inspected belong to red-line behaviors to obtain a first judgment result and its confidence level; and obtaining the red-line behavior score in the multi-dimensional scores according to the first judgment result and its confidence level.
[0011] In an embodiment of the present invention, obtaining the red-line behavior score in the multi-dimensional scores according to the first judgment result and its confidence level includes: if the first judgment result is that all the words of the employee to be quality inspected do not belong to red-line behaviors, the red-line behavior score is the full score; if the first judgment result is that any word of the employee to be quality inspected belongs to a red-line behavior and its confidence level is higher than a preset confidence level, the red-line behavior score is deducted according to a preset score value; if the first judgment result is that any word of the employee to be quality inspected belongs to a red-line behavior and its confidence level is lower than the preset confidence level, manual review is triggered.
[0012] In an embodiment of the present invention, quality inspection scoring is performed on the recorded text data according to a preset scoring rule to obtain multi-dimensional scores, including: determining whether the beginning part of the recorded text data includes a preset opening statement to obtain a second judgment result, and / or determining whether the recorded text data includes a preset sensitive word to obtain a third judgment result; obtaining the service awareness score in the multi-dimensional scores according to the second judgment result, and / or obtaining the sensitive word score in the multi-dimensional scores according to the third judgment result.
[0013] In an embodiment of the present invention, quality inspection scoring is performed on the recorded text data according to a preset scoring rule to obtain multi-dimensional scores, including: using a third large model to process the recorded text data to obtain a driver invitation result; when the driver invitation result is an acceptance of the invitation, further determining whether the recorded text data includes document reminder information to obtain a fourth judgment result; and / or when the driver invitation result is a rejection of the invitation, further determining whether the recorded text data includes persuasion information to obtain a fifth judgment result; obtaining the process notification score in the multi-dimensional scores according to the fourth judgment result, and / or obtaining the invitation skill score in the multi-dimensional scores according to the fifth judgment result.
[0014] In an embodiment of the present invention, quality inspection information data corresponding to the employee to be quality inspected is generated, including: obtaining a lead conversion rate score according to the lead conversion rate and a preset mapping relationship between the lead conversion rate and its score; obtaining the final score of the employee to be quality inspected according to the multi-dimensional scores and the lead conversion rate score.
[0015] To achieve the above and other related purposes, the present invention also provides a multi-modal training device based on intelligent quality inspection, including: a data acquisition unit for acquiring the original conversation recording between the employee to be quality inspected and the user, and the lead conversion rate of the employee to be quality inspected; a preprocessing unit for performing text conversion and cleaning on the original conversation recording to obtain recorded text data; a scoring unit for performing quality inspection scoring on the recorded text data according to a preset scoring rule to obtain multi-dimensional scores, where the multi-dimensional scores are used to characterize the invitation ability of the employee to be quality inspected; a quality inspection information generation unit for generating quality inspection information data corresponding to the employee to be quality inspected when the lead conversion rate and the multi-dimensional scores meet the quality inspection conditions; a training unit for outputting text or audio-visual training tasks according to the quality inspection information data corresponding to the employee to be quality inspected.
[0016] To achieve the above-mentioned purpose and other related purposes, the present invention also provides an electronic device, including a processor, a memory and a communication bus; the communication bus is used to connect the processor and the memory; the processor is used to execute a computer program stored in the memory to implement a method provided in any one of the above embodiments.
[0017] To achieve the above-mentioned object and other related objects, the present invention further provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program is used to enable a computer to execute the method provided in any one of the above-mentioned embodiments.
[0018] Beneficial effects of the present invention: The present invention proposes a multimodal training method and device, equipment, and medium based on intelligent quality inspection. The method uses a large model to perform multi-dimensional intelligent scoring of call data through preset scoring rules, comprehensively evaluates the comprehensive ability of the inviter, and can freely and flexibly adjust the scoring rules, so that specific rules can be set according to needs; in addition, through subsequent training modules, personalized training content and simulation exercises of real scenarios are provided to help inviters quickly improve their invitation capabilities, further enhance employees' practical capabilities and coping skills, and improve training effects and invitation conversion rates. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0020] Figure 1 A flow chart of a training method provided by an embodiment of the present invention;
[0021] Figure 2 A detailed flow chart of step S200 provided in one embodiment of the present invention;
[0022] Figure 3 A flowchart of business accuracy scoring provided by an embodiment of the present invention;
[0023] Figure 4 A red line behavior scoring flow chart provided by an embodiment of the present invention;
[0024] Figure 5 A detailed flow chart of step S322 provided for one embodiment of the present invention;
[0025] Figure 6 A flow chart of service awareness and / or sensitive word scoring provided by an embodiment of the present invention;
[0026] Figure 7 The flowchart of process notification and / or invitation skill scoring provided by an embodiment of the present invention;
[0027] Figure 8 The detailed flowchart of step S400 provided by an embodiment of the present invention;
[0028] Figure 9 The flowchart of another training method provided by an embodiment of the present invention;
[0029] Figure 10 The schematic diagram of the training device provided by an embodiment of the present invention;
[0030] Figure 11 The schematic diagram of the structure of an electronic device provided by an embodiment of the present invention;
[0031] Figure 12 The schematic diagram of the training module provided by an embodiment of the present invention;
[0032] Figure 13 The schematic diagram of the paired practice design algorithm provided by an embodiment of the present invention.
[0033] Explanation of reference numerals: 101, data acquisition unit; 102, preprocessing unit; 103, scoring unit; 104, quality inspection information generation unit; 105, training unit; 201, processor; 202, memory; 301, courseware management sub-unit; 302, learning sub-unit; 303, practice sub-unit; 304, examination sub-unit; 305, evaluation sub-unit. Detailed implementation manners
[0034] The following uses specific specific embodiments to illustrate the implementation manners of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. It should be noted that, without conflict, the following embodiments and the features in the embodiments can be combined with each other. Except for the specific methods, devices, and materials used in the embodiments, according to the knowledge of those skilled in the art in the prior art and the description of the present invention, any methods, devices, and materials similar to or equivalent to those described in the embodiments of the present invention can also be used to implement the present invention.
[0035] It should be understood that the terms used in the embodiments of the present invention are for the purpose of describing specific specific implementation manners, rather than for limiting the protection scope of the present invention. Unless otherwise defined, all technical and scientific terms used in the present invention have the same meaning as commonly understood by those skilled in the art in this technical field.
[0036] In the following description, numerous specific details are explored to provide a more thorough explanation of embodiments of the present invention. However, it will be apparent to those skilled in the art that embodiments of the present invention can be practiced without these specific details. In some of these embodiments, well-known structures and devices are shown in block diagram form rather than in detail to avoid obscuring the embodiments of the present invention.
[0037] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of methods and computer program products that may be implemented according to various embodiments disclosed in the present invention. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a portion of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions noted in the blocks may occur in a different order than that noted in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system that performs the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.
[0038] Please refer to Figure 1 , Figure 1 A multi-modal training method based on intelligent quality inspection provided for an embodiment of the present invention includes steps S100 to S500.
[0039] Step S100: Obtain the original conversation recording between the employee to be quality-inspected and the user, as well as the lead conversion rate of the employee to be quality-inspected. The original conversation recording is the call recording when the employee to be quality-inspected communicates with the user. Taking a freight platform as an example, the employee to be quality-inspected is the invitation clerk, and the user is a person who has registered an account but has not become a driver. The invitation clerk invites these users to register as drivers by phone, and their calls will be saved. In order to quality-inspect this employee, it is first necessary to obtain these original call recordings. The lead conversion rate is an important indicator to measure the effectiveness of marketing or sales activities. It reflects the efficiency from potential customers (leads) to actual transactions (conversions), and is generally calculated directly by the ratio of the number of successfully converted leads to the total number of leads.
[0040] Step S200: Perform text conversion and cleaning on the original conversation recording to obtain the recorded text data. This step is equivalent to preprocessing, converting the original conversation recording into recorded text data that is convenient for subsequent processing.
[0041] Please refer to Figure 2, in a specific embodiment of the present invention, step S200 includes steps S201 to S202.
[0042] Step S201: Use a speech recognition model to convert the original conversation recording into text to obtain the original recording text data. There are many speech recognition models, and existing models in the prior art or self-developed speech recognition models can be used. Generally, these models need to be trained before use. In this embodiment, a trained model is directly used. Only the original conversation recording needs to be input into the trained model, and the model will automatically output the original recording text data.
[0043] Step S202: Perform hot-word error correction on the original recording text data to obtain the recording text data. There are many non-standard places in the original recording text data, resulting in poor data quality. In order to lay a good foundation for subsequent analysis, data cleaning is often required. Data cleaning generally includes removing spaces and special characters, removing stop words (such as common but insignificant words in the text: "um", "of", "is", "in", etc.), spelling correction, and hot-word error correction mentioned in step S202, etc. Hot-word error correction is for industry-specific terms and common errors. It is used to correct common pronunciation errors and misidentifications of industry-specific terms to ensure the accuracy and consistency of the recording text data. For example, when converting in step S201, there may be a situation where "Huolala" is misinterpreted as "Huolala, Huahua la la", and then term error correction needs to be performed through hot-word replacement.
[0044] Step S300: Perform quality inspection scoring on the recording text data according to a preset scoring rule to obtain a multi-dimensional score, and the multi-dimensional score is used to characterize the invitation ability of the employee to be quality inspected. The preset scoring rule can include multiple sub-rules, and the sub-rules can be called quality inspection dimensions. When presetting the scoring rule, multiple quality inspection dimensions are preset, and the corresponding quality inspection standards and deduction rules for each quality inspection dimension are defined. The following is a detailed description.
[0045] In a specific embodiment of the present invention, step S300 includes: judging whether there is a business clue text or a business clue sentence segment in the recording text data according to the recording text data and a preset business knowledge base: if so, obtaining the business accuracy score in the multi-dimensional score according to the business clues included in the recording text data and the responses of the employee to be quality inspected to the business clues; if not, the business accuracy score in the multi-dimensional score is the full score.
[0046] This embodiment targets the dimension of business accuracy. For business leads frequently consulted by users during calls, a business knowledge base can be preset. These business leads can be, for example: the refund path and time limit of the deposit, platform-related fees (such as mileage fees), membership rules, whether car stickers can be pasted, order settlement methods, and so on. In actual application, there may be many business leads. At this time, some common and important business leads can be screened out.
[0047] When scoring this dimension, first, it is necessary to determine whether the recorded text data involves these business leads. If no business leads are involved, it can be directly determined that the employee's answer is accurate, and the business accuracy score is full marks. The value of full marks can be preset, for example, it is 100 points. If business leads are included, specific scoring needs to be carried out according to the employee's reply.
[0048] Please refer to Figure 3 , in a specific embodiment of the present invention, according to the business leads included in the recorded text data and the replies of the employees to be quality inspected for the business leads, the business accuracy score in the multi-dimensional score is obtained, including: S311, using multiple different first large models to evaluate the business leads included in the text data and the replies of the employees to be quality inspected for the business leads, and obtaining the score corresponding to each first large model; S312, determining whether the number of scores greater than the preset score in the scores corresponding to each first large model exceeds the preset number: if so, the business accuracy score is full marks; if not, the business accuracy score is deducted according to the preset score.
[0049] In this embodiment, the first large model and other large models mentioned below can be implemented using common large models, such as the GPT series, BERT large model, DALL-E, etc. The reason for using multiple different first large models and calculating the business accuracy score in this way of preset score and preset number will be more accurate and comprehensive. Specifically, assuming the number of first large models is N, then step S311 will obtain N scores. In step S312, first, it is determined whether each of these N scores is greater than the preset score. If it is greater, the number of scores greater than the preset score is incremented by 1, so that the number M of scores greater than the preset score among the N scores can be obtained. Finally, M is compared with the preset number.
[0050] In this embodiment, when performing deductions, for example, this item's score can be directly deducted to zero, or other more detailed deduction logics can be adopted, such as calculating the final score according to the N scores in the previous paragraph.
[0051] Please refer to Figure 4, in a specific embodiment of the present invention, step S300 includes: S321. Using the second large model, analyze the emotion, semantics, and context information of the words of the employee to be quality-inspected in the recorded text data, and determine whether the words of the employee to be quality-inspected belong to red-line behaviors, so as to obtain the first judgment result and its confidence level; S322. According to the first judgment result and its confidence level, obtain the red-line behavior score in the multi-dimensional score.
[0052] This embodiment is aimed at the red-line behavior dimension. During the call between the employee and the user, general verbal conflicts, provocations, personal attacks, taunts, questioning and contradicting behaviors need to be avoided. To conduct quality inspection on this dimension, not only the text content in the recorded text data needs to be analyzed, but also the emotion and context information of the words should be combined to fully analyze whether the employee's words belong to red-line behaviors. Since this kind of judgment is relatively subjective, when using the second large model for judgment, not only the first judgment result (i.e., "yes" or "no") is given, but also the confidence level of this result needs to be given. Only in this way can the red-line behavior score be obtained more accurately according to the first judgment result and its confidence level.
[0053] The confidence level mentioned here refers to the credibility of the first judgment result. For example, when the results of emotion analysis and semantic understanding are highly consistent and point to red-line behaviors, and there is no contrary evidence in the conversation context, the confidence level score is relatively high; when there are multiple fuzzy factors, the confidence level score is relatively low.
[0054] Taking the freight platform of Huolala as an example, for example, the following three red-line rules can be defined. Rule 1: There is subjective impulsive behavior. When a verbal conflict occurs between the solicitor and the driver, the solicitor shows behaviors such as abusing, provoking, and personally attacking the driver; Rule 2: There is negative and passive dissemination. In the solicitor's reply, it clearly expresses a clear negative / negative evaluation of the Huolala platform. For example, in the solicitor's reply, it clearly and definitely expresses "It's difficult to make money on the Huolala platform", "There are few orders on Huolala", "The unit price on Huolala is low", etc.; Rule 3: The attitude is improper. The solicitor shows clear behaviors such as taunting, contradicting / questioning the driver, etc. For example, words such as "Why did you register" and "Do you want to join after all" appear in the solicitor's reply. These rules can all be set according to requirements. The second large model will give their respective scores and reasons based on the above-defined red-line rules, as well as the comprehensive score and reasons for whether the solicitor violates the red line.
[0055] Please refer to Figure 5 , in a specific embodiment of the present invention, step S322 includes steps S3221 to S3223.
[0056] Step S3221: If the first judgment result is that all the words of the employee to be quality-inspected do not belong to red-line behaviors, the score for red-line behaviors is full marks. In this step, all the words of the employee are normal and do not trigger red-line behaviors, so the score for red-line behaviors is full marks.
[0057] Step S3222: If the first judgment result is that any one of the words of the employee to be quality-inspected belongs to a red-line behavior and its confidence level is higher than the preset confidence level, the score for red-line behaviors is deducted according to the preset score. In this step, because the confidence level is relatively high, it can be considered that the first judgment result is accurate. Therefore, it can be considered that the employee has indeed triggered a red-line behavior and the score can be directly deducted. The specific amount of deduction can be preset or directly deducted to zero.
[0058] Step S3223: If the first judgment result is that any one of the words of the employee to be quality-inspected belongs to a red-line behavior and its confidence level is lower than the preset confidence level, manual review is triggered. In this step, although the first judgment result indicates that the employee has triggered a red-line behavior, considering the relatively low confidence level, if the score is directly deducted, there will be a misjudgment. Therefore, manual review can be triggered and it is up to the manual review to confirm whether to deduct the score.
[0059] Please refer to Figure 6 , in a specific embodiment of the present invention, step S300 includes: S331, determining whether the beginning part of the recorded text data includes a preset opening statement to obtain a second judgment result, and / or determining whether the recorded text data includes a preset sensitive word to obtain a third judgment result; S332, obtaining the service awareness score in the multi-dimensional score according to the second judgment result, and / or obtaining the sensitive word score in the multi-dimensional score according to the third judgment result.
[0060] In this embodiment, the service awareness dimension and the sensitive word dimension are included. Specifically, it includes three scenarios: Scenario 1, only including the service awareness score; Scenario 2, only including the sensitive word score; Scenario 3, including both the service awareness score and the sensitive word score.
[0061] In this embodiment, the service awareness dimension is mainly used to determine whether the employee includes a standard opening statement at the beginning of the call, such as indicating the identity of the Huolala staff, etc. In specific determination, it is determined according to whether the beginning part of the recorded text data includes a preset opening statement. The preset opening statement can be, for example, a sentence or some words. When calculating the score specifically, if the preset opening statement is included, the service awareness score is full marks; if the preset opening statement is not included, the service awareness score is zero.
[0062] In this embodiment, the sensitive word dimension is mainly used to determine whether an employee mentions sensitive words during a call. Taking the Huolala freight scenario as an example, to prevent employees from adding drivers through personal WeChat, "WeChat" can be set as a sensitive word. Sometimes, there are more complex logical judgments between sensitive words. Taking "WeChat" as an example, if an employee mentions "WeChat" alone, it will be regarded as a sensitive word, but if the employee mentions "Enterprise WeChat", it can be considered that no sensitive word is mentioned. When calculating the score specifically, if a sensitive word is mentioned, the score for the sensitive word is zero; if no sensitive word is submitted, the score for the sensitive word is full marks.
[0063] Please refer to Figure 7 , in a specific embodiment of the present invention, step S300 further includes steps S341 to S343.
[0064] Step S341: Use the third large model to process the recorded text data to obtain the driver invitation result. Here, it is mainly to judge whether the driver accepts the invitation, and the result can be accepting the invitation or rejecting the invitation.
[0065] Step S342: When the driver invitation result is accepting the invitation, further judge whether the recorded text data includes document reminder information to obtain a fourth judgment result; and / or when the driver invitation result is rejecting the invitation, further judge whether the recorded text data includes persuasion information to obtain a fifth judgment result.
[0066] Step S343: Obtain the process notification score in the multi-dimensional score according to the fourth judgment result, and / or obtain the invitation skill score in the multi-dimensional score according to the fifth judgment result.
[0067] In this embodiment, it includes a process notification dimension and an invitation skill dimension, which also include three scenarios: Scenario 1, only includes the process notification score; Scenario 2, only includes the invitation skill score; Scenario 3, includes both the process notification score and the invitation skill score.
[0068] In this embodiment, the process notification dimension mainly focuses on whether the subsequent processing process is accurate when the driver clearly accepts the invitation. For example, when a registered user becomes a truck driver and needs to carry 5 documents (ID card, driver's license, vehicle license, annual inspection, compulsory traffic insurance), it is necessary to judge whether the employee's document reminder information completely includes these document information. In addition, it can also be judged whether the employee communicates with the driver about the invitation time. When calculating the score specifically, complete process notification can correspond to full marks, and incomplete process notification (such as missing some information) can correspond to zero points.
[0069] In this embodiment, the invitation skill dimension mainly focuses on whether the employee takes corresponding measures after the driver refuses the invitation. For example, after the user clearly refuses, whether the employee tries to persuade the user to stay, or when the user says they need to consider, whether the employee continues to insist on inviting the driver to join, etc. This dimension mainly emphasizes the ability to actively ask questions when the user refuses and flexibly dig deeper and persuade based on the user's questions / needs scenarios. When calculating the score specifically, if the employee tries to persuade the user to stay, it can correspond to the full score; if not, it can correspond to zero points.
[0070] The full score values of the above multiple dimensions can be set according to the importance of each dimension. Taking the above dimensions as an example, in a specific embodiment of the present invention, the full score of the service awareness score is 10 points, and if it is not satisfied, it is 0 points; the full score of the sensitive word score is 100 points, and if it is not satisfied, it is 0 points; the full score of the process notification score is 100 points, and if it is not satisfied, it is 0 points; the full score of the business accuracy score is 100 points, and if it is not satisfied, it is 0 points; the full score of the invitation skill score is 20 points, and if it is not satisfied, it is 0 points; the full score of the red line behavior score is 100 points, and if it is not satisfied, it is 0 points. Of course, it can also be set to other values, and the values here are for reference only.
[0071] Step S400, when the lead conversion rate and the multi-dimensional scores meet the quality inspection conditions, generate the quality inspection information data corresponding to the employee to be quality inspected. The quality inspection conditions mentioned in this step, for example, are the working period of the employee, because if the working period of the employee to be quality inspected is short, its lead conversion rate is not representative and fluctuates greatly. At this time, even if the quality inspection information data is generated, it is inaccurate. Therefore, only when the quality inspection conditions are met can the quality inspection information data be generated.
[0072] Please refer to Figure 8 , in a specific embodiment of the present invention, step S400 includes step S401 and step S402.
[0073] Step S401, according to the lead conversion rate and the preset mapping relationship between the lead conversion rate and its score, obtain the lead conversion rate score. The lead conversion rate is a value between 0 and 1. The mapping relationship between the lead conversion rate and its score can be preset. For example, when the lead conversion rate is in the range [0, 5%), it corresponds to zero points; when the lead conversion rate is in the range [5%, 10%), it corresponds to 50 points; when the lead conversion rate is greater than or equal to 10%, it corresponds to the full score. With the mapping relationship, the lead conversion rate score can be calculated.
[0074] Step S402, according to the multi-dimensional score and the lead conversion rate score, obtain the final score of the employee to be quality inspected. The final score of the employee to be quality inspected includes two parts: the multi-dimensional score and the lead conversion rate score. Weighting or summing these two parts and other operations are performed to obtain the comprehensive ability score of each employee finally.
[0075] In a specific embodiment of the present invention, the final score can be calculated according to the following formula, for example:
[0076]
[0077] In the formula, score is the final score, α is the weight hyperparameter, D represents the number of dimensions, score i represents the score of the i-th dimension, which can be, for example, the above-mentioned business accuracy score, red line behavior score, service awareness score, sensitive word score, process notification score, invitation skill score, and C is the lead conversion rate score.
[0078] Step S500: Output a text or audio / video training task according to the quality inspection information data corresponding to the employee to be quality inspected. After obtaining the quality inspection information data corresponding to the employee to be quality inspected, appropriate training tasks can be generated for the employee according to some preset rules to improve their invitation skills and capabilities.
[0079] It should be noted that the step division of the above various methods is only for clear description. When implemented, they can be combined into one step or some steps can be split into multiple steps. As long as they contain the same logical relationship, they are all within the protection scope of this patent; adding insignificant modifications or introducing insignificant designs to the algorithm or process, but not changing the core design of its algorithm and process are all within the protection scope of this patent. For example Figure 9 shows the flowchart of another training method.
[0080] Please refer to Figure 10 , Figure 10 which is a multi-modal training device based on intelligent quality inspection provided by an embodiment of the present invention, including a data acquisition unit 101, a preprocessing unit 102, a scoring unit 103, a quality inspection information generation unit 104, and a training unit 105. Among them, the data acquisition unit 101 is used to acquire the original conversation recording between the employee to be quality inspected and the user, as well as the lead conversion rate of the employee to be quality inspected; the preprocessing unit 102 is used to perform text conversion and cleaning on the original conversation recording to obtain the recording text data; the scoring unit 103 is used to process the recording text data according to the preset scoring rules to obtain multi-dimensional scores, and the multi-dimensional scores are used to characterize the invitation ability of the employee to be quality inspected; the quality inspection information generation unit 104 is used to generate the quality inspection information data corresponding to the employee to be quality inspected when the lead conversion rate and the multi-dimensional scores meet the quality inspection conditions; the training unit 105 is used to output a text or audio / video training task according to the quality inspection information data corresponding to the employee to be quality inspected.
[0081] It should be noted that the training device in this embodiment is a device corresponding to the above-mentioned training method, and the functional modules in the training device respectively correspond to the corresponding steps in the training method. The training device in this embodiment can be implemented in cooperation with the training method. That is, without conflict, the relevant technical details mentioned in the training method of the above embodiment can also be applied to the training device in this embodiment.
[0082] Please refer to Figure 11 , Figure 11 which is an electronic device provided by an embodiment of the present invention, including a processor 201, a memory 202, and a communication bus; the communication bus is used to connect the processor 201 and the memory 202; the processor 201 is used to execute the computer program stored in the memory 202 to implement the above-mentioned multi-modal training method based on intelligent quality inspection.
[0083] The above-mentioned electronic device is a device that can automatically perform numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes, but is not limited to, a microprocessor, an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a digital signal processor (DSP), an embedded device, etc.
[0084] The above-mentioned electronic device can be any electronic product that can perform human-computer interaction with users. For example, a personal computer, a tablet computer, a smart phone, a personal digital assistant (PDA), a game console, an Internet Protocol Television (IPTV), a smart wearable device, etc.
[0085] The above-mentioned electronic device may further include a network device and / or a user device. Among them, the network device includes, but is not limited to, a single network server, a server group composed of multiple network servers, or a cloud composed of a large number of hosts or network servers based on cloud computing (Cloud Computing).
[0086] The network where the above-mentioned electronic device is located includes, but is not limited to, the Internet, a wide area network, a metropolitan area network, a local area network, a virtual private network (VPN), etc.
[0087] The above-mentioned processor 201 can be, for example, a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it can also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components; the above-mentioned memory may include a Random Access Memory (RAM), and may also include a non-volatile memory, such as at least one disk memory.
[0088] An embodiment of the present invention also provides a computer-readable storage medium, on which a computer program is stored, and the computer program is used to make a computer execute the above-mentioned multi-modal training method based on intelligent quality inspection.
[0089] Please refer to Figure 12 , in a specific embodiment of the present invention, the training unit 105 includes a courseware management subunit 301, a learning subunit 302, a practice subunit 303, an exam subunit 304, and an evaluation subunit 305.
[0090] The courseware management subunit 301 is used to manage the courseware available for employee training. It supports online management of training courseware, provides online courseware tools for convenient unified management, and its main functions include: (1) Courseware import: supports uploading of courseware files in multiple formats, such as PPT, PDF, video, etc.; (2) Online editing: provides an online editing function, allowing direct modification and update of the courseware content on the platform; (3) Preview function: after the courseware is edited, the courseware effect can be previewed online to ensure the accuracy of the content and the display effect; (4) Version control: automatically saves different versions of the courseware, facilitating backtracking and management of historical records; (5) Permission management: sets different access and editing permissions to ensure the security and confidentiality of the courseware.
[0091] The learning subunit 302 is used for employees to learn courseware and record their learning information. It provides an online learning platform with a user-friendly interface for employees; supports online learning of different types of courseware, enabling employees to equally share training resources; can automatically record employees' learning progress, facilitating employees to view and continue learning at any time, and also facilitating administrators to track the learning progress of all employees. The learning subunit 302 allows employees to actively select courseware for learning or automatically select courseware according to training tasks.
[0092] The practice subunit 303 is used to provide a co-practice robot to have conversations with employees and generate reply suggestions, and record employees' practice information. The co-practice robot can practice specified conversations with employees according to training tasks. It uses the latest AI technology to build a co-practice robot to conduct immersive co-practice with the invitation staff, providing one-on-one accompaniment simulation drills and business guidance services. Online practice currently supports two human-computer interaction modes: voice mode and digital human mode. The design algorithm scheme for the voice and digital human mode co-practice is as Figure 13 shown, where the algorithm process includes 5 process steps: (1) Real-time ASR transcription: The voice input of the customer service or robot is converted into text through the real-time ASR module; (2) Hot word replacement: Replace the problematic words that appear in the ASR transcription result with hot words to improve the accuracy and readability of the text; (3) Vector retrieval: Use the processed text as a query to perform vector retrieval in the corpus to find Q&A pairs (QA pairs) similar to the historical knowledge conversation; (4) Generate reply by large model: According to the retrieved QA pairs, the preset prompt, and the historical conversation memory, the large model generates reply suggestions; (5) Real-time TTS conversion: Convert the text reply generated by the large model into voice through the real-time TTS module and return it to the user.
[0093] In addition, during the use process, employees continuously try and discover problems and submit the feedback to the model, and the system continuously optimizes according to these feedbacks to improve the accuracy and response speed of the model.
[0094] When constructing the evaluation metrics for the dialogue large model, since the large model does not have clear metrics such as AUC like classification or regression models, the present invention combines the general metrics of the dialogue large model and the comprehensive business evaluation metrics to achieve a comprehensive evaluation. The general metrics include but are not limited to semantic consistency, non-repetitiveness, logical coherence, language naturalness, and robustness to extreme questions. For the customer service training robot business, in addition to the basic large model capabilities, the present invention particularly focuses on the following two key metrics: one is process integrity, that is, whether the employee follows the standard opening and closing remarks and whether necessary information such as vehicle type and vehicle owner is confirmed with the driver; the other is assessment perseverance, that is, when the driver has concerns and is reluctant to register, whether the customer service can use effective needs-digging and retention skills to successfully convert the driver into a platform driver. Through these comprehensive evaluation metrics, the present invention can more comprehensively assess and improve the service capabilities of the customer service.
[0095] The examination sub-unit 304 is used to provide corresponding theoretical questions and scenario simulation questions according to the preset level of the employee to conduct an assessment of the employee to obtain the examination results. Theoretical questions can be understood as non-scenario simulation assessments, mainly examining the business knowledge of the inviter in the form of test questions. The scenario simulation assessment can be human-machine simulation clearance, mainly examining the inviter's speech skills in specific invitation scenarios in the form of simulated duels.
[0096] Question bank design: The training questions are divided into non-scenario simulation questions and scenario simulation questions according to the classification of the examination. Among them, the non-scenario simulation questions mainly include single-choice questions, multiple-choice questions, true-false questions, and short-answer questions; the scenario simulation questions are mainly in the form of human-machine multi-round dialogues, and the specific implementation forms are voice and digital human.
[0097] In this embodiment, the reason for making the preset level division of employees is mainly to distinguish different types of employees and design different simulation questions for different types of employees: (1) New employee scenario: Focus on the examination of basic knowledge, mainly covering the basic business processes of Lalamove, common question answers, etc., aiming to ensure that the inviter can master the basic business knowledge proficiently and lay a solid foundation for subsequent complex scenarios. (2) Experienced employee scenario: Focus on invitation skills (such as business scenarios during peak hours, cross-regional, etc.) and the handling of difficult questions. The question types are mainly complex invitation scenarios, customer objection handling, special situation response, etc. (3) Problem employee scenario: Quality inspection can screen out some problem employees, generate special tasks for the dimensions with lack of invitation ability and push relevant courses. After completing the special courses, the inviter will be scored based on the dialogue process and the personal ability of the inviter will be updated.
[0098] The evaluation subunit 305 is used to obtain the learning report of the employee based on the learning information, practice information and test results. Through data analysis and intelligent algorithms, the learning effect of the employees is comprehensively evaluated, and targeted training is carried out for employees with weak foundation in certain dimensions. The main functions include: (1) Learning data analysis: Collect and analyze the learning data of employees, such as learning time, practice results, test results, etc., and generate multi-dimensional learning reports. These belong to quantitative data collection and analysis. Non-quantitative data collection can also be performed. For example, during the learning process, the employee's emotional and attitude information can be collected in various ways. Combine quantitative and non-quantitative data to generate a multi-dimensional learning report. The report content includes not only conventional indicators such as learning results and time, but also emotional and attitude analysis results. (2) Personalized suggestions: According to the learning situation of the employees, personalized learning suggestions and improvement plans are provided to help employees improve their learning effects; for employees who have positive emotional attitudes but slow performance improvement during the learning process, it can also be analyzed whether there are problems with their learning methods; for students who are somewhat negative in emotion but have certain learning potential, incentive and guidance strategies can be formulated. (3) Provide assessment tools to help managers understand the overall performance of employees and make targeted teaching adjustments.
[0099] The present invention proposes a multimodal training method based on intelligent quality inspection, which is mainly aimed at the quality inspection and training of outbound calls of freight inviters of large and small vehicles. It can also be reused in other business lines such as small trucks, moving, etc., and other business scenarios such as customer service, enterprises and automobile sales and other related businesses.
[0100] In general, the present invention uses a large model to perform multi-dimensional intelligent scoring of call data, including voice quality, speaking speed, emotional expression, speech standards and other aspects, to comprehensively evaluate the comprehensive ability of the inviter. The scoring results are visualized through radar charts to intuitively display the performance of each dimension, helping managers and employees to fully understand the scoring results. For low-scoring tail employees and new employees, 1v1 training tasks are automatically generated according to the scoring results, such as digital human sparring, to provide personalized training content to help employees quickly improve their invitation capabilities. In addition, the present invention also includes a scenario simulation question bank design, which provides simulation exercises for a variety of real scenarios to further enhance employees' practical capabilities and coping skills. Through a real-time feedback mechanism, the present invention ensures timely adjustment of training content and strategies, and improves training effectiveness and invitation conversion rates.
[0101] The above embodiments are merely illustrative of the principles and effects of the present invention, and are not intended to limit the present invention. Anyone familiar with the art may modify or alter the above embodiments without departing from the spirit and scope of the present invention. Therefore, all equivalent modifications or alterations made by a person of ordinary skill in the art without departing from the spirit and technical concept disclosed by the present invention shall still be covered by the claims of the present invention.
Claims
1. A multimodal training method based on intelligent quality inspection, characterized in that: include: Obtaining the original conversation recording between the employee to be inspected and the user, as well as the lead conversion rate of the employee to be inspected; Converting and cleaning the original conversation recording to text to obtain recording text data; Performing quality inspection and scoring on the audio recording text data according to a preset scoring rule to obtain a multi-dimensional score, wherein the multi-dimensional score is used to characterize the invitation ability of the employee to be inspected; When the lead conversion rate and the multi-dimensional score meet the quality inspection conditions, generating quality inspection information data corresponding to the employee to be inspected; Output text or audio and video training tasks based on the quality inspection information data corresponding to the employees to be inspected.
2. The multimodal training method based on intelligent quality inspection according to claim 1 is characterized in that: The original conversation recording is converted and cleaned to obtain recording text data, including: Using a speech recognition model to convert the original conversation recording into text to obtain original recording text data; Hot word error correction is performed on the original audio recording text data to obtain the audio recording text data.
3. The multimodal training method based on intelligent quality inspection according to claim 1 is characterized in that: The audio recording data is quality-checked and scored according to the preset scoring rules to obtain multi-dimensional scores, including: According to the recorded text data and a preset business knowledge base, it is determined whether there is a business clue text or a business clue sentence segment in the recorded text data: If yes, then obtaining the business accuracy score in the multi-dimensional score according to the business clues contained in the audio recording text data and the response of the employee to be inspected to the business clues; If not, the business accuracy score in the multi-dimensional score is full marks.
4. The multimodal training method based on intelligent quality inspection according to claim 3 is characterized in that: According to the business clues contained in the recorded text data and the responses of the employees to be inspected to the business clues, the business accuracy score in the multi-dimensional score is obtained, including: Using a plurality of different first models, the business clues contained in the audio recording text data and the responses of the employees to be inspected to the business clues are evaluated to obtain scores corresponding to each of the first models; Determine whether the number of scores corresponding to each of the first largest models that are greater than a preset score exceeds a preset number: If yes, the business accuracy score is full marks; If not, the business accuracy score will be deducted according to the preset score.
5. The multimodal training method based on intelligent quality inspection according to claim 1 is characterized in that: The audio recording data is quality-checked and scored according to the preset scoring rules to obtain multi-dimensional scores, including: Using the second largest model, the emotion, semantics, and context information of the speech of the employee to be inspected in the audio recording text data are analyzed to determine whether the speech of the employee to be inspected belongs to the red line behavior, so as to obtain a first judgment result and its confidence level; According to the first judgment result and its confidence level, a red line behavior score in the multi-dimensional score is obtained.
6. The multimodal training method based on intelligent quality inspection according to claim 5 is characterized in that: According to the first judgment result and its confidence level, a red line behavior score in the multi-dimensional score is obtained, including: If the first judgment result is that all the words of the employee to be inspected do not belong to red line behaviors, then the red line behavior score is full marks; If the first judgment result is that any speech of the employee to be inspected belongs to a red line behavior and its confidence is higher than a preset confidence, then the red line behavior score is deducted according to a preset score; If the first judgment result is that any speech of the employee to be inspected is a red line behavior and its confidence level is lower than a preset confidence level, a manual review is triggered.
7. The multimodal training method based on intelligent quality inspection according to claim 1 is characterized in that: The audio recording data is quality-checked and scored according to the preset scoring rules to obtain multi-dimensional scores, including: Determine whether the beginning of the recorded text data includes a preset opening word to obtain a second determination result, and / or determine whether the recorded text data includes a preset sensitive word to obtain a third determination result; The service awareness score in the multi-dimensional score is obtained according to the second judgment result, and / or the sensitive word score in the multi-dimensional score is obtained according to the third judgment result.
8. The multimodal training method based on intelligent quality inspection according to claim 1 is characterized in that: The audio recording data is quality-checked and scored according to the preset scoring rules to obtain multi-dimensional scores, including: The recording text data is processed using the third model to obtain a driver invitation result; When the driver invitation result is to accept the invitation, further determine whether the recorded text data includes the certificate reminder information to obtain a fourth determination result; and / or when the driver invitation result is to reject the invitation, further determine whether the recorded text data includes the retention information to obtain a fifth determination result; The process notification score in the multi-dimensional score is obtained according to the fourth judgment result, and / or the invitation skill score in the multi-dimensional score is obtained according to the fifth judgment result.
9. The multimodal training method based on intelligent quality inspection according to claim 1 is characterized in that: Generating quality inspection information data corresponding to the employee to be inspected, including: Obtaining a lead conversion rate score according to the lead conversion rate and a preset mapping relationship between the lead conversion rate and its score; The final score of the employee to be quality inspected is obtained according to the multi-dimensional score and the lead conversion rate score.
10. A multimodal training device based on intelligent quality inspection, characterized in that: include: A data acquisition unit, used to acquire the original conversation recording between the employee to be inspected and the user, and the lead conversion rate of the employee to be inspected; A preprocessing unit, used for performing text conversion and cleaning on the original conversation recording to obtain recording text data; A scoring unit, used to perform quality inspection and scoring on the audio recording text data according to a preset scoring rule to obtain a multi-dimensional score, wherein the multi-dimensional score is used to characterize the invitation ability of the employee to be inspected; A quality inspection information generating unit, configured to generate quality inspection information data corresponding to the employee to be inspected when the lead conversion rate and the multi-dimensional score meet the quality inspection conditions; The training unit is used to output text or audio and video training tasks according to the quality inspection information data corresponding to the employees to be inspected.
11. An electronic device, characterized in that: The system comprises a processor, a memory and a communication bus; the communication bus is used to connect the processor and the memory; the processor is used to execute a computer program stored in the memory to implement the method according to any one of claims 1 to 9.
12. A computer-readable storage medium, characterized in that: A computer program is stored thereon, and the computer program is used to make a computer execute the method according to any one of claims 1 to 9.