Voice-based tutoring data analysis methods, devices, computer equipment, and storage media
By conducting multi-dimensional analysis of voice-guided learning data, the problem of limited analytical dimensions in existing technologies has been solved, enabling accurate assessment of users' mastery of knowledge points and improving the quality of learning and teaching.
Patent Information
- Application Number
- CN202210734757.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-27
- Publication Date
- 2025-10-31
- Estimated Expiration
- 2042-06-27
AI Technical Summary
Existing voice-based tutoring systems rely on a single analytical dimension, failing to accurately reflect the user's mastery of knowledge points, resulting in a tedious learning process and an inability to effectively measure teaching quality.
By acquiring the voice data of the practice partner, performing voice recognition and text parsing, and acquiring the text data of the practice partner, we conduct proficiency, similarity, and accuracy analyses to comprehensively evaluate the user's mastery of the knowledge points.
It enables accurate assessment of users' mastery of knowledge points, improves the quantitative assessment capability of learning and teaching quality, and enhances the effectiveness of the training system.
Smart Images

Figure CN115116445B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of speech processing technology, and in particular to a method, apparatus, computer equipment, and storage medium for analyzing speech training data. Background Technology
[0002] Currently, the training system generally uses both online and offline training methods. Regardless of whether it is online or offline, the main focus is on instructors explaining the knowledge points, with a lack of opportunities for students to conduct simulated exercises. This makes the learning process tedious and uninteresting, preventing students from applying what they have learned and deepening their understanding of the knowledge points. Furthermore, it makes it impossible to effectively measure whether students have mastered the knowledge points, which is detrimental to ensuring the teaching quality of instructors and the learning quality of students.
[0003] Existing voice-guided practice systems primarily use client-side applications to publish business scenarios and guiding questions related to business knowledge, enabling users to familiarize themselves with the scripts for various business scenarios and better serve customers. These systems analyze user responses to guiding questions using natural language processing techniques to determine if the responses cover necessary semantic points, thus assigning a practice score. However, this analysis method is relatively singular and cannot accurately reflect the user's level of mastery of the knowledge points. Summary of the Invention
[0004] This invention provides a method, apparatus, computer device, and storage medium for analyzing voice-guided training data, in order to solve the problem that existing voice-guided training data analysis has a single dimension and cannot accurately reflect the user's mastery of knowledge points.
[0005] A method for analyzing voice-based tutoring data includes:
[0006] Obtain the voice data of the coach;
[0007] The training voice data is subjected to speech recognition and text parsing to obtain training text data, which includes scene text data corresponding to N training scenarios, where N≥1;
[0008] Perform proficiency analysis on the scene text data corresponding to each training scenario to obtain the proficiency score corresponding to the training scenario.
[0009] A similarity analysis is performed on the scene text data corresponding to each training scenario and the standard text data corresponding to the training scenario to obtain the similarity score corresponding to the training scenario.
[0010] Accuracy analysis is performed on the scene text data and standard text data corresponding to each training scenario to obtain the accuracy score corresponding to the training scenario.
[0011] A comprehensive analysis is performed on the proficiency score, similarity score, and accuracy score corresponding to the N training scenarios to obtain the target training result.
[0012] A voice-guided training data analysis device, comprising:
[0013] The training voice data acquisition module is used to acquire training voice data;
[0014] The training text data acquisition module is used to perform speech recognition and text parsing on the training voice data to acquire training text data, which includes scene text data corresponding to N training scenarios, where N≥1.
[0015] The proficiency score acquisition module is used to perform proficiency analysis on the scene text data corresponding to each training scenario and acquire the proficiency score corresponding to the training scenario.
[0016] The similarity score acquisition module is used to perform similarity analysis on the scene text data corresponding to each training scenario and the standard text data corresponding to the training scenario, and to obtain the similarity score corresponding to the training scenario.
[0017] The accuracy score acquisition module is used to perform accuracy analysis on the scene text data and the standard text data corresponding to each training scenario to obtain the accuracy score corresponding to the training scenario.
[0018] The target coaching result acquisition module is used to comprehensively analyze the proficiency score, similarity score, and accuracy score corresponding to the N coaching scenarios to obtain the target coaching result.
[0019] A computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the above-described voice training data analysis method.
[0020] A computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described voice training data analysis method.
[0021] The aforementioned voice-guided training data analysis method, device, computer equipment, and storage medium perform speech recognition and parsing on the training voice data to determine the scene text data corresponding to N training scenarios; analyze the proficiency score, similarity score, and accuracy analysis corresponding to each training scenario individually to determine the user's mastery of the knowledge points corresponding to a single training scenario; finally, conduct a comprehensive analysis of the proficiency score, similarity score, and accuracy score corresponding to the N training scenarios to obtain the target training result, ensuring the accuracy of the target training result, so as to quantitatively evaluate the user's mastery of all knowledge points involved in the N training scenarios based on the target training result. Attached Figure Description
[0022] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments of the present invention will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0023] Figure 1 This is a schematic diagram of an application environment for the voice tutoring data analysis method in one embodiment of the present invention;
[0024] Figure 2 This is a flowchart of a voice tutoring data analysis method according to an embodiment of the present invention;
[0025] Figure 3 This is another flowchart of the voice tutoring data analysis method in one embodiment of the present invention;
[0026] Figure 4 This is another flowchart of the voice tutoring data analysis method in one embodiment of the present invention;
[0027] Figure 5 This is another flowchart of the voice tutoring data analysis method in one embodiment of the present invention;
[0028] Figure 6 This is another flowchart of the voice tutoring data analysis method in one embodiment of the present invention;
[0029] Figure 7 This is another flowchart of the voice tutoring data analysis method in one embodiment of the present invention;
[0030] Figure 8 This is another flowchart of the voice tutoring data analysis method in one embodiment of the present invention;
[0031] Figure 9 This is a schematic diagram of a voice tutoring data analysis device according to an embodiment of the present invention;
[0032] Figure 10 This is a schematic diagram of a computer device according to an embodiment of the present invention. Detailed Implementation
[0033] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0034] The voice training data analysis method provided in this embodiment of the invention can be applied to, for example... Figure 1 The application environment is shown. Specifically, this voice tutoring data analysis method is applied in a voice tutoring system, which includes, for example, […]. Figure 1 The client and server shown communicate over a network to comprehensively analyze the practice audio data from dimensions such as proficiency, similarity, and accuracy. This analysis determines the user's mastery of the knowledge points, effectively measuring whether the user has grasped the material. This allows for a quantitative assessment of the instructor's teaching quality and the student's learning quality, contributing to the further improvement of the training system. The client, also known as the user terminal, is the program that provides local services to the client, corresponding to the server. The client can be installed on, but is not limited to, various personal computers, laptops, smartphones, tablets, and portable wearable devices. The server can be implemented using a standalone server or a server cluster consisting of multiple servers.
[0035] In one embodiment, such as Figure 2 As shown, a data analysis method for voice tutoring is provided, and this method is applied to... Figure 1 Taking the server in the example, the following steps are included:
[0036] S201: Obtain training voice data;
[0037] S202: Perform speech recognition and text parsing on the training voice data to obtain training text data. The training text data includes scene text data corresponding to N training scenarios, where N≥1.
[0038] S203: Perform proficiency analysis on the scene text data corresponding to each training scenario and obtain the proficiency score corresponding to the training scenario.
[0039] S204: Perform similarity analysis on the scene text data and the standard text data corresponding to each training scenario to obtain the similarity score for the training scenario.
[0040] S205: Perform accuracy analysis on the scenario text data and standard text data corresponding to each training scenario to obtain the accuracy score for the training scenario.
[0041] S206: Perform a comprehensive analysis of the proficiency score, similarity score, and accuracy score corresponding to N coaching scenarios to obtain the target coaching result.
[0042] Among them, the voice data of the coaching session refers to the voice data formed by the user based on the guidance of the target coaching script.
[0043] As an example, in step S201, the server can control the client to play guidance questions related to the knowledge point using voice, so as to receive the practice voice data of the user answering the guidance questions on the knowledge point.
[0044] In this example, the server pre-creates and stores multiple design practice scripts based on the business experience of senior business personnel. When a user engages in voice practice, they can select one of these scripts as a target script. The client then plays the guiding questions corresponding to the target script and receives the user's voice responses to these questions. In this example, the target script is a question-and-answer script, which can be constructed based on multiple guiding questions corresponding to a specific knowledge point. For example, if a specific knowledge point contains multiple guiding questions Q1, Q2, and Q3, the target script can integrate these questions into a single question-and-answer script using a specific format. The target script is the script selected by the user for voice practice.
[0045] For example, before step S201, the process of pre-creating a training script using the voice training data analysis method includes: acquiring a business knowledge file; determining business keywords, target products, and business experience tags based on the business knowledge file; processing the business keywords, target products, and business experience tags using a conversation template to form a training script, wherein the training script includes standard text data corresponding to N training scenarios; and storing the training script and the standard text data corresponding to the N training scenarios in the system database.
[0046] The business knowledge document is used to record the knowledge points that need to be practiced. For the insurance sector, the business knowledge document is a document containing knowledge points related to the insurance field. Business keywords are keywords extracted from the business knowledge document that are relevant to specific business operations. Target products refer to the specific products addressed in the business knowledge document. Business experience tags are tags representing the operational experience of sales personnel, reflecting that specific content in the business knowledge document is a summary of the sales personnel's operational experience.
[0047] In one example, during insurance business training, it's necessary to combine the communication scripts of sales personnel with written records of conversations by senior sales staff. The business knowledge points mentioned in these records are then tagged with business experience labels. For instance, during the design and creation of training scripts, senior sales staff are organized to analyze the data. Based on this analysis, corresponding business knowledge documents are generated. By processing these documents, business keywords, target products, and business experience tags can be identified. A conversation template is then used to process these keywords, target products, and business experience tags to create training scripts. Each training script includes standard text data corresponding to N training scenarios. These training scripts and the standard text data for each scenario are stored in the system database. This allows other users to select a target training script from multiple scripts and then conduct voice training based on the guiding questions for each scenario within that script. In this example, the training script will contain N training scenarios, and each training scenario can contain multiple conversations. The conversations here can be understood as a question-and-answer communication process, involving the business experience of senior business personnel on when to start a topic, change the topic, and delve into the topic, so as to guide other users to perform corresponding operations through the training script design.
[0048] The accompanying text data refers to the text data corresponding to the accompanying voice data. The accompanying scenario is a pre-set scenario for guiding questions to highlight knowledge points; each accompanying scenario corresponds to at least one guiding question. In this example, during voice accompanying practice, the user can be guided based on the guiding questions corresponding to N accompanying scenarios, so that the resulting accompanying text data includes the scenario text data corresponding to N accompanying scenarios, where N is the number of accompanying scenarios, specifically an integer greater than or equal to 1. The scenario text data is the text data in which the user answers the guiding question of a specific accompanying scenario; that is, it is the text data formed by parsing the user's voice response to the guiding question of the accompanying scenario.
[0049] As an example, in step S202, after the server obtains the training voice data, it uses a voice recognition tool to perform voice recognition and text parsing on the training voice data, and determines the text data output by the voice recognition tool as the training text data, which includes scene text data corresponding to N training scenarios.
[0050] In this example, when providing coaching based on the target coaching script selected by the user, since the target coaching script has N pre-set coaching scenarios, each coaching scenario corresponds to a guidance question, and since the coaching voice data is voice data formed for the target coaching script, the coaching text data determined by its parsing is also related to the target coaching script, including the coaching text data of the guidance questions corresponding to the N coaching scenarios.
[0051] As an example, in step S203, after obtaining the scenario text data corresponding to N coaching scenarios, the server can use pre-set proficiency analysis logic to perform proficiency analysis on the scenario text data corresponding to each coaching scenario. This analysis aims to determine the user's proficiency with the knowledge points contained in the guidance questions for that coaching scenario, thereby obtaining a proficiency score for each coaching scenario. The proficiency analysis logic here is a pre-set logic used to analyze the user's proficiency with the knowledge points. For example, the proficiency score can be determined by analyzing the number of pauses, response time, or other information the user uses when answering guidance questions in a coaching scenario.
[0052] The standard text data consists of the standard answers to a pre-set coaching scenario.
[0053] As an example, in step S204, after obtaining the scenario text data corresponding to N coaching scenarios, the server can first obtain the standard text data corresponding to the coaching scenarios from the system database. Then, it uses pre-set similarity analysis logic to perform similarity analysis on the scenario text data and standard text data corresponding to each coaching scenario, thereby obtaining a similarity score for each coaching scenario. The similarity analysis logic here is a pre-set logic used to analyze the degree of similarity between the scenario text data and the standard text data. For example, the similarity analysis logic can be a processing logic constructed based on existing similarity algorithms. The similarity score determined by the similarity analysis logic can reflect the degree of similarity between the user's answer to the knowledge point corresponding to the guided question and the standard answer.
[0054] As an example, in step S205, after obtaining the scenario text data corresponding to N training scenarios, the server can first retrieve the standard text data corresponding to the training scenarios from the system database. Then, it uses a pre-set accuracy analysis logic to perform accuracy analysis on the scenario text data and standard text data corresponding to each training scenario, thereby obtaining the accuracy score for each training scenario. The accuracy analysis logic here is a pre-set logic that determines the accuracy of the scenario text data relative to the standard text data by analyzing the degree of matching between the keywords in the scenario text data and the keywords in the standard text data. For example, the accuracy analysis logic can determine the accuracy of the user's understanding of the knowledge points corresponding to the guided questions by analyzing the degree of matching between the keywords in the scenario text data and the keywords in the standard text data.
[0055] As an example, in step S206, after obtaining the proficiency scores, similarity scores, and accuracy scores corresponding to N tutoring scenarios, the server can perform a comprehensive analysis of these scores based on pre-set comprehensive analysis logic to determine the target tutoring result. The comprehensive analysis logic is a pre-set logic used to comprehensively analyze the user's learning quality of the knowledge points involved in the N tutoring scenarios. For example, by executing the comprehensive analysis logic, the server can perform a weighted calculation on the proficiency scores, similarity scores, and accuracy scores corresponding to the N tutoring scenarios to determine the final target tutoring result based on the weighted calculation result.
[0056] In the voice-guided practice data analysis method provided in this embodiment, voice recognition and parsing are performed on the practice voice data to determine the scene text data corresponding to N practice scenarios; the proficiency score, similarity score, and accuracy analysis corresponding to each practice scenario are analyzed separately to determine the user's mastery of the knowledge points corresponding to a single practice scenario; finally, a comprehensive analysis is performed on the proficiency score, similarity score, and accuracy score corresponding to the N practice scenarios to obtain the target practice results, ensuring the accuracy of the target practice results, so as to quantitatively evaluate the user's mastery of all knowledge points involved in the N practice scenarios based on the target practice results.
[0057] In one embodiment, such as Figure 3 As shown, step S203 involves performing proficiency analysis on the scenario text data corresponding to each training scenario to obtain the proficiency score for that scenario, including:
[0058] S301: Perform a statistical analysis of modal particles on the scene text data corresponding to each training scenario, and obtain the number of modal particles corresponding to the scene text data.
[0059] S302: Count the number of words in the scene text data corresponding to each training scenario, and obtain the number of scene words corresponding to the scene text data.
[0060] S303: Based on the number of auxiliary words and the number of characters in the scene text data, obtain the proficiency score corresponding to the training scene.
[0061] As an example, in step S301, after obtaining the scenario text data corresponding to N tutoring scenarios, the server can perform a particle count on the scenario text data corresponding to each tutoring scenario, that is, count the number of particle occurrences in the scenario text data to obtain the particle count corresponding to the scenario text data. In this example, the server needs to count the occurrences of "um," "this," "that," or other particle occurrences in the scenario text data that can indicate the user's hesitation in responding to the guided questions, so as to determine the user's proficiency in mastering the knowledge points corresponding to the guided questions based on the number of particles corresponding to these particle occurrences.
[0062] As an example, in step S302, after the server obtains the scene text data corresponding to N coaching scenarios, it can count the number of words in the scene text data corresponding to each coaching scenario, that is, count the number of words the user replied to in the coaching scenario, and obtain the number of scene words corresponding to the scene text data.
[0063] As an example, in step S303, after obtaining the number of auxiliary words and the number of scene texts corresponding to the scene text data, the server can calculate the quotient between the number of auxiliary words and the number of scene texts. This quotient is determined as the proportion of auxiliary words in the scene text data, reflecting the user's proficiency with the knowledge points corresponding to the guided questions. For example, a smaller proportion of auxiliary words indicates a greater proficiency with the knowledge points, and vice versa. Next, the server can normalize the proportion of auxiliary words, converting it into a proficiency score within a specific numerical range. This ensures that the calculated proficiency score is comparable to the similarity score and accuracy score, facilitating comprehensive analysis and obtaining the target practice results.
[0064] In the voice tutoring data analysis method provided in this embodiment, the number of modal particles and words in the scene text data corresponding to each tutoring scenario can be counted to determine the number of modal particles and the number of scene texts respectively. Based on the number of modal particles and the number of scene texts, a proficiency score is determined so that the proficiency score can quantitatively reflect the user's proficiency in mastering the knowledge points corresponding to the tutoring scenario, so as to analyze the user's learning quality and the tutor's teaching quality in the future.
[0065] In one embodiment, the scenario text data corresponding to the training scenario includes M scenario sentences, and the standard text data corresponding to the training scenario includes K standard sentences, where M≥1 and K≥1;
[0066] like Figure 4 As shown, step S204 involves performing a similarity analysis on the scene text data and the standard text data corresponding to each training scenario to obtain a similarity score for each training scenario, including:
[0067] S401: Calculate the similarity between each scene sentence and each standard sentence, and determine the first similarity between the scene sentence and the standard sentence;
[0068] S402: Select the maximum value from the K first similarity scores corresponding to sentences in the same scene and determine it as the second similarity score corresponding to the scene sentences;
[0069] S403: Clean the second similarity scores of the M scene sentences to determine the third similarity scores of the M scene sentences;
[0070] S404: Determine the similarity score of the training scenario based on the third similarity of the sentences in the M scenarios.
[0071] Here, scenario sentences are sentences from the scenario text data corresponding to the training scenario, and M is the number of scenario sentences. Standard sentences are sentences from the standard text data corresponding to the training scenario, and K is the number of standard sentences.
[0072] As an example, in step S401, after obtaining the scenario text data and standard text data corresponding to the training scenario, the server can determine M scenario sentences from the scenario text data and K standard sentences from the standard text data. Then, using a similarity algorithm or other calculation logic, it calculates the similarity between each scenario sentence and each standard sentence to determine the first similarity of the scenario sentence. This first similarity can be understood as the degree of similarity between the scenario sentence in the user's response to the guidance question and the standard sentences in the standard text data. Understandably, since there are K standard sentences, the number of first similarities between the calculated scenario sentences and the K standard sentences is also K.
[0073] As an example, in step S402, when the server obtains the first similarity between each scene sentence and each standard sentence, it can select the maximum value from the K first similarities corresponding to the same scene sentence and determine it as the second similarity corresponding to the scene sentence. Here, the second similarity is the maximum value of the K first similarities, which reflects the degree of similarity between the scene sentence and the most similar standard sentence among the K standard sentences. In other words, the second similarity reflects the degree of similarity between the scene sentence and a certain standard sentence in the standard answer corresponding to the guiding question.
[0074] As an example, in step S403, after determining the second similarity corresponding to the M scene sentences, the server can clean the second similarity of the M scene sentences to remove the second similarity corresponding to scene sentences with large errors, thereby determining the third similarity of the M scene sentences. Here, the third similarity is the similarity after cleaning the second similarity.
[0075] In this example, the server can compare the second similarity of the scene sentence with a pre-set similarity threshold. If the second similarity is greater than the threshold, it indicates a high degree of similarity, and the second similarity is retained as the third similarity. If the second similarity is not greater than the threshold, it indicates a low degree of similarity. Calculating subsequent similarity scores based on this second similarity might lead to significant errors. Therefore, the similarity threshold or a default value (such as 0) can be used as the third similarity. This similarity threshold is a pre-set threshold used to assess whether data cleaning is necessary.
[0076] As an example, in step S404, after determining the third similarity of the M scene sentences, the server can determine the similarity score of the training scenario based on the third similarity of the M scene sentences. In this example, the server can perform weighted processing on the third similarity of the M scene sentences. For example, it can obtain the sentence weight of the most similar standard sentence among the M scene sentences, and perform weighted calculation based on the sentence weight and the third similarity of the M scene sentences to determine the similarity score of the training scenario, so that the similarity score can effectively reflect the degree of similarity between the scene text data and the standard text data of the training scenario.
[0077] In the voice coaching data analysis method provided in this embodiment, the similarity between each scene sentence and each standard sentence is calculated to determine K first similarities corresponding to the same scene sentence. The maximum value among the K first similarities is selected as the second similarity of the scene sentence, so as to determine the degree of similarity between the scene sentence and the most similar standard sentence. Then, the second similarities corresponding to M scene sentences are cleaned to determine the third similarities corresponding to M scene sentences, so as to reduce calculation errors and help ensure the accuracy of the final calculated similarity score. Finally, based on the third similarities corresponding to the M scene sentences, the similarity score corresponding to the coaching scene is determined, so that the similarity score can effectively reflect the degree of similarity between the scene text data and the standard text data of the coaching scene.
[0078] In one embodiment, such as Figure 5 As shown, step S401 involves calculating the similarity between each scene sentence and each standard sentence. The first similarity between the scene sentence and the standard sentence includes:
[0079] S501: Construct a target matrix based on each scenario sentence and each standard sentence, and determine whether the matrix values in the target matrix are the same or different identifiers;
[0080] S502: Determine the minimum span value of the same text based on the matrix positions of the same identifiers in the target matrix;
[0081] S503: Determine the first similarity between the scene sentence and the standard sentence based on the minimum span value of the same text.
[0082] The "identical" identifier is used to indicate that the text in the scenario sentence and the standard sentence are the same, and can be represented by 1. The "different" identifier is used to indicate that the text in the scenario sentence and the standard sentence are different, and can be represented by 0.
[0083] As an example, in step S501, a target matrix A*B is constructed based on the text length A of each scene sentence and the text length B of each standard sentence; each scene sentence and each standard sentence are traversed, and when there are identical characters in the scene sentence and the standard sentence, the same identifier (such as 1) is filled into the corresponding position of the target matrix; otherwise, when there are different characters in the scene sentence and the standard sentence, different identifiers (such as 0) are filled into the corresponding position of the target matrix to construct the target matrix.
[0084] As an example, in step S502, the server can determine the minimum span value of identical text between the scene sentence and the standard sentence based on the matrix positions of the same identifiers in the target matrix. Here, the minimum span value can be understood as the minimum distance between two identical characters.
[0085] In this example, the server can identify two standard keywords from a standard sentence. Based on whether the matrix values in the target matrix are the same or different identifiers, it determines whether the scene sentence contains these two standard keywords. If the scene sentence contains these two standard keywords, the spacing between them is determined. If there are no spacers between the two standard keywords, the span value is 0; if there are spacers between the two standard keywords, the span value is determined based on the number of spacers. Since the same standard keywords may appear simultaneously in any scene sentence, the minimum span value needs to be calculated to determine the minimum span value.
[0086] For example, in the target matrix, since the corresponding matrix positions are filled with the same or different identifiers such as 1 or 0, the minimum span value can be determined by using d[i][j] = Math.min(Math.min(d[i-1][j]+1,d[i][j-1]+1), d[i-1][j-1]+temp). Here, d[i][j] is the minimum span value of the i-th row and j-th column, Math.min is the minimum value function, d[i-1][j] is the minimum span value of the (i-1)-th row and j-th column, d[i][j-1] is the minimum span value of the i-th row and (j-1)-th column, d[i-1][j-1] is the minimum span value of the (i-1)-th row and (j-1)-th column, and temp is the matrix value, which can be set to 1 or 0.
[0087] As an example, in step S503, the server can determine the first similarity between the scene sentence and the standard sentence based on the minimum span value of the same text. The calculation process is as follows: First, based on the text length of each scene sentence and the text length of each standard sentence, the maximum value between the two is selected and determined as the maximum text length; then, based on the minimum span value of the same text and the maximum text length, the quotient between the two is calculated and determined as the first proportion; then, based on the first proportion, the first similarity between the scene sentence and the standard sentence is determined.
[0088] For example, the server can use the formula 1-d[n][m] / Math.max(str.length(), target.length())) to calculate the first similarity between the scene sentence and the standard sentence, where d[n][m] is the minimum span value of the same text, Math.max is the maximum value function, str.length() is the text length of the scene sentence, and target.length() is the text length of the standard sentence.
[0089] In the voice training data analysis method provided in this embodiment, a target matrix is constructed based on scenario sentences and standard sentences. Based on the matrix values in the target matrix, the matrix positions of the same identifiers are determined, and the minimum span value of the same text is determined to determine the closeness of the same text. Then, based on the minimum span value of the same text, the first similarity between the two sentences is determined, which can reflect the closeness of the same text in the sentences, and thus reflect the similarity between the sentences.
[0090] In one embodiment, such as Figure 6 As shown, step S205 involves performing accuracy analysis on the scenario text data and standard text data corresponding to each training scenario to obtain an accuracy score for that training scenario, including:
[0091] S601: Extract keywords from the scene text data corresponding to each training scenario to obtain the training keywords corresponding to the scene text data.
[0092] S602: Obtain the standard keywords corresponding to the standard text data for each training scenario, and determine the number of standard words corresponding to the standard keywords.
[0093] S603: Count the number of matching words between the training coach keywords and the standard keywords;
[0094] S604: Obtain the accuracy score corresponding to the training scenario based on the number of matched words and the number of standard words.
[0095] As an example, in step S601, after the server obtains the scene text data corresponding to N training scenarios, it can extract keywords from the scene text data corresponding to each training scenario and determine the extracted keywords as training keywords in the scene text data. In other words, the training keywords are keywords extracted in real time from the training text data.
[0096] As an example, in step S602, the server can retrieve standard text data corresponding to a specific coaching scenario from the system database, extract keywords from the standard text data, and determine the extracted keywords as standard keywords corresponding to the coaching scenario. In this example, since the coaching text data corresponding to the coaching scenario is pre-set text data by the system, keyword extraction can be performed in advance to store the standard keywords in the system database. This allows the server to directly retrieve the corresponding standard keywords from the system database when accuracy analysis of the scenario text data corresponding to a specific coaching scenario is required, thereby improving the efficiency of standard keyword retrieval.
[0097] In this example, after the server obtains the standard keywords corresponding to the standard text data for each training scenario, it also needs to count the number of all standard keywords and determine them as the number of standard words corresponding to the standard keywords.
[0098] As an example, in step S603, after obtaining the practice keywords corresponding to all scenario text data and the standard keywords corresponding to the standard text data, the server can match each practice keyword with all the standard keywords one by one. If the practice keyword and the standard keyword are the same, the match is considered successful; if the practice keyword and the standard keyword are not the same, the match is considered unsuccessful. Then, the server counts the number of successfully matched practice keywords and determines it as the number of matched words.
[0099] As an example, in step S604, after obtaining the number of matching words and the number of standard words, the server can calculate the quotient between the number of matching words and the number of standard words. This quotient is determined as the matching percentage, which reflects the proportion of successfully matched practice keywords among all standard keywords. This percentage reflects the user's accuracy in understanding the knowledge points corresponding to the guided questions. For example, a higher matching percentage indicates a higher accuracy in understanding the knowledge points, and vice versa. Next, the server can normalize the matching percentage, converting it into an accuracy score within a specific numerical range. This ensures that the calculated accuracy score is comparable to the proficiency score and similarity score, allowing for comprehensive analysis and the acquisition of the target practice results.
[0100] In the voice tutoring data analysis method provided in this embodiment, the tutoring keywords extracted from the scene text data are matched with the standard keywords corresponding to the standard text data to determine the number of matched words. Then, based on the number of matched words and the number of standard words corresponding to the standard keywords, an accuracy score is determined so that the accuracy score can quantitatively reflect the accuracy of the user's mastery of the knowledge points corresponding to the tutoring scene, so as to facilitate subsequent analysis of the user's learning quality and the tutor's teaching quality.
[0101] In one embodiment, such as Figure 7 As shown, step S206 involves comprehensively analyzing the proficiency scores, similarity scores, and accuracy scores corresponding to the N coaching scenarios to obtain the target coaching result, including:
[0102] S701: Based on a comprehensive analysis of the proficiency score, similarity score, and accuracy score corresponding to each training scenario, obtain the scenario training score corresponding to each training scenario.
[0103] S702: Determine the target training result based on the training scores corresponding to N training scenarios.
[0104] Among them, the scenario-based coaching score is the coaching score for a specific coaching scenario, which can be understood as a comprehensive score of the user's mastery of the knowledge points involved in that coaching scenario.
[0105] As an example, in step S701, after obtaining the proficiency score, similarity score, and accuracy score corresponding to N coaching scenarios, the server can perform a weighted analysis on the proficiency score, similarity score, and accuracy score corresponding to each coaching scenario to obtain the scenario coaching score corresponding to each coaching scenario.
[0106] In this example, the server can obtain pre-set proficiency weights, similarity weights, and accuracy weights. Then, it performs a weighted calculation on the proficiency score, proficiency weight, similarity score, similarity weight, accuracy score, and accuracy weight to obtain the scenario-based coaching score for each coaching scenario. Understandably, this scenario-based coaching score comprehensively reflects the user's proficiency and accuracy in mastering the knowledge points involved in the coaching scenario.
[0107] As an example, in step S702, after obtaining the scenario-based tutoring score for each tutoring scenario, the server can perform a comprehensive analysis of the scenario-based tutoring scores for the N tutoring scenarios to determine the target tutoring result. Understandably, this target tutoring result can comprehensively reflect the user's mastery of all knowledge points involved in the N tutoring scenarios.
[0108] In the voice coaching data analysis method provided in this embodiment, the corresponding scenario coaching score is first determined based on the proficiency score, similarity score, and accuracy score of each coaching scenario to assess the user's mastery of the knowledge points involved in each coaching scenario; then, based on the scenario coaching scores corresponding to N coaching scenarios, the final target coaching result is determined to assess the user's mastery of all knowledge points involved in the N coaching scenarios.
[0109] In one embodiment, such as Figure 8 As shown, step S702, which determines the target coaching result based on the coaching scores corresponding to the N coaching scenarios, includes:
[0110] S801: Determine the scenario training result corresponding to each training scenario based on the scenario training score and the passing score threshold.
[0111] S802: Count the number of training scenarios with qualified training results and determine the number of qualified scenarios;
[0112] S803: Determine the target training result based on the number of qualified scenarios and the number of training scenarios.
[0113] Among them, the passing score threshold is a threshold set in advance by the system to evaluate whether the training score of a certain training scenario has reached the passing standard.
[0114] As an example, in step S801, after obtaining the scenario practice scores corresponding to N practice scenarios, the server can compare the scenario practice score for each practice scenario with a pre-set passing score threshold to determine the scenario practice result for that scenario. In this example, the server compares the scenario practice score for each practice scenario with the pre-set passing score threshold. If the scenario practice result is greater than or equal to the passing score threshold, the scenario practice result is determined to be qualified, indicating that the user's mastery of the knowledge points involved in the practice scenario meets the passing standard. If the scenario practice result is less than the passing score threshold, the scenario practice result is determined to be unqualified, indicating that the user's mastery of the knowledge points involved in the practice scenario does not meet the passing standard.
[0115] As an example, in step S802, the server also needs to count the number of training scenarios with qualified training results among the N training scenarios, and determine them as the number of qualified scenarios, so as to comprehensively determine the target training result based on the number of qualified scenarios.
[0116] The number of coaching scenarios refers to the number of coaching scenarios involved in the voice coaching process, and the number of coaching scenarios = N.
[0117] As an example, in step S803, after obtaining the number of qualified scenarios and the number of practice scenarios, the server can also calculate the quotient of the number of qualified scenarios and the number of practice scenarios, and determine the final target practice result based on the quotient of the two, so that the target practice result can comprehensively reflect the user's mastery of the knowledge points involved in the N practice scenarios.
[0118] In the voice tutoring data analysis method provided in this embodiment, the tutoring result of each tutoring scenario is first determined according to the tutoring score and the passing score threshold of each tutoring scenario, so as to determine the user's mastery of the knowledge points involved in a single tutoring scenario; then, the target tutoring result is determined according to the number of passing scenarios and the number of tutoring scenarios, so as to reflect the user's mastery of the knowledge points involved in N tutoring scenarios.
[0119] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0120] In one embodiment, a voice-guided training data analysis device is provided, which corresponds one-to-one with the voice-guided training data analysis method described in the above embodiments. For example... Figure 9As shown, the voice-guided practice data analysis device includes a voice-guided practice data acquisition module 901, a text-guided practice data acquisition module 902, a proficiency score acquisition module 903, a similarity score acquisition module 904, an accuracy score acquisition module 905, and a target practice result acquisition module 906. Detailed descriptions of each functional module are as follows:
[0121] The training voice data acquisition module 901 is used to acquire training voice data;
[0122] The training text data acquisition module 902 is used to perform speech recognition and text parsing on the training voice data to acquire the training text data. The training text data includes scene text data corresponding to N training scenarios, where N≥1.
[0123] The proficiency score acquisition module 903 is used to perform proficiency analysis on the scene text data corresponding to each training scenario and obtain the proficiency score corresponding to the training scenario.
[0124] The similarity score acquisition module 904 is used to perform similarity analysis on the scene text data and the standard text data corresponding to each training scenario to obtain the similarity score corresponding to the training scenario.
[0125] The accuracy score acquisition module 905 is used to perform accuracy analysis on the scene text data and standard text data corresponding to each training scenario to obtain the accuracy score corresponding to the training scenario.
[0126] The target coaching result acquisition module 906 is used to comprehensively analyze the proficiency score, similarity score, and accuracy score corresponding to N coaching scenarios to obtain the target coaching result.
[0127] In one embodiment, the proficiency score acquisition module 903 includes:
[0128] The auxiliary word quantity acquisition unit is used to perform auxiliary word statistics on the scene text data corresponding to each training scenario and obtain the number of auxiliary words corresponding to the scene text data.
[0129] The scene text quantity acquisition unit is used to count the number of words in the scene text data corresponding to each training scene and obtain the scene text quantity corresponding to the scene text data.
[0130] The proficiency score acquisition unit is used to obtain the proficiency score corresponding to the training scenario based on the number of auxiliary words and the number of characters in the scenario text data.
[0131] In one embodiment, the scenario text data corresponding to the training scenario includes M scenario sentences, and the standard text data corresponding to the training scenario includes K standard sentences, where M≥1 and K≥1;
[0132] Similarity score acquisition module 904 includes:
[0133] The first similarity acquisition unit is used to calculate the similarity between each scene sentence and each standard sentence, and the first similarity between the scene sentence and the standard sentence.
[0134] The second similarity acquisition unit is used to select the maximum value from K first similarities and determine it as the second similarity corresponding to the scene sentence;
[0135] The third similarity acquisition unit is used to clean the second similarity corresponding to M scene sentences and determine the third similarity corresponding to M scene sentences.
[0136] The similarity score acquisition unit is used to determine the similarity score of the training scenario based on the third similarity of the sentences in the M scenarios.
[0137] In one embodiment, the first similarity acquisition unit includes:
[0138] The target matrix construction sub-unit is used to construct the target matrix based on each scene sentence and each standard sentence, and to determine whether the matrix values in the target matrix are the same or different identifiers.
[0139] The minimum span value determination sub-unit is used to determine the minimum span value of the same text based on the matrix position of the same identifier in the target matrix.
[0140] The first similarity determination subunit is used to determine the first similarity between the scene sentence and the standard sentence based on the minimum span value of the same characters.
[0141] In one embodiment, the accuracy score acquisition module 905 includes:
[0142] The keyword acquisition unit is used to extract keywords from the scene text data corresponding to each training scenario and obtain the training keywords corresponding to the scene text data.
[0143] The standard word quantity acquisition unit is used to acquire the standard keywords corresponding to the standard text data of each training scenario and determine the standard word quantity corresponding to the standard keywords.
[0144] The matching word count acquisition unit is used to count the number of matching words that match the training keywords and the standard keywords;
[0145] The accuracy score acquisition unit is used to obtain the accuracy score corresponding to the coaching scenario based on the number of matched words and the number of standard words.
[0146] In one embodiment, the target training result acquisition module 906 includes:
[0147] The scenario-based coaching score acquisition unit is used to comprehensively analyze the proficiency score, similarity score, and accuracy score corresponding to each coaching scenario to obtain the scenario-based coaching score for each coaching scenario.
[0148] The target coaching result acquisition unit is used to determine the target coaching result based on the coaching scores corresponding to N coaching scenarios.
[0149] In one embodiment, the target training result acquisition unit includes:
[0150] The scenario-based coaching result acquisition subunit is used to determine the scenario-based coaching result corresponding to each coaching scenario based on the scenario-based coaching score and the passing score threshold.
[0151] The qualified scenario quantity acquisition subunit is used to count the number of training scenarios whose training results are qualified, and determine them as the qualified scenario quantity.
[0152] The target training result determination subunit is used to determine the target training result based on the number of qualified scenarios and the number of training scenarios.
[0153] Specific limitations regarding the voice-guided training data analysis device can be found in the limitations of the voice-guided training data analysis method described above, and will not be repeated here. Each module in the aforementioned voice-guided training data analysis device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in the processor of the computer device in hardware form or independent of it, or stored in the memory of the computer device in software form, so that the processor can call and execute the corresponding operations of each module.
[0154] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 10 As shown, the computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The database stores data used or generated during the execution of a voice-guided training data analysis method. The network interface communicates with external terminals via a network connection. When the computer program is executed by the processor, it implements a voice-guided training data analysis method.
[0155] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the voice-guided training data analysis method described in the above embodiment, for example... Figure 2 As shown in S201-S206, or Figure 3 As shown in Figure 8, to avoid repetition, further details will not be provided here. Alternatively, when the processor executes the computer program, it implements the functions of each module / unit in this embodiment of the voice tutoring data analysis device, for example... Figure 9 The functions of the training voice data acquisition module 901, training text data acquisition module 902, proficiency score acquisition module 903, similarity score acquisition module 904, accuracy score acquisition module 905, and target training result acquisition module 906 shown are not described in detail here to avoid repetition.
[0156] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When executed by a processor, the computer program implements the voice tutoring data analysis method described in the above embodiment, for example... Figure 2 As shown in S201-S206, or Figures 3 to 8 As shown, to avoid repetition, it will not be described again here. Alternatively, when the computer program is executed by the processor, it implements the functions of each module / unit in this embodiment of the voice tutoring data analysis device, for example... Figure 9 The functions of the training voice data acquisition module 901, training text data acquisition module 902, proficiency score acquisition module 903, similarity score acquisition module 904, accuracy score acquisition module 905, and target training result acquisition module 906 shown are not described in detail here to avoid repetition.
[0157] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and RAMbus dynamic RAM (RDRAM), etc.
[0158] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0159] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.
Claims
1. A method for analyzing voice-guided training data, characterized in that, include: Obtain the voice data of the coach; The training voice data is subjected to speech recognition and text parsing to obtain training text data. The training text data includes scenario text data corresponding to N training scenarios, where N≥1; the scenario text data corresponding to each training scenario includes M scenario sentences; and the standard text data corresponding to each training scenario includes K standard sentences, where M≥1 and K≥1. Perform proficiency analysis on the scene text data corresponding to each training scenario to obtain the proficiency score corresponding to the training scenario. Calculate the similarity between each of the scene sentences and each of the standard sentences to determine the first similarity between the scene sentences and the standard sentences; The maximum value is selected from the K first similarity scores and determined as the second similarity score corresponding to the scene sentence; Cleaning the second similarity corresponding to the M scene sentences and determining the third similarity corresponding to the M scene sentences includes: determining the second similarity greater than a preset similarity threshold as the third similarity; Obtain the sentence weights corresponding to the standard sentences that are most similar to the M scene sentences, and perform a weighted calculation based on the sentence weights and the third similarity corresponding to the M scene sentences to determine the similarity score corresponding to the training scene; Accuracy analysis is performed on the scene text data and standard text data corresponding to each training scenario to obtain the accuracy score corresponding to the training scenario. A comprehensive analysis is performed on the proficiency score, similarity score, and accuracy score corresponding to the N training scenarios to obtain the target training result.
2. The voice tutoring data analysis method as described in claim 1, characterized in that, The step of performing proficiency analysis on the scene text data corresponding to each training scenario to obtain the proficiency score corresponding to the training scenario includes: For each training scenario, the number of auxiliary words in the corresponding scenario text data is obtained by counting the auxiliary words in the scenario text data. The number of words in the scene text data corresponding to each training scenario is counted to obtain the number of scene words corresponding to the scene text data. Based on the number of auxiliary words and the number of characters in the scene text data, the proficiency score corresponding to the training scenario is obtained.
3. The voice tutoring data analysis method as described in claim 1, characterized in that, The step of calculating the similarity between each scene sentence and each standard sentence to determine the first similarity between the scene sentence and the standard sentence includes: A target matrix is constructed based on each scenario sentence and each standard sentence, and the matrix values in the target matrix are determined to be the same or different identifiers. Based on the matrix positions of the same identifiers in the target matrix, determine the minimum span value of the same text. The first similarity between the scene sentence and the standard sentence is determined based on the minimum span value of the identical characters.
4. The voice tutoring data analysis method as described in claim 1, characterized in that, The step of performing accuracy analysis on the scene text data and standard text data corresponding to each training scenario to obtain an accuracy score for each training scenario includes: Extract keywords from the scene text data corresponding to each training scenario to obtain the training keywords corresponding to the scene text data. Obtain the standard keywords corresponding to the standard text data for each training scenario, and determine the number of standard words corresponding to the standard keywords. Count the number of matching words between the training keywords and the standard keywords; Based on the number of matched words and the number of standard words, the accuracy score corresponding to the coaching scenario is obtained.
5. The voice tutoring data analysis method as described in claim 1, characterized in that, The process of comprehensively analyzing the proficiency scores, similarity scores, and accuracy scores corresponding to the N training scenarios to obtain the target training result includes: Based on a comprehensive analysis of the proficiency score, similarity score, and accuracy score corresponding to each training scenario, a scenario training score is obtained for each training scenario. The target training result is determined based on the training scores corresponding to the N training scenarios.
6. The voice tutoring data analysis method as described in claim 5, characterized in that, The step of determining the target coaching result based on the coaching scores corresponding to the N coaching scenarios includes: Based on the scenario training score and passing score threshold corresponding to each training scenario, the scenario training result corresponding to the training scenario is determined. The number of training scenarios whose results are deemed qualified is counted, and the number of qualified scenarios is determined. The target training result is determined based on the number of qualified scenarios and the number of training scenarios.
7. A voice-guided training data analysis device, characterized in that, include: The training voice data acquisition module is used to acquire training voice data; The training text data acquisition module is used to perform speech recognition and text parsing on the training voice data to acquire training text data. The training text data includes scenario text data corresponding to N training scenarios, where N≥1; the scenario text data corresponding to each training scenario includes M scenario sentences; and the standard text data corresponding to each training scenario includes K standard sentences, where M≥1 and K≥1. The proficiency score acquisition module is used to perform proficiency analysis on the scene text data corresponding to each training scenario and acquire the proficiency score corresponding to the training scenario. The similarity score acquisition module is used to calculate the similarity between each of the scene sentences and each of the standard sentences, determine the first similarity between the scene sentences and the standard sentences, and select the maximum value from K first similarities to determine the second similarity corresponding to the scene sentence. The process involves cleaning the second similarity scores corresponding to the M scene sentences and determining the third similarity scores corresponding to the M scene sentences. This includes: determining the second similarity scores that are greater than a pre-set similarity threshold as the third similarity scores; obtaining the sentence weights corresponding to the standard sentences that are most similar to the M scene sentences; and performing a weighted calculation based on the sentence weights and the third similarity scores corresponding to the M scene sentences to determine the similarity score corresponding to the training scenario. The accuracy score acquisition module is used to perform accuracy analysis on the scene text data and the standard text data corresponding to each training scenario to obtain the accuracy score corresponding to the training scenario. The target coaching result acquisition module is used to comprehensively analyze the proficiency score, similarity score, and accuracy score corresponding to the N coaching scenarios to obtain the target coaching result.
8. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the voice training data analysis method as described in any one of claims 1 to 6.
9. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the voice training data analysis method as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Intelligent interactive training system
CN110956142A