Speech evaluation data enhancement method, speech evaluation method, device and apparatus
By calculating the speech scores of associated words using phoneme sequences and edit distance, the problem of obtaining training data for speech evaluation models was solved, thus increasing the amount of data and improving model learning.
Patent Information
- Application Number
- CN202310604948.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-05-25
- Publication Date
- 2025-11-11
- Estimated Expiration
- 2043-05-25
AI Technical Summary
In existing technologies, acquiring training data for voice assessment models is difficult and inefficient, requiring manual annotation with specialized knowledge, which makes data acquisition challenging.
By acquiring the speech evaluation data to be enhanced, the associated phoneme sequences and associated words are determined using phoneme sequences and preset phoneme edit distances. The second speech score of the associated words is calculated based on the speech score of the first word and the phoneme edit distance, thereby expanding the speech evaluation data.
This reduces the difficulty of obtaining voice assessment data, increases the amount of data, enriches the training data, and improves the learning effect of the voice assessment model.
Smart Images

Figure CN116597862B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of speech evaluation data processing, specifically to a speech evaluation data enhancement method, speech evaluation method, apparatus, and device. Background Technology
[0002] Speech assessment technology is a subfield of computer-assisted language learning. It aims to efficiently and accurately assess learners' pronunciation accuracy, fluency, and completeness. In improving speech assessment technology, the accuracy of the assessment model's results is directly related to the amount of training data; often, a large amount of training data is required to improve the accuracy of the model's results.
[0003] In the process of realizing this application, the inventors discovered that in related technologies, the training data of voice evaluation models often requires manual annotation and scoring by personnel with professional knowledge, which has technical defects such as high difficulty in obtaining data and low efficiency in obtaining data. Summary of the Invention
[0004] The purpose of this application is to overcome the shortcomings and deficiencies in the prior art and provide a speech evaluation data enhancement method, speech evaluation method, device and equipment, which can reduce the difficulty of obtaining speech evaluation data and achieve the technical effect of efficiently increasing the amount of speech evaluation data.
[0005] The first aspect of this application provides a voice assessment data enhancement method, including:
[0006] Obtain the speech evaluation data to be enhanced; the speech evaluation data includes several first words and the first speech score of each first word, and the phoneme sequence of each first word;
[0007] Based on the phoneme sequence of each first word and the preset phoneme edit distance, determine the associated phoneme sequence of each first word; based on the associated phoneme sequence, determine the associated words of each first word and the corresponding edit distance type.
[0008] The second speech score of each of the associated words is determined based on the first speech score of each of the first words, the phoneme sequence of each of the first words, the phoneme edit distance, and the edit distance type;
[0009] Enhanced speech evaluation data is obtained based on the first word, the first speech score, the associated word, and the second speech score.
[0010] A second aspect of this application provides a voice assessment method, including:
[0011] The initial assessment model is trained based on the speech assessment data obtained by the speech assessment data augmentation method, and the speech assessment model is obtained.
[0012] The speech data and audio text to be evaluated are input into the speech evaluation model to obtain the speech evaluation results.
[0013] A third aspect of this application provides a voice assessment data enhancement device, comprising:
[0014] The speech evaluation data acquisition module is used to acquire speech evaluation data to be enhanced; the speech evaluation data includes several first words and the first speech score of each first word, and the phoneme sequence of each first word;
[0015] The associated word acquisition module is used to determine the associated phoneme sequence of each first word based on the phoneme sequence of each first word and a preset phoneme edit distance; and to determine the associated words of each first word and the corresponding edit distance type based on the associated phoneme sequence.
[0016] The second speech score acquisition module is used to determine the second speech score of each of the associated words based on the first speech score of each of the first words, the phoneme sequence of each of the first words, the phoneme edit distance, and the edit distance type.
[0017] The enhanced speech evaluation data module is used to obtain enhanced speech evaluation data based on the first word, the first speech score, the associated word, and the second speech score.
[0018] A fourth aspect of this application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the method described above.
[0019] A fifth aspect of this application provides a computer device including a storage device, a processor, and a computer program stored in the storage device and executable by the processor, wherein the processor executes the computer program to implement the steps of the method described above.
[0020] Compared to related technologies, this application determines the associated phoneme sequences of each first word and the associated words corresponding to those phoneme sequences based on the phoneme sequence of the first word in the speech evaluation data to be enhanced and a preset phoneme edit distance. Then, based on the first speech score of the first word, the phoneme sequence of each first word, and the phoneme edit distance, the second speech score of each of the associated words is determined. The speech evaluation data is then expanded based on the associated words and the second speech scores to obtain enhanced speech evaluation data. This reduces the difficulty of obtaining speech evaluation data and achieves the technical effect of efficiently increasing the quantity of speech evaluation data. Moreover, the enhanced speech evaluation data is richer in content, thus the speech evaluation model trained based on the enhanced speech evaluation data learns more effectively.
[0021] To provide a clearer understanding of this application, the specific embodiments of this application will be described below in conjunction with the accompanying drawings. Attached Figure Description
[0022] Figure 1 This is a flowchart of a voice evaluation data enhancement method according to an embodiment of this application.
[0023] Figure 2 This is a schematic diagram of associated words in a speech evaluation data enhancement method according to an embodiment of this application.
[0024] Figure 3 This is a flowchart of a voice evaluation method according to an embodiment of this application.
[0025] Figure 4 This is a schematic diagram of the module connections of a voice assessment data enhancement device according to an embodiment of this application.
[0026] 100. Voice assessment data enhancement device; 101. Voice assessment data acquisition module to be enhanced; 102. Related word acquisition module; 103. Second voice score acquisition module; 104. Enhanced voice assessment data module. Detailed Implementation
[0027] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.
[0028] It should be understood that the described embodiments are merely some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of the embodiments of this application.
[0029] In the following description, when referring to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. In the description of this application, it should be understood that the terms "first," "second," "third," etc., are used only to distinguish similar objects and are not necessarily used to describe a specific order or sequence, nor should they be construed as indicating or implying relative importance. Those skilled in the art can understand the specific meaning of the above terms in this application according to the specific circumstances. The singular forms "a," "the," and "the" used in this application and the appended claims are also intended to include the plural forms, unless the context clearly indicates otherwise. The word "if" as used herein can be interpreted as "when," "when," or "in response to determination."
[0030] Furthermore, in the description of this application, unless otherwise stated, "multiple" means two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone. The character " / " generally indicates that the preceding and following related objects have an "or" relationship.
[0031] The speech evaluation data enhancement method of this application can be executed by a speech evaluation data enhancement device. This device can be implemented through software and / or hardware, and can consist of two or more physical entities, or a single physical entity. The hardware referred to as the speech evaluation data enhancement device essentially refers to computer equipment; for example, the speech evaluation data enhancement device can be a computer, mobile phone, tablet, or smart interactive whiteboard, or other intelligent device.
[0032] The voice assessment data enhancement device may include one or more processing cores. It can implement the voice assessment data enhancement method of this application through pure software, or through a combination of software and hardware. For example, it can be implemented using at least one hardware form selected from digital signal processing, field-programmable gate arrays, and programmable logic arrays. It can integrate one or more of the following: a central processing unit, an image processor, and a modem. The voice assessment data enhancement device can run an application program containing the voice assessment data enhancement method of the device. The application program can be presented in a form adapted to the voice assessment data enhancement device, such as an APP application. In some examples, it can also be presented in the form of a system plugin, a web page plugin, etc.
[0033] Please see Figure 1This is a flowchart of a speech evaluation data enhancement method according to an embodiment of this application, including:
[0034] S1: Obtain the speech evaluation data to be enhanced; the speech evaluation data includes several first words and the first speech score of each first word, and the phoneme sequence of each first word.
[0035] The first word refers to a word composed of letter characters, primarily an English word. Optionally, a French or Russian word may also be used.
[0036] The first pronunciation score refers to the score of the test taker's pronunciation of the first word, which can be obtained by a person with English expertise or an English word scoring expert scoring the test taker's pronunciation of the first word.
[0037] A phoneme sequence is the arrangement of the phonemes in a word. A phoneme is the smallest unit of speech that is divided according to the pronunciation of a word. For example, the phoneme sequence of the word "bank" "b ael ng k" (the spaces in the quotation marks are used to separate the phonemes for distinction) includes the four phonemes "b", "ael", "ng", and "k".
[0038] S2: Based on the phoneme sequence of each first word and the preset phoneme edit distance, determine the associated phoneme sequence of each first word; based on each associated phoneme sequence, determine the associated words of each first word and the corresponding edit distance type.
[0039] Phoneme edit distance refers to the number of phonemes in a phoneme sequence that have been altered through editing.
[0040] A related phoneme sequence refers to a new phoneme sequence obtained by editing the phoneme sequence of the first word without exceeding the phoneme edit distance. Therefore, if the phoneme edit distance is 1, it means the phoneme difference between the related phoneme sequence and the phoneme sequence of the first word is caused by one phoneme; if the phoneme edit distance is 2, it means the phoneme difference is caused by two phonemes; and if the phoneme edit distance is 3, it means the phoneme difference is caused by three phonemes. Specifically, if the phoneme edit distance is 1, the related phoneme sequence can be obtained by deleting one phoneme from the phoneme sequence of the first word, inserting one phoneme into the phoneme sequence of the first word, or replacing one phoneme in the phoneme sequence of the first word. Associated words refer to words that correspond to associated phoneme sequences. Taking the word "bank" as an example, deleting the phoneme "k" yields the associated phoneme sequence "bae1 ng", which corresponds to the word "bang"; inserting one phoneme yields the associated phoneme sequence "bl ae1 ng k", which corresponds to the word "blank"; and replacing one phoneme yields the associated phoneme sequence "bae1 ng d", which corresponds to the word "banged".
[0041] The edit distance types for editing operations include delete phoneme type, insert phoneme type, and replace phoneme type.
[0042] The deleted phoneme type refers to the type of related words obtained by deleting a certain number of phonemes based on the phoneme edit distance. The number of phonemes deleted is equal to the phoneme edit distance.
[0043] The phoneme insertion type refers to the type of related words obtained by inserting a number of phonemes based on the phoneme edit distance. The number of phonemes inserted is equal to the phoneme edit distance.
[0044] The replacement phoneme type refers to the type of related word obtained by replacing a number of phonemes based on the phoneme edit distance. The number of phoneme replacements is equal to the phoneme edit distance.
[0045] Taking the first word "bank" and the related word "bang" as examples, the related phoneme sequence "bae1ng" of the related word "bang" is obtained by deleting one phoneme "k". Therefore, the edit distance type of the related word "bang" relative to the first word "bank" is the deleted phoneme type. Taking the first word "bank" and the related word "blank" as examples, the related phoneme sequence "bl ae1 ng k" of the related word "blank" is obtained by inserting one phoneme. Therefore, the edit distance type of the related word "blank" relative to the first word "bank" is the inserted phoneme type. Taking the first word "bank" and the related word "banged" as examples, the related word "banged" is obtained by replacing one phoneme. Therefore, the edit distance type of the related word "banged" relative to the first word "bank" is the replaced phoneme type.
[0046] Step S2 can be implemented by searching a pre-built word library based on the phoneme sequence and phoneme edit distance of the first word. The word library includes the text of at least several words and their corresponding phoneme sequences. For example, taking the word "bank" as an example, searching the pre-built word library based on the phoneme sequence of the first word and a preset phoneme edit distance can yield results such as... Figure 2 The multiple related words shown.
[0047] S3: Determine the second speech score of each associated word based on the first speech score of each first word, the phoneme sequence of each first word, the phoneme edit distance, and the edit distance type.
[0048] Among them, the different phoneme editing distances and the different types of corresponding editing distances will also affect the specific value of the second speech score.
[0049] Specifically, a first word can generate multiple related words based on its phoneme edit distance. These related words may have different edit distance types or different specific phoneme edit contents. Phoneme edit contents refer to the specific phonemes that are edited. Taking the word "bank" as an example, a phoneme edit distance of 1 includes the related phoneme sequence "bae1 ng" for the word "bang" and the related phoneme sequence "bae1 ng d" for the word "banged." The words "bang" and "banged" have the same phoneme edit distance but different edit distance types. Therefore, the specific values of the second phoneme scores for related words of the same first word may differ, and the larger the phoneme edit distance, the smaller the specific value of the second phoneme score for the related word.
[0050] S4: Obtain enhanced speech evaluation data based on the first word, the first speech score, the associated words, and the second speech score.
[0051] By storing the first word, the first speech score, the associated words, and the second speech score, the enhanced speech evaluation data can be obtained. During storage, each first word is associated with its corresponding first speech score, and each associated word is associated with its corresponding second speech score. Specifically, the enhanced speech evaluation data can also be grouped and stored according to the speech scores, which helps users obtain training data from each group according to the training needs of the speech evaluation model.
[0052] Compared to related technologies, this application determines the associated phoneme sequences of each first word and the associated words corresponding to those phoneme sequences based on the phoneme sequence of the first word in the speech evaluation data to be enhanced and a preset phoneme edit distance. Then, based on the first speech score of the first word, the phoneme sequence of each first word, the phoneme edit distance, and the edit distance type, the second speech score of each associated word is determined. This allows the speech evaluation data to be expanded based on the associated words and the second speech score, thereby obtaining enhanced speech evaluation data. This reduces the difficulty of obtaining speech evaluation data and achieves the technical effect of efficiently increasing the quantity of speech evaluation data. Moreover, the enhanced speech evaluation data is richer in content, thus the speech evaluation model trained based on the enhanced speech evaluation data learns more effectively.
[0053] In a feasible embodiment, S3: the step of determining the second speech score of each associated word based on the first speech score of each first word, the phoneme sequence of each first word, the phoneme edit distance, and the edit distance type includes:
[0054] S31: Obtain the speech score calculation method corresponding to the edit distance type of each related word.
[0055] Since different edit distance types directly affect the calculation of the speech score of associated words, it is necessary to obtain the corresponding speech score calculation method according to the edit distance type of each associated word in order to guide the calculation of the speech score of associated words.
[0056] S32: Based on the calculation methods for each speech score, the first speech score of each first word, and the phoneme sequence of each first word, obtain the second speech score of each related word.
[0057] In this embodiment, the speech score calculation method corresponding to the edit distance type is used, and then the second speech score of the associated word is obtained based on the first speech score of the first word in the speech score calculation method. This can quickly obtain the second speech score of the associated word and improve the efficiency of obtaining the second speech score of the associated word.
[0058] In a feasible embodiment, S32: the step of obtaining the second speech score of each associated word based on each speech score calculation method, the first speech score of each first word, and the phoneme sequence of each first word includes:
[0059] S321: If the edit distance type is deleted phoneme type or replaced phoneme type, the ratio of the total score of the deleted or replaced phonemes to the total score of the phonemes in the first word phoneme sequence is determined as the score percentage of the deleted or replaced phonemes.
[0060] The total score of deleted or replaced phonemes refers to the sum of the scores of each deleted or replaced phoneme, while the total score of the phonemes in the first word phoneme sequence refers to the sum of the scores of each phoneme in the first word phoneme sequence.
[0061] S322: Obtain the first deduction value based on the percentage of scores of deleted or replaced phonemes and the first speech score.
[0062] S323: The difference between the first phonetic score and the first deduction score of the first word is determined as the second phonetic score of the associated word corresponding to the first word.
[0063] Steps S321-S323 can be achieved using the following formulas:
[0064]
[0065] Among them, S n S is the second phonetic score of the associated word. o For the first speech score, ∑ω i The score for the deleted or replaced phoneme, ω i For the score of the i-th deleted or replaced phoneme, ∑ω j ω is the total phoneme score of the phoneme sequence of the first word. j This is the score for the j-th phoneme.
[0066] In this embodiment, through steps S321-S323, the second speech score of the associated word with the edit distance type of deletion phoneme type or replacement phoneme type can be accurately obtained.
[0067] In a feasible embodiment, S32: the step of obtaining the second speech score of each associated word based on the speech score calculation method corresponding to each edit distance type, the first speech score of each first word, and the phoneme sequence of each first word includes:
[0068] S324: If the edit distance type is phoneme insertion type, calculate the ratio of the number of inserted phonemes to the total number of phonemes in the associated phoneme sequence.
[0069] S325: Obtain the second deduction value based on the ratio of the number of inserted phonemes to the total number of phonemes in the associated phoneme sequence. The total number of phonemes in the associated phoneme sequence is the sum of the number of inserted phonemes and the total number of phonemes in the first word's phoneme sequence.
[0070] S326: The difference between the first phonetic score and the second deduction score of the first word is determined as the second phonetic score of the related word corresponding to the first word.
[0071] Among them, steps S324-S326 and steps S321-S323 can be independent steps in sequence.
[0072] Steps S324-S326 can be achieved using the following formulas:
[0073]
[0074] Among them, S n S is the second phonetic score of the associated word. o denoted as the first speech score, k as the number of inserted phonemes, m as the total number of phonemes in the first word phoneme sequence, and m+k as the total number of phonemes in the associated phoneme sequence.
[0075] In this embodiment, through steps S324-S326, the second speech score of the associated word with the edit distance type of inserted phoneme can be accurately obtained.
[0076] In one feasible embodiment, the speech evaluation data includes a first phoneme sequence score of a first word phoneme sequence; the first phoneme sequence score includes the scores of all phonemes in the first word phoneme sequence.
[0077] Voice assessment data augmentation methods also include:
[0078] S301: Obtain the second phoneme sequence score of the associated word based on the first phoneme sequence score and the edit distance type.
[0079] In this embodiment, the enhanced speech evaluation data can be further expanded by the second phoneme sequence score of the associated words. This allows the speech evaluation model trained based on the enhanced speech evaluation data to obtain the score data of each word in the evaluation speech data, the score data of the corresponding associated words, and the phoneme sequence score data of the associated words, thus obtaining more data evaluation results.
[0080] In a feasible embodiment, S301: The step of obtaining the second phoneme sequence score of the associated word based on the first phoneme sequence score and the edit distance type includes:
[0081] S3011: If the edit distance type is a deleted phoneme type or a replaced phoneme type, keep the scores of the undeleted or unreplaced phonemes in the associated phoneme sequence, set the score of the replaced phonemes to zero, and obtain the score of the second phoneme sequence based on the scores of each phoneme in the associated phoneme sequence.
[0082] In this embodiment, the second phoneme sequence score of the associated word can be accurately obtained through step S3011.
[0083] In a feasible embodiment, S301: The step of obtaining the second phoneme sequence score of the associated word based on the first phoneme sequence score and the edit distance type includes:
[0084] S3012: If the edit distance type is the inserted phoneme type, the score of the phoneme inserted in the associated phoneme sequence is set to zero, and the score of the second phoneme sequence is obtained based on the scores of each phoneme in the associated phoneme sequence.
[0085] In this embodiment, steps S3011 and S3012 can be independent steps. Step S3012 can accurately obtain the second phoneme sequence score of the associated word.
[0086] In a feasible embodiment, S4: the step of obtaining enhanced speech evaluation data based on the first word, the first speech score, the associated word, and the second speech score includes:
[0087] S41: Based on the first speech score, the second speech score, and multiple preset segments, group the first word and related words to obtain word groups belonging to each segment.
[0088] S42: From each word group, randomly select the same number of grouping elements and the corresponding score data of each grouping element to obtain the enhanced speech evaluation data.
[0089] A grouping element refers to a word in a word group, its corresponding phoneme sequence, its corresponding first and / or second speech scores, and / or its corresponding phoneme sequence score.
[0090] In this embodiment, grouping data according to scores and then randomly selecting the same number of group elements from each group can improve the data balance of the voice assessment data.
[0091] Please see Figure 3 The second embodiment of this application provides a voice assessment method, including:
[0092] S5: Train the initial assessment model based on the speech assessment data obtained by the speech assessment data augmentation method to obtain the speech assessment model.
[0093] The speech evaluation data augmentation method can be any feasible embodiment of the speech evaluation data augmentation method provided in the first embodiment of this application, and the speech evaluation data can be used to train the initial evaluation model using a deep learning model.
[0094] S6: Input the speech data and audio text to be evaluated into the speech evaluation model to obtain the speech evaluation results, which include the score data of each word.
[0095] In this embodiment, enriching the training data of the speech assessment model with enhanced speech assessment data can make the learning of the speech assessment model more complete.
[0096] Please see Figure 4 The third embodiment of this application provides a voice assessment data enhancement device 100, comprising:
[0097] The speech evaluation data acquisition module 101 is used to acquire the speech evaluation data to be enhanced; the speech evaluation data includes several first words and the first speech score of each first word, and the phoneme sequence of each first word;
[0098] The associated word acquisition module 102 is used to determine the associated phoneme sequence of each first word based on the phoneme sequence of each first word and the preset phoneme edit distance; and to determine the associated words of each first word and the corresponding edit distance type based on the associated phoneme sequence.
[0099] The second speech score acquisition module 103 is used to determine the second speech score of each associated word based on the first speech score of each first word, the phoneme sequence of each first word, the phoneme edit distance and the edit distance type.
[0100] The enhanced speech evaluation data module 104 is used to obtain enhanced speech evaluation data based on the first word, the first speech score, the associated words, and the second speech score.
[0101] It should be noted that the speech evaluation data enhancement device 100 provided in the third embodiment of this application is only illustrated by the above-described division of functional modules when executing the speech evaluation data enhancement method. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the speech evaluation data enhancement device 100 provided in the third embodiment of this application and the speech evaluation data enhancement method of the first embodiment of this application belong to the same concept, and its implementation process is detailed in the method embodiment, which will not be repeated here.
[0102] The fourth embodiment of this application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the speech evaluation data enhancement method described above.
[0103] The fifth embodiment of this application provides a computer device, including a storage device, a processor, and a computer program stored in the storage device and executable by the processor. When the processor executes the computer program, it implements the steps of the speech evaluation data enhancement method as described above.
[0104] The device embodiments described above are merely illustrative. The components described as separate parts may or may not be physically separate, and the components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this application according to actual needs. Those skilled in the art can understand and implement this without any inventive effort.
[0105] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0106] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 The computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1The function selected in one or more boxes.
[0107] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function selected in one or more boxes.
[0108] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0109] Memory may include non-persistent memory in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, like read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.
[0110] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.
[0111] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
[0112] The above are merely embodiments of this application and are not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.
Claims
1. A method for augmenting speech assessment data, characterized in that, include: Obtain speech evaluation data that needs improvement; The speech evaluation data includes several first words, the first speech score of each first word, and the phoneme sequence of each first word; Based on the phoneme sequence of each first word and a preset phoneme edit distance, the associated phoneme sequence of each first word is determined; based on the associated phoneme sequence, the associated words of each first word and the corresponding edit distance type are determined; the phoneme edit distance refers to the number of phonemes in the phoneme sequence that are changed due to editing; the edit distance type includes deleted phoneme type, inserted phoneme type, and replaced phoneme type; The second speech score of each of the associated words is determined based on the first speech score of each of the first words, the phoneme sequence of each of the first words, the phoneme edit distance, and the edit distance type; Enhanced speech evaluation data is obtained based on the first word, the first speech score, the associated word, and the second speech score.
2. The voice assessment data enhancement method according to claim 1, characterized in that, The step of determining the second speech score of each of the associated words based on the first speech score of each of the first words, the phoneme sequence of each of the first words, and the phoneme edit distance includes: Obtain the speech score calculation method corresponding to the edit distance type of each of the associated words; Based on the various speech score calculation methods, the first speech score of each first word, and the phoneme sequence of each first word, the second speech score of each associated word is obtained.
3. The voice assessment data enhancement method according to claim 2, characterized in that, The step of obtaining the second speech score of each of the associated words based on the speech score calculation methods, the first speech score of each of the first words, and the phoneme sequence of each of the first words includes: If the edit distance type is a deleted phoneme type or a replaced phoneme type, the ratio of the total score of the deleted or replaced phonemes to the total score of the phonemes in the first word phoneme sequence is determined as the score percentage of the deleted or replaced phonemes. The first deduction value is obtained based on the score ratio of the deleted or replaced phonemes and the first speech score; The difference between the first speech score of the first word and the first deduction score is determined as the second speech score of the associated word corresponding to the first word.
4. The voice assessment data enhancement method according to claim 2, characterized in that, The step of obtaining the second speech score of each of the associated words based on the speech score calculation methods, the first speech score of each of the first words, and the phoneme sequence of each of the first words includes: If the edit distance type is a phoneme insertion type, calculate the ratio of the number of inserted phonemes to the total number of phonemes in the associated phoneme sequence; The second deduction value is obtained based on the ratio of the number of inserted phonemes to the total number of phonemes in the associated phoneme sequence and the first speech score. The difference between the first speech score and the second deduction score of the first word is determined as the second speech score of the associated word corresponding to the first word.
5. The voice assessment data enhancement method according to claim 1, characterized in that, The speech evaluation data includes the first phoneme sequence score of the first word phoneme sequence; The first phoneme sequence score includes the scores of all phonemes in the first word phoneme sequence; The voice assessment data augmentation method also includes: Based on the first phoneme sequence score and the edit distance type, obtain the second phoneme sequence score of the associated word.
6. The voice assessment data enhancement method according to claim 5, characterized in that, The step of obtaining the second phoneme sequence score of the associated word based on the first phoneme sequence score and the edit distance type includes: If the edit distance type is a delete phoneme type or a replace phoneme type, the scores of the undeleted or unreplaced phonemes in the associated phoneme sequence are maintained, the scores of the replaced phonemes are determined to be zero, and the score of the second phoneme sequence is obtained based on the scores of each phoneme in the associated phoneme sequence.
7. The voice assessment data enhancement method according to claim 5, characterized in that, The step of obtaining the second phoneme sequence score of the associated word based on the first phoneme sequence score and the edit distance type includes: If the edit distance type is the inserted phoneme type, the score of the phoneme inserted in the associated phoneme sequence is determined to be zero, and the score of the second phoneme sequence is obtained based on the scores of each phoneme in the associated phoneme sequence.
8. The speech evaluation data enhancement method according to any one of claims 1-7, characterized in that, The step of obtaining enhanced speech evaluation data based on the first word, the first speech score, the associated word, and the second speech score includes: Based on the first speech score, the second speech score, and multiple preset segments, the first word and the associated words are grouped to obtain word groups belonging to each segment; From each of the word groups, the same number of grouping elements and the corresponding score data of each grouping element are randomly selected to obtain the enhanced speech evaluation data.
9. A voice assessment method, characterized in that, include: The speech evaluation data obtained by the speech evaluation data enhancement method according to any one of claims 1-8 is used to train an initial evaluation model to obtain a speech evaluation model; The speech data and audio text to be evaluated are input into the speech evaluation model to obtain the speech evaluation results.
10. A voice assessment data enhancement device, characterized in that, include: The module for acquiring speech evaluation data to be enhanced is used to acquire speech evaluation data to be enhanced. The speech evaluation data includes several first words, the first speech score of each first word, and the phoneme sequence of each first word; The associated word acquisition module is used to determine the associated phoneme sequence of each first word based on the phoneme sequence of each first word and a preset phoneme edit distance; and to determine the associated words of each first word and the corresponding edit distance type based on the associated phoneme sequence. The second speech score acquisition module is used to determine the second speech score of each of the associated words based on the first speech score of each of the first words, the phoneme sequence of each of the first words, the phoneme edit distance, and the edit distance type. The enhanced speech evaluation data module is used to obtain enhanced speech evaluation data based on the first word, the first speech score, the associated word, and the second speech score.
11. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by the processor, it implements the steps of the voice assessment data enhancement method as described in any one of claims 1 to 8.
12. A computer device, characterized in that: It includes a storage device, a processor, and a computer program stored in the storage device and executable by the processor, wherein the processor executes the computer program to implement the steps of the speech evaluation data enhancement method as described in any one of claims 1 to 8.
Citation Information
Patent Citations
Oral language scoring method and device
CN107958673A
Voice data analysis method and system
CN111583908A