Sample set generation method, device and computer equipment for training speech recognition model

By generating heavy accent pinyin sequences through decoding and encoding, and utilizing text-to-speech tools and pinyin sequence filtering rules, the problem of insufficient heavy accent samples was solved, resulting in a high-quality sample set and improving the speech recognition model's ability to recognize heavy accents.

CN121438814BActive Publication Date: 2026-04-14深圳市友杰智新科技有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-12-30
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

Existing technologies suffer from insufficient sample quantity, low accuracy, and lack of diversity when dealing with heavy accent recognition, resulting in insufficient recognition ability of speech recognition models in non-standard pronunciation scenarios.

Method used

By decoding the target command word into a toneless original pinyin sequence, encoding and converting it using a pre-defined heavy accent rule library to generate a heavy accent pinyin sequence, generating audio using a text-to-speech generation tool, and filtering out the heavy accent audio that meets the requirements using pinyin sequence-based filtering rules, a sample set is constructed for training the speech recognition model.

Benefits of technology

It enables the low-cost and efficient generation of high-quality audio sample sets with heavy accents, covering mainstream heavy accent types, improving the recognition accuracy and adaptability of speech recognition models for heavy accents, lowering the technical threshold, and improving the coverage and relevance of the sample set.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121438814B_ABST
    Figure CN121438814B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of speech recognition, and aims to solve the problem of lack of training samples of heavy accent speech recognition model. A method, device and computer equipment for generating a sample set for training a speech recognition model are provided, wherein the method comprises: decoding a target command word into a toneless original pinyin sequence; based on a heavy accent rule library (heavy accent refers to non-standard pronunciation of initial / final change) constructed according to common non-standard pronunciation rules, a heavy accent pinyin sequence is generated through coding conversion; a heavy accent audio is generated through a text-to-speech audio generation tool; an audio input is input into a preset recognition model to obtain a recognition result and convert it into a recognized pinyin sequence; a screening rule is constructed based on the original and / or heavy accent pinyin, and the recognized pinyin sequence is compared to screen the audio; and the audio meeting the requirements is collected to form a training sample set. Through rule-based generation and accurate screening, high-quality heavy accent samples can be efficiently obtained, and the heavy accent recognition performance of the model can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of speech recognition technology, and in particular to a method, apparatus, and computer device for generating a sample set for training a speech recognition model. Background Technology

[0002] Speech recognition technology, as one of the core supporting technologies in the field of human-computer interaction, directly determines the quality of the interactive experience through its recognition accuracy. Current mainstream speech recognition models rely on large-scale audio data training to improve command word recognition capabilities; the higher the consistency between the training samples and the target user's pronunciation, the better the model's recognition performance. However, in real-world applications, some users' Mandarin pronunciation exhibits significant non-standard features—these pronunciations are not simply differences in tone, but rather substantial changes in the initials or finals of the Chinese characters in the command words (this paper defines such non-standard pronunciations as "heavy accents"), such as the common confusion between l and r, or the mixing of retroflex and alveolar consonants. Because these heavy accent pronunciations differ significantly from the standard training samples of existing speech recognition models, the recognition accuracy of these pronunciations drops sharply, severely impacting the technology's universality.

[0003] To address the challenge of heavy accent recognition, existing technologies have developed two main approaches: First, there are specialized dialect recognition models. These models are custom-developed for specific regional dialect characteristics, but they have significant limitations: on the one hand, the model structure needs to be adjusted individually for each region's dialect characteristics, resulting in extremely high adaptation costs; on the other hand, model training requires massive amounts of dialect audio data, leading to long and costly data collection and annotation cycles, making them too difficult to use and unsuitable for heavy accent recognition in common scenarios. Second, there is manual screening of heavy accent audio. Since the demand for heavy accent data corresponding to a single command word is usually small, manual screening becomes a choice in some scenarios. However, this method has significant drawbacks: manual screening is inefficient and difficult to process in batches; furthermore, if not properly controlled during the screening process, the number of heavy accent audio samples used for training can easily exceed the number of standard Mandarin audio samples, leading to overlearning of heavy accent features and consequently reducing the model's ability to recognize standard pronunciation.

[0004] Besides the two types of solutions mentioned above, existing core methods for audio filtering mostly rely on large speech recognition models with language modules (such as Whisper, Funasr, etc.). Their filtering logic involves back-matching the corresponding audio with the Chinese characters output by the model. However, this filtering method has fundamental flaws when dealing with audio with heavy accents: First, the pronunciation rules of heavy accents are diverse and uncertain, lacking a unified pronunciation paradigm as a filtering basis. This means that filtering rules need to be designed separately for each command word, with no reusability between rules, requiring repeated development when adapting to new command words. Second, the recognition results of large speech recognition models for the same audio with heavy accents are unstable, and the Chinese characters output each time may differ, making it difficult to maintain the Chinese character-based filtering rules and ensuring filtering accuracy.

[0005] In summary, existing technologies in the field of heavy accent recognition face a triple dilemma: high barriers to entry for dedicated models, low efficiency of manual screening, and difficulty in reusing existing screening rules. This results in insufficient quantity, low accuracy, and a lack of diversity in the acquisition of heavy accent audio samples, failing to provide a high-quality sample set for training speech recognition models and hindering the improvement of models' ability to recognize heavy accent pronunciation. Therefore, there is an urgent need for a low-cost, efficient, reusable, and stable heavy accent audio screening technology to address the pain points of existing technologies and promote the universal application of speech recognition technology in non-standard pronunciation scenarios. Summary of the Invention

[0006] This invention provides a method, apparatus, and device for generating a sample set for training a speech recognition model, aiming to solve the technical problems of "insufficient quantity, low accuracy, and lack of diversity" in the acquisition of audio samples with heavy accents.

[0007] To achieve the aforementioned objectives, the first aspect of this invention proposes a method for generating a sample set for training a speech recognition model, comprising:

[0008] Decode the target command word into a toneless raw phonetic sequence;

[0009] Based on the pronunciation replacement rules in the preset heavy accent rule library, the original pinyin sequence is encoded and converted to generate at least one set of heavy accent pinyin sequence; wherein, the heavy accent refers to non-standard pronunciation with changes in initials or finals, and the heavy accent rule library is constructed based on common non-standard pronunciation rules;

[0010] Input the heavy accent pinyin sequence into the text into the speech and audio generation tool to generate the corresponding heavy accent audio;

[0011] The audio with the heavy accent is input into a preset speech recognition model to obtain the recognition result corresponding to the audio with the heavy accent, and the recognition result is converted into the corresponding recognition pinyin sequence;

[0012] The audio with a heavy accent is filtered based on preset filtering rules. The filtering rules use the pinyin sequence as the matching basis to construct a matching benchmark. The matching benchmark is generated based on the original pinyin sequence and / or the pinyin sequence with a heavy accent, and is compared with the recognized pinyin sequence.

[0013] Collect audio samples with strong accents that meet the screening requirements to form a sample set for training the speech recognition model.

[0014] Furthermore, in the accent rule base, each set of accent pinyin sequences corresponds to a unique accent pronunciation rule; and a target command word can generate zero or one or more accent pinyin sequences.

[0015] Furthermore, the filtering rules include at least one of the following: full matching of accented pronunciations, matching of standard pronunciations, matching of mixed encodings, and matching of initial and final similarities.

[0016] Furthermore, the specific process of initial consonant-final similarity matching includes:

[0017] Audio samples with heavy accents whose pinyin sequences match the number of characters in the target command word were selected as candidate audio samples.

[0018] For each character in the identified pinyin sequence of the candidate audio, calculate the similarity score between its initial consonant and final vowel and the initial consonant and final vowel of the character corresponding to the target command word.

[0019] Audio with heavy accents that have a total similarity score higher than a preset threshold are retained.

[0020] Furthermore, the specific process of the hybrid encoding matching includes:

[0021] Based on the original pinyin sequence and at least one set of accented pinyin sequences, a candidate list of codes containing multiple pronunciation variations is generated;

[0022] The identified pinyin sequence is compared with the coding candidate list;

[0023] Retain accented audio that exactly matches any of the encoding candidate lists.

[0024] Furthermore, the specific process of the standard pronunciation association matching includes:

[0025] The recognition results are converted into Chinese character sequences;

[0026] If the Chinese character sequence is completely consistent with the target command word, then the recognized pinyin sequence corresponding to the recognition result is compared with the original pinyin sequence;

[0027] Preserve the accented audio where the initial consonant and final vowel are completely identical except for the tone.

[0028] Further, the sample set used to train the speech recognition model includes:

[0029] Obtain an existing sample set of standard Mandarin pronunciation audio containing the target command word;

[0030] The selected audio samples with strong accents that meet the requirements are combined with the existing sample set to form an enhanced training sample set, which is used to train the speech recognition model.

[0031] A second aspect of the present invention provides a sample set generation apparatus for training a speech recognition model, comprising:

[0032] The decoding unit is used to decode the target command word into a toneless original pinyin sequence;

[0033] The conversion unit is used to encode and convert the original pinyin sequence based on the pronunciation replacement rules in the preset heavy accent rule library to generate at least one set of heavy accent pinyin sequences; wherein, the heavy accent refers to non-standard pronunciation with changes in the initial consonant or final vowel, and the heavy accent rule library is constructed based on common non-standard pronunciation rules;

[0034] The generation unit is used to input the heavy accent pinyin sequence into the text into the speech audio generation tool and generate the corresponding heavy accent audio.

[0035] The recognition unit is used to input the accented audio into a preset speech recognition model, obtain the recognition result corresponding to the accented audio, and convert the recognition result into a corresponding recognition pinyin sequence;

[0036] A filtering unit is used to filter the accented audio based on preset filtering rules. The filtering rules use the pinyin sequence as the matching basis to construct a matching benchmark. The matching benchmark is generated based on the original pinyin sequence and / or the accented pinyin sequence and compared with the recognized pinyin sequence.

[0037] The combination unit is used to collect filtered audio samples with strong accents to form a sample set for training the speech recognition model.

[0038] A third aspect of the present invention provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of the sample set generation method for training a speech recognition model as described in any of the preceding claims.

[0039] A fourth aspect of the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the steps of the sample set generation method for training a speech recognition model as described in any of the preceding claims.

[0040] Beneficial effects:

[0041] This invention provides a method, apparatus, and computer device for generating sample sets for training speech recognition models. The method effectively addresses the core pain points of existing heavy accent training sample set generation: high cost, limited coverage, and uncontrollable quality. It eliminates the need for manual recording of heavy accent audio, significantly reducing sample acquisition costs through an automated process of pinyin encoding conversion and text-to-speech generation. Furthermore, it enables batch sample generation, significantly improving efficiency. The sample coverage and targeting are enhanced: the heavy accent rule base is built based on common non-standard pronunciation rules. The generated heavy accent pinyin sequences accurately match the heavy accent features of initial / final changes in real-world scenarios, ensuring sample coverage of mainstream heavy accent types, avoiding invalid sample generation, and improving the adaptability of the sample set to real-world application scenarios. It guarantees sample set quality and training effectiveness: using the original pinyin sequence and / or heavy accent pinyin sequence as matching benchmarks, quantitative screening is achieved through pinyin sequence comparison. This accurately retains samples that reflect heavy accent features and provide effective recognition feedback, while eliminating invalid samples. This provides high-quality training data for the speech recognition model, helping to improve the model's recognition accuracy for heavy accent speech. Attached Figure Description

[0042] Figure 1 A flowchart illustrating a method for generating a sample set for training a speech recognition model according to an embodiment of the invention;

[0043] Figure 2 A schematic diagram of a sample set generation device for training a speech recognition model according to an embodiment of the invention;

[0044] Figure 3 This is a schematic diagram of the structure of a computer device according to an embodiment of the invention.

[0045] The realization of the objective, functional features and advantages of the present invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0046] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0047] Those skilled in the art of the present technology can understand that, unless specifically stated, the singular forms "a", "an", "the above" and "the" used herein may also include the plural forms. It should be further understood that the term "comprising" used in the specification of the present invention means the presence of features, integers, steps, operations, elements, modules and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, modules, components and / or their groups. It should be understood that when we say that an element is "connected" or "coupled" to another element, it can be directly connected or coupled to other elements, or there may also be intermediate elements. In addition, the "connection" or "coupling" used herein may include wireless connection or wireless coupling. The term "and / or" used herein includes all or any module and all combinations of one or more related listed items.

[0048] Those skilled in the art of the present technology can understand that, unless otherwise defined, all terms (including technical terms and scientific terms) used herein have the same meaning as the general understanding of those of ordinary skill in the art to which the present invention pertains. It should also be understood that terms such as those defined in a general dictionary should be understood to have a meaning consistent with the meaning in the context of the prior art, and will not be interpreted with an idealized or overly formal meaning unless specifically defined as here.

[0049] Referring to Figure 1 , an embodiment of the present invention provides a method for generating a sample set for training a speech recognition model, including:

[0050] S1: Decode the target command word into a raw pinyin sequence without tones.

[0051] The target command word refers to a specific vocabulary that the speech recognition model needs to learn to recognize, and is a core instruction preset in a human-computer interaction scenario (such as "start the device", "close the program", "adjust the volume", etc.). The raw pinyin sequence without tones refers to a pinyin combination sequence obtained by stripping the tone information of each Chinese character in the target command word and only retaining the initials and finals, and is the basis for subsequent heavy accent encoding conversion.

[0052] For example, select the target command word "people". First, disassemble it through a pinyin decoding algorithm: the standard pinyin of "人" is "rén", and after stripping the tone, it is "ren"; the standard pinyin of "民" is "mín", and after stripping the tone, it is "min". Finally, the raw pinyin sequence without tones "renmin" is obtained. The decoding process only focuses on the initials and finals, excluding the interference of tone differences on heavy accent recognition - because the core of the heavy accent defined in the present invention is the change of initials or finals, and tone changes do not belong to the category of heavy accents. This decoding method can accurately lock the core dimension of heavy accent screening.

[0053] The standard form for screening heavy accents was clarified, avoiding misjudging tone differences as "heavy accents" and ensuring that subsequent encoding conversion and screening are based on core features, thereby improving the accuracy of the technical solution. At the same time, the unified format of toneless pinyin sequences provides a standardized method for decoding different command words, enhancing the reusability of the solution.

[0054] S2: Based on the pronunciation replacement rules in the preset heavy accent rule library, the original pinyin sequence is encoded and converted to generate at least one set of heavy accent pinyin sequences; wherein, the heavy accent refers to non-standard pronunciation with changes in initials or finals, and the heavy accent rule library is constructed based on common non-standard pronunciation rules.

[0055] The heavy accent rule base is a set of rules built upon common non-standard pronunciation patterns from various regions of China (such as confusion between l and r, mixing of retroflex and alveolar consonants, and confusion of front and back nasal consonants). It includes one-to-one substitution rules between the original phoneme and the target phoneme, with each substitution rule corresponding to only one type of heavy accent. The heavy accent pinyin sequence is a pinyin combination formed by replacing some or all phonemes in the original pinyin sequence using the substitution rules from the heavy accent rule base, corresponding to a specific heavy accent pronunciation morphology.

[0056] For example, the preset heavy accent rule base includes a replacement rule for "l / r indistinguishable" (rule identifier: R→L), which means that all phonemes containing the initial consonant "r" are uniformly replaced with "l". For the original pinyin sequence "renmin", after triggering this replacement rule, the initial consonant "r" in "ren" is replaced with "l", resulting in "len". "Min" remains unchanged as there is no corresponding replacement rule, ultimately generating the heavy accent pinyin sequence "lenmin". This process follows the "unique mapping" principle of the rule base: the "l / r indistinguishable" rule only generates the heavy accent pinyin sequence "lenmin", and will not produce multiple variations due to the same rule; if the target command word is "Beijing" (original pinyin sequence "beijing"), which has no common initial / final change patterns, then no heavy accent pinyin sequence will be generated.

[0057] The standardized generation of heavy accent pinyin sequences is achieved through a heavy accent rule library, which solves the problem of diverse heavy accent pronunciation rules and lack of a unified paradigm in the existing technology; the "unique mapping" rule ensures the consistency between generation and subsequent filtering, avoiding filtering chaos caused by multiple rule conflicts; at the same time, it supports the flexible design of "zero generation", adapting to command words without common heavy accents, and improving the versatility of the solution.

[0058] S3: Input the heavy accent pinyin sequence into the text into the speech audio generation tool to generate the corresponding heavy accent audio.

[0059] Text-to-Speech audio generation tool (TTS tool), a tool with the function of directly driving audio synthesis by pinyin, which supports skipping the Chinese decoding link and directly injecting the pinyin sequence into the synthesis module (synthetic module) to generate audio.

[0060] For example, input the strong-accent pinyin sequence "lenmin" into the TTS tool. When operating, skip the Chinese decoding step自带的中文解码步骤 in the tool and directly inject "lenmin" into the synthesis module. Based on the phoneme characteristics of the pinyin sequence, the TTS tool generates an audio file with the pronunciation of "lenmin" - the pronunciation of this audio corresponds to the strong accent of "confusing l / r", which is exactly the same as the defined strong accent type in the rule library, and there is no pronunciation deviation due to the Chinese decoding of TTS.

[0061] Skipping the Chinese decoding step avoids the problem that the TTS tool misinterprets strong-accent pinyin as standard pinyin, ensuring that the generated audio accurately matches the target strong accent type; the batch generation ability of the TTS tool solves the pain points of low efficiency and high cost in manually recording strong-accent audio, realizing the large-scale output of strong-accent audio.

[0062] S4: Input the strong-accent audio into a preset speech recognition model, obtain the recognition result corresponding to the strong-accent audio, and convert the recognition result into a corresponding recognized pinyin sequence.

[0063] The preset speech recognition model is an open-source Mandarin speech recognition model (such as whisper, funasr), which does not need to be fine-tuned for strong accents, and its built-in language model can achieve unified encoding of words with similar pronunciations. The recognition result refers to the recognition output of the speech recognition model for strong-accent audio, usually presented in the form of a Chinese character sequence. The recognized pinyin sequence refers to decoding the Chinese character sequence of the recognition result into a non-tonal pinyin sequence for subsequent comparison with the benchmark pinyin sequence.

[0064] For example, input the generated "lenmin" strong-accent audio into the open-source Mandarin recognition model whisper. Through the fault-tolerant mechanism of its built-in language model, the model recognizes "lenmin" as the Chinese character sequence "人民"; then decode this Chinese character sequence into the non-tonal recognized pinyin sequence "renmin". In the whole process, no fine-tuning is performed on the whisper model, and only relying on its native Mandarin recognition ability, the recognition and pinyin sequence conversion of strong-accent audio can be completed.

[0065] Reusing the open-source Mandarin recognition model, without customizing a dialect model or fine-tuning the existing model, greatly reduces the technical usage threshold; the step of converting the recognition result to a pinyin sequence changes the "screening based on Chinese characters" in the existing technology to "centering on pinyin", solving the problem of difficult maintenance of screening rules caused by unstable Chinese character recognition results.

[0066] S5: The audio with a heavy accent is filtered based on a preset filtering rule. The filtering rule uses the pinyin sequence as the matching basis to construct a matching benchmark. The matching benchmark is generated based on the original pinyin sequence and / or the heavy accent pinyin sequence and is compared with the recognized pinyin sequence.

[0067] The filtering rules are a multi-dimensional filtering mechanism based on pinyin sequence comparison, including at least one of the following: complete matching of accented pronunciations, matching of standard pronunciations, matching of mixed codes, and matching of initial and final similarities. The matching benchmark is a reference standard used for comparison, which can be constructed based solely on the original pinyin sequence, solely on the pinyin sequence with accented pronunciations, or based on a cross-combination of both.

[0068] For example, this embodiment uses both "exact match for heavy accents" and "associated match for standard pronunciation" as filtering rules to construct corresponding matching benchmarks:

[0069] The benchmark for a perfect match of accented pronunciations is the accented pronunciation sequence “lenmin”.

[0070] Standard pronunciation association matching benchmark: original pinyin sequence "renmin".

[0071] First, a full match of accented words is performed: the recognized pinyin sequence "renmin" is compared with the benchmark "lenmin". Since the initial consonants of "ren" and "len" are different, the comparison result is a mismatch, and the audio will not be included in this type of screening results.

[0072] Next, standard pronunciation association matching is performed: the recognized pinyin sequence "renmin" is compared with the benchmark "renmin". If the initial consonant and final vowel are completely consistent, the comparison result is a match, and the audio is included in this type of filtering results.

[0073] The results of the two filtering rules are independent and returned to the user separately, allowing the user to choose whether to keep the audio.

[0074] The design of multi-dimensional filtering rules ensures both the accuracy of filtering (such as complete matching of accents) and the comprehensiveness of filtering (such as matching of standard pronunciations); the flexible construction method of the matching benchmark covers filtering scenarios with any single benchmark or combination of benchmarks, solving the problem of non-reusability of existing filtering rules; the design of independently returning filtering results avoids data redundancy, gives users the right to choose, and improves the flexibility of the solution.

[0075] S6: Collect audio samples with heavy accents that meet the requirements after screening, and form a sample set for training the speech recognition model.

[0076] The sample set, an enhanced dataset containing standard pronunciation audio and filtered accented audio, is used for training or validation of speech recognition models.

[0077] For example, collect audio clips of the word "lenmin" that have passed the standard pronunciation association matching screening and label them with the heavy accent type "l / r not distinguished" (the label is determined based on the matching screening rules and heavy accent pronunciation rules); obtain the existing standard sample set, which contains audio clips of the standard Mandarin pronunciation of "renmin" (corresponding to the pinyin sequence "renmin"); merge the labeled heavy accent audio clips with the existing standard sample set to form an enhanced training sample set; among them, heavy accent audio clips with large pronunciation deviations (such as "linmin") are separately assigned to the validation set to evaluate the model's heavy accent recognition limit.

[0078] The combination of "standard + heavy accent" in the sample set solves the problem of insufficient diversity in existing training data and provides the model with comprehensive coverage of pronunciation scenarios; the heavy accent type label enriches the sample information, making it easier for the model to learn the features of different types of heavy accents; the separate division of the validation set can guide the model training direction, evaluate the model's capabilities, and further improve the robustness of the speech recognition model.

[0079] In this embodiment, no custom dialect model or fine-tuning of the existing model is required; only the open-source Mandarin recognition model and TTS tool are reused, significantly reducing the technical threshold and cost. The introduction of the heavy accent rule base unifies the paradigm of heavy accent generation and screening, solving the problems of non-reusability and difficulty in maintenance of existing screening rules. Multi-dimensional screening rules and independent return design take into account both the accuracy and comprehensiveness of the samples, giving users flexible choices and avoiding data redundancy. The enhanced sample set includes standard pronunciation and various heavy accent pronunciations, and comes with heavy accent type labels, enriching the diversity and practicality of the training data and effectively improving the speech recognition model's ability to recognize non-standard pronunciations. The division of the validation set realizes the linkage between model training and evaluation, which can guide the direction of model optimization and further improve the robustness and universality of the model.

[0080] In one embodiment, in the accent rule base, each set of accent pinyin sequences corresponds to a unique accent pronunciation rule; and a target command word can generate zero or one or more accent pinyin sequences.

[0081] The unique pronunciation rule for accented sounds means that each set of accented pinyin sequences is generated by only one pronunciation substitution rule. Different accented pinyin sequences correspond to different pronunciation substitution rules, with no rule overlap or repetition mapping.

[0082] For example, the heavy accent rule base contains three independent pronunciation substitution rules: rule 1 (R→L, l / r are not distinguished), rule 2 (ZH→Z, retroflex consonants are confused), and rule 3 (ENG→EN, nasal consonants are confused).

[0083] For the target command word "支持" (original pinyin sequence "zhichi"):

[0084] When trigger rule 2 (ZH→Z), "zhi" is replaced by "zi" and "chi" is replaced by "ci", generating the heavy accent pinyin sequence "zici" - this sequence only corresponds to rule 2 and no other rule can generate the same sequence;

[0085] When trigger rule 3 (ENG→EN), there is no corresponding ENG phoneme for "zhichi", so the sequence corresponding to this rule is not generated;

[0086] Therefore, "支持" only generates a group of heavy accent pinyin sequences "zici".

[0087] For the target command word "城市" (original pinyin sequence "chengshi"):

[0088] When trigger rule 3 (ENG→EN), "cheng" is replaced by "chen", generating the heavy accent pinyin sequence "chenshi";

[0089] When trigger rule 2 (ZH→Z), there is no corresponding ZH phoneme for "shi", so the sequence corresponding to this rule is not generated;

[0090] Therefore, "城市" only generates a group of heavy accent pinyin sequences "chenshi".

[0091] The unique mapping relationship between rules and sequences ensures the accurate distinction of heavy accent types, avoiding screening confusion caused by different rules generating the same sequence; at the same time, the independence of rules facilitates the subsequent maintenance and iteration of the rule library, improving the scalability of the solution.

[0092] The flexibility of generating heavy accent pinyin sequences for target command words. For example:

[0093] Case 1: For the target command word "北京" (original pinyin sequence "beijing"), there is no rule for replacing the initial / final consonants of "bei" and "jing" in the heavy accent rule library, so no heavy accent pinyin sequence is generated (zero generation);

[0094] Case 2: For the target command word "吃饭" (original pinyin sequence "chifan"), only trigger rule 2 (ZH→Z), "chi" is replaced by "ci", generating the heavy accent pinyin sequence "cifan" (one generation);

[0095] Case 3: For the target command word "认真" (original pinyin sequence "renzhen"), rule 1 (R→L) and rule 2 (ZH→Z) can be triggered:

[0096] Rule 1 is triggered: "ren" is replaced with "len", generating the sequence "lenzhen";

[0097] Rule 2 is triggered: "zhen" is replaced with "zen", generating the sequence "renzen";

[0098] Therefore, "seriously" generates two sets of heavy accent pinyin sequences (multiple generation).

[0099] It supports flexible design of "zero, one, or multiple" generation results, adapting to the accent features of different command words. Command words without common accents do not need to be generated, avoiding invalid data redundancy. Command words with multiple accent features can be fully covered, ensuring the diversity of the sample set. At the same time, the controllability of the number of generated words avoids the problem of model overlearning caused by excessive accent audio.

[0100] In this embodiment, the "unique mapping" characteristic of the heavy accent rule base and the "flexible generation" characteristic of the target command word are further defined. The unique mapping rule avoids confusion of heavy accent types, improves the accuracy of heavy accent generation, and reduces the complexity of subsequent screening; the flexible generation design adapts to the heavy accent features of different command words, avoiding invalid data redundancy and the risk of model overlearning; the standardized design of the rule base facilitates the addition, modification, and replacement of rules in the future, improving the scalability and maintainability of the solution; the accurate and diverse heavy accent pinyin sequences provide a reliable guarantee for the subsequent generation of high-quality heavy accent audio and sample sets, further improving the training effect of the speech recognition model.

[0101] In one embodiment, the filtering rules include at least one of the following: full matching of accented pronunciations, matching of standard pronunciations, matching of mixed encodings, and matching of initial and final similarities.

[0102] Complete matching of heavy accents means that the heavy accent pinyin sequence is used as the matching benchmark, and only the heavy accent audio that is completely consistent with the benchmark pinyin sequence is retained.

[0103] Standard pronunciation association matching refers to using the original pinyin sequence as the matching benchmark and retaining the accented audio that is completely consistent with the pinyin sequence of the benchmark.

[0104] Hybrid encoding matching refers to using the cross combination of the original pinyin sequence and the heavy accent pinyin sequence as the matching benchmark, while retaining the heavy accent audio that matches any element in the pinyin sequence and the combination benchmark.

[0105] Initial and final similarity matching refers to first screening and identifying audio sequences with the same number of characters as the target command word, and then scoring the initial and final similarity of each character, retaining audio with a total score higher than the threshold.

[0106] For example, select the target command word "hot water" (original pinyin sequence "reshui"), and generate two sets of accented pinyin sequences based on the accented rule library: Sequence 1 (leshui, rule R→L), Sequence 2 (resui, rule UI→UI, with slight changes in the finals); generate the corresponding two sets of accented audio A (leshui) and audio B (resui) through the TTS tool; input the two sets of audio into the whisper model to obtain the recognition results and recognition pinyin sequences:

[0107] Recognition result of audio A: "hot water", recognition pinyin sequence "reshui";

[0108] Recognition result of audio B: "hot water", recognition pinyin sequence "resui".

[0109] Apply four screening rules respectively:

[0110] Accented exact match: The benchmarks are Sequence 1 "leshui" and Sequence 2 "resui"; the recognized pinyin "reshui" of audio A does not match Sequence 1, and the recognized pinyin "resui" of audio B matches Sequence 2, so retain audio B;

[0111] Standard pronunciation associated match: The benchmark is the original pinyin "reshui"; the recognized pinyin "reshui" of audio A matches the benchmark, so retain audio A;

[0112] Mixed coding match: The benchmark is the cross-combination list {"reshui", "leshui", "resui", "lesui"}; both "reshui" of audio A and "resui" of audio B are in the list, so both are retained;

[0113] Initial and final similarity match: The preset threshold is 0.8; the similarity of each character between the recognized pinyin "reshui" of audio A and the original pinyin "reshui": "re" (1.0), "shui" (1.0), total score 2.0≥0.8, so retain; the similarity of each character between "resui" of audio B and the original pinyin "reshui": "re" (1.0), "sui" and "shui" (0.9), total score 1.9≥0.8, so retain.

[0114] The four screening rules cover the accented screening scenarios from different dimensions: The accented exact match ensures the accuracy of the samples, the standard pronunciation associated match covers common accents, the mixed coding match improves the comprehensiveness of screening, and the initial and final similarity match supplements marginal samples; the multi-rule optional design allows users to flexibly combine according to their needs, solving the problems of single screening dimension and difficulty in balancing quality and quantity in the existing technology.

[0115] In this embodiment, four specific types of screening rules and application methods are defined: through multi-dimensional design of exact match of strong accents, associated match of standard pronunciations, hybrid coding match, and similarity match of initials and finals, a hierarchical screening logic from "precision screening" to "full coverage" and then to "marginal supplementation" is achieved. The multi-dimensional screening rules cover strong-accented audio of different qualities and types, ensuring both the precision of core samples and the diversity and comprehensiveness of samples; the screening rules are all based on phonetic sequence comparison, continuing the "phonetic-core" design and avoiding screening errors caused by unstable Chinese character recognition; the flexible design of the rules can be adapted to different scenario requirements (such as choosing exact match of strong accents when pursuing high quality and choosing hybrid coding match when pursuing quantity), enhancing the practicality of the solution; the hierarchical screening logic can effectively filter out invalid audio, improve the quality of the sample set, and provide strong support for the efficient training of speech recognition models.

[0116] In one embodiment, the specific process of the similarity match of initials and finals includes:

[0117] S511: Screen out strong-accented audio with the same number of characters in the recognized phonetic sequence as the target command word as candidate audio.

[0118] Candidate audio refers to strong-accented audio with the same number of characters in the recognized phonetic sequence as the target command word, excluding invalid audio with deviations in the number of characters in the recognition result caused by acoustic feature interference.

[0119] For example, when the target command word is "mobile phone" (2 characters, original phonetic sequence "shouji"), among the generated strong-accented audio:

[0120] Audio C: Recognition result "mobile phone", recognized phonetic sequence "souji" (2 characters), with the same number of characters, included in the candidate audio;

[0121] Audio D: Recognition result "hand", recognized phonetic sequence "sou" (1 character), with different numbers of characters, excluded;

[0122] Audio E: Recognition result "mobile phone case", recognized phonetic sequence "soujike" (3 characters), with different numbers of characters, excluded.

[0123] Finally, only Audio C is retained as the candidate audio.

[0124] Based on the characteristics of Chinese pronunciation (strong accents usually do not change the number of characters), a large number of irrelevant audio (such as deviations in the number of characters caused by recognition errors) can be quickly excluded through character number screening, reducing the complexity of subsequent similarity calculation and enhancing the screening efficiency.

[0125] S512: For each Chinese initial and final of each character in the recognized pinyin sequence of the candidate audio, calculate the similarity score between it and the Chinese initial and final of the character at the corresponding position in the target command word respectively.

[0126] The similarity score is a quantitative value of the matching degree calculated based on the phoneme pronunciation characteristics, with a value range of 0 - 1. The higher the score, the higher the matching degree of the Chinese initials and finals (for example, the initials "s" and "sh" are similar in pronunciation, and the similarity score is 0.9; the finals "ai" and "ei" have a large pronunciation difference, and the similarity score is 0.3).

[0127] For example, the preset similarity scoring standard for Chinese initials and finals: a perfect match gets 1.0 points, similar pronunciation gets 0.8 - 0.9 points, a large pronunciation difference gets 0.1 - 0.7 points, and completely different gets 0 points.

[0128] For the recognized pinyin sequence "souji" of the candidate audio C and the original pinyin sequence "shouji" of the target command word "mobile phone":

[0129] The first character: The initial "s" of the recognized pinyin "sou" is similar in pronunciation to the initial "sh" of the original pinyin "shou", getting 0.9 points; the final "ou" is exactly the same as "ou", getting 1.0 points; the total score of this character is 1.9 points;

[0130] The second character: The initial "j" of the recognized pinyin "ji" is exactly the same as the initial "j" of the original pinyin "ji", getting 1.0 points; the final "i" is exactly the same as "i", getting 1.0 points; the total score of this character is 2.0 points;

[0131] The total similarity score of audio C is 1.9 + 2.0 = 3.9 points.

[0132] The quantitative method of scoring Chinese initials and finals word by word makes the similarity judgment more objective and accurate, avoiding the screening error caused by subjective judgment; the scoring standard is designed based on Chinese pronunciation characteristics, ensuring that the scoring result fits the actual strong accent scenario and improving the rationality of screening.

[0133] S513: Retain the strong accent audio with a total similarity score higher than the preset threshold.

[0134] The preset total similarity threshold is 3.0 points. The total score of audio C is 3.9 points ≥ 3.0 points, so the audio is initially retained; since the screening of Chinese initials and finals similarity may include a small number of audio with incorrect pronunciations, the initially retained audio C is manually confirmed again: play audio C, and confirm that its pronunciation is "souji" (a strong accent with confusion between flat and retroflex sounds), and there is no incorrect pronunciation, so the audio is finally retained.

[0135] Threshold screening enables precise screening of marginal samples, retaining valid audio with heavy accents while excluding invalid audio with extremely low similarity; the manual secondary confirmation step further improves sample quality, avoids erroneous audio from being mixed into the sample set, and solves the problem of "fuzzy matching error" that may exist in the screening of initials and finals similarity.

[0136] This embodiment details the complete process of initial consonant and final vowel similarity matching: first, invalid audio is eliminated through character count screening; then, objective similarity judgment is achieved through character-by-character initial and final vowel quantification scoring; finally, sample quality is ensured through threshold screening and manual secondary confirmation. This process fully utilizes the characteristics of Chinese pronunciation (stable character count and quantifiable initial and final vowel differences), solving the recognition error problem caused by acoustic feature interference and supplementing marginal accent samples that are difficult to cover by other screening rules. Among them, the character count screening step quickly eliminates irrelevant audio, improving screening efficiency; the quantification scoring method makes the similarity judgment more objective and accurate, avoiding subjective errors; the combination of threshold screening and manual secondary confirmation balances screening efficiency and sample quality, ensuring the effectiveness of marginal samples; this screening rule supplements "slightly heavy accent" and "marginal heavy accent" samples that are difficult to cover by other rules, further enriching the diversity of the sample set and helping the model improve its ability to recognize complex heavy accents.

[0137] In one embodiment, the specific process of hybrid encoding matching includes:

[0138] S521: Based on the original pinyin sequence and at least one set of heavy-accent pinyin sequences, generate a candidate list of codes containing multiple pronunciation variations.

[0139] The candidate encoding list is a set formed by cross-combining the original pinyin sequence and the pinyin sequence with heavy accent. It contains the pinyin sequences corresponding to all possible pronunciation variations, ensuring coverage of mixed scenarios where some syllables have heavy accent and some syllables have standard pronunciation.

[0140] For example, the target command word "supermarket" (original pinyin sequence "chaoshi") generates two sets of heavy accent pinyin sequences based on the heavy accent rule base: sequence A (caoshi, rule CH→C, confusion between retroflex and alveolar consonants) and sequence B (chaosi, rule SH→S, confusion between retroflex and alveolar consonants).

[0141] Generate a candidate encoding list according to the cross-combination rules:

[0142] Original pinyin sequence: “chaoshi”;

[0143] Replace only the first character: "caoshi" (sequence A);

[0144] Replace only the second word: "chaosi" (sequence B);

[0145] Replace the two words: "caosi" (a combination of sequence A and sequence B);

[0146] The final candidate list for encoding is {"chaoshi" "caoshi" "chaosi" "caosi"}.

[0147] The candidate list of codes generated by cross-combination comprehensively covers all pronunciation scenarios of "full standard", "partial heavy accent" and "full heavy accent", solving the problem that single benchmark screening is difficult to cover mixed pronunciation variations and ensuring the comprehensiveness of the screening.

[0148] S522: Compare the identified pinyin sequence with the coding candidate list.

[0149] Generate audio F (caosi, full heavy accent) and audio G (chaosi, partial heavy accent), and input them into the Whisper model:

[0150] The recognition result for audio F is "supermarket", which is the pinyin sequence "caosi".

[0151] The audio G was recognized as "supermarket", and the pinyin sequence "chaosi" was also recognized.

[0152] The two sets of recognized pinyin sequences are compared with the candidate encoding list {"chaoshi" "caoshi" "chaosi" "caosi"}:

[0153] The audio file "caosi" is in the list; the comparison passed.

[0154] The audio G's "chaosi" is in the list, and the comparison is successful.

[0155] The comparison process is based directly on a complete match of the pinyin sequence, ensuring the accuracy of the screening; the comprehensiveness of the coded candidate list enables the effective identification of mixed pronunciation variants, avoiding the omission of effective samples due to accented accents in some syllables.

[0156] S523: Retain the accented audio that exactly matches any of the codes in the candidate encoding list.

[0157] Both audio F and audio G are compared, and the two sets of audio are returned to the user independently. The user can choose to keep all or part of the audio according to their needs. In this embodiment, both sets of audio are kept to supplement the mixed pronunciation variant types of the sample set.

[0158] The independent return design gives users flexible choices and avoids redundancy between mixed pronunciation variants and other types of samples; the retained mixed pronunciation samples enrich the scene coverage of the sample set, enabling the model to learn the recognition rules of some syllable accents, and further improve the robustness of the model.

[0159] This embodiment elaborates in detail the core logic of hybrid coding matching: By cross - combining the original pinyin sequence and the heavy - accented pinyin sequence, a coding candidate list covering all pronunciation variants is constructed, and then the heavy - accented audio containing hybrid pronunciation variants is screened out through exact - match comparison. This design targets the common scenario of "partial syllables with heavy accents", solves the problem that a single benchmark screening cannot cover such variants, and ensures the comprehensiveness of the screening results. Among them, the cross - combination design of the coding candidate list comprehensively covers the scenarios of "all - standard", "partial heavy - accents", and "all heavy - accents", avoiding the omission of valid samples; the exact - match comparison ensures the accuracy of the screening, avoiding misjudgment of hybrid pronunciation variants; the independently returned screening results enhance the user's choice right and adapt to the construction requirements of different sample sets; the addition of hybrid pronunciation variant samples enriches the scenario diversity of the sample set, enabling the model to handle more complex heavy - accented pronunciation scenarios and improving the recognition ability.

[0160] In one embodiment, the specific process of the standard pronunciation association matching includes:

[0161] S531: Convert the recognition result into a Chinese character sequence.

[0162] For the target command word "认识" (original pinyin sequence "renshi"), generate heavy - accented audio H (lenshi, rule R→L), audio I (rensi, rule SH→S), and input the two groups of audio into the funasr model:

[0163] Recognition result of audio H: "认识", converted into the Chinese character sequence "认识";

[0164] Recognition result of audio I: "认识", converted into the Chinese character sequence "认识".

[0165] Utilize the language model fault - tolerance mechanism of the speech recognition model to recognize heavy - accented audio with similar pronunciations as the Chinese character sequence of the target command word, providing a unified Chinese character benchmark for subsequent comparison and ensuring the pertinence of the screening.

[0166] S532: If the Chinese character sequence is exactly the same as the target command word, then compare the recognition pinyin sequence corresponding to the recognition result with the original pinyin sequence.

[0167] Compare the Chinese character sequence "认识" of audio H and I with the target command word "认识", and the two are exactly the same, entering the next - step screening;

[0168] If the recognition result of a certain heavy - accented audio is "认知" (Chinese character sequence "认知"), which is not the same as the target command word "认识", then directly exclude this audio.

[0169] Chinese character sequence comparison can quickly exclude audio whose recognition results are irrelevant to the target command word, reduce the complexity of subsequent pinyin comparison, and improve the screening efficiency; at the same time, using the fault tolerance mechanism of the language model to ensure that common heavy-accent audio is not mis-excluded.

[0170] S533: Retain heavy-accent audio where the initials and finals are exactly the same except for the tones.

[0171] Compare the recognized pinyin sequences of audio H and I that have passed Chinese character sequence comparison with the original pinyin sequence "renshi":

[0172] The recognized pinyin sequence of audio H is "lenshi": for "len" and "ren", the initials are different (confusing l / r), and the finals are the same; for "shi" and "shi", both the initials and finals are the same; except for the tones, the core features of the initials and finals can correspond, meeting the comparison requirements;

[0173] The recognized pinyin sequence of audio I is "rensi": for "ren" and "ren", both the initials and finals are the same; for "si" and "shi", the initials are different (confusing sh / s), and the finals are the same; except for the tones, the core features of the initials and finals can correspond, meeting the comparison requirements.

[0174] Retain audio H and audio I.

[0175] Pinyin sequence comparison focuses on the core features of the initials and finals, excludes tone interference, and accurately matches common heavy-accent types; the comparison standard is consistent with the definition of heavy accents, ensuring that the screening results meet the requirements of sample set construction.

[0176] Furthermore, iterative updates of the heavy-accent rule library will also be carried out.

[0177] During the screening process, it is found that the recognition result of audio J (rinshi) is the Chinese character sequence "认识", and the recognized pinyin sequence "rinshi" meets the requirements when compared with the original pinyin sequence "renshi", but the "r→rin" replacement rule corresponding to "rinshi" is not preset in the heavy-accent rule library;

[0178] Supplement the pronunciation replacement rule of "r→rin" to the heavy-accent rule library to achieve iterative updates of the rule library, and subsequent corresponding heavy-accent pinyin sequences can be automatically generated for similar command words.

[0179] Feeding back the rule library iteration through the screening process enables the rule library to continuously cover new heavy-accent types, improves the scalability and adaptability of the solution, and solves the problem of incomplete coverage that may exist in the initial rule library.

[0180] This embodiment details the complete process of standard pronunciation association matching: utilizing the language model fault-tolerance mechanism of the speech recognition model, it first filters audio related to the target command word through Chinese character sequence comparison, then locks audio that matches the characteristics of heavy accents through Pinyin sequence comparison, and finally achieves continuous optimization through iterative updates of the rule base. This process fully utilizes the capabilities of existing speech recognition models, achieving accurate screening of common heavy accents without additional modifications, while also possessing self-optimization capabilities. Specifically, Chinese character sequence comparison quickly eliminates irrelevant audio, improving screening efficiency; Pinyin sequence comparison focuses on the core features of heavy accents, ensuring screening accuracy; the iterative update mechanism of the rule base enables the solution to continuously cover new heavy accent types, improving scalability and adaptability; the screening rules are specifically designed for common heavy accent scenarios, efficiently screening high-frequency heavy accent samples, providing core data support for model training, and improving the model's recognition accuracy for common heavy accents.

[0181] In one embodiment, the sample set used to train the speech recognition model includes:

[0182] S61: Obtain an existing sample set containing standard Mandarin pronunciation audio of the target command word.

[0183] The existing standard sample set refers to a dataset containing audio recordings of the standard Mandarin pronunciation of the target command word. It serves as the basic training data for speech recognition models, with each audio recording corresponding to a standard pinyin sequence without tone marks.

[0184] For example, if the target command is "play", obtain the existing standard sample set S, which contains 3 standard audio tracks:

[0185] B1: Pronounced "bofang" (standard Mandarin), corresponding to the pinyin sequence "bofang";

[0186] B2: Pronounced "bofang" (standard Mandarin, male voice), corresponding to the pinyin sequence "bofang";

[0187] B3: Pronounced "bofang" (standard Mandarin, female voice), corresponding to the pinyin sequence "bofang".

[0188] The model is expanded based on the existing standard sample set, eliminating the need to build a sample set from scratch, thus reducing the cost and time required for sample set construction. The standard sample set provides the model with basic recognition capabilities, complementing the heavy accent samples.

[0189] S62: The selected audio samples with heavy accents that meet the requirements are merged with the existing sample set to form an enhanced training sample set, which is used for training the speech recognition model.

[0190] For example, in the above embodiments, three audio clips with heavy accents that meet the requirements are generated and filtered out:

[0191] A1: Pronunciation "bofang" (confusion between retroflex and alveolar consonants, "fang" → "fang" without change, "bo" → "bo" without change, actually "pofang", initial consonant p / b confused), corresponding to the pinyin sequence "pofang", tag "p / b indistinguishable";

[0192] A2: Pronunciation "bofang" (confusion between front and back nasal sounds, "fang" → "fan"), corresponding to the pinyin sequence "bofan", tag "ang / an indistinguishable";

[0193] A3: Pronunciation "bofang" (mixed heavy accent, "bo" → "po", "fang" → "fan"), corresponding to the pinyin sequence "pofan", tag "p / b+ang / an indistinguishable".

[0194] Labeling enriches the information of heavy accent audio, making it easier for the model to learn the features of different types of heavy accents; the selected heavy accent audio ensures the quality of the samples and avoids invalid data from affecting the model training effect.

[0195] The three accented audio samples A1, A2, and A3 are merged with the existing standard sample set B to form an enhanced training sample set T (containing B1, B2, B3, A1, A2, and A3).

[0196] Audio A3 (mixed heavy accent) with large pronunciation deviation was selected from the enhanced sample set T as a difficult sample and separately assigned to the validation set V to evaluate the model's ability to recognize complex heavy accents.

[0197] The remaining audio (B1, B2, B3, A1, A2) is used as the training set for daily training of the model.

[0198] The sample set achieves comprehensive coverage of "standard pronunciation + accented pronunciation", solving the problem of insufficient diversity in existing sample sets; the separate division of the validation set enables the linkage between model training and evaluation, and the validation results can guide the direction of model optimization and improve the robustness of the model.

[0199] This embodiment details the construction process of the enhanced training sample set: based on the existing standard sample set, selected and labeled heavy accent audio is added, and a comprehensive sample set is formed by merging them, then divided into training and validation sets. This process makes full use of existing resources, ensuring the continuity and integrity of the sample set; labeling and validation set division improve the practicality of the sample set, providing accurate support for model training and evaluation. Specifically, expanding the existing standard sample set reduces the cost and cycle of sample set construction; the addition of heavy accent audio enriches the diversity of the sample set, enabling the model to learn different types of heavy accent features and improve recognition ability; labeling facilitates targeted model training, improving training efficiency; the division of the training and validation sets enables the linkage between model training and evaluation, guiding the direction of model optimization and further improving the robustness and universality of the model; the enhanced sample set can be directly used for training existing speech recognition models without additional model modifications, significantly improving the practicality and operability of the solution.

[0200] Reference Figure 2 This invention also provides a sample set generation apparatus for training a speech recognition model, used to execute the sample set generation method for training a speech recognition model in any of the above embodiments, including:

[0201] Decoding unit 10 is used to decode the target command word into a toneless original pinyin sequence;

[0202] The conversion unit 20 is used to encode and convert the original pinyin sequence based on the pronunciation replacement rules in the preset heavy accent rule library to generate at least one set of heavy accent pinyin sequences; wherein, the heavy accent refers to non-standard pronunciation with changes in the initial consonant or final vowel, and the heavy accent rule library is constructed based on common non-standard pronunciation rules;

[0203] The generation unit 30 is used to input the heavy accent pinyin sequence into the text into the speech audio generation tool to generate the corresponding heavy accent audio;

[0204] The recognition unit 40 is used to input the accented audio into a preset speech recognition model, obtain the recognition result corresponding to the accented audio, and convert the recognition result into a corresponding recognition pinyin sequence;

[0205] The filtering unit 50 is used to filter the accented audio based on preset filtering rules. The filtering rules are based on the pinyin sequence to construct a matching benchmark. The matching benchmark is generated based on the original pinyin sequence and / or the accented pinyin sequence and compared with the recognized pinyin sequence.

[0206] The combination unit 60 is used to collect filtered audio samples with heavy accents that meet the requirements, forming a sample set for training the speech recognition model.

[0207] Reference Figure 3 The present invention also provides a computer device, the internal structure of which can be as follows: Figure 3 As shown. The computer device includes a processor, memory, network interface, and database connected via a system bus. The processor is designed to provide computing and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores operating devices, computer programs, and a database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The database stores sample sets, etc. The network interface is used to communicate with external terminals via a network connection. Furthermore, the computer device may also include input devices and a display screen, etc. When the computer program is executed by the processor, it implements the sample set generation method for training the speech recognition model in any of the above embodiments. Those skilled in the art will understand that... Figure 3 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer equipment on which the present application is applied.

[0208] One embodiment of this application also provides a computer-readable storage medium storing a computer program thereon. When the computer program is executed by a processor, it implements the sample set generation method for training a speech recognition model in any of the above embodiments. It is understood that the computer-readable storage medium in this embodiment can be a volatile readable storage medium or a non-volatile readable storage medium.

[0209] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in this application and in the embodiments can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual-speed SDRAM (SSRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).

[0210] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, apparatus, article, or method that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, apparatus, article, or method. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, apparatus, article, or method that includes that element.

[0211] The above description is merely a preferred embodiment of the present invention and does not limit the patent scope of the present invention. Any equivalent structural or procedural transformations made based on the content of the present invention's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of the present invention.

Claims

1. A method for generating a sample set for training a speech recognition model, characterized in that, The method includes: Decode the target command word into a toneless raw phonetic sequence; Based on the pronunciation replacement rules in the preset heavy accent rule library, the original pinyin sequence is encoded and converted to generate at least one set of heavy accent pinyin sequence; wherein, the heavy accent refers to non-standard pronunciation with changes in initials or finals, and the heavy accent rule library is constructed based on common non-standard pronunciation rules; The heavy accent pinyin sequence is input into the speech audio generation tool to generate the corresponding heavy accent audio; The audio with the heavy accent is input into a preset speech recognition model to obtain the recognition result corresponding to the audio with the heavy accent, and the recognition result is converted into the corresponding recognition pinyin sequence; The audio with a heavy accent is filtered based on preset filtering rules. The filtering rules use the pinyin sequence as the matching basis to construct a matching benchmark. The matching benchmark is generated based on the original pinyin sequence and / or the pinyin sequence with a heavy accent, and is compared with the recognized pinyin sequence. Collect audio samples with strong accents that meet the screening requirements to form a sample set for training the speech recognition model.

2. The method for generating a sample set for training a speech recognition model according to claim 1, characterized in that, In the heavy accent rule base, each set of heavy accent pinyin sequence corresponds to a unique heavy accent pronunciation rule; and a target command word can generate zero or one or more heavy accent pinyin sequences.

3. The method for generating a sample set for training a speech recognition model according to claim 1, characterized in that, The filtering rules include at least one of the following: full matching of accented pronunciations, matching of standard pronunciations, matching of mixed encodings, and matching of initial and final similarities.

4. The method for generating a sample set for training a speech recognition model according to claim 3, characterized in that, The specific process of matching the similarity between initials and finals includes: Audio samples with heavy accents whose pinyin sequences match the number of characters in the target command word were selected as candidate audio samples. For each character in the identified pinyin sequence of the candidate audio, calculate the similarity score between its initial consonant and final vowel and the initial consonant and final vowel of the character corresponding to the target command word. Audio with heavy accents that have a total similarity score higher than a preset threshold are retained.

5. The method for generating a sample set for training a speech recognition model according to claim 3, characterized in that, The specific process of the hybrid encoding matching includes: Based on the original pinyin sequence and at least one set of accented pinyin sequences, a candidate list of codes containing multiple pronunciation variations is generated; The identified pinyin sequence is compared with the coding candidate list; Retain accented audio that exactly matches any of the encoding candidate lists.

6. The method for generating a sample set for training a speech recognition model according to claim 3, characterized in that, The specific process of standard pronunciation association matching includes: The recognition results are converted into Chinese character sequences; If the Chinese character sequence is completely consistent with the target command word, then the recognized pinyin sequence corresponding to the recognition result is compared with the original pinyin sequence; Preserve the accented audio where the initial consonant and final vowel are completely identical except for the tone.

7. The method for generating a sample set for training a speech recognition model according to any one of claims 1 to 6, characterized in that, The sample set used to train the speech recognition model includes: Obtain an existing sample set of standard Mandarin pronunciation audio containing the target command word; The selected audio samples with strong accents that meet the requirements are combined with the existing sample set to form an enhanced training sample set, which is used to train the speech recognition model.

8. A sample set generation device for training a speech recognition model, characterized in that, include: The decoding unit is used to decode the target command word into a toneless original pinyin sequence; The conversion unit is used to encode and convert the original pinyin sequence based on the pronunciation replacement rules in the preset heavy accent rule library to generate at least one set of heavy accent pinyin sequences; wherein, the heavy accent refers to non-standard pronunciation with changes in the initial consonant or final vowel, and the heavy accent rule library is constructed based on common non-standard pronunciation rules; The generation unit is used to input the heavy accent pinyin sequence into the speech audio generation tool to generate the corresponding heavy accent audio. The recognition unit is used to input the accented audio into a preset speech recognition model, obtain the recognition result corresponding to the accented audio, and convert the recognition result into a corresponding recognition pinyin sequence; A filtering unit is used to filter the accented audio based on preset filtering rules. The filtering rules use the pinyin sequence as the matching basis to construct a matching benchmark. The matching benchmark is generated based on the original pinyin sequence and / or the accented pinyin sequence and compared with the recognized pinyin sequence. The combination unit is used to collect filtered audio samples with strong accents to form a sample set for training the speech recognition model.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method for generating a sample set for training a speech recognition model as described in any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the sample set generation method for training a speech recognition model as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Speech recognition model training method, speech recognition method and device

    CN119832898A

  • Clustering and mining accented speech for inclusive and fair speech recognition

    US20240290322A1