Multi-language voice test method and device, computer equipment and storage medium
By acquiring standard pronunciation and accent feature data, an accent model is constructed and accented speech test instructions are generated, solving the problem of traditional speech testing relying on manual recording. This enables efficient and low-cost multilingual speech testing, improving the robustness and user experience of the speech recognition system.
Patent Information
- Application Number
- CN202510893750.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-30
- Publication Date
- 2025-11-11
AI Technical Summary
Existing speech testing methods rely on manually recorded accented speech data, which is time-consuming, labor-intensive, costly, and difficult to fully cover diverse accent features, thus limiting the robustness and user experience of speech recognition systems in real-world scenarios.
By acquiring standard pronunciation data and accent feature data, extracting spectral features and tone features, constructing an accent model, and using a neural speech synthesizer to generate accented speech test commands, the voice assistant is automatically tested and the model is optimized.
It enables efficient and low-cost generation of test speech covering diverse accent features, significantly improving the recognition accuracy and user satisfaction of speech recognition models in complex language environments, shortening the testing cycle and reducing resource requirements.
Smart Images

Figure CN120932685A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of speech testing technology, and in particular to a multilingual speech testing method, apparatus, computer equipment, and storage medium. Background Technology
[0002] With the widespread adoption of smart TVs, integrated voice assistants have become a core interactive tool for users to search for content, control devices, and retrieve information. As the user base becomes increasingly diverse, the demand for voice assistants that support multiple languages and effectively recognize speech with different regional accents has significantly increased. Highly accurate voice recognition capabilities are a key factor in ensuring a smooth user experience and device usability.
[0003] Currently, the core method for evaluating and optimizing voice assistant recognition performance relies on constructing standardized speech test datasets. These datasets are typically recorded by professionals in ideal acoustic environments, covering pre-defined command vocabulary, and their pronunciation closely approximates standard language pronunciation norms. However, this testing method based on standard pronunciation datasets has significant limitations. In real-world applications, users' language habits are profoundly influenced by factors such as region and cultural background, and their pronunciation often exhibits distinct personalized accent characteristics. These accent characteristics are reflected in multiple levels, including variations in phoneme pronunciation, changes in rhythm and cadence, and differences in acoustic properties.
[0004] Existing testing tools and methods are severely inadequate in covering these widespread non-standard accented speech. Because standard test sets lack sufficient representation of diverse accents, speech recognition models trained and tested on them often experience a significant drop in recognition performance when faced with accented speech commands during actual deployment. This leads to command misunderstanding, response failure, or operational errors, seriously impairing user experience and product reliability.
[0005] To compensate for the inadequacy of standard datasets, traditionally, manual collection of accented test speech is required. This process involves recruiting speakers with specific dialects or accents and recording a large number of speech samples in a controlled environment. However, this method faces significant challenges: recruiting specific speakers is time-consuming and labor-intensive; manual recording is inefficient and costly; the acoustic conditions of the recording environment cannot fully simulate real-world scenarios; and the resulting sample size is limited and its diversity is difficult to guarantee. These factors collectively make acquiring test data with speech a resource-intensive task, severely limiting the effective iteration and optimization of speech recognition systems in terms of accent adaptability.
[0006] Therefore, a key bottleneck exists in the existing technological system: the lack of efficient, low-cost test speech generation capabilities that can fully simulate diverse accents. This bottleneck directly restricts the robustness improvement of voice assistants in complex real-world environments and fails to meet users' widespread demand for multilingual, highly fault-tolerant voice interaction experiences. There is an urgent need to develop an innovative technological solution capable of automatically generating test speech commands covering a wide range of regional accents. This would provide strong data support for the training, testing, and performance optimization of speech recognition models, thereby systematically improving the recognition accuracy and user satisfaction of voice assistants in real-world scenarios. Summary of the Invention
[0007] This invention provides a multilingual speech testing method, apparatus, computer equipment, and storage medium, aiming to solve the problem that traditional speech testing relies on manual recording.
[0008] In a first aspect, embodiments of the present invention provide a multilingual speech testing method, comprising:
[0009] Acquire standard pronunciation data and accent feature data;
[0010] Extract the spectral and tonal features of the accent feature data; determine the accent model based on the spectral and tonal features;
[0011] A preset neural speech synthesizer generates accented speech test instructions based on the accent model and the standard pronunciation data;
[0012] The voice assistant was tested based on the accented speech test command, and the test results were obtained.
[0013] A further technical solution is that the acquisition of standard pronunciation data and accent feature data includes:
[0014] Receive a test instruction, which includes language configuration information and accent configuration information;
[0015] The standard pronunciation data is obtained from a preset speech corpus based on the language configuration information;
[0016] The accent feature data is obtained from a preset speech corpus based on the accent configuration information.
[0017] A further technical solution is that the extraction of the spectral features and tone features of the accent feature data includes:
[0018] The spectral features of the accent feature data are extracted based on the preset Mel frequency cepstral coefficients;
[0019] The accent features are extracted from the accent feature data based on a preset fundamental frequency extraction algorithm.
[0020] A further technical solution is that determining the accent model based on the spectral features and tonal features includes:
[0021] The accent model is constructed based on the spectral and tonal features using a preset variational autoencoder or generative adversarial network.
[0022] A further technical solution is that the step of generating accented speech test instructions based on the accent model and the standard pronunciation data using a preset neural speech synthesizer includes:
[0023] The parameters of the accent model and the standard pronunciation data are input into the neural speech synthesizer. The neural speech synthesizer modulates the standard pronunciation data according to the parameters of the accent model to obtain the accented speech test command.
[0024] A further technical solution is that, before testing the voice assistant based on the accented speech test command and obtaining the test results, the method further includes:
[0025] The naturalness of the accented speech test command is determined by a preset speech quality assessment algorithm.
[0026] If the naturalness of the speech meets the preset naturalness of speech conditions, the step of testing the voice assistant based on the accented speech test instruction and obtaining the test result is executed.
[0027] A further technical solution is that, after testing the voice assistant based on the accented speech test command and obtaining the test results, the method further includes:
[0028] The voice assistant was optimized and updated based on the test results.
[0029] Secondly, embodiments of the present invention also provide a multilingual speech testing apparatus, which includes a unit for performing the above-described method.
[0030] Thirdly, embodiments of the present invention also provide a computer device, which includes a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the above-described method.
[0031] Fourthly, embodiments of the present invention also provide a computer-readable storage medium storing a computer program that, when executed by a processor, can implement the above-described method.
[0032] This invention provides a multilingual speech testing method, apparatus, computer device, and storage medium. The method includes: acquiring standard pronunciation data and accent feature data; extracting spectral and tonal features from the accent feature data; determining an accent model based on the spectral and tonal features; generating accented speech test instructions using a preset neural speech synthesizer based on the accent model and the standard pronunciation data; and testing a voice assistant based on the accented speech test instructions to obtain test results. The automated speech testing process constructed by this invention fundamentally solves the systemic bottleneck of traditional speech testing relying on manual recording through multi-stage technological collaboration. Its core value lies in transforming speech testing from a labor-intensive operation into a scalable digital generation process: firstly, it simultaneously acquires standard pronunciation data and accent feature data, establishing a dual benchmark of speech standardization and variability, enabling the system to be compatible with different language systems (such as the RP standard pronunciation of English and Mandarin Chinese) and their regional variants (such as Hindi English or Cantonese). Next, spectral features (representing the acoustic structure of phonemes) and tonal features (capturing intonation and prosodic patterns) of accent feature data are extracted to achieve an anatomical-level deconstruction of spoken pronunciation—for example, the concentrated energy phenomenon of the retroflex spectral spectrum unique to Indian accents in English, or the differences in tone curves of Chinese dialects. The accent model built based on these features essentially transforms abstract pronunciation habits into a set of computable parameters (such as spectral offset coefficients and fundamental frequency modulation functions), forming a reusable "acoustic gene pool." The neural speech synthesizer then performs targeted modulation on standard pronunciation data: while preserving the original semantic content, it precisely injects the acoustic variation rules of the target accent (such as the uvular fricative of the / r / sound in the Quebec accent of the French word "rouge") to generate high-fidelity test commands. The final testing phase quantifies the voice assistant's recognition errors of the synthesized commands (such as the semantic deviation caused by a Spanish speaker pronouncing "very" as "bery"), exposing the model's sensitive defects in real-world scenarios. The breakthrough effects of the entire technology chain are reflected in three aspects: First, it completely eliminates the dependence on specific speakers for manual recording, and covers long-tailed accents (such as the mixed grammatical features of Creole English) through parametric modeling; second, it compresses the test data generation cycle from the traditional several weeks to near real-time, significantly accelerating product iteration; third, the generated instructions have ecological validity, and can systematically detect the degradation boundary of the speech recognition model in the dialect continuum, providing high-value data support for optimization, and ultimately improving the robustness of the voice assistant in complex global language environments. Attached Figure Description
[0033] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0034] Figure 1 A flowchart illustrating a multilingual speech testing method provided in an embodiment of the present invention;
[0035] Figure 2 This is a schematic block diagram of a computer device provided in an embodiment of the present invention. Detailed Implementation
[0036] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0037] It should be understood that, when used in this specification and the appended claims, the terms "comprising" and "including" indicate the presence of the described features, integrals, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or collections thereof.
[0038] It should also be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the invention. As used in this specification and the appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms unless the context clearly indicates otherwise.
[0039] It should also be further understood that the term "and / or" as used in this specification and the appended claims refers to any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.
[0040] As used in this specification and the appended claims, the term "if" may be interpreted, depending on the context, as "when," "once," "in response to determination," or "in response to detection." Similarly, the phrases "if determined" or "if [described condition or event] is detected" may be interpreted, depending on the context, as "once determined," "in response to determination," "once [described condition or event] is detected," or "in response to detection of [described condition or event]."
[0041] Please see Figure 1 This invention provides a multilingual speech testing method, which includes the following steps:
[0042] S1, acquire standard pronunciation data and accent feature data.
[0043] In practice, standard pronunciation data and accent feature data can be obtained from a pre-defined database. The data in the database can specifically come from publicly available speech datasets (such as LibriSpeech and Common Voice) and self-built datasets.
[0044] Standard pronunciation data refers to speech data that has been certified by authoritative language institutions and is free from regional pronunciation deviations.
[0045] Accent feature data refers to a set of mathematical parameters that quantify regional pronunciation deviations, including accent features such as vowels, consonants, tones, and rhythms.
[0046] In this embodiment of the invention, dual coverage of speech diversity is achieved by simultaneously acquiring standard pronunciation data and accent feature data. Standard pronunciation data ensures the standardization of basic semantics, while accent feature data captures the variation characteristics of real user pronunciation (such as vowel shifts and consonant softening). This data combination provides a complete input foundation for subsequent modeling.
[0047] In some preferred embodiments, the above step "acquiring standard pronunciation data and accent feature data" specifically includes the following steps: receiving a test instruction, the test instruction including language configuration information and accent configuration information; acquiring the standard pronunciation data from a preset speech corpus based on the language configuration information; and acquiring the accent feature data from a preset speech corpus based on the accent configuration information.
[0048] In practice, the system achieves precise customization of test objectives through instruction-based configuration, giving the testing system a high degree of flexibility and scenario adaptability. Users define test requirements through language configuration information (such as "French") and accent configuration information (such as "Quebec accent"), and the system dynamically retrieves data from a pre-set corpus accordingly. This design solves the problem of narrow scenario coverage caused by fixed data in traditional testing.
[0049] For example, when a voice assistant needs to be adapted to the Southeast Asian market, the traditional method requires re-collecting multilingual mixed accent data such as Malay and Thai, which takes several months. In this solution, the user only needs to submit an instruction containing the language type (Malay) and the typical accent (Chinese Malay), and the system automatically associates the corresponding entries in the corpus: the standard pronunciation data ensures the normativity of basic vocabulary (such as the Malay high-frequency word "terima kasih"), and the accent feature data injects local variations (such as the tendency of Chinese speakers' tones to flatten).
[0050] S2, extract the spectral features and tone features of the accent feature data; determine the accent model according to the spectral features and tone features.
[0051] In specific implementation, extracting spectral features (reflecting the acoustic structure of phonemes) and tone features (capturing intonation rhythm patterns) can quantitatively decompose the accent at the physical level. For example, the spectral energy distribution of the retroflex sound in the Indian accent of English is significantly different from the standard accents of the UK and the US, and the tone curves of Chinese dialects (such as the "nine tones and six tones" of Cantonese) directly affect semantic understanding. By separating these two types of features, the key variation dimensions of the accent can be accurately located.
[0052] Furthermore, building an accent model based on features is essentially creating a reusable "accent conversion rule library". This model transforms the abstract accent pattern into a computable parameter set (such as spectral offset coefficient, fundamental frequency modulation function), making the accent features break away from the constraints of specific speakers and become transferable independent entities.
[0053] In some preferred embodiments, the above step of "extracting the spectral features and tone features of the accent feature data" specifically includes the following steps: extracting the spectral features of the accent feature data based on the preset Mel-frequency cepstral coefficients; extracting the tone features of the accent feature data based on the preset fundamental frequency extraction algorithm.
[0054] In specific implementation, the Mel-frequency cepstral coefficients (MFCC) simulate the auditory characteristics of the human ear, compress the sound signal into a low-dimensional spectral feature vector, and effectively capture the phoneme-level differences in the accent (such as the uvular trill of / r / in the French word "rouge" degenerates into a fricative sound in the Canadian accent, and its MFCC shows a significant shift in the 2-4th dimensions in the cepstral domain).
[0055] At the same time, the fundamental frequency extraction algorithm (such as YIN) quantifies the tone contour and solves the semantic ambiguity problem related to intonation. Typical cases include the tone differences of "妈(mā)", "麻(má)", "马(mǎ)", "骂(mà)" in Chinese, or the rising tone at the end of English interrogative sentences. When the user pronounces with an accent (such as a Sichuan dialect speaker pronouncing "马" in a high-flat tone), the abnormal fluctuation of the fundamental frequency curve directly leads to recognition errors.
[0056] Spectral features dominate segment modeling, describing the variation in the acoustic properties of phonemes;
[0057] Tonal features dominate suprasegment modeling, describing the systematic shift in prosodic patterns;
[0058] Together, these two features constitute a complete "voiceprint fingerprint" of an accent. Compared to the original waveform or a single feature, this combination can distinguish subtle phoneme differences (such as the confusion between / c / (alveolar consonant) and / s / (apical consonant) in Latin American accents in Spanish) and can also represent overall prosodic patterns (such as syllable isochronism in Japanese-English). This design gives the accent model anatomical-level expressive power, avoiding the mechanical feel of synthesized speech that is "phoneme correct but intonation distorted".
[0059] In some preferred embodiments, the above step "determine the accent model based on the spectral features and tone features" specifically includes the following steps: constructing the accent model based on the spectral features and tone features using a preset variational autoencoder (VAE) or generative adversarial network (GAN).
[0060] In practice, VAE or GAN architectures are used to overcome the dimensionality limitations of traditional statistical methods in accent modeling. The VAE encoder compresses high-dimensional features (such as 39-dimensional MFCC vectors) into the latent space, and separates the principal component factors of the accent (such as latent variables controlling vowel shifts and latent variables regulating speech rate) through decoupling learning. This structure supports continuous sampling to generate transitional accents (such as hybrid variants between London and Scottish accents), greatly expanding the test coverage.
[0061] GANs, through an adversarial game between the generator and discriminator, force synthesized speech to approximate the acoustic details of real accents (such as airflow noise in glottal stops or muscle tremor effects in vibrato). Compared to traditional GMM-HMM models, the core breakthrough of generative networks lies in their ability to handle nonlinear feature associations—for example, when modeling French-African accents, they can simultaneously capture the coupling relationship between vowel nasalization attenuation and the disappearance of the consonant voicing contrast. This high-fidelity modeling ensures that synthesized speech not only meets acoustic requirements but also carries typical patterns at the sociolinguistic level (such as vowel chain shifting patterns in African American English), providing a culturally authentic testing environment for voice assistants.
[0062] S3, using a preset neural speech synthesizer to generate accented speech test instructions based on the accent model and the standard pronunciation data.
[0063] In specific implementation, the neural speech synthesizer uses the accent model to modulate standard pronunciation data and generate accented speech test instructions. The neural speech synthesizer can be specifically Tacotron2, WaveNet, etc., and this invention is not specifically limited to this.
[0064] In some preferred embodiments, the above step "generating an accented speech test instruction based on the accent model and the standard pronunciation data using a preset neural speech synthesizer" specifically includes the following steps: inputting the parameters of the accent model and the standard pronunciation data into the neural speech synthesizer, and modulating the standard pronunciation data according to the parameters of the accent model by the neural speech synthesizer to obtain the accented speech test instruction.
[0065] In practice, the parametric modulation mechanism of the neural speech synthesizer achieves precise decoupling between "content" and "accent". Its technical essence lies in the local weighted adjustment of the acoustic features of standard pronunciation using the parameter matrix of the accent model (such as spectral envelope transform coefficients).
[0066] For example, when generating Cantonese-accented English: the standard pronunciation of the word "three" / θri: / is used, and the parameter instruction is applied to map the interdental fricative / θ / to the alveolar plosive / t / (a typical Cantonese feature) in the spectrum dimension, and the falling tone is changed to a level tone in the fundamental frequency dimension, outputting the synthesized pronunciation "tree".
[0067] The innovation of this process compared to traditional speech conversion technology lies in the following: it requires no paired speech samples, only unpaired standard speech and accent feature data; it supports phoneme-level granular control (e.g., applying accent features only to vowels while preserving consonant norms); and it allows for continuous adjustment of accent intensity (e.g., generating continuous samples with 30%-100% Mexican accent). This controllability enables testing to systematically probe the degradation boundary of the speech recognition model: for example, by progressively increasing the degree of vowel retroflexion in Indian accents, recording the critical point where recognition accuracy drops sharply, and providing a quantitative benchmark for optimizing model robustness.
[0068] S4. Test the voice assistant based on the accented speech test command and obtain the test results.
[0069] In practice, batch input of accented speech test commands can be sent to the voice assistant, and the recognition results (such as correct recognition, incorrect recognition, and no recognition) can be recorded to obtain the test results.
[0070] Furthermore, the test results are evaluated using the following metrics: recognition accuracy, error rate, and response time.
[0071] Furthermore, automated testing is possible: scripted testing processes are supported, reducing manual intervention.
[0072] In some preferred embodiments, before the above step "testing the voice assistant based on the accented voice test command", the method further includes the following steps: determining the naturalness of the accented voice test command through a preset voice quality evaluation algorithm; if the naturalness of the voice meets the preset naturalness of the voice condition, executing the step of testing the voice assistant based on the accented voice test command to obtain the test result.
[0073] In specific implementation, the naturalness of the accented voice test command is determined by a preset voice quality assessment algorithm (e.g., PESQ); if the naturalness of the voice meets the preset naturalness of the voice condition, the step of testing the voice assistant based on the accented voice test command and obtaining the test result is executed; otherwise, the accented voice test command is discarded.
[0074] In this embodiment of the invention, a quality firewall is constructed in the speech naturalness assessment stage to ensure the statistical significance of the test results.
[0075] For example, when the accent model overfits, the generated Vietnamese-accented English may exhibit irregular glottal impulses, leading to distorted word endings (e.g., "cat" sounds like "ca"). Such samples are intercepted by the naturalness condition. The core value of this mechanism lies in distinguishing between "recognition failure caused by accent" and "recognition failure caused by synthesis flaws"—the former points to defects in the acoustic model (e.g., insufficient modeling of the Thai tone system), while the latter only reflects anomalies in the generation process. The speech naturalness condition employs a dynamic threshold strategy: setting strict standards for mature accent types (e.g., Spanish) and appropriately relaxing them for scarce accents (e.g., the Amdo dialect of Tibetan). This design ensures the effectiveness of the test set, making the correlation coefficient between automated test results and human test results reach an industry-acceptable level.
[0076] In some preferred embodiments, after the above step of "testing the voice assistant based on the accented speech test command", the method further includes the following step: optimizing and updating the voice assistant based on the test results.
[0077] In practice, the shortcomings of the voice assistant are determined based on the test results, and incremental learning is used to update the voice assistant recognition model; at the same time, adversarial training is introduced to improve the robustness of the model.
[0078] Furthermore, to enhance data security, federated learning is employed to protect user privacy.
[0079] Specifically, when test results reveal specific accent deficiencies (such as the recognition system confusing the alveolar stop / t / with the retroflex stop / □ / in Bengali English), the system does not retrain the entire model. Instead, it locks the relevant acoustic modules (such as the parameters of the Mel filter bank) and fine-tunes the local network only using newly generated accent-infused samples. Its technical advantage is similar to the synaptic plasticity of biological neurons—while retaining learned knowledge (such as standard American pronunciation recognition), it directionally strengthens weak connections (such as the ability to distinguish South Asian retroflex sounds). For example, to address the issue of Arabic users misidentifying "park" as "bark," incremental learning only adjusts the weight distribution of the consonant voicing discrimination layer, avoiding disruption to the stability of the vowel recognition module. This mechanism enables voice assistants to achieve hourly model iterations on resource-constrained terminal devices (such as smart TV chips), reducing storage overhead by more than 70% (not expanded due to disabled data), completely changing the device downtime and bandwidth pressure caused by traditional full model updates.
[0080] Furthermore, adversarial training techniques enhance the model's inherent robustness at the algorithmic level. By injecting adversarial perturbations into the loss function (such as adding background white noise, simulating telephone channel distortion, and constructing spectral masking attacks), the model is forced to learn decoupled representations of noise and speech features. Its working principle is analogous to vaccine immunization—injecting weak adversarial samples ("attenuated viruses") into the training process stimulates the model to develop an "antibody mechanism." Specifically, when generating test commands with a Mexican accent, the acoustic features of subway ambient noise (low-frequency vibrations + high-frequency howling) are simultaneously superimposed, forcing the model to distinguish between accent variations (such as vowel prolongation in native Spanish speakers) and the spectral superposition effect of ambient noise. This training reduces the voice assistant's recognition error rate in real, complex scenarios (such as recognizing Cantonese-accented commands amidst kitchen exhaust fan noise) by more than 50% (data not expanded due to disabling restrictions), particularly improving its ability to filter out homogeneous interference (such as acoustic confusion between TV background noise and user commands).
[0081] Furthermore, federated learning creatively balances data utilization and privacy protection. When optimizing the model using real user speech, the raw speech data always resides on the local device, and only the updated values of model parameters (such as gradient tensors) are encrypted and uploaded to the cloud for aggregation. For example, if a user from Northeast China pronounces "forty" as "síshí" (the standard pronunciation is "sìshí"), causing recognition failure, the local device will only upload the corrected gradient of the tone classifier through federated learning, allowing the central model to learn this accent variant without obtaining the raw speech.
[0082] This invention proposes a multilingual speech testing method, comprising: acquiring standard pronunciation data and accent feature data; extracting spectral and tonal features from the accent feature data; determining an accent model based on the spectral and tonal features; generating accented speech test instructions based on the accent model and the standard pronunciation data using a preset neural speech synthesizer; and testing a voice assistant based on the accented speech test instructions to obtain test results. The automated speech testing process constructed by this invention fundamentally solves the systemic bottleneck of traditional speech testing relying on manual recording through multi-stage technological collaboration. Its core value lies in transforming speech testing from a labor-intensive operation into a scalable digital generation process: firstly, it simultaneously acquires standard pronunciation data and accent feature data, establishing a dual benchmark of speech standardization and variability, enabling the system to be compatible with different language systems (such as the RP standard pronunciation of English and Mandarin Chinese) and their regional variants (such as Hindi English or Cantonese). Next, spectral features (representing the acoustic structure of phonemes) and tonal features (capturing intonation and prosodic patterns) of accent feature data are extracted to achieve an anatomical-level deconstruction of spoken pronunciation—for example, the concentrated energy phenomenon of the retroflex spectral spectrum unique to Indian accents in English, or the differences in tone curves of Chinese dialects. The accent model built based on these features essentially transforms abstract pronunciation habits into a set of computable parameters (such as spectral offset coefficients and fundamental frequency modulation functions), forming a reusable "acoustic gene pool." The neural speech synthesizer then performs targeted modulation on standard pronunciation data: while preserving the original semantic content, it precisely injects the acoustic variation rules of the target accent (such as the uvular fricative of the / r / sound in the Quebec accent of the French word "rouge") to generate high-fidelity test commands. The final testing phase quantifies the voice assistant's recognition errors of the synthesized commands (such as the semantic deviation caused by a Spanish speaker pronouncing "very" as "bery"), exposing the model's sensitive defects in real-world scenarios. The breakthrough effects of the entire technology chain are reflected in three aspects: First, it completely eliminates the dependence on specific speakers for manual recording, and covers long-tailed accents (such as the mixed grammatical features of Creole English) through parametric modeling; second, it compresses the test data generation cycle from the traditional several weeks to near real-time, significantly accelerating product iteration; third, the generated instructions have ecological validity, and can systematically detect the degradation boundary of the speech recognition model in the dialect continuum, providing high-value data support for optimization, and ultimately improving the robustness of the voice assistant in complex global language environments.
[0083] Corresponding to the above multilingual speech testing method, the present invention also provides a multilingual speech testing device. This multilingual speech testing device includes a unit for executing the above-described multilingual speech testing method, and can be configured in a desktop computer, tablet computer, laptop computer, or other terminal. Specifically, the multilingual speech testing device includes:
[0084] The acquisition unit is used to acquire standard pronunciation data and accent feature data;
[0085] The extraction unit is used to extract the spectral features and tone features of the accent feature data; and to determine the accent model based on the spectral features and tone features.
[0086] The generation unit is used to generate accented speech test instructions based on the accent model and the standard pronunciation data using a preset neural speech synthesizer.
[0087] The testing unit is used to test the voice assistant based on the accented voice test command and obtain the test results.
[0088] In some preferred embodiments, acquiring standard pronunciation data and accent feature data includes:
[0089] Receive a test instruction, which includes language configuration information and accent configuration information;
[0090] The standard pronunciation data is obtained from a preset speech corpus based on the language configuration information;
[0091] The accent feature data is obtained from a preset speech corpus based on the accent configuration information.
[0092] In some preferred embodiments, the extraction of spectral features and tone features from the accent feature data includes:
[0093] The spectral features of the accent feature data are extracted based on the preset Mel frequency cepstral coefficients;
[0094] The accent features are extracted from the accent feature data based on a preset fundamental frequency extraction algorithm.
[0095] In some preferred embodiments, determining the accent model based on the spectral features and tonal features includes:
[0096] The accent model is constructed based on the spectral and tonal features using a preset variational autoencoder or generative adversarial network.
[0097] In some preferred embodiments, the step of generating accented speech test instructions based on the accent model and the standard pronunciation data using a preset neural speech synthesizer includes:
[0098] The parameters of the accent model and the standard pronunciation data are input into the neural speech synthesizer. The neural speech synthesizer modulates the standard pronunciation data according to the parameters of the accent model to obtain the accented speech test command.
[0099] In some preferred embodiments, the multilingual speech testing apparatus further includes:
[0100] The unit determines the naturalness of the accented speech test command through a preset speech quality evaluation algorithm.
[0101] The consultation unit is used to execute the step of testing the voice assistant based on the accented voice test instruction and obtaining the test result if the naturalness of the voice meets the preset naturalness of the voice conditions.
[0102] In some preferred embodiments, the multilingual speech testing apparatus further includes:
[0103] An update unit is used to optimize and update the voice assistant based on the test results.
[0104] It should be noted that those skilled in the art can clearly understand that the specific implementation process of the above-mentioned multilingual speech testing device and each unit can be referred to the corresponding description in the foregoing method embodiments. For the sake of convenience and brevity, it will not be repeated here.
[0105] The aforementioned multilingual speech testing device can be implemented as a computer program, which can perform tasks such as... Figure 2 It runs on the computer device shown.
[0106] Please see Figure 2 , Figure 2 This is a schematic block diagram of a computer device provided in an embodiment of this application. The computer device 500 can be a terminal or a server. The terminal can be an electronic device with communication functions, such as a smartphone, tablet, laptop, desktop computer, personal digital assistant, or wearable device. The server can be a standalone server or a server cluster composed of multiple servers.
[0107] The computer device 500 includes a processor 502, a memory, and a network interface 505 connected via a system bus 501. The memory may include a non-volatile storage medium 503 and internal memory 504.
[0108] The non-volatile storage medium 503 may store an operating system 5031 and a computer program 5032. When the computer program 5032 is executed, it causes the processor 502 to execute a multilingual speech testing method.
[0109] The processor 502 provides computing and control capabilities to support the operation of the entire computer device 500.
[0110] The internal memory 504 provides an environment for the execution of the computer program 5032 in the non-volatile storage medium 503. When the computer program 5032 is executed by the processor 502, the processor 502 can execute a multilingual speech testing method.
[0111] The network interface 505 is used for network communication with other devices. Those skilled in the art will understand that the above structure is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device 500 to which the present application is applied. A specific computer device 500 may include more or fewer components than shown in the figures, or combine certain components, or have different component arrangements.
[0112] The processor 502 is used to run a computer program 5032 stored in a memory to implement the steps of a multilingual speech testing method provided in any of the above method embodiments.
[0113] It should be understood that in the embodiments of this application, the processor 502 may be a central processing unit (CPU), or it may be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor.
[0114] It will be understood by those skilled in the art that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program may be stored in a storage medium, which is a computer-readable storage medium. The computer program is executed by at least one processor in the computer system to implement the process steps of the embodiments of the above methods.
[0115] Therefore, the present invention also provides a storage medium. This storage medium can be a computer-readable storage medium. The storage medium stores a computer program. When executed by a processor, the computer program causes the processor to perform the steps of a multilingual speech testing method provided in any of the above-described method embodiments.
[0116] The storage medium is a physical, non-transient storage medium, such as a USB flash drive, external hard drive, read-only memory (ROM), magnetic disk, or optical disk, or any other physical storage medium capable of storing program code. The computer-readable storage medium can be non-volatile or volatile.
[0117] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.
[0118] In the several embodiments provided by this invention, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative. For example, the division of each unit is merely a logical functional division, and there may be other division methods in actual implementation. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed.
[0119] The steps in the method of this invention can be adjusted, merged, or reduced in order according to actual needs. The units in the device of this invention can be merged, divided, or reduced according to actual needs. Furthermore, the functional units in the various embodiments of this invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0120] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, a terminal, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention.
[0121] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.
[0122] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Since these modifications and variations fall within the scope of the claims and their equivalents, this invention also intends to include these modifications and variations.
[0123] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and these modifications or substitutions should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A multilingual speech testing method, characterized in that, include: Acquire standard pronunciation data and accent feature data; Extract the spectral features and tone features of the accent feature data; The accent model is determined based on the spectral and tonal characteristics. A preset neural speech synthesizer generates accented speech test instructions based on the accent model and the standard pronunciation data; The voice assistant was tested based on the accented speech test command, and the test results were obtained.
2. The multilingual speech testing method according to claim 1, characterized in that, The acquisition of standard pronunciation data and accent feature data includes: Receive a test instruction, which includes language configuration information and accent configuration information; The standard pronunciation data is obtained from a preset speech corpus based on the language configuration information; The accent feature data is obtained from a preset speech corpus based on the accent configuration information.
3. The multilingual speech testing method according to claim 1, characterized in that, The extraction of spectral features and tone features from the accent feature data includes: The spectral features of the accent feature data are extracted based on the preset Mel frequency cepstral coefficients; The accent features are extracted from the accent feature data based on a preset fundamental frequency extraction algorithm.
4. The multilingual speech testing method according to claim 1, characterized in that, The step of determining the accent model based on the spectral features and tonal features includes: The accent model is constructed based on the spectral and tonal features using a preset variational autoencoder or generative adversarial network.
5. The multilingual speech testing method according to claim 1, characterized in that, The step of generating accented speech test instructions based on the accent model and the standard pronunciation data using a preset neural speech synthesizer includes: The parameters of the accent model and the standard pronunciation data are input into the neural speech synthesizer. The neural speech synthesizer modulates the standard pronunciation data according to the parameters of the accent model to obtain the accented speech test command.
6. The multilingual speech testing method according to claim 5, characterized in that, Before testing the voice assistant based on the accented speech test command and obtaining the test results, the method further includes: The naturalness of the accented speech test command is determined by a preset speech quality assessment algorithm. If the naturalness of the speech meets the preset naturalness of speech conditions, the step of testing the voice assistant based on the accented speech test instruction and obtaining the test result is executed.
7. The multilingual speech testing method according to claim 1, characterized in that, After testing the voice assistant based on the accented speech test command and obtaining the test results, the method further includes: The voice assistant was optimized and updated based on the test results.
8. A multilingual speech testing device, characterized in that, Includes a unit for performing the method as described in any one of claims 1-7.
9. A computer device, characterized in that, The computer device includes a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the method as described in any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The storage medium stores a computer program that, when executed by a processor, can implement the method as described in any one of claims 1-7.
Citation Information
Patent Citations
Intelligent household electrical appliance intelligent level test system and method based on voice interaction
CN112383451A
Equipment performance test method and device, storage medium and electronic device
CN113595811A
End-to-end dialect audio recognition method and system based on multilayer information fusion
CN118737123A
Mandarin tone pronunciation level detection method, system, medium and product oriented to Anmulti-Tibetan districts
CN118737194A
Chinese dialect accent correction method and system adopting acoustic unit
CN119049507A