Cross-language speech recognition conversion method

Through the cross-language speech recognition method of multilingual corpus and neural network architecture, the problems of high recognition complexity and low accuracy in the existing technology are solved, and efficient cross-language speech recognition and text quality improvement are achieved.

CN120544578APending Publication Date: 2025-08-26宋婧婧 +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510784534.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-12
Publication Date
2025-08-26

AI Technical Summary

Technical Problem

In the prior art, in cross-language speech recognition, the complexity of the speech recognition algorithm is increased after data processing, making it difficult to ensure the accuracy of the recognition results.

Method used

Using a matching model containing a multilingual corpus, speech recognition is achieved through denoising, feature extraction and classification of speech data, and a neural network architecture combining deep neural networks and long and short-term memory networks is used for speech recognition, combining semantic analysis and text proofreading to achieve accurate cross-language matching.

Benefits of technology

It effectively removes noise, improves the purity of the audio collection, recognizes the text content and understands the semantic meaning, ensures the accuracy of the text and further improves the text quality through text proofreading.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120544578A_ABST
    Figure CN120544578A_ABST
Patent Text Reader

Abstract

The invention relates to a cross-language speech recognition conversion method, and relates to the technical field of language speech recognition processing, the cross-language speech recognition conversion method comprises a multi-language corpus composed of a matching model containing at least two languages, and the matching model comprises speech data of the languages and corresponding text data; the method comprises the following steps: S01, carding input voice data to obtain a de-noised audio set; s02, extracting and classifying features in the audio set to obtain text data and semantic data; and S03, inputting the text data into the obtained corresponding matching model based on the semantic data so as to obtain language specific text features. Through the steps of multi-language corpus building, fine feature processing, efficient acoustic model, flexible matching model and the like, accuracy and high efficiency of cross-language speech recognition are realized, the complexity of an algorithm is reduced, and the accuracy of a recognition result is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of language speech recognition processing, and in particular to a cross-language speech recognition conversion method. Background Art

[0002] The collected audio data is converted into electrical signals, and then subjected to noise reduction and normalization processing to obtain clearer data. Custom searches are performed on the database (cloud or built-in memory) based on the acquired data to lock the language corresponding to the collected audio data. The data is then converted into audio data for output. This process is the process of speech recognition conversion.

[0003] Chinese patent application number 2017113710958 discloses an intelligent cross-language speech recognition conversion method. This method divides speech data into categories based on language families, establishes distances between language family classes, and then performs detailed language recognition within the language family after preliminarily determining the language family of the speech data to be recognized. If the initial language family recognition is incorrect, the method can further search for adjacent language families based on the established distances between language family classes to confirm the language family. After the language family is recognized, the method converts the recognized language into standardized text and performs word segmentation, word frequency statistics, and other processing to establish a mapping relationship for subsequent speech queries. This invention effectively solves the current problem of unbalanced efficiency and speed in speech recognition. Furthermore, it provides a more reasonable processing of speech-to-text conversion, and the establishment of a mapping relationship increases the efficiency and accuracy of recognition conversion.

[0004] Chinese patent application number CN2019101122992 discloses an intelligent cross-language speech recognition and conversion method, comprising the following steps: S1: speech collection, S2: speech recognition, S3: speech conversion, and S4: speech recognition judgment. The invention is ingeniously designed to achieve speech recognition and conversion. The method facilitates speech denoising to ensure speech quality, facilitating communication between people in different regions. The speech collection module facilitates the collection of speech to be recognized, ensuring the prerequisites for speech conversion. The first and second speech recognition modules facilitate speech recognition and judgment by the first speech recognition module, followed by a determination of the speech language family of the converted speech by the second speech recognition module. This facilitates the classification and storage of speech, and facilitates access to the speech database by other devices.

[0005] In the existing technologies including the above two patents, after data processing, the data to be outputted are directly subjected to basic feature extraction for the speech data to be recognized on the basis of building a corpus, and then the process of text conversion is performed. This process significantly increases the complexity of the speech recognition algorithm and makes it difficult to ensure the accuracy of the recognition results. Summary of the Invention

[0006] In response to the above technical problems, the technical solution adopted by the present invention is a cross-language speech recognition conversion method, comprising a multilingual corpus consisting of a matching model containing at least two or more languages, wherein the matching model includes speech data of the language and corresponding text data;

[0007] The method comprises the following steps:

[0008] S01. Combing the input voice data to obtain a denoised audio set;

[0009] S02. Extracting and classifying features in the audio collection to obtain text data and semantic data;

[0010] S03, inputting the text data into a corresponding matching model based on the semantic data to obtain language-specific text features;

[0011] S04. Processing language-specific text features based on a text generation algorithm to obtain a final recognized text and output it.

[0012] Preferably, the step S01 of sorting out the input voice data includes:

[0013] S11, dividing the extracted voice data into at least ten window data sets of equal period according to the recording time length;

[0014] S12, extracting audio curves of different decibels contained in each window data to obtain an audio curve set;

[0015] S13, extracting the amplitude and frequency of the audio corresponding to the speaking rate and intonation in the speech data, and matching them with the audio curve set under the first window data set to obtain a similar audio curve;

[0016] S14, searching based on the obtained similar audio curves in each of the window data sets, and eliminating the remaining curves to obtain a complete curve data by fitting.

[0017] Preferably, the step of extracting the amplitude and frequency of the audio corresponding to the speaking rate and intonation in the voice data in step S13 includes:

[0018] S131. Read the voice data using a Python audio processing library, and then directly obtain amplitude data using a SciPy function. The amplitude data is a number of discrete scattered points. Connect the plurality of discrete scattered points in chronological order to obtain a visual graph.

[0019] S132. Convert the speech data from the time domain to the frequency domain using a fast Fourier transform, and extract frequency features, where the frequency features include a spectrum centroid and a spectrum roll-off point.

[0020] S133, analyzing the speech data by formant to obtain frequency features related to speech rate and intonation;

[0021] S134 , fusing the frequency features obtained in steps S131 , S132 , and S133 to obtain a frequency range.

[0022] Preferably, the step of extracting and classifying features in the audio set in step S02 includes:

[0023] S21, determining a zero point of frequency in the curve data of the audio set as a cutoff point, and a curve data segment between every two zero points is a frame;

[0024] S22, extracting the feature vector of each frame, and arranging the obtained feature vectors in order to obtain a feature sequence;

[0025] S23. Input the feature vector sequence into a pre-trained acoustic model, and obtain text data through the forward propagation and decoding algorithm of the acoustic model;

[0026] S24. Perform semantic analysis on the text data to extract semantic data.

[0027] Preferably, the acoustic model adopts a neural network architecture that combines a deep neural network and a long short-term memory network.

[0028] Preferably, the text data obtained in step S23 is the feature vector sequence of the recognized part in the pre-trained acoustic model, and the feature vector sequence that cannot be recognized is used as a training set, and the pre-trained acoustic model is trained by the training set, and the steps include:

[0029] S231, importing the training set into the pre-trained acoustic model, and performing forward propagation and backpropagation;

[0030] S232, calculating the model error according to the loss function, and updating the parameters of the pre-trained acoustic model using the gradient descent method to obtain a new pre-trained acoustic model;

[0031] S233. Import the feature sequence obtained in step S22 into the new pre-trained acoustic model. If the feature vector sequence that cannot be recognized is still obtained, record the current number of iterations, calculate the recall rate of the acoustic model, and form a new training set and return to step S231 for execution.

[0032] Preferably, the step S04 further includes performing text proofreading on the final recognized text, the steps of which include:

[0033] S41, performing semantic judgment on the obtained final recognized text, if the judgment result is that the semantics are smooth, then directly output the text; if not, proceed to step S42;

[0034] S42, extracting and removing words that are obviously semantically incorrect from the final recognized text;

[0035] S43: Perform semantic judgment again based on the final recognized text after removing the incorrect words, and generate recommended replacement words based on the judgment result and the final recognized text in the previous section, and replace them;

[0036] The determination is based on the determined matching model and the grammatical rules of the target language.

[0037] The present invention has at least the following beneficial effects:

[0038] 1. By carefully combing the input voice data, the noise is effectively removed and a purer audio collection is obtained.

[0039] 2. By extracting and classifying features from audio collections, we obtain textual and semantic data. This not only helps identify the textual content in speech but also understands the underlying semantic meaning, providing rich information for cross-language recognition.

[0040] 3. Inputting text data into a matching model based on semantic data yields language-specific text features, thereby fully utilizing the resources in the multilingual corpus and achieving accurate cross-language matching.

[0041] 4. Not only does it ensure the accuracy of the text, but it also further improves the quality of the text through text proofreading. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0043] Figure 1 A flow chart of a cross-language speech recognition conversion method provided in Example 1 of the present invention;

[0044] Figure 2 This is a flowchart of S01 provided in the first embodiment of the present invention;

[0045] Figure 3 This is a flowchart of S13 provided in the first embodiment of the present invention;

[0046] Figure 4 This is a flowchart of S02 provided in the first embodiment of the present invention;

[0047] Figure 5 This is a flowchart of S23 provided in the first embodiment of the present invention;

[0048] Figure 6 This is a flowchart of S04 provided in the first embodiment of the present invention. DETAILED DESCRIPTION

[0049] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without making any creative efforts shall fall within the scope of protection of the present invention.

[0050] It should be noted that the terms "first", "second", etc. in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way are interchangeable where appropriate so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations are intended to cover non-exclusive inclusions. For example, a process, method, system, product or server that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0051] Example 1

[0052] This embodiment provides a cross-language speech recognition conversion method, including a multilingual corpus consisting of a matching model comprising at least two or more languages, the matching model including speech data of the languages ​​and corresponding text data;

[0053] The method comprises the following steps:

[0054] S01. Combing the input voice data to obtain a denoised audio set;

[0055] S02. Extract and classify features in the audio collection to obtain text data and semantic data;

[0056] S03, inputting the text data into a corresponding matching model based on the semantic data to obtain language-specific text features;

[0057] S04. Processing language-specific text features based on a text generation algorithm to obtain a final recognized text and output it.

[0058] Specific, combined Figure 2 It can be seen that the step S01 of sorting out the input voice data includes:

[0059] S11, dividing the extracted voice data into at least ten window data sets of equal period according to the recording time length;

[0060] S12, extracting audio curves of different decibels contained in each window data to obtain an audio curve set;

[0061] S13, extracting the amplitude and frequency of the audio corresponding to the speaking rate and intonation in the speech data, and matching them with the audio curve set under the first window data set to obtain a similar audio curve;

[0062] S14. Based on the obtained similar audio curves, search in each window data set, remove the remaining curves, and fit to obtain a complete curve data.

[0063] The above-mentioned fine division and processing of speech data obtains more accurate audio curves and features, which helps to improve the robustness and accuracy of the speech recognition system.

[0064] Specific, combined Figure 3 It can be seen that the step of extracting the amplitude and frequency of the audio corresponding to the speaking rate and intonation in the voice data in step S13 includes:

[0065] S131. Read the voice data using a Python audio processing library, and then directly obtain amplitude data using a SciPy function. The amplitude data is a number of discrete scattered points. Connect the plurality of discrete scattered points in chronological order to obtain a visual graph.

[0066] S132. Convert the speech data from the time domain to the frequency domain using a fast Fourier transform, and extract frequency features, where the frequency features include a spectrum centroid and a spectrum roll-off point.

[0067] S133, analyzing the speech data by formant to obtain frequency features related to speech rate and intonation;

[0068] S134 , fusing the frequency features obtained in steps S131 , S132 , and S133 to obtain a frequency range.

[0069] By extracting features such as amplitude and frequency and fusing them, we obtain a more comprehensive frequency range, which helps the system better recognize speech at different speaking speeds and intonations.

[0070] Specific, combined Figure 4 It can be seen that the steps of extracting and classifying features from the audio set in step S02 include:

[0071] S21, determining a zero point of frequency in the curve data of the audio set as a cutoff point, and the curve data segment between every two zero points is a frame;

[0072] S22, extracting the feature vector of each frame, and arranging the obtained feature vectors in order to obtain a feature sequence;

[0073] S23. Input the feature vector sequence into a pre-trained acoustic model, and obtain text data through the forward propagation and decoding algorithm of the acoustic model;

[0074] S24. Perform semantic analysis on the text data to extract semantic data.

[0075] The acoustic model described above utilizes a neural network architecture that combines a deep neural network with a long short-term memory network. Furthermore, this combined architecture creates a highly efficient acoustic model. This allows for accurate recognition of text data within feature vector sequences and allows for training and optimization of unrecognizable portions.

[0076] Specific, combined Figure 5 It can be seen that the text data obtained in step S23 is the feature vector sequence of the recognized part in the pre-trained acoustic model, and the unrecognizable feature vector sequence is used as a training set, and the pre-trained acoustic model is trained using the training set. The steps include:

[0077] S231, importing the training set into the pre-trained acoustic model and performing forward propagation and backpropagation;

[0078] S232, then calculating the model error according to the loss function, and using the gradient descent method to update the parameters of the pre-trained acoustic model to obtain a new pre-trained acoustic model;

[0079] S233. Import the feature sequence obtained in step S22 into a new pre-trained acoustic model. If an unrecognizable feature vector sequence is still obtained, record the current number of iterations, calculate the recall rate of the acoustic model, and form a new training set and return to step S231 for execution.

[0080] The above-mentioned iterative training method continuously improves the performance of the acoustic model, which not only helps to improve recognition accuracy, but also makes the model more adaptable to different speech environments and conditions.

[0081] Specific, combined Figure 6 It can be seen that step S04 also includes text proofreading processing on the final recognized text, and the steps include:

[0082] S41, performing semantic judgment on the obtained final recognized text. If the judgment result is that the semantics are smooth, the text is directly output; if not, executing step S42;

[0083] S42. Extracting and removing words with obvious semantic errors from the final recognized text;

[0084] S43: Perform semantic judgment again based on the final recognized text after removing the incorrect words, and generate recommended replacement words based on the judgment result and the final recognized text of the previous paragraph, and replace them;

[0085] The judgment is based on the determined matching model and the grammatical rules of the target language.

[0086] As mentioned above, the final recognized text is first subjected to semantic judgment. If the text is semantically coherent, the system will directly output the text without any additional processing. The efficient execution of this step is due to the established matching model and the grammatical rules of the target language, which provide the system with an accurate basis for semantic judgment. Furthermore, if the final recognized text is semantically incoherent, the system will proceed to step S42. In this step, the system uses advanced natural language processing technology to accurately identify and remove words in the text that are obviously semantically incorrect. This processing not only helps to improve the accuracy of the text but also provides a purer text foundation for subsequent replacement operations. Secondly, based on the final recognized text after removing the incorrect words, semantic judgment is performed again. Based on the judgment results and the previous recognized text, the system can intelligently generate recommended replacement words and perform the replacement operation. The implementation of this step is due to the system's in-depth understanding of the grammatical rules of the target language and analysis of a large amount of corpus data. Through the replacement operation, the system can further optimize the quality of the text, making it more consistent with the expression habits of the target language.

[0087] From the above, it can be seen that by carefully combing the input voice data, the noise is effectively removed and a purer audio set is obtained. In addition, by extracting and classifying the features in the audio set, text data and semantic data are obtained. This not only helps to identify the text content in the voice, but also understands the semantic meaning behind it, providing rich information for cross-language recognition. Secondly, the text data is input into the matching model obtained based on the semantic data to obtain language-specific text features, thereby making full use of the resources in the multilingual corpus and achieving accurate cross-language matching. Furthermore, the present invention not only ensures the accuracy of the text, but also further improves the quality of the text through text proofreading.

[0088] Example 2

[0089] An embodiment of the present invention provides a non-transitory computer-readable storage medium, wherein the non-transitory computer-readable storage medium stores at least one instruction or at least one program, and the at least one instruction or at least one program is loaded and executed by a processor to implement the steps:

[0090] A multilingual corpus is created that consists of matching models for at least two or more languages, where the matching models include speech data and corresponding text data for the languages.

[0091] And based on the storage in the readable storage medium, the buffered data in the process of "combing the input voice data to obtain a denoised audio set, then extracting and classifying the features in the audio set to obtain text data and semantic data, and then inputting the text data into the corresponding matching model based on the semantic data to obtain language-specific text features; finally, processing the language-specific text features based on the text generation algorithm to obtain the final recognized text, and outputting it" is executed.

[0092] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiment methods can be implemented by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application may include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0093] Those skilled in the art will clearly understand that for the sake of convenience and brevity in description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0094] Example 3

[0095] An embodiment of the present invention provides an electronic device, including a processor and a memory, wherein the memory stores at least one instruction or at least one program, and the at least one instruction or at least one program is loaded and executed by the processor to implement the following steps:

[0096] The input speech data is sorted to obtain a denoised audio set, and then the features in the audio set are extracted and classified to obtain text data and semantic data. The text data is then input into the corresponding matching model based on the semantic data to obtain language-specific text features; finally, the language-specific text features are processed based on the text generation algorithm to obtain the final recognized text and output it.

[0097] The above description is merely a preferred embodiment of the present invention and does not constitute any form of limitation to the present invention. Although the present invention has been disclosed as a preferred embodiment as above, it is not intended to limit the present invention. Any technician familiar with the present profession can make some changes or modifications to equivalent embodiments of the technical contents disclosed above without departing from the scope of the technical solution of the present invention. However, any simple modifications, equivalent changes and modifications made to the above embodiments based on the technical essence of the present invention without departing from the content of the technical solution of the present invention are still within the scope of the technical solution of the present invention.

Claims

1. A cross-language speech recognition conversion method, characterized in that: A multilingual corpus comprising a matching model comprising at least two or more languages, wherein the matching model comprises speech data of the languages ​​and corresponding text data; The method comprises the following steps: S01. Combing the input voice data to obtain a denoised audio set; S02. Extracting and classifying features in the audio collection to obtain text data and semantic data; S03, inputting the text data into a corresponding matching model based on the semantic data to obtain language-specific text features; S04. Processing language-specific text features based on a text generation algorithm to obtain a final recognized text and output it.

2. A cross-language speech recognition conversion method according to claim 1, characterized in that: The step S01 of sorting out the input voice data includes: S11, dividing the extracted voice data into at least ten window data sets of equal period according to the recording time length; S12, extracting audio curves of different decibels contained in each window data to obtain an audio curve set; S13, extracting the amplitude and frequency of the audio corresponding to the speaking rate and intonation in the speech data, and matching them with the audio curve set under the first window data set to obtain a similar audio curve; S14, searching based on the obtained similar audio curves in each of the window data sets, and eliminating the remaining curves to obtain a complete curve data by fitting.

3. A cross-language speech recognition conversion method according to claim 2, characterized in that: The step of extracting the amplitude and frequency of the audio corresponding to the speaking rate and intonation in the voice data in step S13 includes: S131. Read the voice data using a Python audio processing library, and then directly obtain amplitude data using a SciPy function. The amplitude data is a number of discrete scattered points. Connect the plurality of discrete scattered points in chronological order to obtain a visual graph. S132. Convert the speech data from the time domain to the frequency domain using a fast Fourier transform, and extract frequency features, where the frequency features include a spectrum centroid and a spectrum roll-off point. S133, analyzing the speech data by formant to obtain frequency features related to speech rate and intonation; S134 , fusing the frequency features obtained in steps S131 , S132 , and S133 to obtain a frequency range.

4. The cross-language speech recognition conversion method according to claim 1, characterized in that: The step of extracting and classifying features in the audio set in step S02 includes: S21, determining a zero point of frequency in the curve data of the audio set as a cutoff point, and a curve data segment between every two zero points is a frame; S22, extracting the feature vector of each frame, and arranging the obtained feature vectors in order to obtain a feature sequence; S23. Input the feature vector sequence into a pre-trained acoustic model, and obtain text data through the forward propagation and decoding algorithm of the acoustic model; S24. Perform semantic analysis on the text data to extract semantic data.

5. A cross-language speech recognition conversion method according to claim 4, characterized in that: The acoustic model adopts a neural network architecture that combines a deep neural network and a long short-term memory network.

6. A cross-language speech recognition conversion method according to claim 4, characterized in that: The text data obtained in step S23 is the feature vector sequence of the recognized part in the pre-trained acoustic model, and the feature vector sequence that cannot be recognized is used as a training set, and the pre-trained acoustic model is trained using the training set, the steps of which include: S231, importing the training set into the pre-trained acoustic model, and performing forward propagation and backpropagation; S232, calculating the model error according to the loss function, and updating the parameters of the pre-trained acoustic model using the gradient descent method to obtain a new pre-trained acoustic model; S233. Import the feature sequence obtained in step S22 into the new pre-trained acoustic model. If the feature vector sequence that cannot be recognized is still obtained, record the current number of iterations, calculate the recall rate of the acoustic model, and form a new training set and return to step S231 for execution.

7. The cross-language speech recognition conversion method according to claim 1, characterized in that: The step S04 further includes proofreading the final recognized text, which includes the following steps: S41, performing semantic judgment on the obtained final recognized text, if the judgment result is that the semantics are smooth, then directly output the text; if not, proceed to step S42; S42, extracting and removing words that are obviously semantically incorrect from the final recognized text; S43: Perform semantic judgment again based on the final recognized text after removing the incorrect words, and generate recommended replacement words based on the judgment result and the final recognized text in the previous section, and replace them; The determination is based on the determined matching model and the grammatical rules of the target language.

8. A non-transitory computer-readable storage medium, wherein at least one instruction or at least one program is stored in the non-transitory computer-readable storage medium, characterized in that: The at least one instruction or the at least one program is loaded and executed by the processor to implement the cross-language speech recognition conversion method according to any one of claims 1 to 7.

9. An electronic device, characterized in that: The invention comprises a processor and a memory, wherein the memory stores at least one instruction or at least one program, and the at least one instruction or at least one program is loaded and executed by the processor to implement the cross-language speech recognition conversion method according to any one of claims 1 to 7.