Voice Decoding Method, System, Storage Medium and Terminal
By combining the syllable and time information of the acoustic model and the clear voice classification model for screening and weighting calculations during the speech decoding process, the problem of low decoding accuracy in complex speech environments is solved, and higher decoding accuracy is achieved.
Patent Information
- Application Number
- CN202210499737.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-09
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2042-05-09
AI Technical Summary
Existing speech decoders are difficult to obtain correct decoding results in complex speech environments, mainly due to the inaccuracy of the acoustic model scores.
By inputting the audio data into the acoustic model and the clear and voiced classification model, the initial decoding results are obtained, and the syllable information and time information are used for screening and weighting calculations to obtain the target decoding results.
Improves the decoding accuracy in complex speech environments, reduces the amount of calculation and improves the accuracy of decoding results.
Smart Images

Figure CN114743548B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of speech recognition, and in particular to a speech decoding method, system, storage medium and terminal. Background Art
[0002] In the current speech recognition technology, the existing decoder mainly uses a weighted finite state transducer to combine files such as an acoustic model, context dependency, pronunciation dictionary, and language model through a series of algorithms such as minimization and determinization to form a static decoding graph. The decoding process is to find the optimal path in the decoding graph. However, due to the need to rely on the relative accuracy of the acoustic model score in the above method, in some complex speech environments, when the acoustic model score itself is not very reliable, it is difficult to obtain the correct decoding result.
[0003] Therefore, it is necessary to provide a new type of speech decoding method, system, storage medium and terminal to solve the above problems existing in the prior art. Summary of the Invention
[0004] The purpose of the present invention is to provide a speech decoding method, system, storage medium and terminal, which improve the accuracy of speech decoding in complex speech environments.
[0005] In a first aspect, to achieve the above object, the speech decoding method of the present invention includes:
[0006] Input the audio data to be decoded into an acoustic model, and obtain several groups of initial decoding results through the acoustic model. The initial decoding results include an initial decoding sequence and an initial likelihood score;
[0007] Input the audio data to be decoded into a voiceless / voiced classification model to obtain voiceless / voiced sequence information, where the voiceless / voiced sequence information includes first syllable information and first moment information;
[0008] Screen each group of the initial decoding results according to the first syllable information and the first moment information to obtain a target decoding result to complete the decoding.
[0009] The screening process of each group of the initial decoding results according to the first syllable information and the first moment information to obtain a target decoding result and complete the decoding includes:
[0010] Obtain the initial decoding sequence and the corresponding initial likelihood score in each group of the initial decoding results, where the initial decoding sequence includes second syllable information and second moment information;
[0011] Compare the second syllable information in each of the initial decoding sequences with the first syllable information to obtain syllable difference information;
[0012] Compare the second time information in each of the initial decoding sequences with the first time information to obtain time difference information;
[0013] Perform weighted calculation based on the syllable difference information and the time difference information to obtain the weighted score of each group of the initial decoding results;
[0014] Accumulate and sum the weighted scores of each group of the initial decoding results with the initial likelihood scores to obtain weighted likelihood scores;
[0015] Sort the weighted likelihood scores calculated for each of the initial decoding results, and use the initial decoding result with the highest weighted likelihood score as the target decoding result to complete decoding.
[0016] Optionally, the performing weighted calculation based on the syllable difference information and the time difference information to obtain the weighted score of each group of the initial decoding results includes:
[0017] Obtain the difference value of each syllable between the second syllable information and the first syllable information according to the syllable difference information, and calculate the syllable weighted score based on the difference value of each syllable;
[0018] Obtain the difference value of each syllable time between the second time information and the first time information according to the time difference information, and calculate the time weighted score based on the difference value of each syllable time;
[0019] Obtain the weighted score of each decoding result based on the syllable weighted score and the time weighted score.
[0020] Optionally, the first time information is the time information corresponding to each syllable in the first syllable information, and the second time information is the time information corresponding to each syllable in the second syllable information.
[0021] Optionally, the obtaining several groups of initial decoding results through the acoustic model includes:
[0022] Parse the audio data according to the basic pronunciation dictionary rules to output several groups of initial decoding results.
[0023] Optionally, the training process of the voiceless / voiced classification model includes:
[0024] Determine the voiceless / voiced classification information of each phoneme in the pronunciation dictionary according to the spectrogram, and determine the mapping table of each phoneme corresponding to voiceless / voiced;
[0025] Map the phoneme sequence information in the training data of the acoustic model to voiceless / voiced training information according to the mapping table;
[0026] Select a neural network model as the initial model, and input the voiceless / voiced training information to train the initial model;
[0027] After the coincidence degree between the output result of the trained initial model and the correct result corresponding to the input voiceless / voiced training information reaches a preset threshold, obtain the voiceless / voiced classification model.
[0028] Optionally, the input dimension of the initial model is determined according to the voiceless / voiced training information, and the output dimension of the initial model is determined according to the voiceless / voiced classification information.
[0029] In a second aspect, the present invention also provides a speech decoding system, which includes:
[0030] An initial decoding module, configured to input the audio data to be decoded into an acoustic model, and obtain several groups of initial decoding results through the acoustic model, where the initial decoding results include an initial decoding sequence and an initial likelihood score;
[0031] A voiceless / voiced classification module, configured to input the audio data to be decoded into a voiceless / voiced classification model to obtain voiceless / voiced sequence information, where the voiceless / voiced sequence information includes first syllable information and first moment information;
[0032] A target decoding module, configured to perform screening processing on each group of the initial decoding results according to the first syllable information and the first moment information to obtain a target decoding result, and complete the decoding.
[0033] In a third aspect, the present invention also discloses a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the above-mentioned speech decoding method is implemented.
[0034] In a fourth aspect, the present invention also provides a terminal, including: a processor and a memory;
[0035] The memory is used to store a computer program;
[0036] The processor is used to execute the computer program stored in the memory, so that the terminal executes the above-mentioned speech decoding method. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] Figure 1 It is a schematic diagram of the overall flow of the speech decoding method according to the embodiment of the present invention;
[0038] Figure 2 It is a schematic flowchart of the training process of the voiceless / voiced classification model in the voice decoding method according to the embodiment of the present invention;
[0039] Figure 3 It is a schematic flowchart of step S103 in the terminal upgrade method according to the embodiment of the present invention;
[0040] Figure 4 It is a structural block diagram of the voice decoding system according to the embodiment of the present invention;
[0041] Figure 5 It is a structural block diagram of the device according to the embodiment of the present invention. Detailed implementation manners
[0042] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings of the present invention. Apparently, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention. Unless otherwise defined, the technical terms or scientific terms used herein shall have the ordinary meanings as understood by those of ordinary skill in the art in the field to which the present invention belongs. The words such as "including" used herein mean that the elements or items appearing before this word cover the elements or items listed after this word and their equivalents, without excluding other elements or items.
[0043] In view of the problems existing in the prior art, an embodiment of the present invention provides a voice decoding method. Referring to Figure 1 , the method includes the following steps:
[0044] S101. Input the audio data to be decoded into an acoustic model, and obtain several groups of initial decoding results through the acoustic model. The initial decoding results include an initial decoding sequence and an initial likelihood score.
[0045] In some embodiments, the obtaining several groups of initial decoding results through the acoustic model includes:
[0046] Parse the audio data according to the basic pronunciation dictionary rules to output several groups of initial decoding results, and each group of initial decoding results contains a group of initial decoding sequences and initial likelihood scores, so as to facilitate subsequent screening processing of the initial decoding results.
[0047] In this embodiment, the acoustic model can be an acoustic model directly trained using existing technologies to decode the input audio data to obtain an initial decoding result. Alternatively, it can use the collected training data, follow the rules of the basic pronunciation dictionary, and according to a certain network structure, use the training data to train the acoustic model to obtain an acoustic model that meets the requirements, which will not be elaborated here.
[0048] Among them, the initial decoding sequence includes second syllable information and second time information. The second syllable information includes the composition and order of each syllable, and the second time information is the time information corresponding to each syllable in the second syllable information, so as to determine the time situation of each syllable in the second syllable information according to the second time information.
[0049] S102. Input the audio data to be decoded into the voiceless / voiced sound classification model to obtain voiceless / voiced sound sequence information, where the voiceless / voiced sound sequence information includes first syllable information and first time information.
[0050] After inputting the audio data to be decoded into the voiceless / voiced sound classification model, the voiceless / voiced sound classification model processes the audio data to obtain voiceless / voiced sound sequence information. Among them, the voiceless / voiced sound sequence information includes first syllable information and first time information. The first syllable information includes the composition and order of each syllable, and the first time information is the time information corresponding to each syllable in the first syllable information, so as to determine the time situation of each syllable in the first syllable information according to the first time information.
[0051] In some embodiments, refer to Figure 2 , the training process of the voiceless / voiced sound classification model includes the following steps:
[0052] S201. Determine the voiceless / voiced sound classification information of each phoneme in the pronunciation dictionary according to the spectrogram, and determine the mapping table of each phoneme corresponding to voiceless / voiced sound;
[0053] S202. Map the phoneme sequence information in the training data of the acoustic model to voiceless / voiced sound training information according to the mapping table;
[0054] S203. Select a neural network model as the initial model, and input the voiceless / voiced sound training information to train the initial model;
[0055] S204. After the coincidence degree between the output result of the trained initial model and the correct result corresponding to the input voiceless / voiced sound training information reaches a preset threshold, obtain the voiceless / voiced sound classification model.
[0056] Specifically, first, determine the classification information of voiceless and voiced sounds in the pronunciation dictionary according to the spectrogram, so as to determine the mapping table of each phoneme corresponding to voiceless and voiced sounds. Then, obtain the phoneme sequence information in the training data, map the phoneme sequence information into voiceless / voiced training information according to the mapping table, and input the voiceless / voiced training information into the initial model for training. After that, the initial model generates corresponding output results according to the input information, compares the output results with the correct results corresponding to the voiceless / voiced training information, calculates the coincidence degree between the output results and the correct results corresponding to the input voiceless / voiced training information, and after the coincidence degree reaches the preset threshold, the output of the initial model can be used as the final voiceless / voiced classification model.
[0057] It should be noted that the preset threshold can be selected according to the situation. The preset threshold is at least 90%. In this embodiment, the preset threshold is 90%. When the coincidence degree between the output result of the initial model and the correct result corresponding to the input voiceless / voiced training information reaches more than 90%, the initial model is used as the final voiceless / voiced classification model.
[0058] In some embodiments, the initial model is a feed-forward sequence neural network model.
[0059] In some other embodiments, the input dimension of the initial model is determined according to the voiceless / voiced training information, and the output dimension of the initial model is determined according to the voiceless / voiced classification information, so as to ensure that the input dimension and output dimension of the initial model are corresponding.
[0060] Specifically, when selecting the phoneme sequence information in the training data of the acoustic model as the voiceless / voiced training information, the input dimension is determined according to the composition of the data features in the phoneme sequence information; when selecting the fundamental frequency information in the training data of the acoustic model as the voiceless / voiced training information, the input dimension is determined according to the data feature composition of the fundamental frequency information. And the output dimension of the initial model is determined according to the voiceless / voiced classification information, so as to ensure that output results meeting the requirements can be generated under different input conditions.
[0061] S103. Screen each group of the initial decoding results according to the first syllable information and the first moment information to obtain the target decoding result, thereby completing the decoding.
[0062] In some embodiments, referring to Figure 3 , the above process includes the following steps:
[0063] S301. Obtain the initial decoding sequence and the corresponding initial likelihood score in each group of the initial decoding results;
[0064] S302. Compare the second syllable information in each of the initial decoding sequences with the first syllable information to obtain syllable difference information;
[0065] S303. Compare the second time information in each of the initial decoding sequences with the first time information to obtain time difference information;
[0066] S304. Perform weighted calculation based on the syllable difference information and the time difference information to obtain the weighted score of each group of the initial decoding results;
[0067] S305. Accumulate and sum the weighted scores of each group of initial decoding results and the initial likelihood scores to obtain a weighted likelihood score;
[0068] S306. Sort the weighted likelihood scores calculated for each of the initial decoding results, and use the initial decoding result with the highest weighted likelihood score as the target decoding result to complete decoding.
[0069] Specifically, after obtaining the initial decoding results and the voiceless / voiced sound sequence information including the first syllable information and the first time information, since the initial decoding results include the initial decoding sequences and the initial likelihood scores, and the initial decoding sequences include the second syllable information and the second time information. After obtaining the initial decoding sequences and the corresponding initial likelihood scores in each group of initial decoding results, first compare the second syllable information in each initial decoding sequence with the first syllable information, so as to obtain the difference between the composition and arrangement of syllables in the initial decoding results after acoustic model decoding and the composition and arrangement of syllables obtained after processing by the voiceless / voiced sound classification model, that is, to obtain the syllable difference information. Similarly, compare the second time information in each of the initial decoding sequences with the first time information, so as to obtain the difference between the time jump information of syllables in the initial decoding results after acoustic model decoding and the time jump information of syllables obtained after processing by the voiceless / voiced sound classification model, that is, to obtain the time difference information.
[0070] After obtaining the syllable difference information and the time difference information, perform weighted calculation based on the syllable difference information and the time difference information to obtain the weighted score of each group of the initial decoding results. Among them, the greater the syllable difference information, that is, the greater the difference in the composition and arrangement of syllables between the initial decoding result and the voiceless / voiced sound sequence information, the lower the weighted score; on the contrary, the smaller the syllable difference information, that is, the smaller the difference in the composition and arrangement of syllables between the initial decoding result and the voiceless / voiced sound sequence information, the higher the weighted score.
[0071] Correspondingly, the greater the time difference information is, that is, the greater the difference in the syllable jump time information between the initial decoding result and the voiceless / voiced sound sequence information, the lower the weighted score; on the contrary, the smaller the time difference information is, that is, the smaller the difference in the syllable jump time information between the initial decoding result and the voiceless / voiced sound sequence information, the greater the weighted score. This will not be elaborated here.
[0072] It should be noted that in this embodiment, the weighting coefficients of the syllable difference information and the time difference information corresponding to the weighted scores are set values and can be selected according to the situation. In this embodiment, the weighting coefficient is 1.
[0073] After obtaining the weighted scores, the weighted scores corresponding to each group of initial decoding results are respectively accumulated with the initial likelihood scores in the initial decoding results, so as to obtain the weighted likelihood scores corresponding to each group of initial decoding results. Then, the weighted likelihood scores calculated for each of the initial decoding results are sorted, and the initial decoding result with the highest weighted likelihood score is used as the target decoding result, thus completing the entire decoding process.
[0074] In the above decoding process, according to the composition, arrangement, and jump time of syllables, the initial decoding results of the acoustic model are compared, and certain screening and weighting are performed. The initial decoding results that do not conform to the syllable composition and arrangement are directly filtered out, and those that conform to the number of syllables and whose phoneme jump times are basically consistent with the voiceless / voiced sound classification information of the voiceless / voiced sound classification model are weighted. In this way, more accurate decoding results can be obtained while reducing the amount of calculation, and the decoding process is completed. It can effectively solve the problem of low decoding recognition accuracy of a single acoustic model in a complex speech environment.
[0075] In some embodiments, the obtaining the weighted scores of each group of the initial decoding results by weighted calculation according to the syllable difference information and the time difference information includes:
[0076] Obtaining the difference value of each syllable in the second syllable information and the first syllable information according to the syllable difference information, and calculating the syllable weighted score according to the difference value of each syllable;
[0077] Obtaining the difference value of each syllable time in the second time information and the first time information according to the time difference information, and calculating the time weighted score according to the difference value of each syllable time;
[0078] Obtaining the weighted score of each decoding result according to the syllable weighted score and the time weighted score.
[0079] Specifically, each syllable in the second syllable information and the first syllable information is compared according to the syllable difference information. When the syllable in the initial decoding result exactly corresponds to the syllable in the voiceless / voiced classification information, the current weighted syllable score is obtained as a; conversely, when the syllable in the initial decoding result does not correspond to the syllable in the voiceless / voiced classification information, the current weighted syllable score is obtained as b, where a > b, and both a and b are positive numbers.
[0080] Each syllable transition time in the second time information and the first time information is compared according to the time difference information. When the difference between the syllable transition time in the initial decoding result and the syllable transition time in the voiceless / voiced classification information does not exceed the difference threshold, the current weighted time score is obtained as c; conversely, when the difference between the initial decoding result and the syllable transition time in the voiceless / voiced classification information exceeds the difference threshold, the current weighted syllable score is obtained as d, where c > d, and both c and d are positive numbers. The difference threshold is set according to the first time information and the second time information, which will not be elaborated here.
[0081] After completing the above comparison process, the weighted syllable score of each syllable in the initial decoding result and the weighted time score corresponding to each syllable transition time can be obtained. By adding the weighted syllable score and the weighted time score together, the final weighted score can be obtained to facilitate the subsequent screening of the initial decoding result.
[0082] In the above speech decoding method, compared with the traditional HLCG decoding, when the acoustic model decoding result is not reliable in an environment with a low signal-to-noise ratio, the voiceless / voiced classification information given by the voiceless / voiced classification model can be relied on to jointly infer the initial decoding result to obtain a more accurate decoding result and improve the accuracy of the decoding result. Moreover, in an environment with a high signal-to-noise ratio, the credibility of both the acoustic model and the voiceless / voiced classification model is relatively high and consistent, and it basically does not affect the original performance.
[0083] The present invention also provides a speech decoding system. Refer to Figure 4 , the system includes:
[0084] An initial decoding module 401, configured to input the audio data to be decoded into an acoustic model, and obtain a plurality of groups of initial decoding results through the acoustic model, where the initial decoding results include an initial decoding sequence and an initial likelihood score;
[0085] A voiceless / voiced classification module 402, configured to input the audio data to be decoded into a voiceless / voiced classification model to obtain voiceless / voiced sequence information, where the voiceless / voiced sequence information includes first syllable information and first time information;
[0086] The target decoding module 403 is configured to screen and process each group of the initial decoding results according to the first syllable information and the first moment information to obtain a target decoding result, thereby completing the decoding.
[0087] It should be noted that the structure and principle of the above voice decoding system correspond to the steps in the above voice decoding method one by one, so they will not be elaborated here.
[0088] It should be noted that it should be understood that the division of each module of the above device is only a division of logical functions. In actual implementation, it can be fully or partially integrated into a physical entity, or physically separated. And these modules can all be implemented in the form of software called by a processing element; they can also all be implemented in the form of hardware; or some modules can be implemented in the form of software called by a processing element, and some modules can be implemented in the form of hardware. For example, the selection module can be a separately established processing element, or can be integrated in a certain chip of the above system. In addition, it can also be stored in the memory of the above system in the form of program code, and called and executed by a certain processing element of the above system to perform the functions of the above modules. The implementation of other modules is similar. In addition, all or part of these modules can be integrated together or can be independently implemented. The processing element mentioned here can be an integrated circuit with signal processing capabilities. In the implementation process, each step of the above method or each of the above modules can be completed by the integrated logic circuit in the processor element in hardware or in the form of instructions in software.
[0089] For example, the above modules can be one or more integrated circuits configured to implement the above method, such as: one or more Application Specific Integrated Circuits (ASICs), or, one or more Digital Signal Processors (DSPs), or, one or more Field Programmable Gate Arrays (FPGAs), etc. Again, when a certain module above is implemented in the form of a processing element scheduling program code, the processing element can be a general-purpose processor, such as a Central Processing Unit (CPU) or other processors that can call program code. Again, these modules can be integrated together and implemented in the form of a System-On-a-Chip (SOC).
[0090] In some other embodiments of the present application, the embodiments of the present application disclose a device, such as Figure 5As shown, the device 500 may include: one or more processors 501; a memory 502; a display 503; one or more applications (not shown); and one or more computer programs 504. Each of the above components may be connected via one or more communication buses 505. The one or more computer programs 504 are stored in the memory 502 and configured to be executed by the one or more processors 501. The one or more computer programs 504 include instructions.
[0091] The present invention also discloses a computer-readable storage medium having a computer program stored thereon. When the computer program is run by a processor, it executes the above-described voice decoding method.
[0092] A computer program is stored on the storage medium of the present invention. When the computer program is executed by a processor, it implements the above-described method. The storage medium includes various media that can store program codes, such as a read-only memory (ROM), a random access memory (RAM), a magnetic disk, a USB flash drive, a memory card, or an optical disc.
[0093] In another embodiment disclosed by the present invention, the present invention further provides a chip system. The chip system is coupled to a memory and configured to read and execute program instructions stored in the memory to perform the steps of the above-described voice decoding method.
[0094] Through the description of the above embodiments, those skilled in the art can clearly understand that, for the convenience and simplicity of description, only the above division of each functional module is used as an example. In practical applications, the above functions may be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. The specific working processes of the systems, devices, and units described above may refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0095] In each of the embodiments of the present application, the functional units may be integrated into one processing unit, or each unit may exist physically alone, or two or more units may be integrated into one unit. The above integrated unit may be implemented in the form of hardware or in the form of a software functional unit.
[0096] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiments of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the methods described in the various embodiments of the present application. The aforementioned storage medium includes: various media that can store program codes, such as flash memory, mobile hard disks, read-only memories, random access memories, magnetic disks, or optical discs.
[0097] As described above, the above are only the specific implementation manners of the embodiments of the present application, but the protection scope of the embodiments of the present application is not limited thereto. Any changes or substitutions within the technical scope disclosed in the embodiments of the present application should be covered within the protection scope of the embodiments of the present application. Therefore, the protection scope of the embodiments of the present application shall be subject to the protection scope of the claims.
[0098] Although the embodiments of the present invention have been described in detail above, it is obvious to those skilled in the art that various modifications and changes can be made to these embodiments. However, it should be understood that such modifications and changes are all within the scope and spirit of the present invention described in the claims. Moreover, the present invention described herein may have other embodiments and can be implemented or realized in various ways.
Claims
1. A voice decoding method, characterized in that, The method includes: Input the audio data to be decoded into an acoustic model, and obtain several groups of initial decoding results through the acoustic model. The initial decoding results include an initial decoding sequence and an initial likelihood score; Input the audio data to be decoded into a voiceless / voiced classification model to obtain voiceless / voiced sequence information, where the voiceless / voiced sequence information includes first syllable information and first moment information; Perform screening processing on each group of the initial decoding results according to the first syllable information and the first moment information to obtain a target decoding result to complete decoding; The performing screening processing on each group of the initial decoding results according to the first syllable information and the first moment information to obtain a target decoding result and complete decoding includes: Obtain the initial decoding sequence in each group of the initial decoding results and the corresponding initial likelihood score, where the initial decoding sequence includes second syllable information and second moment information; Compare the second syllable information in each initial decoding sequence with the first syllable information to obtain syllable difference information; Compare the second moment information in each initial decoding sequence with the first moment information to obtain moment difference information; Perform weighted calculation according to the syllable difference information and the moment difference information to obtain the weighted score of each group of the initial decoding results; Accumulatively sum the weighted score of each group of initial decoding results and the initial likelihood score to obtain a weighted likelihood score; Sort the weighted likelihood scores calculated for each initial decoding result, and use the initial decoding result with the highest weighted likelihood score as the target decoding result to complete decoding.
2. The voice decoding method according to claim 1, wherein The performing weighted calculation according to the syllable difference information and the moment difference information to obtain the weighted score of each group of the initial decoding results includes: Obtain the difference value of each syllable between the second syllable information and the first syllable information according to the syllable difference information, and calculate a syllable weighted score according to the difference value of each syllable; Obtain the difference value of each syllable moment between the second moment information and the first moment information according to the moment difference information, and calculate a moment weighted score according to the difference value of each syllable moment; Obtain the weighted score of each decoding result according to the syllable weighted score and the moment weighted score.
3. The voice decoding method according to claim 1, characterized in that, The first moment information is the moment information corresponding to each syllable in the first syllable information, and the second moment information is the moment information corresponding to each syllable in the second syllable information.
4. The voice decoding method according to claim 1, characterized in that The obtaining several groups of initial decoding results through the acoustic model includes: Parse the audio data according to the basic pronunciation dictionary rules to output several groups of initial decoding results.
5. The voice decoding method according to any one of claims 1 to 4, characterized in that The training process of the voiceless / voiced classification model includes: Determine the voiceless / voiced classification information of each phoneme in the pronunciation dictionary according to the spectrogram, and determine the mapping table of each phoneme corresponding to voiceless / voiced; Map the phoneme sequence information in the training data of the acoustic model to voiceless / voiced training information according to the mapping table; Select a neural network model as the initial model, and input the voiceless / voiced training information to train the initial model; After the coincidence degree between the output result of the trained initial model and the correct result corresponding to the input voiceless / voiced training information reaches a preset threshold, obtain the voiceless / voiced classification model.
6. The voice decoding method according to claim 5, characterized in that, The input dimension of the initial model is determined according to the voiceless / voiced training information, and the output dimension of the initial model is determined according to the voiceless / voiced classification information.
7. A voice decoding system, characterized in that, The system includes: An initial decoding module, configured to input the audio data to be decoded into an acoustic model, and obtain several groups of initial decoding results through the acoustic model, where the initial decoding results include an initial decoding sequence and an initial likelihood score; A voiceless / voiced classification module, configured to input the audio data to be decoded into the voiceless / voiced classification model to obtain voiceless / voiced sequence information, where the voiceless / voiced sequence information includes first syllable information and first moment information; A target decoding module, configured to perform screening processing on each group of the initial decoding results according to the first syllable information and the first moment information to obtain a target decoding result and complete decoding, The target decoding module performs screening processing on each group of the initial decoding results according to the first syllable information and the first moment information in the following manner to obtain a target decoding result and complete decoding: Obtain the initial decoding sequence and the corresponding initial likelihood score in each group of the initial decoding results, where the initial decoding sequence includes second syllable information and second moment information; Compare the second syllable information in each initial decoding sequence with the first syllable information to obtain syllable difference information; Compare the second moment information in each initial decoding sequence with the first moment information to obtain moment difference information; Perform weighted calculation according to the syllable difference information and the moment difference information to obtain the weighted score of each group of the initial decoding results; Accumulate and sum the weighted score of each group of initial decoding results and the initial likelihood score to obtain a weighted likelihood score; Sort the weighted likelihood scores calculated for each initial decoding result, and use the initial decoding result with the highest weighted likelihood score as the target decoding result to complete decoding.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the voice decoding method according to any one of claims 1 to 6.
9. A terminal, characterized in that, Including: A processor and a memory; The memory is used to store a computer program; The processor is used to execute the computer program stored in the memory, so that the terminal executes the voice decoding method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Method and device for classifying unvoiced sound and voiced sound
CN102655000A
Voice model adaptive training method, system and device and storage medium
CN111243574A