Speech recognition method and device, equipment and storage medium

By adjusting the loudness of speech data and extracting speech segments, combined with text expression constraints, the problem of recognition accuracy in speech recognition models for speech disorders was solved, improving recognition accuracy and text generation quality, and enhancing the communication efficiency of people with speech disorders.

CN120998205APending Publication Date: 2025-11-21BEIJING ZITIAO NETWORK TECH CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511293713.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-10
Publication Date
2025-11-21

AI Technical Summary

Technical Problem

Existing speech recognition models lack accuracy when dealing with speech impediments, resulting in the generated speech content not matching the input speech content.

Method used

By adjusting the loudness of the speech data to a preset range and extracting speech segments associated with speech activities, text content is generated using a model trained based on the preset loudness range, and the text expression is adjusted through text expression constraints to determine the recognition result.

Benefits of technology

It improves the recognition accuracy of speech data with low or high loudness, reduces interference factors in text content, obtains more accurate and reliable recognition results, and enhances the communication efficiency and interactive enthusiasm of people with speech disorders.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120998205A_ABST
    Figure CN120998205A_ABST
Patent Text Reader

Abstract

The embodiment of the invention relates to a voice recognition method and device, equipment and a storage medium. The method provided by the invention comprises the following steps: adjusting the loudness of acquired voice data to a preset loudness range, wherein the voice data is from an object related to dysarthria; extracting at least one speech segment associated with the speech activity from the speech data; providing audio features of the at least one speech segment to a model to generate first text content, where the model is trained based on training speech associated with a preset loudness range; and based on the at least one text expression constraint, adjusting the text expression of the first text content to determine an identification result corresponding to the voice data. In this way, the embodiments of the present disclosure can improve the accuracy of the text content identified by the voice data (e.g., voice data including dysarthria). In addition, through the embodiment of the invention, the interaction efficiency of dysarthria people can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] Example embodiments of the present disclosure generally relate to the field of computers, and in particular, to a method, an apparatus, a device, and a computer-readable storage medium for speech recognition. BACKGROUND

[0002] With the development of the Internet and computer technology, speech processing has also developed vigorously. In the field of speech processing, speech recognition models have received extensive attention and use. Therefore, the recognition accuracy of the speech recognition model has become the focus of attention. SUMMARY

[0003] In a first aspect of the present disclosure, a method for speech recognition is provided. The method comprises adjusting a loudness of acquired speech data to a preset loudness range, the speech data being from an object related to a speech disorder; extracting at least one speech segment associated with a speech activity from the speech data; providing an audio feature of the at least one speech segment to a model to generate a first text content, wherein the model is trained based on training speech associated with the preset loudness range; and adjusting a text expression of the first text content based on at least one text expression constraint to determine a recognition result corresponding to the speech data.

[0004] In a second aspect of the present disclosure, an apparatus for speech recognition is provided. The apparatus comprises a loudness adjustment module configured to adjust a loudness of acquired speech data to a preset loudness range, the speech data being from an object related to a speech disorder; a segment extraction module configured to extract at least one speech segment associated with a speech activity from the speech data; a content generation module configured to provide an audio feature of the at least one speech segment to a model to generate a first text content, wherein the model is trained based on training speech associated with the preset loudness range; and a text adjustment module configured to adjust a text expression of the first text content based on at least one text expression constraint to determine a recognition result corresponding to the speech data.

[0005] In a third aspect of the present disclosure, an electronic device is provided. The device comprises at least one processor; and at least one memory coupled to the at least one processor and storing instructions for execution by the at least one processor. The instructions, when executed by the at least one processor, cause the device to perform the method of the first aspect.

[0006] In a fourth aspect of the present disclosure, a computer-readable storage medium is provided. The computer-readable storage medium has stored thereon computer-executable instructions that are executable by a processor to implement the method of the first aspect.

[0007] It is to be understood that the particulars shown herein are by way of example and for purposes of illustrative discussion of the embodiments of the present disclosure only and are not intended to limit the scope of the present disclosure to the particular embodiments illustrated. Other BRIEF DESCRIPTION OF DRAWINGS

[0008] The above-mentioned and other features and advantages of various embodiments of the present disclosure will become more apparent by reference to the following detailed description taken in conjunction with the accompanying drawings. In the drawings, like reference numerals designate like elements, wherein:

[0009] FIG. 1 A schematic diagram showing an example environment in which embodiments according to the present disclosure can be implemented is shown;

[0010] FIG. 2 A flow diagram showing an example process for speech recognition according to some embodiments of the present disclosure is shown;

[0011] FIG. 3A A schematic diagram showing an inference process for speech recognition according to some embodiments of the present disclosure is shown;

[0012] FIG. 3B A schematic diagram showing a training process for speech recognition according to some embodiments of the present disclosure is shown;

[0013] FIG. 4 A schematic block diagram showing an example apparatus for speech recognition according to some embodiments of the present disclosure is shown; and

[0014] FIG. 5 A block diagram of an electronic device capable of implementing various embodiments of the present disclosure is shown. DETAILED DESCRIPTION

[0015] Embodiments of the present disclosure will be described hereinafter with reference to the accompanying drawings. While certain embodiments of the present disclosure are shown in the drawings, it is understood that the present disclosure can be embodied in various forms and should not be construed as being limited to the embodiments set forth herein; rather, these embodiments are provided so that the present disclosure will be more thoroughly and completely understood. It is to be understood that the drawings and embodiments of the present disclosure are only for illustrative purposes and should not be construed as limiting the scope of the present disclosure.

[0016] It is noted that the headings provided herein are not limitations of the various embodiments described herein. Various embodiments are described throughout this document and can be included under any heading. Additionally, embodiments described in any heading can be combined with any other embodiment described in the same heading and / or a different heading in any manner.

[0017] In the description of the embodiments of the disclosure, the term "comprising" and its similar terms are understood to be open-ended, i.e., "including but not limited to". The term "based on" is understood to be "based at least in part on". The term "one embodiment" or "the embodiment" is understood to be "at least one embodiment". The term "some embodiments" is understood to be "at least some embodiments". The following can also include other explicit and implicit definitions. The terms "first", "second", etc. can refer to different or the same objects. The following can also include other explicit and implicit definitions.

[0018] Embodiments of the present disclosure can involve data of users, acquisition and / or use of data, etc. These aspects all comply with corresponding laws and regulations and relevant provisions. In embodiments of the present disclosure, all data collection, acquisition, processing, processing, forwarding, use, etc. are carried out on the premise that the user is aware of and confirms. Accordingly, when implementing the embodiments of the present disclosure, the type of data or information that can be involved, the scope of use, the use scenario, etc. should be notified to the user and the authorization of the user should be obtained according to the relevant laws and regulations through appropriate means. The specific notification and / or authorization method can vary according to the actual situation and application scenario, and the scope of the present disclosure is not limited in this regard.

[0019] In the present specification and embodiments, if the scheme involves processing of personal information, it will be processed on the premise of having a legal basis (for example, obtaining the consent of the subject of personal information, or being necessary for the performance of a contract, etc.), and will only be processed within the prescribed or agreed range. The user refuses to process personal information other than the necessary information required for the basic function, which does not affect the user's use of the basic function.

[0020] According to the traditional scheme, when the speech recognition model identifies speech with dysarthria (for example, speech with characteristics such as unclear pronunciation, reduced volume, reduced clarity, slowed speech, and reduced language fluency), the speech recognition model generates a phenomenon that does not match the input speech content, thereby affecting the accuracy of the corresponding recognition result of the speech content.

[0021] Embodiments of the present disclosure propose a speech recognition scheme. The scheme includes: adjusting the loudness of obtained speech data to a preset loudness range, the speech data being from an object related to dysarthria; extracting at least one speech segment associated with speech activity from the speech data; providing an audio feature of the at least one speech segment to a model to generate first text content, wherein the model is trained based on training speech associated with the preset loudness range; and adjusting text expression of the first text content based on at least one text expression constraint to determine a recognition result corresponding to the speech data.

[0022] In this way, the embodiments of the present disclosure can adjust the loudness of the speech data to the preset loudness range, and compared with the speech data without loudness adjustment, the recognition accuracy of the model for the speech data with lower loudness or higher loudness can be improved. In addition, the embodiments of the present disclosure can adjust the text expression of the text content based on at least one text expression constraint, so that the interference factors in the generated text content can be reduced, and a more accurate and reliable recognition result can be obtained.

[0023] Various example implementations of the solution are described in further detail below in conjunction with the accompanying drawings.

[0024] Example Environment

[0025] FIG. 1 A schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented is shown. As shown, the example environment 100 can include an electronic device 110 and a model 120. FIG. 1

[0026] In the example environment 100, the electronic device 110 can obtain speech data and generate a corresponding recognition result by using the model 120. Specifically, the electronic device 110 is at least configured to adjust the loudness of the obtained speech data to a preset loudness range, and can generate text content corresponding to at least one speech segment associated with speech activity in the speech data by using the model 120. Further, the electronic device 110 can adjust the text expression of the text content based on a text expression constraint, so as to determine a recognition result corresponding to the speech data.

[0027] In some embodiments, the model 120 can be deployed on the electronic device 110, or can be deployed on other devices other than the electronic device 110, such as a server. That is, the electronic device 110 can invoke the local or remote model 120.

[0028] In some embodiments, the electronic device 110 can be any type of mobile terminal, fixed terminal, or portable terminal including a mobile handset, a desktop computer, a laptop computer, a notebook computer, a netbook computer, a tablet computer, a media player, a multimedia tablet, a personal communication system (PCS) device, a personal navigation device, a personal digital assistant (PDA), an audio / video player, a digital camera / camcorder, a positioning device, a television receiver, a radio broadcast receiver, an electronic book device, a game device, or any combination thereof, including accessories and peripherals of such devices, or any combination thereof. In some embodiments, the electronic device 110 can also support any type of interface to the user (such as "wearable" circuitry, etc.).

[0029] ​It should be understood that the structure and function of the various elements in environment 100 are described for illustrative purposes only and do not imply any limitation on the scope of this disclosure.

[0030] The following description will continue with reference to the accompanying drawings, which will provide some exemplary embodiments of this disclosure.

[0031] Example Process

[0032] FIG. 2 A flowchart of an example speech recognition process 200 according to some embodiments of the present disclosure is shown. Process 200 can be implemented at electronic device 110. Reference is made below. FIG. 1 Describe the process 200.

[0033] like FIG. 2 As shown in box 210, the electronic device 110 adjusts the loudness of the acquired speech data to a preset loudness range, the speech data being from an object related to speech disorders.

[0034] In some embodiments, such speech data may include speech with articulation disorders, such as speech characterized by unclear pronunciation, reduced volume, decreased clarity, slowed speech rate, and reduced fluency. In some embodiments, such speech data or speech with articulation disorders may originate from an object associated with an articulation disorder, such as a person with an articulation disorder.

[0035] In some embodiments, the electronic device 110 may perform loudness normalization preprocessing on the acquired speech data to adjust the loudness of the speech data to a preset loudness range. As an example, the electronic device 110 may perform loudness normalization preprocessing on the acquired speech data based on a loudness standard to adjust the loudness of the speech data to a preset loudness range.

[0036] To facilitate understanding, the following section will introduce an example process for adjusting the loudness of speech data to a preset loudness range.

[0037] In some embodiments, the electronic device 110 can perform peak normalization processing on the voice data. Specifically, the electronic device 110 can adjust the maximum amplitude of the voice data to a specified preset peak value, for example, -1 dB. In this way, the voice data can be prevented from exceeding the maximum permissible amplitude, thereby avoiding distortion.

[0038] Furthermore, the electronic device 110 can calculate the integrated loudness (or loudness value) of the processed speech data. For example, the electronic device 110 can measure the integrated loudness of the speech data based on a loudness standard.

[0039] Specifically, the electronic device 110 can process the speech data based on a loudness standard using a weighted filter (e.g., a K-weighted filter) to conform to the human ear's perception of loudness at different frequencies. For example, this weighted filter can be represented based on the following function:

[0040]

[0041] Where a1(f), a2(f), and a3(f) represent coefficients related to frequency f, which can be defined based on loudness standards. In some embodiments, the electronic device 110 can perform block calculations according to a single time window, for example, a time window of 0.4 seconds.

[0042] Furthermore, the electronic device 110 can calculate the integrated loudness of the speech data processed by the weighted filter. For example, the calculation process of this integrated loudness can be represented based on the following formula:

[0043]

[0044] Among them, L eq The integrated loudness is represented by , T represents the measurement time window, * represents the filtering operation, H(f) represents the filter, and x(t) represents the input speech data.

[0045] In some embodiments, in response to a mismatch between the integrated loudness and the target loudness of the voice data, the electronic device 110 may perform loudness matching based on dynamic range compression or gain adjustment. Exemplarily, this process can be represented based on the following formula:

[0046]

[0047] Where G represents the adjustment gain, L target L represents the target loudness. eq This indicates the integrated loudness of the speech data.

[0048] In this way, embodiments of the present disclosure can adjust the loudness of the acquired speech data to a preset loudness range based on the above example process.

[0049] In box 220, electronic device 110 extracts at least one speech segment associated with a speech activity from the speech data. In some embodiments, such a speech activity may, for example, indicate human voice activity. In some embodiments, electronic device 110 may identify speech activities in the speech data based on a target model, such as a speech activity detection model, or any other suitable machine learning model.

[0050] To avoid semantic breaks or computational redundancy, in some embodiments, the electronic device 110 may cut based on a minimum cutting algorithm with dynamic threshold search.

[0051] Specifically, the electronic device 110 can determine a set of cut points associated with voice activity based on voice data. In some embodiments, such a set of cut points may include, but is not limited to, a first cut point, a second cut point, and a third cut point. As an example, such a first cut point may be associated with a silent region of the voice data, such a second cut point may be associated with a region of voice activity in the voice data where the loudness is below a threshold, and such a third cut point may be associated with a preset time interval.

[0052] In some embodiments, such a set of cutting points can correspond to different priorities. As an example, the first cutting point corresponds to the first priority, the second cutting point corresponds to the second priority, and the third cutting point corresponds to the third priority, with the first priority being greater than the second priority, and the second priority being greater than the third priority.

[0053] Furthermore, the electronic device 110 can segment the speech data based on such a set of segmentation points and the priority of such a set of segmentation points, so as to extract at least one speech segment from the speech data.

[0054] To reduce operating costs, in some embodiments, the electronic device 110 may also utilize a target model to extract human voice segments from the speech data. As an example, such human voice segments may be determined based on a confidence threshold (e.g., 0.3) that the confidence levels at the start and end of speaking are greater than or equal to the confidence threshold. As an example, such a target model may be a speech activity detection model, but it could also be any other suitable machine learning model.

[0055] In some embodiments, the electronic device 110 may also use the target model to output timestamp information corresponding to each voice segment, such as the start and end times of each voice segment. It is understood that this disclosure is not intended to limit the number of voice segments.

[0056] In some embodiments, such a target model may be, for example, a multi-label classification model, and may also be a model trained based on the permutation-invariant training (PIT) principle.

[0057] For example, suppose a voice clip X corresponds to a reference label y, and the dimension of the reference label y is K. max ×T represents a maximum of K max The active states of a speaker across T time frames are predicted by the target model. The results of the target model can be expressed based on the following formula:

[0058]

[0059] For example, the loss function of the target model can be expressed based on the following formula:

[0060]

[0061] Where, L BCE Let perm(y) represent the binary cross-entropy loss function, and let perm(y) represent the K-value of y. max Possible results after arranging the dimensions.

[0062] Furthermore, the electronic device 110 can determine a corresponding set of segmentation points based on the voice segment, and determine at least one voice segment associated with the voice activity based on the priority of the set of segmentation points. Additionally, such at least one voice segment can also be determined, for example, based on a merging operation.

[0063] In some embodiments, at least one such audio segment may be a segment that satisfies a maximum preset duration constraint (e.g., 30 seconds).

[0064] In this way, embodiments of the present disclosure can segment speech data or human voice segments while ensuring semantic coherence, so as to determine at least one speech segment by identifying the optimal segmentation point.

[0065] In box 230, electronic device 110 provides the model with audio features of at least one speech segment to generate first text content, wherein the model is trained based on training speech associated with a preset loudness range. In some embodiments, such a model includes a recognition unit and a fine-tuning unit; as an example, such a model could be the result of the fine-tuning unit fine-tuning the recognition unit.

[0066] In some embodiments, the electronic device 110 can extract audio features of at least one speech segment. Specifically, the electronic device 110 can perform framing processing on at least one speech segment (e.g., each frame can be 20 milliseconds, and the step size can be 10 milliseconds). Furthermore, the electronic device 110 can extract frequency domain energy through Fast Fourier Transform.

[0067] In some embodiments, the electronic device 110 can convert linear frequencies into Mel scales (simulating human hearing perception) using a Mel filter bank (e.g., 80 filters). In some embodiments, the electronic device 110 can take the logarithmic energy value to generate an 80-dimensional log-Mel spectrogram. Furthermore, the electronic device 110 can perform global mean-variance normalization on the log-Mel spectrogram to improve model robustness.

[0068] In this way, the electronic device 110 can obtain the audio features corresponding to at least one speech segment (e.g., a log-Mel spectrogram normalized to global mean and variance).

[0069] Furthermore, the electronic device 110 can provide the acquired audio features to the model's encoding module. Specifically, the electronic device 110 can process these features via the encoding module to obtain intermediate features of dimensions that can be processed by the subsequent decoding module.

[0070] As an example, the model may include an encoding module consisting of a multi-layer encoder stack (e.g., 12 layers), with each layer containing a multi-head self-attention and feedforward network, and all linear layers having fused fine-tuning units.

[0071] In some embodiments, the electronic device 110 can acquire intermediate features and historical generated text from the encoder output (e.g., mapped to the model dimension via word embeddings) and provide them to the model's decoding module to generate corresponding first text content. In some embodiments, such first text content can be generated, for example, based on an autoregressive approach.

[0072] As an example, the model may include a decoding module consisting of a multi-layered decoder stack (e.g., 12 layers), with each layer containing self-attention and encoder-decoder cross-attention. Self-attention focuses on the context of the generated first text content, cross-attention focuses on audio features with the first text content, and all linear layers have fused fine-tuning units.

[0073] In this way, the electronic device 110 can acquire the first text content corresponding to the audio features of at least one speech segment.

[0074] In box 240, electronic device 110 adjusts the text representation of the first text content based on at least one text representation constraint to determine the recognition result corresponding to the speech data. For ease of description, the following description uses English text as an example of the first text content.

[0075] In some embodiments, such at least one text expression constraint may include a text deduplication constraint. Specifically, in response to the first text content including a plurality of consecutive and identical text elements, the electronic device 110 may, based on the text deduplication constraint, delete some text elements from the plurality of text elements to adjust the text expression of the first text content.

[0076] As an example, such multiple text elements can indicate character elements, such as letters. Exemplarily, in response to the word "leeeeeength" in the first text content containing multiple consecutive and identical letters "e", electronic device 110 can delete portions of the multiple letters "e" based on text deduplication constraints to obtain the corrected word "length". Alternatively, in response to an abnormal word length in the first text content (e.g., the number of letters in the word exceeds a threshold), electronic device 110 can perform deletion processing on the word.

[0077] As another example, such multiple text elements can indicate field elements, such as words, phrases, sentences, paragraphs, etc. For instance, in response to consecutive repetitions of words in the first text content, such as "thethe the", the electronic device 110 can, based on text deduplication constraints, delete portions of multiple repeating words, for example, retaining only one "the".

[0078] In some embodiments, such at least one textual expression constraint may also indicate a semantic constraint. Specifically, in response to the first text content indicating first semantic information and second semantic information, the electronic device 110 may adjust the first text content based on the semantic constraint so that the first text content indicates first speech information but not second semantic information.

[0079] As an example, assuming that the first text content can indicate both first semantic information and second semantic information, the electronic device 110 can adjust it to first speech information that better matches the semantic expression of the first text content based on semantic constraints.

[0080] As another example, suppose the text content “×××” in the first text content can indicate both first semantic information and second semantic information. Then, the electronic device 110 can adjust it to first speech information that is more consistent with the semantic expression of the first text content based on semantic constraints.

[0081] In the above process, the electronic device 110 may, for example, use a pre-trained model to perform semantic disambiguation on the first text content in order to improve the semantic consistency and readability of the first text content.

[0082] In this way, the electronic device 110 can adjust the text expression of the first text content based on at least one text expression constraint to determine the recognition result corresponding to the speech data, thereby reducing interference factors in the generated text content and obtaining a more accurate and reliable recognition result.

[0083] Furthermore, the embodiments of this disclosure can improve the accuracy of speech recognition and the quality of corresponding text generation for people with articulation disorders, thereby enhancing their communication efficiency, increasing their enthusiasm for interaction and participation, and helping them to better realize their social value.

[0084] In some embodiments, the model described above may include a recognition unit and a fine-tuning unit, and such a model can be trained by adjusting the parameters of the fine-tuning unit based on training speech. For ease of understanding, the adjustment process of the fine-tuning unit will be described below.

[0085] In some embodiments, such training speech may include multiple datasets, which may correspond to multiple content sources. Additionally, such training speech may also include speech acquired in real-time by the electronic device 110 based on recording equipment.

[0086] Furthermore, the electronic device 110 can provide training speech to the pre-trained recognition unit to generate second text content.

[0087] In some embodiments, such training speech may be speech obtained through loudness preprocessing and / or enhancement processing. Specifically, electronic device 110 may perform loudness preprocessing on the training speech to adjust its loudness to a preset loudness range. As an example, electronic device 110 may perform loudness preprocessing on the training speech based on a loudness standard. Electronic device 110 may also perform enhancement processing on the training speech. As an example, electronic device 110 may perform temporal stretching on the training speech or a portion of the training speech (e.g., 50% of the speech).

[0088] In some embodiments, electronic device 110 may provide training speech to a recognition unit with a fine-tuning unit to generate third text content.

[0089] In some embodiments, the electronic device 110 may add the low-rank matrix parameters of the fine-tuning unit to the corresponding weights of the recognition unit to apply the fine-tuning unit to the recognition unit. Exemplarily, this process may be represented based on the following formula:

[0090] h t =TransformerDecoder(y <t ,X;W0+ΔW)

[0091] ΔW=AΛB

[0092] Among them, y <t denoted as the text sequence generated at time step t; X represents the extracted features of the input speech; W0 represents the weight matrix of the pre-trained recognition unit; ΔW represents the adaptive low-rank update learned during fine-tuning. and represents a low-rank matrix, whose rank can be adaptively allocated among layers. represents a diagonal matrix, containing r singular values, and r << min(d1, d2). Each element Λ i can correspond to the scaling factor for the low-rank update of the i-th layer.

[0093] In some embodiments, the electronic device 110 can adjust the parameters of the fine-tuning unit based on the comparison between the second text content and the third text content.

[0094] In some embodiments, the training speech involved above can be represented as a first group of training speech. In some embodiments, such a first group of training speech can include a first training speech, and such a first training speech can be, for example, a training speech within a predetermined domain (e.g., in-domain data). Such a first group of training speech can also include a second training speech, and such a second training speech can be, for example, a training speech outside a predetermined domain (e.g., out-of-domain data). As an example, such a first training speech and a second training speech can be determined based on multiple data sets included in the training speech.

[0095] In some embodiments, the electronic device 110 can obtain the annotated text content corresponding to the training speech. Exemplarily, such an annotated text content can be determined based on humans, for example.

[0096] In some embodiments, the electronic device 110 can determine a first text similarity between the second text content and the annotated text content based on the comparison between the second text content and the annotated text content. The electronic device 110 can also determine a second text similarity between the third text content and the annotated text content based on the comparison between the third text content and the annotated text content.

[0097] Furthermore, the electronic device 110 can determine a second group of training speech from the first group of training speech based on the comparison between the first text similarity and the second text similarity.

[0098] In some embodiments, the electronic device 110 can determine whether the first text similarity and the second text similarity corresponding to the first training speech are greater than or equal to a first threshold (e.g., threshold a). In some embodiments, in response to the first text similarity being greater than or equal to the first threshold and the second text similarity being greater than or equal to the first threshold, the electronic device 110 can add the first training speech to the second group of training speech.

[0099] It can be understood that in response to the first training speech having a text similarity greater than or equal to the first threshold after passing through the recognition unit with the fine-tuning unit applied and the recognition unit without the fine-tuning unit applied, the electronic device 110 can add the first training speech to the second group of training speech.

[0100] In this way, simple and difficult samples corresponding to the first training speech can be added to the second set of training speech, thereby removing noisy samples while retaining potentially useful samples from the first training speech.

[0101] To improve the generalization ability and accuracy of the model, in some embodiments, the electronic device 110 can determine whether the first text similarity corresponding to the second training speech and the second text similarity are greater than or equal to a second threshold (e.g., threshold b). The electronic device 110 can also determine whether the second text similarity is greater than the first text similarity.

[0102] Furthermore, in response to a first text similarity greater than or equal to a second threshold and a second text similarity greater than or equal to the second threshold, and the second text similarity being greater than the first text similarity, the electronic device 110 may add the second training speech to the second set of training speech.

[0103] Understandably, in response to the fact that after the second training speech passes through the recognition unit with the applied fine-tuning unit and the recognition unit without the applied fine-tuning unit, the electronic device 110 finds that the text similarity of the second training speech is greater than or equal to the second threshold, and that the recognition effect of the second training speech in the recognition unit with the applied fine-tuning unit is better than the recognition effect corresponding to the recognition unit without the applied fine-tuning unit, the second training speech can be added to the second set of training speech.

[0104] In this way, simple samples corresponding to the second training speech can be added to the second set of training speech, thereby improving the model's generalization ability and accuracy.

[0105] Furthermore, the electronic device 110 can adjust the parameters of the fine-tuning unit based on the second set of training voice obtained above.

[0106] FIG. 3A A schematic diagram of the reasoning process for speech recognition according to some embodiments of the present disclosure is shown.

[0107] In box 301, electronic device 110 can acquire speech data, such speech data may include speech or human voice with speech disorders.

[0108] In box 302, electronic device 110 can adjust the loudness of the acquired speech data to a preset loudness range. As an example, electronic device 110 can perform loudness normalization preprocessing on the acquired speech data based on a loudness standard to adjust the loudness of the speech data to a preset loudness range.

[0109] In box 303, electronic device 110 can perform speech activity detection on speech data to extract “speech segments” associated with speech activity. For example, electronic device 110 can use a speech activity detection model to extract “speech segments” associated with speech activity from speech data.

[0110] In block 304, electronic device 110 can segment the extracted "speech segments" based on a set of segmentation points and a set of priorities corresponding to the segmentation points to obtain at least one speech segment associated with the speech activity. Additionally, electronic device 110 can also merge the extracted "speech segments" to obtain at least one speech segment associated with the speech activity.

[0111] In box 305, electronic device 110 can extract audio features of at least one speech segment.

[0112] In box 306, electronic device 110 can provide the audio feature to the encoding module of the model to obtain intermediate features of dimensions that can be processed by the subsequent decoding module.

[0113] In box 307, electronic device 110 can provide the intermediate features and / or historical generated text to the decoding module of the model to generate the first text content.

[0114] In boxes 308 to 310, the electronic device 110 can adjust the text representation of the first text content based on at least one text representation constraint to determine the recognition result corresponding to the speech data. Such at least one text representation constraint may include text deduplication constraints and semantic constraints.

[0115] FIG. 3B A schematic diagram of a speech recognition training process according to some embodiments of the present disclosure is shown.

[0116] In box 320, electronic device 110 can provide training speech to a pre-trained recognition unit to generate second text content. Such training speech may, for example, be speech that has undergone loudness preprocessing and / or enhancement processing.

[0117] In block 321, electronic device 110 can provide training speech to a recognition unit equipped with fine-tuning units to generate third text content. Additionally, in addition to loudness preprocessing and / or enhancement processing of the training speech, such training speech can also be filtered speech.

[0118] In box 322, electronic device 110 can adjust the parameters of the fine-tuning unit based on a comparison of the second and third text content. Specifically, electronic device 110 can determine a first text similarity based on a comparison of the second text content and the labeled text content; electronic device 110 can determine a second text similarity based on a comparison of the third text content and the labeled text content. Further, electronic device 110 can determine a second set of training speech from the first set of training speech based on a comparison of the first and second text similarities. This second set of speech may include, for example, simple and complex samples from within the domain, and may also include simple samples from outside the domain. Further, electronic device 110 can adjust the parameters of the fine-tuning unit based on the second set of training speech.

[0119] In this way, the embodiments of this disclosure can adjust the loudness of speech data to a preset loudness range, which can improve the accuracy of model recognition compared to speech data without loudness adjustment. Furthermore, the embodiments of this disclosure can adjust the text expression of the text content based on at least one text expression constraint, thereby reducing interference factors in the generated text content and obtaining more accurate and reliable recognition results.

[0120] Example Apparatus and Device

[0121] Embodiments of this disclosure also provide corresponding apparatus for implementing the above methods or processes. FIG. 4 A schematic structural block diagram of an example device 400 for speech recognition according to certain embodiments of the present disclosure is shown. Device 400 may be implemented as or included in electronic device 110. Various modules / components in device 400 may be implemented by hardware, software, firmware, or any combination thereof.

[0122] like FIG. 4 As shown, the device 400 includes a loudness adjustment module 410 configured to adjust the loudness of acquired speech data to a preset loudness range, the speech data being derived from an object associated with speech disorders; a segment extraction module 420 configured to extract at least one speech segment associated with speech activity from the speech data; a content generation module 430 configured to provide an audio feature of at least one speech segment to a model to generate first text content, wherein the model is trained based on training speech associated with a preset loudness range; and a text adjustment module 440 configured to adjust the text expression of the first text content based on at least one text expression constraint to determine a recognition result corresponding to the speech data.

[0123] In some embodiments, the segment extraction module 420 is further configured to determine a set of cut points associated with speech activity based on speech data; and to extract at least one speech segment from the speech data based on the set of cut points.

[0124] In some embodiments, a set of cut points includes at least one of the following: a first cut point associated with a silent region of the speech data; a second cut point associated with a region of speech activity in the speech data where the loudness is below a threshold; and a third cut point associated with a preset time interval.

[0125] In some embodiments, at least one text expression constraint includes a text deduplication constraint, and the text adjustment module 440 is further configured to, in response to the first text content including a plurality of consecutive and identical text elements, delete some text elements from the plurality of text elements based on the text deduplication constraint to adjust the text expression of the first text content.

[0126] In some embodiments, the multiple text elements include character elements and / or field elements.

[0127] In some embodiments, at least one text expression constraint indicates a semantic constraint, and the text adjustment module 440 is further configured to adjust the first text content based on the semantic constraint in response to the first text content indicating first semantic information and second semantic information, such that the first text content indicates first speech information but not second semantic information.

[0128] In some embodiments, the model includes a recognition unit and a fine-tuning unit, wherein the model is trained by adjusting the parameters of the fine-tuning unit based on training speech, wherein the fine-tuning unit is adjusted based on the following process: providing training speech to a pre-trained recognition unit to generate second text content; providing training speech to a recognition unit with the fine-tuning unit applied to generate third text content; and adjusting the parameters of the fine-tuning unit based on a comparison of the second text content and the third text content.

[0129] In some embodiments, the training speech includes a first set of training speech, and adjusting the parameters of the fine-tuning unit based on a comparison of second and third text content includes: determining a first text similarity between the second text content and the labeled text content based on a comparison of the second text content and the labeled text content; determining a second text similarity between the second text content and the labeled text content based on a comparison of the third text content and the labeled text content; determining a second set of training speech from the first set of training speech based on a comparison of the first and second text similarities; and adjusting the parameters of the fine-tuning unit based on the second set of training speech.

[0130] In some embodiments, the first set of training speech includes a first training speech. Determining the second set of training speech from the first set of training speech based on a comparison of a first text similarity and a second text similarity includes: determining whether the first text similarity and the second text similarity corresponding to the first training speech are greater than or equal to a first threshold; and adding the first training speech to the second set of training speech in response to the first text similarity being greater than or equal to the first threshold and the second text similarity being greater than or equal to the first threshold.

[0131] In some embodiments, the first set of training speech includes a second set of training speech. Determining the second set of training speech from the first set of training speech based on a comparison of a first text similarity and a second text similarity includes: determining whether the first text similarity and the second text similarity corresponding to the second set of training speech are greater than or equal to a second threshold; determining whether the second text similarity is greater than the first text similarity; and in response to the first text similarity being greater than or equal to the second threshold and the second text similarity being greater than or equal to the second threshold, and the second text similarity being greater than the first text similarity, adding the second set of training speech to the second set of training speech.

[0132] The modules included in device 400 can be implemented in various ways, including software, hardware, firmware, or any combination thereof. In some embodiments, one or more units can be implemented using software and / or firmware, such as machine-executable instructions stored on a storage medium. In addition to or as an alternative to machine-executable instructions, some or all of the modules in device 400 can be implemented at least partially by one or more hardware logic components. By way of example, and not limitation, exemplary types of hardware logic components that can be used include field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip (SoCs), complex programmable logic devices (CPLDs), and so on.

[0133] like FIG. 5 As shown, electronic device 500 is in the form of a general-purpose electronic device. Components of electronic device 500 may include, but are not limited to, one or more processing units or processors 510, memory 520, storage devices 530, one or more communication units 540, one or more input devices 550, and one or more output devices 560. Processor 510 may be a physical or virtual processor and is capable of performing various processes according to programs stored in memory 520. In a multiprocessor system, multiple processors execute computer-executable instructions in parallel to improve the parallel processing capability of electronic device 500. FIG. 5 The electronic device 500 shown can be used to achieve FIG. 1 Electronic devices 110.

[0134] Electronic device 500 typically includes multiple computer storage media. Such media can be any accessible media that is accessible to electronic device 500, including but not limited to volatile and non-volatile media, removable and non-removable media. Memory 520 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. Storage device 530 can be removable or non-removable media and can include machine-readable media, such as flash drives, disks, or any other media that can be used to store information and / or data and can be accessed within electronic device 500.

[0135] Electronic device 500 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not explicitly stated... FIG. 5 As shown, disk drives for reading from or writing to removable, non-volatile disks (e.g., "floppy disks") and optical disk drives for reading from or writing to removable, non-volatile optical disks can be provided. In these cases, each drive can be connected to a bus (not shown) via one or more data media interfaces. Memory 520 may include computer program product 525 having one or more program modules configured to perform various methods or actions of various embodiments of this disclosure.

[0136] Communication unit 540 enables communication with other electronic devices via a communication medium. Additionally, the functionality of components of electronic device 500 can be implemented using a single computing cluster or multiple computing machines capable of communicating via communication connections. Therefore, electronic device 500 can operate in a networked environment using logical connections to one or more other servers, network personal computers (PCs), or another network node.

[0137] Input device 550 can be one or more input devices, such as a mouse, keyboard, trackball, etc. Output device 560 can be one or more output devices, such as a monitor, speaker, printer, etc. Electronic device 500 can also communicate with one or more external devices (not shown) via communication unit 540 as needed. These external devices include storage devices, display devices, etc., and can communicate with one or more devices that enable user interaction with electronic device 500, or with any device that enables electronic device 500 to communicate with one or more other electronic devices (e.g., network card, modem, etc.). Such communication can be performed via input / output (I / O) interface (not shown).

[0138] According to an exemplary implementation of this disclosure, a computer-readable storage medium is provided that stores computer-executable instructions thereon, wherein the computer-executable instructions are executed by a processor to implement the methods described above. According to an exemplary implementation of this disclosure, a computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, which are executed by a processor to implement the methods described above.

[0139] Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatuses, devices, and computer program products implemented according to this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0140] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processing unit of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.

[0141] Computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions that execute on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.

[0142] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0143] Various implementations of this disclosure have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is chosen to best explain the principles, practical applications, or improvements to technology in the market, or to enable others skilled in the art to understand the various implementations disclosed herein.

Claims

1. A speech recognition method, comprising: The loudness of the acquired speech data is adjusted to a preset loudness range, and the speech data comes from objects related to speech disorders; Extract at least one speech segment associated with the speech activity from the speech data; The model is provided with audio features of at least one speech segment to generate first text content, wherein the model is trained based on training speech associated with the preset loudness range; as well as Based on at least one text expression constraint, the text expression of the first text content is adjusted to determine the recognition result corresponding to the speech data.

2. The method of claim 1, wherein extracting at least one speech segment associated with speech activity from the speech data comprises: Based on the voice data, a set of cut points associated with the voice activity are determined; as well as Based on the set of cutting points, at least one speech segment is extracted from the speech data.

3. The method of claim 2, wherein the set of cutting points includes at least one of the following: The first cutting point is associated with the silent region of the voice data; The second cut point is associated with the speech activity region in the speech data where the loudness is below a threshold. The third cutting point is associated with a preset time interval.

4. The method according to claim 1, wherein the at least one text expression constraint includes a text deduplication constraint, and adjusting the text expression of the first text content based on the at least one text expression constraint includes: In response to the first text content including multiple consecutive and identical text elements, based on the text deduplication constraint, some text elements among the multiple text elements are deleted to adjust the text expression of the first text content.

5. The method of claim 4, wherein the plurality of text elements includes character elements and / or field elements.

6. The method of claim 1, wherein the at least one textual expression constraint indicates a semantic constraint, and adjusting the textual expression of the first text content based on the at least one textual expression constraint comprises: In response to the first text content indicating first semantic information and second semantic information, the first text content is adjusted based on the semantic constraints so that the first text content indicates the first speech information but not the second semantic information.

7. The method of claim 1, wherein the model comprises a recognition unit and a fine-tuning unit, wherein the model is trained by adjusting the parameters of the fine-tuning unit based on the training speech, wherein the fine-tuning unit is adjusted based on the following process: The training speech is provided to the pre-trained recognition unit to generate second text content; The training speech is provided to the recognition unit equipped with the fine-tuning unit to generate third text content; and Based on the comparison between the second text content and the third text content, the parameters of the fine-tuning unit are adjusted.

8. The method of claim 7, wherein the training speech includes a first set of training speech, and adjusting the parameters of the fine-tuning unit based on a comparison of the second text content and the third text content includes: Based on the comparison between the second text content and the labeled text content, the first text similarity between the second text content and the labeled text content is determined; Based on the comparison between the third text content and the labeled text content, the second text similarity between the second text content and the labeled text content is determined; Based on the comparison of the first text similarity and the second text similarity, a second group of training speech is determined from the first group of training speech; as well as Based on the second set of training voices, the parameters of the fine-tuning unit are adjusted.

9. The method according to claim 8, wherein the first set of training speech includes a first set of training speech, and determining the second set of training speech from the first set of training speech based on a comparison of the first text similarity and the second text similarity includes: Determine whether the first text similarity and the second text similarity corresponding to the first training speech are greater than or equal to a first threshold; In response to the first text similarity being greater than or equal to the first threshold and the second text similarity being greater than or equal to the first threshold, the first training speech is added to the second group of training speech.

10. The method of claim 8, wherein the first set of training speech includes a second set of training speech, and determining the second set of training speech from the first set of training speech based on a comparison of the first text similarity and the second text similarity comprises: Determine whether the similarity between the first text and the second text corresponding to the second training speech is greater than or equal to a second threshold; Determine whether the similarity of the second text is greater than the similarity of the first text; as well as In response to the first text similarity being greater than or equal to the second threshold and the second text similarity being greater than or equal to the second threshold, and the second text similarity being greater than the first text similarity, the second training speech is added to the second group of training speech.

11. An apparatus for speech recognition, comprising: A loudness adjustment module is configured to adjust the loudness of acquired speech data to a preset loudness range, wherein the speech data comes from an object related to speech disorders; The segment extraction module is configured to extract at least one speech segment associated with speech activity from the speech data; A content generation module is configured to provide the model with audio features of the at least one speech segment to generate first text content, wherein the model is trained based on training speech associated with the preset loudness range; as well as The text adjustment module is configured to adjust the text expression of the first text content based on at least one text expression constraint in order to determine the recognition result corresponding to the speech data.

12. An electronic device, comprising: At least one processor; as well as At least one memory coupled to the at least one processor and storing instructions for execution by the at least one processor, the instructions causing the electronic device to perform the method according to any one of claims 1 to 10 when executed by the at least one processor.

13. A computer-readable storage medium having stored thereon computer-executable instructions that can be executed by a processor to implement the method according to any one of claims 1 to 10.