Training methods, devices, equipment, and storage media for speech recognition models

By filtering and retraining the speech recognition model to identify incorrectly recognized audio segments, the problem of insufficient accuracy in recognizing consecutive identical text units is solved, thereby improving the accuracy of speech recognition and user satisfaction.

CN115662393BActive Publication Date: 2026-05-26SOUNDAI TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SOUNDAI TECH CO LTD
Filing Date
2022-10-11
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

Existing speech recognition technology is not accurate enough when recognizing consecutive identical text units, leading to a decrease in user satisfaction.

Method used

By filtering out candidate audio samples that were not correctly identified in the sample audio set, audio segments aligned with consecutive identical text units are extracted, and the initial speech recognition model is retrained to form the target speech recognition model.

Benefits of technology

This improved the model's accuracy in recognizing consecutive identical text units, thus enhancing the quality and effectiveness of speech recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115662393B_ABST
    Figure CN115662393B_ABST
Patent Text Reader

Abstract

This application discloses a training method, apparatus, device, and storage medium for a speech recognition model, belonging to the field of artificial intelligence. The method includes: acquiring a sample audio set, the sample audio set including multiple sample audios; filtering candidate sample audios from the sample audio set based on an initial speech recognition model; extracting audio segments from the candidate sample audios; wherein the audio segments include audios aligned with consecutive identical text units from the candidate sample audios; and the initial speech recognition model, when performing speech recognition on the candidate sample audios, failed to correctly recognize the consecutive identical text units; and retraining the initial speech recognition model based on the audio segments to obtain a target speech recognition model. This application can improve speech recognition quality, particularly improving the accuracy of recognizing consecutive identical text units.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence, and in particular to a method, apparatus, device, and storage medium for training a speech recognition model. Background Technology

[0002] Speech recognition is a significant breakthrough in the development of artificial intelligence. Broadly speaking, speech recognition studies speech and aims to enable interaction between humans and machines based on natural language. More specifically, speech recognition is a technology that allows machines to convert speech into text or commands through recognition and understanding.

[0003] Speech recognition technology is now widely used in fields such as industry, home appliances, communications, automobiles, electronics, healthcare, and home services. Among these applications, the accuracy of speech recognition is crucial, as higher accuracy leads to higher user satisfaction. Therefore, how to perform speech recognition accurately to improve its effectiveness has become a hot research topic in this field. Summary of the Invention

[0004] This application provides a method, apparatus, device, and storage medium for training a speech recognition model, which can improve speech recognition quality, especially the accuracy of recognizing consecutive identical text units. The technical solution is as follows:

[0005] On the one hand, a method for training a speech recognition model is provided, the method comprising:

[0006] Obtain a sample audio set, which includes multiple sample audio files;

[0007] Based on the initial speech recognition model, candidate audio samples are selected from the set of sample audio samples.

[0008] Audio segments are extracted from the candidate sample audio; wherein, the audio segments include audio that is aligned with consecutive identical text units in the candidate sample audio; and, the initial speech recognition model does not correctly recognize the consecutive identical text units when performing speech recognition on the candidate sample audio.

[0009] The initial speech recognition model is retrained based on the audio segment to obtain the target speech recognition model.

[0010] In some possible implementations, the sample audio set also includes annotation text aligned with the sample audio; the annotation text includes at least one set of consecutive identical text units;

[0011] The step of filtering candidate audio samples from the sample audio set based on the initial speech recognition model includes: performing speech recognition on the sample audio based on the initial speech recognition model to obtain the predicted text of the sample audio;

[0012] Based on the labeled and predicted text of the sample audio, the candidate sample audio is filtered in the sample audio set.

[0013] In some possible implementations, the audio segment includes adjacent first audio, second audio, and third audio;

[0014] Wherein, the first audio is located before the second audio in time sequence; the third audio is located after the second audio in time sequence; the second audio is the audio in the candidate sample audio that is aligned with the consecutive identical text unit;

[0015] The step of retraining the initial speech recognition model based on the audio segment to obtain the target speech recognition model includes:

[0016] Obtain the labeled text of the audio segment; wherein the labeled text of the audio segment includes the consecutive identical text units and preset tags;

[0017] Align the annotation text of the audio segment with the audio segment itself;

[0018] The preset label is aligned with the first audio and the third audio, and the consecutive identical text units are aligned with the second audio.

[0019] The initial speech recognition model is retrained based on the audio segment and the labeled text aligned with the audio segment to obtain the target speech recognition model.

[0020] In some possible implementations, extracting audio segments from the candidate sample audio includes:

[0021] Obtain the predicted text of the candidate sample audio; wherein the predicted text includes multiple text units and the start and end times of the multiple text units in the candidate sample audio;

[0022] The start and end times of the second audio are determined based on the start and end times of each text unit in the predicted text.

[0023] Wherein, the second audio is the audio in the candidate sample audio that is aligned with the consecutive identical text units;

[0024] Based on the start and end times of the second audio, extract the audio segment from the candidate sample audio;

[0025] The audio segments include adjacent first audio, second audio, and third audio;

[0026] The first audio is time-ordered before the second audio; the third audio is time-ordered after the second audio.

[0027] In some possible implementations, the method further includes:

[0028] Obtain the input audio to be recognized;

[0029] Feature extraction is performed on the input audio to obtain the filter bank features of the input audio and the identity authentication vector corresponding to the speaker;

[0030] The discrete cosine transform of the filter bank features is used to obtain the Mel-frequency cepstral coefficients of the input audio.

[0031] The identity authentication vector and the Mel cepstral coefficients are fused to obtain the audio fusion features of the input audio;

[0032] The audio fusion features are input into the target speech recognition model, and the speech recognition result corresponding to the input audio is output.

[0033] In some possible implementations, the method further includes:

[0034] In response to the speech recognition result including a first number of consecutive digits, the terminal number corresponding to the input audio is determined, and a first selection menu is displayed; wherein, the first selection menu includes multiple options, and one of the options is used to indicate an installed application;

[0035] In response to the selection of a target application based on the plurality of options, communication is performed between the target application and the terminal indicated by the terminal number.

[0036] In some possible implementations, the step of responding to a selection operation of a target application based on the plurality of options, and communicating between the target application and the terminal indicated by the terminal number, includes:

[0037] In response to the target application being an SMS application or a calling application, a second selection menu is displayed; wherein the second selection menu includes multiple options, one of which is used to indicate a user identification card; in response to the selection operation of the target user identification card based on the multiple options, based on the target user identification card, communication is performed with the terminal indicated by the terminal number using the communication method corresponding to the target application;

[0038] Alternatively, in response to the target application being a social application, a third selection menu is displayed; wherein the third selection menu includes multiple options, one of which indicates a communication method of the social application; in response to the selection of a target communication method based on the multiple options, communication is performed with the terminal indicated by the terminal number based on the target communication method;

[0039] The communication methods of the social application include sending application messages, voice calls, or video calls.

[0040] In some possible implementations, the method further includes:

[0041] In response to the voice recognition results, which include the application name of the social application and the username of the target contact in the address book, an information interaction page is displayed through the social application in a top-level display manner.

[0042] The address book corresponds to the social application; the information interaction page is the page for interacting with the target contact.

[0043] In some possible implementations, the method further includes:

[0044] In response to the speech recognition result including a second number of consecutive characters, and the second number of consecutive characters matching a target logistics order number template among multiple logistics order number templates, the system determines the logistics order number corresponding to the input audio and obtains the logistics information of the logistics order indicated by the logistics order number.

[0045] On the other hand, a training device for a speech recognition model is provided, the device comprising:

[0046] The acquisition module is configured to acquire a sample audio set, which includes multiple sample audio files.

[0047] The filtering module is configured to filter candidate audio samples from the set of sample audio samples based on an initial speech recognition model.

[0048] The processing module is configured to extract audio segments from the candidate sample audio.

[0049] The audio segment includes audio segments in the candidate sample audio that are aligned with consecutive identical text units; and the initial speech recognition model fails to correctly recognize the consecutive identical text units when performing speech recognition on the candidate sample audio.

[0050] The training module is configured to retrain the initial speech recognition model based on the audio segment to obtain the target speech recognition model.

[0051] In some possible implementations, the sample audio set also includes annotation text aligned with the sample audio; the annotation text includes at least one set of consecutive identical text units;

[0052] The filtering module is configured to perform speech recognition on the sample audio based on the initial speech recognition model to obtain the predicted text of the sample audio; and to filter the candidate sample audio in the sample audio set according to the labeled text and predicted text of the sample audio.

[0053] In some possible implementations, the audio segment includes adjacent first audio, second audio, and third audio; wherein the first audio is sequentially preceding the second audio; the third audio is sequentially following the second audio; and the second audio is the audio in the candidate sample audio that is aligned with the consecutive identical text units.

[0054] The training module is configured to acquire the labeled text of the audio segment; wherein the labeled text of the audio segment includes consecutive identical text units and preset labels; align the labeled text of the audio segment with the audio segment; wherein the preset labels are aligned with the first audio and the third audio, and the consecutive identical text units are aligned with the second audio; and retrain the initial speech recognition model based on the audio segment and the labeled text aligned with the audio segment to obtain the target speech recognition model.

[0055] In some possible implementations, the processing module is configured to acquire the predicted text of the candidate sample audio; wherein the predicted text includes multiple text units and the start and end times of the multiple text units in the candidate sample audio; determine the start and end times of the second audio based on the start and end times of each text unit in the predicted text; and extract the audio segment from the candidate sample audio based on the start and end times of the second audio.

[0056] In some possible implementations, the process of using the target speech recognition model includes:

[0057] Obtain the input audio to be recognized;

[0058] Feature extraction is performed on the input audio to obtain the filter bank features of the input audio and the identity authentication vector corresponding to the speaker;

[0059] The discrete cosine transform of the filter bank features is used to obtain the Mel-frequency cepstral coefficients of the input audio.

[0060] The identity authentication vector and the Mel cepstral coefficients are fused to obtain the audio fusion features of the input audio;

[0061] The audio fusion features are input into the target speech recognition model, and the speech recognition result corresponding to the input audio is output.

[0062] In some possible implementations, the process of using the target speech recognition model further includes:

[0063] In response to the speech recognition result including a first number of consecutive digits, the terminal number corresponding to the input audio is determined, and a first selection menu is displayed; wherein, the first selection menu includes multiple options, and one of the options is used to indicate an installed application;

[0064] In response to the selection of a target application based on the plurality of options, communication is performed between the target application and the terminal indicated by the terminal number.

[0065] In some possible implementations, the process of using the target speech recognition model further includes:

[0066] In response to the target application being an SMS application or a calling application, a second selection menu is displayed; wherein the second selection menu includes multiple options, one of which is used to indicate a user identification card; in response to the selection operation of the target user identification card based on the multiple options, based on the target user identification card, communication is performed with the terminal indicated by the terminal number using the communication method corresponding to the target application;

[0067] Alternatively, in response to the target application being a social application, a third selection menu is displayed; wherein the third selection menu includes multiple options, one of which indicates a communication method of the social application; in response to the selection of a target communication method based on the multiple options, communication is performed with the terminal indicated by the terminal number based on the target communication method;

[0068] The communication methods of the social application include sending application messages, voice calls, or video calls.

[0069] In some possible implementations, the process of using the target speech recognition model further includes:

[0070] In response to the voice recognition results, which include the application name of the social application and the username of the target contact in the address book, an information interaction page is displayed through the social application in a top-level display manner.

[0071] The address book corresponds to the social application; the information interaction page is the page for interacting with the target contact.

[0072] In some possible implementations, the process of using the target speech recognition model further includes:

[0073] In response to the speech recognition result including a second number of consecutive characters, and the second number of consecutive characters matching a target logistics order number template among multiple logistics order number templates, the system determines the logistics order number corresponding to the input audio and obtains the logistics information of the logistics order indicated by the logistics order number.

[0074] On the other hand, a computer device is provided, the device including a processor and a memory, the memory storing at least one piece of program code, the at least one piece of program code being loaded and executed by the processor to implement the above-described training method for the speech recognition model.

[0075] On the other hand, a computer-readable storage medium is provided, wherein at least one piece of program code is stored in the storage medium, the at least one piece of program code being loaded and executed by a processor to implement the above-described training method for the speech recognition model.

[0076] On the other hand, a computer program product or computer program is provided, which includes computer program code stored in a computer-readable storage medium. A processor of a computer device reads the computer program code from the computer-readable storage medium and executes the computer program code, causing the computer device to perform the above-described training method for the speech recognition model.

[0077] The speech recognition model training scheme provided in this application first filters candidate audio samples that fail to correctly recognize consecutive identical text units from a sample audio set based on an initial speech recognition model. Then, audio segments are extracted from these candidate audio samples; the extracted audio segments include audio aligned with consecutive identical text units from the candidate audio samples. Next, this application retrains the initial speech recognition model based on the extracted audio segments to obtain a target speech recognition model. Because the model relearns the audio features of the incorrectly recognized consecutive identical text units, it improves the model's recognition accuracy on consecutive identical text units. In other words, the newly trained target speech recognition model classifies the audio features of consecutive identical text units more accurately, thus improving the recognition accuracy of consecutive identical text units, resulting in better speech recognition performance and improved speech recognition quality. Attached Figure Description

[0078] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0079] Figure 1 This is a schematic diagram of the implementation environment involved in a speech recognition model training method provided in an embodiment of this application;

[0080] Figure 2 This is a flowchart illustrating a training method for a speech recognition model provided in an embodiment of this application;

[0081] Figure 3 This is a flowchart illustrating a training method for a speech recognition model provided in an embodiment of this application;

[0082] Figure 4 This is a schematic diagram of the structure of a training device for a speech recognition model provided in an embodiment of this application;

[0083] Figure 5 This is a schematic diagram of the structure of a computer device provided in an embodiment of this application;

[0084] Figure 6 This is a schematic diagram of the structure of another computer device provided in an embodiment of this application. Detailed Implementation

[0085] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.

[0086] In this application, the terms "first," "second," etc., are used to distinguish identical or similar items that have essentially the same function. It should be understood that there is no logical or temporal dependency between "first," "second," and "nth," nor does it limit the quantity or execution order. It should also be understood that although the following description uses the terms "first," "second," etc., to describe various elements, these elements should not be limited by the terms.

[0087] These terms are simply used to distinguish one element from another. For example, without departing from the various examples, the first element can be referred to as the second element, and similarly, the second element can be referred to as the first element. Both the first and second elements can be elements, and in some cases, they can be separate and distinct elements.

[0088] "At least one" refers to one or more elements. For example, at least one element can be one element, two elements, three elements, or any integer number of elements greater than or equal to one. "Multiple" refers to two or more elements. For example, multiple elements can be two elements, three elements, or any integer number of elements greater than or equal to two.

[0089] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, data stored, data displayed, etc.) and signals involved in this application are all authorized by the user or fully authorized by all parties, and the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions.

[0090] Figure 1 This is a schematic diagram of the implementation environment involved in a speech recognition model training method provided in an embodiment of this application.

[0091] See Figure 1 The implementation environment includes: model training device 101 and speech recognition device 102.

[0092] The model training device 101 is used to train the speech recognition model, that is, to execute the speech recognition model training method provided in the embodiments of this application; the speech recognition device 102 is used to perform speech recognition based on the trained target speech recognition model, that is, to complete speech recognition based on the target speech recognition model.

[0093] The model training device 101 and the speech recognition device 102 are computer devices with machine learning capabilities. In some possible implementations, the model training device 101 and the speech recognition device 102 can be the same device, or they can be different devices. For example, when the model training device 101 and the speech recognition device 102 are different devices, the model training device 101 can be a fixed computer device such as a personal computer or server, and the speech recognition device 102 can be a mobile computer device such as a tablet computer or smartphone; this application does not impose any limitations on this. Conversely, when the model training device 101 and the speech recognition device 102 are the same device, they can be a personal computer or a server; this application also does not impose any limitations on this.

[0094] The following describes the application scenarios of the speech recognition model training method provided in the embodiments of this application.

[0095] The training method for the speech recognition model provided in this application can improve the model's accuracy in recognizing consecutive identical text units, thereby improving the speech recognition effect. The text unit can be numbers, letters, characters, or a combination of at least two of these, and this application does not impose any limitations on this. Correspondingly, consecutive identical text units can be consecutive identical numbers, consecutive identical letters, consecutive identical characters, etc.

[0096] For example, embodiments of this application can be applied to scenarios involving terminal number recognition or order number recognition. During the recognition process of telephone numbers or order numbers, some consecutive identical digits are easily missed, such as "11" being recognized as "1" and "55" as "5". This is because consecutive identical digits are read quickly and have high feature similarity, making them easy to miss.

[0097] To address this issue, this embodiment first filters candidate audio samples that fail to correctly identify consecutive identical text units from the sample audio set based on the initial speech recognition model. Then, audio segments are extracted from these candidate audio samples; these extracted audio segments include those aligned with consecutive identical text units. Next, this embodiment re-aligns the extracted audio segments and retrains the initial speech recognition model using the re-aligned audio segments to obtain the target speech recognition model. Because the model relearns the audio features of the incorrectly identified consecutive identical text units, its recognition accuracy on consecutive identical text units is improved. In other words, the newly trained target speech recognition model classifies the audio features of consecutive identical text units more accurately, improving the recognition accuracy of consecutive identical text units, resulting in better speech recognition performance and improved speech recognition quality.

[0098] Figure 2 This is a flowchart illustrating a training method for a speech recognition model provided in an embodiment of this application. See also... Figure 2 The method flow provided in this application embodiment includes:

[0099] 201. The model training device acquires a sample audio set; wherein the sample audio set includes multiple sample audios and labeled text aligned with each sample audio; the labeled text includes at least one set of consecutive identical text units.

[0100] The embodiments of this application first require the construction of a sample audio set. Exemplarily, the sample audio is an audio file with a lossless compressed audio format, such as a WAV format audio file, but this application does not impose any limitations on this.

[0101] In some possible implementation manners, sample audio can be crawled from the network through web crawler technology to form a sample audio set; alternatively, sample audio can also be produced through audio recording operations in a quiet environment to form a sample audio set; the present application does not limit this here.

[0102] In an embodiment of the present application, the sample audio in the sample audio set is audio annotated with at least one group of consecutive identical text units. Taking text units as numbers for example, each annotated text includes at least one group of consecutive identical numbers. That is, the sample audio in the sample audio set is audio annotated with at least one group of consecutive identical numbers. For example, the sample audio in the sample audio set is digital audio.

[0103] Among them, the sample audio set for speech recognition includes sample audio and annotated text, but the sample audio and the annotated text are not aligned. For example, for the sample audio containing "Hello", it is not known at which second to start reading "Hello" and at which second to start reading "good", so it is necessary to perform audio-text alignment through audio-text alignment technology. Among them, the audio-text alignment technology refers to aligning audio with the corresponding text to calibrate the pronunciation time of each text unit (such as numbers, letters, words) in the text.

[0104] In some other possible implementation manners, HMM (Hidden Markov Model) + GMM (Gaussian Mixed Model) can be used for audio-text alignment; or, the CTC (Connectionist Temporal Classification) method can be used for audio-text alignment; or, the end2end method can be used for audio-text alignment; the present application does not limit this here.

[0105] 202. The model training device screens candidate sample audio from the sample audio set based on the initial speech recognition model; among them, when the initial speech recognition model performs speech recognition on the candidate sample audio, it fails to correctly recognize the consecutive identical text units therein.

[0106] After obtaining the sample audio set, this embodiment of the application uses a speech recognition model to perform speech recognition on these sample audios. Then, based on the speech recognition results output by the speech recognition model, candidate sample audios are selected from the sample audio set. Here, to distinguish it from the speech recognition model obtained through subsequent retraining, the speech recognition model is referred to herein as the initial speech recognition model. Exemplarily, the speech recognition model used in this embodiment of the application can be an end-to-end speech recognition model, and this application does not impose any limitations on it. Furthermore, before extracting features from the sample audio, preprocessing can be performed on the sample audio; exemplary, preprocessing includes, but is not limited to, frame segmentation, pre-enhancement, windowing, noise reduction, endpoint detection, etc.

[0107] In some possible implementations, candidate audio samples are selected from the sample audio set based on the initial speech recognition model, including but not limited to the following methods: performing speech recognition on each sample audio sample based on the initial speech recognition model to obtain the predicted text of each sample audio sample; and selecting candidate audio samples from the sample audio set based on the labeled text and predicted text of each sample audio sample.

[0108] The number of candidate audio samples selected is multiple. For any given audio sample, since both the labeled and predicted texts contain multiple text units and the start and end times of each text unit within the audio sample, comparing the labeled and predicted texts of the audio sample allows us to determine whether the initial speech recognition model correctly identified consecutive identical text units. Based on the comparison results, candidate audio samples can then be selected from the set of audio samples. In other words, if the initial speech recognition model fails to correctly identify consecutive identical text units when performing speech recognition on the candidate audio samples, and the text units are numbers, then the candidate audio samples are those with consecutive identical digits that were missed.

[0109] 203. The model training device extracts audio segments from candidate sample audio; wherein, the extracted audio segments include audio segments in the candidate sample audio that are aligned with consecutive identical text units.

[0110] For any candidate audio sample, this step uses the start and end times of each text unit in the speech recognition result output by the initial speech recognition model to locate the start and end times of the audio segment containing consecutive identical text units that were not correctly recognized in the candidate audio sample. Furthermore, for fault tolerance and to prevent inaccurate speech recognition results output by the initial speech recognition model, when extracting the aforementioned audio segment, this embodiment of the application also pushes the start time of extraction forward by T audio frames and the end time of extraction backward by T audio frames. The T audio frames can be a dozen or more audio frames, etc., and this application does not impose any limitations on this.

[0111] In other words, extracting audio segments from candidate audio samples includes, but is not limited to, the following methods:

[0112] Obtain the predicted text of the candidate sample audio; determine the start and end times of the second audio based on the start and end times of each text unit in the predicted text; wherein, the second audio is the audio in the candidate sample audio that is aligned with consecutive identical text units; extract audio segments from the candidate sample audio based on the start and end times of the second audio; wherein, the extracted audio segments include adjacent first audio, second audio, and third audio; the first audio is sequentially located before the second audio; the third audio is sequentially located after the second audio.

[0113] 204. The model training device retrains the initial speech recognition model based on the extracted audio segments to obtain the target speech recognition model.

[0114] This step first uses audio-text alignment technology to re-align the extracted audio segments, and then uses the re-aligned audio segments to retrain the initial speech recognition model to obtain the target speech recognition model.

[0115] In some possible implementations, the initial speech recognition model is retrained based on the extracted audio segments to obtain the target speech recognition model, including but not limited to the following methods:

[0116] 2041. Obtain the annotation text of the extracted audio segments; wherein, the annotation text of the extracted audio segments includes consecutive identical text units and preset labels.

[0117] In some other possible implementations, the default label is... <unk>Tags. Correspondingly, the extracted audio segment's labeled text (also called alignment tags) includes consecutive identical text units and... <unk>Tags. For example, if consecutive identical text units are "66", then the alignment tag for the extracted audio segment would be "". <unk> 66 <unk>".

[0118] 2042. Align the annotation text of the extracted audio segments with the extracted audio segments.

[0119] In this embodiment, preset tags are aligned with the first and third audio segments in the extracted audio fragments, while consecutive identical text units are aligned with the second audio segment. Continuing with the example in step 2041, audio frames that do not belong to the "66" pronunciation will be aligned to... <unk>Label.

[0120] 2043. Based on the extracted audio segments and the labeled text aligned with the extracted audio segments, the initial speech recognition model is retrained to obtain the target speech recognition model.

[0121] In this embodiment of the application, during model training, only the audio features corresponding to consecutive identical text units are used to update the model parameters, without using the audio features corresponding to the same text units. <unk>The audio features corresponding to the labels are used to update the model parameters. In this way, the newly trained target speech recognition model can classify the audio features of consecutive identical text units more accurately, which can improve the model's recognition accuracy of consecutive identical text units and greatly reduce the occurrence of missed recognition.

[0122] The speech recognition model training scheme provided in this application first filters candidate audio samples that fail to correctly recognize consecutive identical text units from a sample audio set, based on an initial speech recognition model. Then, audio segments are extracted from these candidate audio samples; the extracted audio segments include those aligned with consecutive identical text units from the candidate audio samples. Next, this application re-aligns the extracted audio segments and retrains the initial speech recognition model using the re-aligned audio segments to obtain the target speech recognition model. Because the model relearns the audio features of the incorrectly recognized consecutive identical text units, it improves the model's recognition accuracy on consecutive identical text units. In other words, the newly trained target speech recognition model classifies the audio features of consecutive identical text units more accurately, improving the recognition accuracy of consecutive identical text units, resulting in better speech recognition performance and improved speech recognition quality.

[0123] Figure 3 This is a flowchart illustrating a training method for a speech recognition model provided in an embodiment of this application. See also... Figure 3 The method flow provided in this application embodiment includes:

[0124] 301. The model training device acquires a sample audio set; wherein the sample audio set includes multiple sample audios and labeled text aligned with each sample audio; the labeled text includes at least one set of consecutive identical text units.

[0125] This step is the same as step 201 above, and will not be repeated here.

[0126] 302. The model training device, based on the initial speech recognition model, filters candidate audio samples from the sample audio set; however, the initial speech recognition model fails to correctly identify consecutive identical text units when performing speech recognition on the candidate audio samples.

[0127] This step is the same as step 202 above, and will not be repeated here.

[0128] 303. The model training device extracts audio segments from candidate sample audio; wherein, the extracted audio segments include audio segments in the candidate sample audio that are aligned with consecutive identical text units.

[0129] This step is the same as step 203 above, and will not be repeated here.

[0130] 304. The model training device retrains the initial speech recognition model based on the extracted audio segments to obtain the target speech recognition model.

[0131] This step is the same as step 204 above, and will not be repeated here.

[0132] 305. In the process of speech recognition, the speech recognition device acquires the input audio to be recognized, and performs speech recognition on the input audio based on the target speech recognition model to obtain the speech recognition result corresponding to the input audio.

[0133] In some possible implementations, speech recognition is performed on the input audio based on a target speech recognition model, including but not limited to the following methods: First, feature extraction is performed on the input audio to obtain the filter bank features and the identity authentication vector corresponding to the speaker; then, discrete cosine transform is performed on the filter bank features of the input audio to obtain the Mel-frequency cepstral coefficients of the input audio; and then, the identity authentication vector and the Mel-frequency cepstral coefficients of the input audio are fused to obtain the audio fusion features of the input audio; finally, the audio fusion features of the input audio are input into the target speech recognition model to obtain the speech recognition result corresponding to the input audio.

[0134] In other possible implementations, the input audio can be the speaker's voice. This input audio can be voice captured by the microphone of the speech recognition device, or voice sent to the speech recognition device from other devices; this application does not impose any limitations on this. Furthermore, before performing feature extraction on the input audio, the speech recognition device typically preprocesses it; for example, preprocessing includes, but is not limited to, framing, pre-enhancement, windowing, noise reduction, and endpoint detection.

[0135] In this embodiment, filter bank features refer to Filter Bank features, also known as FBank features. The authentication vector refers to the i-vector. The i-vector includes not only speaker difference information but also channel difference information. In other words, the i-vector can effectively represent speaker features and channel features; that is, the i-vector is used to represent the features of the speaker and the channel.

[0136] 306. The speech recognition device outputs the speech recognition result corresponding to the input audio.

[0137] For example, the voice recognition device includes a display screen, through which the voice recognition device outputs the voice recognition result corresponding to the input audio.

[0138] In some possible implementations, embodiments of this application can be used for terminal number identification, such as identifying consecutive identical digits.

[0139] In detail, in response to the speech recognition result including a first number of consecutive digits, the speech recognition device determines the terminal number corresponding to the input audio and displays a first selection menu; wherein, the first selection menu includes multiple options, one of which indicates an installed application on the speech recognition device; in response to the user's selection operation of the target application based on the multiple options, the speech recognition device communicates with the terminal indicated by the recognized terminal number based on the target application.

[0140] For example, the target application may be a text messaging application, a calling application, or a social application on the voice recognition device, and this application is not limited thereto. In some other possible implementations, in response to a user's selection of a target application based on multiple options, the voice recognition device communicates with the terminal indicated by the recognized terminal number based on the target application, including but not limited to the following two scenarios.

[0141] Scenario 1: In response to the target application being an SMS application or a calling application, a second selection menu is displayed; wherein the second selection menu includes multiple options, one of which is used to indicate a user identification card; in response to the user's selection operation of the target user identification card based on the multiple options, communication is conducted with the terminal indicated by the identified terminal number based on the target user identification card and using the communication method corresponding to the target application.

[0142] For example, the voice recognition device is a device equipped with a user identification card, and there can be multiple user identification cards. These multiple user identification cards can come from the same operator or from different operators, and this application does not impose any restrictions here. Specifically, in response to the target application being an SMS application, the voice recognition device communicates with the terminal indicated by the identified terminal number by sending an SMS message based on the target user identification card; in response to the target application being a call application, the voice recognition device communicates with the terminal indicated by the identified terminal number by making a phone call based on the target user identification card.

[0143] Scenario 2: In response to the target application being a social application, a third selection menu is displayed; wherein the third selection menu includes multiple options, one of which indicates a communication method of the social application; in response to the user's selection of the target communication method based on the multiple options, the voice recognition device communicates with the terminal indicated by the recognized terminal number based on the target communication method; for example, the communication method of the social application includes, but is not limited to, sending application messages, voice calls, or video calls, and this application does not impose any restrictions.

[0144] In other possible implementations, the embodiments of this application can also be used for order number recognition, such as recognizing consecutive identical characters.

[0145] In detail, in response to the speech recognition result including a second number of consecutive characters, and the second number of consecutive characters matching a target logistics order number template among multiple logistics order number templates, the speech recognition device determines that the input audio corresponds to a logistics order number and obtains the logistics information of the logistics order indicated by the logistics order number.

[0146] Typically, different logistics companies use different order number templates, such as varying character lengths. Some logistics companies use templates that only include numbers, while others use templates that include both numbers and letters. Therefore, it is necessary to match the second consecutive digit of the speech recognition result with these different order number templates. For example, the speech recognition device can retrieve logistics information from the backend server of a cooperating logistics company; this application does not impose any limitations on this.

[0147] In other possible implementations, the embodiments of this application can also be used for text recognition, such as recognizing consecutive identical characters.

[0148] In detail, in response to the voice recognition result including the application name of the social application and the username of the target contact in the address book, the voice recognition device displays an information interaction page through the social application in a top-level display mode; wherein, the address book mentioned here corresponds to the address book of the social application, that is, the address book of the social application; and the information interaction page is the page for interacting with the target contact.

[0149] Additionally, the top-level display mode means that it is displayed on top of all pages, that is, at the very front of the screen, and is not covered by any other pages.

[0150] The speech recognition model training scheme provided in this application first filters candidate audio samples that fail to correctly recognize consecutive identical text units from a sample audio set, based on an initial speech recognition model. Then, audio segments are extracted from these candidate audio samples; the extracted audio segments include those aligned with consecutive identical text units from the candidate audio samples. Next, this application re-aligns the extracted audio segments and retrains the initial speech recognition model using the re-aligned audio segments to obtain the target speech recognition model. Because the model relearns the audio features of the incorrectly recognized consecutive identical text units, it improves the model's recognition accuracy on consecutive identical text units. In other words, the newly trained target speech recognition model classifies the audio features of consecutive identical text units more accurately, improving the recognition accuracy of consecutive identical text units, resulting in better speech recognition performance and improved speech recognition quality.

[0151] Figure 4 This is a schematic diagram of the structure of a training device for a speech recognition model provided in an embodiment of this application.

[0152] See Figure 4 The device includes:

[0153] The acquisition module 401 is configured to acquire a sample audio set, which includes multiple sample audio samples;

[0154] The filtering module 402 is configured to filter candidate sample audios in the sample audio set based on the initial speech recognition model;

[0155] Processing module 403 is configured to extract audio segments from the candidate sample audio;

[0156] The audio segment includes audio segments in the candidate sample audio that are aligned with consecutive identical text units; and the initial speech recognition model fails to correctly recognize the consecutive identical text units when performing speech recognition on the candidate sample audio.

[0157] Training module 404 is configured to retrain the initial speech recognition model based on the audio segment to obtain the target speech recognition model.

[0158] The speech recognition model training scheme provided in this application first filters candidate audio samples that fail to correctly recognize consecutive identical text units from a sample audio set based on an initial speech recognition model. Then, audio segments are extracted from these candidate audio samples; the extracted audio segments include audio aligned with consecutive identical text units from the candidate audio samples. Next, this application retrains the initial speech recognition model based on the extracted audio segments to obtain a target speech recognition model. Because the model relearns the audio features of the incorrectly recognized consecutive identical text units, it improves the model's recognition accuracy on consecutive identical text units. In other words, the newly trained target speech recognition model classifies the audio features of consecutive identical text units more accurately, thus improving the recognition accuracy of consecutive identical text units, resulting in better speech recognition performance and improved speech recognition quality.

[0159] In some possible implementations, the sample audio set also includes annotation text aligned with the sample audio; the annotation text includes at least one set of consecutive identical text units;

[0160] The filtering module 402 is configured to perform speech recognition on the sample audio based on the initial speech recognition model to obtain the predicted text of the sample audio; and to filter the candidate sample audio in the sample audio set according to the labeled text and predicted text of the sample audio.

[0161] In some possible implementations, the audio segment includes adjacent first audio, second audio, and third audio; wherein the first audio is sequentially preceding the second audio; the third audio is sequentially following the second audio; and the second audio is the audio in the candidate sample audio that is aligned with the consecutive identical text units.

[0162] Training module 404 is configured to acquire the labeled text of the audio segment; wherein the labeled text of the audio segment includes the consecutive identical text units and preset labels; align the labeled text of the audio segment with the audio segment; wherein the preset labels are aligned with the first audio and the third audio, and the consecutive identical text units are aligned with the second audio; and retrain the initial speech recognition model based on the audio segment and the labeled text aligned with the audio segment to obtain the target speech recognition model.

[0163] In some possible implementations, the processing module 403 is configured to acquire the predicted text of the candidate sample audio; wherein the predicted text includes multiple text units and the start and end times of the multiple text units in the candidate sample audio; determine the start and end times of the second audio based on the start and end times of each text unit in the predicted text; and extract the audio segment from the candidate sample audio based on the start and end times of the second audio.

[0164] In some possible implementations, the process of using the target speech recognition model includes:

[0165] Obtain the input audio to be recognized;

[0166] Feature extraction is performed on the input audio to obtain the filter bank features of the input audio and the identity authentication vector corresponding to the speaker;

[0167] The discrete cosine transform of the filter bank features is used to obtain the Mel-frequency cepstral coefficients of the input audio.

[0168] The identity authentication vector and the Mel cepstral coefficients are fused to obtain the audio fusion features of the input audio;

[0169] The audio fusion features are input into the target speech recognition model, and the speech recognition result corresponding to the input audio is output.

[0170] In some possible implementations, the process of using the target speech recognition model also includes:

[0171] In response to the speech recognition result including a first number of consecutive digits, the terminal number corresponding to the input audio is determined, and a first selection menu is displayed; wherein, the first selection menu includes multiple options, and one of the options is used to indicate an installed application;

[0172] In response to the selection of a target application based on the plurality of options, communication is performed between the target application and the terminal indicated by the terminal number.

[0173] In some possible implementations, the process of using the target speech recognition model also includes:

[0174] In response to the target application being an SMS application or a calling application, a second selection menu is displayed; wherein the second selection menu includes multiple options, one of which is used to indicate a user identification card; in response to the selection operation of the target user identification card based on the multiple options, based on the target user identification card, communication is performed with the terminal indicated by the terminal number using the communication method corresponding to the target application;

[0175] Alternatively, in response to the target application being a social application, a third selection menu is displayed; wherein the third selection menu includes multiple options, one of which indicates a communication method of the social application; in response to the selection of a target communication method based on the multiple options, communication is performed with the terminal indicated by the terminal number based on the target communication method;

[0176] The communication methods of the social application include sending application messages, voice calls, or video calls.

[0177] In some possible implementations, the process of using the target speech recognition model also includes:

[0178] In response to the voice recognition results, which include the application name of the social application and the username of the target contact in the address book, an information interaction page is displayed through the social application in a top-level display manner.

[0179] The address book corresponds to the social application; the information interaction page is the page for interacting with the target contact.

[0180] In some possible implementations, the process of using the target speech recognition model also includes:

[0181] In response to the speech recognition result including a second number of consecutive characters, and the second number of consecutive characters matching a target logistics order number template among multiple logistics order number templates, the system determines the logistics order number corresponding to the input audio and obtains the logistics information of the logistics order indicated by the logistics order number.

[0182] All of the above-mentioned optional technical solutions can be combined in any way to form optional embodiments of this disclosure, and will not be described in detail here.

[0183] It should be noted that the speech recognition model training device provided in the above embodiments is only illustrated by the division of the above functional modules when training the speech recognition model. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the speech recognition model training device and the speech recognition model training method embodiments provided in the above embodiments belong to the same concept, and the specific implementation process can be found in the method embodiments, which will not be repeated here.

[0184] Figure 5 This is a schematic diagram of the structure of a computer device provided in an embodiment of this application. Exemplarily, the computer 500 can be represented as a voice recognition device.

[0185] Typically, computer device 500 includes a processor 501 and a memory 502. The processor 501 may include one or more processing cores, such as a quad-core processor or an octa-core processor. The processor 501 may be implemented using at least one hardware form selected from DSP (Digital Signal Processing), FPGA (Field Programmable Gate Array), and PLA (Programmable Logic Array). The processor 501 may also include a main processor and a coprocessor. The main processor, also known as a CPU (Central Processing Unit), is used to process data in the wake-up state; the coprocessor is a low-power processor used to process data in the standby state. In one possible implementation, the processor 501 may integrate a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content to be displayed on the screen. In some embodiments, the processor 501 may also include an AI (Artificial Intelligence) processor, which handles computational operations related to machine learning.

[0186] Memory 502 may include one or more computer-readable storage media, which may be non-transitory. Memory 502 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices or flash memory devices. In one possible implementation, the non-transitory computer-readable storage media in memory 502 is used to store at least one program code, which is executed by processor 501 to implement the training method of the speech recognition model provided in the method embodiments of this application.

[0187] In one possible implementation, the computer device 500 may also optionally include a peripheral device interface 503 and at least one peripheral device. The processor 501, memory 502, and peripheral device interface 503 can be connected via a bus or signal lines. Each peripheral device can be connected to the peripheral device interface 503 via a bus, signal lines, or a circuit board. Specifically, the peripheral device includes at least one of the following: a radio frequency circuit 504, a display screen 505, a camera assembly 506, an audio circuit 507, and a power supply 508.

[0188] Peripheral device interface 503 can be used to connect at least one I / O (Input / Output) related peripheral device to processor 501 and memory 502. In one possible implementation, processor 501, memory 502, and peripheral device interface 503 are integrated on the same chip or circuit board; in some other embodiments, any one or two of processor 501, memory 502, and peripheral device interface 503 can be implemented on separate chips or circuit boards, which is not limited in this embodiment.

[0189] The radio frequency (RF) circuit 504 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The RF circuit 504 communicates with communication networks and other communication devices via electromagnetic signals. The RF circuit 504 converts electrical signals into electromagnetic signals for transmission, or converts received electromagnetic signals back into electrical signals. Optionally, the RF circuit 504 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a user identity module card, etc. The RF circuit 504 can communicate with other terminals through at least one wireless communication protocol. This wireless communication protocol includes, but is not limited to: the World Wide Web, metropolitan area networks, intranets, various generations of mobile communication networks (2G, 3G, 4G, and 5G), wireless local area networks, and / or WiFi (Wireless Fidelity) networks. In one possible implementation, the RF circuit 504 may also include circuitry related to NFC (Near Field Communication), which is not limited in this application.

[0190] Display screen 505 is used to display a UI (User Interface). This UI can include graphics, text, icons, videos, and any combination thereof. When display screen 505 is a touch display, it also has the ability to collect touch signals on or above its surface. These touch signals can be input as control signals to processor 501 for processing. In this case, display screen 505 can also be used to provide virtual buttons and / or a virtual keyboard, also known as soft buttons and / or a soft keyboard. In one possible implementation, there can be one display screen 505, located on the front panel of computer device 500; in another possible implementation, there can be at least two display screens, respectively located on different surfaces of computer device 500 or in a folded design; in yet another possible implementation, display screen 505 can be a flexible display screen, located on a curved or folded surface of computer device 500. Furthermore, display screen 505 can be configured as a non-rectangular, irregular shape, i.e., a non-rectangular screen. Display screen 505 can be made of materials such as LCD (Liquid Crystal Display) or OLED (Organic Light-Emitting Diode).

[0191] Camera assembly 506 is used to acquire images or videos. Optionally, camera assembly 506 includes a front-facing camera and a rear-facing camera. Typically, the front-facing camera is located on the front panel of the terminal, and the rear-facing camera is located on the back of the terminal. In one possible implementation, there are at least two rear-facing cameras, which are any one of a main camera, a depth-sensing camera, a wide-angle camera, and a telephoto camera, to achieve background blurring by fusion of the main camera and the depth-sensing camera, panoramic shooting by fusion of the main camera and the wide-angle camera, VR (Virtual Reality) shooting, or other fusion shooting functions. In one possible implementation, camera assembly 506 may also include a flash. The flash can be a single-color temperature flash or a dual-color temperature flash. A dual-color temperature flash refers to a combination of a warm-light flash and a cool-light flash, which can be used for light compensation at different color temperatures.

[0192] The audio circuit 507 may include a microphone and a speaker. The microphone is used to collect sound waves from the user and the environment, converting them into electrical signals that are input to the processor 501 for processing, or to the radio frequency circuit 504 for voice communication. For stereo sound acquisition or noise reduction purposes, multiple microphones may be used, positioned at different locations within the computer device 500. The microphone may also be an array microphone or an omnidirectional microphone. The speaker is used to convert electrical signals from the processor 501 or the radio frequency circuit 504 into sound waves. The speaker may be a traditional film speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can convert electrical signals not only into audible sound waves but also into inaudible sound waves for purposes such as distance measurement. In one possible implementation, the audio circuit 507 may also include a headphone jack.

[0193] Power supply 508 is used to supply power to the various components in computer device 500. Power supply 508 can be AC ​​power, DC power, a disposable battery, or a rechargeable battery. When power supply 508 includes a rechargeable battery, the rechargeable battery can be a wired rechargeable battery or a wireless rechargeable battery. A wired rechargeable battery is a battery that is charged via a wired line, while a wireless rechargeable battery is a battery that is charged via a wireless coil. The rechargeable battery can also be used to support fast charging technology.

[0194] Those skilled in the art will understand that Figure 5 The structure shown does not constitute a limitation on the computer device 500, and may include more or fewer components than shown, or combine certain components, or use different component arrangements.

[0195] Figure 6 This is a schematic diagram of another computer device 600 provided in an embodiment of this application. Exemplarily, the computer 600 can be represented as a model training device.

[0196] The computer device 600 can vary considerably due to differences in configuration or performance. It may include one or more Central Processing Units (CPUs) 601 and one or more memories 602. The memories 602 store at least one line of program code, which is loaded and executed by the processor 601 to implement the training method for the speech recognition model provided in the various method embodiments described above. Of course, the computer device 600 may also have wired or wireless network interfaces, a keyboard, and input / output interfaces for input and output. The computer device 600 may also include other components for implementing device functions, which will not be elaborated upon here.

[0197] In an exemplary embodiment, a computer-readable storage medium is also provided, such as a memory including program code that can be executed by a processor in a computer device to complete the training method of the speech recognition model in the above embodiments. For example, the computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a compact disc read-only memory (CD-ROM), magnetic tape, floppy disk, and optical data storage device, etc.

[0198] In an exemplary embodiment, a computer program product or computer program is also provided, which includes computer program code stored in a computer-readable storage medium. A processor of a computer device reads the computer program code from the computer-readable storage medium and executes the computer program code, causing the computer device to perform the above-described training method for the speech recognition model.

[0199] Those skilled in the art will understand that all or part of the steps of the above embodiments can be implemented by hardware or by a program instructing related hardware. The program can be stored in a computer-readable storage medium, such as a read-only memory, a disk, or an optical disk.

[0200] The above description is merely an optional embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.< / unk> < / unk> < / unk> < / unk> < / unk> < / unk>

Claims

1. A method for training a speech recognition model, characterized in that, The method includes: Obtain a sample audio set, which includes multiple sample audio files; Based on the initial speech recognition model, candidate audio samples are selected from the set of sample audio samples. Obtain the predicted text of the candidate sample audio; the predicted text includes multiple text units and the start and end times of the multiple text units in the candidate sample audio; The start and end times of the second audio are determined based on the start and end times of each text unit in the predicted text; the second audio is the audio in the candidate sample audio that is aligned with consecutive identical text units. Based on the start and end times of the second audio, audio segments are extracted from the candidate sample audio; the audio segments include adjacent first audio, second audio, and third audio, where the first audio is sequentially located before the second audio and the third audio is sequentially located after the second audio; and the initial speech recognition model does not correctly recognize the consecutive identical text units when performing speech recognition on the candidate sample audio. Obtain the labeled text of the audio segment; the labeled text of the audio segment includes the consecutive identical text units and preset tags, the preset tags being located on both sides of the consecutive identical text units; Align the labeled text of the audio segment with the audio segment; wherein, the preset label is aligned with the first audio and the third audio, and the consecutive identical text units are aligned with the second audio; The model parameters are updated using the audio features corresponding to the consecutive identical text units to obtain the target speech recognition model.

2. The method according to claim 1, characterized in that, The sample audio set also includes annotation text aligned with the sample audio; the annotation text includes at least one set of consecutive identical text units; The step of filtering candidate audio samples from the sample audio set based on the initial speech recognition model includes: Based on the initial speech recognition model, speech recognition is performed on the sample audio to obtain the predicted text of the sample audio; Based on the labeled and predicted text of the sample audio, the candidate sample audio is filtered in the sample audio set.

3. The method according to any one of claims 1 to 2, characterized in that, The method further includes: Obtain the input audio to be recognized; Feature extraction is performed on the input audio to obtain the filter bank features of the input audio and the identity authentication vector corresponding to the speaker; The discrete cosine transform of the filter bank features is used to obtain the Mel-frequency cepstral coefficients of the input audio. The identity authentication vector and the Mel cepstral coefficients are fused to obtain the audio fusion features of the input audio; The audio fusion features are input into the target speech recognition model, and the speech recognition result corresponding to the input audio is output.

4. The method according to claim 3, characterized in that, The method further includes: In response to the speech recognition result including a first number of consecutive digits, the terminal number corresponding to the input audio is determined, and a first selection menu is displayed; wherein, the first selection menu includes multiple options, and one of the options is used to indicate an installed application; In response to the selection of a target application based on the plurality of options, communication is performed between the target application and the terminal indicated by the terminal number.

5. The method according to claim 4, characterized in that, The response to selecting a target application based on the plurality of options, and communicating between the target application and the terminal indicated by the terminal number, includes: In response to the target application being an SMS application or a calling application, a second selection menu is displayed; wherein the second selection menu includes multiple options, one of which is used to indicate a user identification card; in response to the selection operation of the target user identification card based on the multiple options, based on the target user identification card, communication is performed with the terminal indicated by the terminal number using the communication method corresponding to the target application; Alternatively, in response to the target application being a social application, a third selection menu is displayed; wherein the third selection menu includes multiple options, one of which indicates a communication method of the social application; in response to the selection of a target communication method based on the multiple options, communication is performed with the terminal indicated by the terminal number based on the target communication method; The communication methods of the social application include sending application messages, voice calls, or video calls.

6. The method according to claim 3, characterized in that, The method further includes: In response to the voice recognition results, which include the application name of the social application and the username of the target contact in the address book, an information interaction page is displayed through the social application in a top-level display manner. The address book corresponds to the social application; the information interaction page is the page for interacting with the target contact.

7. The method according to claim 3, characterized in that, The method further includes: In response to the speech recognition result including a second number of consecutive characters, and the second number of consecutive characters matching a target logistics order number template among multiple logistics order number templates, the system determines the logistics order number corresponding to the input audio and obtains the logistics information of the logistics order indicated by the logistics order number.

8. A training device for a speech recognition model, characterized in that, The device includes: The acquisition module is configured to acquire a sample audio set, which includes multiple sample audio files. The filtering module is configured to filter candidate audio samples from the set of sample audio samples based on an initial speech recognition model. A processing module is configured to acquire predicted text of the candidate sample audio; the predicted text includes multiple text units and the start and end times of the multiple text units in the candidate sample audio; determine the start and end times of a second audio based on the start and end times of each text unit in the predicted text; the second audio is the audio in the candidate sample audio that is aligned with consecutive identical text units; extract audio segments from the candidate sample audio based on the start and end times of the second audio; the audio segments include adjacent first audio, second audio, and third audio, the first audio being sequentially preceding the second audio, and the third audio being sequentially following the second audio; and the initial speech recognition model fails to correctly recognize the consecutive identical text units when performing speech recognition on the candidate sample audio. The training module is configured to acquire the labeled text of the audio segment; the labeled text of the audio segment includes consecutive identical text units and preset labels, the preset labels being located on both sides of the consecutive identical text units; aligning the labeled text of the audio segment with the audio segment; wherein the preset labels are aligned with the first audio and the third audio, and the consecutive identical text units are aligned with the second audio; updating the model parameters using the audio features corresponding to the consecutive identical text units to obtain the target speech recognition model.

9. A computer device, characterized in that, The device includes a processor and a memory, the memory storing at least one line of program code, the at least one line of program code being loaded and executed by the processor to implement the training method of the speech recognition model as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The storage medium stores at least one piece of program code, which is loaded and executed by a processor to implement the training method of the speech recognition model as described in any one of claims 1 to 7.