Speech recognition method and device, electronic equipment, storage medium and program product
By using pre-trained speech recognition models in the speech recognition system, the semantic features and speaker characteristics of the speech encoder in the speech recognition system are extracted, and the problem of low accuracy of traditional speech recognition systems is solved, and higher speech recognition accuracy and model generalization capabilities are achieved.
Patent Information
- Application Number
- CN202510489152.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-17
- Publication Date
- 2025-06-20
AI Technical Summary
Traditional speech recognition systems face the problem of low accuracy of speech recognition in practical applications.
By obtaining the speech to be recognized, the pre-trained speech recognition model, including a speech encoder, extracts the first feature vector of semantic features and speaker features, thereby improving the accuracy of speech recognition. This speech recognition model is obtained by self-supervised pre-training based on single-person speech samples.
The accuracy of speech recognition is improved, especially in scenes where there are noisy voices such as many speakers, the voice content of the target user in the target voice can be more accurately identified, reducing the dependence on the annotated data, and enhancing the generalization ability and adaptability of the model.
Smart Images

Figure CN120183406A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technologies, and in particular, to a speech recognition method, apparatus, electronic device, storage medium, and program product. Background Art
[0002] With the continuous progress of artificial intelligence technologies, speech recognition technology, as an important branch in the field of natural language processing, has achieved remarkable development. However, traditional speech recognition systems still face many challenges in practical applications, and the speech recognition accuracy is relatively low. Summary of the Invention
[0003] To overcome the problems existing in the related art, the present disclosure provides a speech recognition method, apparatus, electronic device, storage medium, and program product.
[0004] According to a first aspect of an embodiment of the present disclosure, there is provided a speech recognition method, including: Obtaining a speech to be recognized, where the speech to be recognized includes a wake-up speech and a target speech; Inputting the speech to be recognized into a pre-trained speech recognition model to obtain a predicted text, where the predicted text is a text for characterizing the speech content of a target user in the target speech, and the target user is the speaker of the wake-up speech; Wherein, the speech recognition model includes a speech encoder, and the speech encoder is configured to obtain a first feature vector including a semantic feature and a speaker feature according to the speech to be recognized, and the first feature vector is used to determine the predicted text, and the speech encoder is obtained by self-supervised pre-training based on at least one single-person speech sample.
[0005] Optionally, the speech recognition method further includes: Obtaining at least one positive sample and / or at least one negative sample; Training an initial speech recognition model by using the positive sample and / or the negative sample until a training end condition is satisfied; Wherein, each positive sample includes a first speech sample, a text for characterizing the content of the single-person speech sample in the first speech sample, and a first wake-up speech sample, and the first wake-up speech sample corresponds to the same speaker as the single-person speech sample in the first speech sample; Each negative sample includes a second speech sample, an empty text sample, and a second wake-up speech sample, and the second wake-up speech sample corresponds to a different speaker from the single-person speech sample in the second speech sample.
[0006] Optionally, the speech recognition method further includes at least one of the following: Superimpose the single-person speech sample and the first interfering sound to obtain the first speech sample; Superimpose the single-person speech sample and the second interfering sound to obtain the second speech sample.
[0007] Optionally, the first interfering sound, and / or the second interfering sound is interfering speech, and the interfering speech corresponds to a different speaker from the single-person speech sample.
[0008] Optionally, training the initial speech recognition model using the positive sample and / or the negative sample until the training end condition is met includes at least one of the following: Train the initial speech recognition model in a first stage using the positive sample until the model loss value of the speech recognition model is less than a first loss threshold; Train the speech recognition model that has completed the first-stage training using the positive sample and the negative sample in a second stage until the model loss value of the speech recognition model is less than a second loss threshold.
[0009] Optionally, training the initial speech recognition model using the positive sample and / or the negative sample includes: Train the initial speech recognition model using the positive sample and / or the negative sample according to a preset prompt instruction, and the prompt instruction is used to indicate the target task of the speech recognition model.
[0010] Optionally, the speech recognition method further includes: Before training the initial speech recognition model using the positive sample and / or the negative sample, train the large language model in the speech recognition model according to at least one single-person speech recognition data, and the single-person speech recognition data includes the single-person speech sample and the text used to characterize the content of the single-person speech sample.
[0011] Optionally, the speech recognition model further includes a linear transformation layer, a Transformer mapping layer, a tokenizer, and a large language model; Inputting the speech to be recognized into the pre-trained speech recognition model to obtain a predicted text includes: Input the speech to be recognized into the speech encoder to obtain the first feature vector; Use the linear transformation layer to map the first feature vector into the feature space of the large language model to obtain a second feature vector; Use the Transformer mapping layer to perform alignment processing on the second feature vector in the time dimension to obtain a third feature vector; Process the preset prompt instruction by using the tokenizer to obtain a prompt instruction feature vector, where the prompt instruction is used to indicate the target task of the speech recognition model; Input the third feature vector and the prompt instruction feature vector into the large language model to generate the predicted text.
[0012] According to a second aspect of the embodiments of the present disclosure, there is provided a speech recognition device, including: An acquisition module, configured to acquire the speech to be recognized, where the speech to be recognized includes a wake-up speech and a target speech; A prediction module, configured to input the speech to be recognized into a pre-trained speech recognition model to obtain a predicted text, where the predicted text is a text used to represent the speech content of the target user in the target speech, and the target user is the speaker of the wake-up speech; Wherein, the speech recognition model includes a speech encoder, and the speech encoder is configured to obtain a first feature vector including semantic features and speaker features according to the speech to be recognized, and the first feature vector is used to determine the predicted text, and the speech encoder is self-supervised pre-trained based on at least one single-person speech sample.
[0013] According to a third aspect of the embodiments of the present disclosure, there is provided an electronic device, including: A processor; A memory for storing processor-executable instructions; Wherein, the processor is configured to execute the executable instructions in the memory to implement the steps of the speech recognition method provided in the first aspect of the present disclosure.
[0014] According to a fourth aspect of the embodiments of the present disclosure, there is provided a computer-readable storage medium, on which computer program instructions are stored, and when the program instructions are executed by a processor, the steps of the speech recognition method provided in the first aspect of the present disclosure are implemented.
[0015] According to a fifth aspect of the embodiments of the present disclosure, there is provided a computer program product, including a computer program, and when the computer program is executed by a processor, the steps of the speech recognition method provided in the first aspect of the present disclosure are implemented.
[0016] The technical solutions provided by the embodiments of the present disclosure may include the following beneficial effects: The speech to be recognized, which includes a wake-up voice and a target voice, is input into a pre-trained speech recognition model. By using the wake-up voice, the speech recognition model can more accurately capture the speech content of the target user in the target voice, improving the accuracy of speech recognition. The speech recognition model can extract a first feature vector that includes both semantic features and speaker features through a speech encoder. Among them, the semantic features help to accurately understand the speech content, and the speaker features are used to distinguish the speech characteristics of different speakers. Thus, in a noisy scenario such as multiple speakers, the speech content of the target user in the target voice can be more accurately recognized using the first feature vector. In addition, self-supervised pre-training of the speech encoder using single-person speech samples can effectively reduce the dependence on labeled data and enhance the generalization ability and adaptability of the model.
[0017] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and do not limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] The accompanying drawings herein are incorporated into the specification and constitute a part of the specification, showing embodiments consistent with the present disclosure and used together with the specification to explain the principles of the present disclosure.
[0019] Figure 1 is a flowchart of a speech recognition method shown according to an exemplary embodiment.
[0020] Figure 2 is a schematic diagram of a speech recognition model training process shown according to an exemplary embodiment.
[0021] Figure 3 is a schematic diagram of a training sample generation shown according to an exemplary embodiment.
[0022] Figure 4 is a schematic diagram of a training sample generation shown according to an exemplary embodiment.
[0023] Figure 5 is a flowchart of a multi-stage fine-tuning of a speech recognition model shown according to an exemplary embodiment.
[0024] Figure 6 is a flowchart of a multi-stage fine-tuning of a speech recognition model shown according to an exemplary embodiment.
[0025] Figure 7 is a schematic diagram of a speech recognition model training process shown according to an exemplary embodiment.
[0026] Figure 8 is a schematic diagram of a speech recognition model structure shown according to an exemplary embodiment.
[0027] Figure 9 It is a schematic diagram showing the generation of predicted text using a speech recognition model according to an exemplary embodiment.
[0028] Figure 10 It is a block diagram of a speech recognition device according to an exemplary embodiment.
[0029] Figure 11 It is a block diagram of an electronic device according to an exemplary embodiment.
[0030] Figure 12 It is a block diagram of an electronic device according to an exemplary embodiment. Detailed implementation manners
[0031] Here, the exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.
[0032] It should be noted that all actions of obtaining signals, information, or data in the present disclosure are carried out on the premise of complying with the corresponding data protection regulations and policies of the country where the location is located and obtaining authorization from the owner of the corresponding device.
[0033] Figure 1 It is a flowchart of a speech recognition method according to an exemplary embodiment. As Figure 1 shown, the method may include step S101 and step S102.
[0034] In step S101, the speech to be recognized is obtained.
[0035] Exemplarily, the speech to be recognized can be obtained through a sound collection device such as a microphone.
[0036] Among them, the speech to be recognized includes a wake-up speech and a target speech. The wake-up speech can be a preset wake-up speech instruction given by the target user, which is an acoustic signal for the user to trigger the device to switch from the sleep or standby state to the active listening and execution of instructions through a preset speech keyword (such as "Hey Siri" or "Xiaomi AI Assistant", etc.). The target speech can be the content other than the wake-up speech in the speech to be recognized, or the target speech can be the speech after the wake-up speech in the speech to be recognized, or the speech of the user collected after the device is awakened.
[0037] In step S102, the speech to be recognized is input into a pre-trained speech recognition model to obtain predicted text.
[0038] Among them, the predicted text is the text used to characterize the speech content of the target user in the target speech, and the target user is the speaker of the wake-up speech.
[0039] Among them, the speech recognition model includes a speech encoder, which is used to obtain a first feature vector including semantic features and speaker features according to the speech to be recognized. The first feature vector is used to determine the predicted text, and the speech encoder is obtained by self-supervised pre-training based on at least one single-person speech sample.
[0040] Exemplarily, the speech encoder included in the speech recognition model can be a Data2Vec2 speech encoder. A large number of unlabeled single-person speech samples can be used to perform self-supervised pre-training on the speech encoder. This pre-training method effectively reduces the need for labeled data and reduces the cost of data collection and annotation. And this pre-training method enables the model to better generalize and adapt when facing new speech.
[0041] The speech recognition model in the present disclosure can use a speech encoder to extract a first feature vector that includes both semantic features and speaker features. The semantic features are used to understand the meaning expressed by the speech, and the speaker features are used to distinguish the speakers in the speech to be recognized. Therefore, using the first feature vector, the speech recognition model can accurately identify the speech content of the target user in the target speech. And by using a single speech encoder to extract the first feature vector, the speech recognition model can have a lower deployment cost.
[0042] It should be noted that the pre-trained speech recognition model applied in step S102 can be trained by using the training method in other embodiments of the present disclosure.
[0043] In the above technical solution, the speech to be recognized including the wake-up speech and the target speech is input into the pre-trained speech recognition model. By using the wake-up speech, the speech recognition model can more accurately capture the speech content of the target user in the target speech and improve the accuracy of speech recognition. The speech recognition model can use the speech encoder to extract a first feature vector that includes both semantic features and speaker features. Among them, the semantic features help to accurately understand the speech content, and the speaker features are used to distinguish the speech characteristics of different speakers. In this way, in a noisy scenario such as multiple speakers, using the first feature vector, the speech content of the target user in the target speech can be more accurately recognized. In addition, using single-person speech samples to perform self-supervised pre-training on the speech encoder can effectively reduce the dependence on labeled data and enhance the generalization ability and adaptability of the model.
[0044] Figure 2is a flowchart of a method for training a speech recognition model shown according to an exemplary embodiment. As Figure 2 shown, the training method may include step S201 and step S202.
[0045] In step S201, at least one positive sample and / or at least one negative sample are obtained.
[0046] Each positive sample includes a first speech sample, text for characterizing the content of the single-speaker speech sample in the first speech sample, and a first wake-up speech sample, where the first wake-up speech sample corresponds to the same speaker as the single-speaker speech sample in the first speech sample; Each negative sample includes a second speech sample, an empty text sample, and a second wake-up speech sample, where the second wake-up speech sample corresponds to a different speaker from the single-speaker speech sample in the second speech sample.
[0047] In an optional embodiment, positive and negative samples may be constructed in the manner as Figure 3 shown: Obtain single-speaker speech samples; According to the single-speaker speech samples, single-speaker speech recognition data are obtained, where the single-speaker speech recognition data include the single-speaker speech samples and text for characterizing the content of the single-speaker speech samples, and among them, some of the single-speaker speech recognition data carry voiceprint tags and some do not carry voiceprint tags; Use the single-speaker speech recognition data carrying voiceprint tags to construct positive and negative samples.
[0048] For example, a positive sample may be constructed in the following manner: Determine any single-speaker speech sample as the first speech sample, determine the text for characterizing the content of this single-speaker speech sample as the text of the corresponding first speech sample, use the voiceprint tag to determine the speech corresponding to the same speaker as this single-speaker speech sample as the first wake-up speech sample, the first speech sample and the first wake-up speech sample in this positive sample can be spliced to obtain the input data when training the speech recognition model, and at this time, the text in this positive sample is the target output when training the speech recognition model. Exemplarily, the first wake-up speech sample may be a preset voice command corresponding to the single-speaker speech sample.
[0049] For example, a negative sample can be constructed in the following way: Determine any single-person speech sample as the second speech sample, and use the voiceprint label to determine the speech of a different speaker corresponding to this single-person speech sample as the second wake-up speech sample. The second speech sample and the second wake-up speech sample in this negative sample can be spliced to obtain the input data when training the speech recognition model. At this time, the empty text sample in this negative sample is the target output when training the speech recognition model. By way of example, the second wake-up speech sample can be a preset voice command of a different speaker corresponding to the single-person speech sample. The second wake-up speech sample and the single-person speech sample in the second speech sample correspond to different speakers. Therefore, constructing a negative sample can be understood as constructing a purely interfering speech sample. Training with negative samples can make the speech recognition model not respond to interfering speech. When the speech input to the speech recognition model does not include the target speech of the same speaker as the wake-up speech, the speech recognition model will not output text or will output an empty text. That is, when performing speech recognition using the trained speech recognition model, it will not output the speech content text of speakers other than the target user, but will output the speech content text of the target user.
[0050] In an optional embodiment, the first speech sample can be determined in the following way: Perform superposition processing on the single-person speech sample and the first interfering sound to obtain the first speech sample. The second speech sample can be determined in the following way: Perform superposition processing on the single-person speech sample and the second interfering sound to obtain the second speech sample.
[0051] In the present disclosure, the first speech sample and the second speech sample may include the same speech sample or may include different speech samples, and no limitation is imposed here. The first interfering sound and the second interfering sound may be the same sound or may be different sounds, and no limitation is imposed here either.
[0052] In one embodiment, the first interfering sound, and / or, the second interfering sound can be music, noise, etc. Among them, the noise can include sounds such as TV background sound, pet barking, tire noise, environmental noise, etc.
[0053] In yet another embodiment, the first interfering sound, and / or, the second interfering sound can be interfering speech of a different speaker corresponding to the single-person speech sample to be superposed. For example, different single-person speech samples correspond to different speakers. As Figure 4 shown, the interfering speech can be a single-person speech sample without a corresponding voiceprint label.
[0054] By performing superposition processing on the single-person speech and the interfering sound, a real speaking environment can be simulated, thereby improving the ability of the speech recognition model to process background noise and enhancing its recognition ability and robustness in a complex acoustic environment.
[0055] Turn back Figure 2, in step S202, the initial speech recognition model is trained using positive samples and / or negative samples until the training end condition is met.
[0056] Exemplarily, the initial speech recognition model can be trained using positive samples or negative samples until the training end condition is met. Another example is that the initial speech recognition model can be trained using positive samples and negative samples until the training end condition is met to reduce the misrecognition of the speech recognition model.
[0057] Exemplarily, the training end condition can include at least one of the following: the output value of the loss function of the model is less than or equal to a preset threshold, and the number of iterations reaches a preset number threshold.
[0058] In one embodiment, the speech recognition model can be trained in a single stage using positive samples and / or negative samples.
[0059] In another embodiment, the speech recognition model can be fine-tuned in multiple stages using positive samples and negative samples to achieve model training, and this process can be as Figure 5 shown in step S301 and step S302.
[0060] In step S301, the initial speech recognition model is trained in the first stage using positive samples until the model loss value of the speech recognition model is less than the first loss threshold.
[0061] In step S302, the speech recognition model that has completed the first stage of training is trained in the second stage using positive samples and negative samples until the model loss value of the speech recognition model is less than the second loss threshold.
[0062] Exemplarily, both the first loss threshold and the second loss threshold can be preset based on actual requirements.
[0063] First, the initial speech recognition model is trained in the first stage using positive samples, which can initially establish an understanding of speech features and improve the ability of the speech recognition model to identify the speaker. On the basis of the completion of the first stage of training, joint training using positive and negative samples in the second stage can further optimize the robustness of the speech recognition model, enabling it to better determine the speaker, improve the ability to remove non-target user speech content, and avoid overfitting of the model to positive samples, thereby improving the generalization ability and anti-interference ability of the model in actual applications.
[0064] In this way, through multi-stage fine-tuning, the parameters of the speech recognition model can be gradually optimized, improving the accuracy and robustness of the speech recognition model, enabling it to have the ability to accurately identify the speech content of the target user in the target speech; at the same time, the risk of overfitting can be reduced, and the generalization ability and adaptability of the speech recognition model can be enhanced.
[0065] It should be noted that if the speech samples in the positive and negative samples used for multi-stage fine-tuning are speech samples with interfering sounds superimposed on single-person speech samples, through the training in the first stage, the speech recognition model can initially acquire the ability to process background noise, and through the training in the second stage, its ability to process background noise can be further enhanced, improving the recognition ability and robustness of the speech recognition model in complex acoustic environments.
[0066] In an optional embodiment, training the initial speech recognition model using positive samples and / or negative samples includes: training the initial speech recognition model using positive samples and / or negative samples according to a preset prompt instruction.
[0067] Among them, the prompt instruction is used to indicate the target task of the speech recognition model. The prompt instruction can be preset by relevant personnel based on the target task of the speech recognition model, and is used to guide the model to generate specific output content. For example, the prompt instruction is used to indicate extracting the speech of the same speaker corresponding to the wake-up speech in the input speech and outputting the text representing the content of this speech. Taking training using positive samples as an example, based on the prompt instruction, the speech recognition model can extract the single-person speech sample of the same speaker corresponding to the first wake-up speech sample in the first speech sample and determine the predicted text representing the content of this single-person speech sample. Taking training using negative samples as an example, based on the prompt instruction, the speech recognition model cannot extract the part of the second speech sample corresponding to the same speaker as the second wake-up speech sample, and at this time, the determined predicted text is an empty text.
[0068] As Figure 6 shown, before performing step S202, step S303 can be executed to train the large language model in the speech recognition model.
[0069] In step S303, the large language model in the speech recognition model is trained according to at least one single-person speech recognition data.
[0070] Among them, the single-person speech recognition data includes a single-person speech sample and the text used to represent the content of the single-person speech sample. Exemplarily, the large language model can be the Qwen2.5 model. Training the large language model in the speech recognition model using single-person speech recognition data can enable the large language model to adapt to the speech modality and improve its ability to recognize single-person speech.
[0071] Figure 7 is a schematic diagram of a speech recognition model training process shown according to an exemplary embodiment. Through this Figure 7 , the implementation process of the speech recognition model training provided by the present disclosure can be more clearly understood. As Figure 7 shown, the method can include the following steps.
[0072] In step S401, a single person voice sample is obtained.
[0073] In step S402, single-person speech recognition data including a single-person speech sample and text for representing the content of the single-person speech sample is constructed.
[0074] In step S403, the single-person voice sample and the first interfering sound are superimposed to obtain a first voice sample; the single-person voice sample and the second interfering sound are superimposed to obtain a second voice sample.
[0075] In step S404, at least one positive sample is constructed, each positive sample includes a first voice sample, a text used to characterize the content of a single voice sample in the first voice sample, and a positive sample of a first wake-up voice sample, and the first wake-up voice sample and the single voice sample in the first voice sample correspond to the same speaker.
[0076] In step S405, at least one negative sample is constructed, each negative sample includes a second voice sample, an empty text sample and a second wake-up voice sample, and the second wake-up voice sample and the single-person voice sample in the second voice sample correspond to different speakers.
[0077] In step S406, a speech encoder of the speech recognition model is self-supervised pre-trained based on at least one single-person speech sample.
[0078] In step S407, the large language model in the speech recognition model is trained according to at least one piece of single-person speech recognition data.
[0079] In step S408, the initial speech recognition model is trained in the first stage according to the preset prompt instruction using positive samples until the model loss value of the speech recognition model is less than the first loss threshold.
[0080] In step S409, the speech recognition model that has completed the first stage training is trained in the second stage according to the preset prompt instructions using positive samples and negative samples until the model loss value of the speech recognition model is less than the second loss threshold.
[0081] Through the above technical solution, a large speech recognition model for the target user can be obtained through gradual fine-tuning, while maintaining the speech recognition performance as much as possible, while introducing the ability to identify the speaker. And the data can be more fully utilized to alleviate the performance problems caused by the lack of speech recognition data with voiceprint tags. The ability of the speech recognition model to handle noise is improved, and its recognition ability and robustness in complex acoustic environments are improved, so that it has the ability to accurately identify the speech content of the target user in the target speech.
[0082] In addition, the specific implementation of the above steps S401 to S409 has been described in detail above, and repeated content will not be elaborated here.
[0083] Figure 8 It is a schematic diagram of a speech recognition model structure shown according to an exemplary embodiment. As Figure 8 shown, the speech recognition model may include a speech encoder, a linear transformation layer (Linear Adapter), a Transformer mapping layer, a text tokenizer, and a large language model. Among them, the parameters in the speech encoder, the linear transformation layer, the Transformer mapping layer, and the large language model can be continuously adjusted as the speech recognition model is trained until the training end condition is met. Combining Figure 9 with the content shown, the implementation process of generating the predicted text by the speech recognition model provided by the present disclosure can be more clearly understood, and the generation of the predicted text can be realized based on the following steps: Input the speech to be recognized into the speech encoder to obtain a first feature vector; Use the linear transformation layer to map the first feature vector into the feature space of the large language model to obtain a second feature vector; Use the Transformer mapping layer to perform alignment processing on the second feature vector in the time dimension to obtain a third feature vector; Use the text tokenizer to process the preset prompt instruction to obtain a prompt instruction feature vector; Input the third feature vector and the prompt instruction feature vector into the large language model to generate a predicted text.
[0084] The speech recognition provided by the present disclosure can be applied to multi-speaker scenarios, such as in-vehicle, conference, shopping mall and other scenarios, or other scenarios with relatively noisy sounds. For example, for the out-of-vehicle scenario, it can enable users to receive more convenient and accurate services when issuing instructions such as opening and closing the trunk and parking outside the vehicle in a high-noise environment.
[0085] In one embodiment, the wake-up speech registered by the user can be obtained in advance and saved. When the speech to be recognized is received, the speech to be recognized can be input into the speech recognition model. The speech recognition model can use the wake-up speech registered by the user to recognize the speech of the same speaker (i.e., the target user) corresponding to the wake-up speech from the speech to be recognized, and output the speech content of the target user in the target speech. If the speech recognition model outputs a non-empty text, it can be determined that the target user is speaking, and the text of the speech content of the target user in the target speech can be output, excluding the influence of possible noise and overlapping human voices in the speech. If the output result of the speech recognition model is empty, it can be determined that the target user is not speaking, and no subsequent reaction will be made.
[0086] It should be noted that the above operations are all carried out under the authorization of the user and strictly comply with relevant laws and regulations such as privacy and security.
[0087] Based on the same inventive concept, the present disclosure also provides a voice recognition device. Figure 10 It is a block diagram of a voice recognition device 500 provided by an exemplary embodiment of the present disclosure. Refer to Figure 10 , the voice recognition device 500 may include: An acquisition module 501, configured to acquire the voice to be recognized, where the voice to be recognized includes a wake-up voice and a target voice; A prediction module 502, configured to input the voice to be recognized into a pre-trained voice recognition model to obtain a prediction text, where the prediction text is a text used to represent the voice content of the target user in the target voice, and the target user is the speaker of the wake-up voice; Wherein, the voice recognition model includes a voice encoder, and the voice encoder is configured to obtain a first feature vector including semantic features and speaker features according to the voice to be recognized, and the first feature vector is used to determine the prediction text, and the voice encoder is self-supervised pre-trained based on at least one single-person voice sample.
[0088] In the above technical solution, the voice to be recognized including the wake-up voice and the target voice is input into the pre-trained voice recognition model. By using the wake-up voice, the voice recognition model can more accurately capture the voice content of the target user in the target voice, improving the accuracy of voice recognition. The voice recognition model can extract a first feature vector including both semantic features and speaker features through the voice encoder. Among them, the semantic features help to accurately understand the voice content, and the speaker features are used to distinguish the voice characteristics of different speakers. Thus, in a noisy scene such as multiple speakers, the voice content of the target user in the target voice can be more accurately recognized by using the first feature vector. In addition, self-supervised pre-training of the voice encoder using single-person voice samples can effectively reduce the dependence on labeled data and enhance the generalization ability and adaptability of the model.
[0089] Optionally, the voice recognition device 500 may further include: A training module, configured to acquire at least one positive sample and / or at least one negative sample; and use the positive sample and / or the negative sample to train the initial voice recognition model until the training end condition is met; Wherein, each positive sample includes a first voice sample, a text used to represent the content of the single-person voice sample in the first voice sample, and a first wake-up voice sample, and the first wake-up voice sample corresponds to the same speaker as the single-person voice sample in the first voice sample; Each of the negative samples includes a second speech sample, an empty text sample, and a second wake-up speech sample, where the second wake-up speech sample corresponds to a different speaker from the single-speaker speech sample in the second speech sample.
[0090] Optionally, the speech recognition device 500 may further include at least one of the following: A first superposition module for superposing the single-speaker speech sample and the first interfering sound to obtain the first speech sample; A second superposition module for superposing the single-speaker speech sample and the second interfering sound to obtain the second speech sample.
[0091] Optionally, the first interfering sound, and / or, the second interfering sound is interfering speech, and the interfering speech corresponds to a different speaker from the single-speaker speech sample.
[0092] Optionally, the training module is used to train the initial speech recognition model with the positive samples and / or the negative samples in the following manner until the training end condition is met: Training the initial speech recognition model in a first stage with the positive samples until the model loss value of the speech recognition model is less than the first loss threshold; Training the speech recognition model that has completed the first stage of training with the positive samples and the negative samples until the model loss value of the speech recognition model is less than the second loss threshold.
[0093] Optionally, the training module is further used to train the initial speech recognition model according to a preset prompt instruction with the positive samples and / or the negative samples, and the prompt instruction is used to indicate the target task of the speech recognition model.
[0094] Optionally, the training module is further used to train the large language model in the speech recognition model according to at least one single-speaker speech recognition data before training the initial speech recognition model with the positive samples and / or the negative samples, and the single-speaker speech recognition data includes the single-speaker speech sample and the text used to characterize the content of the single-speaker speech sample.
[0095] Optionally, the speech recognition model further includes a linear transformation layer, a Transformer mapping layer, a tokenizer, and a large language model; the prediction module 502 is used to obtain the predicted text in the following manner: Inputting the speech to be recognized into the speech encoder to obtain the first feature vector; Using the linear transformation layer to map the first feature vector into the feature space of the large language model to obtain a second feature vector; Align the second feature vector in the time dimension using the Transformer mapping layer to obtain a third feature vector; Process a preset prompt instruction using the tokenizer to obtain a prompt instruction feature vector, where the prompt instruction is used to indicate the target task of the speech recognition model; Input the third feature vector and the prompt instruction feature vector into the large language model to generate the predicted text.
[0096] Regarding the device in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated here.
[0097] The present disclosure also provides a computer-readable storage medium, on which computer program instructions are stored, and when the program instructions are executed by a processor, the steps of the speech recognition method provided by the present disclosure are implemented.
[0098] Figure 11 FIG. is a block diagram of an electronic device 800 shown according to an exemplary embodiment. For example, the electronic device 800 may be a mobile phone, a computer, a messaging device, a tablet device, a personal digital assistant, a vehicle, etc.
[0099] Refer to Figure 11 , the electronic device 800 may include one or more of the following components: a first processing component 802, a first memory 804, a first power component 806, a multimedia component 808, an audio component 810, a first input / output interface 812, a sensor component 814, and a communication component 816.
[0100] The first processing component 802 generally controls the overall operation of the electronic device 800, such as operations associated with display, telephone calls, data communication, camera operations, and recording operations. The first processing component 802 may include one or more first processors 820 to execute instructions to complete all or part of the steps of the above speech recognition method. In addition, the first processing component 802 may include one or more modules to facilitate the interaction between the first processing component 802 and other components. For example, the first processing component 802 may include a multimedia module to facilitate the interaction between the multimedia component 808 and the first processing component 802.
[0101] The first memory 804 is configured to store various types of data to support the operation of the electronic device 800. Examples of such data include instructions for any application or method operating on the electronic device 800, contact data, phone book data, messages, pictures, videos, and the like. The first memory 804 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disk.
[0102] The first power component 806 provides power to various components of the electronic device 800. The first power component 806 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the electronic device 800.
[0103] The multimedia component 808 includes a screen that provides an output interface between the electronic device 800 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can sense not only the boundaries of the touch or swipe actions but also detect the duration and pressure associated with the touch or swipe operations. In some embodiments, the multimedia component 808 includes a front camera and / or a rear camera. When the electronic device 800 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera can receive external multimedia data. Each of the front camera and the rear camera can be a fixed optical lens system or have focal length and optical zoom capabilities.
[0104] The audio component 810 is configured to output and / or input audio signals. For example, the audio component 810 includes a microphone (MIC) that is configured to receive external audio signals when the electronic device 800 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signals can be further stored in the first memory 804 or transmitted via the communication component 816. In some embodiments, the audio component 810 further includes a speaker for outputting audio signals.
[0105] The first input / output interface 812 provides an interface between the first processing component 802 and a peripheral interface module, which can be a keyboard, a click wheel, buttons, etc. These buttons can include, but are not limited to: a home button, a volume button, a power-on button, and a lock button.
[0106] The sensor assembly 814 includes one or more sensors for providing an assessment of the status of various aspects of the electronic device 800. For example, the sensor assembly 814 can detect the on / off state of the electronic device 800, the relative positioning of components, such as the display and keypad of the electronic device 800. The sensor assembly 814 can also detect a change in the position of the electronic device 800 or a component of the electronic device 800, the presence or absence of user contact with the electronic device 800, the orientation or acceleration / deceleration of the electronic device 800, and the temperature change of the electronic device 800. The sensor assembly 814 can include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor assembly 814 can also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor assembly 814 can also include an acceleration sensor, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.
[0107] The communication component 816 is configured to facilitate communication between the electronic device 800 and other devices in a wired or wireless manner. The electronic device 800 can access a wireless network based on communication standards, such as WiFi, 2G, or 3G, or a combination thereof. In an exemplary embodiment, the communication component 816 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 816 further includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.
[0108] In an exemplary embodiment, the electronic device 800 can be implemented by one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components for performing the above-described speech recognition method.
[0109] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a first memory 804 including instructions that can be executed by a first processor 820 of the electronic device 800 to complete the above-described speech recognition method. For example, the non-transitory computer-readable storage medium can be a ROM, a random access first memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device, etc.
[0110] In another exemplary embodiment, a computer program product is also provided. The computer program product includes a computer program executable by a programmable device. The computer program has a code portion for performing the above-described speech recognition method when executed by the programmable device.
[0111] Figure 12 FIG. 4 is a block diagram of an electronic device 1900 shown according to an exemplary embodiment. For example, the electronic device 1900 may be provided as a server. Referring to Figure 12 , the electronic device 1900 includes a second processing component 1922, which further includes one or more processors, and memory resources represented by a second memory 1932 for storing instructions executable by the second processing component 1922, such as application programs. The application programs stored in the second memory 1932 may include one or more modules each corresponding to a set of instructions. In addition, the second processing component 1922 is configured to execute instructions to perform the above-described speech recognition method.
[0112] The electronic device 1900 may further include a second power supply component 1926 configured to perform power management of the electronic device 1900, a wired or wireless network interface 1950 configured to connect the electronic device 1900 to a network, and a second input / output interface 1958. The electronic device 1900 may operate based on an operating system stored in the memory 1932, such as Windows Server TM , Mac OS X TM , Unix TM , Linux TM , FreeBSD TM or the like.
[0113] Those skilled in the art can also understand that the various illustrative logical blocks and steps listed in the embodiments of the present application can be implemented by electronic hardware, computer software, or a combination of both. Whether such a function is implemented by hardware or software depends on the specific application and the design requirements of the entire system. Those skilled in the art can use various methods to implement the described function for each specific application, but such implementation should not be construed as exceeding the scope of protection of the embodiments of the present application.
[0114] It should be understood that unless otherwise specifically stated, the features of some embodiments of the various disclosures described herein can be combined with each other. As used herein, the term "and / or" includes any one of the related listed items and any combination of any two or more of them; similarly, "at least one of... " includes any one of the related listed items and any combination of any two or more of them.
[0115] Although terms such as "first", "second", and "third" may be used herein to describe various components, parts, regions, layers, or sections, these components, parts, regions, layers, or sections are not limited to these terms. On the contrary, these terms are only used to distinguish one component, part, region, layer, or section from another. Thus, the first component, part, region, layer, or section mentioned in the examples described herein may also be referred to as the second component, part, region, layer, or section without departing from the teachings of the various examples. Additionally, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include at least one of such features. In the description herein, the meaning of "a plurality" is at least two, such as two, three, etc., unless otherwise specifically defined.
[0116] Furthermore, the word "exemplary" is used herein to mean serving as an example, instance, or illustration. Any aspect or design described herein as "exemplary" is not necessarily to be construed as advantageous over other aspects or designs. Instead, the use of the word exemplary is intended to present concepts in a concrete fashion. As used herein, the term "or" is intended to mean an inclusive "or" rather than an exclusive "or". That is, unless otherwise specified or clear from the context, "X applies A or B" is intended to mean any of the natural inclusive permutations. That is, if X applies A; X applies B; or X applies both A and B, then "X applies A or B" is satisfied in any of the foregoing instances. Additionally, unless otherwise specified or clear from the context indicating a singular form, the articles "a" and "an" as used in this application and the appended claims are generally understood to mean "one or more".
[0117] Likewise, although the present disclosure has been shown and described with respect to one or more implementations, equivalent variations and modifications will occur to those skilled in the art upon reading and understanding this specification and the drawings. The present disclosure includes all such modifications and variations and is limited only by the scope of the claims. Specifically with respect to the various functions performed by the components described above (e.g., elements, resources, etc.), unless otherwise indicated, the terms used to describe such components are intended to correspond to any component (functionally equivalent) that performs the specific function of the described component, even if not structurally equivalent to the disclosed structure. Additionally, although a particular feature of the present disclosure may have been disclosed with respect to only one of several implementations, such a feature may, as may be desired and advantageous for any given or particular application, be combined with one or more other features of other implementations. Further, with respect to the use of "comprises," "comprising," "has," "having," "includes," or "including" in the detailed description or claims, such terms are intended to be inclusive in a manner similar to the term "including."
[0118] Other embodiments of the present disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include known or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of the present disclosure are pointed out by the appended claims.
[0119] It should be understood that the present disclosure is not limited to the exact structures described above and shown in the drawings, and various modifications and changes may be made without departing from its scope. The scope of the present disclosure is limited only by the appended claims.
Claims
1. A speech recognition method, characterized in that: include: Acquire a speech to be recognized, wherein the speech to be recognized includes a wake-up speech and a target speech; Input the speech to be recognized into a pre-trained speech recognition model to obtain a predicted text, wherein the predicted text is a text used to characterize the speech content of the target user in the target speech, and the target user is the speaker of the wake-up speech; Among them, the speech recognition model includes a speech encoder, which is used to obtain a first feature vector including semantic features and speaker features based on the speech to be recognized, and the first feature vector is used to determine the predicted text. The speech encoder is obtained by self-supervised pre-training based on at least one single-person speech sample.
2. The speech recognition method according to claim 1, characterized in that: The speech recognition method further comprises: Obtain at least one positive sample and / or at least one negative sample; Using the positive samples and / or the negative samples to train the initial speech recognition model until a training end condition is met; Each of the positive samples includes a first voice sample, a text used to characterize the content of a single-person voice sample in the first voice sample, and a first wake-up voice sample, and the first wake-up voice sample and the single-person voice sample in the first voice sample correspond to the same speaker; Each of the negative samples includes a second voice sample, an empty text sample, and a second wake-up voice sample, wherein the second wake-up voice sample and the single-person voice sample in the second voice sample correspond to different speakers.
3. The speech recognition method according to claim 2, characterized in that: The speech recognition method further comprises at least one of the following: Superimposing the single-person voice sample and the first interfering sound to obtain the first voice sample; The single-person voice sample and the second interfering sound are superimposed to obtain the second voice sample.
4. The speech recognition method according to claim 3, characterized in that: The first interfering sound and / or the second interfering sound is an interfering speech, and the interfering speech and the single-person speech sample correspond to different speakers.
5. The speech recognition method according to any one of claims 2 to 4, characterized in that: The using the positive sample and / or the negative sample to train the initial speech recognition model until a training end condition is met includes at least one of the following: Using the positive samples to perform a first phase of training on the initial speech recognition model until a model loss value of the speech recognition model is less than a first loss threshold; The speech recognition model that has completed the first stage of training is trained in the second stage using the positive samples and the negative samples until the model loss value of the speech recognition model is less than a second loss threshold.
6. The speech recognition method according to any one of claims 2 to 5, characterized in that: The using the positive sample and / or the negative sample to train the initial speech recognition model includes: The initial speech recognition model is trained according to a preset prompt instruction and using the positive sample and / or the negative sample, wherein the prompt instruction is used to indicate a target task of the speech recognition model.
7. The speech recognition method according to any one of claims 2 to 6, characterized in that: The speech recognition method further comprises: Before using the positive samples and / or the negative samples to train the initial speech recognition model, the large language model in the speech recognition model is trained according to at least one single-person speech recognition data, wherein the single-person speech recognition data includes the single-person speech sample and text used to characterize the content of the single-person speech sample.
8. The speech recognition method according to any one of claims 1 to 7, characterized in that: The speech recognition model also includes a linear transformation layer, a Transformer mapping layer, a word segmenter and a large language model; The step of inputting the speech to be recognized into a pre-trained speech recognition model to obtain a predicted text includes: Inputting the speech to be recognized into the speech encoder to obtain the first feature vector; Mapping the first feature vector to the feature space of the large language model using the linear transformation layer to obtain a second feature vector; Using the Transformer mapping layer to align the second feature vector in the time dimension to obtain a third feature vector; Processing a preset prompt instruction by using the word segmenter to obtain a prompt instruction feature vector, wherein the prompt instruction is used to indicate a target task of the speech recognition model; The third feature vector and the prompt instruction feature vector are input into the large language model to generate the predicted text.
9. A speech recognition device, characterized in that: include: An acquisition module, used to acquire a speech to be recognized, wherein the speech to be recognized includes a wake-up speech and a target speech; A prediction module, used for inputting the speech to be recognized into a pre-trained speech recognition model to obtain a predicted text, wherein the predicted text is a text used to characterize the speech content of the target user in the target speech, and the target user is the speaker of the wake-up speech; Among them, the speech recognition model includes a speech encoder, which is used to obtain a first feature vector including semantic features and speaker features based on the speech to be recognized, and the first feature vector is used to determine the predicted text. The speech encoder is obtained by self-supervised pre-training based on at least one single-person speech sample.
10. An electronic device, characterized in that: include: processor; a memory for storing processor-executable instructions; The processor is configured to execute the executable instructions in the memory to implement the steps of the speech recognition method according to any one of claims 1 to 8.
11. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the speech recognition method according to any one of claims 1 to 8 are implemented.
12. A computer program product, characterized in that The invention comprises a computer program, which, when executed by a processor, implements the steps of the speech recognition method according to any one of claims 1 to 8.
Citation Information
Cited By
Voiceprint recognition method and system for vehicle-mounted child safety seat
CN122050367A