A speech recognition method and computing device

CN122598652APending Publication Date: 2026-08-18XFUSION DIGITAL TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610578176.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-04-28
Publication Date
2026-08-18

AI Technical Summary

Technical Problem

[0004]但是,现实场景中用户常处于会议讨论、公共交通、商场喧嚣等多人混杂语音环境,此时传统语音识别功能对多人语音交叠与突发噪声的处理性能显著下降,识别准确率大幅波动,且易出现语音丢失、声源混淆及文本错乱等问题

Benefits of technology

[0013]Based on the embodiments of this application, the computing device can effectively distinguish between human voices and ambient sounds, and can further subdivide different speakers in human voices and specific sound source types in ambient sounds, providing a category basis for subsequent audio segment separation, time alignment and differential processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122598652A_ABST
    Figure CN122598652A_ABST
Patent Text Reader

Abstract

Embodiments of the present application relate to the technical field of sound recognition, and provide a speech recognition method and a computing device. The method comprises: acquiring mixed audio containing at least two sound sources; splitting the mixed audio into a plurality of audio segment sequences by using a diffusion model; recognizing a sound type in a non-human voice segment; recognizing a target audio segment based on a speech recognition model and a historical recognition result, to convert the target audio segment into a text sequence; and outputting the text sequence and the sound type corresponding to the non-human voice segment. According to the embodiments of the present application, the computing device introduces a diffusion model to perform fine-grained sound source separation, thereby improving the purity of speech signals in a complex audio environment. In addition, by fusing multi-modal context information such as historical recognition results, adjacent audio segments and environmental sound types, the speech recognition model can enhance the understanding ability and recognition accuracy of context semantics, thereby improving the overall processing efficiency and recognition accuracy of the speech recognition process.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of voice recognition technology, and in particular to a voice recognition method and computing device. Background Technology

[0002] Voice recognition refers to the ability of smart devices to convert speech signals into text. As the core interaction point between users and smart devices, the application scenarios of voice recognition are becoming increasingly widespread and complex with the popularization of smart devices.

[0003] A typical voice recognition function is to recognize a single human voice in a relatively quiet environment, such as when a user gives commands to a smart speaker at home.

[0004] However, in real-world scenarios, users are often in noisy environments with multiple voices, such as meetings, public transportation, and shopping malls. In these situations, the traditional speech recognition function's performance in handling overlapping voices and sudden noises decreases significantly, the recognition accuracy fluctuates greatly, and problems such as speech loss, sound source confusion, and text corruption are prone to occur. Summary of the Invention

[0005] This application provides a speech recognition method and computing device. By introducing a diffusion model to accurately split mixed audio, the method separates human voices from environmental sounds in the mixed audio, effectively improving the accuracy of speech recognition in complex environments.

[0006] To achieve the above objectives, the embodiments of this application adopt the following technical solutions: In a first aspect, embodiments of this application provide a speech recognition method, comprising: acquiring a mixed audio containing at least two sound sources; performing multi-source separation on the mixed audio using a diffusion model to generate multiple temporally complete and mutually independent audio segment sequences; distinguishing between human voice segments and non-human voice segments in the audio segment sequences; identifying a target audio segment based on a speech recognition model and historical recognition results, so as to convert the target audio segment into a text sequence; the target audio segment may be any human voice segment in the mixed audio, and the historical recognition results include the recognized text generated by the speech recognition model based on human voice segments whose start time is located before the start time corresponding to the target audio segment, as well as the acoustic features or mixed audio of human voice segments whose start time is located before the start time corresponding to the target audio segment; and outputting the text sequence.

[0007] Based on the embodiments of this application, the computing device improves the purity of speech signals in complex audio environments by introducing a diffusion model for fine-grained sound source separation. At the same time, by fusing multimodal contextual information such as historical recognition results, adjacent human voice segments, and environmental sound types, the speech recognition model's ability to understand contextual semantics and its recognition accuracy are enhanced. Especially in scenarios with multiple speakers alternating and varying background noise, it can provide more coherent and realistic speech transcription results, improving overall processing efficiency and recognition accuracy.

[0008] In one possible implementation, a diffusion model is used to separate multiple sound sources in the mixed audio, generating multiple temporally complete and independent audio segment sequences. This includes: sampling the mixed audio using the diffusion model to obtain multiple sample segments; determining the sound source category corresponding to each sample segment; the sound source category is used to identify the sound source identity of each sound source corresponding to the mixed audio; and performing a denoising operation on the mixed audio based on each sound source category to obtain at least one audio segment sequence corresponding to each sound source category. Wherein, when there is an overlap of audio segments corresponding to different sound source categories in the mixed audio on the time axis, the denoising operation generates a complete signal for each sound source category within the overlapping time period.

[0009] Based on the embodiments of this application, the computing device can accurately identify the audio features of different speakers and the surrounding environment, enabling it to accurately split mixed audio using a diffusion model, separating clear human voice segments and non-human voice segments from the mixed audio, and providing high-quality input for subsequent speech recognition.

[0010] In another possible implementation, distinguishing between human voice segments and non-human voice segments in an audio segment sequence includes: distinguishing between human voice segments and non-human voice segments in an audio segment sequence based on the sound source category; after distinguishing between human voice segments and non-human voice segments in an audio segment sequence, the method further includes: identifying the sound type in the non-human voice segments; and when outputting the text sequence, outputting the sound type corresponding to the non-human voice segments.

[0011] Based on the embodiments of this application, after the computing device identifies the sound type corresponding to a non-human voice segment, it can output the sound type as additional information synchronously, which makes it easier for users to understand the audio scene context more clearly, and also provides more scene reference information for the speech recognition model, helping the model to understand the current audio content more accurately.

[0012] In another possible implementation, determining the sound source category corresponding to each sample segment includes: extracting audio features corresponding to multiple sample segments; clustering the audio features of multiple sample segments to obtain at least one audio feature cluster; and determining the sound source category based on the center vector of each audio feature cluster.

[0013] Based on the embodiments of this application, the computing device can effectively distinguish between human voices and ambient sounds, and can further subdivide different speakers in human voices and specific sound source types in ambient sounds, providing a category basis for subsequent audio segment separation, time alignment and differential processing.

[0014] In another possible implementation, after using a diffusion model to separate multiple sound sources from the mixed audio and generate multiple temporally complete and independent audio segment sequences, the method further includes: obtaining the first start and end times corresponding to the mixed audio and the second start and end times corresponding to each audio segment sequence; and performing temporal alignment processing on the audio segment sequences based on the first start and end times and the second start and end times.

[0015] Based on the embodiments of this application, the computing device ensures that each audio segment accurately restores the original speaking rhythm and dynamic changes of environmental sound on the timeline by identifying the sound source category, recording time information, and processing time sequence alignment. This provides the speech recognition model with spatiotemporally consistent input, thereby significantly improving the temporal coherence and semantic integrity of the speech recognition results.

[0016] In another possible implementation, identifying the sound type in a non-human voice segment includes: extracting the audio features corresponding to the non-human voice segment; calculating the similarity between the audio features corresponding to the non-human voice segment and reference features in the environmental sound category library based on a preset environmental sound category library; the reference features are used to represent the audio features of environmental sounds in each category in the environmental sound category library; and determining the sound type corresponding to the non-human voice segment based on the similarity.

[0017] Based on the embodiments of this application, the computing device can accurately identify the specific sound type of non-human voice segments, thereby providing detailed environmental sound information for subsequent environmental sound event analysis, scene understanding, or optimization of speech recognition results.

[0018] In another possible implementation, based on a speech recognition model and historical recognition results, a target audio segment is identified to convert the target audio segment into a text sequence. This includes: reading the acoustic features of the target audio segment; retrieving semantic features of text matching the acoustic features from historical recognition results based on the acoustic features; determining non-human voice segments located within the third start and end time of the target audio segment according to the third start and end time; fusing the acoustic features, semantic features, and non-human voice segments located within the third start and end time to obtain the fused features corresponding to the target audio segment; and generating the text sequence corresponding to the target audio segment based on the fused features.

[0019] Based on the embodiments of this application, the computing device utilizes the multi-source information fusion capability formed by the speech recognition model during the training phase to achieve deep interaction between acoustic features, historical semantic features and environmental audio features, thereby achieving accurate recognition of target audio segments and improving the accuracy and anti-interference capability of speech recognition.

[0020] In another possible implementation, the method further includes: acquiring a training audio signal; the training audio signal includes mixed speech samples corresponding to multiple speakers and clean speech samples corresponding to each speaker; performing iterative noise addition processing on the training audio signal to generate a noise-added frequency signal with increasing noise intensity; training an initial diffusion model based on the noise-added frequency signal and the training audio signal; and updating the parameters of the initial diffusion model based on the loss function corresponding to the initial diffusion model to obtain the diffusion model.

[0021] Based on the embodiments of this application, the computing device constructs a training dataset containing mixed speech from multiple speakers and clean speech, and uses an iterative noise-adding method to simulate mixed audio under different noise scenarios. This enables the initial diffusion model to learn during the training process the ability to gradually recover clean audio features from the noise-adding frequency signal with high noise intensity, thereby improving the computing device's ability to separate different speech components in mixed audio.

[0022] In another possible implementation, the method further includes: acquiring a training speech signal, training text corresponding to the training speech signal, a reference speech signal, and a reference text corresponding to the reference speech signal; the reference speech signal includes a speech signal aligned with the contextual features of the training speech signal; using an initial speech recognition model to recognize the training speech signal to generate an initial recognition result; updating the parameters of the initial speech recognition model based on the semantic differences between the initial recognition result and the training text; and performing context-aware training on the updated speech recognition model based on the training speech signal, the reference speech signal, and the reference text.

[0023] Based on the embodiments of this application, the computing device can use the context signal corresponding to the training speech signal as the training input to complete the training and fine-tuning of the speech recognition model, so that it has the ability to convert speech to text and powerful context-aware optimization capabilities, laying the foundation for high-precision speech recognition in practical applications.

[0024] Secondly, embodiments of this application also provide a speech recognition system, comprising: an audio acquisition unit configured to acquire mixed audio containing at least two sound sources; a diffusion model unit configured to perform multi-source separation on the mixed audio to generate multiple temporally complete and mutually independent audio segment sequences; a sound classification unit configured to distinguish between human voice segments and non-human voice segments in the audio segment sequences; a speech recognition unit configured to recognize a target audio segment based on a speech recognition model and historical recognition results, so as to convert the target audio segment into a text sequence; the target audio segment may be any human voice segment in the mixed audio, and the historical recognition results include the recognized text generated by the speech recognition model based on human voice segments whose start time is located before the start time corresponding to the target audio segment, and the acoustic features or mixed audio of human voice segments whose start time is located before the start time corresponding to the target audio segment; and an output unit configured to output the text sequence.

[0025] Thirdly, embodiments of this application also provide a computing device, including: a processor and a memory; the processor and the memory are coupled; the memory is used to store program instructions; the processor is used to execute the program instructions to perform the method as described in any of the first aspects above.

[0026] Fourthly, embodiments of this application provide a chip for performing the methods described in any of the first aspects above.

[0027] Fifthly, embodiments of this application provide a computer-readable storage medium storing computer-executable instructions, which, when executed by a computer, implement the method as described in any of the first aspects.

[0028] In a sixth aspect, embodiments of this application provide a program product including a computer program that, when executed by a processor, implements the method as described in any of the first aspects. Attached Figure Description

[0029] Figure 1 This is a schematic diagram of the structure of a computing device provided in an embodiment of this application; Figure 2 A schematic diagram illustrating a scenario in which a computing device performs a speech recognition task, provided as an embodiment of this application; Figure 3 A schematic diagram of mixed audio obtained by a computing device according to an embodiment of this application; Figure 4 A flowchart illustrating a speech recognition method provided in an embodiment of this application; Figure 5 This is a schematic diagram of a diffusion model for splitting mixed audio, provided in an embodiment of this application. Figure 6A schematic diagram of the training process of a diffusion model provided in an embodiment of this application; Figure 7 A schematic diagram illustrating the principle of a training model provided in an embodiment of this application; Figure 8 This application provides a schematic diagram of a process for splitting and mixing audio in an embodiment. Figure 9 A flowchart illustrating the process of distinguishing between human voice segments and non-human voice segments, provided for an embodiment of this application; Figure 10 A schematic diagram of a process for determining the category of a sound source provided in an embodiment of this application; Figure 11 This application provides a schematic diagram of a process for identifying non-human voice segment sound types in an embodiment of the present application. Figure 12 A schematic diagram illustrating the training process of a speech recognition model provided in an embodiment of this application; Figure 13 A schematic diagram of a speech recognition model provided in an embodiment of this application; Figure 14 A schematic diagram of an audio segment recognition process provided in an embodiment of this application; Figure 15 This is a schematic diagram of the structure of a speech recognition system provided in an embodiment of this application; Figure 16 This is a schematic diagram of a computing device provided in an embodiment of this application. Detailed Implementation

[0030] The technical solutions of the embodiments of this application will now be described with reference to the accompanying drawings. To facilitate a clear description of the technical solutions of the embodiments of this application, the use of terms such as "first," "second," etc., in the embodiments of this application is for illustrative purposes and to distinguish the objects being described. There is no particular order between them, nor does it indicate a specific limitation on the number of devices in the embodiments of this application, and they do not constitute any limitation on the embodiments of this application.

[0031] The application scenarios of this application will be explained below.

[0032] Figure 1 This is a schematic diagram of the structure of a computing device provided in an embodiment of this application. Figure 2 This is a schematic diagram illustrating a scenario in which a computing device performs a speech recognition task, as provided in an embodiment of this application. Figure 3 This is a schematic diagram of mixed audio obtained by a computing device according to an embodiment of this application. The following is based on... Figures 1 to 3 The content shown is an exemplary illustration of the scenarios involved in speech recognition in the embodiments of this application.

[0033] like Figure 1As shown, the computing device 100 may include a processor unit 101 and an audio acquisition device 102. In this embodiment, the processor unit 101 can be used to process speech recognition-related operations; the audio acquisition device 102 is disposed on the housing of the processor unit 101 or detachably connected to the processor unit 101, and is used to acquire mixed audio signals in the actual scene.

[0034] In one example, the audio acquisition device 102 may be a microphone array with multi-channel synchronous sampling capability, which can effectively capture audio information in the scene. The processor unit 101 may be a central processing unit (CPU), a graphics processing unit (GPU), or an application-specific integrated circuit (ASIC). Based on the processor unit 101, the computing device 100 can provide users with automatic speech recognition (ASR) functionality, enabling accurate and efficient understanding and conversion of speech information.

[0035] It should be understood that, Figure 1 The structure of the computing device 100 shown is only an example; in actual applications, in addition to... Figure 1 In addition to the structure of the intelligent assistant device shown, the computing device 100 can be a user's intelligent portable device, such as a smartphone, tablet, or laptop, or it can be an electronic device such as a server or even a server cluster. The specific structure of the computing device 100 is not limited in the embodiments of this application.

[0036] like Figure 2 As shown, in some application scenarios, multiple sound sources may exist around the computing device 100, including voice input from users 110A and 110B, ambient noise source 120, and device operating noise. When performing a speech recognition task, the computing device 100 acquires a mixed audio signal containing the voices of users 110A and 110B and ambient noise source 120 (such as air conditioning operation sound and traffic sound outside the window) through the audio acquisition device 102.

[0037] like Figure 3 As shown, in the mixed audio signal acquired by the audio acquisition device 102, the speech waveforms corresponding to user 110A and user 110B overlap, and environmental noise is randomly distributed throughout the time-frequency domain of the mixed audio. Therefore, when recognizing this mixed audio, the computing device 100 needs to decouple and separate the overlapping speech components and reduce environmental noise interference to improve the accuracy of subsequent speech recognition.

[0038] However, traditional speech recognition functions cannot effectively handle complex scenarios where overlapping speech and noise coexist, leading to a significant increase in recognition error rates. Based on this, embodiments of this application provide a speech recognition method, system, and computing device. The computing device introduces a diffusion model to perform sound separation and context-aware training on the acquired mixed audio, thereby improving speech recognition performance in complex environments, avoiding the speech loss and text corruption problems that occur in traditional speech recognition in noisy environments with multiple speakers, and improving speech recognition accuracy.

[0039] Figure 4 This is a flowchart illustrating a speech recognition method provided in an embodiment of this application. The following is based on... Figure 4 The content shown is used to illustrate the process of speech recognition performed by the computing device in the embodiments of this application.

[0040] like Figure 4 As shown, the process of recognizing voice signals by the computing device in this embodiment may include the following steps S100 to S500.

[0041] S100: Acquire mixed audio containing at least two sound sources.

[0042] In step S100, after receiving a voice recognition command, the computing device can activate an audio acquisition device connected to its processor, such as a microphone array, to acquire mixed audio in the environment where the computing device is located in real time, as the raw audio input that the computing device needs to process.

[0043] In this context, a sound source refers to an object capable of producing sound. In the application scenario of this application, a sound source may include the voices of one or more speakers, as well as various non-human sound sources in the environment, such as the sound of air conditioning running, vehicle traffic, door and window opening and closing, and notification sounds from other electronic devices. Mixed audio refers to a mixed audio signal formed by superimposing sound signals generated by multiple sound sources. This includes the target human voice to be identified, various non-human sound signals as interference, and may also contain the voice signals of multiple speakers simultaneously.

[0044] In this embodiment of the application, the voice recognition command received by the computing device can be triggered by the user's active operation, such as the user pressing a physical button on the computing device, clicking the voice recognition icon in the graphical display interface of the computing device, or requesting the computing device to enable the voice recognition function through a preset wake-up word, such as "XX, please enable the voice recognition function".

[0045] In some embodiments, voice recognition commands can also be initiated by applications in the computing device based on business needs, such as a meeting recording application automatically starting voice recognition when it detects the start of a meeting.

[0046] Upon receiving a voice recognition command, the computing device verifies the user's identity and command permissions to ensure the command's legitimacy and reliability. After successful verification, the computing device activates its audio acquisition module, such as a built-in microphone or an external microphone array, to begin capturing mixed audio from the surrounding environment.

[0047] In some embodiments, because the mixed audio acquired by the computing device through the audio acquisition module is an analog signal, the computing device uses an analog-to-digital converter to convert the analog mixed audio into digital mixed audio after acquiring the mixed audio for subsequent processing. The digital mixed audio is usually stored at a specific sampling rate (such as 16kHz, 44.1kHz, etc.) and sampling bit depth (such as 16 bits) to form corresponding time-series data.

[0048] It should be understood that the acquisition of audio data by the computing device using the audio acquisition module is a continuous process. The acquisition process continues until the speech recognition command is explicitly terminated, or the computing device detects a speech acquisition end signal, such as prolonged silence or a "stop recognition" command issued by the user. In the embodiments of this application, the computing device can immediately stop audio acquisition after detecting the acquisition end signal, and organize all the acquired mixed audio into a continuous digital audio stream. The computing device can then use the built-in programs and models in its built-in processor, such as the central processing unit and / or image processor, to perform speech recognition processing on the received mixed digital audio.

[0049] In another embodiment, while the audio acquisition module is continuously acquiring audio, the computing device can segment the received mixed audio and, after the acquisition time reaches a preset threshold, such as 30 seconds or 1 minute, send the current segment of audio into the processor to perform speech recognition processing.

[0050] It should be noted that the form of the voice recognition command shown in this embodiment is only an example. In practical applications, the voice recognition command can directly carry the corresponding mixed audio, enabling the computing device to obtain the digital audio stream input from an external audio source, such as a recording file from a conference system, smart speaker, or cloud storage, via an interface or network connection. In this way, the computing device can also send the obtained external digital audio stream to the processor for corresponding voice recognition processing without calling its own connected audio acquisition module. This embodiment does not limit the specific method of obtaining the mixed audio.

[0051] S200: Uses a diffusion model to separate multiple sound sources in mixed audio, generating multiple temporally complete and independent audio segment sequences.

[0052] In step S200, after the computing device acquires the mixed audio using the audio acquisition module, it calls the pre-trained diffusion model in the processor to process the obtained mixed audio, thereby separating the multiple sound sources in the mixed audio and splitting the mixed audio into multiple temporally complete and mutually independent audio segment sequences.

[0053] In this context, "temporally complete" means that each separated audio segment sequence covers the complete duration of the original mixed audio from beginning to end, without being truncated into discontinuous segments due to the separation operation; "mutually independent" means that each audio segment sequence contains only the audio signal of a single sound source, thus separating overlapping and mixed audio signals from multiple sound sources. The diffusion model refers to a probabilistic generation-based deep learning model that transforms the original data distribution into a simple prior distribution by progressively adding Gaussian noise, and then reconstructs the target data through an inverse denoising process.

[0054] The computing device uses a diffusion model to split the audio into multiple audio segments, which can include human voice segments and non-human voice segments. Human voice segments refer to audio signals containing the user's speech content, while non-human voice segments include background noise, music, mechanical sounds, and other non-speech components.

[0055] In this embodiment, the pre-trained diffusion model configured in the processor of the computing device can first sample the input mixed audio in the time domain, extracting multiple sample segments with a certain time length from the continuous audio stream. These sample segments may contain mixed information from different sound sources, such as simultaneous speaker speech and background environmental noise.

[0056] Then, the diffusion model will determine the sound source category for each sample segment. By analyzing the spectral characteristics, energy distribution, and timbre features of the sample segment, it will determine the identity of its main corresponding sound source, such as distinguishing different sound source categories such as the speech of speaker A, the speech of speaker B, the noise of air conditioner operation, and the sound of keyboard typing.

[0057] After identifying the sound source category of each sample segment, the diffusion model performs specific denoising and separation operations in parallel for each sound source category. The denoising and separation process performed by the diffusion model is equivalent to gradually removing noise and interference components unrelated to the target sound source category from the original mixed audio, thereby extracting the signal of each sound source category in the original mixed audio signal into an independent audio segment sequence, with each audio segment sequence corresponding to a specific sound source category.

[0058] S300: Distinguishes between human voice segments and non-human voice segments in an audio segment sequence.

[0059] In step S300, after the computing device completes the splitting of the mixed audio using the diffusion model, the computing device can distinguish between human voice segments and non-human voice segments in the audio segment sequence, so that the computing device can obtain the audio segments that need to be processed differently.

[0060] Among them, a human voice segment refers to an audio segment containing the voice of a speaker, whose spectral characteristics and energy distribution conform to the physiological characteristics of human vocalization, and carries the speech content that needs to be identified; a non-human voice segment is any independent audio segment that does not contain human speech and exists only as environmental interference.

[0061] In this embodiment, the computing device performs a recognition operation on the human voice segments in the audio segment sequence in a subsequent process to obtain the corresponding text sequence; while the non-human voice segments in the audio segment sequence are regarded as auxiliary parameters for recognizing human voice segments, as well as environmental noise signals. Therefore, the computing device can distinguish between human voice segments and non-human voice segments in the audio segment sequence by differentiating the acoustic features of audio segments corresponding to different sound sources in the audio segment sequence.

[0062] Figure 5 This is a schematic diagram of a diffusion model for splitting mixed audio, provided in an embodiment of this application.

[0063] In one example, a computing device can distinguish between human and non-human voice segments during the process of splitting mixed audio using a diffusion model. For example... Figure 5 As shown, the diffusion model is used to... Figure 2 Taking the mixed audio shown in the figure as an example, the diffusion model, through sampling and recognition, determines that there are three types of sound sources in the mixed audio: the voice of user 110A, the voice of user 110B, and environmental noise. The diffusion model can then perform three denoising processes in parallel based on the original mixed audio to extract the clean voice segments of user 110A, the clean voice segments of user 110B, and the environmental noise segments, respectively, as the audio segment sequence obtained from the original mixed audio.

[0064] Furthermore, after obtaining multiple audio segment sequences, the computing device can also split audio signals in the same audio segment sequence whose silence time exceeds a preset threshold into two audio segments based on the start and end times of the silence time, according to the silence time in the audio segment sequence corresponding to the same sound source.

[0065] The silence period refers to the duration during which the audio energy within the segment is below a preset energy threshold, corresponding to the gaps in the speaker's speech. This segmentation allows for multiple segments of speech from the same speaker at different times, facilitating separate recognition processing for each segment and improving recognition accuracy. In this embodiment, for each segment of speech obtained through segmentation, the computing device can label its corresponding sound source identity based on its corresponding audio segment sequence, facilitating the subsequent output of recognition results organized by sound source.

[0066] For example, such as Figure 5 As shown, the computing device can further split the audio segment sequence corresponding to user 110A into two independent human voice segments, thereby removing the gaps caused by user 110A pausing to speak from the audio segments, in order to optimize the subsequent audio segment recognition process.

[0067] In some embodiments, after obtaining the audio segment sequence, the computing device can also identify the start and end times of the silence time in each audio segment sequence, thereby determining the start and end times of the audio segments contained in each audio segment sequence based on the overall start and end times of the audio segment sequence, and then combining the corresponding sound source types of each audio segment sequence to obtain the corresponding human voice segments and non-human voice segments.

[0068] In this way, when the computing device uses the diffusion model to perform audio decomposition, it can avoid simply and crudely separating the time-frequency aliasing effect between different sound sources, and avoid excessive smoothing or loss of the feature information of different sound sources. Thus, while preserving the naturalness of the speech and the realism of the environment, it can achieve high-fidelity, fine-grained sound source separation.

[0069] In this embodiment, after the computing device completes audio segmentation using a diffusion model, it can determine the semantic role of the audio segment based on the sound source categories classified by the diffusion model, thereby providing corresponding semantic labels for subsequent speech recognition. For example, the computing device can, based on the sound source categories classified by the diffusion model, identify audio segments containing user speech as human voice segments and audio segments not containing user speech as non-human voice segments, so that human voice segments contain the speaker's voice information, while non-human voice segments contain other environmental sounds besides human voice. The computing device can perform different processing methods on human voice segments and non-human voice segments in subsequent processes to optimize the subsequent speech recognition effect.

[0070] S400: Based on a speech recognition model and historical recognition results, it identifies target audio segments to convert them into text sequences.

[0071] In step S400, after the computing device uses a diffusion model to separate human voice segments and non-human voice segments from the mixed audio, it can sequentially perform speech recognition processing on each human voice segment to obtain the text transcription result corresponding to each human voice segment. In this embodiment, when the computing device uses the speech recognition model to recognize human voice segments, it can combine the historical recognition results of the speech recognition model to dynamically model the contextual semantics of the current human voice segment to improve the recognition accuracy.

[0072] The target audio segment identified by the speech recognition model may include any human voice segment in the mixed audio. The historical recognition results include the recognized text generated by the speech recognition model based on the human voice segments whose start time is earlier than the start time corresponding to the target audio segment, as well as the acoustic features or corresponding mixed audio signals of the human voice segments whose start time is earlier than the start time corresponding to the target audio segment.

[0073] For example, when a computing device uses a speech recognition model to recognize human voice segments, it determines the corresponding speech recognition order and the contextual position of each human voice segment in the mixed audio based on the time sequence of each human voice segment. In this way, when recognizing a target audio segment, it calls the speech recognition results before and after the target audio segment to construct a dynamic context and improve the accuracy and coherence of contextual understanding.

[0074] In this embodiment, when the computing device uses a speech recognition model to recognize a target audio segment, in addition to calling historical recognition results, it can also determine non-human voice segments within the duration of the target audio segment, as well as other voice segments before and after the target audio segment, based on the start and end times corresponding to the target audio segment. The computing device can input the non-human voice segments, other voice segments before and after the target audio segment, and historical recognition results together as multimodal context into the speech recognition model to help the speech recognition model more accurately infer the current speaker's intention and emotional tendency, thereby improving the semantic accuracy of speech recognition.

[0075] In this way, computing devices can use human voice segments and non-human voice segments adjacent to the target audio segment as contextual clues, and combine them with historical recognition results to construct context, thereby improving semantic coherence in long speech and multi-speaker scenarios.

[0076] It should be noted that the computing device can process different types of sound segments in parallel when processing human voice segments and non-human voice segments, i.e., when executing steps S300 and S400. For example, while the computing device uses the sound classification unit to identify the sound type in the non-human voice segment, it can call the speech recognition model to process the human voice segment to improve the overall processing efficiency.

[0077] Once the non-human voice segment is identified, the computing device can feed back the environmental sound type information to the ongoing or subsequent speech recognition model as supplementary contextual information to further optimize the recognition process of the target audio segment. For example, if the non-human voice segment is identified as a "car horn," the speech recognition model, when processing human voice segments from the same or similar time periods, can anticipate potential distortion caused by sudden high-decibel noise from the outside world and perform targeted compensation or error-tolerant processing on the corresponding acoustic features during the decoding process.

[0078] S500: Output text sequence.

[0079] In step S500, after identifying the text sequence corresponding to each target audio segment, the computing device integrates this information and outputs it to the user or relevant application in a preset format. For example, the computing device can organize the text sequence according to the time sequence of the voice segments and the corresponding identifiers of the speakers, such as marking the text sequence with identifiers like Speaker A, Speaker B, etc., or distinguishing them by different colors, fonts, etc., so that the user can clearly understand the content of different speakers' speech.

[0080] In some embodiments, the computing device can also store the integrated text sequence and sound type information as a text file, send it to a designated storage location, or display it in real time on the graphical display interface of the computing device to meet the output needs of different application scenarios.

[0081] Through steps S100 to S500, the computing device can achieve a complete speech recognition process, from receiving speech recognition commands, acquiring and processing mixed audio, using a diffusion model for sound source separation, identifying environmental sound types, combining multimodal context for speech-to-text conversion, to finally integrating and outputting a text sequence and environmental sound type information. This speech recognition process improves the purity of speech signals in complex audio environments by introducing a diffusion model for fine-grained sound source separation. Simultaneously, by fusing historical recognition results, adjacent voice segments, and multimodal contextual information such as environmental sound types, it enhances the speech recognition model's ability to understand contextual semantics and its recognition accuracy. Especially in scenarios with multiple speakers alternating and varying background noise, it can provide more coherent speech-to-text results that are closer to the real-world context, improving overall processing efficiency and recognition accuracy.

[0082] Figure 6 This is a schematic diagram of the training process of a diffusion model provided in an embodiment of this application. Figure 7 This is a schematic diagram illustrating the principle of a training model provided in an embodiment of this application. The following is in conjunction with... Figure 6 and Figure 7The content shown here is an exemplary illustration of the training process of the diffusion model preset by the processor in the computing device in the embodiments of this application.

[0083] like Figure 6 As shown in the embodiments of this application, the training process of the diffusion model by the computing device may include the following steps S610 to S640.

[0084] S610: Acquire training audio signals.

[0085] In step S610, before training the diffusion model, the computing device can first acquire the corresponding training audio signal as training data for the initial diffusion model to learn and optimize. The training audio signal includes mixed speech samples corresponding to multiple speakers and clean speech samples corresponding to each speaker. The initial diffusion model refers to a diffusion model structure that has not been trained or has only undergone minimal training, and whose parameters have not yet been optimized to a stable state.

[0086] like Figure 7 As shown, during the training phase, the diffusion model includes a forward diffusion process and a reverse diffusion process. The forward diffusion process is a noise-adding process, while the reverse diffusion process is a noise-removing process. For example, from X0 to X... T The process is a forward diffusion process, caused by X T The process to X0 is a reverse diffusion process.

[0087] Where X0 represents the initial clean speech signal, which in this embodiment can be the training audio signal obtained by the computing device; X T This represents the fully noise-added mixed audio, where T refers to the total number of steps in the forward diffusion process, i.e., the maximum number of iterations in the diffusion process, and X... t This represents the intermediate mixed audio state after adding noise at step t, where t can be any integer between 1 and T.

[0088] After obtaining the training audio signals, the computing device can use each signal as X0 and complete the training of the diffusion model by executing subsequent steps S620 to S640.

[0089] It should be noted that when the computing device performs the forward diffusion process, it can select any speaker's clean or mixed speech sample from the training audio signal as X0, and gradually add Gaussian noise until it reaches X0. T In this embodiment of the application, no limitation is made on the number of speakers included in the training audio signal or the degree of speech overlap between multiple speakers.

[0090] S620: Performs iterative noise addition processing on the training audio signal to generate a noise-added audio signal with increasing noise intensity.

[0091] In step S620, the computing device performs a forward diffusion process on the training audio signal according to a preset noise scheduling strategy, that is, Gaussian noise is added to the training audio signal X0 layer by layer to obtain the noise-added frequency signal X at time t. t The noise-added audio signal refers to the audio signal obtained by iteratively adding Gaussian noise to the training audio signal.

[0092] For example, during the forward diffusion process, the computing device can add Gaussian noise ε to X0 layer by layer until it completely degenerates into pure noise X. T For example, the noise addition operation at each step in the forward diffusion process can be implemented using a preset noise addition formula. For instance, at step t, the computing device can add noise based on the current diffusion step number t and a preset noise figure β. t Gaussian noise ε is mixed into the audio signal X of the current state according to a certain proportion. t-1 In the process, the noise-added frequency signal X is obtained. t Among them, the noise figure β t The noise level can gradually increase with the number of diffusion steps t, resulting in lower noise intensity in the early stages, preserving more features of the training audio signal, and gradually increasing noise intensity in the later stages until the signal is completely covered by noise. Through this gradual noise addition method, the diffusion model can learn the complete degradation process from clean speech to pure noise, laying the foundation for subsequent reverse denoising and separation.

[0093] Based on noise figure β t Given the mixed audio X0, the computing device can derive the expression for the noise-added frequency signal Xt with any number of noise-added steps t, i.e., the following equation (1): (1); in, This refers to the signal preservation coefficient, which is related to the noise figure β. t They are related, and their sum is 1; This refers to the noise accumulation factor, which represents the proportion of noise introduced up to step t. It can be obtained through the signal preservation factor. The result is obtained by multiplying the products together. and These refer to signal weights and noise weights, respectively; while It is standard Gaussian noise.

[0094] Based on equation (1), X t It is a weighted superposition of the training audio signal X0 and the accumulated noise, with the weights dynamically determined by the temporal characteristics of the diffusion process. As t increases, Continuous decay Gradually approaching 1, X t Increasingly close to pure noise X T .

[0095] Based on the above process, the computing device can add noise to the training audio signal to generate a noisy sample set corresponding to the initial diffusion model, providing sufficient input parameters for subsequent inverse denoising modeling. This noisy sample set covers the full-scale degradation trajectory from slight distortion to complete noise, improving the generalization ability of the diffusion model in various noise environments.

[0096] S630: Train the initial diffusion model based on the noise-added frequency signal and the training audio signal.

[0097] In step S630, after the computing device performs the forward diffusion process and obtains the noisy sample set, it can add the noise frequency signal X from the noisy sample set. t The original training audio signal X0 is used as a supervision pair and input into the initial diffusion model. In this embodiment, the initial diffusion model uses X... t Given the conditional input, the objective is to predict the Gaussian noise ε added at the corresponding step t, and to optimize the parameters by minimizing the error between the predicted noise and the actual noise.

[0098] It should be understood that the training process involves inputting the supervised pairs from the noisy sample set one by one into the initial diffusion model, comparing the output of the initial diffusion model with the real noise labels, and then executing step S640 to continuously update the model parameters until the loss function converges to the preset threshold.

[0099] S640: Based on the loss function corresponding to the initial diffusion model, update the parameters of the initial diffusion model to obtain the diffusion model.

[0100] In step S640, the computing device can use the backpropagation algorithm and a preset loss function, such as the mean squared error loss function, to calculate the error value of the initial diffusion model when predicting noise, and adjust the model's weight parameters, bias parameters, etc., according to the error value.

[0101] The loss function is used to measure the difference between the predicted noise output by the initial diffusion model and the actual noise added during the forward diffusion process. The smaller the difference, the lower the loss value. The computing device iterates through training, that is, it repeatedly executes steps S620 to S640. In each iteration, samples in the training audio signal are randomly selected for noise addition, model prediction, and parameter updates until the value of the loss function stabilizes within a preset range or reaches a preset number of training rounds. At this time, the parameters of the initial diffusion model have been optimized to a better state, thus obtaining the trained diffusion model. For example, in the embodiments of this application, the mean square error corresponding to the initial diffusion model can be expressed as the following formula (2): (2); in, This refers to the initial diffusion model, whose parameter θ is continuously updated through the back diffusion process; L is the loss value, which is the difference between the model output and the actual noise. This refers to the output of the initial diffusion model, that is, the model's prediction of the noise added at the current time t; This represents the expectation for the training audio signal X0; This refers to the sum of squares of the differences between predicted noise and actual noise; this loss function guides the model to gradually approximate the actual noise distribution, thereby accurately removing noise at each stage during the reverse denoising process.

[0102] Through continuous optimization of equation (2), the model parameter θ gradually converges, enabling the model output to more accurately fit the real noise structure, and finally achieving the transformation from pure noise X to pure noise X in the back diffusion stage. T A stable mapping to the high-quality reconstructed training audio signal X0 is achieved.

[0103] Through steps S610 to S640, the computing device completes the training process of the diffusion model. The trained diffusion model can learn the complete degradation path from pure speech signal to pure noise and has the ability to infer the original pure speech signal from the noisy signal. In this embodiment, when the computing device receives mixed audio containing mixed speech and environmental noise, the diffusion model can perform reverse diffusion denoising processing on the input mixed audio based on the noise distribution pattern and denoising strategy learned during its training process, gradually stripping away noise components and separating clear human voice segments and non-human voice segments, providing high-quality input data for subsequent speech recognition and environmental sound type recognition.

[0104] Figure 8 This is a schematic diagram illustrating a process for splitting and mixing audio, provided as an embodiment of this application. The following is based on... Figure 8 The content shown is used to illustrate, by way of example, the process by which a computing device calls a diffusion model to process mixed audio in the embodiments of this application.

[0105] like Figure 8 As shown, the process of splitting mixed audio using a diffusion model by the computing device provided in this application embodiment may include the following steps S210 to S250.

[0106] S210: Use a diffusion model to sample from the mixed audio to obtain multiple sample segments.

[0107] In step S210, the computing device can use the diffusion model trained in steps S610 to S640 to sample the received target mixed audio, and collect multiple sample segments with preset durations, such as 2s, 3s, etc., each sample segment containing certain speech or environmental sound information.

[0108] In the embodiments of this application, when the computing device selects sample segments using the diffusion model, the obtained sample segments can be continuous or non-uniformly sampled based on the energy changes or other characteristics of the mixed audio, so as to ensure that the sound content with different time periods and different characteristics in the mixed audio can be covered, and to avoid missing some sound sources that only appear at specific times while only analyzing local segments.

[0109] For example, for a meeting recording that includes multiple speakers taking turns speaking and background music, the computing device can perform intensive sampling near the start and end points of the detected speech activity, while performing sparse sampling during the periods when the music is continuous and there is no speech. This improves sampling efficiency and reduces redundant data while ensuring that key speech information is fully collected.

[0110] S220: Determine the sound source category corresponding to each sample segment.

[0111] In step S220, after obtaining the sample segments, the computing device can determine the sound source category corresponding to each sample segment, thereby determining the sound source attributes of different segments in the mixed audio. In this embodiment, the sound source category is used to identify the sound source identity of each sound source corresponding to the mixed audio. The sound source category can be pre-divided into human voice category and ambient sound category. The human voice category can be further subdivided into subcategories of different speakers, while the ambient sound category can be natural or artificial sounds that are not human voices.

[0112] When identifying the sound source category corresponding to each sample segment, the computing device can perform acoustic feature extraction and classification modeling on the sample segments, and analyze them in conjunction with acoustic features such as time-spectrum diagrams. For example, the computing device can extract acoustic features such as Mel-frequency cepstral coefficients (MFCC), spectral centroid, and zero-crossing rate of the sample segment to determine the probability value of whether the sample segment belongs to the human voice category or the ambient sound category. When the probability value exceeds a preset threshold, the sound source category of the sample segment can be determined.

[0113] For example, if a sample segment has obvious formant structures in its MFCC features and its zero-crossing rate is within the range of human voice features, and its probability of belonging to the human voice category is over 90%, then the sound source category of the sample segment is determined to be human voice; if the spectral centroid of a sample segment is concentrated in the high-frequency region and has no obvious speech prosody features, and its probability of belonging to the environmental sound category is over 85%, then its sound source category is determined to be environmental sound.

[0114] In this embodiment of the application, for sample segments identified as human voice categories, the computing device can further compare them with the voiceprint features of known speakers using voiceprint recognition technology, thereby determining the specific speaker sub-category, such as "speaker A" or "speaker B", and providing corresponding sound source category labels for subsequent speech separation and transcription.

[0115] S230: Perform a noise reduction operation on the mixed audio based on each sound source category to obtain at least one audio segment corresponding to each sound source category.

[0116] In step S230, the computing device may perform a corresponding back diffusion process based on the sound source categories identified from the sample segments, for each sound source category, to split the mixed audio into audio segments corresponding to various sound source categories.

[0117] In this embodiment of the application, for a signal belonging to the human voice category in the mixed audio, the computing device can invoke the reverse diffusion process of the diffusion model, using the sample segment corresponding to the category as a guiding condition, to gradually remove noise and other sound source components that are unrelated to the human voice feature from the original mixed audio signal.

[0118] For example, when the sound source category is "Speaker A", the diffusion model uses the clean speech feature distribution of "Speaker A" obtained during the sound source category identification process. In each step of the reverse denoising, it predicts and subtracts noise components that do not match the speech features of "Speaker A", while retaining and enhancing the signal parts that match its features. In this process, the diffusion model dynamically adjusts the denoising intensity and direction to ensure that while removing interference, it preserves the integrity and naturalness of the target human voice to the greatest extent possible.

[0119] Similarly, for sound source category labeled as ambient sound, the computing device also initiates a reverse diffusion denoising process, using the acoustic characteristics of the ambient sound sample fragments as a reference, to separate the audio segment sequence that conforms to the ambient sound characteristics from the mixed audio.

[0120] Through this targeted denoising process for different sound source categories, the original mixed audio signal is split into multiple independent audio segment sequences corresponding to human voices and ambient sounds, respectively. Each audio segment has a high signal-to-noise ratio and retains the core acoustic characteristics of its corresponding sound source category.

[0121] It should be noted that the computing device can invoke the diffusion model to perform the same number of denoising operations in parallel based on the number of sound source categories, so as to synchronously generate an audio segment sequence corresponding to each sound source category. For example, if the mixed audio contains three sound sources: "speaker A", "speaker B" and "background music", the computing device can simultaneously start three back-diffusion denoising processes, using the voiceprint features of "speaker A", the voiceprint features of "speaker B" and the spectral features of "background music" as guiding conditions, respectively, to process the original mixed audio in parallel, thereby efficiently separating three independent audio segment sequences and avoiding the loss of target sound source features or processing delays that may be caused by serial processing.

[0122] In some embodiments of this application, when the computing device performs back-diffusion denoising on each sound source category using a diffusion model, it also detects the continuity of audio segments corresponding to the same sound source category. Adjacent audio segments with the same sound source category but an interval exceeding a preset threshold are separated to form independent audio sub-segments. This facilitates the computing device to accurately segment and index each audio sub-segment based on timestamps, ensuring that the final generated text sequence is strictly aligned with each audio sub-segment, and ensuring that the speech recognition result is accurately mapped to the corresponding speaker or sound source event in the time dimension.

[0123] For example, if the interval between two sample segments of the sound source category "speaker A" is greater than 5 seconds and the acoustic feature similarity exceeds 95%, they are determined to be different rounds of speech by the same speaker. The computing device can independently label them as different audio segments and generate a time-stamped text line for each segment in the subsequent speech recognition process.

[0124] S240: Obtain the first start and end times corresponding to the mixed audio and the second start and end times corresponding to each audio segment sequence.

[0125] In step S240, after the computing device calls the diffusion model to complete the audio splitting of different sound sources in the mixed audio, it records the overall time range of the original mixed audio, namely the first start and end time. This time range takes the start time of the mixed audio as the starting point and the end time as the ending point. For example, if the total duration of the mixed audio is 60s, then the first start and end time can be represented as [0s, 60s].

[0126] Meanwhile, for each individual voice segment and non-voice segment obtained from the splitting, the computing device will accurately record its specific time position in the original mixed audio, that is, the second start and end time corresponding to each audio segment. For example, there are two audio segments with the sound source category "speaker A". One of them starts from 5s and ends at 15s in the original audio, and the other starts from 35s and ends at 55s in the original audio. Then its second start and end time can be [5s, 15s] and [35s, 55s]; the audio segment of "air conditioner running sound" may run through the entire original audio, and its second start and end time is [0s, 60s].

[0127] When the computing device acquires this time information, it combines the timestamps of the sample segments recorded during the sampling process with the diffusion model's tracking of the mixed audio time dimension during processing. This ensures that the second start and end times accurately reflect the time distribution of each audio segment in the original audio, providing precise time coordinates for the subsequent audio segment recognition process.

[0128] S250: Perform timing alignment processing on the audio segment sequence based on the first start and end times and the second start and end times.

[0129] In step S250, the computing device can associate and match the second start and end times of each audio segment in each audio segment sequence with the first start and end times of the original mixed audio to establish a unified time axis coordinate system. For example, taking the start time 0s of the original mixed audio as the origin of the time axis, the second start and end times such as [5s, 15s] of "speaker A", [20s, 35s] of "speaker B" and [0s, 60s] of "air conditioner operation sound" are mapped onto this time axis to form the position markers of each audio segment in the audio segment sequence in the time dimension.

[0130] In this embodiment, the computing device also checks whether there is overlap or gap between the audio segments on the timeline and makes adjustments according to the needs of the actual application scenario. For audio segments that partially overlap in time, such as a 5-second gap between the end time (15s) of "speaker A" and the start time (20s) of "speaker B", the computing device retains the gap to reflect the actual silence or unidentified sound source periods in the original audio.

[0131] For minor time overlaps caused by sampling or processing errors, such as two audio segments whose start and end times overlap by only 0.1 seconds, the computing device can automatically determine the overlapping part as an extension of one of the audio segments by setting a time tolerance threshold, such as 0.5 seconds, or adjust the classification according to the priority of the sound source category.

[0132] In scenarios with partial overlap, the computing device prioritizes ensuring the temporal integrity of human voice segments and prioritizes them according to the start time of the audio segments to ensure that the continuity of the speaker's speech is not interrupted. When human voice overlaps with ambient sound, the complete segment is retained based on the time interval corresponding to the human voice segment, while non-human voice segments are cropped or labeled as background layers.

[0133] Through this temporal alignment process, all audio segments form an orderly and clear distribution structure on the timeline, enabling computing devices to perform subsequent speech recognition or environmental sound analysis on each audio segment based on the time sequence. This ensures that the processing results are highly consistent with the original mixed audio in the time dimension, providing an input basis for computing devices to generate recognition text or environmental sound event reports with accurate timestamps in subsequent processes.

[0134] Through steps S210 to S250, the computing device can accurately split the mixed audio with the help of the diffusion model. It not only separates the audio segments of different sound sources in the mixed audio, but also provides an input basis for subsequent speech recognition and environmental sound analysis through sound source category recognition, time information recording and time sequence alignment processing.

[0135] Figure 9 This is a schematic diagram illustrating a process for distinguishing between human voice segments and non-human voice segments, provided as an embodiment of this application. The following is based on... Figure 9 The content shown illustrates, by way of example, the process by which a computing device distinguishes between human voice segments and non-human voice segments in an audio segment sequence in this application embodiment.

[0136] like Figure 9 As shown in the embodiments of this application, the process by which the computing device distinguishes between human voice segments and non-human voice segments in an audio segment may include the following steps S310 to S330.

[0137] S310: Based on the sound source category, distinguish between human voice segments and non-human voice segments in an audio segment sequence.

[0138] In step S310, after obtaining the corresponding audio segment sequence, the computing device can clearly divide each audio segment into human voice segments and non-human voice segments according to the sound source category label corresponding to each audio segment sequence determined by the diffusion model in step S220.

[0139] For example, for audio segment sequences labeled as “human voice category” and each speaker subcategory, such as audio segment sequences corresponding to “speaker A” and “speaker B”, the computing device can uniformly classify the corresponding audio segments as human voice segments.

[0140] Audio segments labeled as "ambient sound category," such as "background music," "keyboard tapping," and "outdoor rain," can be categorized by the computing device as non-human voice segments.

[0141] In this embodiment of the application, the computing device can split audio signals in the same audio segment sequence whose silence time exceeds a preset threshold into two audio segments based on the start and end times of the silence time, according to the silence time in the silence time in the audio segment sequence corresponding to the same sound source.

[0142] The computing device can identify different audio segments within the same audio segment sequence as belonging to the category corresponding to the sound source category label of the audio segment sequence. For example, if the sound source category label of the audio segment sequence is human voice, then all audio segments contained in that audio segment sequence can be considered human voice segments. In this way, by classifying the separated audio segments by attribute, the computing device can provide a basis for subsequent differentiated processing of different types of audio segments.

[0143] For example, if three audio segment sequences, namely “speaker A”, “speaker B” and “air conditioner operation sound”, are separated in step S230, then in this step, the audio segments in the audio segment sequences corresponding to “speaker A” and “speaker B” will be identified as human voice segments, while “air conditioner operation sound” will be identified as a non-human voice segment.

[0144] S320: Identifies the sound type in non-human voice segments.

[0145] In step S320, after the computing device completes audio segmentation using a diffusion model and labels the corresponding sound source categories, it uses different processing strategies to identify and process human voice segments and non-human voice segments. In this embodiment, the computing device can execute the processing flow corresponding to human voice segments through the aforementioned step S400. When there are non-human voice segments in the audio segment sequence separated from the mixed audio, the computing device can simultaneously execute step S320 to identify the sound type in the non-human voice segments while executing S400 to identify human voice segments.

[0146] In this embodiment, for non-human voice segments, the computing device invokes the sound classification unit in the processor to further identify the sound type of the non-human voice segments. For example, the sound classification unit has preset feature templates or classification models for various common environmental sounds. For instance, it uses trained convolutional neural networks or recurrent neural networks to extract and analyze the acoustic features of non-human voice segments, such as spectral characteristics, Mel-frequency cepstral coefficients, and energy.

[0147] By comparing and matching the extracted features with the features of each category in the preset feature template or classification model, the computing device can identify which type of sound is contained in the non-human voice segment, such as the sound of an air conditioner running, the sound of a keyboard typing, the sound of a car horn outside the window, music, the sound of wind, the sound of rain, or the sound of a specific device alarm.

[0148] The computing device can record the sound type information identified by the sound classification unit, so that it can be output together with the speech recognition results, providing more comprehensive audio scene information for the subsequent speech recognition process and speech recognition results.

[0149] S330: When outputting a text sequence, output the sound type corresponding to the non-human voice segment.

[0150] In step S330, after obtaining the sound type corresponding to the non-human voice segment, the computing device can combine the recognition results of each human voice segment with the results of the computing device and output them to the user in a preset format. For example, the computing device will append the sound type information of the recognized non-human voice segments, such as "air conditioner running sound" or "car horn sound," as text annotations next to the corresponding text sequence, or list them together at the beginning or end of the overall recognition result, thereby providing the user with richer audio scene background information.

[0151] For example, in a meeting recording scenario, the output of the computing device after recognizing the speech can be displayed as: "[Ambient sound: Air conditioner running sound] Speaker A: The main topic we are discussing today is...; Speaker B: I think we need to consider... [Ambient sound: Brief keyboard tapping sound]", allowing users to intuitively understand the audio context corresponding to the recognized text.

[0152] In some embodiments, if the non-human voice segment is a sound type that requires special attention, such as a device alarm sound, the computing device can also highlight the sound type with a special marker to remind the user to pay attention to abnormal sound events during that period. This processing method of simultaneously outputting non-human voice sound types can completely restore the scene information of the original mixed audio, avoiding the loss of scene information caused by only outputting human voice text, and improving the completeness and reference value of the speech recognition results.

[0153] Through the above steps S310 to S330, the computing device can distinguish between human voice segments and non-human voice segments, and perform type identification and result labeling on non-human voice segments. This clarifies the core processing object for subsequent speech recognition and completely preserves the scene information of the original audio, providing support for generating comprehensive and complete recognition results.

[0154] Figure 10 This is a schematic flowchart illustrating a method for determining the category of a sound source, provided in an embodiment of this application. The following is based on... Figure 10The content shown illustrates, by way of example, the process by which the computing device determines the sound source category of a sample segment in the embodiments of this application.

[0155] like Figure 10 As shown, in this embodiment of the application, the process of the computing device identifying the sound source type of the sample segment may include the following steps S221 to S223.

[0156] S221: Extract audio features corresponding to multiple sample segments.

[0157] In step S221, after obtaining the sample segments, the computing device can extract audio features for each sample segment. These audio features can describe the sound characteristics of the sample segments from different dimensions. For example, the audio features that the computing device can extract include, but are not limited to, Mel frequency cepstral coefficients (MFCC), Mel spectrogram, spectral centroid, spectral bandwidth, zero-crossing rate, short-time energy, fundamental frequency (F0), and voice activity detection (VAD) features.

[0158] Among them, MFCC can effectively reflect the nonlinear perception characteristics of human ears to sound frequencies, and contains the spectral envelope information of sound, which is a commonly used feature in speech recognition; Mel spectrogram, on the other hand, converts linear frequencies into Mel frequencies, which is more in line with the characteristics of human hearing and can intuitively show the energy distribution of sound at different frequencies and times.

[0159] The spectral centroid reflects the frequency position where sound energy is concentrated; high-frequency sounds have a higher spectral centroid, while low-frequency sounds have a lower one. The zero-crossing rate indicates the number of times the signal waveform crosses the zero level per unit time, which is helpful in distinguishing between unvoiced and voiced sounds. Short-time energy reflects the changes in sound intensity and can be used to detect the start and end of speech.

[0160] By extracting these multi-dimensional audio features, computing devices provide rich criteria for subsequent sound source classification. For example, for a sample segment containing human voice, its MFCC features will exhibit a clear formant structure, the fundamental frequency will fluctuate within a certain range, and the short-term energy will change periodically with the rhythm of the speech; while for a sample segment of ambient sound from pure music, its Mel spectrogram may have a richer high-frequency overtone structure, the spectral centroid may fluctuate significantly with the rhythm of the music, and the zero-crossing rate is relatively stable.

[0161] S222: Cluster the audio features of multiple sample segments to obtain at least one audio feature cluster.

[0162] In step S222, the computing device may use a clustering algorithm to group the audio features of the extracted multiple sample segments, aggregating sample segments with high feature similarity into the same audio feature cluster. For example, the computing device may use the K-means clustering algorithm to perform audio feature clustering.

[0163] The computing device can preset the number of clusters based on the number of sample segments and the initial feature distribution. For example, based on the human voice and ambient sound categories already divided in step S220, the device can perform sub-clustering on the sample segments under each category. Taking the human voice category as an example, the computing device combines the MFCC, fundamental frequency, and voiceprint features of all human voice sample segments into a high-dimensional feature vector. By calculating the Euclidean distance or cosine similarity between vectors, the device groups sample segments that are close in distance in the feature vector space into the same audio feature cluster. Each audio feature cluster represents a potential speaker sub-category.

[0164] During clustering, the computing device dynamically adjusts the cluster centers, iteratively optimizing to minimize the feature variance of samples within the same cluster and maximize the feature differences between different clusters. For example, if there are two speakers, the clustering algorithm will group sample segments with similar voiceprint features, such as spectral envelope and formant frequency distribution, into two clusters, with each cluster corresponding to one speaker.

[0165] For ambient sound categories, the computing device also clusters the sample segments based on features such as spectral centroid, spectral bandwidth, and zero-crossing rate, grouping ambient sound samples with similar acoustic properties into the same feature cluster, thereby achieving a more detailed classification of ambient sound categories.

[0166] Through clustering, computing devices can organize originally disordered sample fragments into several audio feature clusters with clear common characteristics, providing a structured feature set for subsequent category determination.

[0167] S223: Determine the sound source category based on the center vector of each audio feature cluster.

[0168] In step S223, the computing device can compare the center vector of each audio feature cluster with a preset sound source category feature library to determine the sound source category corresponding to the audio feature cluster. The sound source category feature library contains standard feature vectors of various known sound source categories, such as voiceprint feature templates of different speakers and spectral feature models of common environmental sounds (such as car horns, telephone rings, wind sounds, etc.), which are used to perform similarity matching with the center vectors of the audio feature clusters identified by the computing device, thereby determining the most matching sound source category.

[0169] For example, the computing device calculates the similarity between the center vector of an audio feature cluster and each standard feature vector in the feature library, and determines the sound source category corresponding to the standard feature vector with the highest similarity as the sound source category of the audio feature cluster. For instance, if the cosine similarity between the center vector of an audio feature cluster and the voiceprint feature template of "speaker C" in the human voice category reaches 0.92, which is much higher than the similarity with other speaker or ambient sound categories, then the sound source category of the sample segment of that cluster is determined to be "speaker C".

[0170] For each ambient sound feature cluster, the computing device matches its center vector with standard ambient sound models in the feature library, such as "keyboard typing," "air conditioner operation," and "background conversation." If the cluster matches the spectral feature model of "keyboard typing" best, the corresponding sound source category is labeled as "keyboard typing." If the similarity between the center vector of an audio feature cluster and all known categories in the feature library is below a preset threshold, such as 0.7, the computing device can temporarily label it as an "unknown sound source category" and supplement the classification through manual annotation or further feature learning in subsequent processing.

[0171] Through this center vector-based sound source category matching, the computing device can accurately assign a corresponding sound source category label to each audio feature cluster, thereby completing the sound source category determination of the sample segment.

[0172] Through steps S221 to S223, the computing device can extract multi-dimensional audio features from sample segments, aggregate sample segments with similar features into feature clusters using a clustering algorithm, and accurately determine the sound source category of each sample segment based on the comparison between the center vector of the feature cluster and a preset sound source category feature library. By executing this process, the computing device can effectively distinguish between human voices and ambient sounds, and can further subdivide the categories, such as different speakers in human voices and specific sound source types in ambient sounds. This provides crucial category basis for subsequent audio segment separation, time alignment, and differential processing, ensuring the speech recognition system accurately grasps sound source information in complex audio environments.

[0173] Figure 11 This is a schematic diagram illustrating a process for identifying non-human voice segment sound types, provided as an embodiment of this application. The following is based on... Figure 11 The content shown illustrates, by way of example, the process by which a computing device identifies the sound type in a non-human voice segment in an embodiment of this application.

[0174] like Figure 11 As shown, in the embodiments of this application, the process by which the computing device identifies the sound type in a non-human voice segment may include the following steps S321 to S323.

[0175] S321: Extract audio features corresponding to non-human voice segments.

[0176] In step S321, after acquiring the non-human voice segment, the computing device extracts its audio features. These audio features represent the unique acoustic properties of the ambient sound, enabling accurate sound type identification in the subsequent process. Similar to extracting the audio features of the sample segment in step S221, the audio features extracted here for the non-human voice segment can also include the audio features from step S221.

[0177] Furthermore, considering the diversity and complexity of environmental sounds, computing devices can also extract dynamic features such as spectral flux, chroma features, and first and second-order differences of MFCC to better describe the characteristics of environmental sounds over time. For example, for a non-human sound segment like a car horn, its Mel spectrum may show obvious energy peaks in specific high-frequency regions, a high spectral centroid, and short-term energy exhibiting sudden, strong pulse characteristics; while a non-human sound segment like flowing water typically has a high zero-crossing rate, a relatively low spectral centroid, a more uniform energy distribution, and a smoother change in spectral flux.

[0178] The computing device extracts these multi-dimensional static and dynamic audio features to construct a feature vector that can comprehensively characterize the acoustic properties of non-human voice segments.

[0179] S322: Based on a preset ambient sound category library, calculate the similarity between the audio features corresponding to non-human voice segments and the reference features in the ambient sound category library.

[0180] In step S322, the computing device calculates the similarity between the extracted audio feature vector of the non-human voice segment and each reference feature vector in the preset environmental sound category library. The reference features are used to represent the audio features of each category of environmental sound in the environmental sound category library. The preset environmental sound category library contains reference feature vectors of various common environmental sound types, such as "car horn", "telephone ring", "keyboard typing", "air conditioner operation", "wind sound", "rain sound", "background music", etc. Each sound type corresponds to one or more trained and optimized reference feature templates.

[0181] Computational devices can use various metrics such as cosine similarity and Euclidean distance to calculate the similarity between feature vectors. Taking cosine similarity as an example, the closer the value is to 1, the more consistent the direction of the feature vector of the non-human voice segment with the reference feature vector, meaning that their acoustic characteristics are more similar. For example, when the cosine similarity between the reference feature vector of "car horn" and the feature vector of a certain non-human voice segment is 0.85, while the cosine similarity with the reference feature vector of "telephone ringing" is 0.32, it indicates that the non-human voice segment is more likely to belong to the "car horn" type.

[0182] The computing device will traverse all reference features in the ambient sound category library, calculate and record the similarity value between the feature vector of non-human voice segments and each reference feature, and provide a quantitative basis for subsequent determination of sound type.

[0183] S323: Determine the sound type corresponding to the non-human voice segment based on similarity.

[0184] In step S323, the computing device can determine the ambient sound category corresponding to the reference feature with the highest similarity as the sound type of the non-human voice segment. For example, if a non-human voice segment has the highest similarity to the reference feature vector of "keyboard typing sound" in the ambient sound category library, reaching 0.88, and the similarity value is higher than the preset confidence threshold, then the computing device determines that the sound type of the non-human voice segment is "keyboard typing sound".

[0185] If multiple reference features have similarities that are close and all are above the confidence threshold, the computing device can further combine the time information of non-human voice segments, their correlation with other audio segments, or the context of the scene. For example, in a meeting scene, "speaking" or "keyboard typing" is more likely to occur than "car horn" to assist in the judgment and determine the most likely sound type.

[0186] For non-human voice segments where the similarity of all reference features is below the confidence threshold, the computing device can mark them as "unrecognized ambient sound" and temporarily store their feature vectors for later re-recognition by updating the ambient sound category library or introducing more training samples.

[0187] Through the steps S321 to S323 described above, the computing device can accurately identify the specific sound type of non-human voice segments, thereby providing detailed environmental sound information for subsequent environmental sound event analysis, scene understanding, or optimization of speech recognition results.

[0188] Figure 12 This is a schematic diagram of the training process of a speech recognition model provided in an embodiment of this application. Figure 13 This is a schematic diagram of a speech recognition model provided in an embodiment of this application. The following is in conjunction with... Figure 12 and Figure 13 The content shown here is an exemplary illustration of the training process of the speech recognition model pre-installed in the processor of the computing device in the embodiments of this application.

[0189] like Figure 12 As shown in the embodiments of this application, the training process of the speech recognition model by the computing device may include the following steps S710 to S740.

[0190] S710: Acquire training speech signals, training text corresponding to training speech signals, reference speech signals, and reference text corresponding to reference speech signals.

[0191] In step S710, before training the speech recognition model, the computing device needs to acquire the speech dataset required for training, including the training speech signal, the training text corresponding to the training speech signal, and the reference speech signal and the reference text corresponding to the reference speech signal. The training speech signal refers to the original speech samples used for model parameter initialization and supervised learning; the training text is the standard transcribed text corresponding to the training speech signal, used to construct labels for supervised learning; the reference speech signal includes a speech signal aligned with the contextual features of the training speech signal, and the reference text is the standard transcribed text corresponding to the reference speech signal. Together, they constitute the context-aware supervision signal.

[0192] In this embodiment of the application, the training process of the speech recognition model by the computing device includes a training phase and a fine-tuning phase. The training phase refers to the initial parameter learning of the model using training speech signals and training text to build a basic speech-to-text mapping capability. The fine-tuning phase is based on the training phase, combining reference speech signals and reference text to finely adjust the model parameters so that the model can better utilize contextual information to optimize the recognition results.

[0193] For example, during the training phase, computing devices can input a large number of training speech signals containing different speakers and accents into the model. By comparing the differences between the predicted text output by the model and the training text, the weight parameters of the model are continuously adjusted, so that the model gradually learns the correspondence between speech features and text sequences.

[0194] During the fine-tuning phase, the computing device introduces reference speech signals as context for the training speech signals. For example, in a speech recognition task within a meeting setting, the reference speech signals could be speeches from the same meeting occurring in a time period adjacent to the training speech signals, and the reference text would be an accurate transcription of those speeches. By fusing the features of the reference speech signals (such as voiceprint features and semantic features) with the features of the training speech signals, the model can consider the semantic coherence of the context and the speaker's speech habits during the recognition process, thereby reducing recognition errors caused by homonyms, accent differences, or short-term noise interference.

[0195] S720: Uses the initial speech recognition model to recognize the training speech signal in order to generate the initial recognition result.

[0196] In step S720, as Figure 13 As shown in (a), the computing device can input training speech signals and training text into the initial speech recognition model. The initial speech recognition model outputs the corresponding recognition results and compares them with the training text, thereby realizing the training process of the initial speech recognition model.

[0197] In one example, the initial speech recognition model may include structures such as a feature extraction layer, an encoder, and a decoder. After receiving the training speech signal, the feature extraction layer in the initial speech recognition model can preprocess the input training speech signal, such as pre-emphasis, framing, and windowing, and then extract acoustic features such as Mel-frequency cepstral coefficients (MFCC) and filter bank features (Fbank).

[0198] The encoder maps these acoustic features into a high-dimensional sequence of hidden states, capturing temporal information and contextual dependencies in the speech signal. For example, it can use a long short-term memory network or a Transformer structure as the encoder and perform weighted processing on features at different time steps through a self-attention mechanism.

[0199] The decoder then generates the corresponding initial recognition text sequence, i.e., the initial recognition result, based on the hidden state sequence output by the encoder and combined with the model's prior knowledge. For example, for the training speech signal "What will the weather be like tomorrow?", the initial speech recognition model may output an initial recognition result with slight deviations, such as "How will the weather be tomorrow?" or "How will the air be tomorrow?".

[0200] S730: Update the parameters of the initial speech recognition model based on the semantic differences between the initial recognition results and the training text.

[0201] In step S730, the computing device can quantify the semantic difference between the initial recognition result and the training text by means of semantic similarity calculation or minimizing the negative log-likelihood loss function, and update the parameters of the initial speech recognition model based on the backpropagation of the difference.

[0202] In one example, the computing device can convert the initial recognition result and the training text into word vectors or sentence embeddings, and measure the semantic deviation between the two by calculating metrics such as cosine similarity and edit distance. For example, if the initial recognition result is "How is the gas station tomorrow?" and the training text is "How is the weather tomorrow?", the edit distance between the two is 1, indicating low semantic similarity. In this case, the initial speech recognition model will calculate the loss value corresponding to this difference and adjust the weights, biases, and other parameters of each layer in the encoder and decoder using the gradient descent algorithm to reduce the probability of similar errors occurring in subsequent recognition processes.

[0203] In some embodiments, the computing device may also use minimizing the negative log-likelihood loss function as the objective to directly optimize the model's ability to fit the real speech distribution. The expression for this loss function can be given by the following equation (3): (3); Here, X represents the training speech signal, and Y represents the corresponding standard text label; P(Y|X) represents the conditional probability that the model generates the correct text sequence Y given the speech signal X. By maximizing this probability, the model gradually improves the accuracy of its modeling of the speech-to-text mapping relationship.

[0204] It should be understood that during the training phase, the computing device will iteratively execute steps S720 to S730 until the model's recognition accuracy on the validation set reaches a preset threshold or the loss value converges, thus completing the parameter optimization of the initial model.

[0205] S740: Context-aware training is performed on the updated speech recognition model based on the training speech signal, the reference speech signal, and the reference text.

[0206] In step S740, after the training phase of the initial speech recognition model is completed, the computing device can introduce a context-aware joint modeling mechanism by combining reference speech signals and reference text. By fusing speech features and text semantic priors through a cross-modal attention module, the model's ability to dynamically capture context features is enhanced.

[0207] like Figure 13 As shown in (b) and (c), the speech recognition model in this embodiment may include an audio frame embedding unit, multiple audio feature encoding layers, a text embedding unit, multiple text feature encoding layers, and a cross-modal attention fusion layer. The audio frame embedding unit can receive training speech signals and reference speech signals and convert them into corresponding audio frame embedding vectors. These vectors can be generated by superimposing Mel spectrogram features or MFCC features with positional encoding to preserve the temporal positional information of the speech signals.

[0208] Multiple audio feature encoding layers adopt a Transformer structure. Each layer contains a multi-head self-attention mechanism and a feedforward neural network to encode the audio frame embedding vector layer by layer, extracting deep acoustic features and contextual dependencies from the training speech signal and the reference speech signal.

[0209] The text embedding unit converts the reference text into text embedding vectors, obtains the semantic representation of the reference text through a pre-trained language model, and incorporates positional encoding to distinguish the positions of different words in the text sequence. Multiple text feature encoding layers can also be based on the Transformer structure to encode the text embedding vectors, capturing the semantic logic and contextual information of the reference text.

[0210] As the core module of the model, the cross-modal attention fusion layer can receive the speech feature sequence output by the audio feature encoding layer and the text feature sequence output by the text feature encoding layer. By calculating the cross-attention weights between the speech features and the text features, it realizes the deep interaction and fusion of speech information and text semantics. For example, when the word "apple" appears in the training speech signal, if there are related words such as "fruit" and "edible" in the reference text, the cross-modal attention fusion layer will enhance the influence of these semantically related features on the speech recognition of "apple", thereby reducing the possibility of misrecognizing "apple" as "pingguo" or "pan".

[0211] During the context-aware training process, the computing device inputs the fused feature vector into the decoder to generate the predicted text that integrates the context information, and further adjusts the parameters of the audio encoding layer, text encoding layer, and cross-modal attention layer in the model by comparing the differences between the predicted text and the reference text, enabling the model to dynamically utilize the context clues in the reference speech and reference text to optimize the recognition results.

[0212] In some embodiments, the computing device can also introduce environmental sound features as additional inputs to the audio frame embedding unit, and splice or weight-fuse the corresponding features, such as energy distribution features and spectral characteristic vectors, with the audio frame embedding vector. In this way, the model can learn the potential impact of different environmental sounds on speech recognition during the training process. For example, in an environment with strong "keyboard tapping sounds", the model can automatically enhance the attention to the high-frequency components in the speech signal or suppress the noise in a specific frequency band, thereby improving the speech recognition accuracy in complex environments.

[0213] Correspondingly, the computing device can also optimize the model parameters using the loss function during the fine-tuning stage, focusing on adjusting the cross-modal attention weights and the environmental feature fusion coefficients. The computing device can add two constraint terms, namely the semantic consistency of the front and back segments and the recognition and alignment of environmental sound events, to the original loss function, enabling the model to not only focus on the accurate recognition of single-frame speech but also strengthen the modeling ability of semantic coherence and the temporal alignment of environmental sound events.

[0214] Through this training method of multi-source information fusion, the speech recognition model can not only learn the mapping relationship between speech and text but also combine context speech, text semantics, and environmental sound features to achieve more accurate and intelligent speech recognition.

[0215] Through the above steps S710 to S740, the computing device can complete the training and fine-tuning of the speech recognition model, enabling it to have the speech-to-text ability and a powerful context-aware optimization ability, laying a foundation for high-precision speech recognition in practical applications.

[0216] Figure 14This is a schematic diagram illustrating a process for recognizing audio segments, provided as an embodiment of this application. The following is based on... Figure 14 The content shown is used to illustrate, by way of example, the process by which a computing device calls a speech recognition model to process audio segments in the embodiments of this application.

[0217] like Figure 14 As shown, the process of the computing device provided in this application embodiment using a speech recognition model to process audio segments may include the following steps S410 to S450.

[0218] S410: Read the acoustic features of the target audio segment.

[0219] In step S410, the computing device can invoke the speech recognition model trained in the preceding steps to process the input audio segment to be recognized, thereby obtaining the corresponding speech recognition result, i.e., the text sequence. In one example, the target audio segment is the human voice segment output by the diffusion model in the preceding steps.

[0220] The computing device first preprocesses the target audio segment, similar to the preprocessing of the training speech signal during the training phase. This includes pre-emphasis to boost high-frequency signal energy, dividing the mixed audio into frames of 20-30 milliseconds in length, and applying a Hamming window to each frame to reduce spectral leakage. Subsequently, the computing device uses the audio frame embedding unit in the speech recognition model to extract acoustic features, such as Mel-frequency cepstral coefficients (MFCCs) or filter bank features (Fbank), from the preprocessed audio frames. These acoustic features are then superimposed with positional coding information to generate the audio frame embedding vector corresponding to the target audio segment, thus preserving the temporal order information of the mixed audio.

[0221] In this embodiment of the application, the computing device can also determine the preceding and following audio segments adjacent to the target audio segment through the third start and end time corresponding to the target audio segment, input their acoustic features into the speech recognition model, construct a cross-segment context-aware window, and provide context for the subsequent recognition process.

[0222] S420: Based on acoustic features, retrieve the semantic features of texts that match the acoustic features from historical recognition results.

[0223] In step S420, the computing device can retrieve the context fragment with the highest semantic consistency from the historical recognition results based on acoustic features, and fuse the environmental sound features with the highest matching degree in the environmental sound category library to construct a multi-dimensional context-aware input.

[0224] In one example, the computing device first inputs the acoustic features of the target audio segment and adjacent audio segments into the audio feature encoding layer of the speech recognition model to generate a deep feature representation of the current speech.

[0225] Simultaneously, the computing device accesses a pre-set historical recognition result database, which stores recently recognized text sequences and their corresponding semantic feature vectors. The computing device calculates the cosine similarity between the current deep speech features and the historical text semantic feature vectors, selecting the top N historical recognition results with the highest similarity as contextual references.

[0226] For example, if the current target audio segment is "The meeting will be held at 3 p.m.", and the historical recognition results contain "This project discussion meeting is scheduled for this week", then the semantic features of the historical text may have a high similarity to the current speech features, and it will be selected as the context.

[0227] In some embodiments, the computing device can also use the text sequence corresponding to adjacent audio segments in the historical recognition results as text that matches the acoustic features, and then perform context modeling based on the semantic features of the text sequence. For example, if the recognized text corresponding to the preceding audio segment of the target audio segment is "Please prepare your presentation materials in advance", the computing device will extract the semantic features of the text as an important contextual basis for the recognition of the current target audio segment, so as to enhance the recognition accuracy of "meeting" related words.

[0228] S430: Based on the third start and end time of the target audio segment, determine the non-human voice segment located within the third start and end time.

[0229] In step S430, the computing device can extract a non-human voice segment from the original mixed audio that completely overlaps with the third start and end time of the target audio segment by timestamp matching.

[0230] For example, if the third start and end time of the target audio segment is from the 30th to the 105th second of the original mixed audio, the computing device will extract the non-human voice segments within the corresponding time period from the original mixed audio as input to the speech recognition model, and perform temporal alignment and fusion with the acoustic features of the target speech segment so that the model can refer to the environmental sound characteristics within that time period when recognizing the target speech.

[0231] S440: Fusion of acoustic features, semantic features, and non-human voice segments located in the third start and end time interval to obtain the fused features corresponding to the target audio segment.

[0232] In step S440, the computing device can use the cross-modal feature fusion module in the speech recognition model to perform layer-by-layer alignment and weighted fusion of acoustic features, historical semantic features and time-spectrum maps of non-human voice segments to generate unified context-aware fusion features.

[0233] In one example, the cross-modal feature fusion module first converts acoustic features, semantic features, and non-human voice segments into feature vectors of the same dimension. Then, it calculates the correlation weights among the three through a multi-head cross-attention mechanism. For example, when a non-human voice segment contains "conference room background noise," the model reduces the interference weight of that frequency band noise on the acoustic features while increasing the weight of semantic features related to "conference," making the recognized text more consistent with the context.

[0234] During the fusion process, the model dynamically adjusts the contribution ratio of each feature. For scenarios with clear speech signals and well-defined historical semantics, the weights of acoustic and semantic features are increased accordingly. Conversely, in situations with complex environmental noise, the feature weights of non-human voice segments are increased to help the model distinguish between valid speech and noise. Through this multi-feature fusion strategy, the speech recognition model can comprehensively utilize the acoustic information of the speech itself, the semantic logic of the historical context, and the sound characteristics of the current environment, significantly improving the accuracy of target audio segment recognition.

[0235] S450: Generate the text sequence corresponding to the target audio segment based on the fusion features.

[0236] In step S450, the computing device inputs the fused features into the decoder corresponding to the speech recognition model. The decoder adopts an autoregressive decoding method based on Transformer and decodes the fused feature sequence through a multi-head self-attention mechanism to generate the text sequence corresponding to the target audio segment word by word.

[0237] During the decoding process, the decoder combines the generated historical text sequence to dynamically predict the next most likely word. For example, when the fused features contain key information such as "meeting" and "3 p.m.", and the historical text sequence has already generated "the meeting will be held", the decoder will prioritize predicting "afternoon" as the next word based on contextual semantics and acoustic features, and then generate the complete text sequence "the meeting will be held at 3 p.m."

[0238] In some embodiments, the decoder also uses strategies such as beam search to retain multiple candidate text sequences during the generation process, and selects the optimal output result based on the combined language model score and acoustic model score, ensuring that the generated text sequence not only matches the target audio segment acoustically, but also has good semantic coherence and grammatical correctness.

[0239] In addition, for ambiguous sounds or homophones that may appear during the recognition process, the model will combine contextual semantic features and environmental audio features to resolve ambiguities. For example, in the recognition of "apple" and "pingguo", if the contextual semantic features contain information such as "fruit" and "edible", the model will prioritize "apple" as the correct recognition result, so as to finally output an accurate and context-appropriate text sequence.

[0240] After recognizing a target audio segment using a speech recognition model, the computing device can time-stamp the recognized content. This involves associating and storing each word or sentence in the text sequence with the third start and end time of the target audio segment. This allows for quick location of the specific audio time segment when it's necessary to trace the correspondence between audio and text later. Furthermore, this time-stamping mechanism not only facilitates simultaneous listening to the corresponding audio content while viewing the recognized text but also enables more refined audio-text linkage operations in subsequent text editing, information retrieval, or voice interaction applications.

[0241] Through steps S410 to S450, the computing device can accurately identify the target audio segment and generate a text sequence with time stamps. This process fully utilizes the multi-source information fusion capability built during the training phase, effectively improving the accuracy and anti-interference capability of speech recognition through deep interaction of acoustic features, historical semantic features, and environmental audio features.

[0242] Figure 15 This is a schematic diagram of the structure of a speech recognition system provided in an embodiment of this application.

[0243] Corresponding to the embodiments of the aforementioned speech recognition methods, this application also provides embodiments of a speech recognition system. For example... Figure 15 As shown, the speech recognition system 1500 may include an audio acquisition unit 1510, a diffusion model unit 1520, a sound classification unit 1530, a speech recognition unit 1540, and an output unit 1550.

[0244] The audio acquisition unit 1510 is configured to acquire mixed audio containing at least two sound sources.

[0245] The diffusion model unit 1520 is configured to perform multi-source separation on mixed audio to generate multiple temporally complete and independent audio segment sequences.

[0246] The sound classification unit 1530 is configured to distinguish between human voice segments and non-human voice segments in an audio segment sequence.

[0247] The speech recognition unit 1540 is configured to recognize a target audio segment based on a speech recognition model and historical recognition results, so as to convert the target audio segment into a text sequence; the target audio segment can be any human voice segment in mixed audio, and the historical recognition results include the recognized text generated by the speech recognition model based on human voice segments whose start time is located before the start time corresponding to the target audio segment, as well as the acoustic features or mixed audio of human voice segments whose start time is located before the start time corresponding to the target audio segment.

[0248] Output unit 1550 is configured to output a text sequence.

[0249] Figure 16 This is a schematic diagram of a computing device provided in an embodiment of this application.

[0250] like Figure 16 As shown, the computing device 1600 includes a processor 1601 and a memory 1602. Exemplarily, the computing device 1600 may also include a communications interface 1603 and a communications bus 1604.

[0251] The processor 1601, memory 1602, and communication interface 1603 communicate with each other via communication bus 1604. The communication interface 1603 may include a transmitter and receiver for communicating with other devices or communication networks, and may be a wired interface (port), such as a fiber distributed data interface (FDDI) or a gigabit Ethernet interface (GE).

[0252] In this embodiment, the communication interface 1603 can be used to enable communication between the computing device 1600 and the terminal device operated by the user, so that the computing device 1600 can receive user input operations, provide a visual interface for the terminal device and execute operations to generate corresponding operation results, thereby configuring the interface of the management system deployed in the computing device 1600.

[0253] In some embodiments, the processor 1601 is used to execute program 1605, which specifically performs the relevant steps in the above-described speech recognition method embodiments. Specifically, program 1605 may include program code, which includes computer-executable instructions.

[0254] For example, processor 1601 may be a central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement some embodiments of this application. Computing device 1600 may include one or more processors, which may be processors of the same type, such as one or more CPUs; or processors of different types, such as one or more CPUs and one or more ASICs. The CPU may be a single-core CPU or a multi-core CPU.

[0255] In some embodiments, memory 1602 is used to store program 1605. Memory 1602 may include high-speed random access memory (RAM) or non-volatile memory (NVM), such as at least one disk storage device.

[0256] Specifically, program 1605 can be called by processor 1601 to cause computing device 1600 to perform speech recognition method operations.

[0257] Some embodiments of this application provide a computer-readable storage medium storing at least one executable instruction that, when executed on a computing device 1600, causes the computing device 1600 to perform the speech recognition method described in the above embodiments.

[0258] For example, the computer-readable storage medium can be a read-only memory (ROM), a random access memory (RAM), a compact disc read-only memory (CD-ROM), magnetic tape, a floppy disk, and an optical data storage device.

[0259] This application provides a chip system in some embodiments, which is applied to a server. The chip system includes one or more interface circuits and one or more processors. The interface circuits and processors are interconnected via lines. The interface circuits are used to receive signals from the server's memory and send signals to the processors, the signals including computer instructions stored in the memory. When the processor executes the computer instructions, the server performs various steps in the speech recognition method shown in the above-described method embodiments.

[0260] The beneficial effects that the readable storage medium provided in some embodiments of this application can achieve can be referred to the beneficial effects in the corresponding speech recognition methods provided above, and will not be repeated here.

[0261] The embodiments described above are merely specific embodiments of this application and are not intended to limit the scope of protection of this application. Any modifications, equivalent substitutions, improvements, etc., made based on the technical solution of this application should be included within the scope of protection of this application.

Claims

1. A speech recognition method, characterized in that, include: Obtain a mixed audio file containing at least two sound sources; The mixed audio is separated into multiple sound sources using a diffusion model, generating multiple temporally complete and mutually independent audio segment sequences. Distinguish between human voice segments and non-human voice segments in the audio segment sequence; Based on the speech recognition model and historical recognition results, a target audio segment is identified to convert the target audio segment into a text sequence; the target audio segment can be any of the human voice segments in the mixed audio; the historical recognition results include the recognized text generated by the speech recognition model based on the human voice segments whose start time is located before the start time corresponding to the target audio segment, as well as the acoustic features or mixed audio of the human voice segments whose start time is located before the start time corresponding to the target audio segment. Output the text sequence.

2. The method according to claim 1, characterized in that, The method of using a diffusion model to separate multiple sound sources in the mixed audio generates multiple temporally complete and independent audio segment sequences, including: A diffusion model is used to sample the mixed audio to obtain multiple sample segments; Determine the sound source category corresponding to each sample segment; the sound source category is used to identify the sound source identity of each sound source corresponding to the mixed audio. Based on each of the sound source categories, a denoising operation is performed on the mixed audio to obtain at least one audio segment sequence corresponding to each of the sound source categories; wherein, when there is an overlap of audio segments corresponding to different sound source categories on the time axis of the mixed audio, the denoising operation generates a complete signal for each sound source category within the overlapping time period.

3. The method according to claim 2, characterized in that, The distinction between human voice segments and non-human voice segments in the audio segment sequence includes: Based on the sound source category, distinguish between human voice segments and non-human voice segments in the audio segment sequence; After distinguishing between human voice segments and non-human voice segments in the audio segment sequence, the method further includes: Identify the sound type in the non-human voice segment; When outputting the text sequence, output the sound type corresponding to the non-human voice segment.

4. The method according to claim 2, characterized in that, Determining the sound source category corresponding to each sample segment includes: Extract the audio features corresponding to the multiple sample segments; The audio features of the multiple sample segments are clustered to obtain at least one audio feature cluster; The sound source category is determined based on the center vector of each of the audio feature clusters.

5. The method according to any one of claims 1 to 4, characterized in that, After performing multi-source separation on the mixed audio using a diffusion model to generate multiple temporally complete and independent audio segment sequences, the method further includes: Obtain the first start and end times corresponding to the mixed audio and the second start and end times corresponding to each audio segment sequence; Based on the first start and end times and the second start and end times, time alignment processing is performed on the audio segment sequence.

6. The method according to claim 3, characterized in that, The identification of the sound type in the non-human voice segment includes: Extract the audio features corresponding to the non-human voice segments; Based on a preset ambient sound category library, the similarity between the audio features corresponding to the non-human voice segment and the reference features in the ambient sound category library is calculated; the reference features are used to represent the audio features of each category of ambient sound in the ambient sound category library. Based on the similarity, the sound type corresponding to the non-human voice segment is determined.

7. The method according to any one of claims 1 to 6, characterized in that, The process of identifying target audio segments based on a speech recognition model and historical recognition results, and converting the target audio segments into text sequences, includes: Read the acoustic features of the target audio segment; Based on the acoustic features, semantic features of texts that match the acoustic features are retrieved from the historical recognition results; Based on the third start and end time of the target audio segment, determine the non-human voice segment located within the third start and end time. The acoustic features, the semantic features, and the non-human voice segments located within the third start and end time period are fused to obtain the fused features corresponding to the target audio segment; Based on the fusion features, a text sequence corresponding to the target audio segment is generated.

8. The method according to any one of claims 1 to 7, characterized in that, The method further includes: Acquire training audio signals; the training audio signals include mixed speech samples corresponding to multiple speakers and clean speech samples corresponding to each speaker; The training audio signal is subjected to iterative noise addition processing to generate a noise-added audio signal with increasing noise intensity; Based on the noise-added audio signal and the training audio signal, train the initial diffusion model; Based on the loss function corresponding to the initial diffusion model, the parameters of the initial diffusion model are updated to obtain the diffusion model.

9. The method according to any one of claims 1 to 8, characterized in that, The method further includes: Acquire a training speech signal, training text corresponding to the training speech signal, a reference speech signal, and reference text corresponding to the reference speech signal; the reference speech signal includes a speech signal aligned with the context features of the training speech signal. The training speech signal is recognized using the initial speech recognition model to generate an initial recognition result; Based on the semantic difference between the initial recognition result and the training text, the parameters of the initial speech recognition model are updated; Based on the training speech signal, the reference speech signal, and the reference text, the updated speech recognition model is trained with context awareness.

10. A computing device, characterized in that, include: Processor and memory; The processor and the memory are coupled together; The memory is used to store program instructions; The processor is used to execute the program instructions to perform the speech recognition method as described in any one of claims 1 to 9.