Voice signal processing method and device, electronic equipment, medium and program product

By performing separate processing and feature comparison on the voice acquisition array, the problem of intelligent devices being unable to accurately distinguish the voice signals of the target object and the interference object in multi-turn dialogues was solved, and high-precision voice interaction was achieved when the target object's position moved.

CN122135732APending Publication Date: 2026-06-02BEIJING XIAOMI MOBILE SOFTWARE CO LTD +1

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING XIAOMI MOBILE SOFTWARE CO LTD
Filing Date
2024-12-02
Publication Date
2026-06-02

Smart Images

  • Figure CN122135732A_ABST
    Figure CN122135732A_ABST
Patent Text Reader

Abstract

This disclosure relates to a speech signal processing method, apparatus, electronic device, medium, and program product. The speech signal processing method includes: separating and processing speech observation signals acquired by a speech acquisition array to obtain independent source speech signals; when human speech is detected, extracting a first audio mixing feature of each independent source speech signal; determining a target independent source speech signal based on the first audio mixing feature of each independent source speech signal and the target audio mixing feature of the target object; and performing speech recognition on the target independent source speech signal. Thus, even if the location of the target object moves during voice interaction, the target independent source speech signal containing the target object's speech can be accurately determined, thereby accurately distinguishing the target object's speech signal from the speech signal of interfering objects. This improves the accuracy of voice interaction between smart devices and target objects, enhancing the user experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of signal processing technology, and in particular to a speech signal processing method, apparatus, electronic device, medium, and program product. Background Technology

[0002] With the rapid development of artificial intelligence technology, voice interaction between smart devices and users has become an important branch. This interaction method not only greatly enhances the user experience but also brings convenience and intelligence to various application scenarios. In the application of voice interaction in smart devices, smart cockpits, smart speakers, and smart TVs have become important tools for home entertainment and information retrieval. Users only need to speak their needs, and these devices can respond quickly, playing music, providing news, setting reminders, and so on.

[0003] With the application of large-scale model technology, the intelligence level of voice interaction systems has been significantly enhanced, and users are more willing to communicate with smart devices. Driven by this trend, multi-turn dialogue technology has become increasingly practical. Multi-turn dialogue refers to multiple interactions after a single wake-up call, making the voice interaction experience smoother and more natural. In multi-turn dialogue, the front-end signal processing module of the smart device determines the location of the speaker based on the wake-up word, recognizing and interacting only with the speech at that location. That is, the front-end signal processing module distinguishes between the speaker and interfering objects based solely on the speaker's location when the wake-up word is issued. However, because the speaker may move during interaction with the smart device, the algorithm that determines the speaker's location solely based on the wake-up word in the front-end signal processing is no longer applicable, affecting the accuracy of voice interaction between the smart device and the user. Summary of the Invention

[0004] To overcome the problems existing in related technologies, this disclosure provides a speech signal processing method, apparatus, electronic device, medium, and program product.

[0005] According to a first aspect of the present disclosure, a speech signal processing method is provided, the method comprising: The voice observation signals acquired by the voice acquisition array are separated and processed to obtain independent source voice signals corresponding to each voice acquisition device in the voice acquisition array; When human speech is detected, a first audio mixing feature is extracted from each of the independent source speech signals; The target independent source speech signal is determined based on the first audio mixing feature of each independent source speech signal and the target audio mixing feature of the target object; Speech recognition is performed on the target independent source speech signal.

[0006] Optionally, determining the target independent source speech signal based on the first audio mixing feature of each of the independent source speech signals and the target audio mixing feature of the target object includes: Based on the first audio mixing feature of each independent source speech signal and the target audio mixing feature of the target object, determine the similarity between the first audio mixing feature of each independent source speech signal and the target audio mixing feature; The target independent source speech signal is determined based on the similarity of each of the independent source speech signals.

[0007] Optionally, determining the target independent source speech signal based on the similarity of each of the independent source speech signals includes: If there are multiple candidate independent source speech signals with a similarity greater than or equal to a preset threshold, then the energy characteristics of each candidate independent source speech signal are determined. The candidate independent source speech signal with the largest energy feature is determined as the target independent source speech signal.

[0008] Optionally, determining the target independent source speech signal based on the similarity of each of the independent source speech signals further includes: If the number of candidate independent source speech signals with a similarity greater than or equal to a preset threshold is one, then the candidate independent source speech signal is determined as the target independent source speech signal.

[0009] Optionally, determining the target independent source speech signal based on the similarity of each of the independent source speech signals further includes: If the similarity of each independent source speech signal is less than a preset threshold, then the energy characteristics of each independent source speech signal are determined. The independent source speech signal with the smallest energy feature is identified as the target independent source speech signal.

[0010] Optionally, when detecting human speech, extracting a first audio mixing feature for each of the independent source speech signals includes: When human voice is detected, the first audio mixing feature of each independent source speech signal is extracted within a preset time period, wherein the start time of the preset time period is the time when human voice is detected.

[0011] Optionally, the target audio mixing features of the target object are determined by the following methods: In response to the detection of a wake-up event, a target audio mixing feature of the target object is determined, wherein the target audio mixing feature includes at least the voiceprint feature and / or fundamental frequency feature of the target object.

[0012] Optionally, the method for determining the target audio mixing features further includes: Extract and cache the second audio mixing features of each of the independent source speech signals; In response to detecting a wake-up event, the target audio mixing characteristics of the target object are determined, including: In response to the detection of a target wake word, the independent source speech signal containing the target wake word is determined as the independent source speech signal corresponding to the target object that triggered the wake-up event; Determine the time period of the target wake word in the independent source speech signal corresponding to the target object; The target audio mixing features of the target object are determined based on the second audio mixing features of the independent source speech signal corresponding to the target object and the time period.

[0013] Optionally, determining the target audio mixing features of the target object based on the second audio mixing features of the independent source speech signals corresponding to the target object and the time period includes: In the second audio mixing features of the independent source speech signal corresponding to the target object, a third audio mixing feature located within the time period is determined; Based on the third audio mixing feature, the target audio mixing feature of the target object is determined.

[0014] Optionally, in response to detecting a wake-up event, determining the target audio mixing features of the target object includes: In response to the detection of a key wake-up event, human voice is detected in each independent source speech signal; In response to the detection of the human voice, the audio mixing features of the human voice are determined as the target audio mixing features of the target object.

[0015] Optionally, the method further includes: For each independent source speech signal, determine the signal-to-noise ratio of the independent source speech signal, and / or the confidence level of the audio mixing feature of the independent source speech signal; based on the signal-to-noise ratio and / or the confidence level of the audio mixing feature, determine whether there is a human voice signal in the independent source speech signal. When it is determined that the human voice signal is present in at least one of the independent source speech signals, it is determined that the human voice signal has been detected.

[0016] Optionally, determining whether a human voice signal exists in the independent source speech signal based on the signal-to-noise ratio and / or the confidence level of the audio mixing features includes: If the signal-to-noise ratio is greater than or equal to a first threshold, and / or the confidence level is greater than or equal to a second threshold, then it is determined that a human voice signal exists in the independent source speech signal.

[0017] Optionally, the separation process includes at least blind source separation; the separation process further includes echo cancellation and / or noise reverberation suppression.

[0018] According to a second aspect of the present disclosure, a speech signal processing apparatus is provided, the apparatus comprising: The processing module is configured to separate and process the speech observation signals acquired by the speech acquisition array to obtain independent source speech signals corresponding to each speech acquisition device in the speech acquisition array; The first extraction module is configured to extract a first audio mixing feature of each of the independent source speech signals when human speech is detected; The first determining module is configured to determine the target independent source speech signal based on the first audio mixing feature of each of the independent source speech signals and the target audio mixing feature of the target object; The recognition module is configured to perform speech recognition on the target independent source speech signal.

[0019] According to a third aspect of the present disclosure, an electronic device is provided, comprising: a processor; Memory used to store processor-executable instructions; The processor is configured to execute the instructions to enable the electronic device to implement the speech signal processing method as described in any one of the first aspects of the embodiments of this disclosure.

[0020] According to a fourth aspect of the present disclosure, a computer-readable storage medium is provided that stores computer program instructions thereon, which, when executed by a processor, implement the steps of the speech signal processing method provided in the first aspect of the present disclosure.

[0021] According to a fifth aspect of the present disclosure, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps of the speech signal processing method provided in the first aspect of the present disclosure.

[0022] By employing the above technical solution, upon detecting human voice, the first audio mixing feature of each independent source speech signal is extracted. Based on the first audio mixing feature of each independent source speech signal and the target audio mixing feature of the target object, the target independent source speech signal is determined, and speech recognition is performed on the target independent source speech signal. Since the first audio mixing feature of each independent source speech signal is compared with the target audio mixing feature of the target object upon detecting human voice to determine the target independent source speech signal requiring speech recognition, even if the target object's position moves during voice interaction, the target independent source speech signal containing the target object's voice can be accurately determined. This allows for accurate differentiation between the target object's speech signal and the speech signal of interfering objects. Thus, the accuracy of voice interaction between smart devices and target objects is improved, enhancing the user experience.

[0023] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description

[0024] The accompanying drawings, which are incorporated in and form a part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure.

[0025] Figure 1 This is a flowchart illustrating a speech signal processing method according to an exemplary embodiment.

[0026] Figure 2 This is a flowchart illustrating another speech signal processing method according to an exemplary embodiment.

[0027] Figure 3 This is a block diagram illustrating a speech signal processing apparatus according to an exemplary embodiment.

[0028] Figure 4 This is a block diagram illustrating an electronic device according to an exemplary embodiment. Detailed Implementation

[0029] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numerals in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this disclosure as detailed in the appended claims.

[0030] It should be noted that all actions involving the acquisition of signals, information, or data in this disclosure are carried out in compliance with the relevant data protection laws and policies of the country where the location is situated, and with authorization from the owner of the relevant device.

[0031] In related technologies, the common approach is to incorporate directional information into the speaker separation algorithm, combining beamforming and blind source separation algorithms. This allows speaker signals from a specific direction to be fixed in a specific channel after blind source separation, thus solving the channel ambiguity problem inherent in blind source separation algorithms. However, this method is suitable for scenarios with relatively fixed speaker positions, such as multi-zone smart cockpits. For scenarios where the target speaker's position moves, or even while moving and speaking, this method cannot distinguish between the target and the interfering object.

[0032] In view of this, this disclosure provides a speech signal processing method, apparatus, electronic device, medium, and program product. Upon detecting human speech, it extracts a first audio mixing feature of each independent source speech signal. Based on the first audio mixing feature of each independent source speech signal and the target audio mixing feature of the target object, it determines the target independent source speech signal and performs speech recognition on the target independent source speech signal. Because the first audio mixing feature of each independent source speech signal is compared with the target audio mixing feature of the target object upon detecting human speech to determine the target independent source speech signal requiring speech recognition, even if the target object's position moves during voice interaction, the target independent source speech signal containing the target object's speech can be accurately determined. This allows for accurate differentiation between the target object's speech signal and the speech signal of interfering objects. Thus, the accuracy of voice interaction between smart devices and target objects is improved, enhancing the user experience.

[0033] Figure 1 This is a flowchart illustrating a speech signal processing method according to an exemplary embodiment. The speech signal processing method can be used in smart devices that can use voice interaction, such as smart cockpits, speakers, and televisions. Figure 1 As shown, the speech signal processing method may include the following steps.

[0034] In step S11, the voice observation signals acquired by the voice acquisition array are separated to obtain independent source voice signals corresponding to each voice acquisition device in the voice acquisition array.

[0035] In the field of signal processing technology, observed signal refers to the signal acquired by a data acquisition device (such as a sensor). For example, in the field of speech signal processing technology, speech observed signal refers to the sound signal acquired by a speech acquisition device (such as a microphone or other sound pickup device).

[0036] In some implementations, a voice acquisition array is disposed in a smart device and may include multiple voice acquisition devices. Each voice acquisition device may be a microphone. When the voice acquisition array includes multiple voice acquisition devices, the voice observation signal acquired by the voice acquisition array may include the voice observation signals acquired by each individual voice acquisition device.

[0037] In some embodiments, the separation process includes at least blind source separation, that is, blind source separation is performed on the speech observation signal acquired by the speech acquisition array to obtain an independent source speech signal corresponding to each speech acquisition device. For example, if the speech acquisition array includes two speech acquisition devices, blind source separation of the speech observation signal can obtain two independent source speech signals.

[0038] By employing the above technical solution, blind source separation of the speech signal is performed to obtain the independent source speech signal corresponding to each speech acquisition device. In this way, the speech signals of multiple objects can be distinguished, or the speech of the target object and environmental noise can be distinguished.

[0039] In some embodiments, the separation process includes echo cancellation and / or noise reverberation suppression in addition to blind source separation. For example, the speech observation signal acquired by the speech acquisition array is first subjected to echo cancellation and / or noise reverberation suppression to obtain a processed speech signal. Then, blind source separation is performed on the processed speech signal to obtain the independent source speech signal corresponding to each speech acquisition device.

[0040] Thus, echo cancellation and / or noise reverberation suppression processing of the speech observation signals acquired by the speech acquisition array can eliminate noise to a certain extent and improve the quality of the speech signal.

[0041] It should be understood that the technique of separating and processing speech observation signals to obtain independent source speech signals is a relatively mature technique, and this disclosure does not make any specific limitations on it.

[0042] In step S12, when human voice is detected, the first audio mixing feature of each independent source speech signal is extracted.

[0043] In this disclosure, each independent source speech signal can be detected. When human voice is detected in at least one independent source speech signal, it is determined that human voice has been detected. At this time, the first audio mixing feature of each independent source speech signal is extracted.

[0044] In this disclosure, the audio mixing features may include voiceprint features and / or fundamental frequency features. For example, the first audio mixing feature, the second audio mixing feature, the third audio mixing feature, and the target audio mixing feature are all voiceprint features; or, the first audio mixing feature, the second audio mixing feature, the third audio mixing feature, and the target audio mixing feature are all fundamental frequency features; or, the first audio mixing feature, the second audio mixing feature, the third audio mixing feature, and the target audio mixing feature are all voiceprint features and fundamental frequency features.

[0045] Voiceprint characteristics include features such as pitch, timbre, intensity, sound wavelength, frequency, and rhythmic variations that reflect the unique speaking characteristics of different people. Because different people have different oral cavity, vocal cords, and other vocal organs, and different speaking habits, each person also has different voiceprint characteristics. Fundamental frequency characteristics refer to the lowest frequency component in a sound or vibration waveform, which determines the pitch of the sound.

[0046] In some embodiments, audio mixing features can be extracted using various techniques such as time-domain autocorrelation, cepstral analysis, frequency-domain methods, and deep learning-based methods. For example, each independent source speech signal can be input into a deep learning-based neural network to obtain the first audio mixing feature of each independent source speech signal.

[0047] It should be understood that the methods for extracting audio mixing features mentioned above are relatively mature technologies, and this disclosure does not impose any specific limitations on them.

[0048] In step S13, the target independent source speech signal is determined based on the first audio mixing feature of each independent source speech signal and the target audio mixing feature of the target object.

[0049] In step S14, speech recognition is performed on the target independent source speech signal.

[0050] The target object refers to the object that interacts with the smart device via voice. For example, the target object could be the object that triggers a wake-up event, such as by entering a wake word or by performing a button press to wake up.

[0051] In this disclosure, a first audio mixing feature of each independent source speech signal and a target audio mixing feature of the target object are used to determine the target independent source speech signal in real time or periodically from multiple independent source speech signals. Thus, during the interaction process, regardless of the location of the target object, the target independent source speech signal is obtained by comparing the first audio mixing feature of each independent source speech signal with the target audio mixing feature of the target object. Therefore, even if the location of the object interacting with the smart device changes during the interaction process, the target independent source speech signal can still be determined and speech recognition can be performed on the target independent source speech signal.

[0052] For example, after identifying the target independent source speech signal, the target independent source speech signal can be input into an Automatic Speech Recognition (ASR) module, which will then convert the speech into computer-readable text.

[0053] By employing the above technical solution, upon detecting human voice, the first audio mixing feature of each independent source speech signal is extracted. Based on the first audio mixing feature of each independent source speech signal and the target audio mixing feature of the target object, the target independent source speech signal is determined, and speech recognition is performed on the target independent source speech signal. Since the first audio mixing feature of each independent source speech signal is compared with the target audio mixing feature of the target object upon detecting human voice to determine the target independent source speech signal requiring speech recognition, even if the target object's position moves during voice interaction, the target independent source speech signal containing the target object's voice can be accurately determined. This allows for accurate differentiation between the target object's speech signal and the speech signal of interfering objects. Thus, the accuracy of voice interaction between smart devices and target objects is improved, enhancing the user experience.

[0054] To facilitate a better understanding of the speech signal processing method provided in this disclosure by those skilled in the art, a complete embodiment is described below.

[0055] The following explains the methods for detecting human speech.

[0056] In some embodiments, the method may further include: For each independent source speech signal, determine the signal-to-noise ratio of the independent source speech signal, and / or the confidence level of the audio mixing features of the independent source speech signal, and determine whether there is a human voice signal in the independent source speech signal based on the signal-to-noise ratio and / or the confidence level of the audio mixing features; A human voice signal is determined to be detected when it is determined that a human voice signal is present in at least one independent source speech signal.

[0057] In one implementation, the presence of a human voice signal in an independent source speech signal can be determined solely based on the signal-to-noise ratio (SNR) of the independent source speech signal. For example, the presence of a human voice signal in the independent source speech signal is determined when the SNR is greater than or equal to a first threshold.

[0058] In another implementation, the presence of a human voice signal in an independent source speech signal can be determined solely based on the confidence level of the audio mixing features of the independent source speech signals. Here, the confidence level represents the probability of the human voice signal being present. For example, if the confidence level of the audio mixing features is greater than or equal to a first threshold, it is determined that a human voice signal is present in the independent source speech signal.

[0059] However, to accurately identify the presence of human voice signals in independent source speech signals and reduce the influence of sudden noises or ambient whispers, another implementation method can employ a more rigorous Voice Activity Detection (VAD) approach. This VAD method combines energy-based VAD detection with a confidence assessment of audio mixture features. While energy-based VAD detection is generally accurate for speech segments, it is sensitive to noise, particularly sudden noise, which may lead to misidentification of sudden noise as speech. Speech signals typically have prominent fundamental frequency characteristics and relatively high confidence, while noise signals often have less pronounced fundamental frequency characteristics and lower confidence. Therefore, combining energy-based VAD detection with a confidence assessment of audio mixture features can accurately determine the presence of human voice signals.

[0060] In this implementation, for each independent source speech signal, the signal-to-noise ratio (SNR) of the independent source speech signal and the confidence level of the audio mixing features of the independent source speech signal are determined separately. For example, for each frame of speech signal, the noise energy of that frame is estimated using the minimum tracking method. Then, the ratio of the original audio energy to the noise energy of that frame is determined as the SNR of that frame. The SNRs of multiple frames of speech signals are smoothed to obtain the final SNR. Furthermore, the confidence level of the audio mixing features can be obtained during the extraction of the audio mixing features; for example, the audio mixing features and their confidence levels can be obtained using a deep learning-based neural network.

[0061] After obtaining the confidence level of the signal-to-noise ratio and audio mixing features of each independent source speech signal, the presence of human voice signals in the independent source speech signals is determined based on the confidence level of the signal-to-noise ratio and audio mixing features.

[0062] For example, when the signal-to-noise ratio is greater than or equal to a first threshold and the confidence level is greater than or equal to a second threshold, it is determined that a human voice signal exists in the independent source speech signal. The thresholds can be set empirically, and this disclosure does not impose specific limitations on them.

[0063] By adopting the above technical solution, the presence of human voice signals in independent source speech signals can be determined by the confidence level of signal-to-noise ratio and / or audio mixing features, thereby improving the accuracy of human voice signal detection and thus improving the accuracy of speech signal processing and subsequent speech interaction.

[0064] After detecting human speech in the manner described above or otherwise, the first audio mixing feature of each independent source speech signal can be extracted.

[0065] Considering that the fundamental frequency fluctuation becomes more significant as the speech length increases, and that the fundamental frequency of the initial part of the speech is closer to that of the wake word, it can more accurately reflect the actual fundamental frequency of the target object. Furthermore, by buffering only the initial short segment of the speech signal, and then outputting the entire target independent source speech signal for speech recognition after determining the target independent source speech signal, and outputting subsequent audio in real time except for the initial short segment, the real-time performance and efficiency of the subsequent speech recognition stage can be guaranteed, resulting in a smoother and more natural presentation of the recognition results on the screen. Therefore, in some embodiments, when human voice is detected, the first audio mixing feature of each independent source speech signal within a preset time period is extracted, where the start time of the preset time period is the moment human voice is detected.

[0066] For example, when detecting human speech, the first audio mixing feature of the speech signal within an initial short time period (e.g., within 60 milliseconds) is extracted from each independent source speech signal.

[0067] By employing the above technical solution, upon detecting human speech, the first audio mixing feature of each independent source speech signal within a preset time period is extracted, with the start time of the preset time period being the moment human speech is detected. This improves the accuracy of determining the target independent source speech signal when subsequently based on the first audio mixing feature and the target audio mixing feature of the target object. Furthermore, it ensures real-time and efficient processing by the subsequent automatic speech recognition module, resulting in a smoother and more natural presentation of the recognition results on the screen, thus enhancing the user experience.

[0068] The method for determining the target audio mixing features of a target object is described below. In this disclosure, the target audio mixing features of a target object are determined by determining the target audio mixing features.

[0069] In this disclosure, the target audio mixing features of a target object are determined by the following method: in response to the detection of a wake-up event, the target audio mixing features of the target object are determined, wherein the target audio mixing features include at least the voiceprint features and / or fundamental frequency features of the target object.

[0070] Considering that the user performing the wake-up operation is usually the target object for voice interaction with the smart device, this disclosure determines the target audio mixing features of the target object when a wake-up event is detected. Here, detecting a wake-up event can refer to detecting a target wake-up word or detecting other operations such as button presses.

[0071] In some embodiments, determining the target audio mixing features of a target object may further include: extracting and caching the second audio mixing features of each independent source speech signal; and determining the target audio mixing features of a target object in response to detecting a wake-up event, which may include: determining the independent source speech signal containing the target wake-up word as the independent source speech signal corresponding to the target object that triggered the wake-up event in response to detecting a target wake-up word; determining the time period of the target wake-up word in the independent source speech signal corresponding to the target object; and determining the target audio mixing features of the target object based on the second audio mixing features and the time period of the independent source speech signal corresponding to the target object.

[0072] It should be understood that the first audio mixing feature and the second audio mixing feature are audio mixing features from different stages. In this embodiment, when the voice observation signal is acquired and the independent source voice signal corresponding to each voice acquisition device is obtained, the audio mixing feature of each independent source voice signal is extracted and cached. The audio mixing feature of each independent source voice signal extracted between the time of acquiring the voice observation signal and the time of detecting the target wake-up word is denoted as the second audio mixing feature. That is to say, the second audio mixing feature is the audio mixing feature of the voice signal acquired by the voice acquisition array before the target wake-up word is detected; the second audio mixing feature is the audio mixing feature of the voice signal before the smart device is woken up. The first audio mixing feature refers to the audio mixing feature of each independent source voice signal extracted after the target wake-up word is detected, i.e., the electronic device is woken up, and then human voice is detected again. That is to say, the first audio mixing feature refers to the audio mixing feature of the voice signal when human voice is detected again after the smart device is woken up.

[0073] The target wake-up word refers to a preset word used to wake up a smart device. A target wake-up word typically includes multiple characters, for example, two to four characters. In this disclosure, detecting a target wake-up word can mean detecting a portion of the target wake-up word, such as detecting the first two characters, or it can mean detecting the complete target wake-up word. To avoid misidentifying the target wake-up word and improve the accuracy of the target audio mixing features of the identified target object, this disclosure uses the example of detecting the complete target wake-up word for description.

[0074] In this embodiment, upon detecting a target wake-up word, the independent source speech signal containing the target wake-up word is determined, and this independent source speech signal is identified as the independent source speech signal corresponding to the target object that triggered the wake-up event. Furthermore, since the independent source speech signal containing the target wake-up word includes not only the speech signal corresponding to the target wake-up word but also other speech information such as speech signals corresponding to environmental noise, after determining the independent source speech signal corresponding to the target object, the time period of the target wake-up word within the independent source speech signal corresponding to the target object can be further determined. Then, based on the second audio mixing feature and the time period of the independent source speech signal corresponding to the target object, the target audio mixing feature of the target object is determined.

[0075] For example, determining the target audio mixing features of the target object based on the second audio mixing features and time period of the independent source speech signal corresponding to the target object may include: determining a third audio mixing feature located within the time period in the second audio mixing features of the independent source speech signal corresponding to the target object; and determining the target audio mixing features of the target object based on the third audio mixing feature.

[0076] For example, assuming the time period for the target wake word in the independent source speech signal corresponding to the target object is t1 to t2, then the audio mixing features of the speech signal in the independent source speech signal corresponding to the target object within the time period t1 to t2 are determined as the third audio mixing feature. That is, the second audio mixing feature located within the time period t1 to t2 is determined as the third audio mixing feature.

[0077] The time period from t1 to t2 typically includes multiple frames of speech signals; that is, the third audio mixing feature includes the audio mixing features of multiple frames of speech signals. A specific implementation method for determining the target audio mixing feature of the target object based on the third audio mixing feature can be: determining the average or median of the third audio mixing features of the multiple frames of speech signals as the target audio mixing feature of the target object.

[0078] By employing the above technical solution, during each voice interaction, the audio mixing features of the speech signal corresponding to the target wake-up word can be determined as the target audio mixing features of the target object, eliminating the need to pre-store the target audio mixing features of the target object. Furthermore, since the audio mixing features of the speech signal corresponding to the target wake-up word during the current voice interaction are determined as the target audio mixing features of the target object, the accuracy of determining the target audio mixing features is improved, thereby enhancing the accuracy of determining the target independent source speech signal based on the target audio mixing features.

[0079] In addition, when a target wake word is detected, the start and end times of the target wake word can be determined, that is, the time period of the target wake word can be determined. Then, the speech signal within that time period can be determined in the independent source speech signal where the target wake word is located, and the audio mixing features of the speech signal within that time period can be extracted as the target audio mixing features of the target object.

[0080] In other embodiments, determining the target audio mixing features of the target object in response to detecting a wake-up event may further include: detecting human voice speech in each independent source speech signal in response to detecting a key wake-up event; and determining the audio mixing features of the human voice speech as the target audio mixing features of the target object in response to detecting human voice speech.

[0081] In this embodiment, after the smart device is woken up by a button, the first user to interact is identified as the target object. Therefore, when the button wake-up event is detected, human voice is detected in each independent source speech signal. Regardless of which independent source speech signal the human voice is detected in, the audio mixing features of the first detected human voice are identified as the target audio mixing features of the target object.

[0082] It should be understood that, in this embodiment, after a key wake-up event is detected, in addition to determining the audio mixing features of the first detected human voice as the target audio mixing features of the target object, the human voice can also be recognized.

[0083] By adopting the above technical solution, in the scenario of waking up intelligence by pressing a button, the audio mixing features of the detected human voice can be identified as the target audio mixing features of the target object, which improves the flexibility of determining the target audio mixing features.

[0084] After determining the target audio mixing features of the target object according to the above scheme, human voice is detected again. When human voice is detected again while the smart device is in a wake-up state, the first audio mixing feature of each independent source speech signal is extracted, or the first audio mixing feature of each independent source speech signal within a preset time period is extracted. Then, based on the first audio mixing features and the target audio mixing features of the target object, the target independent source speech signal is determined.

[0085] In some embodiments, Figure 1In step S13, determining the target independent source speech signal based on the first audio mixing feature of each independent source speech signal and the target audio mixing feature of the target object can include: determining the similarity between the first audio mixing feature of each independent source speech signal and the target audio mixing feature of the target object; and determining the target independent source speech signal based on the similarity between each independent source speech signal. Here, similarity refers to the degree of similarity between the first audio mixing feature of the independent source speech signal and the target audio mixing feature. The higher the similarity, the higher the probability that the independent source speech signal is the speech signal of the target object.

[0086] In one implementation, the method for determining the target independent source speech signal based on the similarity of each independent source speech signal is as follows: if there are multiple candidate independent source speech signals with a similarity greater than or equal to a preset threshold, then the energy characteristics of each candidate independent source speech signal are determined; the candidate independent source speech signal with the largest energy characteristic is determined as the target independent source speech signal. The preset threshold can be set empirically.

[0087] When there are multiple candidate independent source speech signals, representing speech signals that all multiple candidate independent source speech signals belong to the target object, the target independent source speech signal can be determined based on the energy of each candidate independent source speech signal. For example, the candidate independent source speech signal with the largest energy feature can be determined as the target independent source speech signal.

[0088] In another implementation, if the number of candidate independent source speech signals with a similarity greater than or equal to a preset threshold is one, then the candidate independent source speech signal is determined as the target independent source speech signal.

[0089] In this embodiment, if the similarity of only one candidate independent source speech signal is greater than or equal to a preset threshold, then the candidate independent source speech signal is determined to be the speech signal of the target object. At this time, the candidate independent source speech signal is determined to be the target independent source speech signal.

[0090] In another implementation, if the similarity of each independent source speech signal is less than a preset threshold, then the energy characteristics of each independent source speech signal are determined; the independent source speech signal with the smallest energy characteristics is determined as the target independent source speech signal.

[0091] In this embodiment, if the similarity of each independent source speech signal is less than a preset threshold, then each independent source speech signal is determined not to belong to the target object's speech signal. In this case, to avoid identifying incorrect information, the independent source speech signal with the lowest energy characteristic can be identified as the target independent source speech signal, and this target independent source speech signal can be identified. This reduces the false recognition rate and the risk of incorrect information being identified.

[0092] Figure 2 This is a flowchart illustrating another speech signal processing method according to an exemplary embodiment. For example... Figure 2 As shown, the method may include the following steps.

[0093] In step S21, the voice observation signals acquired by the voice acquisition array are separated to obtain independent source voice signals corresponding to each voice acquisition device in the voice acquisition array.

[0094] Steps S22 to S25 can be executed after step S21, or steps S26 and S27 can be executed after step S21.

[0095] In step S22, the second audio mixing features of each independent source speech signal are extracted and cached respectively.

[0096] In step S23, in response to detecting the target wake-up word, the independent source speech signal containing the target wake-up word is determined as the independent source speech signal corresponding to the target object that triggered the wake-up event.

[0097] In step S24, the time period of the target wake word in the independent source speech signal corresponding to the target object is determined.

[0098] In step S25, the target audio mixing features of the target object are determined based on the second audio mixing features and time period of the independent source speech signal corresponding to the target object.

[0099] In step S26, in response to the detection of a key wake-up event, human voice is detected in each independent source speech signal.

[0100] In step S27, in response to the detection of human voice speech, the audio mixing features of the human voice speech are determined as the target audio mixing features of the target object.

[0101] After completing step S25 or step S27, proceed to steps S28 through S210.

[0102] In step S28, after the device is woken up, when human voice is detected, the first audio mixing feature of each independent source speech signal is extracted.

[0103] In step S29, the target independent source speech signal is determined based on the first audio mixing feature of each independent source speech signal and the target audio mixing feature of the target object.

[0104] In step S210, speech recognition is performed on the target independent source speech signal.

[0105] about Figure 2The specific methods of each step in the process have been described in detail in the above embodiments, and will not be elaborated here.

[0106] Based on the same inventive concept, this disclosure also provides a speech signal processing device. Figure 3 This is a block diagram illustrating a speech signal processing apparatus according to an exemplary embodiment. Figure 3 As shown, the speech signal processing device 300 may include: The processing module 301 is configured to separate and process the voice observation signal acquired by the voice acquisition array to obtain an independent source voice signal corresponding to each voice acquisition device in the voice acquisition array. The first extraction module 302 is configured to extract a first audio mixing feature of each of the independent source speech signals when human voice is detected. The first determining module 303 is configured to determine the target independent source speech signal based on the first audio mixing feature of each of the independent source speech signals and the target audio mixing feature of the target object; The recognition module 304 is configured to perform speech recognition on the target independent source speech signal.

[0107] Optionally, the first determining module 303 may include: The first determining submodule is configured to determine the similarity between the first audio mixing feature of each independent source speech signal and the target audio mixing feature based on the first audio mixing feature of each independent source speech signal and the target audio mixing feature of the target object. The second determining submodule is configured to determine the target independent source speech signal based on the similarity of each of the independent source speech signals.

[0108] Optionally, the second determining submodule is configured as follows: If there are multiple candidate independent source speech signals with a similarity greater than or equal to a preset threshold, then the energy characteristics of each candidate independent source speech signal are determined. The candidate independent source speech signal with the largest energy feature is determined as the target independent source speech signal.

[0109] Optionally, the second determining submodule is configured to: if the number of candidate independent source speech signals with a similarity greater than or equal to a preset threshold is one, then determine the candidate independent source speech signal as the target independent source speech signal.

[0110] Optionally, the second determining submodule is configured as follows: If the similarity of each independent source speech signal is less than a preset threshold, then the energy characteristics of each independent source speech signal are determined. The independent source speech signal with the smallest energy feature is identified as the target independent source speech signal.

[0111] Optionally, the first extraction module 302 is configured to: when human voice is detected, extract the first audio mixing feature of each independent source speech signal within a preset time period, wherein the start time of the preset time period is the time when human voice is detected.

[0112] Optionally, the target audio mixing features of the target object are determined by the following methods: In response to the detection of a wake-up event, a target audio mixing feature of the target object is determined, wherein the target audio mixing feature includes at least the voiceprint feature and / or fundamental frequency feature of the target object.

[0113] Optionally, the method for determining the target audio mixing features further includes: Extract and cache the second audio mixing features of each of the independent source speech signals; In response to detecting a wake-up event, the target audio mixing characteristics of the target object are determined, including: In response to the detection of a target wake word, the independent source speech signal containing the target wake word is determined as the independent source speech signal corresponding to the target object that triggered the wake-up event; Determine the time period of the target wake word in the independent source speech signal corresponding to the target object; The target audio mixing features of the target object are determined based on the second audio mixing features of the independent source speech signal corresponding to the target object and the time period.

[0114] Optionally, determining the target audio mixing features of the target object based on the second audio mixing features of the independent source speech signals corresponding to the target object and the time period includes: In the second audio mixing features of the independent source speech signal corresponding to the target object, a third audio mixing feature located within the time period is determined; Based on the third audio mixing feature, the target audio mixing feature of the target object is determined.

[0115] Optionally, in response to detecting a wake-up event, determining the target audio mixing features of the target object includes: In response to the detection of a key wake-up event, human voice is detected in each independent source speech signal; In response to the detection of the human voice, the audio mixing features of the human voice are determined as the target audio mixing features of the target object.

[0116] Optionally, the speech signal processing device 300 may include: The second determining module is configured to, for each independent source speech signal, determine the signal-to-noise ratio of the independent source speech signal, and / or the confidence level of the audio mixing features of the independent source speech signal, and determine whether there is a human voice signal in the independent source speech signal based on the signal-to-noise ratio and / or the confidence level of the audio mixing features; The third determining module is configured to determine that the human voice signal is detected when it is determined that the human voice signal is present in at least one of the independent source speech signals.

[0117] Optionally, the second determining module is configured to determine that a human voice signal exists in the independent source speech signal when the signal-to-noise ratio is greater than or equal to a first threshold and / or the confidence level is greater than or equal to a second threshold.

[0118] Optionally, the separation process includes at least blind source separation; the separation process further includes echo cancellation and / or noise reverberation suppression.

[0119] Regarding the apparatus in the above embodiments, the specific manner in which each module performs its operation has been described in detail in the embodiments related to the method, and will not be elaborated upon here.

[0120] This disclosure also provides a computer-readable storage medium having stored thereon computer program instructions that, when executed by a processor, implement the steps of the speech signal processing method provided in this disclosure.

[0121] Figure 4 This is a block diagram illustrating an electronic device according to an exemplary embodiment. The electronic device can be a smart device or a vehicle. For example, electronic device 800 can be a mobile phone, computer, digital broadcasting terminal, messaging device, game console, tablet device, medical device, fitness equipment, personal digital assistant, etc.

[0122] Reference Figure 4 The electronic device 800 may include one or more of the following components: processing component 802, memory 804, power supply component 806, multimedia component 808, audio component 810, input / output interface 812, sensor component 814, and communication component 816.

[0123] Processing component 802 typically controls the overall operation of electronic device 800, such as operations associated with display, telephone calls, data communication, camera operation, and recording. Processing component 802 may include one or more processors 820 to execute instructions to complete all or part of the steps of the aforementioned voice signal processing method. Furthermore, processing component 802 may include one or more modules to facilitate interaction between processing component 802 and other components. For example, processing component 802 may include a multimedia module to facilitate interaction between multimedia component 808 and processing component 802.

[0124] Memory 804 is configured to store various types of data to support the operation of electronic device 800. Examples of such data include instructions for any application or method operating on electronic device 800, contact data, phonebook data, messages, pictures, videos, etc. Memory 804 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.

[0125] Power supply component 806 provides power to various components of electronic device 800. Power supply component 806 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to electronic device 800.

[0126] Multimedia component 808 includes a screen that provides an output interface between the electronic device 800 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touchscreen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors may sense not only the boundaries of the touch or swipe action but also the duration and pressure associated with the touch or swipe operation. In some embodiments, multimedia component 808 includes a front-facing camera and / or a rear-facing camera. When the electronic device 800 is in an operating mode, such as a shooting mode or a video mode, the front-facing camera and / or the rear-facing camera may receive external multimedia data. Each front-facing camera and rear-facing camera may be a fixed optical lens system or have focal length and optical zoom capabilities.

[0127] Audio component 810 is configured to output and / or input audio signals. For example, audio component 810 includes a microphone (MIC) configured to receive external audio signals when electronic device 800 is in an operating mode, such as call mode, recording mode, and voice recognition mode. The received audio signals may be further stored in memory 804 or transmitted via communication component 816. In some embodiments, audio component 810 also includes a speaker for outputting audio signals.

[0128] Input / output interface 812 provides an interface between processing component 802 and peripheral interface modules, such as keyboards, click wheels, buttons, etc. These buttons may include, but are not limited to, home buttons, volume buttons, power buttons, and lock buttons.

[0129] Sensor assembly 814 includes one or more sensors for providing state assessments of various aspects of electronic device 800. For example, sensor assembly 814 can detect the on / off state of electronic device 800, the relative positioning of components such as the display and keypad of electronic device 800, changes in position of electronic device 800 or a component of electronic device 800, the presence or absence of user contact with electronic device 800, orientation or acceleration / deceleration of electronic device 800, and temperature changes of electronic device 800. Sensor assembly 814 may include a proximity sensor configured to detect the presence of nearby objects without any physical contact. Sensor assembly 814 may also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, sensor assembly 814 may also include an accelerometer, gyroscope, magnetometer, pressure sensor, or temperature sensor.

[0130] Communication component 816 is configured to facilitate wired or wireless communication between electronic device 800 and other devices. Electronic device 800 can access wireless networks based on communication standards, such as WiFi, 2G, or 3G, or combinations thereof. In one exemplary embodiment, communication component 816 receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In one exemplary embodiment, communication component 816 also includes a near-field communication (NFC) module to facilitate short-range communication. For example, the NFC module may be implemented based on radio frequency identification (RFID) technology, Infrared Data Association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.

[0131] In an exemplary embodiment, the electronic device 800 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the above-described speech signal processing method.

[0132] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 804 including instructions, which can be executed by a processor 820 of an electronic device 800 to complete the aforementioned voice signal processing method. For example, the non-transitory computer-readable storage medium may be a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc.

[0133] In another exemplary embodiment, a computer program product is also provided, the computer program product comprising a computer program executable by a programmable device, the computer program having a code portion for performing the above-described speech signal processing method when executed by the programmable device.

[0134] It should be understood that, unless otherwise specifically indicated, features of various embodiments of this disclosure described herein can be combined with each other. As used herein, the term “and / or” includes any one of the relevant listed items and any combination of any two or more; similarly, “at least one of…” includes any one of the relevant listed items and any combination of any two or more.

[0135] The terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In this description, "a plurality of" means at least two, such as two, three, etc., unless otherwise explicitly specified.

[0136] Furthermore, the term “exemplary” is used herein to mean serving as an example, instance, or illustration. Any aspect or design described herein as “exemplary” is not necessarily to be construed as advantageous compared to other aspects or designs. Rather, the use of the term “exemplary” is intended to present the concept in a concrete manner. As used herein, the term “or” is intended to mean an inclusive “or” rather than an exclusive “or.” That is, unless otherwise specified or clear from the context, “X applies A or B” is intended to mean any of the natural inclusive arrangements. That is, “X applies A or B” satisfies any of the foregoing instances if X applies A; X applies B; or both X applies A and B. Additionally, unless otherwise specified or clear from the context to refer to the singular form, the articles “a” and “an” as used in this application and the appended claims are generally understood to mean “one or more.”

[0137] Similarly, although this disclosure has been shown and described with respect to one or more implementations, equivalent variations and modifications will occur to those skilled in the art upon reading and understanding this specification and the accompanying drawings. This disclosure includes all such modifications and variations and is limited only by the scope of the claims. In particular, with respect to the various functions performed by the components described above (e.g., elements, resources, etc.), unless otherwise indicated, the terminology used to describe such components is intended to correspond to any component (functionally equivalent) that performs the specific function of the described component, even if structurally not equivalent to the disclosed structure. Furthermore, although specific features of this disclosure may have been disclosed with respect to only one of several implementations, such features may be combined with one or more other features of other implementations, as may be desired and advantageous to any given or particular application. Moreover, with regard to the terms “comprising,” “owning,” “having,” “having,” or variations thereof as used in the detailed description or claims, such terms are intended to be inclusive in a manner similar to the term “including.”

[0138] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the appended claims.

[0139] It should be understood that this disclosure is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this disclosure is limited only by the appended claims.

Claims

1. A speech signal processing method, characterized in that, The method includes: The voice observation signals acquired by the voice acquisition array are separated and processed to obtain independent source voice signals corresponding to each voice acquisition device in the voice acquisition array; When human speech is detected, a first audio mixing feature is extracted from each of the independent source speech signals; The target independent source speech signal is determined based on the first audio mixing feature of each independent source speech signal and the target audio mixing feature of the target object; Speech recognition is performed on the target independent source speech signal.

2. The method according to claim 1, characterized in that, The step of determining the target independent source speech signal based on the first audio mixing feature of each of the independent source speech signals and the target audio mixing feature of the target object includes: Based on the first audio mixing feature of each independent source speech signal and the target audio mixing feature of the target object, determine the similarity between the first audio mixing feature of each independent source speech signal and the target audio mixing feature; The target independent source speech signal is determined based on the similarity of each of the independent source speech signals.

3. The method according to claim 2, characterized in that, Determining the target independent source speech signal based on the similarity of each of the independent source speech signals includes: If there are multiple candidate independent source speech signals with a similarity greater than or equal to a preset threshold, then the energy characteristics of each candidate independent source speech signal are determined. The candidate independent source speech signal with the largest energy feature is determined as the target independent source speech signal.

4. The method according to claim 2, characterized in that, The step of determining the target independent source speech signal based on the similarity of each of the independent source speech signals further includes: If the number of candidate independent source speech signals with a similarity greater than or equal to a preset threshold is one, then the candidate independent source speech signal is determined as the target independent source speech signal.

5. The method according to claim 2, characterized in that, The step of determining the target independent source speech signal based on the similarity of each of the independent source speech signals further includes: If the similarity of each independent source speech signal is less than a preset threshold, then the energy characteristics of each independent source speech signal are determined. The independent source speech signal with the smallest energy feature is identified as the target independent source speech signal.

6. The method according to claim 1, characterized in that, When human speech is detected, extracting the first audio mixing feature of each independent source speech signal includes: When human voice is detected, the first audio mixing feature of each independent source speech signal is extracted within a preset time period, wherein the start time of the preset time period is the time when human voice is detected.

7. The method according to any one of claims 1-6, characterized in that, The target audio mixing features of the target object are determined using the following method: In response to the detection of a wake-up event, a target audio mixing feature of the target object is determined, wherein the target audio mixing feature includes at least the voiceprint feature and / or fundamental frequency feature of the target object.

8. The method according to claim 7, characterized in that, The method for determining the target audio mixing features also includes: Extract and cache the second audio mixing features of each of the independent source speech signals; In response to detecting a wake-up event, the target audio mixing characteristics of the target object are determined, including: In response to the detection of a target wake word, the independent source speech signal containing the target wake word is determined as the independent source speech signal corresponding to the target object that triggered the wake-up event; Determine the time period of the target wake word in the independent source speech signal corresponding to the target object; The target audio mixing features of the target object are determined based on the second audio mixing features of the independent source speech signal corresponding to the target object and the time period.

9. The method according to claim 8, characterized in that, The step of determining the target audio mixing features of the target object based on the second audio mixing features of the independent source speech signal corresponding to the target object and the time period includes: In the second audio mixing features of the independent source speech signal corresponding to the target object, a third audio mixing feature located within the time period is determined; Based on the third audio mixing feature, the target audio mixing feature of the target object is determined.

10. The method according to claim 7, characterized in that, In response to detecting a wake-up event, the target audio mixing characteristics of the target object are determined, including: In response to the detection of a key wake-up event, human voice is detected in each independent source speech signal; In response to the detection of the human voice, the audio mixing features of the human voice are determined as the target audio mixing features of the target object.

11. The method according to any one of claims 1-6, characterized in that, The method further includes: For each independent source speech signal, determine the signal-to-noise ratio of the independent source speech signal, and / or the confidence level of the audio mixing feature of the independent source speech signal; based on the signal-to-noise ratio and / or the confidence level of the audio mixing feature, determine whether there is a human voice signal in the independent source speech signal. When it is determined that the human voice signal is present in at least one of the independent source speech signals, it is determined that the human voice signal has been detected.

12. The method according to claim 11, characterized in that, Determining whether a human voice signal exists in the independent source speech signal based on the confidence level of the signal-to-noise ratio and / or the audio mixing features includes: If the signal-to-noise ratio is greater than or equal to a first threshold, and / or the confidence level is greater than or equal to a second threshold, then it is determined that a human voice signal exists in the independent source speech signal.

13. The method according to claim 1, characterized in that, The separation process includes at least blind source separation; the separation process also includes echo cancellation and / or noise reverberation suppression.

14. A speech signal processing device, characterized in that, The device includes: The processing module is configured to separate and process the speech observation signals acquired by the speech acquisition array to obtain independent source speech signals corresponding to each speech acquisition device in the speech acquisition array; The first extraction module is configured to extract a first audio mixing feature of each of the independent source speech signals when human speech is detected; The first determining module is configured to determine the target independent source speech signal based on the first audio mixing feature of each of the independent source speech signals and the target audio mixing feature of the target object; The recognition module is configured to perform speech recognition on the target independent source speech signal.

15. An electronic device, characterized in that, include: processor; Memory used to store processor-executable instructions; The processor is configured to execute the instructions to cause the electronic device to implement the method as described in any one of claims 1-13.

16. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the computer program performs the steps of the method described in any one of claims 1-13.

17. A computer program product, characterized in that, Includes a computer program that, when executed by a processor, implements the steps of the method according to any one of claims 1-13.