Method and apparatus for detecting wearer audio, device and product
By extracting the wake-up features of the wake-up segment audio and the detection features of the detection segment audio, the problem of distinguishing the wearer's voice from environmental noise in wearable devices is solved, achieving accurate wearer audio detection and improving the user experience.
Patent Information
- Application Number
- PCT/CN2025/101308
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-06-27
- Filing Date
- 2025-06-17
- Publication Date
- 2026-01-02
AI Technical Summary
Existing wearable devices struggle to effectively distinguish between the wearer's voice and ambient noise in audio signal processing, leading to high hardware costs and increased computational complexity, which negatively impacts user experience.
By extracting the wake-up features of the wake-up segment audio and combining them with the detection segment audio and wake-up features, the detection features of the detection segment audio are determined to achieve wearer audio detection. The energy-related features of the wake-up segment audio are used to identify external interference, enabling accurate and real-time detection of wearer audio.
Without requiring additional hardware assistance, it accurately identifies external interference, improving the accuracy and real-time performance of wearer audio detection, providing a reliable foundation for subsequent personalized audio processing and applications, and enhancing the user experience.
Smart Images

Figure CN2025101308_02012026_PF_FP_ABST
Abstract
Description
Method, apparatus, device and product for detecting wearer audio
[0001] The present application claims priority from the Chinese patent application No. 202410851558.4 and titled "Method, apparatus, device and product for detecting wearer audio", filed on June 27, 2024, the content of which is incorporated herein by reference in its entirety. TECHNICAL FIELD
[0002] The present disclosure relates generally to the field of computers, and more specifically, to a method, apparatus, electronic device and program product for detecting wearer audio. BACKGROUND
[0003] A wearable device is a portable device that can be directly worn on the body or integrated into the user's clothes or accessories. Wearable devices usually exist in the form of portable accessories with partial computing functions and the ability to connect to mobile phones and various terminals. Mainstream product forms include watches, shoes, glasses, and other product forms such as smart clothing, backpacks, canes, accessories, etc.
[0004] In the application scenario of wearable devices, the audio function as an important component of wearable devices is committed to providing personalized services for users. To achieve this goal, advanced audio processing technologies such as noise suppression technology are widely used. With the continuous development of technology, future wearable devices will be able to more intelligently process the acquired audio signals, thereby being able to provide more personalized and convenient services. SUMMARY
[0005] Embodiments of the present disclosure provide a method, apparatus, electronic device and program product for detecting wearer audio.
[0006] According to a first aspect of the present disclosure, a method for detecting wearer audio is provided. The method comprises determining a wake-up feature of a wake-up segment audio based on an energy of the wake-up segment audio. The method further comprises determining a detection feature of a detection segment audio based on the detection segment audio and the wake-up feature. In addition, the method further comprises detecting wearer audio in the detection segment audio based on the detection feature.
[0007] In a second aspect of the present disclosure, an apparatus for detecting wearer audio is provided. The apparatus comprises a wake-up feature determination module configured to determine a wake-up feature of a wake-up segment audio based on an energy of the wake-up segment audio. The apparatus further comprises a detection feature determination module configured to determine a detection feature of a detection segment audio based on the detection segment audio and the wake-up feature. In addition, the apparatus further comprises a wearer audio detection module configured to detect wearer audio in the detection segment audio based on the detection feature.
[0008] In a third aspect of the disclosure, an electronic device is provided. The electronic device includes a processor and a memory coupled with the processor, the memory having stored therein instructions which, when executed by the processor, cause the electronic device to perform the method according to the first aspect.
[0009] In a fourth aspect of the disclosure, a computer program product having stored thereon instructions including computer executable instructions, wherein the computer executable instructions are executed by a processor to implement the method of the first aspect.
[0010] The summary is intended to provide a simplified form of the concepts selected to introduce the subject matter, which will be further described below in the detailed description. It is not intended to identify key or essential features of the claimed subject matter or to limit the scope of the claimed subject matter. BRIEF DESCRIPTION OF DRAWINGS
[0011] The above and other features, aspects, and advantages of various embodiments of the present disclosure will become more apparent from the following detailed description taken in conjunction with the accompanying drawings. In the drawings, like reference numerals denote like elements, wherein:
[0012] FIG. 1 illustrates a schematic diagram of an example environment in which some embodiments of the present disclosure can be implemented;
[0013] FIG. 2 illustrates a flowchart of a method for detecting wearer audio of some embodiments of the present disclosure;
[0014] FIG. 3 illustrates a schematic diagram of an architecture for detecting wearer audio of some embodiments of the present disclosure;
[0015] FIG. 4A illustrates a schematic diagram for determining a wake-up feature in a registration phase of some embodiments of the present disclosure;
[0016] FIG. 4B illustrates a schematic diagram for determining a detection feature in a detection phase of some embodiments of the present disclosure;
[0017] FIG. 5A illustrates a schematic diagram for determining a first wake-up feature of some embodiments of the present disclosure;
[0018] FIG. 5B illustrates a schematic diagram for determining a second wake-up feature of some embodiments of the present disclosure;
[0019] FIG. 6A illustrates a schematic diagram for determining a first detection feature of some embodiments of the present disclosure;
[0020] FIG. 6B illustrates a schematic diagram for determining a second wake-up feature of some embodiments of the present disclosure;
[0021] FIGS. 7A-7D show a schematic diagram of detecting wearer audio according to the wake-up feature and the detection feature, according to some embodiments of the present disclosure;
[0022] FIG. 8 shows a block diagram of an apparatus for detecting wearer audio, according to some embodiments of the present disclosure; and
[0023] FIG. 9 shows a block diagram of an electronic device, according to some embodiments of the present disclosure.
[0024] In all the drawings, same or similar reference numerals indicate same or similar elements. DETAILED DESCRIPTION
[0025] It can be understood that the data involved in the technical solutions of the present disclosure (including but not limited to the data itself, the acquisition or use of the data) should comply with the requirements of the relevant laws and regulations and relevant provisions.
[0026] It can be understood that, before using the technical solutions disclosed in the embodiments of the present disclosure, the type of personal information involved in the present disclosure, the scope of use, the use scenario, etc. should be informed to the user and the authorization of the user should be obtained through appropriate means according to relevant laws and regulations.
[0027] For example, when receiving the active request of the user, the user is sent a prompt information to explicitly prompt the user that the operation requested to be executed will need to acquire and use the personal information of the user. Thus, the user can voluntarily choose whether to provide the personal information to the electronic device, application program, server or storage medium, etc. software or hardware that executes the operation of the technical solutions of the present disclosure according to the prompt information.
[0028] As an optional but non-limiting implementation manner, in response to receiving the active request of the user, the manner of sending the prompt information to the user may, for example, be a pop-up window manner, and the prompt information can be presented in the pop-up window in the form of text. In addition, the pop-up window can also carry a selection control for the user to select "agree" or "disagree" to provide the personal information to the electronic device.
[0029] It can be understood that the above notification and user authorization process is only illustrative, and does not limit the implementation manner of the present disclosure, and other manners that meet the relevant laws and regulations can also be applied to the implementation manner of the present disclosure.
[0030] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms, and should not be interpreted as being limited to the embodiments set forth herein, but rather these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes, and are not intended to limit the scope of protection of the present disclosure.
[0031] In the description of embodiments of the disclosure, the term "comprising" and its similar words are understood to be open-ended, i.e., "including but not limited to". The term "based on" is understood to be "based at least in part on". The term "one embodiment" or "the embodiment" is understood to be "at least one embodiment". The terms "first", "second", and so on can refer to different or same objects unless explicitly stated otherwise. Other explicit and implicit definitions can also be included below.
[0032] As mentioned before, in the application scenario of wearable devices, the audio function is very important. In order to provide personalized user experience for the wearer, it is usually necessary to distinguish the wearer's voice and the environmental noise, so as to realize the processing and application of the wearer's voice signal. In the related art, the wearer's audio signal can be picked up by using a hardware device such as a bone conduction microphone, however, this method by means of the bone conduction microphone has strict requirements on the pickup quality and deployment position of the bone conduction microphone, and undoubtedly also increases the hardware cost. Another technology that uses voiceprints to verify the wearer's audio needs to obtain the voiceprint information of the user in advance, which increases the complexity of the user's use. In addition, the calculation amount and parameter amount of the voiceprint verification technology are also high, and it is difficult to balance the computing power and accuracy on a platform with limited resources.
[0033] In an embodiment of the disclosure, the wake-up features of the user wake-up segment audio are extracted, and the detection features of the detection segment audio are determined by means of the wake-up features extracted from the wake-up segment and the detection segment audio, and then the wearer audio detection is realized based on the detection features. This method of detecting the wearer audio by referring to the energy-related features of the wake-up segment audio can effectively identify external interference, accurately and in real time detect the wearer audio, and provide a reliable basis for subsequent personalized audio processing and application, thereby improving the user experience.
[0034] FIG. 1 shows a schematic diagram of an example environment 100 in which some embodiments of the disclosure can be implemented. As shown in FIG. 1, the wearer audio detection system 120 receives audio 110 from a wearable device. The wearable device can be smart AR glasses, earphones, etc., which are configured with multi-channel air conduction microphones to obtain sound, and the audio 110 can be mixed audio containing multiple sound sources and environmental noise, etc.
[0035] With continued reference to FIG. 1, the start point 132 and the end point 134 of the wearer audio 130 in the audio 110 can be obtained through the analysis of the wearer audio detection system 120. These two time points can provide the exact range of the wearer audio 130 in the whole audio, so as to facilitate the subsequent more detailed analysis and application of the wearer's voice. In some embodiments, the wearer audio detection system 120 can obtain the wake-up segment audio according to the time stamp of the audio 110, so that the wearer audio detection system 120 can refer to the energy-related features of the wake-up segment audio to effectively and timely detect the wearer audio.
[0036] It can be understood that before the analysis of the audio 110, the audio 110 can also be pre-processed as necessary, such as sampling rate conversion, noise filtering, etc., to improve the accuracy of subsequent analysis.
[0037] The processes according to embodiments of the present disclosure will be described in detail below with reference to FIGS. 2-9. For ease of understanding, the specific data mentioned in the following description are exemplary and are not intended to limit the protection scope of the present disclosure. It can be understood that the embodiments described below can also include additional actions not shown and / or can omit the actions shown, and the scope of the present disclosure is not limited in this respect.
[0038] FIG. 2 shows a flowchart of a method 200 for detecting wearer audio according to some embodiments of the present disclosure. The method 200 can be performed by an apparatus for detecting wearer audio. The apparatus can be implemented by software and / or hardware. Next, the method 200 will be described schematically with the apparatus for detecting wearer audio as the execution subject (such as smart AR glasses or earphones, etc.) as an example. Referring to FIG. 2, the method 200 can include block 202, block 204, and block 206.
[0039] At block 202, a wake-up feature of the wake-up segment audio is determined based on the energy of the wake-up segment audio. In some embodiments, the audio of the wake-up segment refers to the audio that wakes up the wearable device, and there can be different or the same wake-up words for a certain or certain type of device. In some embodiments, the time stamp of the wake-up segment can be detected by a wake-up algorithm, so that the wake-up segment audio can be obtained. In some embodiments, the energy of the audio refers to the strength or amplitude of the sound contained in the audio signal. The wake-up feature of the wake-up segment audio is determined according to the energy of the wake-up segment audio, and the wake-up feature can include a first wake-up feature and a second wake-up feature. Before the wake-up feature is obtained, the wake-up segment audio in the time domain can be converted to the time-frequency domain.
[0040] At block 204, based on the detected segment audio and the wake-up feature, a detection feature of the detected segment audio is determined. In some embodiments, the detected segment audio refers to the audio acquired from the wearable device in the process of the user's use, other than the wake-up segment audio. In some embodiments, the detected segment audio can also be acquired according to a timestamp, that is, part of the audio 110 shown in FIG. 1 is the detected segment audio. In some embodiments, the detection feature of the detected segment audio can be obtained according to the acquired detected segment audio and the aforementioned acquired wake-up feature, and the detection feature can include a first detection feature and a second detection feature. In some embodiments, the detected segment audio in the time domain can be converted to the time-frequency domain.
[0041] At block 206, based on the detection feature, wearer audio is detected in the detected segment audio. In some embodiments, the start point and the end point of the wearer audio can be detected in the detected segment audio according to the first detection feature and / or the second detection feature, that is, 132 shown in FIG. 1 is the start point of the wearer audio, and 134 shown in FIG. 1 is the end point of the wearer audio. The two time points can provide the exact range of the wearer audio in the whole segment audio, thereby facilitating more detailed analysis and application of the wearer's sound.
[0042] In embodiments of the present disclosure, the wake-up feature of the user's wake-up segment audio is extracted based on the energy of the wake-up segment audio, and the detection feature of the detected segment audio is determined by means of the wake-up feature extracted from the detected segment audio and the wake-up segment, and then the wearer audio detection is realized based on the detection feature. This method of detecting the wearer audio by referring to the energy-related features of the wake-up segment audio can effectively identify external interference, accurately and in real time detect the wearer audio, provide a reliable basis for subsequent personalized audio processing and application, and thus improve the user experience.
[0043] FIG. 3 shows a schematic diagram of an architecture 300 for detecting wearer audio according to some embodiments of the present disclosure. As shown in FIG. 3, at 310, audio is acquired. The audio can be acquired by a wearable device configured with a multi-channel air conduction microphone, so it can be an audio signal mixed with multiple sound sources, and the wearable device can be a smart AR glasses or a headset device.
[0044] Referring to FIG. 3, it can be understood that the present disclosure mainly involves two stages, one is a registration stage (as shown in 320 of FIG. 3) and the other is a detection stage (as shown in 330 and 340 of FIG. 3). In the registration stage, the relevant parameters of the wearer of the wake-up segment audio are calculated, and in the detection stage, the detection features of the detection segment audio are determined according to the calculated relevant parameters of the wearer of the wake-up segment audio and in combination with the audio of the detection segment, so that the starting point and the ending point of the wearer audio can be detected according to the detection features. Through this method of detecting the wearer according to the audio energy, the external interference can be accurately identified without the assistance of additional hardware such as bone conduction earphones, so as to detect the wearer audio.
[0045] In order to more clearly illustrate the process of calculating the relevant parameters of the wearer of the wake-up segment audio at 320, the following will be described in combination with FIG. 4A. FIG. 4A shows a schematic diagram of a wake-up feature 400A for determining the wake-up feature in the registration stage according to some embodiments of the present disclosure.
[0046] Referring to FIG. 4A, at 410A, the wake-up segment audio is extracted. In some embodiments, the timestamp of the wake-up segment can be detected by a wake-up algorithm, so that the audio data of the wake-up segment can be obtained. For example, the wake-up algorithm can be a wake-up algorithm based on speech features or a data packet wake-up algorithm, and the present disclosure does not limit it. At 420A, the time-frequency domain positive transform is performed on the wake-up segment audio. In some embodiments, the time-frequency domain positive transform can be performed on the time domain audio signals of each channel air conduction microphone of the wake-up segment audio to obtain the time-frequency domain audio signals of each channel air conduction microphone. In some embodiments, the optional transform methods include short-time Fourier transform, sub-band decomposition and the like, and the present disclosure does not limit it.
[0047] Continuing to refer to FIG. 4A, at 430A, the audio is separated to determine the first wake-up feature. In order to more clearly illustrate the process of separating the audio to determine the first wake-up feature at 430A, the following will be described in combination with FIG. 5A. FIG. 5A shows a schematic diagram of 500A for determining the first wake-up feature according to some embodiments of the present disclosure. Referring to FIG. 5A, at 510A, the separation matrix of the wake-up segment is determined. In some embodiments, the separation matrix information of the wake-up segment can be extracted by using a separation algorithm, and the optional separation algorithm includes but is not limited to an independent vector analysis algorithm and the like, and the present disclosure does not limit it.
[0048] As shown in FIG. 5A, at 520A, the wake-up segment audio is separated. In some embodiments, the audio mixed with multiple sound sources can be separated into individual signal sources by the separation algorithm mentioned at 510A. For example, the audio acquired by the multi-channel air conduction microphone can be separated into multiple channel individual audio, such as the wearer audio channel and the wearer audio suppressed channel. At 530A, the energy of the wearer audio suppressed channel is determined, which is relatively low compared to the energy of the wearer audio channel, so that the energy of the wearer audio suppressed channel can be clearly determined, given that the microphone is configured on the wearable device. At 540A, the energy difference is calculated to determine the first wake-up feature. In some embodiments, for an original microphone, the first wake-up feature can be the energy difference between the original microphone audio channel and the wearer signal suppressed channel, assuming that the audio energy of the original microphone channel is EA, and after separation processing, assuming that the audio energy of the wearer suppressed is EB, the energy difference is EA-EB. It can also be understood that the first wake-up feature is the energy of the wearer audio of the wake-up segment.
[0049] Referring back to FIG. 4A, at 440A, a beam algorithm is applied to determine the second wake-up feature. To better illustrate the process of separating the audio to determine the first wake-up feature at 440A, the following will be described in conjunction with FIG. 5B. FIG. 5B shows a schematic diagram of 500B for determining the second wake-up feature according to some embodiments of the present disclosure. Referring to FIG. 5, at 510B, a beam algorithm with multiple target orientations is applied to the wake-up segment audio. In some embodiments, the beam algorithm with multiple target orientations is used to scan the microphone time-frequency domain audio signal, and it can be understood that the number and direction of the target orientations can be determined according to the computing power and application scenario. For example, one orientation can be set as a target orientation every 45 degrees in the [0°, 180°] and [180°, 360°] directions to apply the beam algorithm. If the computing power is strong and the application scenario is complex, one orientation can be selected as a target orientation every 10° to apply the beam algorithm. In some embodiments, an enhanced beam algorithm or a suppressed beam algorithm can be selected, which is not limited by the present disclosure.
[0050] With continued reference to FIG. 5, at 520B, the relative energy difference of the original microphone and the audio after applying the beam algorithm in different orientations is calculated. For example, if 5 orientations are set, the relative energy difference can be calculated as [A1, A2, A3, A3, A5]. At 530B, the relative energy difference is determined as the second wake-up feature. In this way, the second wake-up feature is [A1, A2, A3, A3, A5]. It can also be understood that, for the effectiveness and simplicity of data processing, a certain number of frequency points within the wake-up segment duration can be selected at 522B before calculating the relative energy difference, and these frequency points are averaged.
[0051] Referring back to FIG. 3, after the relevant parameters of the wearer of the wake-up segment audio are determined in the registration stage, the detection stage can be entered to detect the wearer audio of the detection segment. At 330, the energy and other features of the audio are calculated with the relevant parameters obtained from the wake-up segment. To better illustrate the process of calculating the energy and other features of the audio at 330 with the relevant parameters obtained from the wake-up segment, the following will be described in conjunction with FIG. 4B. FIG. 4B shows a schematic diagram of 400B for determining detection features in the detection stage according to some embodiments of the present disclosure. Referring to FIG. 4B, at 410B, a time domain positive transform is performed on the detection segment audio. The time-frequency domain positive transform is performed on the time domain audio signals of each channel of the detection segment audio to obtain the time-frequency domain audio signals of each channel of the detection segment. The optional transform methods include short-time Fourier transform, sub-band decomposition, etc.
[0052] With reference back to FIG. 4B, at 420B, the first wake-up features are matched to obtain the first detection features. To better illustrate the process of matching the first wake-up features to obtain the first detection features at 420B, the following will be described in conjunction with FIG. 6A. FIG. 6A shows a schematic diagram of 600A for determining the first detection features according to some embodiments of the present disclosure. Referring to FIG. 6A, at 610A, the audio of the detection segment is separated with the separation matrix of the wake-up segment. The separation matrix can effectively separate the audio signals from different sound sources. In some embodiments, the separation matrix of the wake-up segment can be fixed to separate the time-frequency domain audio signals of each channel of the detection segment. In some embodiments, the separation matrix of the wake-up segment can also be used to initialize the online separation algorithm parameters of the detection segment audio, wherein the separation matrix parameters corresponding to the channel numbers of the wearer signal in the separation matrix of the wake-up segment can be fixed, and the separation matrix parameters of other channel numbers can be updated in real time, so that the time-frequency domain audio signals of each channel of the detection segment can be separated.
[0053] By separating the detection segment audio with the separation matrix of the wake-up segment, the efficiency of audio processing can be improved, the errors caused by recalculating the separation matrix can be reduced, and the consistency of the processing of the wake-up segment audio and the detection segment audio can also be ensured.
[0054] With continued reference to FIG. 6A, after obtaining the audio of each separated channel, the energy difference can be calculated at 620A. In some embodiments, the energy difference between the original microphone channel energy of the detection segment and the channel number channel energy of the corresponding suppressed wearer signal after the separation processing can be calculated. Before the separation processing, assume the audio energy of the original microphone channel is EC, and after the separation processing, assume the audio energy of the suppressed wearer is ED, the energy difference is EC-ED. Then at 630A, the ratio of the energy difference and the first wake-up feature is recorded as the first detection feature. In connection with the description of 540 in FIG. 5A, the first detection feature can be (EC-ED) / (EA-EB), which can be understood as the ratio of the energy difference change of the detection segment relative to the wake-up segment. The larger the ratio, the greater the relative strength of the wearer audio is increased or the lower the relative strength of the wearer audio is decreased.
[0055] Returning to FIG. 4B, at 430B, the second wake-up feature is matched to obtain the second detection feature. To better illustrate the process of matching the second wake-up feature to obtain the second detection feature, the process will be described in connection with FIG. 6B. FIG. 6B illustrates a schematic diagram of 600B for determining the second wake-up feature in accordance with some embodiments of the present disclosure. Referring to FIG. 6B, at 610B, a beamforming algorithm of multiple target orientations is applied to the audio of the detection segment. In some embodiments, the microphone time-frequency domain audio signal of the detection segment can be scanned using the same target orientation beamforming algorithm as the wake-up segment. If 5 orientations are set in the wake-up segment, then 5 orientations are also set in the detection segment, so that the relative energy difference between the original microphone and the audio after the beamforming algorithm is calculated at 620B, for example, the energy difference [B1, B2, B3, B3, B5] is calculated.
[0056] With continued reference to FIG. 6B, then at 630B, the difference between the relative energy difference of each orientation and the second wake-up feature of the corresponding orientation is calculated. In connection with the second wake-up feature [A1, A2, A3, A3, A5] obtained at 520B in FIG. 5B, the difference [|B1-A1|, |B2-A2|, |B3-A3|, |B4-A4|, |B5-A5|] can be obtained. Then at 640B, the difference between the maximum value and the minimum value of the difference is recorded as the second detection feature. Assuming B1-A1 is the maximum value and B5-A5 is the minimum value, then the second detection feature is ||B1-A1|-|B5-A5||.
[0057] Returning to FIG. 3, at 340, audio of the wearer is detected based on the acquired features. To better illustrate the process of detecting audio of the wearer based on the acquired features, the process will be described below in conjunction with FIGS. 7A-7B. FIG. 7A illustrates a diagram of detecting wearer audio 700A based on a first detection feature, according to some embodiments of the present disclosure. As shown in FIG. 7A, at 710A, it is determined whether the first detection feature is greater than or equal to a first threshold. If it is greater than or equal to the first threshold, then the frame signal can be determined to be a wearer frame at 712A. If the first detection feature is less than the first threshold, then the frame signal can be determined to be a non-wearer frame at 714A. In particular, the first threshold can be set to 1, and if the first detection feature is greater than or equal to 1, the frame signal can be considered to be a wearer frame. If the first detection feature is less than 1, the frame signal can be considered to be a non-wearer frame (ambient noise frame).
[0058] With continued reference to FIG. 7A, after determining whether the frame is a wearer frame or a non-wearer frame, the start point and the end point of the wearer frame need to be determined. In particular, at 722A, it is determined whether M frames out of N consecutive frames are wearer frames. If so, then the start point of the N frames can be recorded as the start point of the wearer audio at 724A. At 732A, it is determined whether Q frames out of P consecutive frames are non-wearer frames (ambient noise frames). If so, then the end point of the P frames can be recorded as the end point of the wearer audio at 734A. In some embodiments, the frames are arranged in chronological order.
[0059] FIG. 7B illustrates a diagram of detecting wearer audio 700B based on a second detection feature, according to some embodiments of the present disclosure. As shown in FIG. 7B, the wearer audio can be detected based on the rising edge and the falling edge of the second detection feature. In particular, if the height of the falling edge of the falling edge of the second detection feature exceeds a second threshold, the start point of the falling edge 710B can be recorded as the start point of the wearer audio. If the height of the rising edge of the second detection feature exceeds a third threshold, the end point of the rising edge 720B can be recorded as the end point of the wearer audio.
[0060] FIG. 7C illustrates a diagram of determining whether a frame is a wearer frame or a non-wearer frame 700C based on a second detection feature, according to some embodiments of the present disclosure. As shown in FIG. 7C, at 710C, the first rising edge or the first falling edge of the second detection feature is determined. At 720C, the second detection feature belonging to the wearer audio and the second detection feature belonging to the non-wearer audio (ambient noise) are initialized to obtain a fourth threshold. In some embodiments, the fourth threshold can be set to the average of the second detection feature belonging to the wearer audio and the second detection feature belonging to the non-wearer audio (ambient noise).
[0061] With reference back to FIG. 7C, at 730C, it is determined whether the second detection feature is less than or equal to the fourth threshold value. If yes, at 740C, the frame signal is determined as a wearer frame. Then at 750C, the second detection feature belonging to the wearer audio is updated.
[0062] With reference back to FIG. 7C, at 730C, it is determined whether the second detection feature is less than or equal to the fourth threshold value. If yes, at 740C, the frame signal is determined as a wearer frame. Then at 750C, the second detection feature belonging to the wearer audio is updated.
[0063] As shown in FIG. 7C, at 780C, the fourth threshold value is further updated according to the updated second detection feature belonging to the wearer and the second detection feature belonging to the suppressed wearer audio (ambient noise).
[0064] FIG. 7D shows a schematic diagram of detecting wearer audio 700D in combination with the first detection feature and the second detection feature according to some embodiments of the present disclosure. As shown in FIG. 7D, at 710D, the start point and the end point of the wearer audio are determined in combination with the first detection feature and the second detection feature. At 722D, it is determined whether there are Ml frames in the continuous N frames that are determined as wearer frames based on the first detection feature. If yes, at 724D, it is determined whether there are M2 frames in the continuous N frames that are determined as wearer frames based on the second detection feature. If yes, at 726D, the start point of the N frames can be recorded as the start point of the wearer audio.
[0065] As shown in FIG. 7D, at 732D, it is determined whether there are Ql frames in the continuous Pl frames that are determined as suppressed wearer frames (ambient noise frames) based on the first detection feature. If yes, at 734D, the end point of the Pl frames can be recorded as the end point of the detected wearer audio.
[0066] With reference back to FIG. 7D, at 742D, there are Q2 frames in the continuous P2 frames that are determined as suppressed wearer frames (ambient noise frames) based on the second detection feature. If yes, at 744D, the end point of the Q2 frames can be recorded as the end point of the detected wearer audio.
[0067] With this detection method based on the energy-related feature of the wake-up audio, external interference can be accurately identified without relying on additional hardware such as bone conduction earphones and other auxiliary devices, so that the wearer's voice signal can be effectively detected, providing a basis for subsequent application and processing of the wearer's voice signal, and thus personalized service experience can be provided for the user.
[0068] FIG. 8 shows a block diagram of an apparatus 800 for detecting wearer audio according to some embodiments of the present disclosure. As shown in FIG. 8, the apparatus 800 includes a wake-up feature determination module 802 configured to determine a wake-up feature of the wake-up segment audio based on an energy of the wake-up segment audio. The apparatus 800 further includes a detection feature determination module 804 configured to determine a detection feature of the detection segment audio based on the detection segment audio and the wake-up feature. In addition, the apparatus 800 further includes a wearer audio detection module 806 configured to detect the wearer audio in the detection segment audio based on the detection feature.
[0069] FIG. 9 shows a block diagram of an electronic device 900 according to some embodiments of the present disclosure. The device 900 can be a device or apparatus described in embodiments of the present disclosure. As shown in FIG. 9, the device 900 includes a central processing unit (CPU) and / or a graphics processing unit (GPU) 901, which can perform various appropriate actions and processes according to computer program instructions stored in a read-only memory (ROM) 902 or loaded from a storage unit 908 into a random access memory (RAM) 903. Various programs and data required for operation of the device 900 can also be stored in the RAM 903. The CPU / GPU 901, the ROM 902, and the RAM 903 are connected to each other through a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904. Although not shown in FIG. 9, the device 900 can also include a coprocessor.
[0070] Various components in the device 900 are connected to the I / O interface 905, including an input unit 906, e.g., a keyboard, a mouse, etc.; an output unit 907, e.g., various types of displays, speakers, etc.; the storage unit 908, e.g., a magnetic disk, a magneto-optical disk, etc.; and a communication unit 909, e.g., a network card, a modem, a wireless communication transceiver, etc. The communication unit 909 allows the device 900 to exchange information / data with other devices through a computer network, such as the Internet, and / or various telecommunication networks.
[0071] The various methods or processes described above can be performed by the CPU / GPU 901. For example, in some embodiments, the methods can be implemented as a computer software program tangibly embodied in a machine-readable medium, e.g., the storage unit 908. In some embodiments, part or all of the computer program can be loaded and / or installed on the device 900 via the ROM 902 and / or the communication unit 909. When the computer program is loaded into the RAM 903 and executed by the CPU / GPU 901, one or more steps or actions of the methods or processes described above can be performed.
[0072] In some embodiments, the methods and processes described above can be tied to a computer program product. The computer program product can include a computer readable storage medium having computer readable program instructions thereon for performing various aspects of the present disclosure.
[0073] The computer readable storage medium can be a tangible device that can retain and store instructions for use by an instruction execution device. The computer readable storage medium can be, for example, but is not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer readable storage medium include the following: a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanically encoded device such as punch-cards or raised structures in a groove having instructions recorded thereon, and any suitable combination of the foregoing. A computer readable storage medium, as used herein, is not to be construed as being transitory signals per se, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through a waveguide or other transmission media (e.g., light pulses passing through a fiber-optic cable), or electrical signals transmitted through a wire.
[0074] Computer readable program instructions described herein can be downloaded to respective computing / processing devices from a computer readable storage medium or to an external computer or external storage device via a network, for example, the Internet, a local area network, a wide area network and / or a wireless network. The network can comprise copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers and / or edge servers. A network adapter card or network interface in each computing / processing device receives computer readable program instructions from the network and forwards the computer readable program instructions for storage in a computer readable storage medium within the respective computing / processing device.
[0075] Computer readable program instructions for carrying out operations of the present disclosure can be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine dependent instructions, microcode, firmware instructions, state-setting data, or either source code or object code written in any combination of one or more programming languages, including object oriented programming languages and conventional procedural programming languages. The computer readable program instructions can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider). In some embodiments, electronic circuitry including, for example, programmable logic circuitry, field-programmable gate array (FPGA), or programmable logic array (PLA) can execute the computer readable program instructions by utilizing state information of the computer readable program instructions to personalize the electronic circuitry, in order to perform aspects of the present disclosure.
[0076] These computer readable program instructions can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks. These computer readable program instructions can also be stored in a computer readable storage medium that can include a non-transitory computer readable storage medium that can direct a computer, a programmable data processing apparatus, and / or other devices to function in a particular manner, such that the computer readable storage medium having instructions stored therein comprises an article of manufacture including instructions which implement aspects of the function / act specified in the flowchart and / or block diagram block or blocks.
[0077] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer, other programmable data processing apparatus, or other device implement the functions / acts specified in the flowchart and / or block diagram block or blocks.
[0078] The computer program product of the first aspect can include a computer readable storage medium. The computer readable storage medium can include instructions embedded or transcribed in it. The instructions can include at least one of: a plurality of instructions that, when executed by a processor of an electronic device, cause the electronic device to perform the method of the first aspect; or a plurality of instructions that, when executed by a processor of an electronic device, cause the electronic device to perform the method of the second aspect. The computer readable storage medium can include at least one of: a magnetic storage medium; an optical storage medium; an electronic storage medium; or an atomic storage medium.
[0079] Embodiments of the present disclosure have been described above, with the understanding that the foregoing description is that of the embodiments and is not exhaustive of or limited to the embodiments disclosed. Many modifications and variations, in addition to those described, will be apparent to those skilled in the art from this disclosure. The above description is exemplary only, and not exhaustive of or limiting to the embodiments disclosed. Constructions, embodiments, and methods described herein are presented by way of example only and are not intended to limit the scope of the disclosure. The various embodiments described herein can be implemented in a convenient computer-based system or programmable logic, for example, with a computer system. The constructed system can include a processor, a storage medium readable by the processor (including volatile and non-volatile memory and / or storage elements), at least one input device, and at least one output device. The volatile and non-volatile memory and / or storage elements can be a computer-readable medium, a floppy disk, a RAM, a ROM, an erasable programmable read only memory (EPROM), an electrically erasable programmable read only memory (EEPROM), a flash memory, etc. The floppy disk can be a diskette, a small diskette, a mini diskette, etc. The system and various components thereof can be implemented with various software
[0080] Some example implementations of the present disclosure are listed below.
[0081] Example 1. A method for detecting wearer audio, comprising:
[0082] determining a wake-up feature of the wake-up segment audio based on an energy of the wake-up segment audio;
[0083] determining a detection feature of the detection segment audio based on the detection segment audio and the wake-up feature; and
[0084] detecting wearer audio in the detection segment audio based on the detection feature.
[0085] Example 2. The method of example 1, wherein determining a wake-up feature of the wake-up segment audio based on an energy of the wake-up segment audio comprises:
[0086] determining a first wake-up feature based on the energy of the wake-up segment audio and an energy of the wake-up segment that suppresses wearer audio.
[0087] Example 3. The method of any one of examples 1-2, wherein determining a first wake-up feature based on the energy of the wake-up segment audio and an energy of the wake-up segment that suppresses wearer audio comprises:
[0088] determining a separation matrix for the wake-up segment audio;
[0089] separating the wake-up segment audio into a plurality of audios, the plurality of audios including the wearer audio of the wake-up segment and the suppress wearer audio of the wake-up segment; and
[0090] determining the first wake-up feature based on a difference between the energy of the wake-up segment audio and the energy of the suppress wearer audio.
[0091] Example 4. The method of any one of examples 1-3, wherein determining a wake-up feature for the wake-up segment audio based on the energy of the wake-up segment audio further comprises:
[0092] determining a plurality of energies after applying a beamforming algorithm for the wake-up segment audio at a plurality of target orientations; and
[0093] determining a second wake-up feature based on the plurality of energies of the wake-up segment audio at the plurality of target orientations and the plurality of energies after applying the beamforming algorithm for the wake-up segment audio at the plurality of target orientations.
[0094] Example 5. The method of any one of examples 1-4, wherein determining a detection feature for the detection segment audio based on the detection segment audio and the wake-up feature comprises:
[0095] determining an energy of a suppress wearer audio of the detection segment based on the separation matrix for the wake-up segment audio;
[0096] determining a difference between the energy of the detection segment audio and the energy of the suppress wearer audio of the detection segment; and
[0097] determining a first detection feature based on the difference and the first wake-up feature.
[0098] Example 6. The method of any one of examples 1-5, wherein determining a detection feature for the detection segment audio based on the detection segment audio and the wake-up feature further comprises:
[0099] determining a plurality of energies after applying a beamforming algorithm for the detection segment audio at a plurality of target orientations;
[0100] determining a plurality of energy differences between the plurality of energies of the detection segment audio at the plurality of target orientations and the plurality of energies after applying the beamforming algorithm for the detection segment audio at the plurality of target orientations;
[0101] determining a plurality of difference values between the plurality of energy differences and a second wake-up feature; and
[0102] determine, based on a maximum difference value and a minimum difference value of the plurality of difference values, a second detection feature.
[0103] Example 7. The method of any one of examples 1-6, wherein detecting, based on the detection feature, wearer audio in the detection segment audio comprises:
[0104] determining whether a frame corresponding to the first detection feature is a wearer frame or a suppress-wearer frame;
[0105] in response to a plurality of consecutive wearer frames in the plurality of frame sets satisfying a first condition, determining a first-in-time frame in the plurality of frame sets as a start point of the wearer audio of the detection segment; and
[0106] in response to a plurality of consecutive suppress-wearer frames in the plurality of frame sets satisfying a second condition, determining a last-in-time frame in the plurality of frame sets as an end point of the wearer audio of the detection segment.
[0107] Example 8. The method of any one of examples 1-7, determining whether a frame corresponding to the first detection feature is a wearer frame or a suppress-wearer frame comprises:
[0108] in response to the first detection feature being greater than or equal to a first threshold value, determining the corresponding frame as the wearer frame; or
[0109] in response to the first detection feature being less than the first threshold value, determining the corresponding frame as the suppress-wearer frame.
[0110] Example 9. The method of any one of examples 1-8, wherein detecting, based on the detection feature, wearer audio in the detection segment audio further comprises:
[0111] determining a falling edge and a rising edge of the second detection feature;
[0112] in response to a height of the falling edge satisfying a second threshold value, determining a start point of the falling edge as a start point of the wearer audio of the detection segment; and
[0113] in response to a height of the falling edge satisfying a third threshold value, determining an end point of the rising edge as an end point of the wearer audio of the detection segment.
[0114] Example 10. The method of any one of examples 1-9, wherein detecting, based on the detection feature, wearer audio in the detection segment audio further comprises:
[0115] determining a fourth threshold value based on a first rising edge or a first falling edge of the second detection feature of the wearer audio of the detection segment and the second detection feature of the non-wearer audio of the detection segment;
[0116] updating the second detection feature of the wearer audio of the detection segment by determining a frame corresponding to the second detection feature as the wearer frame in response to the second detection feature being less than or equal to the fourth threshold value;
[0117] updating the second detection feature of the non-wearer audio of the detection segment by determining a frame corresponding to the second detection feature as the non-wearer frame in response to the second detection feature being greater than the fourth threshold value; and
[0118] updating the fourth threshold value based on the updated second detection feature of the wearer audio of the detection segment and the updated second detection feature of the non-wearer audio of the detection segment.
[0119] Example 11. The method of any one of examples 1-10, wherein detecting wearer audio in the detection segment audio based on the detection feature further comprises:
[0120] determining a first frame of the plurality of frame sets as a start point of the wearer audio of the detection segment in response to a consecutive plurality of wearer frames of the plurality of frame sets simultaneously satisfying a second condition and a third condition; and
[0121] determining a last frame of the plurality of frame sets as an end point of the wearer audio of the detection segment in response to a consecutive plurality of non-wearer frames of the plurality of frame sets satisfying a fourth condition; or
[0122] determining a last frame of the plurality of frame sets satisfying a fifth condition as the end point of the wearer audio of the detection segment in response to a consecutive plurality of non-wearer frames of the plurality of frame sets satisfying the fifth condition.
[0123] Example 12. The method of any one of examples 1-11, further comprising:
[0124] obtaining the wake-up segment audio and the detection segment audio based on timestamps; and
[0125] converting the wake-up segment audio and the detection segment audio from a time domain to a time-frequency domain.
[0126] Example 13. An apparatus for detecting wearer audio, comprising:
[0127] a wake-up feature determination module configured to determine a wake-up feature of the wake-up segment audio based on an energy of the wake-up segment audio;
[0128] a detection feature determination module configured to determine a detection feature of the detection segment audio based on the detection segment audio and the wake-up feature; and
[0129] a wearer audio detection module configured to detect wearer audio in the detection segment audio based on the detection feature.
[0130] Example 14. The apparatus of Example 13, wherein the wake-up feature determination module comprises:
[0131] a first wake-up feature determination module configured to determine a first wake-up feature based on the energy of the wake-up segment audio and an energy of the wearer- suppressed audio of the wake-up segment.
[0132] Example 15. The apparatus of any one of Examples 13-14, wherein the first wake-up feature determination module comprises:
[0133] a separation matrix determination module configured to determine a separation matrix for the wake-up segment audio;
[0134] an audio separation module configured to separate the wake-up segment audio into a plurality of audios, the plurality of audios comprising the wearer audio of the wake-up segment and the wearer-suppressed audio of the wake-up segment; and
[0135] a first determination module configured to determine the first wake-up feature based on a difference between the energy of the wake-up segment audio and the energy of the wearer-suppressed audio.
[0136] Example 16. The apparatus of any one of Examples 13-15, wherein the wake-up feature determination module further comprises:
[0137] a first energy determination module configured to determine a plurality of energies after applying a beamforming algorithm to the wake-up segment audio at a plurality of target orientations; and
[0138] a second wake-up feature determination module configured to determine a second wake-up feature based on the plurality of energies of the wake-up segment audio at the plurality of target orientations and the plurality of energies after applying the beamforming algorithm to the wake-up segment audio at the plurality of target orientations.
[0139] Example 17. The apparatus of any one of Examples 13-16, wherein the detection feature determination module comprises:
[0140] a second energy determination module configured to determine an energy of the wearer-suppressed audio of the detection segment based on the separation matrix for the wake-up segment audio;
[0141] a first difference determining module configured to determine a difference between the energy of the detected segment audio and the energy of the detected segment of the suppressor wearer audio; and
[0142] a first detection feature determining module configured to determine a first detection feature based on the difference and a first wake-up feature.
[0143] Example 18. The apparatus of any one of examples 13-17, wherein the detection feature determining module further comprises:
[0144] a third energy determining module configured to determine a plurality of energies after applying a beamforming algorithm to the detected segment audio at a plurality of target orientations;
[0145] an energy difference determining module configured to determine a plurality of energy differences between the plurality of energies of the detected segment audio at the plurality of target orientations and the plurality of energies after applying the beamforming algorithm to the detected segment audio at the plurality of target orientations;
[0146] a second difference determining module configured to determine a plurality of differences between the plurality of energy differences and a second wake-up feature; and
[0147] a second detection feature determining module configured to determine a second detection feature based on a maximum difference and a minimum difference in the plurality of differences.
[0148] Example 19. The apparatus of any one of examples 13-18, wherein the wearer audio detecting module comprises:
[0149] a second determining module configured to determine whether a frame corresponding to the first detection feature is a wearer frame or a suppressor wearer frame;
[0150] a first start point determining module configured to determine a first frame in a plurality of frame set as a start point of the wearer audio of the detected segment in response to a plurality of consecutive wearer frames in the plurality of frame set satisfying a first condition; and
[0151] a first end point determining module configured to determine a last frame in the plurality of frame set as an end point of the wearer audio of the detected segment in response to a plurality of consecutive suppressor wearer frames in the plurality of frame set satisfying a second condition.
[0152] Example 20. The apparatus of any one of examples 13-19, wherein the second determining module comprises:
[0153] a first wearer frame determining module configured to determine the frame as the wearer frame in response to the first detection feature being greater than or equal to a first threshold; or
[0154] a first suppressor frame determination module configured to determine, in response to the first detection feature being less than the first threshold, a frame corresponding to the first detection feature as the suppressor wearer frame.
[0155] Example 21. The apparatus of any one of examples 13-20, wherein the wearer audio detection module further comprises:
[0156] a third determination module configured to determine a falling edge and a rising edge of the second detection feature;
[0157] a second onset determination module configured to determine, in response to a height of the falling edge satisfying a second threshold, an onset of the falling edge as an onset of the wearer audio of the detection segment; and
[0158] a second tail determination module configured to determine, in response to a height of the falling edge satisfying a third threshold, a tail of the rising edge as a tail of the wearer audio of the detection segment.
[0159] Example 22. The apparatus of any one of examples 13-21, wherein the wearer audio detection module further comprises:
[0160] a fourth threshold determination module configured to determine, based on a first rising edge or a first falling edge of the second detection feature, a fourth threshold by initializing the second detection feature of the wearer audio of the detection segment and the second detection feature of the suppressor wearer audio of the detection segment;
[0161] a first update module configured to update, in response to the second detection feature being less than or equal to the fourth threshold, the second detection feature of the wearer audio of the detection segment by determining a frame corresponding to the second detection feature as the wearer frame;
[0162] a second update module configured to update, in response to the second detection feature being greater than the fourth threshold, the second detection feature of the suppressor wearer audio of the detection segment by determining a frame corresponding to the second detection feature as the suppressor wearer frame; and
[0163] a third update module configured to update the fourth threshold based on the updated second detection feature of the wearer audio of the detection segment and the updated second detection feature of the suppressor wearer audio of the detection segment.
[0164] Example 23. The apparatus of any one of examples 13-22, wherein the wearer audio detection module further comprises:
[0165] a third start point determination module configured to determine, in response to a plurality of consecutive wearer frames in the plurality of frame sets satisfying a second condition and a third condition simultaneously, a first frame in the plurality of frame sets as a start point of the wearer audio of the detection segment; and
[0166] a third end point determination module configured to determine, in response to a plurality of consecutive wearer-suppressed frames in the plurality of frame sets satisfying a fourth condition, an end frame in the plurality of frame sets as an end point of the wearer audio of the detection segment; or
[0167] a fourth end point determination module configured to determine, in response to a plurality of consecutive wearer-suppressed frames in the plurality of frame sets satisfying a fifth condition, an end frame in the plurality of frame sets satisfying the fifth condition as the end point of the wearer audio of the detection segment.
[0168] Example 24. The apparatus of any of Examples 13-23, further comprising:
[0169] an obtaining module configured to obtain the wake-up segment audio and the detection segment audio based on timestamps; and
[0170] a converting module configured to convert the wake-up segment audio and the detection segment audio from a time domain to a time-frequency domain.
[0171] Example 25. An electronic device, comprising:
[0172] a processor; and
[0173] a memory coupled with the processor, the memory having instructions stored therein that, when executed by the processor, cause the electronic device to perform actions comprising:
[0174] determining, based on an energy of a wake-up segment audio, a wake-up feature of the wake-up segment audio;
[0175] determining, based on a detection segment audio and the wake-up feature, a detection feature of the detection segment audio; and
[0176] detecting, based on the detection feature, wearer audio in the detection segment audio.
[0177] Example 26. The electronic device of Example 25, wherein determining, based on an energy of a wake-up segment audio, a wake-up feature of the wake-up segment audio comprises:
[0178] determining, based on the energy of the wake-up segment audio and an energy of wearer-suppressed audio of the wake-up segment, a first wake-up feature.
[0179] Example 27. The electronic device of any of examples 25-26, wherein determining a first wake-up feature based on the energy of the wake-up segment audio and the energy of the wake- wearer audio of the wake-up segment comprises:
[0180] determining a separation matrix for the wake-up segment audio;
[0181] separating the wake-up segment audio into a plurality of audios comprising the wake- wearer audio of the wake-up segment and the wake-wearer audio of the wake-up segment; and
[0182] determining the first wake-up feature based on a difference between the energy of the wake-up segment audio and the energy of the wake-wearer audio.
[0183] Example 28. The electronic device of any of examples 25-27, wherein determining a wake-up feature for the wake-up segment audio based on the energy of the wake-up segment audio further comprises:
[0184] determining a plurality of energies after applying a beamforming algorithm for the wake-up segment audio at a plurality of target orientations; and
[0185] determining a second wake-up feature based on the plurality of energies of the wake-up segment audio at the plurality of target orientations and the plurality of energies after applying a beamforming algorithm for the wake-up segment audio at a plurality of target orientations.
[0186] Example 29. The electronic device of any of examples 25-28, wherein determining a detection feature for the detection segment audio based on the detection segment audio and the wake-up feature comprises:
[0187] determining an energy of a wake-wearer audio of the detection segment based on the separation matrix for the wake-up segment audio;
[0188] determining a difference between the energy of the detection segment audio and the energy of the wake-wearer audio of the detection segment; and
[0189] determining a first detection feature based on the difference and the first wake-up feature.
[0190] Example 30. The electronic device of any of examples 25-29, wherein determining a detection feature for the detection segment audio based on the detection segment audio and the wake-up feature further comprises:
[0191] determining a plurality of energies after applying a beamforming algorithm for the detection segment audio at a plurality of target orientations;
[0192] determining a plurality of energy differences between the plurality of energies of the detection segment audio at the plurality of target orientations and the plurality of energies after applying a beamforming algorithm for the detection segment audio at a plurality of target orientations.
[0193] determining a plurality of differences of the plurality of energy differences and the second wake-up feature; and
[0194] determining a second detection feature based on a maximum difference and a minimum difference of the plurality of differences.
[0195] Example 31. The electronic device of any of examples 25-30, wherein detecting wearer audio in the detection segment audio based on the detection feature comprises:
[0196] determining whether a frame corresponding to the first detection feature is a wearer frame or a suppress-wearer frame;
[0197] in response to a plurality of consecutive wearer frames in the plurality of frame sets satisfying a first condition, determining a first-in-time frame in the plurality of frame sets as a start point of the wearer audio of the detection segment; and
[0198] in response to a plurality of consecutive suppress-wearer frames in the plurality of frame sets satisfying a second condition, determining a last-in-time frame in the plurality of frame sets as an end point of the wearer audio of the detection segment.
[0199] Example 32. The electronic device of any of examples 25-31, wherein determining whether a frame corresponding to the first detection feature is a wearer frame or a suppress-wearer frame comprises:
[0200] in response to the first detection feature being greater than or equal to a first threshold, determining the corresponding frame as the wearer frame; or
[0201] in response to the first detection feature being less than the first threshold, determining the corresponding frame as the suppress-wearer frame.
[0202] Example 33. The electronic device of any of examples 25-32, wherein detecting wearer audio in the detection segment audio based on the detection feature further comprises:
[0203] determining a falling edge and a rising edge of the second detection feature;
[0204] in response to a height of the falling edge satisfying a second threshold, determining a start point of the falling edge as a start point of the wearer audio of the detection segment; and
[0205] in response to a height of the falling edge satisfying a third threshold, determining an end point of the rising edge as an end point of the wearer audio of the detection segment.
[0206] Example 34. The electronic device of any of examples 25-33, wherein detecting wearer audio in the detection segment audio based on the detection feature further comprises:
[0207] determining a fourth threshold value based on a first rising edge or a first falling edge of the second detection feature by initializing the second detection feature of the wearer audio of the detection segment and the second detection feature of the non-wearer audio of the detection segment;
[0208] updating the second detection feature of the wearer audio of the detection segment by determining a frame corresponding to the second detection feature as the wearer frame in response to the second detection feature being less than or equal to the fourth threshold value;
[0209] updating the second detection feature of the non-wearer audio of the detection segment by determining a frame corresponding to the second detection feature as the non-wearer frame in response to the second detection feature being greater than the fourth threshold value; and
[0210] updating the fourth threshold value based on the updated second detection feature of the wearer audio of the detection segment and the updated second detection feature of the non-wearer audio of the detection segment.
[0211] Example 35. The electronic device of any of examples 25-34, wherein detecting wearer audio in the detection segment audio based on the detection feature further comprises:
[0212] determining a first frame of the plurality of frame sets as a start point of the wearer audio of the detection segment in response to consecutive ones of the plurality of wearer frames in the plurality of frame sets simultaneously satisfying a second condition and a third condition; and
[0213] determining a last frame of the plurality of frame sets as an end point of the wearer audio of the detection segment in response to consecutive ones of the plurality of non-wearer frames in the plurality of frame sets satisfying a fourth condition; or
[0214] determining a last frame of the plurality of frame sets satisfying a fifth condition as the end point of the wearer audio of the detection segment in response to consecutive ones of the plurality of non-wearer frames in the plurality of frame sets satisfying the fifth condition.
[0215] Example 36. The electronic device of any of examples 25-35, the actions further comprising:
[0216] obtaining the wake-up segment audio and the detection segment audio based on timestamps; and
[0217] converting the wake-up segment audio and the detection segment audio from a time domain to a time-frequency domain.
[0218] Example 37. A computer-readable storage medium having stored thereon computer- executable instructions, wherein the computer-executable instructions, when executed by a processor, implement a method according to any of examples 1-12.
[0219] Example 38. A computer program product tangibly stored in a computer-readable medium and comprising computer-executable instructions that, when executed by a device, cause the device to perform a method according to any of examples 1-12.
[0220] Although the present disclosure has been described in some detail with specific reference to structure features and / or method logical actions, it is understood that the subject matter defined in the appended claims is not necessarily limited to the particular features or acts described above. Rather, the particular features and acts described above are merely illustrative examples of implementing the claims.
Claims
1. A method for detecting wearer audio, comprising: The wake-up characteristics of the wake-up segment audio are determined based on its energy. Based on the detected audio segment and the wake-up features, the detection features of the detected audio segment are determined; and Based on the detection features, the wearer's audio is detected in the detection segment audio.
2. The method according to claim 1, wherein determining the wake-up characteristics of the wake-up segment audio based on the energy of the wake-up segment audio includes: A first wake-up feature is determined based on the energy of the wake-up segment audio and the energy of the wake-up segment's suppression of the wearer's audio.
3. The method of claim 2, wherein determining the first wake-up feature based on the energy of the wake-up segment audio and the energy of the wake-up segment's suppressive wearer audio comprises: Determine the separation matrix for the wake-up segment audio; The wake-up segment audio is separated into multiple audio segments, including the wearer audio segment of the wake-up segment and the wearer-suppressing audio segment of the wake-up segment; as well as The first wake-up feature is determined based on the difference between the energy of the wake-up segment audio and the energy of the suppressor wearer audio.
4. The method according to claim 3, wherein determining the wake-up characteristics of the wake-up segment audio based on the energy of the wake-up segment audio further includes: Determine multiple energies after applying a beamforming algorithm to the wake-up segment audio at multiple target locations; as well as The second wake-up feature is determined based on the multiple energies of the wake-up segment audio at multiple target locations and the multiple energies after applying a beamforming algorithm to the wake-up segment audio at multiple target locations.
5. The method according to claim 1, wherein determining the detection features of the detection segment audio based on the detection segment audio and the wake-up features includes: Based on the separation matrix for the wake-up segment audio, the energy of the suppressed wearer audio in the detection segment is determined; Determine the difference between the energy of the detected audio segment and the energy of the suppressed wearer audio segment; as well as Based on the difference and the first wake-up feature, the first detection feature is determined.
6. The method according to claim 1, wherein determining the detection features of the detection segment audio based on the detection segment audio and the wake-up features further comprises: Determine multiple energies after applying a beamforming algorithm to the detected audio segment at multiple target locations; Determine multiple energy differences between the detected audio segment at multiple target locations and multiple energies after applying a beamforming algorithm to the detected audio segment at multiple target locations; Determine multiple differences between the multiple energy differences and multiple differences between the second wake-up characteristics; as well as The second detection feature is determined based on the maximum and minimum differences among the plurality of differences.
7. The method of claim 6, wherein detecting wearer audio in the detection segment audio based on the detection feature comprises: Determine whether the frame corresponding to the first detected feature is a wearer frame or a suppressed wearer frame; In response to the presence of multiple consecutive wearer frames in a set of multiple frames satisfying a first condition, the frame with the highest temporal order in the set of multiple frames is determined as the starting point of the wearer audio in the detection segment; as well as In response to the fact that there are consecutive wearer-suppressing frames in the plurality of frame sets that satisfy the second condition, the frame with the last time order in the plurality of frame sets is determined as the end point of the wearer audio of the detection segment.
8. The method of claim 7, wherein determining whether the frame corresponding to the first detection feature is a wearer frame or a suppressed wearer frame comprises: In response to the first detected feature being greater than or equal to the first threshold, the corresponding frame is determined to be the wearer frame; or In response to the first detection feature being less than the first threshold, the corresponding frame is determined to be the inhibited wearer frame.
9. The method of claim 6, wherein detecting wearer audio in the detection segment audio based on the detection feature further comprises: Determine the falling edge and rising edge of the second detection feature; In response to the height of the falling edge satisfying a second threshold, the starting point of the falling edge is determined as the starting point of the wearer's audio in the detection segment; as well as In response to the falling edge height satisfying a third threshold, the tail point of the rising edge is determined as the tail point of the wearer's audio in the detection segment.
10. The method of claim 7, wherein detecting wearer audio in the detection segment audio based on the detection feature further comprises: Based on the first rising edge or the first falling edge of the second detection feature, a fourth threshold is determined by initializing the second detection feature of the wearer audio of the detection segment and the second detection feature of the suppressed wearer audio of the detection segment. In response to the second detection feature being less than or equal to the fourth threshold, the second detection feature of the wearer audio in the detection segment is updated by determining that the frame corresponding to the second detection feature is the wearer frame; In response to the second detection feature being greater than the fourth threshold, the second detection feature of the suppressed wearer audio in the detection segment is updated by determining that the frame corresponding to the second detection feature is the suppressed wearer frame; as well as The fourth threshold is updated based on the second detection feature of the wearer audio in the updated detection segment and the second detection feature of the suppressed wearer audio in the updated detection segment.
11. The method of claim 6, wherein detecting wearer audio in the detection segment audio based on the detection feature further comprises: In response to the fact that there are consecutive wearer frames in the plurality of frame sets that simultaneously satisfy the second condition and the third condition, the first frame in the plurality of frame sets is determined as the starting point of the wearer audio of the detection segment; as well as In response to the fact that there are consecutive wearer-suppressing frames in the plurality of frame sets that satisfy the fourth condition, the last frame in the plurality of frame sets is determined to be the end point of the wearer audio of the detection segment. or In response to the occurrence of consecutive wearer-suppressing frames satisfying a fifth condition in the plurality of frame sets, the last frame in the plurality of frame sets that satisfies the fifth condition is determined as the tail point of the wearer audio of the detection segment.
12. The method according to claim 1, further comprising: The wake-up segment audio and the detection segment audio are obtained based on the timestamp; as well as The wake-up segment audio and the detection segment audio are converted from the time domain to the time-frequency domain.
13. A device for detecting wearer audio, comprising: The wake-up feature determination module is configured to determine the wake-up features of the wake-up segment audio based on the energy of the wake-up segment audio. The detection feature determination module is configured to determine the detection features of the detection segment audio based on the detection segment audio and the wake-up features; and The wearer audio detection module is configured to detect wearer audio in the detection segment audio based on the detection features.
14. An electronic device comprising: processor; as well as A memory coupled to the processor, the memory having instructions stored therein, which, when executed by the processor, cause the electronic device to perform the method according to any one of claims 1 to 12.
15. A computer program product comprising computer-executable instructions, wherein the computer-executable instructions are executed by a processor to implement the method according to any one of claims 1 to 12.
Citation Information
Patent Citations
Voice wake-up method and device, chip, electronic equipment and storage medium
CN112951243A
Voice signal processing method and device, storage medium, electronic equipment and vehicle
CN114783458A
Voice wake-up method and device, electronic equipment and storage medium
CN117198342A
Voice processing method, electronic device, storage medium and computer program product
CN118155618A
Pre-wakeword speech processing
US10192546B1