Method, device, equipment and product for detecting audio of wearer

By extracting the wake-up features of the wake-up segment audio and the detection features of the detection segment audio, the problem of distinguishing between the wearer's voice and environmental noise in wearable devices is solved, achieving accurate wearer audio detection and improving the user experience.

CN121237132APending Publication Date: 2025-12-30BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410851558.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-06-27
Publication Date
2025-12-30

AI Technical Summary

Technical Problem

Existing wearable devices struggle to effectively distinguish between the wearer's voice and ambient noise in audio signal processing, leading to high hardware costs and increased computational complexity, which negatively impacts user experience.

Method used

By extracting the wake-up features of the wake-up segment audio and combining them with the wake-up features extracted from the detection segment audio, the detection features of the detection segment audio are determined to achieve wearer audio detection. The energy-related features of the wake-up segment audio are used to identify external interference, enabling accurate and real-time detection of wearer audio.

Benefits of technology

It requires no additional hardware assistance, accurately identifies external interference, provides personalized audio processing and applications, and enhances the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121237132A_ABST
    Figure CN121237132A_ABST
Patent Text Reader

Abstract

The embodiment of the invention relates to a method, a device, equipment and a product for detecting audio of a wearer. The method includes determining wake-up features of the wake-up segment audio based on energy of the wake-up segment audio, and determining detection features of the detection segment audio based on the detection segment audio and the wake-up features. The method further includes detecting a wearer audio in the detection segment audio based on the detection feature. According to the embodiment of the invention, the audio of the wearer can be effectively detected in real time by referring to the energy related characteristics of the audio of the wake-up segment, so that the accuracy of audio detection is improved, a reliable basis is provided for subsequent audio processing and application, and the user experience is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure generally relates to the field of computers, and more specifically to methods, apparatus, electronic devices, and program products for detecting the audio of a wearer. Background Technology

[0002] Wearable devices are portable devices that can be worn directly on the body or integrated into a user's clothing or accessories. Wearable devices typically exist in the form of portable accessories with some computing capabilities, capable of connecting to mobile phones and various other terminals. Mainstream product forms include watches, shoes, glasses, and other product forms such as smart clothing, backpacks, canes, and accessories.

[0003] In wearable device applications, audio functionality, a crucial component, aims to provide personalized services to users. To achieve this, advanced audio processing technologies such as noise suppression are widely used. With continuous technological advancements, future wearable devices will be able to process acquired audio signals more intelligently, thereby offering more personalized and convenient services. Summary of the Invention

[0004] Embodiments of this disclosure provide a method, apparatus, electronic device, and program product for detecting wearer audio.

[0005] According to a first aspect of this disclosure, a method for detecting wearer audio is provided. The method includes determining wake-up features of a wake-up segment audio based on the energy of that wake-up segment audio. The method also includes determining detection features of the detection segment audio based on the detection features and the wake-up features. Furthermore, the method includes detecting wearer audio in the detection segment audio based on the detection features.

[0006] In a second aspect of this disclosure, an apparatus for detecting wearer audio is provided. The apparatus includes a wake-up feature determination module configured to determine wake-up features of a wake-up segment audio based on the energy of that segment. The apparatus also includes a detection feature determination module configured to determine detection features of a detection segment audio based on the detection segment audio and the wake-up features. Furthermore, the apparatus includes a wearer audio detection module configured to detect wearer audio in the detection segment audio based on the detection features.

[0007] In a third aspect of this disclosure, an electronic device is provided. The electronic device includes a processor and a memory coupled to the processor, the memory having instructions stored therein, which, when executed by the processor, cause the electronic device to perform the method according to the first aspect.

[0008] In a fourth aspect of this disclosure, a computer program product is provided that stores computer-executable instructions thereon, wherein the computer-executable instructions are executed by a processor to implement the method of the first aspect.

[0009] The summary section is intended to present the chosen concepts in a simplified form, which will be further described in the detailed description below. The summary section is not intended to identify key or principal features of the claimed subject matter, nor is it intended to limit the scope of the claimed subject matter. Attached Figure Description

[0010] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. In the drawings, the same or similar reference numerals denote the same or similar elements, wherein:

[0011] Figure 1 A schematic diagram of an example environment in which some embodiments of this disclosure may be implemented is shown;

[0012] Figure 2 Flowcharts of methods for detecting wearer audio according to some embodiments of this disclosure are shown;

[0013] Figure 3 A schematic diagram of an architecture for detecting wearer audio, representing some embodiments of this disclosure, is shown.

[0014] Figure 4A Schematic diagrams illustrating some embodiments of the present disclosure for determining wake-up characteristics during the registration phase are shown;

[0015] Figure 4B Schematic diagrams illustrating some embodiments of the present disclosure for determining detection features during the detection phase are shown;

[0016] Figure 5A Schematic diagrams illustrating some embodiments of the present disclosure for determining a first wake-up feature are shown;

[0017] Figure 5B Schematic diagrams illustrating some embodiments of the present disclosure for determining a second wake-up feature are shown;

[0018] Figure 6A Schematic diagrams illustrating some embodiments of the present disclosure for determining a first detection feature are shown;

[0019] Figure 6B Schematic diagrams illustrating some embodiments of the present disclosure for determining a second wake-up feature are shown;

[0020] Figures 7A-7D The illustration shows schematic diagrams of some embodiments of the present disclosure for detecting wearer audio based on wake-up features and detection features;

[0021] Figure 8 Block diagrams of apparatus for detecting wearer audio according to some embodiments of the present disclosure are shown; and

[0022] Figure 9 Block diagrams of electronic devices according to some embodiments of the present disclosure are shown.

[0023] In all the accompanying figures, the same or similar reference numerals denote the same or similar elements. Detailed Implementation

[0024] It is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and related provisions.

[0025] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.

[0026] For example, upon receiving a user's proactive request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as the electronic device, application, server, or storage medium performing the operations of this disclosed technical solution, based on the prompt message.

[0027] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.

[0028] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.

[0029] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0030] In the description of embodiments of this disclosure, the term "comprising" and similar terms should be understood as open-ended inclusion, i.e., "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The terms "first", "second", etc., may refer to different or the same objects unless explicitly stated. Other explicit and implicit definitions may also be included below.

[0031] As mentioned earlier, audio functionality is crucial in wearable device applications. To provide a personalized user experience, it's typically necessary to distinguish between the wearer's voice and ambient noise, enabling targeted processing and application of the wearer's audio signal. One approach involves using hardware devices like bone conduction microphones to specifically pick up the wearer's audio signal. However, this method places strict requirements on the microphone's pickup quality and placement, and undoubtedly increases hardware costs. Another technique uses voiceprint verification to authenticate the wearer's audio. This method requires prior acquisition of the user's voiceprint information, increasing user complexity. Furthermore, voiceprint verification technology has high computational and parameter overhead, making it difficult to balance computational power and accuracy on resource-constrained platforms.

[0032] In the embodiments of this disclosure, wake-up features of the user wake-up segment audio are extracted, and detection features of the detection segment audio are determined by using the detection segment audio and the wake-up features extracted from the wake-up segment. Wearer audio detection is then achieved based on these detection features. This method of detecting wearer audio by referencing the energy-related features of the wake-up segment audio can effectively identify external interference and accurately and in real-time detect wearer audio, providing a reliable foundation for subsequent personalized audio processing and applications, thereby improving the user experience.

[0033] Figure 1 A schematic diagram of an example environment 100 in which some embodiments of this disclosure may be implemented is shown. For example... Figure 1 As shown, the wearer audio detection system 120 receives audio 110 from a wearable device. The wearable device can be smart AR glasses, headphones, etc., equipped with multi-channel air conduction microphones to acquire sound. The audio 110 can be a mixed audio stream containing multiple sound sources and environmental noise.

[0034] Continue to refer to Figure 1After analysis and processing by the wearer audio detection system 120, the start point 132 and end point 134 of the wearer audio 130 portion in the audio 110 can be obtained. These two time points can provide the precise range of the wearer audio 130 within the entire audio segment, thereby facilitating more detailed analysis and application of the wearer's voice. In some embodiments, the wearer audio detection system 120 can obtain the wake-up segment audio based on the timestamp of the audio 110, thereby enabling the wearer audio detection system 120 to effectively and in real-time detect the wearer audio by referring to the energy-related characteristics of the wake-up segment audio.

[0035] Understandably, before performing analysis on audio 110, necessary preprocessing, such as sampling rate conversion and noise filtering, can be performed on audio 110 to improve the accuracy of subsequent analysis.

[0036] The following will combine Figures 2 to 9 The process according to embodiments of this disclosure is described in detail. For ease of understanding, the specific data mentioned in the following description are exemplary and not intended to limit the scope of this disclosure. It is understood that the embodiments described below may also include additional actions not shown and / or actions shown may be omitted, and the scope of this disclosure is not limited in this respect.

[0037] Figure 2 A flowchart of a method 200 for detecting wearer audio, according to some embodiments of this disclosure, is shown. Method 200 can be executed by a device for detecting wearer audio. This device can be implemented in software and / or hardware. The method 200 will now be illustrated illustratively using a device for detecting wearer audio (such as smart AR glasses or headphones) as an example. Reference Figure 2 Method 200 may include boxes 202, 204 and 206.

[0038] In box 202, wake-up characteristics of the wake-up segment audio are determined based on its energy. In some embodiments, the wake-up segment audio refers to the audio used to wake up a wearable device; for a specific device or a specific type of device, there may be different or the same wake-up word. In some embodiments, the wake-up segment audio can be obtained by detecting the timestamp of the wake-up segment using a wake-up algorithm. In some embodiments, the energy of the audio refers to the sound intensity or amplitude contained in the audio signal. The wake-up characteristics of the wake-up segment audio are determined based on its energy, and these characteristics may include a first wake-up characteristic and a second wake-up characteristic. Before obtaining the wake-up characteristics, the wake-up segment audio in the time domain can be converted to the time-frequency domain.

[0039] In box 204, detection features of the detection segment audio are determined based on the detection segment audio and wake-up features. In some embodiments, the detection segment audio refers to audio other than the wake-up segment audio acquired from the wearable device during user use. In some embodiments, the detection segment audio can also be acquired based on a timestamp, i.e. Figure 1 A portion of the audio 110 shown is a detection segment audio. In some embodiments, detection features of the detection segment audio can be obtained based on the acquired detection segment audio and the aforementioned acquired wake-up features. The detection features may include a first detection feature and a second detection feature. In some embodiments, the detection segment audio in the time domain can be converted to the time-frequency domain.

[0040] In box 206, wearer audio is detected in the detection segment audio based on detection features. In some embodiments, the start and end points of the wearer audio can be detected in the detection segment audio based on a first detection feature and / or a second detection feature, i.e., in Figure 1 As shown, 132 is the starting point of the wearer's audio. Figure 1 As shown in Figure 134, these are the end points of the wearer's audio. These two time points can provide the exact range of the wearer's audio within the entire audio segment, thus facilitating more detailed analysis and application of the wearer's voice later on.

[0041] In the embodiments of this disclosure, wake-up features of the user's wake-up segment audio are extracted based on the energy of the wake-up segment audio. Detection features of the detection segment audio are determined using the detection segment audio and the wake-up features extracted from the wake-up segment, and then wearer audio detection is achieved based on these detection features. This method of detecting wearer audio by referencing the energy-related features of the wake-up segment audio can effectively identify external interference and accurately and in real-time detect wearer audio, providing a reliable foundation for subsequent personalized audio processing and applications, thereby improving the user experience.

[0042] Figure 3 A schematic diagram of an architecture 300 for detecting wearer audio, according to some embodiments of this disclosure, is shown. Figure 3 As shown, at 310, audio is acquired. The audio can be acquired through a wearable device configured with a multi-channel air conduction microphone, and therefore can be an audio signal that mixes multiple sound sources. The wearable device can be a smart AR glasses device or a headset, etc.

[0043] refer to Figure 3 It is understandable that this disclosure mainly involves two stages, one of which is the registration stage (such as...). Figure 3 As shown in 320), one is the detection stage (such as...). Figure 3(As shown in 330 and 340). During the registration phase, relevant parameters of the wearer for the wake-up segment audio are calculated. During the detection phase, the calculated parameters of the wearer for the wake-up segment audio are combined with the audio of the detection segment to determine the detection characteristics of the detection segment audio. This allows for the detection of the start and end points of the wearer's audio based on these characteristics. This method of detecting the wearer based on audio energy can accurately identify external interference and thus detect the wearer's audio without the need for additional hardware such as bone conduction headphones.

[0044] To more clearly illustrate the process of calculating the wearer's relevant parameters for the wake-up audio segment at 320, the following will combine... Figure 4A To describe. Figure 4A A schematic diagram of some embodiments of the present disclosure for determining wake-up feature 400A during the registration phase is shown.

[0045] refer to Figure 4A At 410A, the wake-up segment audio is extracted. In some embodiments, the wake-up segment's timestamp can be detected using a wake-up algorithm to obtain the wake-up segment's audio data. For example, the wake-up algorithm can be a speech feature-based wake-up algorithm or a data packet wake-up algorithm, etc., and this disclosure does not limit this. At 420A, a time-frequency domain forward transform is performed on the wake-up segment audio. In some embodiments, a time-frequency domain forward transform can be performed on the time-domain audio signals of each channel's air conduction microphone to obtain the time-frequency domain audio signals of each channel's air conduction microphone. In some embodiments, optional transform methods include short-time Fourier forward transform, sub-band decomposition, etc., and this disclosure does not limit this.

[0046] Continue to refer to Figure 4A In the 430A, audio is separated to determine the first wake-up feature. To better illustrate the process of separating audio to determine the first wake-up feature in the 430A, the following will combine... Figure 5A To describe. Figure 5A A schematic diagram of a 500A for determining a first wake-up characteristic is shown, representing some embodiments of this disclosure. (See reference...) Figure 5A In 510A, the separation matrix of the wake-up segment is determined. In some embodiments, a separation algorithm can be used to extract the separation matrix information of the wake-up segment. Optional separation algorithms include, but are not limited to, independent vector analysis algorithms, etc., and this disclosure does not limit them.

[0047] like Figure 5AAs shown in 520A, the wake-up segment audio is separated. In some embodiments, the audio mixed with multiple sound sources can be separated into individual signal sources using the separation algorithm mentioned in 510A. For example, the audio acquired through a multi-channel air conduction microphone can be separated into multiple individual audio channels, such as a wearer audio channel and a suppressed wearer audio channel. In 530A, the energy of the suppressed wearer audio is determined. Given that the microphone is configured on a wearable device, the energy of the suppressed wearer audio channel is lower than that of the wearer audio channel, thus clearly determining the energy of the suppressed wearer audio channel. In 540A, the energy difference is calculated to determine the first wake-up feature. In some embodiments, for a single original microphone, the first wake-up feature can be the energy difference between the original microphone audio channel and the channel suppressing the wearer signal. Assuming the audio energy of the original microphone channel is EA, and after separation processing, assuming the energy of the suppressed wearer audio is EB, the energy difference is EA-EB. This can also be understood as the first wake-up feature being the energy of the wearer audio in the wake-up segment.

[0048] Return to reference Figure 4A In the 440A, a beamforming algorithm is applied to determine the second wake-up feature. To better illustrate the process of separating audio to determine the first wake-up feature in the 440A, the following will combine... Figure 5B To describe. Figure 5B A schematic diagram of 500B for determining a second wake-up feature according to some embodiments of the present disclosure is shown. Referring to FIG5, in 510B, a beamforming algorithm with multiple target orientations is applied to the wake-up segment audio. In some embodiments, the microphone time-frequency domain audio signal is scanned using a beamforming algorithm with multiple target orientations. It is understood that the number and direction of target orientations can be determined according to computing power and application scenario. For example, an orientation can be set as a target orientation every 45 degrees in the [0°, 180°] and [180°, 360°] directions to apply the beamforming algorithm. If the computing power is strong and the application scenario is complex, an orientation can be selected as a target orientation every 10° to apply the beamforming algorithm. In some embodiments, an enhanced beamforming algorithm or a suppressed beamforming algorithm can be selected, and the present disclosure does not limit this.

[0049] Referring back to Figure 5, at 520B, the relative energy difference between the original microphone and the audio after beamforming is calculated for different orientations. For example, if five orientations are set, the relative energy difference can be calculated as [A1, A2, A3, A3, A5]. At 530B, the relative energy difference is determined as the second wake-up feature. Thus, the second wake-up feature is [A1, A2, A3, A3, A5]. It is also understandable that, for the sake of data processing efficiency and simplicity, before calculating the relative energy difference, at 522B, a certain number of frequency points within the wake-up duration can be selected, and these frequency points can be averaged.

[0050] Back Figure 3 After determining the wearer's relevant parameters for the wake-up segment audio during the registration phase, the detection phase can proceed to detect the wearer's audio during the detection segment. At step 330, the relevant parameters obtained from the wake-up segment are used to calculate characteristics such as audio energy. To better illustrate the process of calculating audio energy and other characteristics using the relevant parameters obtained from the wake-up segment at step 330, the following will combine... Figure 4B To describe. Figure 4B A schematic diagram of a 400B for determining detection features during the detection phase, according to some embodiments of this disclosure, is shown. (See reference...) Figure 4B In 410B, a forward time-domain transform is performed on the audio signal of the detection segment. A forward time-frequency transform is performed on the time-domain audio signals of each channel of the air conduction microphone in the detection segment to obtain the time-frequency audio signals of each channel of the air conduction microphone in the detection segment. Optional transform methods include short-time Fourier transform, sub-band decomposition, etc.

[0051] Continue to refer to Figure 4B In 420B, the first wake-up feature is matched to obtain the first detection feature. To better illustrate the process of obtaining the first detection feature by matching the first wake-up feature in 420B, the following will combine... Figure 6A To describe. Figure 6A A schematic diagram of a 600A for determining a first detection feature, according to some embodiments of this disclosure, is shown. (See reference...) Figure 6A In the 610A, the audio of the detection segment is separated using a separation matrix of the wake-up segment. The separation matrix effectively separates audio signals from different sound sources. In some embodiments, the wake-up segment separation matrix can be fixed to separate the time-frequency domain audio signals of each channel of the air conduction microphone in the detection segment. In some embodiments, the wake-up segment separation matrix can also be used to initialize the online separation algorithm parameters of the detection segment audio, wherein the separation matrix parameters corresponding to the channel number that suppresses the wearer's signal in the wake-up segment separation matrix can be fixed, and the separation matrix parameters of other channel numbers can be updated in real time, thereby enabling the separation of the time-frequency domain audio signals of each channel of the air conduction microphone in the detection segment.

[0052] By using the wake-up segment separation matrix to separate the detection segment audio, the efficiency of audio processing can be improved, the error that may be caused by recalculating the separation matrix can be reduced, and the consistency of processing for the wake-up segment audio and the detection segment audio can be ensured.

[0053] Continue to refer to Figure 6AAfter obtaining the audio from each separated channel, the energy difference can be calculated at 620A. In some embodiments, the energy difference between the original microphone channel energy and the energy of the corresponding channel number that suppresses the wearer's signal after separation processing can be calculated. Before separation processing, it is assumed that the audio energy of the original microphone channel is EC; after separation processing, it is assumed that the audio energy suppressing the wearer's signal is ED, and the energy difference is EC-ED. Then, at 630A, the ratio of this energy difference to the first wake-up feature is recorded as the first detection feature. Combined with... Figure 5A As shown in Figure 540, the first detection feature can be (EC-ED) / (EA-EB), which can be understood as the proportion of the energy difference change between the detection segment and the wake-up segment. The larger the ratio, the greater the relative intensity of the wearer's audio or the less the relative intensity of the suppressed wearer's audio.

[0054] Back Figure 4B In 430B, the second wake-up feature is matched to obtain the second detection feature. To better illustrate the process of matching the second wake-up feature to obtain the second detection feature, the following will combine... Figure 6B To describe. Figure 6B A schematic diagram of a 600B for determining a second wake-up characteristic, according to some embodiments of this disclosure, is shown. (See reference...) Figure 6B In 610B, beamforming algorithms with multiple target orientations are applied to the audio in the detection segment. In some embodiments, the time-frequency domain audio signal of the microphone in the detection segment can be scanned using a beamforming algorithm with the same target orientation as the wake-up segment. If five orientations are set in the wake-up segment, then five orientations are also set in the detection segment. Thus, in 620B, the relative energy difference between the original microphone and the audio after applying the beamforming algorithm in different orientations can be calculated, for example, the energy difference can be calculated as [B1, B2, B3, B3, B5].

[0055] Continue to refer to Figure 6B Next, at 630B, the difference between the relative energy difference in each direction and the difference between the second wake-up feature in the corresponding direction is calculated. Combining the second wake-up feature obtained at 520B as shown in Figure 5, which is [A1,A2,A3,A3,A5], the difference can be obtained as [|B1-A1|,|B2-A2|,|B3-A3|,|B4-A4|,|B5-A5|]. Then, at 640B, the difference between the maximum and minimum values ​​of the difference is recorded as the second detection feature. Assuming that B1-A1 is the maximum value and B5-A5 is the minimum value, then the second detection feature is ||B1-A1|-|B5-A5||.

[0056] return Figure 3 At 340, the wearer's audio is detected based on the acquired features. To better illustrate the process of detecting the wearer's audio based on the acquired features, the following will combine... Figures 7A-7B To describe. Figure 7A A schematic diagram illustrating some embodiments of the present disclosure of detecting wearer audio 700A based on a first detection feature is shown. For example... Figure 7A As shown, at 710A, it is determined whether the first detection feature is greater than or equal to the first threshold. If it is greater than or equal to the first threshold, then at 712A, the signal frame can be identified as a wearer frame. If the first detection feature is less than the first threshold, then at 714A, the signal can be identified as a suppressed wearer frame. Specifically, the first threshold can be set to 1. If the first detection feature is greater than or equal to 1, the signal frame can be considered a wearer frame. If the first detection feature is less than 1, the signal frame can be considered a suppressed wearer frame (ambient noise frame).

[0057] Continue to refer to Figure 7A After determining whether a frame is a wearer frame or a suppressed wearer frame, it is necessary to determine the start and end points of the wearer frame. Specifically, in 722A, it can be determined whether there are M frames in a consecutive N frame sequence that are all wearer frames. If so, then in 724A, the start point in N can be recorded as the start point of the wearer audio. In 732A, it can be determined whether there are Q frames in a consecutive P frame sequence that are all suppressed wearer frames (ambient noise frames). If so, then in 734A, the end point of the P frame can be recorded as the end point of the wearer audio. In some embodiments, these frames are arranged in chronological order.

[0058] Figure 7B A schematic diagram illustrating some embodiments of the present disclosure of detecting wearer audio 700B based on a second detection feature is shown. For example... Figure 7B As shown, the wearer's audio can be detected based on the rising and falling edges of the second detection feature. Specifically, if the height of the falling edge of the second detection feature exceeds a second threshold, the starting point of the falling edge 710B can be recorded as the starting point of the wearer's audio. If the height of the rising edge of the second detection feature exceeds a third threshold, the ending point of the rising edge 720B can be recorded as the ending point of the wearer's audio.

[0059] Figure 7C A schematic diagram illustrating some embodiments of this disclosure is shown, illustrating how a wearer frame or a suppressed wearer frame 700C is determined based on a second detection feature. For example... Figure 7C As shown, at 710C, the first rising edge or the first falling edge of the second detection feature is determined. At 720C, the second detection feature belonging to the wearer's audio and the second detection feature belonging to the suppressed wearer's audio (ambient noise) are initialized to obtain the fourth threshold. In some embodiments, the fourth threshold can be set as the average of the second detection feature belonging to the wearer's audio and the second detection feature belonging to the suppressed wearer's audio (ambient noise).

[0060] Continue to refer to Figure 7CAt 730C, it is determined whether the second detection feature is less than or equal to the fourth threshold. If so, at 740C, this frame is identified as a wearer frame. Then at 750C, the second detection feature belonging to the wearer audio is updated.

[0061] Continue to refer to Figure 7C At 730C, it is determined whether the second detection feature is less than or equal to the fourth threshold. If not, then at 760C, this frame signal is identified as a suppressed wearer frame. Then at 770C, the second detection feature belonging to suppressed wearer audio is updated.

[0062] like Figure 7C As shown, at 780C, the fourth threshold is updated based on the updated second detection feature belonging to the wearer and the second detection feature belonging to the suppression of wearer audio (ambient noise).

[0063] Figure 7D A schematic diagram illustrating some embodiments of this disclosure of detecting a wearer's audio 700D by combining a first detection feature and a second detection feature is shown. Figure 7D As shown, in 710D, the start and end points of the wearer's audio are determined by combining the first and second detection features. In 722D, it is determined whether there is a frame M1 in N consecutive frames that is a wearer frame identified based on the first detection feature. If so, in 724D, it is determined whether there is a frame M2 in N consecutive frames that is a wearer frame identified based on the second detection feature. If so, in 726D, the start point of the N frames can be recorded as the start point of the wearer's audio.

[0064] like Figure 7D As shown in 732D, a Q1 frame is determined to be a wearer-suppressed frame (ambient noise frame) based on the first detection feature in a consecutive P1 frame. If so, in 734D, the end point of the P1 frame can be recorded as the end point of the detected wearer audio segment.

[0065] Continue to refer to Figure 7D In the 742D, among consecutive P2 frames, there is a Q2 frame that is a wearer suppression frame (ambient noise frame) based on the second detection feature. If so, in the 744D, the end point of the Q2 frame can be recorded as the end point of the wearer audio in the detection segment.

[0066] This detection method, based on the energy-related characteristics of wake-up audio, can accurately identify external interference without relying on additional hardware such as bone conduction headphones, thereby effectively detecting the wearer's voice signal. This provides a foundation for subsequent applications and processing of the wearer's voice signal, enabling personalized service experiences for users.

[0067] Figure 8 A block diagram of a device 800 for detecting wearer audio, according to some embodiments of the present disclosure, is shown. Figure 8 As shown, device 800 includes a wake-up feature determination module 802, configured to determine wake-up features of the wake-up segment audio based on the energy of the wake-up segment audio. Device 800 also includes a detection feature determination module 804, configured to determine detection features of the detection segment audio based on the detection segment audio and the wake-up features. Furthermore, device 800 includes a wearer audio detection module 806, configured to detect wearer audio in the detection segment audio based on the detection features.

[0068] Figure 9 A block diagram of an electronic device 900 according to some embodiments of the present disclosure is shown. Device 900 may be the device or apparatus described in the embodiments of the present disclosure. Figure 9 As shown, device 900 includes a central processing unit (CPU) and / or a graphics processing unit (GPU) 901, which can perform various appropriate actions and processes according to computer program instructions stored in read-only memory (ROM) 902 or loaded from storage unit 908 into random access memory (RAM) 903. The RAM 903 can also store various programs and data required for the operation of device 900. The CPU / GPU 901, ROM 902, and RAM 903 are interconnected via bus 904. Input / output (I / O) interface 905 is also connected to bus 904. Although not shown in... Figure 9 As shown, device 900 may also include a coprocessor.

[0069] Multiple components in device 900 are connected to I / O interface 905, including: input unit 906, such as keyboard, mouse, etc.; output unit 907, such as various types of monitors, speakers, etc.; storage unit 908, such as disk, optical disk, etc.; and communication unit 909, such as network card, modem, wireless transceiver, etc. Communication unit 909 allows device 900 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0070] The various methods or processes described above can be executed by CPU / GPU 901. For example, in some embodiments, the methods can be implemented as computer software programs tangibly contained in a machine-readable medium, such as storage unit 908. In some embodiments, part or all of the computer program can be loaded and / or installed on device 900 via ROM 902 and / or communication unit 909. When the computer program is loaded into RAM 903 and executed by CPU / GPU 901, one or more steps or actions in the methods or processes described above can be performed.

[0071] In some embodiments, the methods and processes described above can be implemented as a computer program product. The computer program product may include a computer-readable storage medium having computer-readable program instructions loaded thereon for performing various aspects of this disclosure.

[0072] Computer-readable storage media can be tangible devices capable of holding and storing instructions for use by an instruction execution device. Computer-readable storage media can be, for example, but not limited to, electrical storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or any suitable combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disc read-only memory (CD-ROM), digital multifunction disc (DVD), memory sticks, floppy disks, mechanical encoding devices, such as punch cards or recessed protrusions storing instructions thereon, and any suitable combination thereof. The computer-readable storage media used herein are not to be construed as transient signals themselves, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides or other transmission media (e.g., light pulses through fiber optic cables), or electrical signals transmitted through wires.

[0073] The computer-readable program instructions described herein can be downloaded from computer-readable storage media to various computing / processing devices, or downloaded via a network, such as the Internet, a local area network (LAN), a wide area network (WAN), and / or a wireless network, to an external computer or external storage device. The network may include copper cables, fiber optic cables, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards them to the computer-readable storage media in the respective computing / processing device.

[0074] Computer program instructions used to perform the operations of this disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, status setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages ​​and conventional procedural programming languages. The computer-readable program instructions may execute entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, electronic circuitry, such as programmable logic circuitry, field-programmable gate arrays (FPGAs), or programmable logic arrays (PLAs), is personalized by utilizing the status information of the computer-readable program instructions to implement various aspects of this disclosure.

[0075] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processing unit of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.

[0076] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.

[0077] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of devices, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction containing one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than those marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0078] The various embodiments of this disclosure have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is chosen to best explain the principles, practical applications, or technical improvements to the technology in the market, or to enable others skilled in the art to understand the embodiments disclosed herein.

[0079] The following are some example implementations of this disclosure.

[0080] Example 1. A method for detecting wearer audio, comprising:

[0081] The wake-up characteristics of the wake-up segment audio are determined based on its energy.

[0082] Based on the detected audio segment and the wake-up features, the detection features of the detected audio segment are determined; and

[0083] Based on the detection features, the wearer's audio is detected in the detection segment audio.

[0084] Example 2. According to the method described in Example 1, determining the wake-up characteristics of the wake-up segment audio based on its energy includes:

[0085] A first wake-up feature is determined based on the energy of the wake-up segment audio and the energy of the wake-up segment's suppression of the wearer's audio.

[0086] Example 3. The method according to any one of Examples 1-2, wherein determining the first wake-up feature based on the energy of the wake-up segment audio and the energy of the wake-up segment's suppressive wearer audio includes:

[0087] Determine the separation matrix for the wake-up segment audio;

[0088] The wake-up segment audio is separated into multiple audio segments, including the wearer audio segment of the wake-up segment and the wearer-suppressing audio segment of the wake-up segment; and

[0089] The first wake-up feature is determined based on the difference between the energy of the wake-up segment audio and the energy of the suppressor wearer audio.

[0090] Example 4. The method according to any one of Examples 1-3, wherein determining the wake-up characteristics of the wake-up segment audio based on the energy of the wake-up segment audio further includes:

[0091] Determine multiple energies after applying beamforming algorithms to the wake-up segment audio at multiple target locations; and

[0092] The second wake-up feature is determined based on the multiple energies of the wake-up segment audio at multiple target locations and the multiple energies after applying a beamforming algorithm to the wake-up segment audio at multiple target locations.

[0093] Example 5. The method according to any one of Examples 1-4, wherein determining the detection features of the detection segment audio based on the detection segment audio and the wake-up features includes:

[0094] Based on the separation matrix for the wake-up segment audio, the energy of the suppressed wearer audio in the detection segment is determined;

[0095] Determine the difference between the energy of the detected audio segment and the energy of the suppressed wearer audio segment; and

[0096] Based on the difference and the first wake-up feature, the first detection feature is determined.

[0097] Example 6. The method according to any one of Examples 1-5, wherein determining the detection features of the detection segment audio based on the detection segment audio and the wake-up features further includes:

[0098] Determine multiple energies after applying a beamforming algorithm to the detected audio segment at multiple target locations;

[0099] Determine multiple energy differences between the detected audio segment at multiple target locations and multiple energies after applying a beamforming algorithm to the detected audio segment at multiple target locations;

[0100] Determine multiple differences between the plurality of energy differences and multiple differences between the second wake-up characteristics; and

[0101] The second detection feature is determined based on the maximum and minimum differences among the plurality of differences.

[0102] Example 7. The method according to any one of Examples 1-6, wherein detecting wearer audio in the detection segment audio based on the detection feature comprises:

[0103] Determine whether the frame corresponding to the first detected feature is a wearer frame or a suppressed wearer frame;

[0104] In response to the presence of multiple consecutive wearer frames satisfying a first condition in a set of multiple frames, the frame with the highest temporal order in the set of multiple frames is determined as the starting point of the wearer audio in the detection segment; and

[0105] In response to the fact that there are consecutive wearer-suppressing frames in the plurality of frame sets that satisfy the second condition, the frame with the last time order in the plurality of frame sets is determined as the end point of the wearer audio of the detection segment.

[0106] Example 8. Determining whether the frame corresponding to the first detection feature is a wearer frame or a suppressed wearer frame according to any one of Examples 1-7 includes:

[0107] In response to the first detected feature being greater than or equal to a first threshold, the corresponding frame is determined to be the wearer frame; or

[0108] In response to the first detection feature being less than the first threshold, the corresponding frame is determined to be the inhibited wearer frame.

[0109] Example 9. The method according to any one of Examples 1-8, wherein detecting wearer audio in the detection segment audio based on the detection feature further comprises:

[0110] Determine the falling edge and rising edge of the second detection feature;

[0111] In response to the falling edge height satisfying a second threshold, the starting point of the falling edge is determined as the starting point of the wearer's audio in the detection segment; and

[0112] In response to the falling edge height satisfying a third threshold, the tail point of the rising edge is determined as the tail point of the wearer's audio in the detection segment.

[0113] Example 10. The method according to any one of Examples 1-9, wherein detecting wearer audio in the detection segment audio based on the detection feature further comprises:

[0114] Based on the first rising edge or the first falling edge of the second detection feature, a fourth threshold is determined by initializing the second detection feature of the wearer audio of the detection segment and the second detection feature of the suppressed wearer audio of the detection segment.

[0115] In response to the second detection feature being less than or equal to the fourth threshold, the second detection feature of the wearer audio in the detection segment is updated by determining that the frame corresponding to the second detection feature is the wearer frame;

[0116] In response to the second detection feature being greater than the fourth threshold, the second detection feature of the suppressed-wearer audio in the detection segment is updated by determining that the frame corresponding to the second detection feature is the suppressed-wearer frame; and

[0117] The fourth threshold is updated based on the second detection feature of the wearer audio in the updated detection segment and the second detection feature of the suppressed wearer audio in the updated detection segment.

[0118] Example 11. The method according to any one of Examples 1-10, wherein detecting wearer audio in the detection segment audio based on the detection feature further comprises:

[0119] In response to the occurrence of consecutive wearer frames in the plurality of frame sets simultaneously satisfying both the second and third conditions, the first frame in the plurality of frame sets is determined as the starting point of the wearer audio in the detection segment; and

[0120] In response to the occurrence of consecutive wearer-suppressing frames satisfying a fourth condition in the plurality of frame sets, the last frame in the sorted plurality of frame sets is determined as the end point of the wearer audio in the detection segment; or

[0121] In response to the occurrence of consecutive wearer-suppressing frames satisfying a fifth condition in the plurality of frame sets, the last frame in the plurality of frame sets that satisfies the fifth condition is determined as the tail point of the wearer audio of the detection segment.

[0122] Example 12. The method according to any one of Examples 1-11 further includes:

[0123] The wake-up segment audio and the detection segment audio are obtained based on timestamps; and

[0124] The wake-up segment audio and the detection segment audio are converted from the time domain to the time-frequency domain.

[0125] Example 13. An apparatus for detecting wearer audio, comprising:

[0126] The wake-up feature determination module is configured to determine the wake-up features of the wake-up segment audio based on the energy of the wake-up segment audio.

[0127] The detection feature determination module is configured to determine the detection features of the detection segment audio based on the detection segment audio and the wake-up features; and

[0128] The wearer audio detection module is configured to detect wearer audio in the detection segment audio based on the detection features.

[0129] Example 14. The apparatus according to Example 13, wherein the wake-up feature determination module includes:

[0130] The first wake-up feature determination module is configured to determine a first wake-up feature based on the energy of the wake-up segment audio and the energy of the wake-up segment suppressing the wearer's audio.

[0131] Example 15. The apparatus according to any one of Examples 13-14, wherein the first wake-up feature determining module comprises:

[0132] A separation matrix determination module is configured to determine a separation matrix for the wake-up segment audio;

[0133] An audio separation module is configured to separate the wake-up segment audio into multiple audio streams, the multiple audio streams including the wearer audio of the wake-up segment and the wearer-suppressing audio of the wake-up segment; and

[0134] A first determining module is configured to determine the first wake-up feature based on the difference between the energy of the wake-up segment audio and the energy of the suppressor wearer audio.

[0135] Example 16. The apparatus according to any one of Examples 13-15, wherein the wake-up feature determination module further comprises:

[0136] The first energy determination module is configured to determine multiple energies after applying a beamforming algorithm to the wake-up segment audio at multiple target locations; and

[0137] The second wake-up feature determination module is configured to determine the second wake-up feature based on the multiple energies of the wake-up segment audio at multiple target locations and the multiple energies after applying a beamforming algorithm to the wake-up segment audio at multiple target locations.

[0138] Example 17. The apparatus according to any one of Examples 13-16, wherein the detection feature determination module comprises:

[0139] The second energy determination module is configured to determine the energy of the suppressed wearer audio in the detection segment based on a separation matrix for the wake-up segment audio.

[0140] The first difference determination module is configured to determine the difference between the energy of the detection segment audio and the energy of the suppressed wearer audio of the detection segment; and

[0141] The first detection feature determination module is configured to determine the first detection feature based on the difference and the first wake-up feature.

[0142] Example 18. The apparatus according to any one of Examples 13-17, wherein the detection feature determination module further comprises:

[0143] The third energy determination module is configured to determine multiple energies after applying a beamforming algorithm to the detected audio segment at multiple target locations.

[0144] An energy difference determination module is configured to determine multiple energies of the detected segment audio at multiple target locations and multiple energies of the detected segment audio after applying a beamforming algorithm at multiple target locations;

[0145] The second difference determination module is configured to determine multiple differences between the plurality of energy differences and the second wake-up features; and

[0146] The second detection feature determination module is configured to determine the second detection feature based on the maximum and minimum differences among the plurality of differences.

[0147] Example 19. The apparatus according to any one of Examples 13-18, wherein the wearer audio detection module comprises:

[0148] The second determining module is configured to determine whether the frame corresponding to the first detection feature is a wearer frame or a suppressed wearer frame.

[0149] The first starting point determination module is configured to, in response to a plurality of consecutive wearer frames satisfying a first condition in a plurality of frame sets, determine the frame with the highest temporal order in the plurality of frame sets as the starting point of the wearer audio of the detection segment; and

[0150] The first tail point determination module is configured to determine the last frame in the multiple frame sets in time order as the tail point of the wearer audio of the detection segment in response to a second condition where there are consecutive multiple suppressed wearer frames in the multiple frame sets.

[0151] Example 20. The apparatus according to any one of Examples 13-19, wherein the second determining module comprises:

[0152] The first wearer frame determination module is configured to determine the corresponding frame as the wearer frame in response to the first detected feature being greater than or equal to a first threshold; or

[0153] The first suppressor frame determination module is configured to determine the corresponding frame as the suppressor wearer frame in response to the first detection feature being less than the first threshold.

[0154] Example 21. The apparatus according to any one of Examples 13-20, wherein the wearer audio detection module further comprises:

[0155] The third determining module is configured to determine the falling edge and rising edge of the second detection feature;

[0156] The second starting point determination module is configured to determine the starting point of the falling edge as the starting point of the wearer's audio in the detection segment in response to the falling edge height satisfying a second threshold; and

[0157] The second tail point determination module is configured to determine the tail point of the rising edge as the tail point of the wearer's audio in the detection segment in response to the height of the falling edge satisfying a third threshold.

[0158] Example 22. The apparatus according to any one of Examples 13-21, wherein the wearer audio detection module further comprises:

[0159] The fourth threshold determination module is configured to determine the fourth threshold based on the first rising edge or the first falling edge of the second detection feature by initializing the second detection feature of the wearer audio of the detection segment and the second detection feature of the suppressed wearer audio of the detection segment.

[0160] The first update module is configured to update the second detection feature of the wearer audio in the detection segment in response to the second detection feature being less than or equal to the fourth threshold by determining that the frame corresponding to the second detection feature is the wearer frame;

[0161] The second update module is configured to update the second detection feature of the suppressed-wearer audio in the detection segment in response to the second detection feature being greater than the fourth threshold by determining that the frame corresponding to the second detection feature is the suppressed-wearer frame; and

[0162] The third update module is configured to update the fourth threshold based on the second detection feature of the wearer audio of the updated detection segment and the second detection feature of the suppressed wearer audio of the updated detection segment.

[0163] Example 23. The apparatus according to any one of Examples 13-22, wherein the wearer audio detection module further comprises:

[0164] The third starting point determination module is configured to, in response to the presence of consecutive wearer frames in the plurality of frame sets simultaneously satisfying both the second and third conditions, determine the first ordered frame in the plurality of frame sets as the starting point of the wearer audio of the detection segment; and

[0165] The third tail point determination module is configured to determine the last frame in the sorted sequence of the plurality of frame sets as the tail point of the wearer audio of the detection segment in response to a fourth condition where consecutive plurality of suppressed wearer frames in the plurality of frame sets satisfy the fourth condition; or

[0166] The fourth tail point determination module is configured to determine the last sorted frame in the plurality of frames that satisfies the fifth condition as the tail point of the wearer audio of the detection segment in response to a plurality of consecutive suppressed wearer frames in the plurality of frame sets satisfying the fifth condition.

[0167] Example 24. The apparatus according to any one of Examples 13-23 further includes:

[0168] The acquisition module is configured to acquire the wake-up segment audio and the detection segment audio based on timestamps; and

[0169] The conversion module is configured to convert the wake-up segment audio and the detection segment audio from the time domain to the time-frequency domain.

[0170] Example 25. An electronic device comprising:

[0171] Processor; and

[0172] A memory coupled to the processor, the memory having instructions stored therein, which, when executed by the processor, cause the electronic device to perform actions, the actions including:

[0173] The wake-up characteristics of the wake-up segment audio are determined based on its energy.

[0174] Based on the detected audio segment and the wake-up features, the detection features of the detected audio segment are determined; and

[0175] Based on the detection features, the wearer's audio is detected in the detection segment audio.

[0176] Example 26. The electronic device according to Example 25, wherein determining the wake-up characteristics of the wake-up segment audio based on the energy of the wake-up segment audio includes:

[0177] A first wake-up feature is determined based on the energy of the wake-up segment audio and the energy of the wake-up segment's suppression of the wearer's audio.

[0178] Example 27. An electronic device according to any one of Examples 25-26, wherein determining the first wake-up feature based on the energy of the wake-up segment audio and the energy of the wake-up segment's suppressive wearer audio includes:

[0179] Determine the separation matrix for the wake-up segment audio;

[0180] The wake-up segment audio is separated into multiple audio segments, including the wearer audio segment of the wake-up segment and the wearer-suppressing audio segment of the wake-up segment; and

[0181] The first wake-up feature is determined based on the difference between the energy of the wake-up segment audio and the energy of the suppressor wearer audio.

[0182] Example 28. An electronic device according to any one of Examples 25-27, wherein determining the wake-up characteristics of the wake-up segment audio based on the energy of the wake-up segment audio further includes:

[0183] Determine multiple energies after applying beamforming algorithms to the wake-up segment audio at multiple target locations; and

[0184] The second wake-up feature is determined based on the multiple energies of the wake-up segment audio at multiple target locations and the multiple energies after applying a beamforming algorithm to the wake-up segment audio at multiple target locations.

[0185] Example 29. An electronic device according to any one of Examples 25-28, wherein determining the detection features of the detection segment audio based on the detection segment audio and the wake-up features includes:

[0186] Based on the separation matrix for the wake-up segment audio, the energy of the suppressed wearer audio in the detection segment is determined;

[0187] Determine the difference between the energy of the detected audio segment and the energy of the suppressed wearer audio segment; and

[0188] Based on the difference and the first wake-up feature, the first detection feature is determined.

[0189] Example 30. An electronic device according to any one of Examples 25-29, wherein determining the detection features of the detection segment audio based on the detection segment audio and the wake-up features further includes:

[0190] Determine multiple energies after applying a beamforming algorithm to the detected audio segment at multiple target locations;

[0191] Determine multiple energy differences between the detected audio segment at multiple target locations and multiple energies after applying a beamforming algorithm to the detected audio segment at multiple target locations;

[0192] Determine multiple differences between the plurality of energy differences and multiple differences between the second wake-up characteristics; and

[0193] The second detection feature is determined based on the maximum and minimum differences among the plurality of differences.

[0194] Example 31. An electronic device according to any one of Examples 25-30, wherein detecting wearer audio in the detection segment audio based on the detection feature comprises:

[0195] Determine whether the frame corresponding to the first detected feature is a wearer frame or a suppressed wearer frame;

[0196] In response to the presence of multiple consecutive wearer frames satisfying a first condition in a set of multiple frames, the frame with the highest temporal order in the set of multiple frames is determined as the starting point of the wearer audio in the detection segment; and

[0197] In response to the fact that there are consecutive wearer-suppressing frames in the plurality of frame sets that satisfy the second condition, the frame with the last time order in the plurality of frame sets is determined as the end point of the wearer audio of the detection segment.

[0198] Example 32. An electronic device according to any one of Examples 25-31, wherein determining whether the frame corresponding to the first detection feature is a wearer frame or a wearer-suppressing frame includes:

[0199] In response to the first detected feature being greater than or equal to a first threshold, the corresponding frame is determined to be the wearer frame; or

[0200] In response to the first detection feature being less than the first threshold, the corresponding frame is determined to be the inhibited wearer frame.

[0201] Example 33. An electronic device according to any one of Examples 25-32, wherein detecting wearer audio in the detection segment audio based on the detection feature further includes:

[0202] Determine the falling edge and rising edge of the second detection feature;

[0203] In response to the falling edge height satisfying a second threshold, the starting point of the falling edge is determined as the starting point of the wearer's audio in the detection segment; and

[0204] In response to the falling edge height satisfying a third threshold, the tail point of the rising edge is determined as the tail point of the wearer's audio in the detection segment.

[0205] Example 34. An electronic device according to any one of Examples 25-33, wherein detecting wearer audio in the detection segment audio based on the detection feature further includes:

[0206] Based on the first rising edge or the first falling edge of the second detection feature, a fourth threshold is determined by initializing the second detection feature of the wearer audio of the detection segment and the second detection feature of the suppressed wearer audio of the detection segment.

[0207] In response to the second detection feature being less than or equal to the fourth threshold, the second detection feature of the wearer audio in the detection segment is updated by determining that the frame corresponding to the second detection feature is the wearer frame;

[0208] In response to the second detection feature being greater than the fourth threshold, the second detection feature of the suppressed-wearer audio in the detection segment is updated by determining that the frame corresponding to the second detection feature is the suppressed-wearer frame; and

[0209] The fourth threshold is updated based on the second detection feature of the wearer audio in the updated detection segment and the second detection feature of the suppressed wearer audio in the updated detection segment.

[0210] Example 35. An electronic device according to any one of Examples 25-34, wherein detecting wearer audio in the detection segment audio based on the detection feature further includes:

[0211] In response to the occurrence of consecutive wearer frames in the plurality of frame sets simultaneously satisfying both the second and third conditions, the first frame in the plurality of frame sets is determined as the starting point of the wearer audio in the detection segment; and

[0212] In response to the occurrence of consecutive wearer-suppressing frames satisfying a fourth condition in the plurality of frame sets, the last frame in the sorted plurality of frame sets is determined as the end point of the wearer audio in the detection segment; or

[0213] In response to the occurrence of consecutive wearer-suppressing frames satisfying a fifth condition in the plurality of frame sets, the last frame in the plurality of frame sets that satisfies the fifth condition is determined as the tail point of the wearer audio of the detection segment.

[0214] Example 36. The electronic device according to any one of Examples 25-35, wherein the operation further includes:

[0215] The wake-up segment audio and the detection segment audio are obtained based on timestamps; and

[0216] The wake-up segment audio and the detection segment audio are converted from the time domain to the time-frequency domain.

[0217] Example 37. A computer-readable storage medium having stored thereon computer-executable instructions, wherein the computer-executable instructions are executed by a processor to implement the method according to any one of Examples 1 to 12.

[0218] Example 38. A computer program product tangibly stored on a computer-readable medium and comprising computer-executable instructions that, when executed by a device, cause the device to perform the method according to any one of Examples 1 to 12.

[0219] Although this disclosure has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative examples of implementing the claims.

Claims

1. A method for detecting wearer audio, comprising: determining a wake-up feature of a wake-up segment audio based on an energy of the wake-up segment audio; determining a detection feature of a detection segment audio based on the detection segment audio and the wake-up feature; and detecting wearer audio in the detection segment audio based on the detection feature.

2. The method of claim 1, wherein determining a wake-up feature of a wake-up segment audio based on an energy of the wake-up segment audio comprises: determining a first wake-up feature based on the energy of the wake-up segment audio and an energy of a suppress-wearer-audio of the wake-up segment.

3. The method of claim 2, wherein determining a first wake-up feature based on the energy of the wake-up segment audio and an energy of a suppress-wearer-audio of the wake-up segment comprises: determining a separation matrix for the wake-up segment audio; separating the wake-up segment audio into a plurality of audios, the plurality of audios comprising the wearer audio of the wake-up segment and the suppress-wearer-audio of the wake-up segment; and determining the first wake-up feature based on a difference between the energy of the wake-up segment audio and the energy of the suppress-wearer-audio.

4. The method of claim 3, wherein determining a wake-up feature of a wake-up segment audio based on an energy of the wake-up segment audio further comprises: determining a plurality of energies after applying a beamforming algorithm for the wake-up segment audio at a plurality of target orientations; and determining a second wake-up feature based on the plurality of energies of the wake-up segment audio at the plurality of target orientations and the plurality of energies after applying the beamforming algorithm for the wake-up segment audio at the plurality of target orientations.

5. The method of claim 1, wherein determining a detection feature of a detection segment audio based on the detection segment audio and the wake-up feature comprises: determining an energy of a suppress-wearer-audio of a detection segment based on a separation matrix for the wake-up segment audio; determining a difference between the energy of the detection segment audio and the energy of the suppress-wearer-audio of the detection segment; and determining a first detection feature based on the difference and a first wake-up feature.

6. The method of claim 1, wherein determining a detection feature of a detection segment audio based on the detection segment audio and the wake-up feature further comprises: determining a plurality of energies after applying a beamforming algorithm for the detection segment audio at a plurality of target orientations; determining a plurality of energy differences between the plurality of energies of the detection segment audio at the plurality of target orientations and the plurality of energies after applying the beamforming algorithm for the detection segment audio at the plurality of target orientations; determining a plurality of difference values between the plurality of energy differences and a second wake-up feature; and determining a second detection feature based on a maximum difference value and a minimum difference value in the plurality of difference values.

7. The method of claim 6, wherein detecting wearer audio in the detection segment audio based on the detection feature comprises: determining whether a frame corresponding to the first detection feature is a wearer frame or a suppress-wearer-frame; determining a first frame in time in a plurality of frame set as a start point of the wearer audio of the detection segment in response to a plurality of consecutive wearer frames in the plurality of frame set satisfying a first condition; and ​ ​ ​ ​ ​ in response to a plurality of consecutive wearer-suppression frames in the plurality of frame sets satisfying a second condition, determining a frame in the plurality of frame sets that is temporally last in the order to be an end point of the wearer audio of the detection segment.

8. The method of claim 7, wherein determining whether a frame corresponding to a first detection feature is a wearer frame or a wearer-suppression frame comprises: in response to the first detection feature being greater than or equal to a first threshold, determining that the corresponding frame is the wearer frame; or in response to the first detection feature being less than the first threshold, determining that the corresponding frame is the wearer-suppression frame.

9. The method of claim 6, wherein detecting wearer audio in the detection segment audio based on the detection features further comprises: determining a falling edge and a rising edge of the second detection feature; in response to a height of the falling edge satisfying a second threshold, determining a start point of the falling edge to be a start point of the wearer audio of the detection segment; and in response to a height of the falling edge satisfying a third threshold, determining an end point of the rising edge to be an end point of the wearer audio of the detection segment.

10. The method of claim 7, wherein detecting wearer audio in the detection segment audio based on the detection features further comprises: based on a first rising edge or a first falling edge of the second detection feature, determining a fourth threshold by initializing the second detection feature of the wearer audio of the detection segment and the second detection feature of the wearer-suppression audio of the detection segment; in response to the second detection feature being less than or equal to the fourth threshold, updating the second detection feature of the wearer audio of the detection segment by determining that a frame corresponding to the second detection feature is the wearer frame; in response to the second detection feature being greater than the fourth threshold, updating the second detection feature of the wearer-suppression audio of the detection segment by determining that the frame corresponding to the second detection feature is the wearer-suppression frame; and based on the updated second detection feature of the wearer audio of the detection segment and the updated second detection feature of the wearer-suppression audio of the detection segment, updating the fourth threshold.

11. The method of claim 6, wherein detecting wearer audio in the detection segment audio based on the detection features further comprises: in response to a plurality of consecutive wearer frames in the plurality of frame sets simultaneously satisfying a second condition and a third condition, determining a first frame in the plurality of frame sets that is temporally first in the order to be a start point of the wearer audio of the detection segment; and in response to a plurality of consecutive wearer-suppression frames in the plurality of frame sets satisfying a fourth condition, determining a frame in the plurality of frame sets that is temporally last in the order to be an end point of the wearer audio of the detection segment; or in response to a plurality of consecutive wearer-suppression frames in the plurality of frame sets satisfying a fifth condition, determining a frame in the plurality of frame sets that is temporally last in the order that satisfies the fifth condition to be the end point of the wearer audio of the detection segment.

12. The method of claim 1, further comprising: acquiring the wake-up segment audio and the detection segment audio based on the time stamp; and converting the wake-up segment audio and the detection segment audio from time domain to time-frequency domain.

13. An apparatus for detecting wearer audio, comprising: a wake-up feature determination module configured to determine a wake-up feature of the wake-up segment audio based on an energy of the wake-up segment audio; a detection feature determination module configured to determine a detection feature of the detection segment audio based on the detection segment audio and the wake-up feature; and a wearer audio detection module configured to detect wearer audio in the detection segment audio based on the detection feature.

14. An electronic device, comprising: a processor; and a memory coupled with the processor, the memory having stored therein instructions which, when executed by the processor, cause the electronic device to perform the method according to any one of claims 1 to 12.

15. A computer program product comprising computer-executable instructions, wherein the computer-executable instructions are executed by a processor to implement the method according to any one of claims 1 to 12.