Audio detection method and device, electronic device, and storage medium
By dynamically adjusting the sliding window and extracting time-frequency features, combined with an audio detection model, the problem of detection accuracy in noisy and diverse environments for voice wake-up systems was solved, achieving higher accuracy in voice wake-up word detection and task execution.
Patent Information
- Application Number
- CN202510101514.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-11-11
- Estimated Expiration
- 2045-01-22
AI Technical Summary
In situations with high environmental noise, diverse user accents, and limitations in the hardware of the data acquisition equipment, existing technologies struggle to accurately detect voice wake-up words, leading to a decrease in the accuracy of voice wake-up systems.
By dynamically adjusting the sliding window based on the energy information of audio data, the time-frequency features of the audio sliding window segment are extracted, and the voice wake-up word is detected using an audio detection model.
It improves the accuracy of voice wake-up word detection and the execution accuracy of the voice wake-up system, enhancing the flexibility and precision of voice wake-up tasks.
Smart Images

Figure CN119920259B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of computer technology, and in particular to an audio detection method and apparatus, electronic equipment, computer-readable storage medium, and computer program product. Background Technology
[0002] With the development of voice technology, voice wake-up, as one of the core technologies of smart devices, has been widely used in various fields such as smart homes, smart assistants, and wearable devices. In current practical applications, voice recognition technology is usually used to detect the user's voice commands. However, in situations with high environmental noise, diverse user accents, and limitations in the hardware of the acquisition equipment, it is difficult to accurately detect the wake-up word from the user's voice commands, thus reducing the accuracy of the voice wake-up system. Summary of the Invention
[0003] This disclosure provides an audio detection method and apparatus, an electronic device, a computer-readable storage medium, and a computer program product.
[0004] Firstly, this disclosure provides an audio detection method, including:
[0005] Based on the energy information of the collected audio data, the sliding window to be updated of the audio data is adjusted to obtain the adjusted target sliding window. The energy information includes the audio energy values of multiple audio frames in the audio data, and the sliding window to be updated is determined based on the generation order of the existing sliding windows of the audio data.
[0006] Based on the target sliding window and the target step size corresponding to the target sliding window, the audio data is processed by sliding window to obtain an audio sliding window segment;
[0007] Extract the time-frequency features of the audio sliding window segment;
[0008] Detect whether a voice wake-up word exists in the audio sliding window segment based on the time-frequency characteristics;
[0009] If the voice wake-up word is present in the audio sliding window segment, the voice wake-up task corresponding to the audio data is executed.
[0010] Secondly, this disclosure provides an audio detection device, an adjustment module configured to adjust a sliding window to be updated in the audio data based on energy information of the collected audio data, to obtain an adjusted target sliding window, wherein the energy information includes audio energy values of multiple audio frames in the audio data, and the sliding window to be updated is determined based on the generation order of existing sliding windows in the audio data;
[0011] The sliding window module is configured to perform sliding window processing on the audio data according to the target sliding window and the target step size corresponding to the target sliding window, so as to obtain an audio sliding window segment;
[0012] The extraction module is configured to extract the time-frequency features of the audio sliding window segment;
[0013] The detection module is configured to detect whether a voice wake-up word exists in the audio sliding window segment based on the time-frequency features;
[0014] The execution module is configured to execute the voice wake-up task corresponding to the audio data when the voice wake-up word is present in the audio sliding window segment.
[0015] Thirdly, this disclosure provides an electronic device comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores one or more computer programs executable by the at least one processor, the one or more computer programs being executed by the at least one processor to enable the at least one processor to perform the above-described audio detection method.
[0016] Fourthly, this disclosure provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the above-described audio detection method.
[0017] Fifthly, this disclosure provides a computer program product comprising computer-readable code, or a non-volatile computer-readable storage medium carrying computer-readable code, wherein when the computer-readable code is executed in a processor of an electronic device, the processor in the electronic device performs the above-described audio detection method.
[0018] The audio detection method provided in this disclosure adjusts the sliding window to be updated in the audio data based on the energy information of the collected audio data to obtain an adjusted target sliding window. Since the energy information of the audio data is constantly changing in practical applications, adjusting the sliding window to be updated based on the energy information allows for dynamic adjustment of the sliding window, meaning the adjusted target sliding window is constantly changing. Sliding window processing of the audio data according to the adjusted target sliding window and its corresponding target step size allows for more flexible acquisition of different audio features in the audio data, resulting in richer audio features in the audio sliding window segment. This improves the accuracy of time-frequency feature extraction in subsequent processes based on the audio sliding window segment. Detecting voice wake-up words based on highly accurate time-frequency features further improves the accuracy of voice wake-up word detection and the accuracy of waking up the voice wake-up system and executing the corresponding voice wake-up task.
[0019] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0020] The accompanying drawings are provided to further illustrate the present disclosure and form part of the specification. They are used together with the embodiments of the present disclosure to explain the disclosure and do not constitute a limitation thereof. The above and other features and advantages will become more apparent to those skilled in the art from the detailed description of exemplary embodiments with reference to the accompanying drawings, in which:
[0021] Figure 1 This is an application scenario diagram of the audio detection method and apparatus provided in the embodiments of this disclosure;
[0022] Figure 2 A flowchart of an audio detection method provided in this disclosure embodiment;
[0023] Figure 3 This is a schematic diagram illustrating the determination of a target energy value according to an embodiment of the present disclosure;
[0024] Figure 4 This is a schematic diagram illustrating another method for determining a target energy value, as provided in an embodiment of this disclosure.
[0025] Figure 5 This is a schematic diagram of the structure of an audio detection model provided in an embodiment of the present disclosure;
[0026] Figure 6 A block diagram of an audio detection device provided in an embodiment of this disclosure;
[0027] Figure 7 This is a block diagram of an electronic device provided in an embodiment of the present disclosure. Detailed Implementation
[0028] To enable those skilled in the art to better understand the technical solutions of this disclosure, exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments of this disclosure to aid understanding. These should be considered merely exemplary. Therefore, those skilled in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0029] Where there is no conflict, the various embodiments of this disclosure and the features thereof in the embodiments may be combined with each other.
[0030] As used herein, the term “and / or” includes any and all combinations of one or more related enumerated entries.
[0031] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit this disclosure. As used herein, the singular forms “a” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will also be understood that when the terms “comprising” and / or “made of” are used in this specification, the presence of the stated feature, integral, step, operation, element, and / or component is specified, but the presence or addition of one or more other features, integrals, steps, operations, elements, components, and / or groups thereof is not excluded. Words such as “connected” or “linked” are not limited to physical or mechanical connections but can include electrical connections, whether direct or indirect.
[0032] Unless otherwise specified, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by one of ordinary skill in the art. It will also be understood that terms such as those defined in commonly used dictionaries should be interpreted as having a meaning consistent with their meaning in the context of the relevant art and this disclosure, and will not be interpreted as having an idealized or overly formal meaning, unless expressly so defined herein.
[0033] The collection, storage, use, processing, transmission, provision, and disclosure of user personal information in this technical solution comply with relevant laws and regulations and do not violate public order and good morals. The use of user data in this technical solution follows relevant national laws and regulations (e.g., the "Information Security Technology - Personal Information Security Specification"). For example, appropriate measures are taken for personal information access control; restrictions are imposed on the display of personal information; the purpose of using personal information does not exceed the scope of direct or reasonable association; and explicit identity targeting is eliminated when using personal information to avoid precisely identifying specific individuals.
[0034] With the development of voice technology, voice wake-up, as one of the core technologies of smart devices, has been widely used in various fields such as smart homes, smart assistants, and wearable devices. In current practical applications, voice recognition technology is usually used to detect the user's voice commands. However, in situations with high environmental noise, diverse user accents, and limitations in the hardware of the acquisition equipment, it is difficult to accurately detect the wake-up word from the user's voice commands, thus reducing the accuracy of the voice wake-up system.
[0035] Audio signals are continuous and changeable in time, which makes it difficult to fully capture the temporal and frequency domain features of audio data when the audio signal changes rapidly. However, temporal and frequency domain features are important features for detecting the presence of voice wake words in audio data. Therefore, it is also crucial to know how to detect the presence of voice wake words in audio data based on the captured temporal and frequency domain features.
[0036] Based on this, the present disclosure provides an audio detection method, an audio detection device, an electronic device, a computer-readable storage medium, and a computer program product, which will be described in detail in the following embodiments.
[0037] Figure 1 This diagram illustrates an application scenario of the audio detection method and apparatus provided in the embodiments of this disclosure.
[0038] like Figure 1 As shown, the application scenario of this disclosure embodiment may include terminal device 101, network 103, and server 102. Network 103 is used as a medium to provide a communication link between terminal device 101 and server 102. Network 103 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.
[0039] Users can use terminal device 101 to interact with server 102 via network 103 to receive or send messages, etc. Various communication client applications can be installed on terminal device 101, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, etc. (for example only).
[0040] Terminal device 101 can be various electronic devices with a display screen and support web browsing, including but not limited to smartphones, tablets, laptops, and desktop computers.
[0041] Server 102 can be a server that provides various services, such as a backend management server that supports websites browsed by users using terminal device 101 (for example only). The backend management server can analyze and process data such as received user requests, and feed back the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal device.
[0042] It should be noted that the audio detection method and apparatus provided in this disclosure embodiment can be executed by server 102. Accordingly, the audio detection method and apparatus provided in this disclosure embodiment can be set in server 102. The audio detection method and apparatus provided in this disclosure embodiment can also be executed by a server or server cluster that is different from server 102 and is capable of communicating with terminal device 101 and / or server 102. Accordingly, the audio detection method and apparatus provided in this disclosure embodiment can also be set in a server or server cluster that is different from server 102 and is capable of communicating with terminal device 101 and / or server 102.
[0043] It should be understood that Figure 1 The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included.
[0044] Reference Figure 2 , Figure 2 A flowchart of an audio detection method according to an embodiment of this disclosure is shown. The method specifically includes the following steps:
[0045] Step 202: Based on the energy information of the collected audio data, adjust the sliding window to be updated of the audio data to obtain the adjusted target sliding window.
[0046] Because audio signals are continuous and variable in time, they are typically segmented into multiple audio segments for easier analysis and processing. These segments are then analyzed and processed separately. In practice, sliding windows are often used to segment audio data, applying a fixed window length and step size to obtain the corresponding audio segments. However, fixed-length sliding windows often fail to fully capture the audio characteristics at different time scales, making it difficult to adapt to the continuity and variability of audio signals.
[0047] Based on this, embodiments of this disclosure adaptively capture the audio features of audio data by dynamically adjusting the size of the sliding window. To determine the intensity or amplitude changes of the acquired audio data over time, the energy information of the audio data is typically calculated. The audio detection method provided in this disclosure dynamically adjusts the sliding window of the audio data using the energy information, and then performs sliding window processing on the audio data based on the adjusted sliding window to obtain audio sliding window segments of the audio data.
[0048] Among them, the energy information includes the audio energy values of multiple audio frames in the audio data. The audio frame is specifically the audio frame obtained by dividing the audio data into frames during the process of calculating the energy information of the audio data. The sliding window to be updated is determined based on the generation order of the existing sliding windows in the audio data, specifically the last sliding window in the existing sliding windows of the audio data. The target sliding window refers to the sliding window after the sliding window to be updated has been adjusted.
[0049] Specifically, after acquiring audio data, the energy information of the audio data is obtained, and the sliding window to be updated is adjusted based on the energy information to obtain the target sliding window. The specific implementation method of adjusting the sliding window to be updated based on the energy information of the audio data is as follows:
[0050] In one specific embodiment provided in this disclosure, based on the energy information of the collected audio data, the sliding window to be updated of the audio data is adjusted to obtain the adjusted target sliding window, including:
[0051] Determine the end time corresponding to the sliding window to be updated, and obtain the target energy value from multiple audio energy values based on the end time;
[0052] Based on the target energy value and the audio energy value corresponding to the sliding window to be updated, the sliding window to be updated is adjusted to obtain the adjusted target sliding window.
[0053] The end time refers to the moment when the sliding window to be updated ends its sliding motion within the time range of the audio data. The target energy value refers to the audio energy value used to adjust the sliding window to be updated. The target energy value can be the audio energy value of a single audio frame in the audio data, or it can be the average audio energy value across multiple audio frames.
[0054] Specifically, within the time frame of the audio data, the end time corresponding to the sliding window to be updated is determined. Based on the end time of the sliding window to be updated, the target energy value for adjusting the sliding window to be updated is obtained from the energy information of the audio data (i.e., the audio energy values of multiple audio frames). Based on the target energy value and the audio energy value corresponding to the sliding window to be updated, the sliding window to be updated is adjusted to obtain the adjusted target sliding window.
[0055] As mentioned above, the target energy value can be the audio energy value of a single audio frame or the average audio energy value between multiple audio frames. Therefore, the corresponding audio frame can be determined in the audio data based on the end time of the sliding window to be updated, and the target energy value can be obtained based on the audio frame.
[0056] Based on this, in a specific embodiment provided in this disclosure, obtaining the target energy value from multiple audio energy values according to the end time includes:
[0057] The target audio frame is determined based on the said end time;
[0058] Obtain the target energy value of the target audio frame from the plurality of audio energy values.
[0059] Among them, the target audio frame refers to the audio frame acquired from the end time; the target energy value refers to the audio energy value of the target audio frame.
[0060] Specifically, after determining the end time corresponding to the sliding window to be updated, the target audio frame is obtained from the audio data based on the end time, and the audio energy value of the target audio frame is obtained from the audio energy values of multiple audio frames, i.e., the target energy value, so that the sliding window to be updated can be adjusted according to the target energy value in the subsequent process to obtain the target sliding window.
[0061] Furthermore, referring to Figure 3 , Figure 3 A schematic diagram illustrating a method for determining a target energy value according to an embodiment of this disclosure is shown. Figure 3 As shown, the end time of the sliding window to be updated is 0.1 seconds of the audio data. Figure 3 If the audio signal changes in the audio data are not shown, then the target audio frame is determined within the time range of the audio data starting from 0.1 seconds; that is, the first audio frame starting from 0.1 seconds is the target audio frame. From the energy information of the audio data, the audio energy value of the target audio frame is obtained as the target energy value.
[0062] Furthermore, in practical applications, adjusting the sliding window to be updated based on the audio energy frame of each audio frame will increase the computing resources of the terminal device. Therefore, in order to reduce the consumption of computing resources of the terminal device and reduce the number of adjustments to the sliding window to be updated, multiple audio frames within a preset time interval can be obtained through a preset time interval, and the sliding window to be updated can be adjusted according to the audio energy value corresponding to the multiple audio frames to obtain the target sliding window.
[0063] In one specific embodiment provided in this disclosure, obtaining a target energy value from multiple audio energy values based on the end time includes:
[0064] The target time interval is determined based on the end time and the preset time interval;
[0065] Determine the set of target audio frames within the target time interval;
[0066] Obtain multiple audio energy values corresponding to the target audio frame set from the multiple audio energy values;
[0067] The target energy value is determined based on the multiple audio energy values.
[0068] The preset time interval refers to a pre-defined time range used to determine the target audio frame set, such as 0.05 seconds or 0.1 seconds. The preset time interval can be set according to actual application conditions, and this disclosure does not impose any limitations on it. The target audio frame set includes multiple audio frames. The target time interval refers to the time interval within the time range of the audio data that corresponds to the preset time interval. The target time interval is determined based on the end time of the sliding window to be updated and the preset time interval. For example, if the end time of the sliding window to be updated is 0.1 seconds and the preset time interval is 0.05 seconds, then the target time interval can be determined to be 0.1 seconds to 0.15 seconds of the audio data.
[0069] Specifically, after determining the end time of the sliding window to be updated, a target time interval can be determined within the time range of the audio data based on the end time and a preset time interval. Multiple audio frames located within the target time interval are then identified as target audio frames (i.e., the target audio frame set). The audio energy value of each target audio frame in the target audio frame set is obtained from the energy information of the audio data. The average audio energy value among the audio energy values of each audio frame is determined as the target energy value, or the sum of the audio energy values among the audio energy values of each audio frame is determined as the target energy value. The method for determining the target energy value corresponding to the target audio frame set can be determined according to the actual application situation, and this disclosure does not limit it here.
[0070] Reference Figure 4 , Figure 4 A schematic diagram illustrating another method for determining a target energy value according to an embodiment of this disclosure is shown. For example... Figure 4 As shown, the end time of the sliding window to be updated is 0.1 seconds of the audio data. Figure 4 (The audio signal changes in the audio data are not shown in the image). The preset time interval is 0.1 seconds. Based on the end time of the sliding window to be updated being 0.1 seconds and the preset time interval being 0.1 seconds, the target time interval can be determined to be 0.1 seconds to 0.2 seconds of the audio data. Audio frames located between 0.1 seconds and 0.2 seconds of the audio data are identified as target audio frames to obtain a target audio frame set. From the energy information of the audio data, the audio energy value of each target audio frame in the target audio frame set is obtained, and the target energy value is determined based on the audio energy value of each target audio frame.
[0071] This embodiment of the disclosure determines the target audio frame by the end time of the sliding window to be updated, and adjusts the sliding window to be updated according to the target energy value of the target audio frame. This allows for adjustment of the sliding window to be updated for each acquired audio frame, improving the accuracy of dynamic adjustment of the sliding window. Alternatively, a preset time interval can be set to acquire a set of target audio frames within the target time interval. The target energy value is determined according to the audio energy value corresponding to each audio frame in the target audio frame set, reducing the number of times the sliding window to be updated is adjusted and reducing the consumption of computing resources.
[0072] After determining the target energy value, the fluctuation of the audio signal of the audio data can be determined based on the target energy value and the audio energy value corresponding to the sliding window to be updated, so as to determine whether the sliding window to be updated needs to be adjusted.
[0073] In one specific embodiment provided in this disclosure, adjusting the sliding window to be updated based on the target energy value and the audio energy value corresponding to the sliding window to be updated to obtain the adjusted target sliding window includes:
[0074] If the target energy value is greater than the audio energy value corresponding to the sliding window to be updated, the sliding window to be updated is reduced by a preset size to obtain the target sliding window corresponding to the target energy value.
[0075] If the target energy value is less than the audio energy value corresponding to the sliding window to be updated, the sliding window to be updated is increased by the preset size to obtain the target sliding window corresponding to the target energy value;
[0076] If the target energy value is equal to the audio energy value corresponding to the sliding window to be updated, the sliding window to be updated is kept unchanged, and the sliding window to be updated is determined as the target sliding window corresponding to the target energy value.
[0077] The preset size refers to the size of the sliding window to be updated that is set in advance, such as 0.01 seconds, 0.02 seconds, etc.
[0078] Specifically, after determining the target energy value, it is compared with the audio energy value corresponding to the sliding window to be updated. If the target energy value is greater than the audio energy value corresponding to the sliding window to be updated, it indicates that the energy of the audio signal in the audio data has increased. In this case, the sliding window to be updated can be reduced in size by a preset amount to obtain the target sliding window. For example, if the original size of the sliding window to be updated (i.e., the window length of the sliding window to be updated is 0.05 seconds) and the preset size is 0.01 seconds, then the adjusted target sliding window size is 0.04 seconds.
[0079] If the target energy value is less than the audio energy value corresponding to the sliding window to be updated, it means that the energy of the audio signal in the audio data has decreased. In this case, the sliding window to be updated can be enlarged. Specifically, the sliding window to be updated can be enlarged by a preset size to obtain the target sliding window. Taking the original size of the sliding window to be updated (i.e., the window length of the sliding window to be updated is 0.05 seconds) and the preset size of 0.01 seconds as an example, a target sliding window with a window length of 0.06 seconds can be obtained.
[0080] If the target energy value is equal to the audio energy value corresponding to the sliding window to be updated, it means that the energy of the audio signal of the audio data is stable and there is no significant fluctuation. At this time, the sliding window to be updated can be kept unchanged and the sliding window to be updated can be determined as the target sliding window.
[0081] By comparing the target energy value with the audio energy value corresponding to the sliding window to be updated, and adjusting the sliding window to be updated based on the preset size, the target sliding window can dynamically adapt to the energy changes of the audio signal in the audio data, thereby improving the accuracy of segmenting audio sliding window segments in subsequent processes.
[0082] To further improve the accuracy of dynamically adjusting the sliding window to be updated, the size of the sliding window to be updated can be determined based on the difference between the target energy value and the audio energy value corresponding to the sliding window to be updated. Then, the sliding window to be updated is adjusted based on the determined size. The specific implementation method is as follows:
[0083] In one specific embodiment provided in this disclosure, adjusting the sliding window to be updated based on the target energy value and the audio energy value corresponding to the sliding window to be updated to obtain the adjusted target sliding window includes:
[0084] Determine the energy difference between the target energy value and the audio energy value corresponding to the sliding window to be updated;
[0085] The adjustment size is determined based on the energy difference;
[0086] Adjust the sliding window to be updated based on the adjusted size to obtain the adjusted target sliding window.
[0087] Among them, "adjust size" refers to adjusting the size of the sliding window to be updated.
[0088] Specifically, after determining the target energy value, the target energy value is compared with the audio energy value corresponding to the sliding window to be updated, and the energy difference between the target energy value and the audio energy value corresponding to the sliding window to be updated is determined. Based on the energy difference, the adjustment size of the sliding window to be updated needs to be determined. For example, a smaller energy difference results in a smaller adjustment size, and a larger energy difference results in a larger adjustment size. The sliding window to be updated is adjusted according to the determined adjustment size to obtain the adjusted target sliding window. The implementation method for adjusting the sliding window to be updated based on the adjustment size is the same as the implementation method for adjusting the sliding window to be updated based on the preset size described above. The implementation method for adjusting the sliding window to be updated based on the adjustment size can be found in the implementation method for adjusting the sliding window to be updated based on the preset size described above, and will not be repeated here.
[0089] By determining the energy difference between the target energy value and the audio energy value corresponding to the sliding window to be updated, the adjustment size for adjusting the sliding window to be updated is determined. The sliding window to be updated is then adjusted according to the adjustment size, further improving the accuracy of the adjustment of the sliding window to be updated.
[0090] In practical applications, the correspondence between energy information and sliding window size can be preset, so as to directly adjust the sliding window to be updated according to the determined target energy value, thereby simplifying the process of adjusting the sliding window to be updated.
[0091] Based on this, in a specific embodiment provided in this disclosure, the target sliding window of the audio data to be updated is adjusted based on the energy information of the collected audio data to obtain the adjusted target sliding window, including:
[0092] Determine the end time corresponding to the sliding window to be updated, and obtain the target energy value from multiple audio energy values based on the end time;
[0093] Determine the energy value range to which the target energy value belongs;
[0094] Obtain the sliding window size corresponding to the energy value range, adjust the size of the sliding window to be updated to the sliding window size, and obtain the target sliding window corresponding to the target energy value.
[0095] Here, "energy value range" refers to the range of energy values in audio data, divided according to their magnitude. "Sliding window size" refers to the sliding window size corresponding to the energy value range, specifically the adjusted target sliding window size.
[0096] Specifically, using the same method as described above for obtaining the target energy value, the end time corresponding to the sliding window to be updated is determined, and the target energy value is obtained from multiple audio energy values in the energy information based on the end time. After obtaining the target energy value, the energy value range to which the target energy value belongs is determined, and the sliding window size corresponding to the energy value range to which the target energy value belongs is obtained based on the correspondence between the energy value range and the sliding window size. The size of the sliding window to be updated is adjusted to the determined sliding window size to obtain the adjusted target sliding window.
[0097] By pre-dividing the energy information of audio data into multiple energy value ranges and setting the correspondence between energy value ranges and sliding window sizes, after obtaining the target energy value, the size of the target sliding window can be determined by identifying the energy value range to which the target energy value belongs, based on the correspondence between the energy value range and the sliding window size. The size of the sliding window to be updated can then be adjusted to the sliding window size to obtain the target sliding window. This greatly simplifies the process of adjusting the sliding window to be updated and improves processing efficiency.
[0098] The audio detection method provided in this disclosure adjusts the sliding window to be updated based on the energy information of the audio data, thereby obtaining the adjusted target sliding window and realizing dynamic adjustment of the sliding window to be updated.
[0099] Step 204: Perform sliding window processing on the audio data according to the target sliding window and the target step size corresponding to the target sliding window to obtain an audio sliding window segment.
[0100] After adjusting the sliding window to be updated to obtain the target sliding window, the step size of the target sliding window can be determined. Based on the target sliding window and the step size of the target sliding window, the audio data is processed to obtain the audio sliding window segment.
[0101] The target step size is the movement step size corresponding to the target sliding window. The target step size can be set to a length smaller than the target sliding window length, such as half or one-third of the target sliding window length. The audio sliding window segment is the audio segment obtained based on the target sliding window.
[0102] In practical applications, the target step size of the target sliding window is determined by the window length of the target sliding window, and the audio data is processed by sliding window based on the target sliding window and the target step size to obtain the audio sliding window segment corresponding to the target sliding window.
[0103] By setting the target step size to a length smaller than the window length of the target sliding window, data loss can be avoided when performing sliding window processing on audio data.
[0104] Step 206: Extract the time-frequency features of the audio sliding window segment.
[0105] After obtaining the audio sliding window segment, the time-frequency features of the audio sliding window segment are extracted to detect whether a voice wake-up word exists in the audio sliding window segment in subsequent processes. The time-frequency features include time features and frequency features.
[0106] In one specific embodiment provided in this disclosure, extracting the time-frequency features of the audio sliding window segment includes:
[0107] Perform a Fourier transform on the audio sliding window segment to obtain the time characteristics and first frequency characteristics of the audio sliding window segment;
[0108] The first frequency feature is mapped to the Mel scale based on the Mel spectrum to obtain the second frequency feature of the audio sliding window segment;
[0109] The time feature and the second frequency feature are determined as the time-frequency features of the audio sliding window segment.
[0110] Specifically, the Fourier transform can be a short-time Fourier transform (STFT) or a wavelet transform, etc., and in this embodiment, the short-time Fourier transform is preferred. The time characteristic refers to the characteristic distribution of the audio signal in the audio sliding window segment along the time dimension; the first frequency characteristic refers to the characteristic distribution of the audio signal in the audio sliding window segment along the frequency dimension obtained based on the Fourier transform; the second frequency characteristic refers to the characteristic distribution of the audio signal in the audio sliding window segment along the frequency dimension obtained based on the Mel-frequency spectrum. The time-frequency characteristics can be specifically represented by a two-dimensional time-frequency diagram.
[0111] Specifically, performing a Fourier transform (short-time Fourier transform or wavelet transform) on the audio sliding window segment yields the characteristic distribution of the audio signal in the time dimension, i.e., the time characteristics, and the characteristic distribution of the audio signal in the frequency dimension, i.e., the first frequency characteristics. The process of performing a short-time Fourier transform on the audio sliding window segment can be found in Equation 1 below:
[0112]
[0113] Where X(t, f) is the short-time Fourier transform result at time t and frequency f; x(τ) is the input audio signal; w(t-τ) is a window function used to limit the effective range of the short-time Fourier transform; e -j2πfτ It is a complex exponential function used to transform audio signals from the time domain to the frequency domain.
[0114] The short-time Fourier transform of the audio signal in the audio sliding window segment is performed using Formula 1 above. After obtaining the time characteristics and the first frequency characteristics of the audio sliding window segment, the first frequency characteristics are mapped to the Mel scale based on the Mel spectrum to obtain the second frequency characteristics of the audio sliding window segment, as detailed in Formula 2 below:
[0115]
[0116] Among them, f mel Let f be the second frequency feature, i.e., the frequency feature under the Mel scale; and let f be the first frequency feature. Equation 2 can be used to map the first frequency feature after Fourier transform to the Mel scale, thus obtaining the second frequency feature of the audio sliding window segment. The obtained time feature and second frequency feature are the time-frequency features of the audio sliding window segment. Furthermore, a two-dimensional time-frequency diagram of the audio sliding window segment can be generated based on the time-frequency features to intuitively understand the characteristic distribution of the audio signal in the audio sliding window segment.
[0117] In practical applications, to improve the stability of time-frequency characteristics and reduce the influence of noise, the time-frequency characteristics can be normalized to obtain normalized time-frequency characteristics. See Formula 3 below for details:
[0118]
[0119] Among them, X norm X represents the normalized time-frequency feature; X represents the unnormalized time-frequency feature, specifically including time features and second frequency features; μ and σ are the normalization parameters, respectively. Formula 3 above can be used to normalize the time-frequency features of the audio sliding window segment, scaling them to a standard range to obtain the normalized time-frequency features.
[0120] Furthermore, if the environment in which the audio data is collected is noisy, noise reduction processing can be performed on the time-frequency features to reduce noise in the audio sliding window segment. See Formula 4 below for details:
[0121]
[0122] Where Y(f) is the spectrum of the audio sliding window segment after noise reduction; S(f) is the original spectrum of the audio sliding window segment; X(f) is the spectrum of the second frequency feature of the audio sliding window segment; and N(f) is the noise spectrum of the noise in the audio sliding window segment. Noise reduction of the time-frequency features in the audio sliding window segment is achieved through the above formula 4.
[0123] In practical applications, other noise reduction methods can also be used to reduce noise in audio sliding window segments, and are not limited to the noise reduction methods mentioned above.
[0124] The audio detection method provided in this disclosure improves the accuracy of the obtained time-frequency features by combining Fourier transform and Mel spectrum extraction after obtaining an audio sliding window segment.
[0125] Step 208: Detect whether there is a voice wake-up word in the audio sliding window segment based on the time-frequency features.
[0126] After extracting the time-frequency features of the obtained audio sliding window segment, the audio sliding window segment can be detected based on the time-frequency features to determine whether a voice wake-up word exists in the audio sliding window segment. Here, a voice wake-up word refers to a keyword used to activate the voice wake-up system.
[0127] In one specific embodiment provided in this disclosure, detecting whether a voice wake-up word exists in the audio sliding window segment based on the time-frequency features includes:
[0128] The time-frequency features are input into the audio detection model to obtain the probability value of the presence of a voice wake-up word in the audio sliding window segment;
[0129] If the probability value reaches a preset probability value, it is determined that a voice wake-up word exists in the audio sliding window segment;
[0130] If the probability value does not reach the preset probability value, it is determined that there is no voice wake-up word in the audio sliding window segment.
[0131] The audio detection model is used to detect whether a wake-up word exists in an audio sliding window segment. The audio detection model can be a convolutional neural network model, a residual convolutional network model, etc. The probability value represents the probability that a wake-up word exists in the audio sliding window segment. The preset probability value refers to a pre-set probability threshold used to measure whether a wake-up word exists in the audio sliding window segment.
[0132] Specifically, by inputting the obtained time-frequency features into the audio detection model, the probability value output by the audio detection model can be obtained. If the probability value reaches the preset probability value, it is determined that there is a voice wake-up word in the audio sliding window segment. If the probability value does not reach the preset probability value, it is determined that there is no voice wake-up word in the audio sliding window segment.
[0133] To improve the accuracy of identifying whether a voice wake-up word exists in an audio sliding window segment, an audio detection model is used to classify time-frequency features. The time-frequency features are processed based on multi-scale convolutional layers to adapt to the continuity and variability of audio signals over time. Therefore, the audio detection model includes a first convolutional layer, a second convolutional layer, and a fully connected layer. The time-frequency features are processed through convolutional layers of different scales, and the time-frequency features are classified based on the fully connected layer and output probability values to determine whether a voice wake-up word exists in the audio sliding window segment.
[0134] Based on this, in a specific embodiment provided in this disclosure, the time-frequency features are input into an audio detection model to obtain the probability value of the presence of a voice wake-up word in the audio sliding window segment, including:
[0135] The time-frequency features are input into the first convolutional layer to obtain the first convolutional information of the time-frequency features;
[0136] The first convolutional information is input into the second convolutional layer, and the first convolutional information is subjected to dilated convolution processing to obtain the second convolutional information;
[0137] The second convolutional information is input into the fully connected layer to obtain the probability value corresponding to the audio sliding window segment.
[0138] Specifically, the first convolutional layer is a two-dimensional convolutional layer without dilated convolution, such as a 1×1 convolutional layer; the second convolutional layer includes a dilated convolutional layer with dilated convolution, such as a 3×3 convolutional layer with a dilation rate of 2. The first convolutional information refers to the information obtained after convolutional processing of time-frequency features; the second convolutional information refers to the information obtained after dilated convolution processing of the first convolutional information. Specifically, inputting time-frequency features into the first convolutional layer of the audio detection model yields the first convolutional information; inputting the first convolutional information into the second convolutional layer of the audio detection model and performing dilated convolution processing yields the second convolutional information; inputting the second convolutional information into a fully connected layer and processing it yields the probability value corresponding to the audio sliding window segment.
[0139] Furthermore, in order to alleviate the gradient vanishing problem in convolutional networks and avoid information loss in convolutional networks, the audio detection model provided in this embodiment uses residual connections between each convolutional layer to improve the stability of the audio detection model.
[0140] Reference Figure 5 , Figure 5 A schematic diagram of the structure of an audio detection model provided according to an embodiment of this disclosure is shown. Figure 5 As shown, the audio detection model includes an input layer, a first convolutional layer, an average pooling layer, a second convolutional layer, a fully connected layer, and an output layer. The second convolutional layer comprises multiple dilated convolutional layers and a normalization layer. Figure 5(Taking three dilated convolutional layers and a normalization layer as an example). In practical applications, the second convolutional layer can include multiple dilated convolutional layers and normalization layers. The number of kernels and the corresponding dilation rate of multiple dilated convolutional layers can be determined according to the actual application. Multiple dilated convolutional layers can obtain information about time-frequency features at different time scales, which helps improve the accuracy of voice wake-up word recognition. Furthermore, to enable the audio detection model to run on low-power devices with limited computing resources, an average pooling layer can be deployed in the audio detection model to reduce the spatial size of the first convolutional information, thereby reducing the consumption of computing resources and memory usage. Whether an average pooling layer needs to be deployed can be determined based on the computing resources of the device in the actual application.
[0141] Before using an audio detection model to detect wake words based on time-frequency features, the audio detection model needs to be trained first, so that wake word detection can be performed based on the trained audio detection model. Therefore, in a specific embodiment provided in this disclosure, the audio detection model is trained using the following method:
[0142] Acquire sample training data, wherein the sample training data includes sample time-frequency features and sample labels corresponding to the sample time-frequency features;
[0143] The time-frequency features of the sample are input into the audio detection model to obtain the predicted probability value;
[0144] The loss value of the audio detection model is calculated based on the predicted probability value and the sample label;
[0145] The audio detection model is trained based on the loss value.
[0146] Among them, the sample training data refers to the training data used to train the audio detection model, including the sample time-frequency features and the sample labels corresponding to the sample time-frequency features; the sample time-frequency features refer to the time-frequency features used to train the audio detection model; the sample label refers to whether the audio sliding window segment corresponding to the sample time-frequency features contains the actual label of the voice wake-up word; the predicted probability value refers to the probability value obtained by inputting the sample time-frequency features into the audio detection model; and the loss value is used to measure the difference between the predicted probability value of the sample time-frequency features and the actual label.
[0147] Specifically, audio data covering different background noise, accents, and speech rates are collected from multiple environments. Audio data containing a wake-up word is labeled as positive samples, and audio data not containing a wake-up word is labeled as negative samples, generating corresponding sample labels. Based on the same method used to extract the time-frequency features of audio sliding window segments, the time-frequency features of the samples are obtained. The time-frequency features of the samples are input into the audio detection model to obtain the predicted probability value output by the audio detection model. Since the audio detection model is not yet trained, the accuracy of the predicted probability value is low. It is necessary to adjust the model parameters (such as model learning rate, other hyperparameters, etc.) of the audio detection model accordingly. Specifically, the loss value of the audio detection model is calculated based on the predicted probability value and sample labels. The loss function used to calculate the loss value of the audio detection model can be cross-entropy loss, squared loss function, etc. In the embodiments provided in this disclosure, the cross-entropy loss function is preferred as the loss function used to calculate the loss value of the audio detection model. The model parameters of the audio detection model are adjusted according to the calculated loss value. The adjusted model parameters are used to continue training the audio detection model with the next batch of sample training data until the stopping condition of model training is reached.
[0148] To prevent overfitting during training, a validation dataset can be used to periodically evaluate the audio detection model. To ensure the accuracy of the trained model, a test set can be used to evaluate its performance. Specifically, the audio detection model can be evaluated by calculating its accuracy, recall, and F1 score, and the model structure can be adjusted based on the evaluation results, such as increasing or decreasing the depth of convolutional layers or adjusting the residual connection method.
[0149] The audio detection method provided in this disclosure can determine the probability of a voice wake-up word in an audio sliding window segment by training an audio detection model based on the time-frequency features of the audio sliding window segment, and detect whether a voice wake-up word exists in the audio sliding window segment by combining a preset probability value, thereby improving the accuracy of voice wake-up word recognition.
[0150] Step 210: If the voice wake-up word exists in the audio sliding window segment, execute the voice wake-up task corresponding to the audio data.
[0151] After detecting the presence of a wake-up word in an audio sliding window segment, the system can determine whether to execute the corresponding voice wake-up task based on the detection result. The voice wake-up task refers to the task of waking up or starting the voice wake-up system based on audio data. If a wake-up word is detected in the audio sliding window segment, it means that the audio data to which the audio sliding window segment belongs contains a wake-up word, and the corresponding voice wake-up task can be executed.
[0152] By detecting audio sliding window segments, it can be determined whether a voice wake-up word exists in the audio sliding window segment. If it exists, the voice wake-up task corresponding to the audio data can be executed; if it does not exist, the voice wake-up task corresponding to the audio data will not be executed, thereby improving the accuracy of voice wake-up task execution.
[0153] The audio detection method disclosed herein includes: adjusting a sliding window to be updated in the audio data based on energy information of the collected audio data to obtain an adjusted target sliding window, wherein the energy information includes audio energy values of multiple audio frames in the audio data, and the sliding window to be updated is determined based on the generation order of existing sliding windows in the audio data; performing sliding window processing on the audio data according to the target sliding window and a target step size corresponding to the target sliding window to obtain an audio sliding window segment; extracting time-frequency features of the audio sliding window segment; detecting whether a voice wake-up word exists in the audio sliding window segment based on the time-frequency features; and executing a voice wake-up task corresponding to the audio data if the voice wake-up word exists in the audio sliding window segment.
[0154] This embodiment of the disclosure implements the adjustment of the sliding window to be updated in audio data based on the energy information of the collected audio data to obtain an adjusted target sliding window. Adjusting the sliding window to be updated based on energy information enables dynamic adjustment of the sliding window. Sliding window processing of the audio data according to the adjusted target sliding window and its corresponding target step size allows for more flexible acquisition of different audio features in the audio data, resulting in richer audio features in the audio sliding window segment. This improves the accuracy of time-frequency feature extraction in subsequent processes based on the audio sliding window segment. Detecting voice wake-up words based on highly accurate time-frequency features improves the accuracy of voice wake-up word detection. By training an audio detection model based on the time-frequency features of the audio sliding window segment, the probability of a voice wake-up word existing in the audio sliding window segment is determined. Combining this with a preset probability value further improves the accuracy of voice wake-up word recognition, thereby improving the accuracy of waking up the voice wake-up system and executing the corresponding voice wake-up task.
[0155] It is understood that the various method embodiments mentioned above in this disclosure can be combined with each other to form combined embodiments without violating the principle and logic. Due to space limitations, this disclosure will not elaborate further. Those skilled in the art will understand that in the above methods of specific implementation, the specific execution order of each step should be determined by its function and possible internal logic.
[0156] In addition, this disclosure also provides an audio detection device, an electronic device, a computer-readable storage medium, and a computer program product, all of which can be used to implement any of the audio detection methods provided in this disclosure. The corresponding technical solutions and descriptions are described in the corresponding descriptions in the method section and will not be repeated here.
[0157] Figure 6 A block diagram of an audio detection apparatus according to an embodiment of the present disclosure is shown. (Refer to...) Figure 6 This disclosure provides an audio detection device, which includes:
[0158] The adjustment module 602 is configured to adjust the sliding window to be updated in the audio data based on the energy information of the collected audio data to obtain the adjusted target sliding window. The energy information includes the audio energy values of multiple audio frames in the audio data, and the sliding window to be updated is determined based on the generation order of the existing sliding windows in the audio data.
[0159] The sliding window module 604 is configured to perform sliding window processing on the audio data according to the target sliding window and the target step size corresponding to the target sliding window to obtain an audio sliding window segment;
[0160] Extraction module 606 is configured to extract the time-frequency features of the audio sliding window segment;
[0161] The detection module 608 is configured to detect whether a voice wake-up word exists in the audio sliding window segment based on the time-frequency features;
[0162] The execution module 610 is configured to execute the voice wake-up task corresponding to the audio data when the voice wake-up word is present in the audio sliding window segment.
[0163] Optionally, the adjustment module 602 is further configured to:
[0164] Determine the end time corresponding to the sliding window to be updated, and obtain the target energy value from multiple audio energy values based on the end time;
[0165] Based on the target energy value and the audio energy value corresponding to the sliding window to be updated, the sliding window to be updated is adjusted to obtain the adjusted target sliding window.
[0166] Optionally, the adjustment module 602 is further configured to:
[0167] The target audio frame is determined based on the said end time;
[0168] Obtain the target energy value of the target audio frame from the plurality of audio energy values.
[0169] Optionally, the adjustment module 602 is further configured to:
[0170] The target time interval is determined based on the end time and the preset time interval;
[0171] Determine the set of target audio frames within the target time interval;
[0172] Obtain multiple audio energy values corresponding to the target audio frame set from the multiple audio energy values;
[0173] The target energy value is determined based on the multiple audio energy values.
[0174] Optionally, the adjustment module 602 is further configured to:
[0175] If the target energy value is greater than the audio energy value corresponding to the sliding window to be updated, the sliding window to be updated is reduced by a preset size to obtain the target sliding window corresponding to the target energy value.
[0176] If the target energy value is less than the audio energy value corresponding to the sliding window to be updated, the sliding window to be updated is increased by the preset size to obtain the target sliding window corresponding to the target energy value;
[0177] If the target energy value is equal to the audio energy value corresponding to the sliding window to be updated, the sliding window to be updated is kept unchanged, and the sliding window to be updated is determined as the target sliding window corresponding to the target energy value.
[0178] Optionally, the adjustment module 602 is further configured to:
[0179] Determine the energy difference between the target energy value and the audio energy value corresponding to the sliding window to be updated;
[0180] The adjustment size is determined based on the energy difference;
[0181] Adjust the sliding window to be updated based on the adjusted size to obtain the adjusted target sliding window.
[0182] Optionally, the adjustment module 602 is further configured to:
[0183] Determine the end time corresponding to the sliding window to be updated, and obtain the target energy value from multiple audio energy values based on the end time;
[0184] Determine the energy value range to which the target energy value belongs;
[0185] Obtain the sliding window size corresponding to the energy value range, adjust the size of the sliding window to be updated to the sliding window size, and obtain the target sliding window corresponding to the target energy value.
[0186] Optionally, the extraction module 606 is further configured to:
[0187] Perform a Fourier transform on the audio sliding window segment to obtain the time characteristics and first frequency characteristics of the audio sliding window segment;
[0188] The first frequency feature is mapped to the Mel scale based on the Mel spectrum to obtain the second frequency feature of the audio sliding window segment;
[0189] The time feature and the second frequency feature are determined as the time-frequency features of the audio sliding window segment.
[0190] Optionally, the detection module 608 is further configured to:
[0191] The time-frequency features are input into the audio detection model to obtain the probability value of the presence of a voice wake-up word in the audio sliding window segment;
[0192] If the probability value reaches a preset probability value, it is determined that a voice wake-up word exists in the audio sliding window segment;
[0193] If the probability value does not reach the preset probability value, it is determined that there is no voice wake-up word in the audio sliding window segment.
[0194] Optionally, the audio detection model includes a first convolutional layer, a second convolutional layer, and a fully connected layer;
[0195] The detection module 608 is further configured as follows:
[0196] The time-frequency features are input into the first convolutional layer to obtain the first convolutional information of the time-frequency features;
[0197] The first convolutional information is input into the second convolutional layer, and the first convolutional information is subjected to dilated convolution processing to obtain the second convolutional information;
[0198] The second convolutional information is input into the fully connected layer to obtain the probability value corresponding to the audio sliding window segment.
[0199] Optionally, the audio detection device further includes a training module configured to:
[0200] Acquire sample training data, wherein the sample training data includes sample time-frequency features and sample labels corresponding to the sample time-frequency features;
[0201] The time-frequency features of the sample are input into the audio detection model to obtain the predicted probability value;
[0202] The loss value of the audio detection model is calculated based on the predicted probability value and the sample label;
[0203] The audio detection model is trained based on the loss value.
[0204] The audio detection device provided in this disclosure includes: an adjustment module configured to adjust a sliding window to be updated in the audio data based on energy information of the acquired audio data to obtain an adjusted target sliding window, wherein the energy information includes audio energy values of multiple audio frames in the audio data, and the sliding window to be updated is determined based on the generation order of existing sliding windows in the audio data; a sliding window module configured to perform sliding window processing on the audio data according to the target sliding window and a target step size corresponding to the target sliding window to obtain an audio sliding window segment; an extraction module configured to extract time-frequency features of the audio sliding window segment; a detection module configured to detect whether a voice wake-up word exists in the audio sliding window segment based on the time-frequency features; and an execution module configured to execute a voice wake-up task corresponding to the audio data if the voice wake-up word exists in the audio sliding window segment.
[0205] This embodiment of the disclosure implements the adjustment of the sliding window to be updated in audio data based on the energy information of the collected audio data to obtain an adjusted target sliding window. Adjusting the sliding window to be updated based on energy information enables dynamic adjustment of the sliding window. Sliding window processing of the audio data according to the adjusted target sliding window and its corresponding target step size allows for more flexible acquisition of different audio features in the audio data, resulting in richer audio features in the audio sliding window segment. This improves the accuracy of time-frequency feature extraction in subsequent processes based on the audio sliding window segment. Detecting voice wake-up words based on highly accurate time-frequency features improves the accuracy of voice wake-up word detection. By training an audio detection model based on the time-frequency features of the audio sliding window segment, the probability of a voice wake-up word existing in the audio sliding window segment is determined. Combining this with a preset probability value further improves the accuracy of voice wake-up word recognition, thereby improving the accuracy of waking up the voice wake-up system and executing the corresponding voice wake-up task.
[0206] Each module in the aforementioned audio detection device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the corresponding operations of each module.
[0207] Figure 7 This is a block diagram of an electronic device provided in an embodiment of the present disclosure.
[0208] Reference Figure 7This disclosure provides an electronic device, which includes: at least one processor 701; at least one memory 702; and one or more I / O interfaces 703 connected between the processor 701 and the memory 702; wherein the memory 702 stores one or more computer programs that can be executed by the at least one processor 701, and the one or more computer programs are executed by the at least one processor 701 to enable the at least one processor 701 to perform the above-described audio detection method.
[0209] The modules in the aforementioned electronic devices can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in the processor of a computer device in hardware form or independent of it, or stored in the memory of a computer device in software form, so that the processor can call and execute the operations corresponding to each module.
[0210] This disclosure also provides a computer-readable storage medium storing a computer program thereon, wherein the computer program, when executed by a processor, implements the above-described audio detection method. The computer-readable storage medium may be volatile or non-volatile.
[0211] This disclosure also provides a computer program product, including computer-readable code, or a non-volatile computer-readable storage medium carrying computer-readable code, wherein when the computer-readable code is run in the processor of an electronic device, the processor in the electronic device executes the above-described audio detection method.
[0212] Those skilled in the art will understand that all or some of the steps, systems, and apparatuses disclosed above, and their functional modules / units, can be implemented as software, firmware, hardware, or suitable combinations thereof. In hardware implementations, the division between functional modules / units mentioned above does not necessarily correspond to the division of physical components; for example, a physical component may have multiple functions, or a function or step may be performed collaboratively by several physical components. Some or all physical components may be implemented as software executed by a processor, such as a central processing unit, digital signal processor, or microprocessor, or as hardware, or as an integrated circuit, such as an application-specific integrated circuit (ASIC). Such software can be distributed on a computer-readable storage medium, which may include computer storage media (or non-transitory media) and communication media (or transient media).
[0213] As is known to those skilled in the art, the term computer storage medium includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information (such as computer-readable program instructions, data structures, program modules, or other data). Computer storage media includes, but is not limited to, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), static random access memory (SRAM), flash memory or other memory technologies, portable compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical disc storage, magnetic cartridges, magnetic tape, disk storage or other magnetic storage devices, or any other medium that can be used to store desired information and is accessible to a computer. Furthermore, it is known to those skilled in the art that communication media typically contain computer-readable program instructions, data structures, program modules, or other data in modulated data signals such as carrier waves or other transmission mechanisms, and may include any information delivery medium.
[0214] The computer-readable program instructions described herein can be downloaded from computer-readable storage media to various computing / processing devices, or downloaded via a network, such as the Internet, local area network, wide area network, and / or wireless network, to an external computer or external storage device. The network may include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards them to the computer-readable storage media in the respective computing / processing device.
[0215] Computer program instructions used to perform the operations of this disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, status setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk, C++, etc., and conventional procedural programming languages such as the "C" language or similar programming languages. The computer-readable program instructions may execute entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, electronic circuitry, such as programmable logic circuitry, field-programmable gate arrays (FPGAs), or programmable logic arrays (PLAs), is personalized by utilizing the status information of the computer-readable program instructions to implement various aspects of this disclosure.
[0216] The computer program product described herein can be implemented specifically through hardware, software, or a combination thereof. In one alternative embodiment, the computer program product is specifically embodied in a computer storage medium; in another alternative embodiment, the computer program product is specifically embodied in a software product, such as a software development kit (SDK), etc.
[0217] Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0218] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processor of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner; thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0219] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.
[0220] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction containing one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than those shown in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.
[0221] Example embodiments have been disclosed herein, and while specific terminology has been used, it is for illustrative purposes only and should be construed as such, and is not intended to be limiting. In some instances, it will be apparent to those skilled in the art that features, characteristics, and / or elements described in connection with particular embodiments may be used alone, or in combination with features, characteristics, and / or elements described in connection with other embodiments, unless otherwise expressly indicated. Therefore, those skilled in the art will understand that various changes in form and detail may be made without departing from the scope of this disclosure as set forth by the appended claims.
Claims
1. An audio detection method, characterized in that, include: Based on the energy information of the collected audio data, the sliding window to be updated of the audio data is adjusted to obtain the adjusted target sliding window. The energy information includes the audio energy values of multiple audio frames in the audio data, and the sliding window to be updated is determined based on the generation order of the existing sliding windows of the audio data. Based on the target sliding window and the target step size corresponding to the target sliding window, the audio data is processed by sliding window to obtain an audio sliding window segment; Extract the time-frequency features of the audio sliding window segment; Detect whether a voice wake-up word exists in the audio sliding window segment based on the time-frequency characteristics; If the voice wake-up word is present in the audio sliding window segment, the voice wake-up task corresponding to the audio data is executed.
2. The method as described in claim 1, characterized in that, Based on the energy information of the collected audio data, the sliding window to be updated of the audio data is adjusted to obtain the adjusted target sliding window, including: Determine the end time corresponding to the sliding window to be updated, and obtain the target energy value from multiple audio energy values based on the end time; Based on the target energy value and the audio energy value corresponding to the sliding window to be updated, the sliding window to be updated is adjusted to obtain the adjusted target sliding window.
3. The method as described in claim 2, characterized in that, Based on the stated end time, the target energy value is obtained from multiple audio energy values, including: The target audio frame is determined based on the said end time; Obtain the target energy value of the target audio frame from the plurality of audio energy values.
4. The method as described in claim 2, characterized in that, Based on the stated end time, the target energy value is obtained from multiple audio energy values, including: The target time interval is determined based on the end time and the preset time interval; Determine the set of target audio frames within the target time interval; Obtain multiple audio energy values corresponding to the target audio frame set from the multiple audio energy values; The target energy value is determined based on the multiple audio energy values.
5. The method as described in claim 2, characterized in that, Based on the target energy value and the audio energy value corresponding to the sliding window to be updated, the sliding window to be updated is adjusted to obtain the adjusted target sliding window, including: If the target energy value is greater than the audio energy value corresponding to the sliding window to be updated, the sliding window to be updated is reduced by a preset size to obtain the target sliding window corresponding to the target energy value. If the target energy value is less than the audio energy value corresponding to the sliding window to be updated, the sliding window to be updated is increased by the preset size to obtain the target sliding window corresponding to the target energy value; If the target energy value is equal to the audio energy value corresponding to the sliding window to be updated, the sliding window to be updated is kept unchanged, and the sliding window to be updated is determined as the target sliding window corresponding to the target energy value.
6. The method as described in claim 2, characterized in that, Based on the target energy value and the audio energy value corresponding to the sliding window to be updated, the sliding window to be updated is adjusted to obtain the adjusted target sliding window, including: Determine the energy difference between the target energy value and the audio energy value corresponding to the sliding window to be updated; The adjustment size is determined based on the energy difference; Adjust the sliding window to be updated based on the adjusted size to obtain the adjusted target sliding window.
7. The method as described in claim 1, characterized in that, Based on the energy information of the collected audio data, the sliding window to be updated of the audio data is adjusted to obtain the adjusted target sliding window, including: Determine the end time corresponding to the sliding window to be updated, and obtain the target energy value from multiple audio energy values based on the end time; Determine the energy value range to which the target energy value belongs; Obtain the sliding window size corresponding to the energy value range, adjust the size of the sliding window to be updated to the sliding window size, and obtain the target sliding window corresponding to the target energy value.
8. The method as described in claim 1, characterized in that, Extracting the time-frequency features of the audio sliding window segment includes: Perform a Fourier transform on the audio sliding window segment to obtain the time characteristics and first frequency characteristics of the audio sliding window segment; The first frequency feature is mapped to the Mel scale based on the Mel spectrum to obtain the second frequency feature of the audio sliding window segment; The time feature and the second frequency feature are determined as the time-frequency features of the audio sliding window segment.
9. The method as described in claim 1, characterized in that, Detecting whether a voice wake-up word exists in the audio sliding window segment based on the time-frequency features includes: The time-frequency features are input into the audio detection model to obtain the probability value of the presence of a voice wake-up word in the audio sliding window segment; If the probability value reaches a preset probability value, it is determined that a voice wake-up word exists in the audio sliding window segment; If the probability value does not reach the preset probability value, it is determined that there is no voice wake-up word in the audio sliding window segment.
10. The method as described in claim 9, characterized in that, The audio detection model includes a first convolutional layer, a second convolutional layer, and a fully connected layer; The time-frequency features are input into the audio detection model to obtain the probability value of the presence of a voice wake-up word in the audio sliding window segment, including: The time-frequency features are input into the first convolutional layer to obtain the first convolutional information of the time-frequency features; The first convolutional information is input into the second convolutional layer, and the first convolutional information is subjected to dilated convolution processing to obtain the second convolutional information; The second convolutional information is input into the fully connected layer to obtain the probability value corresponding to the audio sliding window segment.
11. The method as described in claim 9, characterized in that, The audio detection model was trained using the following method: Acquire sample training data, wherein the sample training data includes sample time-frequency features and sample labels corresponding to the sample time-frequency features; The time-frequency features of the sample are input into the audio detection model to obtain the predicted probability value; The loss value of the audio detection model is calculated based on the predicted probability value and the sample label; The audio detection model is trained based on the loss value.
12. An audio detection device, characterized in that, include: The adjustment module is configured to adjust the sliding window to be updated in the audio data based on the energy information of the collected audio data, so as to obtain the adjusted target sliding window. The energy information includes the audio energy values of multiple audio frames in the audio data, and the sliding window to be updated is determined based on the generation order of the existing sliding windows in the audio data. The sliding window module is configured to perform sliding window processing on the audio data according to the target sliding window and the target step size corresponding to the target sliding window, so as to obtain an audio sliding window segment; The extraction module is configured to extract the time-frequency features of the audio sliding window segment; The detection module is configured to detect whether a voice wake-up word exists in the audio sliding window segment based on the time-frequency features; The execution module is configured to execute the voice wake-up task corresponding to the audio data when the voice wake-up word is present in the audio sliding window segment.
13. An electronic device, characterized in that, include: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores one or more computer programs that can be executed by the at least one processor, the one or more computer programs being executed by the at least one processor to enable the at least one processor to perform the audio detection method as described in any one of claims 1-11.
14. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the audio detection method as described in any one of claims 1-11.
Citation Information
Patent Citations
Intelligent awakening system based on electroencephalogram signal
CN102671276A
Radar target recognition method based on attention mechanism and bidirectional stacked recurrent neural network
CN111736125A