Audio processing method and device, electronic equipment and readable storage medium

CN117636872BActive Publication Date: 2026-09-25BEIJING DIDI INFINITY TECH & DEV CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202210951620.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-09
Publication Date
2026-09-25
Estimated Expiration
2042-08-09

AI Technical Summary

Benefits of technology

[0018]第五方面,本申请实施例提供了一种计算机程序产品,包括计算机程序/指令,所述计算机程序/指令被处理器执行时实现如第一方面所述的方法。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117636872B_ABST
    Figure CN117636872B_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide an audio processing method and device, electronic equipment and readable storage medium, relating to the technical field of computer. When a user authorizes to open the function of voice waking up a target device, embodiments of the present application can directly determine whether to wake up the target device according to the matching degree between the collected audio and the pre-recorded wake-up audio. In this process, the collected audio does not need to be converted into text, and the collected audio does not need to be compared with the text. Therefore, through the embodiments of the present application, the error generated in the process of converting the audio into the video is avoided, and the accuracy of waking up the target device is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular to an audio processing method, apparatus, electronic device, and readable storage medium. Background Technology

[0002] Currently, more and more devices can be woken up by the user's voice to achieve device intelligence.

[0003] In related technologies, after the device collects audio, it performs speech recognition on the audio and compares the audio with pre-stored reference text. If the recognition result of the audio matches the reference text, the device corresponding to the audio will be woken up to execute subsequent instructions.

[0004] In this process, since the relevant technologies need to convert audio into text and compare it, there will be some errors in the audio and text conversion process, which may lead to problems such as failure to wake up the device or accidental wake-up of the device. Therefore, how to improve the accuracy of waking up the device is an urgent problem to be solved. Summary of the Invention

[0005] In view of this, embodiments of this application provide an audio processing method, apparatus, electronic device, and readable storage medium to improve the accuracy of wake-up devices.

[0006] Firstly, an audio processing method is provided, the method comprising:

[0007] Acquire the captured audio.

[0008] Read a pre-recorded wake-up audio, which includes at least the audio corresponding to the target wake-up word.

[0009] Determine the audio similarity between the acquired audio and the wake-up audio.

[0010] The target device is woken up in response to the audio similarity meeting the wake-up condition.

[0011] Secondly, an audio processing apparatus is provided, the apparatus comprising:

[0012] The audio acquisition module is configured to acquire audio.

[0013] The wake-up audio reading module is configured to read pre-recorded wake-up audio, which includes at least the audio corresponding to the target wake-up word.

[0014] An audio similarity determination module is configured to determine the audio similarity between the acquired audio and the wake-up audio.

[0015] The wake-up module is configured to wake up the target device in response to the audio similarity meeting the wake-up condition.

[0016] Thirdly, embodiments of this application provide an electronic device, including a memory and a processor, wherein the memory is used to store one or more computer program instructions, wherein the one or more computer program instructions are executed by the processor to implement the method as described in the first aspect.

[0017] Fourthly, embodiments of this application provide a computer-readable storage medium storing computer program instructions thereon, which, when executed by a processor, implement the method described in the first aspect.

[0018] Fifthly, embodiments of this application provide a computer program product, including a computer program / instructions, which, when executed by a processor, implement the method described in the first aspect.

[0019] Through the embodiments of this application, the determination of whether to wake up the target device can be made directly based on the degree of matching between the captured audio and the pre-recorded wake-up audio. In this process, there is no need to convert the captured audio into text, nor is there a need to compare the captured audio with text, thus avoiding the errors that occur during the audio-to-video conversion process and improving the accuracy of waking up the target device. Attached Figure Description

[0020] The above and other objects, features and advantages of the present application will become clearer from the following description of embodiments of the present application with reference to the accompanying drawings, in which:

[0021] Figure 1 This is a flowchart illustrating the audio processing method in an embodiment of this application;

[0022] Figure 2 This is a flowchart of the audio processing method in the embodiments of this application;

[0023] Figure 3 This is a flowchart illustrating the segmented processing procedure in an embodiment of this application;

[0024] Figure 4 A flowchart illustrating the process of determining the average feature in an embodiment of this application;

[0025] Figure 5 This is a flowchart of another audio processing method in the embodiments of this application;

[0026] Figure 6 This is a flowchart of another audio processing method in the embodiments of this application;

[0027] Figure 7 This is a flowchart of another audio processing method in the embodiments of this application;

[0028] Figure 8 This is a flowchart of another audio processing method in the embodiments of this application;

[0029] Figure 9 This is a schematic diagram of the audio processing device in the embodiments of this application;

[0030] Figure 10 This is a schematic diagram of the structure of the electronic device in the embodiments of this application. Detailed Implementation

[0031] The present application is described below based on embodiments, but it is not limited to these embodiments. In the detailed description of the present application below, certain specific details are described in detail. Those skilled in the art can fully understand the present application without these details. To avoid obscuring the substance of the present application, well-known methods, processes, flows, elements, and circuits are not described in detail.

[0032] Furthermore, those skilled in the art should understand that the accompanying drawings provided herein are for illustrative purposes only and are not necessarily drawn to scale.

[0033] Unless the context explicitly requires it, words such as "including" or "contains" in the instruction manual should be interpreted as including rather than exclusive or exhaustive; that is, meaning "including but not limited to".

[0034] In the description of this application, it should be understood that the terms "first," "second," etc., are used for descriptive purposes only and should not be construed as indicating or implying relative importance. Furthermore, in the description of this application, unless otherwise stated, "multiple" means two or more. Additionally, in this application, prior authorization from the user is required when obtaining information that needs to be authorized and when activating the voice wake-up function of the target device.

[0035] Currently, in human-computer interaction scenarios, users waking up devices via voice is an essential part. For example, after authorization, users can wake up devices such as speakers, mobile phones, home appliances, and in-vehicle terminals via voice. After being woken up, these devices can perform corresponding operations (such as playing songs, starting navigation services, and opening car windows) according to the user's control (such as touch control, voice control, etc.).

[0036] In related technologies, after the aforementioned device acquires audio, it performs speech recognition on the audio and compares the audio with pre-stored reference text. If the recognition result of the audio matches the reference text, the device corresponding to the audio will be woken up to execute subsequent instructions.

[0037] However, in practical applications, due to the different speaking styles, tones, and timbres of different users, when a user's voice differs from standard Mandarin (for example, when a user wakes up the device using a dialect), there will be a large error in the audio and text conversion process, which may result in problems such as the inability to wake up the device or the device being woken up by mistake. Therefore, how to improve the accuracy of waking up the device is an urgent problem to be solved.

[0038] To address the aforementioned issues, this application provides an audio processing method applicable to electronic devices. The electronic device can be a terminal or a server. The terminal can be a smartphone, tablet, or personal computer (PC), etc. The server can be a single server, a server cluster configured in a distributed manner, or a cloud server.

[0039] Through the audio processing method of this application embodiment, the electronic device can acquire the collected audio, then match the collected audio with the pre-recorded wake-up audio, and determine whether to wake up the target device based on the collected audio and the wake-up audio.

[0040] like Figure 1 As shown, Figure 1 This is a flowchart illustrating the audio processing method in an embodiment of this application. The flowchart includes an audio acquisition 11 and an electronic device 12.

[0041] The audio to be acquired 11 can be the user's voice (i.e., the user's speaking voice) or the audio emitted by a sound-emitting device (such as a microphone). The electronic device 12 can be a terminal or a server, which can acquire the audio to be acquired 11 through its own audio acquisition unit or through an external audio acquisition device.

[0042] After acquiring the audio 11, the electronic device 12 can determine whether to wake up the target device based on the degree of matching (i.e., audio similarity) between the acquired audio 11 and the pre-recorded wake-up audio. If the acquired audio 11 and the wake-up audio match, the target device is woken up; otherwise, the process ends. The target device can be the electronic device 12 itself (i.e., the electronic device 12 can determine whether to wake up its corresponding functional module based on the matching result), or it can be other devices that are wired or wirelessly connected to the electronic device 12.

[0043] Through the embodiments of this application, the determination of whether to wake up the target device can be made directly based on the degree of matching between the captured audio and the pre-recorded wake-up audio. In this process, there is no need to convert the captured audio into text, nor is there a need to compare the captured audio with text, thus avoiding the errors that occur during the audio-to-video conversion process and improving the accuracy of waking up the target device.

[0044] Specifically, such as Figure 2 As shown, the audio processing method of this application embodiment may include the following steps:

[0045] In step 21, the audio is acquired.

[0046] The audio can be acquired by the electronic device through its own audio acquisition unit, or by an external audio acquisition device.

[0047] In step 22, the pre-recorded wake-up audio is read.

[0048] The wake-up audio is a pre-recorded audio track that serves as a reference standard for whether to wake up the target device. When the acquired audio matches the wake-up audio, this embodiment of the application can wake up the target device. In practical applications, users can record one or more wake-up audio tracks into the electronic device using its own built-in audio acquisition unit or an external audio acquisition device. During the process of determining whether to wake up the target device, the electronic device can use one or more of the aforementioned wake-up audio tracks to determine whether the acquired audio matches the wake-up audio.

[0049] It should be noted that when a user records multiple wake-up audios into an electronic device, each wake-up audio must correspond to the same content. In other words, for the same content, a user can repeatedly record multiple wake-up audios into the electronic device.

[0050] The wake-up audio includes at least the audio corresponding to the target wake-up word, which can be a word in Chinese, English, etc. (e.g., opening the car window, opening the car door, turning on the air conditioner, etc.). Furthermore, since this embodiment determines whether the collected audio and the wake-up audio match based on the similarity between audio files, this embodiment does not consider the semantics of the collected audio when waking up the target device. Therefore, the target wake-up word in this embodiment can also be a word without semantic meaning, to expand the types of target wake-up words, such as onomatopoeia, etc.

[0051] In an optional implementation, the present application embodiments may further preprocess the acquired audio and wake-up audio.

[0052] The preprocessing includes at least one of endpoint detection and noise detection. Endpoint detection can be Voice Activity Detection (VAD). In this embodiment, VAD can extract the pronunciation parts from the acquired audio and wake-up audio to improve the efficiency of audio processing. Noise detection in this embodiment can detect and remove noise in the acquired audio and wake-up audio to improve the accuracy of audio similarity.

[0053] In step 23, the audio similarity between the acquired audio and the wake-up audio is determined.

[0054] The audio similarity can be used to characterize the degree of matching between the speaker's voice (i.e., the acquired audio) and the pre-recorded wake-up audio. In this embodiment, the degree of matching may include the degree of matching between the features of the acquired audio and the wake-up audio, or the degree of matching between the object of the acquired audio and the object of the wake-up audio (i.e., the degree of matching between speakers).

[0055] In one optional implementation, audio similarity may include feature similarity. Feature similarity can be used to characterize the feature distance between the acquired audio features and the wake-up audio features. The shorter the feature distance, the higher the feature similarity between the acquired audio and the wake-up audio; conversely, the larger the feature distance, the lower the feature similarity between the acquired audio and the wake-up audio. Furthermore, the features of the acquired audio and the wake-up audio can be determined using a pre-trained neural network model.

[0056] Furthermore, step 23 above can be specifically executed as follows: segmenting the acquired audio according to the predetermined window length and predetermined window shift, determining each test segment corresponding to the acquired audio, and determining the feature similarity between each test segment and the wake-up audio.

[0057] The predetermined window length and predetermined window shift can be set according to actual conditions. The predetermined window length can be a fixed value (e.g., 50ms, 60ms, etc.) or a window length determined based on one or more wake-up audios (e.g., the predetermined window length can be the average length of multiple wake-up audios). The predetermined window shift is generally smaller than the predetermined window length (e.g., if the predetermined window length is 100ms, the predetermined window shift can be 20ms, 30ms, etc.). The smaller the predetermined window shift, the more test segments are determined in this embodiment; the larger the predetermined window shift, the fewer test segments are determined in this embodiment.

[0058] like Figure 3 As shown, Figure 3 This is a flowchart illustrating the segmented processing procedure in an embodiment of this application. The flowchart includes segments to be tested 31, 32, and 33, and a predetermined window shift S. The length of the segment to be tested is the predetermined window length, and the distance moved between adjacent segments is the predetermined window shift (i.e.,...). Figure 3 The predetermined window shift S is shown in the figure.

[0059] Through the embodiments of this application, the acquired audio can be segmented using a predetermined window length and a predetermined window shift to determine test segment 31, test segment 32, and test segment 33. It should be noted that since the length of the acquired audio is determined based on the actual acquisition situation, the last test segment determined in this embodiment (e.g., in...) Figure 3The length of the segment to be tested (33) can be less than the predetermined window length. In another case, embodiments of this application may also add a blank frame at the end of the acquired audio so that the length of the last segment to be tested is equal to the predetermined window length.

[0060] After determining each segment to be tested, the embodiments of this application can determine the feature similarity between each segment to be tested and the wake-up audio, thereby determining whether the collected audio contains the audio corresponding to the target wake-up word.

[0061] Furthermore, the process of determining the feature similarity between each test segment and the wake-up audio in this embodiment can be performed as follows: determining the bottleneck features of each test segment, and determining the feature similarity between each test segment and the wake-up audio based on the feature distance between the bottleneck features of each test segment and the average features corresponding to the wake-up audio.

[0062] In one alternative implementation, the bottleneck features can be determined based on the bottleneck feature layer in a pre-trained speech recognition model.

[0063] Among them, the bottleneck feature layer can reduce or increase the dimensionality of the segment under test, thereby reducing the number of parameters of the segment under test and reducing the amount of computation during data processing, thus reducing the data processing pressure on electronic devices.

[0064] In practical applications, embodiments of this application can first train a speech recognition model, and then use the bottleneck feature layer in the trained speech recognition model to determine the bottleneck features of each test segment. The speech recognition model can be a Time Delay Neural Network Factorized (TDNNF) or other suitable network models.

[0065] The average feature can represent the average value of multiple features. In the embodiments of this application, the average feature can represent the average feature among multiple wake-up audios corresponding to the same content.

[0066] Specifically, in one optional implementation, the process of determining the average feature corresponding to the wake-up audio can be performed as follows: for each target wake-up word, determine at least one wake-up audio corresponding to the target wake-up word, determine the bottleneck feature corresponding to each wake-up audio, and perform feature averaging on the bottleneck feature corresponding to each wake-up audio to determine the average feature corresponding to the wake-up audio.

[0067] For example, such as Figure 4 As shown, Figure 4 This is a flowchart illustrating the process of determining the average feature in an embodiment of this application.

[0068] When determining the average feature 45, this embodiment of the application can first determine each wake-up audio (i.e., wake-up audio 42, wake-up audio 43, and wake-up audio 44) corresponding to the target wake-up word 41, wherein, Figure 4 This is merely an example of an embodiment of this application. In practical applications, the number of wake-up audios is not limited to 3. The number of wake-up audios can be any natural number greater than or equal to 1.

[0069] Furthermore, embodiments of this application can determine the bottleneck features of each wake-up audio (i.e., bottleneck feature 421, bottleneck feature 431 and bottleneck feature 441), and perform feature averaging on each bottleneck feature to determine the average feature 45.

[0070] After determining the bottleneck features of each test segment and the average features corresponding to the wake-up audio, embodiments of this application can determine the feature similarity between each test segment and the wake-up audio based on the feature distance between the bottleneck features of each test segment and the average features corresponding to the wake-up audio.

[0071] For example, embodiments of this application can determine the aforementioned feature similarity using the SLN-DTW (segmental local normalized DTW) algorithm. SLN-DTW is an optimized version of the standard DTW algorithm, adding average distance as a distance metric.

[0072] Specifically, in this embodiment of the application, after preprocessing the acquired audio and the wake-up audio, the acquired audio can be segmented to determine each segment to be tested.

[0073] The frame sequence of the wake-up audio can be represented as q = (q1, q2) / (q3) / (q4) / (q5) / (q6) / (q7) / (q8) / (q9) / (q1) / (q1) / (q2 ... 2,…. q m Each frame sequence of the segment to be tested can be represented as s = (s1, s2) 2,…. s n ), where m represents the number of frames in the corresponding wake-up audio, and n represents the number of frames in the corresponding segment to be tested.

[0074] Therefore, in the embodiments of this application, q = (q1, q2) can be determined. 2,…. q m ) and s=(s1,s 2,…. s n The distance between each frame in the test segment is calculated and a distance matrix is ​​established, thereby determining the distance (i.e., feature similarity) between the test segment and the wake-up audio based on the distance between each frame.

[0075] Specifically, in this embodiment, the cumulative distance a(i,j) and path length l(i,j) in the distance matrix can be initialized. Here, a(i,j) represents the cumulative distance traveled from the starting point (1,e) to (i,j), which can be represented by the normalized distance from the i-th frame of the wake-up audio to the j-th frame of the segment under test. l(i,j) represents the path length traveled from the starting point (1,e) to (i,j), which can be represented by the number of frames between the starting point (1,e) and point (i,j) in the distance matrix. The initialization process of a(i,j) and l(i,j) can be specifically represented by the following formula:

[0076]

[0077] a(1,j)=dist(1,j) norm

[0078] l(i,1)=i

[0079] l(1,j)=1

[0080] Here, dist(k,l) is used to characterize the normalized distance between the k-th frame of the wake-up audio and the l-th frame of the segment to be tested.

[0081] Furthermore, embodiments of this application can perform iterative processing, selecting a point (u, v) from {(i-1, j), (i, j-1), (i-1, j-1)} such that... The calculated result is minimized, thus yielding the following result:

[0082] a(i,j)=a(u,v)+dist(i,j) norm

[0083] l(i,j)=l(u,v)+1

[0084] Furthermore, embodiments of this application can determine a minimum matching path min with an average cumulative distance cost(i,j) = a(i,j)l(i,j) in the distance matrix. j=1,2,…n (cost(m,j)). Here, cost(i,j) represents the average cumulative distance, and the minimum matching path is min. j=1,2,…n The smaller the value of (cost(m,j)), the higher the similarity between the test segment and the wake-up audio, and the greater the feature similarity.

[0085] Through the embodiments of this application, the determination of whether to wake up the target device can be made directly based on the degree of matching between the captured audio and the pre-recorded wake-up audio. In this process, there is no need to convert the captured audio into text, nor is there a need to compare the captured audio with text, thus avoiding the errors that occur during the audio-to-video conversion process and improving the accuracy of waking up the target device.

[0086] In an optional implementation, audio similarity may further include object similarity. Object similarity characterizes the probability that the object corresponding to the captured audio is the same as the object corresponding to the wake-up audio. A higher probability indicates higher object similarity between the captured audio and the wake-up audio; a lower probability indicates lower object similarity. Furthermore, object similarity can be determined using a pre-trained probabilistic model or a neural network model.

[0087] In other words, object similarity is used to characterize the degree of matching between the voice object corresponding to the collected audio and the voice object corresponding to the wake-up audio. The greater the object similarity, the greater the probability that the collected audio and the wake-up audio are from the same voice object.

[0088] Furthermore, in this embodiment of the application, the process of determining the audio similarity between the acquired audio and the wake-up audio can be performed as follows: inputting the acquired audio into a pre-set audio object recognition model, and determining the object similarity output by the audio object recognition model.

[0089] In practical applications, since most of the collected audio and wake-up audio are the user's voice, the audio object recognition model in this application embodiment can be a speaker recognition model, such as a voiceprint matching model.

[0090] In this embodiment, the similarity between the captured audio and the wake-up audio content can be determined by feature similarity, and the similarity between the captured audio and the wake-up audio object can be determined by an audio object recognition model. Therefore, this embodiment can wake up the target device only when the content and object of the captured audio and the wake-up audio match, improving the targeting of the device wake-up process.

[0091] In step 24, the target device is woken up in response to the audio similarity meeting the wake-up condition.

[0092] Combination Figure 1 The content shown can be targeted at a device that is... Figure 1 The electronic device 12 in the text can also be other devices associated with the electronic device 12.

[0093] Taking the target device as the window control unit in a vehicle as an example, the electronic device 12 can be an in-vehicle terminal installed in the vehicle. The in-vehicle terminal can acquire the voice of the user in the vehicle (i.e., collect audio) through its own audio acquisition unit or an external audio acquisition device.

[0094] Users can pre-record one or more "open the window" wake-up audio messages to the in-vehicle terminal via an audio acquisition unit or device. When the user is driving or riding in the vehicle, the in-vehicle terminal can acquire the acquired audio and match it with the wake-up audio ("open the window"). If the audio similarity between the acquired audio and the wake-up audio meets the wake-up condition, indicating that the acquired audio is also "open the window," the in-vehicle terminal can wake up the window control unit and send a control command to the window control unit to open the vehicle's windows.

[0095] It should be noted that in the example above, since the wake-up audio is "open the window," the vehicle terminal can directly send a control command to the window control unit after waking it up, causing the vehicle's windows to open. In another scenario, if the wake-up audio does not include semantics related to controlling a device (e.g., the wake-up audio is "hello" or a verbal phrase), then after waking the target device, the target device can switch to an awakened state, standby state, etc. Furthermore, the electronic device can then control the target device based on subsequent audio input.

[0096] Therefore, through the embodiments of this application, the determination of whether to wake up the target device can be made directly based on the degree of matching between the captured audio and the pre-recorded wake-up audio. In this process, there is no need to convert the captured audio to text, nor to compare the captured audio with text, thus avoiding errors generated during the audio-to-video conversion process and improving the accuracy of waking up the target device.

[0097] In an optional implementation, step 24 above can be specifically executed as follows: in response to the feature similarity being greater than or equal to a first similarity threshold and the object similarity being greater than or equal to a second similarity threshold, the target device is woken up.

[0098] The first similarity threshold and the second similarity threshold can be set according to the actual situation. For example, the first similarity threshold and the second similarity threshold can be 90%, 93%, 95%, etc.

[0099] In other words, under this condition, the embodiments of this application determine both feature similarity and object similarity. The wake-up condition at this time is that the content of the collected audio and the wake-up audio match, and the objects of the collected audio and the wake-up audio also match.

[0100] For example, such as Figure 5 As shown, Figure 5This is a flowchart of an audio processing method according to an embodiment of this application, which specifically includes the following steps:

[0101] In step 51, the audio is acquired.

[0102] The audio can be acquired by the electronic device through its own audio acquisition unit, or by an external audio acquisition device.

[0103] In step 52, endpoint detection and noise detection are performed on the acquired audio.

[0104] In step 53, the feature similarity and object similarity between the acquired audio and the wake-up audio are determined.

[0105] In this application embodiment, endpoint detection and noise detection can be performed on the wake-up audio in advance, or during the wake-up process of the target device. If the wake-up audio is recorded during the wake-up audio recording process, the computational load when waking up the target device can be reduced.

[0106] In step 54, it is determined whether the feature similarity is greater than or equal to the first similarity threshold and the object similarity is greater than or equal to the second similarity threshold. If the feature similarity is greater than or equal to the first similarity threshold and the object similarity is greater than or equal to the second similarity threshold, then proceed to step 55; otherwise, the process ends.

[0107] In step 55, the target device is woken up.

[0108] Taking the target device as the window control unit in a vehicle as an example, the electronic device can be an in-vehicle terminal installed in the vehicle. The in-vehicle terminal can acquire the voice of the user in the vehicle (i.e., collect audio) through its own audio acquisition unit or an external audio acquisition device.

[0109] User A pre-records one or more "open the window" wake-up audio messages to the in-vehicle terminal via an audio acquisition unit or device. When User A is driving or riding in the vehicle, the in-vehicle terminal can acquire the acquired audio and match it with the wake-up audio ("open the window"). If the feature similarity between the acquired audio and the wake-up audio is greater than or equal to a first similarity threshold, and the object similarity is greater than or equal to a second similarity threshold, indicating that the acquired audio also indicates "open the window" and the object corresponding to the acquired audio is User A, then the in-vehicle terminal can wake up the window control unit and send a control command to the window control unit to open the vehicle's windows.

[0110] In other words, through the embodiments of this application, the vehicle terminal will only wake up the window control unit when user A says "open the window", avoiding interference from other users in the process of waking up the target device. This not only improves the accuracy of waking up the target device, but also improves driving safety.

[0111] Therefore, in this embodiment, the similarity between the captured audio and the wake-up audio content can be determined by feature similarity, and the similarity between the captured audio and the wake-up audio object can be determined by an audio object recognition model. Thus, this embodiment can wake up the target device only when the content and object of the captured audio and the wake-up audio match, improving the targeting of the device wake-up process.

[0112] In another alternative implementation, step 24 above can be further executed as follows: in response to a feature similarity greater than or equal to a first similarity threshold, wake up the target device.

[0113] In this case, the embodiments of this application can determine whether to wake up the target device based solely on feature similarity, thereby improving the speed of waking up the target device.

[0114] For example, such as Figure 6 As shown, Figure 6 This is a flowchart of another audio processing method according to an embodiment of this application, which specifically includes the following steps:

[0115] In step 61, the audio is acquired.

[0116] The audio can be acquired by the electronic device through its own audio acquisition unit, or by an external audio acquisition device.

[0117] In step 62, endpoint detection and noise detection are performed on the acquired audio.

[0118] In step 63, the feature similarity between the acquired audio and the wake-up audio is determined.

[0119] In this application embodiment, endpoint detection and noise detection can be performed on the wake-up audio in advance, or during the wake-up process of the target device. If the wake-up audio is recorded during the wake-up audio recording process, the computational load when waking up the target device can be reduced.

[0120] In step 64, determine whether the feature similarity is greater than or equal to the first similarity threshold. If the feature similarity is greater than or equal to the first similarity threshold, proceed to step 55; otherwise, end the process.

[0121] In step 65, the target device is woken up.

[0122] This application provides an embodiment that allows for direct determination of whether to wake up a target device based on the matching degree between the captured audio and a pre-recorded wake-up audio. This process eliminates the need to convert the captured audio to text or compare it with text, avoiding errors introduced during audio-to-video conversion and improving the accuracy of waking up the target device. Furthermore, since this application only determines whether to wake up the target device based on feature similarity, it also improves the wake-up speed of the target device in addition to increasing accuracy.

[0123] In one alternative implementation, such as Figure 7 As shown, if the audio length of the acquired audio is too long, step 24 above can also be performed as follows:

[0124] In step 71, in response to the audio length of the acquired audio being greater than a length threshold, test segments whose corresponding feature similarity is less than a first similarity threshold are deleted, and at least one candidate segment is determined.

[0125] In practical applications, if the audio length of the collected audio is too long, the problem of repeated wake-up may occur. For example, if the same target wake-up word appears 3 times in the collected audio, the target device may be woken up 3 times.

[0126] Therefore, to solve the aforementioned problem of repeated wake-up, this application embodiment can set a length threshold and process the acquired audio with a length greater than the length threshold as a whole. The length threshold can be set according to actual conditions; for example, the length threshold can be 300ms, 400ms, 450.6ms, etc.

[0127] Furthermore, since the same target wake-up word may appear multiple times in the collected audio, this embodiment of the application can retain the test segment with a feature similarity greater than or equal to the first similarity threshold as a candidate segment after determining the feature similarity, that is, retain the test segment that may contain the target wake-up word as a candidate segment.

[0128] In step 72, each candidate segment is filtered based on the feature similarity corresponding to each candidate segment and the overlap (Intersection over Union, IoU) between each candidate segment to determine at least one target segment.

[0129] In this embodiment, the overlap between candidate segments can be calculated using a non-maximum suppression (NMS) algorithm. When the overlap is too large (e.g., the overlap is greater than or equal to a preset overlap threshold), it indicates that the two candidate segments corresponding to this overlap are candidate segments corresponding to the target keyword that appears in the same instance. Furthermore, this embodiment can retain candidate segments with smaller overlap as target segments.

[0130] For example, in this embodiment of the application, test segments with a corresponding feature similarity less than a first similarity threshold are deleted, and five candidate segments (candidate segments A, B, C, D, and E) are determined. Among them, the feature similarity corresponding to candidate segment A is 0.78, the feature similarity corresponding to candidate segment B is 0.98, the feature similarity corresponding to candidate segment C is 0.83, the feature similarity corresponding to candidate segment D is 0.68, and the feature similarity corresponding to candidate segment E is 0.81.

[0131] In this embodiment, the candidate segment with the highest feature similarity among the five candidate segments (i.e., candidate segment B) can be used as the target segment, and the overlap between candidate segment B and the other four candidate segments can be determined. If the overlap is greater than or equal to a preset overlap threshold, the corresponding candidate segment is deleted; if the overlap is less than the preset overlap threshold, the corresponding candidate segment is retained.

[0132] If the overlap between candidate segment B and candidate segments A and C is greater than or equal to the preset overlap threshold, then candidate segments A and C are deleted, while candidate segments D and E are retained.

[0133] Furthermore, embodiments of this application can repeat the above steps to determine all target segments. For example, embodiments of this application can use the candidate segment with the highest corresponding feature similarity among the remaining candidate segments (i.e., candidate segment E) as the target segment, and determine the overlap between candidate segment E and the other 1 candidate segment. If the overlap is greater than or equal to a preset overlap threshold, the corresponding candidate segment is deleted; if the overlap is less than the preset overlap threshold, the corresponding candidate segment is retained.

[0134] If the overlap between candidate segment E and candidate segment D is greater than or equal to a preset overlap threshold, then candidate segment D is deleted, and candidate segment E is retained. At this time, all candidate segments are either deleted or retained (candidate segments A, C, and D are deleted, and candidate segments B and E are retained).

[0135] In step 73, in response to the number of target fragments being greater than or equal to the number threshold, the target device is woken up.

[0136] The quantity threshold can be a natural number greater than or equal to 1. If the quantity threshold is equal to 1, it means that the target device is activated when the target wake-up word appears in the collected audio. If the quantity threshold is greater than 1, it means that the target device is activated only when the target wake-up word appears multiple times in the collected audio. This avoids the target device being woken up multiple times in a short period of time.

[0137] In one alternative implementation, such as Figure 8 As shown, if the target wake-up word is a reduplicated word, the process of determining the feature similarity between each test segment and the wake-up audio in this embodiment of the application may include the following steps:

[0138] In step 81, the average similarity between the repeated segments in the audio corresponding to the reduplicated word is determined.

[0139] Here, reduplicated words are words that include consecutive repetitions, such as "ABAB". In this embodiment, since the wake-up audio is a pre-recorded and accurate audio, this embodiment can determine the average similarity between the repetitive segments in the audio corresponding to the reduplicated word, and use this average similarity as a criterion for determining whether to wake up the target device.

[0140] It should be noted that when the number of repeated segments in the audio corresponding to a reduplicated word is 2, the feature similarity between these 2 repeated segments is the average similarity. When the number of repeated segments in the audio corresponding to a reduplicated word is greater than 2, the embodiments of this application can determine the feature similarity between each repeated segment, and then determine the average similarity of each feature similarity.

[0141] In step 82, for each segment to be tested, the feature similarity between each repeated segment and the corresponding part in the segment to be tested is determined.

[0142] To improve the accuracy of waking up the target device, embodiments of this application can determine the similarity of each repeated segment in the audio corresponding to the reduplicated word individually.

[0143] For example, if the wake-up audio "ABAB" is represented as (t begin ,t end If the first half of the wake-up audio is then characterized as... The latter half can be characterized as Therefore, the embodiments of this application can be determined and The feature similarity between them is used as the average similarity mentioned above.

[0144] Furthermore, for each segment to be tested, embodiments of this application can determine the feature similarity between the first half of the segment to be tested and the first half of the wake-up audio, and the feature similarity between the second half of the segment to be tested and the second half of the wake-up audio.

[0145] Furthermore, in an optional implementation, step 24 above can be performed as follows: in response to the fact that the similarity of each feature corresponding to the segment to be tested is greater than or equal to the average similarity, the target device is woken up.

[0146] In this embodiment, each repeated segment is compared individually. Therefore, the target device will only be woken up when each repeated segment matches the wake-up audio, thereby improving the accuracy of waking up the target device and avoiding false wake-ups.

[0147] Based on the same technical concept, embodiments of this application also provide an audio processing device, such as... Figure 9 As shown, the device includes: an audio acquisition module 91, a wake-up audio reading module 92, an audio similarity determination module 93, and a wake-up module 94.

[0148] The audio acquisition module 91 is configured to acquire audio.

[0149] The wake-up audio reading module 92 is configured to read pre-recorded wake-up audio, which includes at least the audio corresponding to the target wake-up word.

[0150] The audio similarity determination module 93 is configured to determine the audio similarity between the acquired audio and the wake-up audio.

[0151] The wake-up module 94 is configured to wake up the target device in response to the audio similarity meeting the wake-up condition.

[0152] In some embodiments, the audio similarity includes feature similarity.

[0153] The audio similarity determination module 93 is specifically configured as follows:

[0154] The acquired audio is segmented according to a predetermined window length and a predetermined window shift to determine each segment to be tested corresponding to the acquired audio.

[0155] Determine the feature similarity between each of the test segments and the wake-up audio.

[0156] In some embodiments, the audio similarity also includes object similarity;

[0157] The audio similarity determination module 93 is specifically configured as follows:

[0158] The collected audio is input into a pre-set audio object recognition model, and the object similarity output by the audio object recognition model is determined. The object similarity is used to characterize the probability that the object corresponding to the collected audio is the same as the object corresponding to the wake-up audio.

[0159] In some embodiments, the wake-up module 94 is specifically configured as follows:

[0160] In response to the feature similarity being greater than or equal to a first similarity threshold and the object similarity being greater than or equal to a second similarity threshold, the target device is woken up.

[0161] In some embodiments, the audio similarity determination module 93 is specifically configured to:

[0162] Determine the bottleneck characteristics of each of the tested segments.

[0163] The feature similarity between each test segment and the wake-up audio is determined based on the feature distance between the bottleneck features of each test segment and the average features corresponding to the wake-up audio.

[0164] In some embodiments, the average characteristics corresponding to the wake-up audio are determined based on at least the following modules:

[0165] The wake-up audio determination module is configured to determine at least one wake-up audio corresponding to each target wake-up word.

[0166] The bottleneck feature determination module is configured to determine the bottleneck features corresponding to each of the wake-up audios.

[0167] The average feature determination module is configured to perform feature averaging on the bottleneck features corresponding to each of the wake-up audios in order to determine the average features corresponding to the wake-up audios.

[0168] In some embodiments, the wake-up module 94 is specifically configured as follows:

[0169] In response to the feature similarity being greater than or equal to a first similarity threshold, the target device is woken up.

[0170] In some embodiments, the wake-up module 94 is specifically configured as follows:

[0171] In response to the audio length of the acquired audio being greater than a length threshold, segments of the test that have a corresponding feature similarity less than a first similarity threshold are deleted, and at least one candidate segment is determined.

[0172] Based on the feature similarity corresponding to each candidate segment and the overlap between each candidate segment, the candidate segments are filtered to determine at least one target segment.

[0173] The target device is woken up in response to the number of target fragments being greater than or equal to a number threshold.

[0174] In some embodiments, in response to the target wake word being a reduplicated word, the audio similarity determination module 93 is specifically configured to:

[0175] Determine the average similarity between the repeated segments in the audio corresponding to the reduplicated word.

[0176] For each segment to be tested, the feature similarity between each repeated segment and the corresponding part of the segment to be tested is determined.

[0177] In some embodiments, the wake-up module 94 is specifically configured as follows:

[0178] In response to the fact that the similarity of each feature corresponding to the segment to be tested is greater than or equal to the average similarity, the target device is woken up.

[0179] In some embodiments, the apparatus further includes:

[0180] A preprocessing module is configured to preprocess the acquired audio and the wake-up audio, the preprocessing including at least one of endpoint detection and noise detection.

[0181] In some embodiments, the bottleneck features are determined based on a bottleneck feature layer in a pre-trained speech recognition model.

[0182] Through the embodiments of this application, the determination of whether to wake up the target device can be made directly based on the degree of matching between the captured audio and the pre-recorded wake-up audio. In this process, there is no need to convert the captured audio into text, nor is there a need to compare the captured audio with text, thus avoiding the errors that occur during the audio-to-video conversion process and improving the accuracy of waking up the target device.

[0183] Figure 10 This is a schematic diagram of an electronic device according to an embodiment of this application. For example... Figure 10 As shown, Figure 10The illustrated electronic device is a general address lookup device, comprising a general computer hardware architecture, including at least a processor 101 and a memory 102. The processor 101 and memory 102 are connected via a bus 103. The memory 102 is adapted to store instructions or programs executable by the processor 101. The processor 101 can be a standalone microprocessor or a collection of one or more microprocessors. Thus, the processor 101 executes the instructions stored in the memory 102, thereby performing the method flow described in the embodiments of this application to process data and control other devices. The bus 103 connects the aforementioned components together, and also connects these components to a display controller 104, a display device, and an input / output (I / O) device 105. The input / output (I / O) device 105 can be a mouse, keyboard, modem, network interface, touch input device, motion-sensing input device, printer, and other devices known in the art. Typically, the input / output device 105 is connected to the system via an input / output (I / O) controller 106.

[0184] Those skilled in the art will understand that embodiments of this application can be provided as methods, apparatus (devices), or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-readable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0185] This application is described with reference to flowchart illustrations of methods, apparatus (devices), and computer program products according to embodiments of this application. It should be understood that each step in the flowchart can be implemented by computer program instructions.

[0186] These computer program instructions may be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including an instruction means, the implementation process of which is described in the instruction means. Figure 1 The function specified in one or more processes.

[0187] These computer program instructions may also be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing device to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing device, produce instructions for implementing processes. Figure 1 A device for a function specified in one or more processes.

[0188] Another embodiment of this application relates to a non-volatile storage medium for storing a computer-readable program for use by a computer to execute some or all of the above-described method embodiments.

[0189] That is, those skilled in the art will understand that all or part of the steps in the methods of the above embodiments can be implemented by a program specifying the relevant hardware. This program is stored in a storage medium and includes several instructions to cause a device (which may be a microcontroller, chip, etc.) or processor to execute all or part of the steps of the methods described in the embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, a portable hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0190] Another embodiment of this application relates to a computer program product, including a computer program / instructions that, when executed by a processor, can implement some or all of the above-described method embodiments.

[0191] That is, those skilled in the art will understand that the embodiments of this application can specify related hardware (including the processor itself) by having the processor execute a computer program product (computer program / instructions) to implement all or part of the steps in the methods of the above embodiments.

[0192] The above description is merely a preferred embodiment of this application and is not intended to limit this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.

Claims

1. An audio processing method, characterized in that, The method includes: Acquire the captured audio; Read a pre-recorded wake-up audio, wherein the wake-up audio includes at least the audio corresponding to the target wake-up word; Determine the audio similarity between the acquired audio and the wake-up audio; and In response to the audio similarity meeting the wake-up condition, the target device is woken up; The audio similarity includes feature similarity; Determining the audio similarity between the acquired audio and the wake-up audio includes: The acquired audio is segmented according to a predetermined window length and a predetermined window shift to determine the corresponding test segments; and Determine the feature similarity between each of the test segments and the wake-up audio; or, in response to the target wake-up word being a reduplicated word, determine the average similarity between each repeated segment in the audio corresponding to the reduplicated word; and for each of the test segments, determine the feature similarity between each repeated segment and the corresponding part of the test segment. The step of waking up the target device in response to the audio similarity meeting the wake-up condition includes: In response to the number of target segments being greater than or equal to a number threshold among the tested segments, the target device is woken up; or, in response to the similarity of each feature corresponding to the tested segment being greater than or equal to the average similarity, the target device is woken up; the target segment is the tested segment among the tested segments whose feature similarity with the wake-up audio is greater than a first similarity threshold, and whose overlap with other tested segments is less than a preset overlap threshold.

2. The method according to claim 1, characterized in that, The audio similarity also includes object similarity; Determining the audio similarity between the acquired audio and the wake-up audio includes: The collected audio is input into a pre-set audio object recognition model, and the object similarity output by the audio object recognition model is determined. The object similarity is used to characterize the probability that the object corresponding to the collected audio is the same as the object corresponding to the wake-up audio.

3. The method according to claim 2, characterized in that, The step of waking up the target device in response to the audio similarity meeting the wake-up condition includes: In response to the feature similarity being greater than or equal to a first similarity threshold and the object similarity being greater than or equal to a second similarity threshold, the target device is woken up.

4. The method according to claim 1, characterized in that, Determining the feature similarity between each of the test segments and the wake-up audio includes: Determine the bottleneck characteristics of each of the tested segments; and The feature similarity between each test segment and the wake-up audio is determined based on the feature distance between the bottleneck features of each test segment and the average features corresponding to the wake-up audio.

5. The method according to claim 4, characterized in that, The average features corresponding to the wake-up audio are determined based on at least the following steps: For each target wake word, determine at least one wake-up audio corresponding to the target wake word; Determine the bottleneck features corresponding to each of the aforementioned wake-up audios; and The bottleneck features corresponding to each of the wake-up audios are subjected to feature averaging to determine the average features corresponding to the wake-up audios.

6. The method according to claim 1, characterized in that, The step of waking up the target device in response to the audio similarity meeting the wake-up condition includes: In response to the feature similarity being greater than or equal to a first similarity threshold, the target device is woken up.

7. The method according to claim 1, characterized in that, The step of waking up the target device in response to the audio similarity meeting the wake-up condition includes: In response to the audio length of the collected audio being greater than a length threshold, segments of the test that have a corresponding feature similarity less than a first similarity threshold are deleted, and at least one candidate segment is determined. Based on the feature similarity corresponding to each candidate segment and the overlap between each candidate segment, the candidate segments are filtered to determine at least one target segment; and The target device is woken up in response to the number of target fragments being greater than or equal to a number threshold.

8. The method according to claim 1, characterized in that, The method further includes: The acquired audio and the wake-up audio are preprocessed, and the preprocessing includes at least one of endpoint detection and noise detection.

9. The method according to claim 4 or 5, characterized in that, The bottleneck features are determined based on the bottleneck feature layer in a pre-trained speech recognition model.

10. An audio processing apparatus, characterized in that, The device includes: The audio acquisition module is configured to acquire audio. The wake-up audio reading module is configured to read pre-recorded wake-up audio, which includes at least the audio corresponding to the target wake-up word; An audio similarity determination module is configured to determine the audio similarity between the acquired audio and the wake-up audio; and The wake-up module is configured to wake up the target device in response to the audio similarity meeting the wake-up condition; The audio similarity includes feature similarity; the audio similarity determination module is further configured to segment the acquired audio according to a predetermined window length and a predetermined window shift to determine each test segment corresponding to the acquired audio; and to determine the feature similarity between each test segment and the wake-up audio; or, in response to the target wake-up word being a reduplicated word, to determine the average similarity between each repeated segment in the audio corresponding to the reduplicated word; and, for each test segment, to determine the feature similarity between each repeated segment and the corresponding part of the test segment. The wake-up module is further configured to wake up the target device in response to the number of target segments being greater than or equal to a number threshold among the test segments; or, in response to the similarity of each feature corresponding to the test segment being greater than or equal to the average similarity; the target segment is the test segment among the test segments whose feature similarity with the wake-up audio is greater than a first similarity threshold, and whose overlap with other test segments is less than a preset overlap threshold.

11. An electronic device comprising a memory and a processor, characterized in that, The memory is used to store one or more computer program instructions, wherein the one or more computer program instructions are executed by the processor to implement the method as described in any one of claims 1-9.

12. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the method of any one of claims 1-9.

13. A computer program product comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, they implement the method of any one of claims 1-9.

Citation Information

Patent Citations

  • Target language detection method and device

    CN110491375A

  • Voiceprint recognition method, singer authentication method, electronic equipment and storage medium

    CN113366567A

  • Method for voice wake-up, electronic equipment, storage medium, and program

    CN114038457A