Method, device and electronic device for identifying nearby wake-up device

By performing targeted noise reduction processing on user wake-up voice, the problem of inaccurate recognition of nearby wake-up devices in a noise environment is solved in the prior art, accurate identification of the distance between the device and the user and accurate identification of nearby wake-up devices is achieved, and user experience is improved.

CN118748014BActive Publication Date: 2025-05-06MIDEA GRP (SHANGHAI) CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410963592.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-17
Publication Date
2025-05-06
Estimated Expiration
2044-07-17

AI Technical Summary

Technical Problem

The existing method of identifying nearby wake-up device cannot accurately identify the distance between the device and the user in a noisy environment, resulting in inaccurate wake-up and affecting the user experience.

Method used

By determining whether the received user wake-up voice comes from a multi-channel or single-channel voice device, the target voice is reduced based on the noise suppression system or signal-to-noise ratio, the noise-reduced voice is obtained, and the nearby wake-up device is calculated based on its energy.

Benefits of technology

It effectively avoids noise interference, accurately characterizes the user's wake-up voice energy, recognizes the accurate distance between the device and the user, and ultimately accurately identifys the nearest wake-up device closest to the user, improving the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118748014B_ABST
    Figure CN118748014B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of nearby wake-up technology, and provides a nearby wake-up device identification method, device and electronic device. The method includes: determining whether the received user wake-up voice comes from a multi-channel voice device; if so, reducing the noise of the first target voice based on the noise suppression degree to obtain a first noise-reduced voice; if not, reducing the noise of the second target voice based on the signal-to-noise ratio to obtain a second noise-reduced voice; according to the first noise-reduced voice and / or the second noise-reduced voice, identifying the nearby wake-up device. Since the present invention performs targeted noise reduction on the target voice of any type of voice device, it can avoid noise interference on the energy information of each target voice, so that it can accurately characterize the user wake-up voice energy, thereby accurately identifying the distance between the device and the user, and finally accurately identifying the nearest nearby wake-up device to the user for wake-up, thereby improving the user experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of nearby wake-up, and in particular to a nearby wake-up device identification method, device and electronic device. Background Art

[0002] With the popularity of smart voice homes, there will be multiple smart voice devices in the same room or area at the same time. When the user shouts the wake-up word, it is necessary to identify the voice device closest to the user from multiple smart voice devices and wake it up to execute the user's subsequent voice tasks.

[0003] However, the current method for identifying nearby wake-up devices is mainly based on the energy information of the user's wake-up voice. However, when there is noise near the voice device, the energy information is interfered by the noise. The current nearby wake-up device identification method based on energy information cannot accurately identify the distance between the device and the user based on the energy information. It will prioritize the voice device closest to the noise rather than the voice device closest to the user as the wake-up device, resulting in inaccurate nearby wake-up and affecting the user experience. Summary of the invention

[0004] The present invention aims to solve at least one of the technical problems existing in the related art. To this end, the present invention proposes a nearby wake-up device identification method, which can prevent the energy information of the target voice of each voice device type from being interfered by noise, so that it can accurately characterize the user's wake-up voice energy, thereby accurately identifying the distance between the device and the user, and finally accurately identifying the nearest nearby wake-up device to the user for wake-up, thereby improving the user experience.

[0005] The present invention also provides a nearby wake-up device identification device, electronic device, storage medium and program product.

[0006] According to the first aspect of the present invention, a nearby wake-up device identification method includes:

[0007] Determine whether the received user wake-up voice comes from a multi-channel voice device;

[0008] If yes, then denoising the first target speech based on the noise suppression degree to obtain a first denoised speech;

[0009] If not, then denoising the second target speech based on the signal-to-noise ratio to obtain a second denoised speech;

[0010] Identify a nearby wake-up device according to the first noise-reduced voice and / or the second noise-reduced voice;

[0011] The first target voice is a voice obtained by processing a user wake-up voice received by a multi-channel voice device, and the second target voice is a voice obtained by processing the user wake-up voice received by a single-channel voice device.

[0012] According to the nearby wake-up device identification method of an embodiment of the present invention, for a multi-channel voice device, the target voice is denoised based on the noise suppression degree to obtain a first noise-reduced voice; for a single-channel voice device, the target voice is denoised based on the signal-to-noise ratio to obtain a second noise-reduced voice, and then the nearby wake-up device is identified based on the first noise-reduced voice and / or the second noise-reduced voice. Since the target voice of any type of voice device is specifically denoised, it is possible to avoid noise interference in the energy information of each target voice, so that it can accurately characterize the user's wake-up voice energy, thereby accurately identifying the distance between the device and the user, and finally accurately identifying the nearest wake-up device closest to the user for wake-up, thereby improving the user experience.

[0013] According to an embodiment of the present invention, the noise suppression degree is determined based on the following method:

[0014] Acquire the number of sound sources and a first signal-to-noise ratio of the first target speech;

[0015] If the number of sound sources is greater than 1, or the first signal-to-noise ratio is less than a signal-to-noise ratio threshold, determining the noise suppression degree to be a high suppression degree greater than the first suppression degree;

[0016] If the number of sound sources is equal to 1, and the first signal-to-noise ratio is greater than or equal to a signal-to-noise ratio threshold, determining the noise suppression degree to be a low suppression degree that is less than the second suppression degree;

[0017] The first suppression degree is greater than the second suppression degree.

[0018] According to an embodiment of the present invention, the step of reducing noise on the first target speech based on the noise suppression degree to obtain the first noise-reduced speech includes:

[0019] If the noise suppression degree is the high suppression degree, a high suppression degree noise reduction algorithm is used to reduce noise on the first target speech to obtain a first noise-reduced speech;

[0020] If the noise suppression degree is the low suppression degree, a low suppression degree noise reduction algorithm is used to reduce noise on the first target speech to obtain a first noise-reduced speech.

[0021] According to an embodiment of the present invention, the step of using a high suppression noise reduction algorithm to reduce noise on the first target speech to obtain a first noise-reduced speech includes:

[0022] Performing sound source localization on the first target speech to obtain a direction of arrival of the first target speech;

[0023] The first target speech is enhanced in the arrival direction to obtain a first noise-reduced speech.

[0024] According to an embodiment of the present invention, obtaining the number of sound sources of the first target speech includes:

[0025] Performing sound source localization on the first target speech to obtain the number of peaks of the signal waveform of the first target speech in the frequency domain;

[0026] The peak number is determined as the sound source number of the first target speech.

[0027] According to an embodiment of the present invention, the step of performing noise reduction on the second target speech based on the signal-to-noise ratio to obtain the second noise-reduced speech includes:

[0028] Acquire a second signal-to-noise ratio of the second target speech;

[0029] If the second signal-to-noise ratio is less than the signal-to-noise ratio threshold, single-channel noise reduction is performed on the second target speech to obtain a second noise-reduced speech.

[0030] According to an embodiment of the present invention, the identifying a nearby wake-up device according to the first noise-reduced voice and / or the second noise-reduced voice includes:

[0031] Selecting the first noise-reduced speech and / or the second noise-reduced speech in a preset frequency band to obtain a first speech in a preset frequency band and / or a second speech in a preset frequency band;

[0032] Performing energy calculation on the first speech in the preset frequency band and / or the second speech in the preset frequency band to obtain first speech energy and / or second speech energy;

[0033] The voice device corresponding to the maximum value of the first voice energy and / or the second voice energy is determined as the nearby wake-up device.

[0034] According to an embodiment of the present invention, a target signal-to-noise ratio is determined based on the following method, where the target signal-to-noise ratio includes the first signal-to-noise ratio and / or the second signal-to-noise ratio:

[0035] Acquire a first signal power spectrum between a target time and a starting time of a target speech; the target time is a preset time before the starting time of the target speech;

[0036] Acquire a second signal power spectrum between the start time and the end time of the target speech;

[0037] Obtaining a target signal-to-noise ratio corresponding to the target speech according to the first signal power spectrum and the second signal power spectrum;

[0038] The target speech includes the first target speech and / or the second target speech.

[0039] According to an embodiment of the present invention, the first target speech is obtained by processing in the following manner:

[0040] Multi-channel echo cancellation is performed on the user wake-up voice received by the multi-channel voice device to obtain the first target voice.

[0041] According to an embodiment of the present invention, the second target speech is obtained by processing in the following manner:

[0042] Single-channel echo cancellation is performed on the user wake-up voice received by the single-channel voice device to obtain the second target voice.

[0043] According to a second aspect of the present invention, a nearby wake-up device identification device includes:

[0044] The voice source determination module is used to determine whether the received user wake-up voice comes from a multi-channel voice device;

[0045] A first noise reduction module is used to: if yes, perform noise reduction on the first target speech based on the noise suppression degree to obtain a first noise-reduced speech;

[0046] A second noise reduction module is used to: if not, perform noise reduction on the second target speech based on the signal-to-noise ratio to obtain a second noise-reduced speech;

[0047] A wake-up device identification module, used to: identify a nearby wake-up device according to the first noise-reduced voice and / or the second noise-reduced voice;

[0048] The first target voice is a voice obtained by processing a user wake-up voice received by a multi-channel voice device, and the second target voice is a voice obtained by processing the user wake-up voice received by a single-channel voice device.

[0049] According to the apparatus for identifying nearby wake-up devices in an embodiment of the present invention, for multi-channel voice devices, noise reduction is performed on the target voice based on the noise suppression degree to obtain a first noise-reduced voice; for single-channel voice devices, noise reduction is performed on the target voice based on the signal-to-noise ratio to obtain a second noise-reduced voice, and then the nearby wake-up device is identified based on the first noise-reduced voice and / or the second noise-reduced voice. Since the target voice of any type of voice device is specifically noise-reduced, it is possible to prevent the energy information of each target voice from being interfered by noise, so that it can accurately characterize the energy of the user's wake-up voice, thereby accurately identifying the distance between the device and the user, and finally accurately identifying the nearest wake-up device to the user for wake-up, thereby improving the user experience.

[0050] According to an embodiment of the third aspect of the present invention, an electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, and when the processor executes the program, the method for identifying a nearby wake-up device described in the embodiment of the first aspect is implemented.

[0051] According to the non-transitory computer-readable storage medium of the fourth aspect embodiment of the present invention, a computer program is stored thereon, and when the computer program is executed by a processor, the method for identifying a nearby wake-up device described in the first aspect embodiment is implemented.

[0052] According to a computer program product of an embodiment of the fifth aspect of the present invention, it includes a computer program, and when the computer program is executed by a processor, it implements the nearby wake-up device identification method described in the embodiment of the first aspect.

[0053] The above one or more technical solutions in the embodiments of the present invention have at least one of the following technical effects:

[0054] For multi-channel voice devices, the present invention performs noise reduction on the target voice based on the noise suppression degree to obtain the first noise-reduced voice. For single-channel voice devices, the present invention performs noise reduction on the target voice based on the signal-to-noise ratio to obtain the second noise-reduced voice. Then, based on the first noise-reduced voice and / or the second noise-reduced voice, the nearby wake-up device is identified. Since the target voice of any type of voice device is specifically noise-reduced, it is possible to avoid noise interference on the energy information of each target voice, so that the energy of the user's wake-up voice can be accurately characterized, thereby accurately identifying the distance between the device and the user, and finally accurately identifying the nearest wake-up device to the user for wake-up, thereby improving the user experience.

[0055] Additional aspects and advantages of the present invention will be given in part in the following description and in part will be obvious from the following description, or will be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or related technologies, the drawings required for use in the embodiments or related technical descriptions are briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0057] Figure 1 This is one of the flow charts of the nearby wake-up device identification method provided by an embodiment of the present invention;

[0058] Figure 2 This is a second flow chart of the nearby wake-up device identification method provided by an embodiment of the present invention;

[0059] Figure 3 This is a third flow chart of the nearby wake-up device identification method provided by an embodiment of the present invention;

[0060] Figure 4 This is a fourth flow chart of the nearby wake-up device identification method provided by an embodiment of the present invention;

[0061] Figure 5 This is a fifth flow chart of the method for identifying a nearby wake-up device provided in an embodiment of the present invention;

[0062] Figure 6 It is a structural diagram of a nearby wake-up device identification device provided by an embodiment of the present invention;

[0063] Figure 7 It is a schematic diagram of the structure of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0064] The following embodiments of the present invention are described in further detail in conjunction with the accompanying drawings and examples. The following examples are used to illustrate the present invention, but are not intended to limit the scope of the present invention.

[0065] In the description of the embodiments of the present invention, it should be noted that the terms "center", "longitudinal", "lateral", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside" and the like indicate positions or positional relationships based on the positions or positional relationships shown in the accompanying drawings, and are only for the convenience of describing the embodiments of the present invention and simplifying the description, and do not indicate or imply that the devices or elements referred to must have a specific orientation, be constructed and operate in a specific orientation, and therefore cannot be understood as limitations on the embodiments of the present invention. In addition, the terms "first", "second", and "third" are used for descriptive purposes only and cannot be understood as indicating or implying relative importance.

[0066] In the description of the embodiments of the present invention, it should be noted that, unless otherwise clearly specified and limited, the terms "connected" and "connection" should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection or an indirect connection through an intermediate medium. For ordinary technicians in this field, the specific meanings of the above terms in the embodiments of the present invention can be understood according to specific circumstances.

[0067] In the embodiments of the present invention, unless otherwise clearly specified and limited, the first feature being "above" or "below" the second feature may mean that the first and second features are in direct contact, or the first and second features are in indirect contact through an intermediate medium. Moreover, the first feature being "above", "above" and "above" the second feature may mean that the first feature is directly above or obliquely above the second feature, or simply means that the first feature is higher in level than the second feature. The first feature being "below", "below" and "below" the second feature may mean that the first feature is directly below or obliquely below the second feature, or simply means that the first feature is lower in level than the second feature.

[0068] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the embodiment of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples, without contradiction.

[0069] Figure 1 This is one of the flow charts of the nearby wake-up device identification method provided by an embodiment of the present invention. Figure 1 , an embodiment of the present invention provides a method for identifying a nearby wake-up device, which may include:

[0070] 101. Determine whether the received user wake-up voice comes from a multi-channel voice device;

[0071] 102. If yes, perform noise reduction on the first target speech based on the noise suppression degree to obtain a first noise-reduced speech;

[0072] 103. If not, perform noise reduction on the second target speech based on the signal-to-noise ratio to obtain a second noise-reduced speech;

[0073] 104. Identify a nearby wake-up device according to the first noise-reduced voice and / or the second noise-reduced voice.

[0074] The first target voice is a voice obtained by processing the user wake-up voice received by the multi-channel voice device, and the second target voice is a voice obtained by processing the user wake-up voice received by the single-channel voice device.

[0075] In this embodiment, the execution subject can be one of the multi-channel voice devices, one of the single-channel voice devices, or a third-party central control device, which is not limited here.

[0076] In step 101, other voice devices other than the execution subject will send the received user wake-up voice to the execution subject. When the execution subject is not the central control device, it will also receive the user wake-up voice. At this time, the execution subject can first judge the channel category of itself and other voice devices, and specifically process and classify the user wake-up voices sent by devices of different channel types together with the user wake-up voices received by itself to obtain the first target voice and the second target voice.

[0077] In step 104, since the execution subject may only receive the user wake-up voice from the multi-channel voice device, may only receive the user wake-up voice from the single-channel voice device, or may receive the user wake-up voice from both the multi-channel voice device and the single-channel voice device, different methods need to be used to identify the nearest wake-up device for these three situations:

[0078] 1. If only the user wake-up voice is received from a multi-channel voice device, only the nearest device needs to be woken up based on the first noise-reduced voice recognition;

[0079] 2. If only the user wake-up voice is received from a single-channel voice device, only the nearest device needs to be woken up based on the second noise-reduced voice recognition;

[0080] 3. If the user wake-up voice is received from a multi-channel voice device and a single-channel voice device at the same time, it is necessary to identify the nearest wake-up device based on the first noise reduction voice and the second noise reduction voice.

[0081] The method for identifying nearby wake-up devices provided in this embodiment performs noise reduction on the target voice based on the noise suppression degree to obtain the first noise-reduced voice. For single-channel voice devices, the target voice is noise-reduced based on the signal-to-noise ratio to obtain the second noise-reduced voice, and then the nearby wake-up device is identified based on the first noise-reduced voice and / or the second noise-reduced voice. Since the target voice of any type of voice device is noise-reduced in a targeted manner, it is possible to avoid noise interference on the energy information of each target voice, so that it can accurately characterize the user's wake-up voice energy, thereby accurately identifying the distance between the device and the user, and finally accurately identifying the nearest wake-up device closest to the user for wake-up, thereby improving the user experience.

[0082] Figure 2 This is a second flow chart of the nearby wake-up device identification method provided by an embodiment of the present invention. Figure 2 In one embodiment, the noise suppression degree can be determined based on the following method:

[0083] 201. Obtain the number of sound sources and a first signal-to-noise ratio of a first target speech;

[0084] 202. If the number of sound sources is greater than 1, or the first signal-to-noise ratio is less than the signal-to-noise ratio threshold, determine the noise suppression degree to be a high suppression degree greater than the first suppression degree;

[0085] 203. If the number of sound sources is equal to 1, and the first signal-to-noise ratio is greater than or equal to the signal-to-noise ratio threshold, determine that the noise suppression degree is a low suppression degree that is less than the second suppression degree.

[0086] The first suppression degree is greater than the second suppression degree.

[0087] When there is a stable noise source near the voice device corresponding to the first target voice, there will be corresponding stable noise in the first target voice. Therefore, its sound source comes not only from the user but also from the noise source, that is, the number of sound sources corresponding to the first target voice is at least two; when there is no stable noise source near the voice device corresponding to the first target voice, there will be no stable noise in the first target voice. Therefore, its sound source only comes from the user, that is, the number of sound sources corresponding to the first target voice is only one;

[0088] The signal-to-noise ratio can represent the level of noise. When the first signal-to-noise ratio is high, it means that the first target speech has low noise; when the first signal-to-noise ratio is low, it means that the first target speech has high noise.

[0089] Therefore, if the number of sound sources is greater than 1, or the first signal-to-noise ratio is less than the signal-to-noise ratio threshold, it means that there is high noise or a stable noise source near the multi-channel voice device, and the noise suppression degree needs to be set to a higher first suppression degree; if the number of sound sources is equal to 1, and the first signal-to-noise ratio is greater than or equal to the signal-to-noise ratio threshold, it means that there is low noise and no stable noise source near the multi-channel voice device, and the noise suppression degree needs to be set to a lower second suppression degree. Among them, the signal-to-noise ratio threshold, the first suppression degree and the second suppression degree can be set according to actual conditions, and are not limited here. In this embodiment, the signal-to-noise ratio threshold can be set to 5 decibels.

[0090] In this embodiment, since the number of sound sources and the signal-to-noise ratio can make a relatively accurate judgment on the stability and level of noise, the comprehensive judgment of the number of sound sources and the signal-to-noise ratio can be used to screen out situations with high noise or stable noise sources and low noise and no stable noise sources, so as to set the noise suppression degree in a targeted manner, so that targeted noise reduction can be performed in various situations later.

[0091] In one embodiment, performing noise reduction on the first target speech based on the noise suppression degree to obtain the first noise-reduced speech may include:

[0092] If the noise suppression degree is high, a high suppression degree noise reduction algorithm is used to reduce noise on the first target speech to obtain a first noise-reduced speech;

[0093] If the noise suppression degree is a low suppression degree, a low suppression degree noise reduction algorithm is used to reduce the noise of the first target speech to obtain a first noise-reduced speech.

[0094] Among them, high-suppression noise reduction algorithms include array enhancement algorithms such as MVDR (Minimum Variance Distortionless Response) algorithm and GSC (Generalized Sidelobe Canceller) algorithm, which can achieve high-suppression noise reduction; low-suppression noise reduction algorithms include DSB (Delay and Sum Beamforming) algorithm, which can eliminate certain reverberation effects.

[0095] In this embodiment, when there is high noise or a stable noise source near the multi-channel voice device, a high-suppression noise reduction algorithm is used to suppress the noise energy as much as possible. When there is low noise and no stable noise source near the multi-channel voice device, a low-suppression noise reduction algorithm is used to retain the energy of the target voice as much as possible. This ensures that the energy information of the target voice of the multi-channel voice device can accurately represent the actual wake-up voice energy of the user regardless of the degree of noise interference.

[0096] In one embodiment, using a high suppression noise reduction algorithm to reduce noise on the first target speech to obtain a first noise-reduced speech may include:

[0097] The first target speech is localized to obtain a direction of arrival of the first target speech, and the first target speech is enhanced in the direction of arrival to obtain a first noise-reduced speech.

[0098] The sound source localization may adopt the MVDR algorithm, the MUSIC (Multiple Signal Classification) algorithm, the SRP-PHAT (Spectral Ratio with Phase Transform) algorithm or the multi-sound source localization fusion algorithm, which is not limited here.

[0099] Specifically, the first target speech can be enhanced in the wave direction through the microphone array, thereby achieving a high degree of noise reduction and obtaining a first noise-reduced speech.

[0100] The high-suppression noise reduction algorithm of this embodiment performs targeted enhancement on the target speech at its arrival party to reversely suppress noise energy, thereby achieving high-suppression noise reduction.

[0101] In one embodiment, obtaining the number of sound sources of the first target speech may include:

[0102] The sound source of the first target speech is localized to obtain the number of peaks of the signal waveform of the first target speech in the frequency domain, and the number of peaks is determined as the number of sound sources of the first target speech.

[0103] Generally speaking, a speech signal waveform formed by a sound source has only one peak value in the frequency domain. Therefore, the number of corresponding sound sources can be determined according to the number of peak values ​​of the signal waveform of the first target speech in the frequency domain.

[0104] The same sound source localization method as mentioned above can be used to obtain the number of sound sources of the first target voice while obtaining the arrival direction of the first target voice. Alternatively, a method different from the above method can be used to obtain the number of sound sources of the first target voice, which is not limited here.

[0105] This embodiment can obtain the accurate number of sound sources quickly and conveniently according to the number of peak values ​​of the speech signal waveform within the frequency range through the one-to-one correspondence between the number of sound sources.

[0106] In one embodiment, performing noise reduction on the second target speech based on the signal-to-noise ratio to obtain the second noise-reduced speech may include:

[0107] A second signal-to-noise ratio of the second target speech is obtained. If the second signal-to-noise ratio is less than the signal-to-noise ratio threshold, single-channel noise reduction is performed on the second target speech to obtain a second noise-reduced speech.

[0108] The single-channel voice device corresponding to the second target voice may be a single-microphone device.

[0109] Since single-channel voice devices lack spatial information, they cannot directly use spatial filtering technologies such as beamforming to distinguish between the user wake-up voice and noise in the corresponding second target voice. Therefore, they need to rely on traditional signal processing methods such as time domain filtering, frequency domain filtering, or deep neural networks to reduce the noise of the second target voice.

[0110] Based on the characteristics of a single-channel voice device, this embodiment performs targeted noise reduction on its second target voice, which can effectively suppress noise energy, so that no matter what degree of noise interference the target voice of the single-channel voice device is subject to, its energy information can accurately represent the user's actual wake-up voice energy.

[0111] Figure 3 This is a flowchart of the method for identifying a nearby wake-up device provided by an embodiment of the present invention. Figure 3 In one embodiment, identifying a nearby wake-up device according to the first noise-reduced voice and / or the second noise-reduced voice may include:

[0112] 301. Selecting a first noise-reduced speech and / or a second noise-reduced speech in a preset frequency band to obtain a first speech in the preset frequency band and / or a second speech in the preset frequency band;

[0113] 302. Perform energy calculation on a first speech in a preset frequency band and / or a second speech in a preset frequency band to obtain first speech energy and / or second speech energy;

[0114] 303. Determine a voice device corresponding to a maximum value of the first voice energy and / or the second voice energy as a nearby wake-up device.

[0115] In step 301, the preset frequency band can be selected according to actual needs and is not limited here. In this embodiment, a bandpass filter can be used to select the first noise-reduced speech and / or the second noise-reduced speech in a frequency band between 200 Hz and 4000 Hz.

[0116] In step 303, under ideal conditions without noise, the closer the voice device is to the user, the greater the energy of the user's wake-up voice it receives. Therefore, after the noise influence in each target voice is eliminated to the maximum extent, its voice energy can be close to the voice energy under ideal conditions, so it can be considered that the voice device corresponding to the maximum voice energy is the voice device closest to the user.

[0117] It should be noted that, corresponding to the three situations in which the execution subject receives the source of the user's wake-up voice, different methods are used to identify the nearby wake-up device, which can be specifically:

[0118] 1. If only the user wake-up voice from the multi-channel voice device is received, the voice of the first noise-reduced voice in the preset frequency band is selected to obtain the first voice in the preset frequency band, and the energy of the first voice in the preset frequency band is calculated to obtain the first voice energy, and the voice device corresponding to the maximum value of the first voice energy is determined as the nearest wake-up device;

[0119] 2. If only the user wake-up voice is received from a single-channel voice device, the second noise-reduced voice in the preset frequency band is selected to obtain the second voice in the preset frequency band, and the energy of the second voice in the preset frequency band is calculated to obtain the second voice energy. The voice device corresponding to the maximum value of the second voice energy is determined as the nearest wake-up device;

[0120] 3. If the user wake-up voice is received from a multi-channel voice device and a single-channel voice device at the same time, the first noise-reduced voice and the voice of the first noise-reduced voice in the preset frequency band are selected to obtain the first voice in the preset frequency band and the second voice in the preset frequency band, and the energy of the first voice in the preset frequency band and the second voice in the preset frequency band are calculated to obtain the first voice energy and the second voice energy, and the voice device corresponding to the maximum value of the first voice energy and the second voice energy is determined as the nearest wake-up device.

[0121] In this embodiment, by calculating the speech energy of a specific frequency band after noise reduction, the energy of each speech can be made close to the ideal noise-free speech energy. Therefore, based on the relationship between the size of the speech energy and the distance between the user and the voice device, the voice device closest to the user can be accurately identified and used as the nearby wake-up device.

[0122] Figure 4 This is a fourth flow chart of the nearby wake-up device identification method provided by an embodiment of the present invention. Figure 4 In one embodiment, determining a target signal-to-noise ratio based on the following method, where the target signal-to-noise ratio includes the first signal-to-noise ratio and / or the second signal-to-noise ratio, may include:

[0123] 401. Acquire a first signal power spectrum between a target time and a start time of a target speech;

[0124] The target time is a preset time before the onset time of the target speech;

[0125] 402. Acquire a second signal power spectrum between a start time and an end time of the target speech;

[0126] 403. Obtain a target signal-to-noise ratio corresponding to the target speech according to the first signal power spectrum and the second signal power spectrum.

[0127] The target speech includes a first target speech and / or a second target speech.

[0128] In step 401, the target time can be set according to actual conditions and is not limited here. In this embodiment, the target time can be set to a time 2 seconds before the start time of the target speech.

[0129] In step 402, the start time and the stop time of the target speech are first detected, the speech between the two times is intercepted, and the signal power spectrum corresponding to the speech is obtained.

[0130] Intercepting the speech according to the start and stop times of the target speech can prevent the speech outside the period between the start and stop times from being mistakenly included. On the one hand, the signal-to-noise ratio can be calculated more accurately. On the other hand, the influence of other noises can be reduced in the subsequent energy calculation.

[0131] In step 403, the target signal-to-noise ratio can be calculated according to the following formula: :

[0132] ;

[0133] in, is the first power spectrum, is the second power spectrum.

[0134] It should be noted that the above steps include three situations:

[0135] 1. If the target signal-to-noise ratio only includes the first signal-to-noise ratio, then the target speech only includes the first target speech. In this embodiment, the first signal-to-noise ratio is obtained based on the signal power spectrum of the first target speech related time period;

[0136] 2. If the target signal-to-noise ratio only includes the second signal-to-noise ratio, then the target speech only includes the second target speech, and this embodiment obtains the second signal-to-noise ratio based on the signal power spectrum of the second target speech related time period;

[0137] 3. If the target signal-to-noise ratio includes both the first signal-to-noise ratio and the second signal-to-noise ratio, then the target speech includes both the first target speech and the second target speech. In this embodiment, the first signal-to-noise ratio is obtained based on the signal power spectrum of the first target speech-related time period, and the second signal-to-noise ratio is obtained based on the signal power spectrum of the second target speech-related time period.

[0138] This embodiment calculates the target signal-to-noise ratio through the power spectrum of the target speech period and the power spectrum of the specific period before the target speech, that is, the target signal-to-noise ratio is calculated through the first power spectrum in which only noise exists and the second power spectrum in which both noise and user wake-up speech exist. This can more accurately evaluate the ratio of useful signals to interference signals in the target speech, thereby obtaining an accurate target signal-to-noise ratio.

[0139] In one embodiment, the first target speech may be obtained by processing in the following manner:

[0140] Multi-channel echo cancellation is performed on the user wake-up voice received by the multi-channel voice device to obtain a first target voice.

[0141] Similarly, the second target speech can be obtained by processing in the following way:

[0142] Single-channel echo cancellation is performed on the user wake-up voice received by the single-channel voice device to obtain a second target voice.

[0143] Whether it is a single-channel voice device or a multi-channel voice device, when it has its own broadcast, the energy information of the user's wake-up voice it receives will also be interfered by its own broadcast. The nearby wake-up method based on energy information will give priority to the voice device with the broadcast as the wake-up device, resulting in inaccurate nearby wake-up and affecting the user experience.

[0144] In this embodiment, targeted echo cancellation is performed for both single-channel voice devices and multi-channel voice devices, which can minimize the nonlinear noise generated by the broadcast when each voice device has its own broadcast, thereby reducing the impact of the broadcast echo on nearby wake-up.

[0145] Figure 5 This is a fifth flow chart of the nearby wake-up device identification method provided by an embodiment of the present invention. Figure 5 In one embodiment, the whole process of the nearby wake-up device identification method can be described as follows:

[0146] 1. Determine whether the received user wake-up voice comes from a multi-channel voice device;

[0147] 2. If yes, perform multi-channel echo cancellation on the user wake-up voice to obtain the first target voice; if no, perform single-channel echo cancellation on the user wake-up voice to obtain the second target voice;

[0148] 3. Determine whether to enter the awakening state;

[0149] 4. If yes, for a multi-channel speech device, monitor the start and end time of the first target speech, and for a single-channel speech device, monitor the start and end time of the second target speech; if no, return to step 1;

[0150] 5. For the first target voice, perform sound source localization and signal-to-noise ratio calculation before and after wake-up. For the second target voice, only the signal-to-noise ratio calculation before and after wake-up is required.

[0151] 6. According to the calculation result of the first target voice, determine whether it is necessary to suppress the noise; according to the calculation result of the second target voice, determine whether the signal-to-noise ratio before and after wake-up is less than 5 decibels;

[0152] 7. For the first target voice, if high noise suppression is required, a high suppression noise reduction algorithm is used for noise reduction. If high noise suppression is not required, a low suppression noise reduction algorithm is used for noise reduction. For the second target voice, if the signal-to-noise ratio before and after wake-up is less than 5 decibels, single-channel noise reduction is used.

[0153] 8. Regardless of high suppression noise reduction, low suppression noise reduction or single-channel noise reduction, for all the voices after noise reduction, a bandpass filter is used to select a unified target frequency band;

[0154] 9. Calculate the energy of the speech in each target frequency band;

[0155] 10. If there is only a multi-channel voice device in the device domain and high suppression noise reduction is used for its first target voice, the multi-channel voice device corresponding to the maximum value of all voice energy in the target frequency band after high suppression noise reduction is selected as the nearest wake-up device;

[0156] When there is only a multi-channel voice device in the device domain and low suppression noise reduction is used for its first target voice, the multi-channel voice device corresponding to the maximum value of all voice energy in the target frequency band after low suppression noise reduction is selected as the nearest wake-up device;

[0157] When there is only a single-channel voice device in the device domain, the single-channel voice device corresponding to the maximum value of all voice energy in the target frequency band after single-channel noise reduction is selected as the nearest wake-up device;

[0158] When there are both multi-channel voice devices and single-channel voice devices in the device domain, the target voices corresponding to each voice device are selected and noise reduction is performed according to their respective noise reduction algorithms. The voice device corresponding to the maximum voice energy of each noise-reduced voice in the target frequency band is used as the nearest wake-up device.

[0159] This embodiment can be applicable to scenarios where only single-channel voice devices or multi-channel voice devices exist in the same space, and can also be applicable to scenarios where both single-channel voice devices and multi-channel voice devices exist in the same space, thereby realizing nearby wake-up device recognition in all scenarios.

[0160] The following is a description of a nearby wake-up device identification device provided in an embodiment of the present invention. The nearby wake-up device identification device described below and the nearby wake-up device identification method described above can be referenced to each other.

[0161] Figure 6 Schematic diagram of the structure of the nearby wake-up device identification device provided by an embodiment of the present invention. Figure 6 , an embodiment of the present invention provides a nearby wake-up device identification device, which may include:

[0162] The voice source determination module 601 is used to determine whether the received user wake-up voice comes from a multi-channel voice device;

[0163] A first noise reduction module 602 is configured to: if yes, perform noise reduction on the first target speech based on the noise suppression degree to obtain a first noise-reduced speech;

[0164] A second noise reduction module 603 is used to: if not, perform noise reduction on the second target speech based on the signal-to-noise ratio to obtain a second noise-reduced speech;

[0165] The wake-up device identification module 604 is used to: identify a nearby wake-up device according to the first noise-reduced voice and / or the second noise-reduced voice;

[0166] The first target voice is a voice obtained by processing a user wake-up voice received by a multi-channel voice device, and the second target voice is a voice obtained by processing the user wake-up voice received by a single-channel voice device.

[0167] The nearby wake-up device identification device provided in this embodiment performs noise reduction on its target voice based on the noise suppression degree to obtain a first noise-reduced voice. For a single-channel voice device, the target voice is noise-reduced based on the signal-to-noise ratio to obtain a second noise-reduced voice, and then the nearby wake-up device is identified based on the first noise-reduced voice and / or the second noise-reduced voice. Since the target voice of any type of voice device is noise-reduced in a targeted manner, the energy information of each target voice can be prevented from being interfered by noise, so that it can accurately characterize the user's wake-up voice energy, thereby accurately identifying the distance between the device and the user, and finally accurately identifying the nearest wake-up device closest to the user for wake-up, thereby improving the user experience.

[0168] In one embodiment, a noise suppression degree determination module (not shown in the figure) is further included, which is used to:

[0169] Acquire the number of sound sources and a first signal-to-noise ratio of the first target speech;

[0170] If the number of sound sources is greater than 1, or the first signal-to-noise ratio is less than a signal-to-noise ratio threshold, determining the noise suppression degree to be a high suppression degree greater than the first suppression degree;

[0171] If the number of sound sources is equal to 1, and the first signal-to-noise ratio is greater than or equal to a signal-to-noise ratio threshold, determining the noise suppression degree to be a low suppression degree that is less than the second suppression degree;

[0172] The first suppression degree is greater than the second suppression degree.

[0173] In one embodiment, the first noise reduction module 602 is specifically configured to:

[0174] If the noise suppression degree is the high suppression degree, a high suppression degree noise reduction algorithm is used to reduce noise on the first target speech to obtain a first noise-reduced speech;

[0175] If the noise suppression degree is the low suppression degree, a low suppression degree noise reduction algorithm is used to reduce noise on the first target speech to obtain a first noise-reduced speech.

[0176] In one embodiment, the first noise reduction module 602 is specifically configured to:

[0177] Performing sound source localization on the first target speech to obtain a direction of arrival of the first target speech;

[0178] The first target speech is enhanced in the arrival direction to obtain a first noise-reduced speech.

[0179] In one embodiment, the noise suppression degree determination module is specifically configured to:

[0180] Performing sound source localization on the first target speech to obtain the number of peaks of the signal waveform of the first target speech in the frequency domain;

[0181] The peak number is determined as the sound source number of the first target speech.

[0182] In one embodiment, the second noise reduction module 603 is specifically configured to:

[0183] Acquire a second signal-to-noise ratio of the second target speech;

[0184] If the second signal-to-noise ratio is less than the signal-to-noise ratio threshold, single-channel noise reduction is performed on the second target speech to obtain a second noise-reduced speech.

[0185] In one embodiment, the wake-up device identification module 604 is specifically used to:

[0186] Selecting the first noise-reduced speech and / or the second noise-reduced speech in a preset frequency band to obtain a first speech in a preset frequency band and / or a second speech in a preset frequency band;

[0187] Performing energy calculation on the first speech in the preset frequency band and / or the second speech in the preset frequency band to obtain first speech energy and / or second speech energy;

[0188] The voice device corresponding to the maximum value of the first voice energy and / or the second voice energy is determined as the nearby wake-up device.

[0189] In one embodiment, a target signal-to-noise ratio determination module (not shown in the figure) is further included, which is used to:

[0190] Acquire a first signal power spectrum between a target time and a starting time of a target speech; the target time is a preset time before the starting time of the target speech;

[0191] Acquire a second signal power spectrum between the start time and the end time of the target speech;

[0192] Obtaining a target signal-to-noise ratio corresponding to the target speech according to the first signal power spectrum and the second signal power spectrum;

[0193] The target speech includes the first target speech and / or the second target speech.

[0194] In one embodiment, an echo cancellation module (not shown in the figure) is further included, which is used to:

[0195] Multi-channel echo cancellation is performed on the user wake-up voice received by the multi-channel voice device to obtain the first target voice.

[0196] In one embodiment, the echo cancellation module is further configured to:

[0197] Single-channel echo cancellation is performed on the user wake-up voice received by the single-channel voice device to obtain the second target voice.

[0198] Figure 7 is a schematic diagram of the structure of an electronic device provided by an embodiment of the present invention, such as Figure 7 As shown, the electronic device may include: a processor 710, a communication interface 720, a memory 730 and a communication bus 740, wherein the processor 710, the communication interface 720 and the memory 730 communicate with each other through the communication bus 740. The processor 710 may call the logic instructions in the memory 730 to execute the following method:

[0199] Determine whether the received user wake-up voice comes from a multi-channel voice device;

[0200] If yes, then denoising the first target speech based on the noise suppression degree to obtain a first denoised speech;

[0201] If not, denoising the second target speech based on the signal-to-noise ratio to obtain a second denoised speech;

[0202] Identify a nearby wake-up device according to the first noise-reduced voice and / or the second noise-reduced voice;

[0203] The first target voice is a voice obtained by processing a user wake-up voice received by a multi-channel voice device, and the second target voice is a voice obtained by processing the user wake-up voice received by a single-channel voice device.

[0204] In addition, the logic instructions in the above-mentioned memory 730 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when it is sold or used as an independent product. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the relevant technology or the part of the technical solution, can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk and other media that can store program codes.

[0205] On the other hand, an embodiment of the present invention discloses a computer program product, wherein the computer program product includes a computer program stored on a non-transitory computer-readable storage medium, wherein the computer program includes program instructions. When the program instructions are executed by a computer, the computer can perform the methods provided by the above-mentioned method embodiments, for example, including:

[0206] Determine whether the received user wake-up voice comes from a multi-channel voice device;

[0207] If yes, then denoising the first target speech based on the noise suppression degree to obtain a first denoised speech;

[0208] If not, then denoising the second target speech based on the signal-to-noise ratio to obtain a second denoised speech;

[0209] Identify a nearby wake-up device according to the first noise-reduced voice and / or the second noise-reduced voice;

[0210] The first target voice is a voice obtained by processing a user wake-up voice received by a multi-channel voice device, and the second target voice is a voice obtained by processing the user wake-up voice received by a single-channel voice device.

[0211] In another aspect, an embodiment of the present invention further provides a non-transitory computer-readable storage medium having a computer program stored thereon, which is implemented when the computer program is executed by a processor to perform the transmission method provided in the above embodiments, for example, including:

[0212] Determine whether the received user wake-up voice comes from a multi-channel voice device;

[0213] If yes, then denoising the first target speech based on the noise suppression degree to obtain a first denoised speech;

[0214] If not, then denoising the second target speech based on the signal-to-noise ratio to obtain a second denoised speech;

[0215] Identify a nearby wake-up device according to the first noise-reduced voice and / or the second noise-reduced voice;

[0216] The first target voice is a voice obtained by processing a user wake-up voice received by a multi-channel voice device, and the second target voice is a voice obtained by processing the user wake-up voice received by a single-channel voice device.

[0217] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative labor.

[0218] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the relevant technology can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiment.

[0219] Finally, it should be noted that the above embodiments are only used to illustrate the present invention, rather than to limit the present invention. Although the present invention is described in detail with reference to the embodiments, it should be understood by those skilled in the art that various combinations, modifications or equivalent substitutions of the technical solutions of the present invention do not depart from the spirit and scope of the technical solutions of the present invention, and should be included in the scope of the claims of the present invention.

Claims

1. A method for identifying a nearby wake-up device, characterized in that: include: Determine whether the received user wake-up voice comes from a multi-channel voice device; If yes, then the first target speech is denoised based on the noise suppression degree to obtain a first denoised speech, including: The noise suppression degree includes a high suppression degree and a low suppression degree. Based on different suppression degrees, a corresponding noise reduction algorithm is used to reduce noise on the first target speech to obtain a first noise-reduced speech; If not, the second target speech is denoised based on the signal-to-noise ratio to obtain a second denoised speech, including: Acquire a second signal-to-noise ratio of the second target speech; If the second signal-to-noise ratio is less than the signal-to-noise ratio threshold, performing single-channel noise reduction on the second target speech to obtain a second noise-reduced speech; Identifying a nearby wake-up device according to the first noise-reduced voice and / or the second noise-reduced voice includes: If only the user wake-up voice is received from the multi-channel voice device, the voice of the first noise-reduced voice in the preset frequency band is selected to obtain the first voice in the preset frequency band, the energy of the first voice in the preset frequency band is calculated to obtain the first voice energy, and the voice device corresponding to the maximum value of the first voice energy is determined as the nearest wake-up device; If only the user wake-up voice is received from a single-channel voice device, the voice of the second noise-reduced voice in the preset frequency band is selected to obtain the second voice in the preset frequency band, and the energy of the second voice in the preset frequency band is calculated to obtain the second voice energy, and the voice device corresponding to the maximum value of the second voice energy is determined as the nearest wake-up device; If a user wake-up voice is received from a multi-channel voice device and a single-channel voice device at the same time, the first noise-reduced voice and the voice of the first noise-reduced voice in a preset frequency band are selected to obtain a first voice in the preset frequency band and a second voice in the preset frequency band, energy calculation is performed on the first voice in the preset frequency band and the second voice in the preset frequency band to obtain a first voice energy and a second voice energy, and a voice device corresponding to the maximum value of the first voice energy and the second voice energy is determined as the nearest wake-up device; The first target voice is a voice obtained by processing a user wake-up voice received by a multi-channel voice device, and the second target voice is a voice obtained by processing the user wake-up voice received by a single-channel voice device.

2. The method for identifying a nearby wake-up device according to claim 1, characterized in that: The noise suppression degree is determined based on the following method: Acquire the number of sound sources and a first signal-to-noise ratio of the first target speech; If the number of sound sources is greater than 1, or the first signal-to-noise ratio is less than a signal-to-noise ratio threshold, determining the noise suppression degree to be a high suppression degree greater than the first suppression degree; If the number of sound sources is equal to 1, and the first signal-to-noise ratio is greater than or equal to a signal-to-noise ratio threshold, determining the noise suppression degree to be a low suppression degree that is less than the second suppression degree; The first suppression degree is greater than the second suppression degree.

3. The method for identifying a nearby wake-up device according to claim 2, characterized in that: The step of reducing the noise of the first target speech based on the noise suppression degree to obtain the first noise-reduced speech includes: If the noise suppression degree is the high suppression degree, a high suppression degree noise reduction algorithm is used to reduce noise on the first target speech to obtain a first noise-reduced speech; If the noise suppression degree is the low suppression degree, a low suppression degree noise reduction algorithm is used to reduce noise on the first target speech to obtain a first noise-reduced speech.

4. The method for identifying a nearby wake-up device according to claim 3, characterized in that: The step of using a high suppression noise reduction algorithm to reduce the noise of the first target speech to obtain a first noise-reduced speech includes: Performing sound source localization on the first target speech to obtain a direction of arrival of the first target speech; The first target speech is enhanced in the arrival direction to obtain a first noise-reduced speech.

5. The method for identifying a nearby wake-up device according to claim 2, characterized in that: The obtaining the number of sound sources of the first target speech includes: Performing sound source localization on the first target speech to obtain the number of peaks of the signal waveform of the first target speech in the frequency domain; The peak number is determined as the sound source number of the first target speech.

6. The method for identifying a nearby wake-up device according to claim 2, characterized in that: Determining a target signal-to-noise ratio based on the following manner, where the target signal-to-noise ratio includes the first signal-to-noise ratio and / or the second signal-to-noise ratio, includes: Acquire a first signal power spectrum between a target time and a starting time of a target speech; the target time is a preset time before the starting time of the target speech; Acquire a second signal power spectrum between the start time and the end time of the target speech; Obtaining a target signal-to-noise ratio corresponding to the target speech according to the first signal power spectrum and the second signal power spectrum; The target speech includes the first target speech and / or the second target speech.

7. The method for identifying a nearby wake-up device according to any one of claims 1 to 6, characterized in that: The first target speech is obtained by processing in the following manner: Multi-channel echo cancellation is performed on the user wake-up voice received by the multi-channel voice device to obtain the first target voice.

8. The method for identifying a nearby wake-up device according to any one of claims 1 to 6, characterized in that: The second target speech is obtained by processing in the following manner: Single-channel echo cancellation is performed on the user wake-up voice received by the single-channel voice device to obtain the second target voice.

9. A nearby wake-up device identification device, characterized in that: include: The voice source determination module is used to determine whether the received user wake-up voice comes from a multi-channel voice device; The first noise reduction module is used to: if yes, then perform noise reduction on the first target speech based on the noise suppression degree to obtain a first noise-reduced speech, including: The noise suppression degree includes a high suppression degree and a low suppression degree. Based on different suppression degrees, a corresponding noise reduction algorithm is used to reduce noise on the first target speech to obtain a first noise-reduced speech; The second noise reduction module is used to: if not, then perform noise reduction on the second target speech based on the signal-to-noise ratio to obtain a second noise-reduced speech, including: Acquire a second signal-to-noise ratio of the second target speech; If the second signal-to-noise ratio is less than the signal-to-noise ratio threshold, performing single-channel noise reduction on the second target speech to obtain a second noise-reduced speech; The wake-up device identification module is used to identify the nearby wake-up device according to the first noise-reduced voice and / or the second noise-reduced voice, including: If only the user wake-up voice is received from the multi-channel voice device, the voice of the first noise-reduced voice in the preset frequency band is selected to obtain the first voice in the preset frequency band, the energy of the first voice in the preset frequency band is calculated to obtain the first voice energy, and the voice device corresponding to the maximum value of the first voice energy is determined as the nearest wake-up device; If only the user wake-up voice is received from a single-channel voice device, the voice of the second noise-reduced voice in the preset frequency band is selected to obtain the second voice in the preset frequency band, and the energy of the second voice in the preset frequency band is calculated to obtain the second voice energy, and the voice device corresponding to the maximum value of the second voice energy is determined as the nearest wake-up device; If a user wake-up voice is received from a multi-channel voice device and a single-channel voice device at the same time, the first noise-reduced voice and the voice of the first noise-reduced voice in a preset frequency band are selected to obtain a first voice in the preset frequency band and a second voice in the preset frequency band, energy calculation is performed on the first voice in the preset frequency band and the second voice in the preset frequency band to obtain a first voice energy and a second voice energy, and a voice device corresponding to the maximum value of the first voice energy and the second voice energy is determined as the nearest wake-up device; The first target voice is a voice obtained by processing a user wake-up voice received by a multi-channel voice device, and the second target voice is a voice obtained by processing the user wake-up voice received by a single-channel voice device.

10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the method for identifying a nearby wake-up device as described in any one of claims 1 to 8 is implemented.

11. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for identifying a nearby wake-up device as described in any one of claims 1 to 8 is implemented.

12. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the method for identifying a nearby wake-up device as described in any one of claims 1 to 8 is implemented.

Citation Information

Patent Citations

  • Wake-up control method and device and computer readable storage medium

    CN112037787A

  • Voice equipment awakening method and device, electronic equipment, storage medium and chip

    CN115966204A