Intelligent device screening method and apparatus, storage medium, and electronic device

By detecting and identifying the target speech parameters of audio data packets, and using coherence and attenuation functions to calculate the speech parameters, the problem of low wake-up accuracy of smart devices is solved, and high-accuracy wake-up is achieved in complex environments.

CN116504242BActive Publication Date: 2026-05-19HAIER YOUJIA INTELLIGENT TECH (BEIJING) CO LTD +2
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HAIER YOUJIA INTELLIGENT TECH (BEIJING) CO LTD
Filing Date
2023-03-31
Publication Date
2026-05-19

AI Technical Summary

Technical Problem

In existing technologies, the accuracy of waking up smart devices is relatively low, especially in environments with high reverberation or background noise, where the signal-to-noise ratio of the voice signal decreases, resulting in insufficient discrimination accuracy.

Method used

By acquiring audio data packets collected from multiple smart devices in the target scene, the target voice parameters are detected, the data information of the audio data packets is identified, and the target smart devices that respond to the wake-up command are selected based on the voice parameters. The voice parameters are calculated using coherence and attenuation functions to improve the signal-to-noise ratio and thus improve the wake-up accuracy.

Benefits of technology

It improves the accuracy of waking up smart devices, enhances the signal discrimination capability in complex environments, and ensures that the target smart device can correctly respond to the user's wake-up command.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116504242B_ABST
    Figure CN116504242B_ABST
Patent Text Reader

Abstract

The application discloses a screening method and device of an intelligent device, a storage medium and an electronic device, relates to the technical field of smart homes, and the screening method of the intelligent device comprises the following steps: acquiring audio data packets collected by each intelligent device in a plurality of intelligent devices in a target scene, wherein the audio data packets carry wake-up instructions for waking up the plurality of intelligent devices; detecting target voice parameters of the audio data packets, wherein the target voice parameters are used for indicating the probability that each audio data in the audio data packets belongs to voice data; identifying data information corresponding to the audio data packets according to the target voice parameters; determining voice parameters of voice carried in the audio data packets according to the data information; and screening a target intelligent device from the plurality of intelligent devices according to the voice parameters, wherein the target intelligent device is used for responding to the wake-up instructions. By adopting the technical scheme, the problems, such as low accuracy of intelligent device wake-up in related technologies, are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of smart home technology, and more specifically, to a method and apparatus for screening smart devices, a storage medium, and an electronic device. Background Technology

[0002] With the promotion of the smart home concept, the scenarios that enable voice interaction are becoming increasingly diverse. When a user speaks a wake-up word in a scenario with multiple smart devices, multiple smart devices equipped with voice wake-up modules may respond simultaneously. However, in reality, only one smart device needs to be selected to respond in order to achieve voice interaction between the user and the smart device. Currently, a distributed wake-up method is generally used to select the target smart device to respond to the user. When multiple smart devices are woken up at the same time, the smart devices will first upload the feature values ​​of the voice signals obtained by their own voice wake-up modules to the cloud. The cloud will then determine the target smart device according to the discrimination criteria and send the discrimination results to each smart device.

[0003] Most distributed wake-up criteria are based on energy methods, which determine the accuracy of speech signals received by the voice wake-up module deployed on the smart device. However, using the energy of the speech signal as a criterion may be affected by a variety of factors. For example, when there is a lot of reverberation or background noise in the environment, the signal-to-noise ratio of the speech signal received by the microphone device will decrease, and the accuracy of using the energy of the speech signal as a feature will decrease.

[0004] There is still no effective solution to the problem of low accuracy in waking up smart devices in related technologies. Summary of the Invention

[0005] This application provides a method and apparatus for screening smart devices, a storage medium, and an electronic device to at least solve the problem of low accuracy in waking up smart devices in related technologies.

[0006] According to one embodiment of this application, a method for screening smart devices is provided, including:

[0007] Acquire audio data packets collected by each of multiple smart devices in a target scene, wherein the audio data packets carry a wake-up command for waking up the multiple smart devices;

[0008] Detect the target speech parameters of the audio data packet, wherein the target speech parameters are used to indicate the probability that each audio data in the audio data packet belongs to speech data;

[0009] The data information corresponding to the audio data packet is identified based on the target speech parameters, wherein the data information is used to indicate the data attributes of the speech data packet present in the audio data packet;

[0010] The speech parameters of the speech carried in the audio data packet are determined based on the data information.

[0011] A target smart device is selected from the plurality of smart devices based on the voice parameters, wherein the target smart device is used to respond to the wake-up command.

[0012] Optionally, in one exemplary embodiment, detecting the target speech parameters of the audio data packet includes:

[0013] Obtain first audio data collected by the first voice acquisition device and second audio data collected by the second voice acquisition device deployed on each smart device from the audio data packet;

[0014] Calculate the coherence function between the first audio data and the second audio data at each time-frequency point, wherein the coherence function is used to indicate the frequency domain correlation between the first audio data and the second audio data at each time-frequency point;

[0015] The target ratio of the audio data packet relative to the scattered noise field of the target scene at each time frequency point is calculated based on the coherence function and the attenuation function corresponding to the target scene, to obtain a ratio sequence, wherein the attenuation function is used to indicate the degree of attenuation of audio in the target scene;

[0016] The ratio sequence is converted into a speech parameter sequence based on the reverberation parameters of the smart device to obtain the target speech parameters, wherein the speech parameter corresponding to each time frequency point in the speech parameter sequence is positively correlated with the probability of speech existing at each time frequency point.

[0017] Optionally, in an exemplary embodiment, identifying the data information corresponding to the audio data packet based on the target speech parameters includes:

[0018] Construct target audio data using the audio data packet and the target speech parameters;

[0019] The target audio data is input into the target recognition model, and the amplitude ratio output by the target recognition model is obtained as the data information. The amplitude ratio is used to indicate the proportional relationship between the speech amplitude of the speech signal formed by the speech data packet and the audio amplitude of the audio signal formed by the audio data packet. The target recognition model is obtained by training an initial recognition model using audio data samples labeled with amplitude ratio.

[0020] Optionally, in one exemplary embodiment, constructing the target audio data using the audio data packet and the target speech parameters includes:

[0021] Obtain the first frequency domain data of the first audio data collected by the first voice acquisition device deployed on each smart device in the audio data packet, and the second frequency domain data of the second audio data collected by the second voice acquisition device.

[0022] The first frequency domain data, the second frequency domain data, and the target speech parameters are concatenated to form the target audio data.

[0023] Optionally, in an exemplary embodiment, the step of inputting the target audio data into a target recognition model and obtaining the amplitude ratio output by the target recognition model as the data information includes:

[0024] The target audio data is input into the convolutional layer of the target recognition model to obtain the initial data features output by the convolutional layer;

[0025] The initial data features are input into the long short-term memory layer included in the target recognition model to obtain the target data features output by the long short-term memory layer;

[0026] The target data features are input into the fully connected layer of the target recognition model to obtain the amplitude ratio output by the fully connected layer.

[0027] Optionally, in an exemplary embodiment, determining the speech parameters of the speech carried in the audio data packet based on the data information includes:

[0028] Determine the signal amplitude of the audio signal formed in the audio data packet;

[0029] The product of the amplitude ratio included in the data information and the signal amplitude is determined as the speech parameter, wherein the amplitude ratio is used to indicate the proportional relationship between the speech amplitude of the speech signal formed by the speech data packet and the audio amplitude of the audio signal formed by the audio data packet.

[0030] Optionally, in an exemplary embodiment, the step of filtering target smart devices from the plurality of smart devices based on the voice parameters includes:

[0031] The speech parameters are converted into energy parameters, wherein the energy parameters are used to indicate the speech energy of the speech included in the audio data packet;

[0032] The smart device corresponding to the audio data packet with the largest energy parameter among the plurality of smart devices is identified as the target smart device.

[0033] According to another embodiment of the present application, a screening device for smart devices is also provided, comprising:

[0034] The acquisition module is used to acquire audio data packets collected by each of the multiple smart devices in the target scene, wherein the audio data packets carry a wake-up command for waking up the multiple smart devices;

[0035] A detection module is used to detect the target speech parameters of the audio data packet, wherein the target speech parameters are used to indicate the probability that each audio data in the audio data packet belongs to speech data;

[0036] The recognition module is used to recognize the data information corresponding to the audio data packet according to the target speech parameters, wherein the data information is used to indicate the data attributes of the speech data packet present in the audio data packet;

[0037] The determining module is used to determine the speech parameters of the speech carried in the audio data packet based on the data information;

[0038] A filtering module is used to filter target smart devices from the plurality of smart devices according to the voice parameters, wherein the target smart device is used to respond to the wake-up command.

[0039] According to another aspect of the embodiments of this application, a computer-readable storage medium is also provided, wherein a computer program is stored in the computer-readable storage medium, and the computer program is configured to execute the above-described smart device screening method when running.

[0040] According to another aspect of the embodiments of this application, an electronic device is also provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the above-described intelligent device selection method through the computer program.

[0041] In this embodiment, audio data packets carrying wake-up commands are acquired from each of multiple smart devices in a target scene. These wake-up commands are used to wake up the multiple smart devices in the target scene. A target speech parameter is detected, indicating the probability that each audio data packet in the audio data packet belongs to speech data. Based on the target speech parameter, data information corresponding to the audio data packet, indicating the data attributes of the speech data packets present in the audio data packet, is identified. Based on the data information, the speech parameters of the speech carried in the audio data packet are determined. Based on the speech parameters, the target smart device responding to the wake-up command is selected from the multiple smart devices. In other words, the probability that each audio data packet belongs to speech data is detected from the acquired audio data packet as the target speech parameter. The target speech parameter is used to identify the data attributes of the speech data packets present in the audio data packet, i.e., the data information corresponding to the audio data packet. The speech parameters of the speech carried in the audio data packet are determined through the data information, so that the speech parameters used to select the target smart device focus on the speech data packet portion of the audio data packet, improving the signal-to-noise ratio of the signal in the audio data packet, thereby improving the accuracy of distributed discrimination in determining the target smart device responding to the user. By adopting the above technical solution, the problem of low accuracy in waking up smart devices in related technologies has been solved, and the technical effect of improving the accuracy of waking up smart devices has been achieved. Attached Figure Description

[0042] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0043] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0044] Figure 1 This is a schematic diagram of the hardware environment for a smart device screening method according to an embodiment of this application;

[0045] Figure 2 This is a flowchart of a smart device screening method according to an embodiment of this application;

[0046] Figure 3 This is an example diagram illustrating the detection of target speech parameters in audio data packets according to embodiments of this application;

[0047] Figure 4 This is an example diagram illustrating the data information corresponding to the audio data packets identified according to an embodiment of this application;

[0048] Figure 5 This is a structural block diagram of a screening device for a smart device according to an embodiment of this application. Detailed Implementation

[0049] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.

[0050] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0051] According to one aspect of the embodiments of this application, a method for selecting smart devices is provided. This method is widely applicable to whole-house intelligent digital control application scenarios such as smart homes, smart home ecosystems, and intelligence house ecosystems. Optionally, in this embodiment, the above-mentioned method for selecting smart devices can be applied to, for example... Figure 1 The hardware environment shown consists of terminal device 102 and server 104. For example... Figure 1 As shown, server 104 is connected to terminal device 102 via a network and can be used to provide services (such as application services) to the terminal or clients installed on the terminal. A database can be set up on the server or independently of the server to provide data storage services for server 104. Cloud computing and / or edge computing services can be configured on the server or independently of the server to provide data processing services for server 104.

[0052] The aforementioned network may include, but is not limited to, at least one of the following: wired network, wireless network. The aforementioned wired network may include, but is not limited to, at least one of the following: wide area network, metropolitan area network, local area network. The aforementioned wireless network may include, but is not limited to, at least one of the following: Wi-Fi (Wireless Fidelity), Bluetooth. The terminal device 102 may not be limited to PC, mobile phone, tablet computer, smart air conditioner, smart range hood, smart refrigerator, smart oven, smart stove, smart washing machine, smart water heater, smart washing equipment, smart dishwasher, smart projector, smart TV, smart clothes rack, smart curtains, smart audio-visual equipment, smart socket, smart speaker, smart speaker box, smart fresh air equipment, smart kitchen and bathroom equipment, smart bathroom equipment, smart robot vacuum cleaner, smart window cleaning robot, smart mopping robot, smart air purifier, smart steam oven, smart microwave oven, smart water heater, smart air purifier, smart water dispenser, smart door lock, etc.

[0053] This embodiment provides a method for screening smart devices, applied to the aforementioned device terminals. Figure 2 This is a flowchart of a smart device screening method according to an embodiment of this application, the process including the following steps:

[0054] Step S202: Obtain audio data packets collected by each of the multiple smart devices in the target scene, wherein the audio data packets carry a wake-up command for waking up the multiple smart devices;

[0055] Step S204: Detect the target speech parameter of the audio data packet, wherein the target speech parameter is used to indicate the probability that each audio data in the audio data packet belongs to speech data;

[0056] Step S206: Identify the data information corresponding to the audio data packet according to the target speech parameters, wherein the data information is used to indicate the data attributes of the speech data packet present in the audio data packet;

[0057] Step S208: Determine the speech parameters of the speech carried in the audio data packet based on the data information;

[0058] Step S210: Select a target smart device from the plurality of smart devices according to the voice parameters, wherein the target smart device is used to respond to the wake-up command.

[0059] Through the above steps, the probability of each audio data point belonging to speech data is detected from the collected audio data packets and used as the target speech parameter. This target speech parameter is then used to identify the data attributes of the speech data points within the audio data packets, i.e., the corresponding data information. The speech parameters of the speech carried in the audio data packets are determined based on this data information. This ensures that the speech parameters used to filter target smart devices focus on the speech data portion of the audio data packets, improving the signal-to-noise ratio of the audio data packets and thus enhancing the accuracy of distributed discrimination in determining the target smart device for responding to users. This technical solution solves the problem of low accuracy in smart device wake-up in related technologies, achieving the technical effect of improving the accuracy of smart device wake-up.

[0060] Optionally, in this embodiment, the aforementioned smart device may be, but is not limited to, a device with voice acquisition function and capable of voice interaction. For example, the smart device can acquire nearby sound information in real time, and when it recognizes that the sound information contains a wake-up command and the smart device is identified as the target smart device, the smart device may, but is not limited to, respond to the wake-up command.

[0061] Optionally, in this embodiment, the aforementioned smart device may include, but is not limited to, devices that can perform corresponding operations according to instructions, and may include, but is not limited to, smart home devices, smart vehicle devices, etc., with network connectivity. The aforementioned smart device may be, but is not limited to, a device capable of querying and updating data through a network connection and adjusting its own state.

[0062] For example, smart home devices can include, but are not limited to, smart air conditioners, smart range hoods, smart refrigerators, smart ovens, smart stoves, smart washing machines, smart water heaters, smart laundry equipment, smart dishwashers, smart projectors, smart TVs, smart clothes racks, smart curtains, smart sockets, smart speakers, smart loudspeakers, smart ventilation systems, smart kitchen and bathroom equipment, smart bathroom fixtures, smart robot vacuums, smart window cleaning robots, smart mopping robots, smart air purifiers, smart steam ovens, smart microwave ovens, smart water heaters, smart air purifiers, smart water dispensers, smart door locks, etc. Smart in-vehicle devices can include, but are not limited to, smart car air conditioners, smart windshield wipers, smart car speakers, smart car refrigerators, etc.

[0063] In the technical solution provided in step S202 above, the wake-up command can be used, but is not limited to, to control the smart devices deployed in the target scene. It can be, but is not limited to, the device that the user directly wants to control, or a statement from the user expressing their current state. For example, the wake-up command can be, but is not limited to, "turn on the light," which directly expresses that the user wants the light device to be turned on. Or the wake-up command can be, but is not limited to, "I'm home," which expresses that the user wants to control the devices in their home to achieve the state of "coming home," and so on.

[0064] Optionally, in this embodiment, the target scenario may include, but is not limited to, any type of indoor or outdoor space, such as: rooms, warehouses, factories, offices, parking lots, etc.

[0065] Optionally, in this embodiment, the smart device may be used, but is not limited to, to collect audio data packets. For example, taking a target scene that includes multiple smart devices as an example, each smart device in the target scene may, but is not limited to, have its corresponding audio data packets.

[0066] Optionally, in this embodiment, the aforementioned smart device may, but is not limited to, collect audio data through a voice acquisition device, which may, but is not limited to, a device deployed on the smart device, such as a microphone deployed on the smart device.

[0067] Optionally, in this embodiment, the aforementioned audio data packet may include, but is not limited to, audio data collected by the smart device through the voice acquisition device. For example, taking a smart device A in the target scenario, and a voice acquisition device A and a voice acquisition device B deployed on the smart device A as an example, the audio data packet may include, but is not limited to, audio data A collected by the voice acquisition device A deployed on the smart device A and audio data B collected by the voice acquisition device B.

[0068] Optionally, in this embodiment, the audio data packet can be used, but is not limited to, to wake up multiple smart devices in the target scene, and the audio data packet can include, but is not limited to, wake-up commands, noise in the target scene, etc.

[0069] In the technical solution provided in step S204 above, the target speech parameters of the audio data packet can be obtained by determining, but is not limited to, the probability that each audio data in the audio data packet belongs to speech data.

[0070] Optionally, in this embodiment, the target speech parameters of the audio data packet can be determined based on, but not limited to, the correlation between each audio data in the audio data packet and the scattered noise field in the target scene. For example, the coherence function between audio data can be determined based on the cross spectrum and self spectrum between each audio data to obtain the correlation between the audio data; the target speech parameters of the audio data packet can be calculated based on, but not limited to, the coherence function between the scattered noise field in the target scene and the audio data.

[0071] In one exemplary embodiment, the target speech parameters of the audio data packet can be detected, but is not limited to, by: obtaining first audio data collected by a first speech acquisition device and second audio data collected by a second speech acquisition device deployed on each smart device from the audio data packet; calculating a coherence function between the first audio data and the second audio data at each time-frequency point, wherein the coherence function is used to indicate the frequency domain correlation between the first audio data and the second audio data at each time-frequency point; calculating a target ratio of the audio data packet relative to the scattered noise field of the target scene at each time-frequency point based on the coherence function and the attenuation function corresponding to the target scene, to obtain a ratio sequence, wherein the attenuation function is used to indicate the degree of attenuation of audio in the target scene; converting the ratio sequence into a speech parameter sequence based on the reverberation parameters of the smart device to obtain the target speech parameters, wherein the speech parameter corresponding to each time-frequency point in the speech parameter sequence is positively correlated with the probability of speech existing at each time-frequency point.

[0072] Optionally, in this embodiment, the aforementioned voice acquisition device may be, but is not limited to, a device capable of acquiring audio data, or the aforementioned voice acquisition device may be, but is not limited to, a device capable of converting acquired sound information into audio data, such as a microphone deployed on a smart device, etc.

[0073] Optionally, in this embodiment, the aforementioned smart device may, but is not limited to, deploy multiple voice acquisition devices. Any two voice acquisition devices deployed on the smart device may, but are not limited to, serve as the first voice acquisition device and the second voice device. For example, taking voice acquisition device A, voice acquisition device B, and voice acquisition device C deployed on the smart device as an example, voice acquisition device A and voice acquisition device B may, but are not limited to, serve as the first voice acquisition device and the second voice device. Alternatively, voice acquisition device B and voice acquisition device C may also be used as the first voice acquisition device and the second voice device. Alternatively, voice acquisition device A and voice acquisition device C may also be used as the first voice acquisition device and the second voice device.

[0074] Optionally, in this embodiment, the audio data collected by the aforementioned voice acquisition device may include, but is not limited to, data collected by the voice acquisition device at multiple time and frequency points.

[0075] Optionally, in this embodiment, the coherence function between the first audio data and the second audio data at each time-frequency point can be calculated based on the cross spectrum and autospectrum of the first audio data and the second audio data, for example, but not limited to, using a formula. Calculate the coherence function between the first audio data and the second audio data at each time-frequency point, where P xy (l, f) represents the cross spectrum of the first audio data and the second audio data, P x (l, f) and P y (l, f) represent the autospectral signatures of the first and second audio data, respectively.

[0076] Optionally, in this embodiment, the attenuation function corresponding to the target scene can be, but is not limited to, used to indicate the scattered noise field of the target scene. It can be, but is not limited to, determined by the distance between the first voice acquisition device and the second voice device in the target scene. For example, taking the distance between the first voice acquisition device and the second voice device as d, the attenuation function corresponding to the target scene can be, but is not limited to, R. n (f) = sinc(2fd / c), where c can be, but is not limited to, the speed of sound.

[0077] Optionally, in this embodiment, the target ratio can be, but is not limited to, the ratio of audio data to the scattered noise field, to obtain a ratio sequence, for example, it can be, but is not limited to, based on the formula Determine the ratio of the audio data to the scattered noise field, where R e Used to indicate taking real numbers, R n R is used to indicate the decay function corresponding to the target scene. x This is used to indicate the coherence function between audio data at each time-frequency point, resulting in a ratio sequence.

[0078] Optionally, in this embodiment, the reverberation parameter of the aforementioned smart device can be used, but is not limited to, to indicate the effect of reverberation on audio data in the target scene. It can be used, but is not limited to, to adjust the target speech parameter of the audio data packet through the reverberation parameter. For example, when the effect of reverberation in the audio data is relatively small, the reverberation parameter of the smart device can be, but is not limited to, close to 1, indicating that not much gain is needed, and the target speech parameter of the audio data packet can be, but is not limited to, reduced. Or, when the effect of reverberation in the audio data is relatively large, the reverberation parameter of the smart device can be, but is not limited to, greater than 1, indicating that more gain is needed, and the target speech parameter of the audio data packet can be, but is not limited to, increased to weaken the effect of reverberation on the audio data packet.

[0079] In one exemplary embodiment, an example of detecting target speech parameters of an audio data packet is provided. Figure 3 This is an example diagram illustrating the detection of target speech parameters in audio data packets according to embodiments of this application, such as... Figure 3 As shown, taking two microphones deployed on the same device as a first voice acquisition device and a second voice acquisition device, with the first voice acquisition device used to acquire first audio data x(t) and the second voice acquisition device used to acquire second audio data y(t) as an example, the target voice parameters of the audio data packet can be detected through, but is not limited to, the following steps:

[0080] S302: Perform STFT (Short-time Fourier Transform) operation on the first audio data x(t) and the second audio data y(t) to obtain the frequency domain data X(l, f) of the first audio data x(t) and the frequency domain data Y(l, f) of the second audio data y(t), where l can be, but is not limited to, a frame index, and f can be, but is not limited to, a frequency band;

[0081] S304: Calculate the cross spectrum and autospectrum of the first audio data and the second audio data, where α can be, but is not limited to, a smoothing factor:

[0082] P xy (l, f) = αP xy (l, f)+(1-α)X(l, f)Y*(l, f);

[0083] P x (l, f) = αP x (l, f)+(1-α)X(l, f)X*(l, f);

[0084] P y (l, f) = αP y (l, f)+(1-α)Y(l, f)Y*(l, f);

[0085] S306: Calculate the coherence function between the first audio data and the second audio data at each time-frequency point:

[0086]

[0087] S308: Based on the attenuation function R corresponding to the target scene n (f) = sinc(2fd / c) calculates the target ratio (which can be, but is not limited to, the ratio of the audio data packet to the scattered noise field at each time frequency point), resulting in a ratio sequence of audio data packets. Here, d in the attenuation function indicates the distance between the first and second voice acquisition devices, and c in the attenuation function indicates the speed of sound, thus obtaining the ratio sequence.

[0088] S310: Determine the target speech parameters of the audio data packet based on the reverberation parameter μ of the smart device.

[0089]

[0090] In the technical solution provided in step S206 above, the data attributes of the aforementioned voice data packet in the audio data packet can be, but are not limited to, used to indicate the proportional relationship between the amplitude of the voice data packet present in the audio data packet and the total amplitude of the audio data packet.

[0091] Optionally, in this embodiment, the aforementioned voice data packet may, but is not limited to, exist in the audio data packet. The data attributes of the voice data packet may, but are not limited to, be determined based on the proportion of the voice data packet in the audio data packet. For example, taking the noisy signal amplitude obtained from the audio data packet and the clean voice amplitude obtained from the voice data packet as an example, the ratio between the clean voice amplitude and the noisy signal amplitude may, but is not limited to, be determined as the data attribute of the voice data packet in the audio data packet.

[0092] In one exemplary embodiment, the data information corresponding to the audio data packet can be identified based on the target speech parameters in the following manner, but not limited to: constructing target audio data using the audio data packet and the target speech parameters; inputting the target audio data into a target recognition model to obtain the amplitude ratio output by the target recognition model as the data information, wherein the amplitude ratio is used to indicate the proportional relationship between the speech amplitude of the speech signal formed by the speech data packet and the audio amplitude of the audio signal formed by the audio data packet, and the target recognition model is obtained by training an initial recognition model using audio data samples labeled with amplitude ratio tags.

[0093] Optionally, in this embodiment, the target audio data may be, but is not limited to, a set of audio data packets and target speech parameters. For example, the target audio data may be, but is not limited to, (audio data packets, target speech parameters). Alternatively, the target audio data may be, but is not limited to, (target speech parameters, audio data packets).

[0094] Optionally, in this embodiment, the amplitude ratio can be determined based on, but is not limited to, the ratio between the voice amplitude of the voice signal formed by the voice data packet and the audio amplitude of the audio signal formed by the audio data packet.

[0095] Optionally, in this embodiment, the audio data samples labeled with amplitude ratios may be, but are not limited to, audio data that determines the ratio between the speech amplitude of the speech signal and the audio amplitude of the audio signal. The target recognition model may be trained by inputting the audio data that determines the ratio between the speech amplitude of the speech signal and the audio amplitude of the audio signal into the target recognition model.

[0096] Optionally, in this embodiment, audio data including audio data samples can be used as input to the initial recognition model to train the target recognition model, for example, the target speech parameters corresponding to the audio data samples can be used together with the audio data samples as input to the initial recognition model, and the target speech parameters of the audio data samples can help the model learn more historical information and spectral information.

[0097] Optionally, in this embodiment, the target speech parameter corresponding to the audio data sample can be, but is not limited to, used to indicate the weight of the audio data in the audio data sample. For example, when the probability of speech presence in the audio data sample is high, the target speech parameter corresponding to the time-frequency point in the audio data sample can be, but is not limited to, closer to 1. Alternatively, when noise or reverberation is high, the target speech parameter corresponding to the time-frequency point in the audio data sample can be, but is not limited to, smaller. The initial recognition model can be controlled to focus on learning information about time-frequency points with a high probability of speech presence by specifying the magnitude of the target speech parameter in the audio data sample learned by the initial recognition model.

[0098] In one exemplary embodiment, target audio data may be constructed using the audio data packet and the target speech parameters in the following manner, but not limited to: acquiring first frequency domain data of first audio data acquired by a first speech acquisition device deployed on each smart device and second frequency domain data of second audio data acquired by a second speech acquisition device in the audio data packet; and concatenating the first frequency domain data, the second frequency domain data, and the target speech parameters to form the target audio data.

[0099] Optionally, in this embodiment, the aforementioned voice acquisition device may be, but is not limited to, a device capable of acquiring audio data, or the aforementioned voice acquisition device may be, but is not limited to, a device capable of converting the acquired audio data into frequency domain data, such as a microphone deployed on a smart device, etc.

[0100] Optionally, in this embodiment, the aforementioned smart device may, but is not limited to, deploy multiple voice acquisition devices. Any two voice acquisition devices deployed on the smart device may, but are not limited to, serve as the first voice acquisition device and the second voice device. For example, taking voice acquisition device A, voice acquisition device B, and voice acquisition device C deployed on the smart device as an example, voice acquisition device A and voice acquisition device B may, but are not limited to, serve as the first voice acquisition device and the second voice device. Alternatively, voice acquisition device B and voice acquisition device C may also be used as the first voice acquisition device and the second voice device. Alternatively, voice acquisition device A and voice acquisition device C may also be used as the first voice acquisition device and the second voice device.

[0101] Optionally, in this embodiment, the frequency domain data mentioned above may be, but is not limited to, obtained by performing a short-time Fourier transform on the audio data acquired by the voice acquisition device. For example, taking the audio data x(t) acquired by the voice acquisition device as an example, the frequency domain data X(l, f) of the audio data x(t) may be obtained by performing a short-time Fourier transform on the audio data x(t), where l may be, but is not limited to, a frame index, and f may be, but is not limited to, a frequency band.

[0102] Optionally, in this embodiment, the first frequency domain data of the first audio data x(t) collected by the first voice acquisition device deployed on the smart device is X(l,f), the second frequency domain data of the second audio data y(t) collected by the second voice acquisition device is Y(l,f), and the target voice parameter is G(l,f). The modulus abs(X(l,f)) of the first frequency domain data X(l,f) and the modulus abs(Y(l,f)) of the second frequency domain data Y(l,f) are concatenated with the target voice parameter G(l,f) to obtain [abs(X(l,f)), abs(Y(l,f)), G(l,f)] as the target audio data, which is then input into the target recognition model.

[0103] In one exemplary embodiment, the target audio data may be input into a target recognition model in the following manner, but not limited to, to obtain the amplitude ratio output by the target recognition model as the data information: inputting the target audio data into a convolutional layer included in the target recognition model to obtain initial data features output by the convolutional layer; inputting the initial data features into a long short-term memory layer included in the target recognition model to obtain target data features output by the long short-term memory layer; and inputting the target data features into a fully connected layer included in the target recognition model to obtain the amplitude ratio output by the fully connected layer.

[0104] Optionally, in this embodiment, the target recognition model may include, but is not limited to, convolutional layers, LSTM (Long Short-Term Memory) layers, fully connected layers, etc. For example, it may, but is not limited to, inputting target audio data into the convolutional layer of the target recognition model to extract the first feature information of the target audio data, obtaining initial data features; it may, but is not limited to, inputting the initial data features into the LSTM layer of the target recognition model to further extract the second feature information of the initial data features, obtaining target data features; it may, but is not limited to, inputting the target data features into the fully connected layer of the target recognition model to map the obtained distributed feature representation with the audio data samples used during the training of the initial recognition model, obtaining the amplitude ratio output by the fully connected layer. Alternatively, the target recognition model may also include, but is not limited to, an activation layer, which may include, but is not limited to, an activation function, and may, but is not limited to, linearly mapping the output amplitude ratio through the activation layer of the target recognition model.

[0105] In one exemplary embodiment, an example is provided for identifying data information corresponding to audio data packets. Figure 4 This is an example diagram illustrating the data information corresponding to the identified audio data packets according to embodiments of this application, such as... Figure 4 As shown, taking the first audio data x(t) and the second audio data y(t) as examples, a short-time Fourier transform operation is performed to obtain a frame length l of 512 and an effective frequency band f of 257. The modulus abs(X(l,f)) of the first frequency domain data X(l,f) and the modulus abs(Y(l,f)) of the second frequency domain data Y(l,f) are concatenated with the target speech parameter G(l,f) to obtain [abs(X(l,f)), abs(Y(l,f)), G(l,f)] as the target audio data input to the target recognition model. The target recognition model can obtain the proportional relationship between the speech amplitude of the speech signal formed by the speech data packet and the audio amplitude of the audio signal formed by the audio data packet through the following methods, and obtain the data information corresponding to the audio data packet:

[0106] Based on the first audio data x(t) and the second audio data y(t), a short-time Fourier transform operation is performed to obtain a frame length l of 512 and an effective frequency band f of 257. The frequency domain data X(l, f) of the first audio data x(t) and the frequency domain data Y(l, f) of the second audio data y(t) are determined to be complex arrays of 257, and the target speech parameter G(l, f) is a real array of 257.

[0107] Input the target audio data [abs(X(l,f)), abs(Y(l,f)), G(l,f)] into the one-dimensional convolutional layer conv1d of the target recognition model to obtain the initial data features of the convolutional layer output: in_channels=257*3, out_channels=1024, kernel_size=4, stride=1, padding=1.

[0108] The initial data features are input into the LSTM layer of the target recognition model to obtain the target data features output by the LSTM layer: Input_size = 1024, hidden_size = 512, num_layers = 1.

[0109] The target data features are input into the activation layer and fully connected layer of the target recognition model, and the amplitude ratio of the fully connected layer output is obtained: Input_size = 512, hidden_size = 257.

[0110] Finally, the vector of the ratio (amplitude ratio) of the clean speech amplitude to the noisy signal amplitude output by the target recognition model is obtained as 257*1.

[0111] The vector of 257*1, which is the ratio of the clean speech amplitude to the noisy signal amplitude output by the above target recognition model, can be used as the data information corresponding to the audio data packet, but is not limited to this.

[0112] In the technical solution provided in step S208 above, the speech parameters of the speech carried in the audio data packet can be, but are not limited to, used to indicate the frequency domain value of the speech carried in the audio data packet.

[0113] In one exemplary embodiment, the speech parameters of the speech carried in the audio data packet can be determined based on the data information in the following manner, but not limited to: determining the signal amplitude of the audio signal formed in the audio data packet; and determining the speech parameters by multiplying the amplitude ratio included in the data information by the signal amplitude, wherein the amplitude ratio is used to indicate the proportional relationship between the speech amplitude of the speech signal formed by the speech data packet and the audio amplitude of the audio signal formed by the audio data packet.

[0114] Optionally, in this embodiment, the above-mentioned speech parameters may be used, but are not limited to, to estimate the clean signal frequency domain value of the speech carried in the audio data packet.

[0115] Optionally, in this embodiment, the above-mentioned signal amplitude can be used, but is not limited to, to indicate the signals included in the audio data packet.

[0116] Optionally, in this embodiment, the aforementioned amplitude ratio can be, but is not limited to, used to indicate the ratio between the signal amplitude of the speech signal and the signal amplitude of the audio signal including the speech signal. It can be, but is not limited to, calculating the speech parameters of the speech carried in the audio data packet based on the ratio between the signal amplitude of the speech signal and the signal amplitude of the audio signal including the speech signal. For example, taking the audio amplitude of the audio signal formed by the audio data packet as the signal amplitude of the noisy signal in the audio data packet, and the speech amplitude of the speech signal formed by the speech data packet as the signal amplitude of the clean speech in the audio data packet as an example, it can be, but is not limited to, using the ratio between the signal amplitude of the clean speech output by the target recognition model and the signal amplitude of the noisy signal as the data information corresponding to the audio data packet. It can be, but is not limited to, using the product of the data information output by the target recognition model and the signal amplitude of the noisy signal in the audio data packet as the speech parameters of the speech carried in the audio data packet.

[0117] In the technical solution provided in step S210 above, the target smart device can be determined from multiple smart devices based on the voice parameters of the smart device. For example, taking the voice parameters of the smart device as an example that can be used to indicate the energy of the audio data collected by the smart device, the smart device corresponding to the audio data with the largest energy can be determined as the target smart device based on the energy of the audio data.

[0118] In one exemplary embodiment, a target smart device may be selected from the plurality of smart devices based on the voice parameters in the following manner, but not limited to: converting the voice parameters into energy parameters, wherein the energy parameters are used to indicate the voice energy of the voice included in the audio data packet; and determining the smart device corresponding to the audio data packet with the largest corresponding energy parameter among the plurality of smart devices as the target smart device.

[0119] Optionally, in this embodiment, the energy level of the audio data collected by the smart device can be determined based on the energy parameters obtained from the conversion of voice parameters, but is not limited to the smart device corresponding to the audio data with the largest energy parameter being identified as the target smart device.

[0120] Optionally, in this embodiment, taking the use of energy as a discrimination criterion and the distributed wake-up method to determine the target smart device from the environment as an example, since the sensitivity of each voice acquisition device deployed on multiple smart devices in the target scene is inconsistent, it is possible, but not limited to, selecting one voice acquisition device deployed on one smart device in the target scene as a standard, and performing coefficient calibration operations on other voice acquisition devices.

[0121] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods of the various embodiments of this application.

[0122] Figure 5 This is a structural block diagram of a screening device for a smart device according to an embodiment of this application; as shown below. Figure 5 As shown, it includes:

[0123] The acquisition module 502 is used to acquire audio data packets collected by each of the multiple smart devices in the target scene, wherein the audio data packets carry a wake-up command for waking up the multiple smart devices;

[0124] The detection module 504 is used to detect the target speech parameters of the audio data packet, wherein the target speech parameters are used to indicate the probability that each audio data in the audio data packet belongs to speech data;

[0125] The recognition module 506 is used to recognize the data information corresponding to the audio data packet according to the target speech parameters, wherein the data information is used to indicate the data attributes of the speech data packet present in the audio data packet;

[0126] The determining module 508 is used to determine the speech parameters of the speech carried in the audio data packet based on the data information.

[0127] The filtering module 510 is used to filter target smart devices from the plurality of smart devices according to the voice parameters, wherein the target smart device is used to respond to the wake-up command.

[0128] Through the above embodiments, the probability of each audio data point belonging to speech data is detected from the collected audio data packets and used as the target speech parameter. The target speech parameter is then used to identify the data attributes of the speech data points within the audio data packets, i.e., the corresponding data information. This data information is used to determine the speech parameters carried in the audio data packets. This ensures that the speech parameters used to filter target smart devices focus on the speech data portion of the audio data packets, improving the signal-to-noise ratio of the audio data packets and thus enhancing the accuracy of distributed discrimination in determining the target smart device for responding to users. By adopting the above technical solution, the problem of low accuracy in smart device wake-up in related technologies is solved, achieving the technical effect of improving the accuracy of smart device wake-up.

[0129] In one exemplary embodiment, the detection module includes:

[0130] The acquisition unit is used to acquire, from the audio data packet, first audio data acquired by the first voice acquisition device deployed on each smart device and second audio data acquired by the second voice acquisition device;

[0131] The first calculation unit is used to calculate the coherence function between the first audio data and the second audio data at each time-frequency point, wherein the coherence function is used to indicate the frequency domain correlation between the first audio data and the second audio data at each time-frequency point;

[0132] The second calculation unit is used to calculate the target ratio of the audio data packet relative to the scattered noise field of the target scene at each time frequency point according to the coherence function and the attenuation function corresponding to the target scene, and obtain a ratio sequence, wherein the attenuation function is used to indicate the degree of attenuation of audio in the target scene;

[0133] The first conversion unit is used to convert the ratio sequence into a speech parameter sequence based on the reverberation parameters of the smart device to obtain the target speech parameters, wherein the speech parameter corresponding to each time frequency point in the speech parameter sequence is positively correlated with the probability of speech existing at each time frequency point.

[0134] In one exemplary embodiment, the identification module includes:

[0135] A construction unit is used to construct target audio data using the audio data packet and the target speech parameters;

[0136] The processing unit is used to input the target audio data into the target recognition model and obtain the amplitude ratio output by the target recognition model as the data information. The amplitude ratio is used to indicate the proportional relationship between the speech amplitude of the speech signal formed by the speech data packet and the audio amplitude of the audio signal formed by the audio data packet. The target recognition model is obtained by training an initial recognition model using audio data samples labeled with amplitude ratio.

[0137] In an exemplary embodiment, the construction unit is configured to: acquire first frequency domain data of first audio data acquired by a first voice acquisition device deployed on each smart device and second frequency domain data of second audio data acquired by a second voice acquisition device in the audio data packet; and concatenate the first frequency domain data, the second frequency domain data and the target voice parameters to form the target audio data.

[0138] In an exemplary embodiment, the processing unit is configured to: input the target audio data into a convolutional layer included in the target recognition model to obtain initial data features output by the convolutional layer; input the initial data features into a long short-term memory layer included in the target recognition model to obtain target data features output by the long short-term memory layer; and input the target data features into a fully connected layer included in the target recognition model to obtain the amplitude ratio output by the fully connected layer.

[0139] In one exemplary embodiment, the determining module includes:

[0140] The first determining unit is used to determine the signal amplitude of the audio signal formed in the audio data packet;

[0141] The second determining unit is used to determine the product of the amplitude ratio included in the data information and the signal amplitude as the speech parameter, wherein the amplitude ratio is used to indicate the proportional relationship between the speech amplitude of the speech signal formed by the speech data packet and the audio amplitude of the audio signal formed by the audio data packet.

[0142] In one exemplary embodiment, the filtering module includes:

[0143] The second conversion unit is used to convert the speech parameters into energy parameters, wherein the energy parameters are used to indicate the speech energy of the speech included in the audio data packet;

[0144] The third determining unit is used to determine the smart device corresponding to the audio data packet with the largest energy parameter among the plurality of smart devices as the target smart device.

[0145] Embodiments of this application also provide a storage medium including a stored program, wherein the program executes any of the methods described above when it is run.

[0146] Optionally, in this embodiment, the storage medium may be configured to store program code for performing the following steps:

[0147] S1, acquire audio data packets collected by each of the multiple smart devices in the target scene, wherein the audio data packets carry a wake-up command for waking up the multiple smart devices;

[0148] S2, Detect the target speech parameter of the audio data packet, wherein the target speech parameter is used to indicate the probability that each audio data in the audio data packet belongs to speech data;

[0149] S3, Identify the data information corresponding to the audio data packet according to the target speech parameters, wherein the data information is used to indicate the data attributes of the speech data packet present in the audio data packet;

[0150] S4, determine the speech parameters of the speech carried in the audio data packet based on the data information;

[0151] S5, select a target smart device from the plurality of smart devices according to the voice parameters, wherein the target smart device is used to respond to the wake-up command.

[0152] Embodiments of this application also provide an electronic device including a memory and a processor, wherein the memory stores a computer program and the processor is configured to run the computer program to perform the steps in any of the above method embodiments.

[0153] Optionally, the electronic device may further include a transmission device and an input / output device, wherein the transmission device is connected to the processor and the input / output device is connected to the processor.

[0154] Optionally, in this embodiment, the processor can be configured to perform the following steps via a computer program:

[0155] S1, acquire audio data packets collected by each of the multiple smart devices in the target scene, wherein the audio data packets carry a wake-up command for waking up the multiple smart devices;

[0156] S2, Detect the target speech parameter of the audio data packet, wherein the target speech parameter is used to indicate the probability that each audio data in the audio data packet belongs to speech data;

[0157] S3, Identify the data information corresponding to the audio data packet according to the target speech parameters, wherein the data information is used to indicate the data attributes of the speech data packet present in the audio data packet;

[0158] S4, determine the speech parameters of the speech carried in the audio data packet based on the data information;

[0159] S5, select a target smart device from the plurality of smart devices according to the voice parameters, wherein the target smart device is used to respond to the wake-up command.

[0160] Optionally, in this embodiment, the storage medium may include, but is not limited to, various media capable of storing program code, such as USB flash drives, read-only memory (ROM), random access memory (RAM), portable hard drives, magnetic disks, or optical disks.

[0161] Optionally, specific examples in this embodiment can refer to the examples described in the above embodiments and optional implementations, and will not be repeated here.

[0162] Obviously, those skilled in the art should understand that the modules or steps of this application described above can be implemented using general-purpose computing devices. They can be centralized on a single computing device or distributed across a network of multiple computing devices. Optionally, they can be implemented using computer-executable program code, thereby storing them in a storage device for execution by a computing device. In some cases, the steps shown or described can be performed in a different order than those presented here, or they can be fabricated as separate integrated circuit modules, or multiple modules or steps can be fabricated as a single integrated circuit module. Thus, this application is not limited to any particular combination of hardware and software.

[0163] The above description is only a preferred embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this application, and these improvements and modifications should also be considered within the scope of protection of this application.

Claims

1. A method for screening intelligent devices, characterized in that, include: Acquire audio data packets collected by each of the multiple smart devices in the target scene, wherein the audio data packets carry a wake-up command for waking up the multiple smart devices; Detect the target speech parameters of the audio data packet, wherein the target speech parameters are used to indicate the probability that each audio data in the audio data packet belongs to speech data; The data information corresponding to the audio data packet is identified according to the target speech parameters, wherein the data information is used to indicate the data attributes of the speech data packet in the audio data packet, and the data attributes include: the ratio between the clean speech amplitude and the noisy signal amplitude, wherein the clean speech amplitude is obtained from the speech data packet, and the noisy signal amplitude is obtained from the audio data packet; The speech parameters of the speech carried in the audio data packet are determined based on the data information. A target smart device is selected from the plurality of smart devices based on the voice parameters, wherein the target smart device is used to respond to the wake-up command; The step of determining the speech parameters of the speech carried in the audio data packet based on the data information includes: determining the signal amplitude of the audio signal formed in the audio data packet; and determining the speech parameter by multiplying the amplitude ratio included in the data information by the signal amplitude, wherein the amplitude ratio is used to indicate the proportional relationship between the speech amplitude of the speech signal formed by the speech data packet and the audio amplitude of the audio signal formed by the audio data packet.

2. The method according to claim 1, characterized in that, The detection of the target speech parameters of the audio data packet includes: Obtain first audio data collected by the first voice acquisition device and second audio data collected by the second voice acquisition device deployed on each smart device from the audio data packet; Calculate the coherence function between the first audio data and the second audio data at each time-frequency point, wherein the coherence function is used to indicate the frequency domain correlation between the first audio data and the second audio data at each time-frequency point; The target ratio of the audio data packet relative to the scattered noise field of the target scene at each time frequency point is calculated based on the coherence function and the attenuation function corresponding to the target scene, to obtain a ratio sequence, wherein the attenuation function is used to indicate the degree of attenuation of audio in the target scene; The ratio sequence is converted into a speech parameter sequence based on the reverberation parameters of the smart device to obtain the target speech parameters, wherein the speech parameter corresponding to each time frequency point in the speech parameter sequence is positively correlated with the probability of speech existing at each time frequency point.

3. The method according to claim 1, characterized in that, The step of identifying the data information corresponding to the audio data packet based on the target speech parameters includes: Construct target audio data using the audio data packet and the target speech parameters; The target audio data is input into the target recognition model, and the amplitude ratio output by the target recognition model is obtained as the data information. The amplitude ratio is used to indicate the proportional relationship between the speech amplitude of the speech signal formed by the speech data packet and the audio amplitude of the audio signal formed by the audio data packet. The target recognition model is obtained by training an initial recognition model using audio data samples labeled with amplitude ratio.

4. The method according to claim 3, characterized in that, The step of constructing target audio data using the audio data packet and the target speech parameters includes: Obtain the first frequency domain data of the first audio data collected by the first voice acquisition device deployed on each smart device in the audio data packet, and the second frequency domain data of the second audio data collected by the second voice acquisition device. The first frequency domain data, the second frequency domain data, and the target speech parameters are concatenated to form the target audio data.

5. The method according to claim 3, characterized in that, The step of inputting the target audio data into the target recognition model and obtaining the amplitude ratio output by the target recognition model as the data information includes: The target audio data is input into the convolutional layer of the target recognition model to obtain the initial data features output by the convolutional layer; The initial data features are input into the long short-term memory layer included in the target recognition model to obtain the target data features output by the long short-term memory layer; The target data features are input into the fully connected layer of the target recognition model to obtain the amplitude ratio output by the fully connected layer.

6. The method according to claim 1, characterized in that, The step of filtering target smart devices from the plurality of smart devices based on the voice parameters includes: The speech parameters are converted into energy parameters, wherein the energy parameters are used to indicate the speech energy of the speech included in the audio data packet; The smart device corresponding to the audio data packet with the largest energy parameter among the plurality of smart devices is identified as the target smart device.

7. A screening device for intelligent devices, characterized in that, include: The acquisition module is used to acquire audio data packets collected by each of the multiple smart devices in the target scene, wherein the audio data packets carry a wake-up command for waking up the multiple smart devices; A detection module is used to detect the target speech parameters of the audio data packet, wherein the target speech parameters are used to indicate the probability that each audio data in the audio data packet belongs to speech data; The recognition module is used to recognize the data information corresponding to the audio data packet according to the target speech parameters, wherein the data information is used to indicate the data attributes of the speech data packet in the audio data packet, and the data attributes include: the ratio between the clean speech amplitude and the noisy signal amplitude, wherein the clean speech amplitude is obtained from the speech data packet, and the noisy signal amplitude is obtained from the audio data packet; The determining module is used to determine the speech parameters of the speech carried in the audio data packet based on the data information; A filtering module is used to filter target smart devices from the plurality of smart devices according to the voice parameters, wherein the target smart device is used to respond to the wake-up command; The determining module includes: a first determining unit, configured to determine the signal amplitude of the audio signal formed in the audio data packet; and a second determining unit, configured to determine the product of the amplitude ratio included in the data information and the signal amplitude as the speech parameter, wherein the amplitude ratio is used to indicate the proportional relationship between the speech amplitude of the speech signal formed by the speech data packet and the audio amplitude of the audio signal formed by the audio data packet.

8. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored program, wherein the program, when executed, performs the method of any one of claims 1 to 6.

9. An electronic device comprising a memory and a processor, characterized in that, The memory stores a computer program, and the processor is configured to execute the method of any one of claims 1 to 6 through the computer program.