Wake-up method of distributed voice interaction device, storage medium and electronic device
By using an arbitration model based on the sound source orientation and directional gain of the wake-up audio, the problem of false wake-ups in distributed wake-up is solved, enabling precise device selection in complex environments and improving wake-up accuracy.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- QINGDAO HAIER TECH
- Filing Date
- 2025-12-25
- Publication Date
- 2026-04-17
AI Technical Summary
In multi-device smart homes, under distributed wake-up scenarios, existing technologies assume that the voice of the wake-up target is an isotropic sound source, leading to false wake-up problems in "nearby wake-up" or "orientation wake-up" schemes, and failing to accurately identify the most suitable wake-up device.
The arbitration model arbitrates the equivalent distance between the wake-up audio source location and the voice interaction device based on the directional gain of the wake-up audio source orientation. By combining deep learning and physical laws, it outputs a score of monotonic mapping relationship to determine the most suitable target device to respond.
It improves the accuracy of distributed wake-up and reduces false wake-ups, making it suitable for smart home devices in complex environments.
Smart Images

Figure CN121393443B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of smart home technology, and more specifically, to a wake-up method, storage medium, and electronic device for a distributed voice interaction device. Background Technology
[0002] In multi-device smart home scenarios, under distributed wake-up conditions, the same wake-up word is often triggered simultaneously by multiple devices (i.e., multiple voice interaction devices, referred to simply as devices). Related technologies' "proximity wake-up" (energy / amplitude) and "orientation wake-up" (direction) schemes mostly implicitly assume that the speaker is an isotropic (omnidirectional) sound source, optimizing only between "nearest / loudest" and "forward." This leads to the following drawbacks:
[0003] 1) The “nearest device” under different positions / postures is not necessarily the “loudest device” or the “most satisfactory response device”: When a user speaks with their back to the device, the high frequencies are significantly pointed forward, and the device behind may receive weaker effective voice energy even if it is close.
[0004] 2) Lack of perception of frequency band differences: The directionality of human voice is approximately omnidirectional below 1kHz, and gradually increases above 1kHz. Existing energy or single-directional wake-up criteria are difficult to stably reflect the impact of this frequency-dependent directionality on "perceiving the closest (hearing the closest)".
[0005] Therefore, in related technologies, the voice of the wake-up object (such as a person) in the distributed wake-up is not an isotropic sound source, which leads to the problem of false wake-up in the "nearby wake-up" (energy / amplitude) or "orientation wake-up" (direction) schemes.
[0006] In related technologies, the speech of the wake-up object in distributed wake-up is not an isotropic sound source, which leads to false wake-up problems in "nearby wake-up" (energy / amplitude) or "orientation wake-up" (direction) schemes. No effective solution has yet been proposed.
[0007] Therefore, it is necessary to improve the relevant technology to overcome the aforementioned defects. Summary of the Invention
[0008] This application provides a wake-up method, storage medium, and electronic device for a distributed voice interaction device, to at least solve the problem in the related art where the voice of the wake-up object in a distributed wake-up is not an isotropic sound source, leading to false wake-up in "nearby wake-up" (energy / amplitude) or "orientation wake-up" (direction) schemes.
[0009] According to one aspect of the embodiments of this application, a wake-up method for a distributed voice interaction device is provided, comprising: arbitrating the equivalent distance between the sound source location of the wake-up audio and the voice interaction device based on the directional gain of the sound source orientation of the wake-up audio using an arbitration model, and obtaining a first score that has a monotonic mapping relationship with the equivalent distance; wherein the arbitration model is deployed in the voice interaction device, the equivalent distance has a first approximate relationship with the geometric distance and the directional gain, the geometric distance is the physical geometric distance between the voice interaction device and the sound source location, and the arbitration model is pre-trained to learn the monotonic mapping relationship; determining a target device with the smallest equivalent distance to the sound source location among a plurality of voice interaction devices based on the first score, so as to respond to the wake-up audio through the target device.
[0010] According to another aspect of the embodiments of this application, a wake-up device for a distributed voice interaction device is also provided, comprising: an arbitration module, configured to arbitrate the equivalent distance between the sound source location of the wake-up audio and the voice interaction device based on the directional gain of the sound source orientation of the wake-up audio using an arbitration model, and obtain a first score that has a monotonic mapping relationship with the equivalent distance; wherein, the arbitration model is deployed in the voice interaction device, the equivalent distance has a first approximate relationship with the geometric distance and the directional gain, the geometric distance is the physical geometric distance between the voice interaction device and the sound source location, and the arbitration model is pre-trained to learn the monotonic mapping relationship; and a determination module, configured to determine, based on the first score, a target device among a plurality of voice interaction devices with the smallest equivalent distance to the sound source location, so as to respond to the wake-up audio through the target device.
[0011] According to another aspect of the embodiments of this application, a computer-readable storage medium is also provided, wherein a computer program is stored in the computer-readable storage medium, and the computer program is configured to execute the above-described wake-up method of the distributed voice interaction device when it is run.
[0012] According to another aspect of the embodiments of this application, an electronic device is also provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the wake-up method of the distributed voice interaction device through the computer program.
[0013] According to another aspect of the embodiments of this application, a computer program product is also provided, including a computer program that, when executed by a processor, implements the above-described method for waking up a distributed voice interaction device.
[0014] This application arbitrates the equivalent distance between the sound source location of the wake-up audio and the voice interaction device using an arbitration model based on the directional gain of the sound source orientation of the wake-up audio, obtaining a first score that has a monotonic mapping relationship with the equivalent distance. The arbitration model is deployed in the voice interaction device, and the equivalent distance has a first approximate relationship with the geometric distance and the directional gain. The geometric distance is the physical geometric distance between the voice interaction device and the sound source location. The arbitration model is pre-trained to learn the monotonic mapping relationship. The first score is used to determine the target device with the smallest equivalent distance to the sound source location among multiple voice interaction devices, so that the target device responds to the wake-up audio. Therefore, the above technical solution solves the problem in related technologies where the voice of the wake-up object is not an isotropic sound source, leading to false wake-ups in "nearby wake-up" (energy / amplitude) or "orientation wake-up" (direction) schemes; thus improving wake-up accuracy. Attached Figure Description
[0015] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0016] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0017] Figure 1 This is a schematic diagram of the hardware environment for an optional wake-up method for a distributed voice interaction device according to an embodiment of this application;
[0018] Figure 2 This is a flowchart of an optional wake-up method for a distributed voice interaction device according to an embodiment of this application;
[0019] Figure 3 This is an architecture diagram of a distributed wake-up arbitration system according to an embodiment of this application;
[0020] Figure 4 This is a schematic diagram of the arbitration model of an optional wake-up method for a distributed voice interaction device according to an embodiment of this application;
[0021] Figure 5 This is a schematic diagram of the time-frequency lightweight convolutional stack in the arbitration model of an optional wake-up method for a distributed voice interaction device according to an embodiment of this application.
[0022] Figure 6This is a structural block diagram of a wake-up device for an optional distributed voice interaction device according to an embodiment of this application. Detailed Implementation
[0023] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.
[0024] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0025] According to one aspect of the embodiments of this application, a method for waking up a distributed voice interaction device is provided. This method for waking up a distributed voice interaction device is widely used in whole-house intelligent digital control application scenarios such as smart homes, smart home ecosystems, and intelligencehouse ecosystems. Optionally, in this embodiment, the above-mentioned method for waking up a distributed voice interaction device can be applied to, for example... Figure 1 The hardware environment shown consists of multiple terminal devices 102 and a server 104. For example... Figure 1 As shown, server 104 is connected to multiple terminal devices 102 via a network and can be used to provide services (such as application services) to terminals or clients installed on terminals. A database can be set up on the server or independently of the server to provide data storage services for server 104. Cloud computing and / or edge computing services can be configured on the server or independently of the server to provide data processing services for server 104.
[0026] The aforementioned network may include, but is not limited to, at least one of the following: wired network, wireless network. The aforementioned wired network may include, but is not limited to, at least one of the following: wide area network, metropolitan area network, local area network. The aforementioned wireless network may include, but is not limited to, at least one of the following: Wi-Fi (Wireless Fidelity), Bluetooth. The terminal device 102 may not be limited to PC, mobile phone, tablet computer, smart air conditioner, smart range hood, smart refrigerator, smart oven, smart stove, smart washing machine, smart water heater, smart washing equipment, smart dishwasher, smart projector, smart TV, smart clothes rack, smart curtains, smart audio-visual equipment, smart socket, smart speaker, smart speaker box, smart fresh air equipment, smart kitchen and bathroom equipment, smart bathroom equipment, smart robot vacuum cleaner, smart window cleaning robot, smart mopping robot, smart air purifier, smart steam oven, smart microwave oven, smart water heater, smart air purifier, smart water dispenser, smart door lock, etc.
[0027] This embodiment provides a wake-up method for a distributed voice interaction device, including but not limited to the terminal device described above. The terminal device can be a voice interaction device, such as a smart speaker or smart audio system. Figure 2 This is a flowchart of an optional wake-up method for a distributed voice interaction device according to an embodiment of this application. The process includes the following steps:
[0028] Step S202: Arbitrate the equivalent distance between the sound source location of the wake-up audio and the voice interaction device using an arbitration model based on the directional gain of the sound source orientation of the wake-up audio, and obtain a first score that has a monotonic mapping relationship with the equivalent distance; wherein, the arbitration model is deployed in the voice interaction device, the equivalent distance has a first approximate relationship with the geometric distance and the directional gain, the geometric distance is the physical geometric distance between the voice interaction device and the sound source location, and the arbitration model has learned the monotonic mapping relationship through pre-training;
[0029] Step S204: Determine the target device with the smallest equivalent distance to the sound source location among the plurality of voice interaction devices based on the first score, so as to respond to the wake-up audio through the target device.
[0030] Through the above steps, an arbitration model arbitrates the equivalent distance between the sound source location of the wake-up audio and the voice interaction device based on the directional gain of the sound source orientation of the wake-up audio, obtaining a first score that has a monotonic mapping relationship with the equivalent distance. The arbitration model is deployed in the voice interaction device, and there is a first approximate relationship between the equivalent distance, the geometric distance, and the directional gain. The geometric distance is the physical geometric distance between the voice interaction device and the sound source location. The arbitration model is pre-trained to learn the monotonic mapping relationship. The first score determines the target device with the smallest equivalent distance to the sound source location among multiple voice interaction devices, allowing the target device to respond to the wake-up audio. Therefore, this technical solution solves the problem in related technologies where the voice of the wake-up object is not an isotropic sound source, leading to false wake-ups in "nearby wake-up" (energy / amplitude) or "orientation wake-up" (direction) schemes; thus improving wake-up accuracy.
[0031] In an exemplary embodiment, before arbitrating the equivalent distance between the sound source location of the wake-up audio and the voice interaction device based on the directional gain of the sound source orientation of the wake-up audio using an arbitration model to obtain a first score that has a monotonic mapping relationship with the equivalent distance, the method further includes: acquiring training samples, wherein the training samples are a set of devices responding to each wake-up event in the same space, and the observation values of the training samples are audio features collected by each device in the set of devices; arbitrating the equivalent distance of the audio features using a target model to output a second score corresponding to each device; determining the soft target distribution of the training samples using the geometric information corresponding to the training samples; determining the response probability of each device using the second score; determining the total loss of the target model using the response probability and the soft target distribution, and performing a backpropagation gradient update on the target model using the total loss to obtain the arbitration model.
[0032] Optionally, the response probability of each device can be calculated based on the second score using the Softmax function, and the sum of the response probabilities of all devices in the device set should be 1.
[0033] Through the steps described above, guided by the distribution of soft targets, the arbitration model learns how to determine the suitability of a device's response based on audio characteristics under complex environmental conditions (including changes in the relative positions of the device and the sound source, the orientation of the sound source, etc.), ultimately achieving the goal of accurate arbitration based on equivalent distance. This process cleverly utilizes the combination of physical laws and deep learning to realize a distributed arbitration strategy that does not require explicit distance measurement.
[0034] Optionally, determining the response probability of each device using the second score includes: wherein determining the response probability includes: ;in, This represents the response probability. This refers to the set of devices; This represents the device index for each of the aforementioned devices. This represents the device index of any device in the device set other than each of the aforementioned devices; For temperature parameters, , Indicates the second score. Indicates equipment The score is divided into two parts, with the second part including the score of the actual response device.
[0035] Optionally, determining the total loss of the target model using the response probability and the soft target distribution includes: determining the interval ranking loss corresponding to the target model using the second score; and determining the total loss of the target model using the interval ranking loss, the response probability, and the soft target distribution; wherein determining the interval ranking loss includes: ;in, This represents the interval sorting loss. This represents the device index of the actual responding device in the device set, where m represents the actual responding device in the device set. With any device The interval between Indicates equipment The score, This represents the score of the actual response device output by the target model; where, This represents the device index of any device in the device set other than each of the aforementioned devices.
[0036] The interval ranking loss is used to ensure that the model can correctly distinguish the real responding device from the other devices, especially among devices that are close in equivalent distance. The model's scores should also have a significant interval, with the scores of real responding devices being significantly higher than those of non-responding devices, at least by a fixed interval m, to avoid false wake-ups or wake-up uncertainty.
[0037] The determination of the total loss of the target model using the interval sorting loss, the response probability, and the soft target distribution includes: determining the distillation loss of the target device using the response probability and the soft target distribution; determining the cross-entropy loss of the target model using the response probability; and performing a weighted summation of the cross-entropy loss, the distillation loss, and the interval sorting loss to obtain the total loss.
[0038] Furthermore, determining the total loss of the target model using the response probability and the soft target distribution includes: determining the distillation loss of the target device using the response probability and the soft target distribution; and determining the total loss of the target model using the distillation loss.
[0039] The determination of the total loss of the target model through the distillation loss includes: determining the interval ranking loss corresponding to the target model through the second score, determining the cross-entropy loss of the target model through the response probability; and performing a weighted summation of the cross-entropy loss, the distillation loss, and the interval ranking loss to obtain the total loss.
[0040] Furthermore, determining the soft target distribution of the training samples using the geometric information corresponding to the training samples includes: determining the directional gain corresponding to the audio feature using a preset directional model; determining the equivalent distance corresponding to the audio feature using the directional gain corresponding to the audio feature, the geometric information, and the first approximation relationship; wherein the geometric information includes: the sound source location of each wake-up event, the sound source orientation of each wake-up event, and the device location of each device; and determining the soft target distribution using the equivalent distance corresponding to the audio feature.
[0041] Among them, the preset directional model can be an nth-order cardioid model, which corresponds to the directional gain of the audio features. The approximate determination formulas include:
[0042] .
[0043] in, This represents the angle between the sound source axis and the sound source of each wake-up event. =0 indicates the front (=0 means directly in front). This indicates the frequency of the sound source in each wake-up event. , This represents the frequency-dependent directional factor. With frequency Monotonically increasing.
[0044] Optionally, the directional gain can also be tested in the same space and determined using test data.
[0045] By incorporating the directional gain corresponding to the audio features and the geometric information into the first approximation relationship, the equivalent distance corresponding to the audio features can be obtained.
[0046] For a speech segment, the equivalent distance corresponding to the audio features is calculated. Equivalent distance weighted by harmonic average aggregation to frequency band This aligns more closely with the intuition of "energy superposition." Specifically, the formula for calculating the equivalent distance using frequency band weighting is defined as follows:
[0047] .
[0048] in, Let b represent a subset of representative frequency bands (e.g., the mid-to-high frequency subset in Mel-bands), where b represents the index of the frequency band within the subset of representative frequency bands. It is a frequency characteristic point of frequency band b. This represents the weight of frequency band b.
[0049] The formula for calculating the distribution of soft targets is as follows:
[0050] .
[0051] in, Represents the distribution of soft targets. These are the distribution parameters.
[0052] Optionally, determining the distillation loss of the target device using the response probability and the soft target distribution includes: determining the KL divergence between the response probability and the soft target distribution of each device to obtain the distillation loss, wherein the distillation loss is used to enable the target model to learn a monotonic mapping relationship between the second score and the equivalent distance corresponding to the audio feature.
[0053] Among them, distillation loss The calculation formula is:
[0054] .
[0055] Furthermore, determining the total loss of the target model using the response probability and the soft target distribution includes: determining the cross-entropy loss of the target model using the response probability; and determining the total loss of the target model using the cross-entropy loss, the response probability, and the soft target distribution.
[0056] The determination of the total loss of the target model using the cross-entropy loss, the response probability, and the soft target distribution includes: determining the interval ranking loss corresponding to the target model using the second score; determining the distillation loss of the target device using the response probability and the soft target distribution; and performing a weighted summation of the cross-entropy loss, the distillation loss, and the interval ranking loss to obtain the total loss.
[0057] Optionally, determining the cross-entropy loss of the target model using the response probability includes: wherein determining the cross-entropy loss includes: ;in, This represents the device index of the actual responding device in the device set. This represents the probability of a device's actual response. This represents the cross-entropy loss.
[0058] In an exemplary embodiment, determining the total loss of the target model using the response probability and the soft target distribution includes: determining the interval ranking loss corresponding to the target model using the second score; determining the cross-entropy loss of the target model using the response probability; determining the distillation loss of the target model using the soft target distribution; and performing a weighted summation of the cross-entropy loss, the distillation loss, and the interval ranking loss to obtain the total loss.
[0059] The formula for determining the total loss includes:
[0060] .
[0061] in, The weight representing the distillation loss, The weights represent the loss of the interval sorting.
[0062] In an exemplary embodiment, before arbitrating the equivalent distance between the sound source location of the wake-up audio and the voice interaction device based on the directional gain of the wake-up audio's sound source orientation using an arbitration model to obtain a first score that has a monotonic mapping relationship with the equivalent distance, the method further includes: determining a second approximate relationship between the geometric distance, the directional gain, and the relative intensity of the wake-up audio; and obtaining a third approximate relationship with the audio intensity of the wake-up audio; determining the first approximate relationship between the equivalent distance, the geometric distance, and the directional gain through the second approximate relationship and the third approximate relationship; wherein, the first approximate relationship is: ;in, This represents the equivalent distance. Represents the geometric distance, This refers to the directional gain.
[0063] Optionally, the approximate formula for the second approximation relation is:
[0064] .
[0065] in, Indicates the relative intensity of the wake-up audio. The air absorption coefficient (related to frequency ω, temperature, and humidity) represents the degree of sound attenuation in air; generally, higher frequency sound waves attenuate faster in air.
[0066] The approximate formula for the third approximation relation is: The third approximation relation refers to the sound intensity of a sound source in a free field. (Equivalent to the audio intensity of the wake-up audio in the above embodiments) and distance It approximately satisfies the inverse square law .
[0067] This application embodiment converts directionality and absorption into "equivalent geometric distance" based on the second and third approximation relations, so that the same intensity is achieved at the axial equivalent point (θ=0). The conversion process includes:
[0068] .
[0069] In short and medium distances, In smaller home environments, the impact of air absorption can be initially ignored, i.e., the index term can be disregarded. Or incorporate the exponent term. The effective terms (engineering approximations) are used to obtain the first approximation relation.
[0070] In an exemplary embodiment, an arbitration model arbitrates the equivalent distance between the sound source location of the wake-up audio and the voice interaction device based on the directional gain of the sound source orientation of the wake-up audio, obtaining a first score that has a monotonic mapping relationship with the equivalent distance. This includes: recalibrating the feature matrix of the target audio feature of the wake-up audio band-wise using a band-selection gate, wherein the recalibrated feature matrix highlights the high-frequency directivity of the target audio feature; processing the recalibrated feature matrix using a time-frequency lightweight convolutional stack to obtain a time-frequency encoding matrix corresponding to the recalibrated feature matrix; performing weighted pooling on the time-frequency encoding matrix using a global context pooling layer to obtain a global representation corresponding to the time-frequency encoding matrix; and monotonically projecting the global representation using a physical consistency projector to output the first score corresponding to the target audio feature. The arbitration model includes: the band-selection gate, the time-frequency lightweight convolutional stack, the global context pooling layer, and the physical consistency projector.
[0071] Optionally, the arbitration model used in this embodiment can be a Directionality-aware Equivalent-distance Arbitrator (DEAR) model. After training the arbitration model using the total loss described in the above embodiments, if a voice interaction device with the arbitration model receives a wake-up audio, the voice interaction device can extract the target audio features of the wake-up audio and input these target audio features into the arbitration model for processing. The arbitration model processes the target audio features sequentially through a frequency band selection gate, a time-frequency lightweight convolutional stack, a global context pooling layer, and a physically consistent projection head to obtain the first score corresponding to the target audio features.
[0072] Furthermore, the feature matrix of the target audio features of the wake-up audio is recalibrated band-by-band using a band selection gate, including: normalizing the features of each band in the feature matrix to obtain normalized feature values; generating initial band weights for each band using the normalized feature values, wherein the initial band weights increase with the frequency of the wake-up audio; normalizing the initial band weights to obtain normalized band weights; and recalibrating the feature matrix band-by-band using the normalized band weights. The initial band weights increase slowly with the frequency of the wake-up audio.
[0073] Specifically, the feature matrix of the target audio features undergoes band-level normalization. This step standardizes the feature values in each band, eliminating the inherent differences in feature intensity between different bands and laying the foundation for subsequent operations. Then, initial band weights are generated using the normalized feature values. Importantly, these weights increase with frequency, a design that fully reflects the dominant position of high-frequency sound components in direction perception—that is, high-frequency components have stronger directionality in human voices. Next, these initial weights are normalized, and the resulting normalized band weights are used for band-by-band recalibration of the feature matrix. This process essentially assigns greater weight to high-frequency components, allowing the recalibrated feature matrix to more clearly highlight the high-frequency directional features of the wake-up audio.
[0074] The feature matrix processed by the frequency band selection gate is further fed into a lightweight time-frequency convolution stack. This stack employs depthwise separable convolution and temporally dilated convolution, which can capture the short-term spectral features and medium-to-long-term modulation patterns of the wake-up audio while maintaining low computational cost. These patterns are crucial for identifying the sound source direction and estimating the equivalent distance. The time-frequency coding matrix is then weighted by a global context pooling layer. This layer automatically selects the most decisive spatiotemporal segments in response to the wake-up event, forming a global representation. This representation integrates key information from the wake-up audio, providing strong feature support for the final equivalent distance estimation.
[0075] Finally, the global representation is monotonically projected using a physically consistent projection head to generate a first score that is monotonically negatively correlated with the equivalent distance. This step ensures that the score output by the model can accurately reflect the relative proximity between the device and the sound source, and can achieve correct device arbitration even in complex scenarios (such as when the user is facing away from the device or there are multiple device responses).
[0076] The arbitration model's processing of audio features not only considers the spectral characteristics of the audio features but also fully integrates the directional gain of the sound source orientation. The final output score accurately reflects the equivalent distance between the device and the sound source, greatly enhancing the accuracy and robustness of device selection in distributed voice wake-up scenarios. It demonstrates superior performance, particularly in handling high-frequency directivity and complex indoor environmental effects, effectively addressing the limitations of traditional energy / amplitude arbitration or simple orientation arbitration in practical applications. This provides a more intelligent and user-friendly solution for IoT applications such as smart homes.
[0077] In an exemplary embodiment, determining the target device with the smallest equivalent distance to the sound source location among a plurality of voice interaction devices using the first score includes at least one of the following: sending the first score to an arbitration device to instruct the arbitration device to determine the target device with the smallest equivalent distance to the sound source location among a plurality of voice interaction devices using the first score, wherein the arbitration device includes one of the following: any of the other voice interaction devices among the plurality of voice interaction devices, or a cloud arbitration device; obtaining a third score sent by the other voice interaction devices, and determining the target device with the smallest equivalent distance to the sound source location among a plurality of voice interaction devices using the third score and the first score.
[0078] In other words, any device (including but not limited to the wake-up device itself) can act as an arbitration device, responsible for collecting the first scores from all potential responding devices. Specifically, determining the target device with the smallest equivalent distance to the sound source location among the plurality of voice interaction devices using the first scores includes: identifying the device corresponding to the highest score among the plurality of first scores as the target device. Determining the target device with the smallest equivalent distance to the sound source location among the plurality of voice interaction devices using the third score and the first score includes: identifying the device corresponding to the highest score between the first score and the third score as the target device.
[0079] Obviously, the embodiments described above are only some embodiments of this application, and not all embodiments. To better understand the wake-up method of the distributed voice interaction device described above, the process is explained below with reference to embodiments, but this is not intended to limit the technical solutions of the embodiments of this application. Specifically:
[0080] In multi-device smart homes, the same wake word is often triggered simultaneously by multiple devices. Existing "proximity wake-up" (energy / amplitude) and "orientation wake-up" (direction) solutions mostly implicitly assume that the speaker is an isotropic (omnidirectional) sound source, or only optimize between "nearest / loudest" and "forward," without uniformly incorporating the "frequency-related directionality" of human speech radiation into the arbitration criteria. This leads to:
[0081] 1) "Nearest ≠ Loudest ≠ Most Desired Response Device" under different placement / postures: Human voice radiation exhibits frequency-dependent directionality (high frequencies are more forward-oriented). Mid-to-high frequencies are more easily attenuated by air and sound-absorbing materials; if a user speaks with their back to a device, high frequencies are blocked / diffused by their head and significantly directed forward. Even if the device behind is close, it may receive weaker effective speech energy, meaning the received effective speech characteristics will be inferior to those of the device in front.
[0082] 2) Lack of perception of frequency band differences: The directionality of human voice is approximately omnidirectional below 1kHz, and gradually increases above 1kHz. Existing energy or single-directional criteria are difficult to stably reflect the impact of this frequency-dependent directionality on "perceiving the closest (hearing the closest)".
[0083] 3) The gap between training labels and physical quantities: In engineering, the labels that can be collected are mostly "final response device index (who responds)", which makes it difficult to directly label "equivalent distance" (physical quantity), making the two-stage route of "estimating physical quantity → mapping to arbitration decision" not feasible in terms of data.
[0084] This raises a technical problem that needs to be solved: given the constraint that each device can only calculate its own score locally during the inference phase, and that it only has the available label of "response device index", how can each device output a score that is monotonically negatively correlated with the "equivalent distance value relative to the sound source", and fully utilize the directionality of human voice frequency correlation without explicitly regressing the distance, to achieve deployable distributed voice wake-up arbitration.
[0085] Among related technologies, end-to-end device arbitration (Alexa, 2021) proposes modeling the problem of "which device to choose to respond" as "choosing the nearest one from multiple concurrent triggers," replacing explicit localization with end-to-end deep learning, and using large-scale indoor simulation data for training. It outperforms traditional signal processing baselines under noise / reverberation conditions. This work emphasizes that it is not necessary to fully compute 3D positions; relative proximity embeddings and aggregation decisions can be learned directly. However, this approach focuses on end-to-end "nearest device" learning, but has not yet uniformly mapped "voice frequency-related directionality" to a single-machine deployable model with "equivalent distance values" as "distributable scoring." Furthermore, the path of training to monotonic scoring using only "response device indexes" without explicit distance regression lacks systematic description.
[0086] In summary, the existing arbitration criteria (energy / amplitude or direction) do not uniformly consider the multiplicative effect of "frequency × direction", which leads to inconsistencies among the nearest / loudest / forward three in complex scenarios, resulting in defects such as "one call, multiple responses / incorrect responses / delays".
[0087] To address the aforementioned problems, the objectives of this application's embodiments are as follows:
[0088] 1) A compact neural network with single-device input and single scalar output is proposed. The output score is monotonically negatively correlated with the "equivalent distance". Arbitration can be achieved by local comparison after multi-device parallel inference.
[0089] 2) By utilizing physical priors (voice directionality + propagation model) during training, the target is distilled into a soft target. Without explicit distance labels, it can still learn a score consistent with the equivalent distance.
[0090] 3) Design feasible data construction and simulation enhancement (including directional sound sources, heterogeneous devices, and synchronous alignment), as well as training strategies for group-wide normalization targets, to ensure closed-loop availability in multi-device group supervision and single-device independent inference.
[0091] 4) It meets the requirements of edge device deployment. The device side only needs local audio and does not need to pull the original audio from other devices in real time, resulting in low power consumption and low latency.
[0092] Specifically, the "equivalent distance" is defined in this application embodiment. Let the distance from the sound source (user's mouth) to the device be... The geometric distance is The angle between the sound source axis and the sound source axis is ( (Indicates directly in front). Let the angular frequency be... , , Indicates the frequency point index. The source directional gain (relative to the axial front, linear power domain) is: ,satisfy , (The back direction approaches 0). Therefore, under the free field approximation, the frequency point... The relative intensity reaching the equipment is approximately:
[0093] .
[0094] in Air absorption coefficient (with frequency) Air absorption (related to temperature and humidity) indicates the degree of sound attenuation in air; generally, higher frequency sound waves attenuate faster in air. Air absorption is frequency-dependent: mid-to-high frequencies have stronger molecular absorption in air. ISO 9613-1 provides formulas for calculating the air absorption coefficient under frequency, temperature, humidity, and air pressure conditions, which is the engineering basis for simulation and propagation estimation.
[0095] Equivalent distance is defined as converting directionality and absorption into "equivalent geometric distance", such that for axial equivalent points ( They have the same strength:
[0096] .
[0097] It should be noted that the conversion process incorporates sound propagation and the inverse square law. Specifically, the sound intensity I of a free-field point sound source approximately satisfies the inverse square law with respect to distance r. The sound pressure level satisfies p∝1 / r. When early reflections / reverberation are present in the room, the approximate relationship will be corrected, but the overall trend of "distance attenuation" still holds.
[0098] In short and medium distances, In smaller home environments, the exponential term can be ignored or incorporated into the effective terms of G (engineering approximation):
[0099] .
[0100] The corresponding physical meaning is that deviating from the axis (G becomes smaller) will effectively "distance" the sound, explaining the phenomenon that "the sound is less pleasing when facing away from a nearby device than when facing forward to a distant device." The directionality of human voices increases with frequency (G is smaller at high frequencies), therefore the equivalent distance is more sensitive to high-frequency components, which is consistent with experience.
[0101] It should be noted that human voice directivity is important. Human voices are not isotropic sound sources; high-frequency energy is concentrated in front of the mouth. Directivity increases with frequency, a consensus found in numerous measurement papers and reviews. Related technologies have systematically studied the frequency-dependent directivity index (DI) and changes in head orientation, reporting individual differences. Human voice directivity is the physical basis for the "equivalent distance value" concept proposed in this application's embodiments: the greater the angle of deviation from the sound source axis, the more equivalently, "a longer geometric distance is needed to obtain the same effective energy."
[0102] Secondly, the embodiments of this application combine the equivalent distance with the monotonicity of the scoring function and the learnable target.
[0103] In actual equipment arbitration, the embodiments of this application do not explicitly output... Take a monotonically decreasing mapping. Compress "equivalent distance" into a score :
[0104] .
[0105] in, This represents the aggregation of equivalent distances across multiple frequency bands for a given speech segment. As long as... Monotonous, then and Monotonically negatively correlated. This application embodiment uses... The approximation is learned from LogFBank (or MFCC) by a neural network (equivalent to the target model in the above embodiments), without the need for explicit construction. or The tags, which only utilize the supervision of "who responded" within the group, encourage the network to learn a scoring order consistent with "equivalent proximity".
[0106] This application embodiment includes a distributed wake-up arbitration system (used to execute the wake-up method of the above-mentioned distributed voice interaction device), and the architecture diagram of the system is as follows. Figure 3 As shown. The system as a whole consists of the device side and the arbitration layer. Each device collects its own wake-up audio, extracts the audio features of the wake-up audio (equivalent to the target audio features in the above embodiment), sends them into the DEAR model, and outputs a local score S (in the case of devices A, B, C..., S includes...). Figure 3 The S_A, S_B, S_C, etc., shown in the example. Optionally, in this embodiment, S represents the first score output by the arbitration model, and... This represents the second score in the process of training the target model to obtain the arbitration model. In a local or edge bus, multiple devices only need to exchange a scalar S (or aggregated by the gateway) to obtain the decision result of the responding device, whichever has the largest value. There is no need to exchange raw audio (privacy-friendly) or explicitly estimate the orientation or distance.
[0107] Among these, audio characteristics: Traditional small devices and streaming wake-up applications extensively use Log-Mel filter-bank energies (LogFBank) or Mel-Frequency Cepstral Coefficients (MFCC) as front-end acoustic representations. Both are transformations of the short-time power spectrum on the Mel frequency standard, where MFCC is the cepstral coefficient obtained from LogFBank through Discrete Cosine Transform (DCT).
[0108] The target model trained in this embodiment can be a Direction-Aware Isometric Arbiter (DEAR) model (hereinafter referred to as the model). The architecture of the DEAR model satisfies single-ended input / single scalar output, physical interpretability, and lightweight inference. The model structure is as follows: Figure 4 As shown.
[0109] Model input: Feature matrix collected by a single device (i.e., each device) (Extract LogFBank or MFCC features from the wake-up audio received by each device, with a frame shift of 10ms, a frame length of 25ms, and 64–80 Mel filters, aiming to make the energy-frequency structure explicit, facilitating subsequent learning of "high-frequency directionality" related patterns. Feature matrix) (equivalent to the target audio features in the above embodiments), T represents the number of frames, and F represents the number of frequency bins.
[0110] Model output: scalar score It is used for comparison of multiple devices (equivalent to multiple voice interaction devices).
[0111] The feature transfer process for the model backbone and its modules is as follows: .
[0112] Among them, frequency band selection gate processing (Equivalent to the target audio features in the above embodiments), to obtain (Equivalent to the recalibrated feature matrix in the above embodiments), time-frequency lightweight convolution stack processing ,get (Equivalent to the time-frequency coding matrix in the above embodiments), global context pooling layer processing. ,get (Equivalent to the global representation in the above embodiments), Physically Consistent Projection Head Processing , and got the first score s.
[0113] The main module structure includes the following modules:
[0114] 1) Frequency Gate (F-Gate): Learnable gating or attention along the frequency dimension ,right Frequency band recalibration ,in, This is equivalent to the recalibrated feature matrix in the above embodiments. For frame index, This is the index of the frequency bin, and the corresponding frequency bin is... This module is designed to emphasize the contribution of the >1kHz band (which is more directional) to the scoring; W is initialized with a slight high-pass bias and trained for adaptive fine-tuning.
[0115] Specifically, the F-Gate module aims to highlight the contribution of high-frequency (>1kHz) directivity enhancement to the equivalent distance, constructing interpretable and trainable band-wise weights. To address this, physical guidance at the feature level is used to constrain the frequency band gate F-Gate, including: a) initializing W as a slightly rising function (emphasis >1kHz) to match the human voice pattern that "high frequencies are more directional"; b) applying non-negative constraints and total variation (TV) regularization to W during training to ensure "smoothness and interpretability".
[0116] (1) Statistics and parameterization include:
[0117] a) Frequency band normalization:
[0118] .
[0119] .
[0120] , The first Variance at each frequency bin Use a small value to avoid the denominator being 0.
[0121] b) Gated generation (band-by-band sensing, including high-frequency initialization), including:
[0122] Choosing a layer of affine activation + monotonic activation yields non-negative weights:
[0123] .
[0124] .
[0125] in and These are learnable parameters. For high-frequency initialization, their initial values can be:
[0126] .
[0127] Make the initial As frequency gradually increases ( , (Controlling the turning point), which aligns with the physical prior that "high-frequency frequencies are more directional".
[0128] c) Normalization:
[0129] or .
[0130] in, This is the index of the frequency bin.
[0131] d) Gating process: .
[0132] (2) Regularity and constraints, including: a) Nonnegativity: through Natural satisfaction b) Smoothness (total variation TV): Suppressing excessive oscillations and maintaining interpretable frequency band trends. c) Physical prior closure: In conjunction with Physics-KD, making gating more closely fit soft targets of "frequency × directivity".
[0133] The F-Gate module helps ensure that the equivalent distance aligns with intuitive recognizability. In "positive example devices" vs other devices "Expected feature difference" Subscript This means taking all values of that dimension, therefore (This means taking all time frames). If the final score is approximately linearly weighted, then the optimal distinguishing weight, in the Fisher discriminant sense, is similar to... A positive correlation exists. The high-frequency differences in human voice (significantly attenuated when facing away) cause... At higher frequencies, the difference is even greater, so F-Gate learns that "high-frequency bias" has both physical and statistical basis.
[0134] 2) Lightweight Temporal Convolutional Stack (DW-TCN): It uses depthwise separable convolution (MobileNet style) + temporal dilated convolution to capture amplitude envelope, spectral tilt, formants and short-term dynamics while maintaining low computational power.
[0135] Specifically, the DW-TCN module is designed to capture short-time spectral morphology and medium-time modulation with very few parameters.
[0136] (1) First, perform channel enhancement using 1x1 convolution: .
[0137] (2) such as Figure 5 As shown, L residual blocks are stacked, each block containing:
[0138] a) Depthwise Conv (along time frequency 2D):
[0139] .
[0140] in Convolution is performed on each channel separately; time-dimensional dilatancy. (Exponential inflation). , These represent the size of the convolutional kernel in the time dimension and the frequency dimension, respectively. Generally taken .
[0141] b) Pointwise Conv Fusion Channel: .
[0142] c) Residuals: Then reactivate.
[0143] (3) Activation function Use SiLU or ReLU; use GroupNorm for normalization (number of groups G=4 or 8) to adapt to small batches.
[0144] (4) Effective receptive field (time dimension derivation):
[0145] No. Layer time-receptive field increment Overall feeling field: It can effectively cover the wake word and the surrounding context.
[0146] 3) Global Attention Pooling (GAP-Attn): For... Time-frequency coding (Equivalent to the time-frequency coding matrix in the above embodiment) is weighted pooled (C represents the number of channels), with weights generated by the attention subnet, and the output is a global representation. (D represents the dimension of h, where h is equivalent to the global representation in the above embodiment). This module aims to aggregate stable features of a wake word fragment and suppress silence / noise moments.
[0147] Specifically, the GAP-Attn module aims to integrate time-frequency characteristics. Convergence into fixed-dimensional embedding Highlight the effective moments / frequency bands of the voice.
[0148] (1) Decompose attention (time × frequency):
[0149] Constructing time weights With frequency weight (all ,and ):
[0150] .
[0151] .
[0152] Where u and v are the query vectors of the attention heads in the time and frequency dimensions, respectively, used to calculate attention scores, enabling the model to learn which features are more important. Weights , and deviation , These are all model parameters to be learned.
[0153] (2) Joint weighting and projection:
[0154] .
[0155] in, Suppressing silence / noise moments, Emphasis on high-frequency bands (complementing F-Gate), weighting Perform the final linear projection.
[0156] 4) Physically Consistent Projection (Phys-Proj): A single layer of small MLP output scalars. The projection head incorporates monotonicity constraints (guaranteed by non-negative weights and bias control in the last layer or by Softplus activation, and negatively correlated with the "equivalent distance" based on the overall loss). No other devices are required during inference; only the local output is used. .
[0157] Specifically, the output score s and the "equivalent distance" "Monotonic negative correlation. The Phys-Proj module aims to achieve implicit monotonicity through parameter nonnegation and distillation objectives."
[0158] (1) Monotonic projection:
[0159] make After passing through a small MLP layer (1–2 layers), it becomes linearly scalar:
[0160] .
[0161] Where 'a' is a non-negative weight parameter (which can be used) (The weight 'a' must be non-negative). If For component-wise monotonically increasing activation (ReLU / SiLU), then s for each Monotonically increasing. Combined with Physics-KD, z will learn a representation positively correlated with proximity, thus making s, overall, more closely related to proximity. Negative correlation.
[0162] (2) Consistency with physical distillation:
[0163] Target distribution (equivalent to the soft target distribution in the above embodiments):
[0164] .
[0165] in represents the target distribution parameters.
[0166] During training (Equivalent to the KL divergence in the above embodiments) is used as the loss term, making
[0167] .
[0168] therefore and They exhibit a monotonically negative correlation, thus eliminating the need for explicit regression distance. Temperature parameter (learnable or hyperparameter).
[0169] It should be noted that group normalization (only during training) refers to the same room / same statement and N devices. The probabilities are obtained through a list-based Softmax algorithm. Cross-entropy training is performed using the response device index. Each of the above modules performs an irreplaceable function, with no parallel redundant branches, and the total number of network parameters can be controlled within <1–3M, making it friendly for edge deployment.
[0170] Furthermore, the training strategy for the direction-aware isometric arbiter DEAR model includes: setting a training sample as a set of devices in the same room experiencing a single wake-up event. The observations are for the characteristics of each device. Network output scalar The specific training process includes:
[0171] 1) List-based cross-entropy (within-group Softmax):
[0172] .
[0173] .
[0174] in Indexing for real-responding devices. This allows the network to learn the order of "maximum positive scores".
[0175] 2) Physical-guided knowledge distillation (Physics-KD).
[0176] Using the geometric information (sound source / device location), sound source orientation, and a human voice directionality model available in the training data, a soft target distribution is given. .
[0177] (1) Choose a simple directional approximation, such as the nth-order heart shape model:
[0178] .
[0179] With frequency Monotonically increasing (its trend / piecewise fit DI value can be estimated by the latest measurement studies).
[0180] (2) For the speech segment, put Aggregate by harmonic average This aligns more closely with the intuition of "energy superposition." Define the equivalent distance for frequency band weighting:
[0181] .
[0182] These are representative frequency bands (e.g., the mid-to-high frequency subset in Mel-bands).
[0183] (3) Distribution of soft targets (the closer the target, the higher the probability):
[0184] .
[0185] in These are the distribution parameters.
[0186] Distillation loss:
[0187] .
[0188] It is used only during training; inference does not require geometry / orientation. It injects the physical laws of human voice directionality and propagation into the network, enabling the learned scores to be... It is spontaneously monotonically negatively correlated with "equivalent distance".
[0189] 3) Contrast / Ranking Enhancement:
[0190] For positive examples With negative examples Add the interval sorting loss:
[0191] .
[0192] ( ) represents the margin. Margin ranking loss is used to improve the margin robustness of hard-nearest neighbor samples.
[0193] Total loss: .in, As weight.
[0194] Equivalence between single-device output and intra-group supervision: During training, intra-group Softmax actually learns a list-based sorting target. During inference, it only requires comparing the unnormalized scores of each device. Theoretically, at the same temperature The Softmax Top-1 values are consistent because:
[0195] .
[0196] For the set of devices involved in the aforementioned wake-up event, the dataset construction and simulation process includes:
[0197] 1) Real data collection:
[0198] Room and equipment layout: Arrange N heterogeneous devices (different microphone types / gains / array sizes) in several real home / laboratory rooms and mark their three-dimensional positions;
[0199] Sound source location and orientation: Mark the three-dimensional location of the sound source and record the orientation of the sound source (the horizontal orientation needs to be recorded accurately, while the vertical orientation can be recorded approximately).
[0200] Tags: "Response Device Index" indicating a system operating online ”;
[0201] Synchronization: NTP synchronization + fragment cross-correlation fine alignment is adopted to align the same wake word to the frame of each device with an accuracy of ±10 pm 10–20 ms.
[0202] Privacy: Only features are retained, the original voice is not stored; if it needs to be retained, it is anonymized and encrypted locally.
[0203] 2) Supplementing simulation data:
[0204] To cover rare geometries / orientations and extreme reverberation / noise, room impulse responses (RIRs) are generated using Pyroomacoustics’ directional sound sources / microphones and mirror source method, and then convolved onto near-field / far-field speech libraries, adding household noise (air conditioner, TV, kitchen, etc.) and interfering speech.
[0205] Human voice directionality model: The frequency trend of the latest DI data is fitted using a cross-band cardioid / order-boosted cosine model.
[0206] Air absorption: Set temperature and humidity according to ISO9613-1, and approximately overlay it onto G or explicitly model it in distant scenes.
[0207] Device response randomization: Domain randomization is applied to frequency response / automatic gain control (AGC) / quantization / noise floor to improve generalization.
[0208] Synchronization jitter: Simulate 5–30ms drift to train the model to be robust to mild misalignment.
[0209] 3) Data organization:
[0210] A single sample = a set of devices that are "the same wake-up event, the same room, and N concurrent device segments";
[0211] Input is , tag as ;
[0212] Training / verification / testing is divided into multiple sections based on room, speaker, and device to prevent leaks;
[0213] Mix real data with simulation data in small batches at a ratio of 1:1 or 2:1, prioritizing real data and supplementing with simulation data to complete the coverage.
[0214] Optionally, some alternative solutions in the embodiments of this application include:
[0215] 1) The last layer of Phys-Proj can be replaced with a monotonic neural network (Monotonic NN) or an equally spaced piecewise linear monotonic function to further explicitly guarantee monotonicity.
[0216] 2) Soft distribution You can also not use it Instead of using the inverse square law or inverse equivalent distance linear / power law normalization:
[0217] .in For exponents.
[0218] 3) F-Gate can be replaced with Squeeze-Excitation frequency domain attention or piecewise spline gating to improve interpretability.
[0219] 4) The model input can be changed to a multi-resolution LogFBank (short / medium / long windows side by side), or phase / group delay features can be added; if the device is a small array, the energy spectrum at the back end of the array can be added as an additional channel.
[0220] 5) Replace the list-based cross-entropy with a learning ranking loss such as ListNet / ListMLE; or use temperature-controlled soft labels (labeling positive examples...). Increasing the temperature factor (by one) leads to stable convergence.
[0221] 6) Where bandwidth allows, only 2–3 light features (such as high-frequency energy ratio / local signal-to-noise ratio (SNR) estimates) are swapped as arbitration secondary features to further improve the stability of boundary samples.
[0222] Through the above scheme, this application provides a unified arbitration concept for "equivalent distance value": the frequency-related directionality of human voice and its propagation / absorption are equated to a "scalar equivalent to distance," serving as the sole implicit physical factor for multi-device arbitration, unifying "proximity / orientation / energy" into a single metric. It provides supervised learning of monotonic scoring with "response index only": through in-group Softmax list-style target training, inference can stably output scores monotonically negatively correlated with "equivalent distance" without requiring group information. Utilizing the position / orientation and directionality models / DI statistics available during training, a soft target distribution is generated, approximated by KL distillation, enabling the model to achieve interpretable monotonicity without explicit distance labels. The F-Gate uses "high frequencies are more directional" as a prior for shape constraints and regularization, with a clear function. Training uses group sample supervision, and inference is independent at one end; arbitration is achieved by exchanging only the scoring scalar, without transmitting the original audio, satisfying privacy and bandwidth constraints.
[0223] Therefore, the embodiments of this application have the following advantages:
[0224] 1) Physically Interpretable: The model's output score is monotonically negatively correlated with the equivalent distance, directly reflecting the physical essence of "closest to hear" (directivity × distance × absorption). 2) Robust to Back-to-Near Devices: When the user's back is to the device, the sound source's directional gain decreases, the equivalent distance increases, the score decreases, and erroneous responses are naturally suppressed. 3) Training-Friendly: Only a "response device index" is needed, avoiding expensive distance labeling / full-field geometry reconstruction; physical distillation further improves generalization with limited real-world data. 4) Lightweight Deployment: A small, single-end model, inference only produces scalars, resulting in extremely low computational power / bandwidth / latency, suitable for the Internet of Things (IoT) edge. 5) Privacy and Security: No raw audio is exchanged / uploaded; only scores can be uploaded at the edge, satisfying the data minimization principle. 6) Compatible with Existing Systems: Seamlessly replaces the "maximum energy" or "closest to hear" strategy; can also be fused in parallel with context arbitration (task type / device capability). Related context arbitration patents / methods can be used as a secondary selector in the subsequent stage.
[0225] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods of the various embodiments of this application.
[0226] This embodiment also provides a wake-up device for a distributed voice interaction device, which is used to implement the above embodiments and preferred embodiments; details already described will not be repeated. As used below, the term "module" can refer to a combination of software and / or hardware that implements a predetermined function. Although the device described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.
[0227] Figure 6 This is a structural block diagram of an optional wake-up device for a distributed voice interaction device according to an embodiment of this application. The device includes:
[0228] Arbitration module 62 is used to arbitrate the equivalent distance between the sound source location of the wake-up audio and the voice interaction device based on the directional gain of the sound source orientation of the wake-up audio using an arbitration model, and obtain a first score that has a monotonic mapping relationship with the equivalent distance; wherein, the arbitration model is deployed in the voice interaction device, the equivalent distance has a first approximate relationship with the geometric distance and the directional gain, the geometric distance is the physical geometric distance between the voice interaction device and the sound source location, and the arbitration model learns the monotonic mapping relationship through pre-training;
[0229] The determining module 64 is used to determine, based on the first score, the target device among the plurality of voice interaction devices with the smallest equivalent distance to the sound source location, so as to respond to the wake-up audio through the target device.
[0230] Using the aforementioned device, an arbitration model arbitrates the equivalent distance between the sound source location of the wake-up audio and the voice interaction device based on the directional gain of the sound source orientation of the wake-up audio, obtaining a first score that has a monotonic mapping relationship with the equivalent distance. The arbitration model is deployed in the voice interaction device, and there is a first approximate relationship between the equivalent distance, the geometric distance, and the directional gain. The geometric distance is the physical geometric distance between the voice interaction device and the sound source location. The arbitration model is pre-trained to learn the monotonic mapping relationship. The first score determines the target device among multiple voice interaction devices with the smallest equivalent distance to the sound source location, allowing the target device to respond to the wake-up audio. Therefore, this technical solution solves the problem in related technologies where the voice source of the wake-up object in distributed wake-up is not isotropic, leading to false wake-ups in "nearby wake-up" (energy / amplitude) or "orientation wake-up" (direction) schemes; thus improving wake-up accuracy.
[0231] In an exemplary embodiment, the device further includes a training module, configured to: arbitrate the equivalent distance between the sound source location of the wake-up audio and the voice interaction device based on the directional gain of the sound source orientation of the wake-up audio using an arbitration model, and obtain a first score that has a monotonic mapping relationship with the equivalent distance; acquire training samples, wherein the training samples are a set of devices responding to each wake-up event in the same space, and the observation values of the training samples are audio features collected by each device in the set of devices; arbitrate the equivalent distance of the audio features using a target model, and output a second score corresponding to each device; determine the soft target distribution of the training samples using the geometric information corresponding to the training samples; determine the response probability of each device using the second score; determine the total loss of the target model using the response probability and the soft target distribution, and perform a backpropagation gradient update on the target model using the total loss to obtain the arbitration model.
[0232] In one exemplary embodiment, the training module is further configured to determine the response probability, including: ;in, This represents the response probability. This refers to the set of devices; This represents the device index for each of the aforementioned devices. This represents the device index of any device in the device set other than each of the aforementioned devices; For temperature parameters, , Indicates the second score. Indicates equipment The score is divided into two parts, with the second part including the score of the actual response device.
[0233] In an exemplary embodiment, the training module is further configured to determine the interval ranking loss corresponding to the target model using the second score; and to determine the total loss of the target model using the interval ranking loss, the response probability, and the soft target distribution; wherein determining the interval ranking loss includes: ;in, This represents the interval sorting loss. This represents the device index of the actual responding device in the device set, where m represents the actual responding device in the device set. With any device The interval between Indicates equipment The score, This represents the score of the actual response device output by the target model; where, This represents the device index of any device in the device set other than each of the aforementioned devices.
[0234] In an exemplary embodiment, the training module is further configured to determine the distillation loss of the target device using the response probability and the soft target distribution; and to determine the total loss of the target model using the distillation loss.
[0235] In an exemplary embodiment, the training module is further configured to determine the KL divergence between the response probability of each device and the soft target distribution to obtain the distillation loss, wherein the distillation loss is used to enable the target model to learn a monotonic mapping relationship between the second score and the equivalent distance corresponding to the audio feature.
[0236] In an exemplary embodiment, the training module is further configured to determine the cross-entropy loss of the target model using the response probability; and to determine the total loss of the target model using the cross-entropy loss, the response probability, and the soft target distribution.
[0237] In one exemplary embodiment, the training module is further configured to determine the cross-entropy loss, including: ;in, This represents the device index of the actual responding device in the device set. This represents the probability of a device's actual response. This represents the cross-entropy loss.
[0238] In an exemplary embodiment, the training module is further configured to determine the directional gain corresponding to the audio feature through a preset directional model; determine the equivalent distance corresponding to the audio feature through the directional gain corresponding to the audio feature, the geometric information, and the first approximation relationship; wherein the geometric information includes: the sound source location of each wake-up event, the sound source orientation of each wake-up event, and the device location of each device; and determine the soft target distribution through the equivalent distance corresponding to the audio feature.
[0239] In an exemplary embodiment, the training module is further configured to determine the interval ranking loss corresponding to the target model through the second score; determine the cross-entropy loss of the target model through the response probability; determine the distillation loss of the target model through the soft target distribution; and perform a weighted summation of the cross-entropy loss, the distillation loss, and the interval ranking loss to obtain the total loss.
[0240] In an exemplary embodiment, the device further includes an approximation module, configured to arbitrate the equivalent distance between the sound source location of the wake-up audio and the voice interaction device based on the directional gain of the wake-up audio's sound source orientation using an arbitration model, and before obtaining a first score that has a monotonic mapping relationship with the equivalent distance, determine a second approximation relationship between the geometric distance, the directional gain, and the relative intensity of the wake-up audio; and obtain a third approximation relationship with the audio intensity of the wake-up audio; and determine the first approximation relationship between the equivalent distance, the geometric distance, and the directional gain through the second approximation relationship and the third approximation relationship; wherein, the first approximation relationship is: ;in, This represents the equivalent distance. Represents the geometric distance, This refers to the directional gain.
[0241] In an exemplary embodiment, the arbitration module 62 is further configured to perform band-wise recalibration of the feature matrix of the target audio feature of the wake-up audio through a band selection gate, wherein the recalibrated feature matrix can highlight the high-frequency directivity of the target audio feature; process the recalibrated feature matrix through a time-frequency lightweight convolutional stack to obtain a time-frequency encoding matrix corresponding to the recalibrated feature matrix; perform weighted pooling on the time-frequency encoding matrix through a global context pooling layer to obtain a global representation corresponding to the time-frequency encoding matrix; and perform monotonic projection on the global representation through a physical consistency projection head to output the first score corresponding to the target audio feature; wherein the arbitration model includes: the band selection gate, the time-frequency lightweight convolutional stack, the global context pooling layer, and the physical consistency projection head.
[0242] In an exemplary embodiment, the arbitration module 62 is further configured to normalize the features of each frequency band in the feature matrix to obtain normalized feature values; generate an initial frequency band weight for each frequency band using the normalized feature values, wherein the initial frequency band weight increases with the frequency of the wake-up audio; normalize the initial frequency band weight to obtain normalized frequency band weights; and recalibrate the feature matrix band-by-band using the normalized frequency band weights.
[0243] In an exemplary embodiment, the determining module 64 is further configured to: send the first score to the arbitration device to instruct the arbitration device to determine, based on the first score, the target device among the plurality of voice interaction devices with the smallest equivalent distance to the sound source location, wherein the arbitration device includes one of the following: any of the other voice interaction devices among the plurality of voice interaction devices, or a cloud arbitration device; obtain a third score sent by the other voice interaction devices, and determine, based on the third score and the first score, the target device among the plurality of voice interaction devices with the smallest equivalent distance to the sound source location.
[0244] Embodiments of this application also provide a computer-readable storage medium storing a computer program, wherein the computer program is configured to execute the steps in any of the above method embodiments when run.
[0245] Optionally, in this embodiment, the storage medium may be configured to store a computer program for performing the following steps:
[0246] S1, the equivalent distance between the sound source location of the wake-up audio and the voice interaction device is arbitrated by an arbitration model based on the directional gain of the sound source orientation of the wake-up audio, and a first score with a monotonic mapping relationship with the equivalent distance is obtained; wherein, the arbitration model is deployed in the voice interaction device, the equivalent distance has a first approximate relationship with the geometric distance and the directional gain, the geometric distance is the physical geometric distance between the voice interaction device and the sound source location, and the arbitration model has learned the monotonic mapping relationship through pre-training;
[0247] S2, using the first score, determine the target device with the smallest equivalent distance to the sound source location among the plurality of voice interaction devices, so as to respond to the wake-up audio through the target device.
[0248] In one exemplary embodiment, the aforementioned computer-readable storage medium may include, but is not limited to, various media capable of storing computer programs, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard disk, magnetic disk, or optical disk.
[0249] Specific examples in this embodiment can be found in the examples described in the above embodiments and exemplary implementations, and will not be repeated here.
[0250] Embodiments of this application also provide an electronic device including a memory and a processor, the memory storing a computer program and the processor being configured to run the computer program to perform the steps in any of the above method embodiments.
[0251] Optionally, in this embodiment, the processor can be configured to perform the following steps via a computer program:
[0252] S1, the equivalent distance between the sound source location of the wake-up audio and the voice interaction device is arbitrated by an arbitration model based on the directional gain of the sound source orientation of the wake-up audio, and a first score with a monotonic mapping relationship with the equivalent distance is obtained; wherein, the arbitration model is deployed in the voice interaction device, the equivalent distance has a first approximate relationship with the geometric distance and the directional gain, the geometric distance is the physical geometric distance between the voice interaction device and the sound source location, and the arbitration model has learned the monotonic mapping relationship through pre-training;
[0253] S2, using the first score, determine the target device with the smallest equivalent distance to the sound source location among the plurality of voice interaction devices, so as to respond to the wake-up audio through the target device.
[0254] In one exemplary embodiment, the electronic device may further include a transmission device and an input / output device, wherein the transmission device is connected to the processor and the input / output device is connected to the processor.
[0255] Embodiments of this application also provide a computer program product, which includes a computer program that, when executed by a processor, implements the steps in any of the above method embodiments.
[0256] Embodiments of this application also provide another computer program product, including a non-volatile computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps in any of the above method embodiments.
[0257] Embodiments of this application also provide a computer program that includes computer instructions stored in a computer-readable storage medium; a processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the steps in any of the above method embodiments.
[0258] Specific examples in this embodiment can be found in the examples described in the above embodiments and exemplary implementations, and will not be repeated here.
[0259] Obviously, those skilled in the art should understand that the modules or steps of this application described above can be implemented using general-purpose computing devices. They can be centralized on a single computing device or distributed across a network of multiple computing devices. They can be implemented using computer-executable program code, and thus can be stored in a storage device for execution by a computing device. In some cases, the steps shown or described can be performed in a different order than those presented here, or they can be fabricated as separate integrated circuit modules, or multiple modules or steps can be fabricated as a single integrated circuit module. Thus, this application is not limited to any particular combination of hardware and software.
[0260] The above description is only a preferred embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this application, and these improvements and modifications should also be considered within the scope of protection of this application.
Claims
1. A wake-up method for a distributed voice interaction device, characterized in that, include: An arbitration model arbitrates the equivalent distance between the sound source location of the wake-up audio and the voice interaction device based on the directional gain of the sound source orientation of the wake-up audio, obtaining a first score that has a monotonic mapping relationship with the equivalent distance; wherein, the arbitration model is deployed in the voice interaction device, and there is a first approximate relationship between the equivalent distance, the geometric distance, and the directional gain, wherein the geometric distance is the physical geometric distance between the voice interaction device and the sound source location, and the arbitration model is pre-trained to learn the monotonic mapping relationship; wherein, the target audio features of the wake-up audio are input into the arbitration model; The method determines a target device with the smallest equivalent distance to the sound source location among a plurality of voice interaction devices based on the first score, so as to respond to the wake-up audio through the target device. The method further includes, before obtaining the first score which has a monotonic mapping relationship with the equivalent distance: Acquire training samples, wherein the training samples are a set of devices in the same space that respond to each wake-up event, and the observation values of the training samples are the audio features collected by each device in the set of devices; The audio features are arbitrated using an equivalent distance using a target model to output a second score for each device; the soft target distribution of the training samples is determined using the geometric information corresponding to the training samples; and the response probability of each device is determined using the second score. The total loss of the target model is determined by the response probability and the soft target distribution, and the target model is updated by reverse gradient using the total loss to obtain the arbitration model.
2. The wake-up method for a distributed voice interaction device according to claim 1, characterized in that, The response probability of each device is determined by the second score. include: Determining the response probability includes: ; in, This represents the response probability. This refers to the set of devices; This represents the device index for each of the aforementioned devices. This represents the device index of any device in the device set other than each of the aforementioned devices; For temperature parameters, , Indicates the second score. Indicates equipment The score is divided into two parts, with the second part including the score of the actual response device.
3. The wake-up method for a distributed voice interaction device according to claim 1, characterized in that, The total loss of the target model is determined by the response probability and the soft target distribution, including: The second score is used to determine the interval ranking loss corresponding to the target model; The total loss of the target model is determined by the interval sorting loss, the response probability, and the soft target distribution. Determining the interval sorting loss includes: ; in, This represents the interval sorting loss. This represents the device index of the actual responding device in the device set, where m represents the actual responding device in the device set. With any device The interval between Indicates equipment The score, This represents the score of the actual response device output by the target model; where, This represents the device index of any device in the device set other than each of the aforementioned devices.
4. The wake-up method for a distributed voice interaction device according to claim 1, characterized in that, The total loss of the target model is determined by the response probability and the soft target distribution, including: The distillation loss of the target device is determined by the response probability and the soft target distribution; The total loss of the target model is determined by the distillation loss.
5. The wake-up method for a distributed voice interaction device according to claim 4, characterized in that, Determining the distillation loss of the target device using the response probability and the soft target distribution includes: The KL divergence between the response probability of each device and the soft target distribution is determined to obtain the distillation loss, wherein the distillation loss is used to enable the target model to learn a monotonic mapping relationship between the second score and the equivalent distance corresponding to the audio feature.
6. The wake-up method for a distributed voice interaction device according to claim 1, characterized in that, The total loss of the target model is determined by the response probability and the soft target distribution, including: The cross-entropy loss of the target model is determined by the response probability; The total loss of the target model is determined by the cross-entropy loss, the response probability, and the soft target distribution.
7. The wake-up method for a distributed voice interaction device according to claim 6, characterized in that, The cross-entropy loss of the target model is determined by the response probability. include: Determining the cross-entropy loss includes: ; in, This represents the device index of the actual responding device in the device set. This represents the probability of a device's actual response. This represents the cross-entropy loss.
8. The wake-up method for a distributed voice interaction device according to claim 1, characterized in that, Determining the soft target distribution of the training samples using the geometric information corresponding to the training samples includes: The directional gain corresponding to the audio feature is determined by a preset directional model; The equivalent distance corresponding to the audio feature is determined by the directional gain corresponding to the audio feature, the geometric information, and the first approximation relationship; wherein, the geometric information includes: the sound source location of each wake-up event, the sound source orientation of each wake-up event, and the device location of each device; The distribution of soft targets is determined by the equivalent distance corresponding to the audio features.
9. The wake-up method for a distributed voice interaction device according to claim 1, characterized in that, The total loss of the target model is determined by the response probability and the soft target distribution, including: The second score is used to determine the interval ranking loss corresponding to the target model; the response probability is used to determine the cross-entropy loss of the target model; and the soft target distribution is used to determine the distillation loss of the target model. The total loss is obtained by weighted summing of the cross-entropy loss, the distillation loss, and the interval sorting loss.
10. The wake-up method for a distributed voice interaction device according to claim 1, characterized in that, Before arbitrating the equivalent distance between the sound source location of the wake-up audio and the voice interaction device using an arbitration model based on the directional gain of the sound source orientation of the wake-up audio, and obtaining a first score that has a monotonic mapping relationship with the equivalent distance, the method further includes: Determine a second approximate relationship between the geometric distance, the directional gain, and the relative intensity of the wake-up audio; and obtain a third approximate relationship with the audio intensity of the wake-up audio; The first approximation relationship between the equivalent distance, the geometric distance, and the directional gain is determined by the second approximation relationship and the third approximation relationship; The first approximation relationship is: ; in, This represents the equivalent distance. Represents the geometric distance, This represents the directional gain.
11. The wake-up method for a distributed voice interaction device according to claim 1, characterized in that, The arbitration model arbitrates the equivalent distance between the wake-up audio source location and the voice interaction device based on the directional gain of the wake-up audio source orientation, obtaining a first score that has a monotonic mapping relationship with the equivalent distance, including: The feature matrix of the target audio feature of the wake-up audio is recalibrated band by band by band using a frequency band selection gate, wherein the recalibrated feature matrix can highlight the high-frequency directivity of the target audio feature. The recalibrated feature matrix is processed by a time-frequency lightweight convolution stack to obtain the time-frequency coding matrix corresponding to the recalibrated feature matrix. The time-frequency coding matrix is weighted and pooled through a global context pooling layer to obtain the global representation of the time-frequency coding matrix; The global representation is monotonically projected using a physically consistent projection head, and the first score corresponding to the target audio feature is output. The arbitration model includes: the frequency band selection gate, the time-frequency lightweight convolution stack, the global context pooling layer, and the physical consistency projection head.
12. The wake-up method for a distributed voice interaction device according to claim 11, characterized in that, The feature matrix of the target audio features of the wake-up audio is recalibrated band-by-band using a frequency band selection gate, including: The features of each frequency band in the feature matrix are normalized to obtain normalized feature values; The initial frequency band weights for each frequency band are generated using the normalized feature values, wherein the initial frequency band weights increase with the frequency of the wake-up audio. The initial frequency band weights are normalized to obtain normalized frequency band weights; The feature matrix is recalibrated band-by-band using the normalized band weights.
13. The wake-up method for a distributed voice interaction device according to claim 1, characterized in that, The target device with the smallest equivalent distance to the sound source location among the plurality of voice interaction devices is determined by the first score, including at least one of the following: The first score is sent to the arbitration device to instruct the arbitration device to determine the target device with the smallest equivalent distance to the sound source location among the plurality of voice interaction devices based on the first score, wherein the arbitration device includes one of the following: any device among the other voice interaction devices among the plurality of voice interaction devices, or a cloud arbitration device; Obtain the third score sent by the other voice interaction devices, and determine the target device with the smallest equivalent distance to the sound source location among the multiple voice interaction devices based on the third score and the first score.
14. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored program, wherein the program, when executed, performs the method of any one of claims 1 to 13.
15. An electronic device comprising a memory and a processor, characterized in that, The memory stores a computer program, and the processor is configured to execute the method of any one of claims 1 to 13 through the computer program.
Citation Information
Patent Citations
Systems and methods for selective wake word detection using neural network models
CA3067776A1
Device waking up method and system for acoustic networking
CN110288997A