Wake-up device identification method, device, equipment, storage medium and program product
By adopting the method of phase adaptation calculation and gain difference fusion in intelligent voice devices, gain calibration is performed on the wake-up voice received by the voice device, which solves the problem of inaccurate recognition caused by gain differences in traditional methods, and improves the accuracy and user experience of device recognition.
Patent Information
- Application Number
- CN202411142240.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-19
- Publication Date
- 2025-05-02
- Estimated Expiration
- 2044-08-19
AI Technical Summary
The traditional wake-up device recognition method has different gains under the influence of various factors, resulting in the voice energy after gain being unable to accurately measure the distance between the device and the user, reducing the accuracy of recognition and affecting the user experience.
The target gain difference between the target voice device and the reference voice device is calculated based on the stage adaptation method of the voice device life cycle stage, and the gain difference in different stages is fused to obtain the fused gain difference, and then the voice energy of the user's wake-up voice received by the target voice device is calibrated to eliminate the gain difference.
The voice energy gain of each target voice device is achieved at the same level as the reference voice device, accurately measure the distance between the device and the user, and improve the accuracy of wake-up device recognition and user experience.
Smart Images

Figure CN119068876B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of nearby wake-up technology, and in particular to a wake-up device identification method, device, equipment, storage medium and program product. Background Art
[0002] With the popularity of smart voice homes, there will be multiple smart voice devices in the same room or area at the same time. When the user shouts the wake-up word, it is necessary to identify the voice device closest to the user from multiple smart voice devices and wake it up in order to receive and execute the user's subsequent voice commands.
[0003] Since the energy of the user's wake-up voice received by the voice device closer to the user is usually larger, the traditional wake-up device identification method can identify the wake-up device by the energy of the user's wake-up voice received by each voice device. However, each voice device has a gain effect on the received user wake-up voice, and the voice energy used to identify the wake-up device is actually the voice energy after gain. The gain of different voice devices under the influence of various factors is different, resulting in the voice energy after gain cannot accurately measure the distance between the corresponding voice device and the user, reducing the accuracy of wake-up device recognition and affecting the user experience. Summary of the invention
[0004] The present invention aims to solve at least one of the technical problems existing in the related art. To this end, the present invention proposes a wake-up device identification method, which can minimize the gap between the speech energy gain of each target voice device and the speech energy gain of the reference voice device under the interference of various factors, so that the speech energy gain of each target voice device is kept at the same level as the speech energy gain of the reference voice device, and then the speech energy after gain of each target voice device can accurately measure the distance between the corresponding voice device and the user, thereby improving the accuracy of wake-up device identification and simultaneously improving the user experience.
[0005] The present invention also provides a wake-up device identification apparatus, a device, a storage medium and a program product.
[0006] According to a first aspect of the present invention, a method for waking up a device for identifying an object includes:
[0007] Based on the life cycle stage of the voice device, a target gain difference between each target voice device and a reference voice device in the target network is calculated by a stage adaptation method; the reference voice device is any voice device in the target network, and the target voice device is any voice device in the target network except the reference voice device;
[0008] The target gain differences corresponding to each target voice device at multiple life cycle stages are respectively fused to obtain the fused gain differences corresponding to each target voice device;
[0009] Based on the fusion gain difference corresponding to each target voice device, the voice energy of the user wake-up voice received by each target voice device is gain calibrated to obtain the calibrated voice energy of each target voice device;
[0010] Determine a maximum target speech energy from each target speech energy, and identify a speech device corresponding to the maximum target speech energy as a wake-up device;
[0011] The target speech energy includes the calibrated speech energy of each target speech device.
[0012] According to the wake-up device identification method of the embodiment of the present invention, based on the different life cycle stages of the voice device, different calculation methods adapted to each stage are used to calculate the target gain difference between each target voice device and the reference voice device in the same network. By covering the life cycle stages, various factors that cause the target gain difference between each target voice device and the reference voice device can be covered as much as possible, and the accurate target gain difference under the interference of various factors can be obtained through targeted calculation methods. The target gain difference of the target voice device in different stages is then fused to obtain the fused gain difference under the interference of various factors. Based on the fused gain difference, the voice energy of the user wake-up voice received by the target voice device is gain calibrated to minimize the gap between the voice energy gain of each target voice device and the voice energy gain of the reference voice device under the interference of various factors, so that the voice energy gain of each target voice device is kept at the same level as the voice energy gain of the reference voice device, so that the voice energy gain of each target voice device after gain can accurately measure the distance between the corresponding voice device and the user, thereby improving the accuracy of wake-up device identification and simultaneously improving user experience.
[0013] According to an embodiment of the present invention, the target speech energy also includes speech energy of the user wake-up speech received by the reference speech device.
[0014] According to an embodiment of the present invention, the target gain difference corresponding to any target voice device in the user use stage is determined based on the following method:
[0015] When each voice device is in a gain stable state, a target gain difference corresponding to the target voice device is calculated.
[0016] According to one embodiment of the present invention, whether any voice device is in a gain stable state is determined based on the following method:
[0017] Calculating the gains of multiple first preset duration background noises collected by the voice device during the late night period to obtain multiple gains;
[0018] Obtaining a gain variation coefficient of the speech device based on a gain mean of the multiple gains and a gain standard deviation of the multiple gains;
[0019] If the gain variation coefficient of the voice device is less than the coefficient threshold, it is determined that the voice device is in a gain stable state.
[0020] According to an embodiment of the present invention, the calculating the target gain difference corresponding to the target voice device includes:
[0021] Calculating the difference between the gain mean of the target voice device and the gain mean of the reference voice device to obtain a current gain difference corresponding to the target voice device;
[0022] Based on a weighted sum of the current gain difference and the historical gain differences corresponding to the target voice device, a target gain difference corresponding to the target voice device is obtained.
[0023] According to an embodiment of the present invention, the gain of any first preset duration background noise is determined based on the following method:
[0024] Calculating power spectra of multiple second preset duration background noises in the first preset duration background noise to obtain multiple power spectra;
[0025] Based on an average value of the multiple power spectra, a gain of the first preset duration background noise is determined.
[0026] According to one embodiment of the present invention, the gain mean and the gain standard deviation are determined based on the following method:
[0027] Eliminating the maximum value and the minimum value among the multiple gains to obtain multiple gains to be processed;
[0028] The average value and the standard deviation of the plurality of gains to be processed are calculated to obtain the gain mean value and the gain standard deviation.
[0029] According to an embodiment of the present invention, the target gain difference corresponding to any target voice device in the user use stage is determined based on the following method:
[0030] The difference between the historical gain of the target voice device and the historical gain of the reference voice device is calculated to obtain a target gain difference corresponding to the target voice device.
[0031] According to one embodiment of the present invention, the historical gain of any voice device is determined based on the following method:
[0032] Calculating the gain of the target frequency band speech in the wake-up speech collected by the speech device in multiple historical wake-up time periods to obtain multiple gains to be processed;
[0033] The historical gain is determined based on an average value of the plurality of gains to be processed.
[0034] According to one embodiment of the present invention, the target gain difference corresponding to any target speech device in the target stage is determined based on the following method:
[0035] Calculating the difference between the frequency gain of the target voice device and the frequency gain of the reference voice device to obtain a target gain difference corresponding to the target voice device;
[0036] The target stage includes a design stage and / or a factory delivery stage.
[0037] According to an embodiment of the present invention, the frequency gain of any voice device is determined based on the following method:
[0038] Calculating the gain of each frequency point in the swept-frequency sound signal collected by the voice device in the target stage to obtain a plurality of gains to be processed; the swept-frequency sound signal is played by the speaker at at least one angular position relative to the voice device and after reaching equilibrium;
[0039] The frequency point gain is determined based on an average value of the multiple gains to be processed.
[0040] According to an embodiment of the present invention, the target gain differences corresponding to each target voice device at multiple life cycle stages are respectively fused to obtain the fused gain differences corresponding to each target voice device, including:
[0041] The target gain differences corresponding to each target voice device at multiple life cycle stages are weighted and summed to obtain the fusion gain difference corresponding to each target voice device;
[0042] The determining factors of the weight of the target gain difference include at least one of the stage selection, the networking time of each voice device, or the wake-up times of each voice device.
[0043] According to a second aspect of the present invention, a wake-up device identification device includes:
[0044] A target gain difference calculation module is used to: calculate the target gain difference between each target voice device and a reference voice device in a target network by adopting a stage adaptation method based on the life cycle stage of the voice device; the reference voice device is any voice device in the target network, and the target voice device is any voice device in the target network except the reference voice device;
[0045] The target gain difference fusion module is used to: fuse the target gain differences corresponding to each target voice device at multiple life cycle stages respectively to obtain the fused gain differences corresponding to each target voice device;
[0046] A gain calibration module is used to: perform gain calibration on the speech energy of the user wake-up speech received by each target speech device based on the fusion gain difference corresponding to each target speech device, so as to obtain the calibrated speech energy of each target speech device;
[0047] A wake-up device identification module is used to: determine the maximum target voice energy from each target voice energy, and identify the voice device corresponding to the maximum target voice energy as the wake-up device;
[0048] The target speech energy includes the calibrated speech energy of each target speech device.
[0049] According to the wake-up device recognition device of the embodiment of the present invention, based on the different life cycle stages of the voice device, different calculation methods adapted to each stage are used to calculate the target gain difference between each target voice device and the reference voice device in the same network. By covering the life cycle stages, various factors that cause the target gain difference between each target voice device and the reference voice device can be covered as much as possible, and the accurate target gain difference under the interference of various factors can be obtained through targeted calculation methods. The target gain difference of the target voice device in different stages is then fused to obtain the fused gain difference under the interference of various factors. Based on the fused gain difference, the voice energy of the user wake-up voice received by the target voice device is gain calibrated to minimize the gap between the voice energy gain of each target voice device and the voice energy gain of the reference voice device under the interference of various factors, so that the voice energy gain of each target voice device is kept at the same level as the voice energy gain of the reference voice device, so that the voice energy gain of each target voice device after gain can accurately measure the distance between the corresponding voice device and the user, thereby improving the accuracy of wake-up device recognition and simultaneously improving user experience.
[0050] According to an embodiment of the third aspect of the present invention, an electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, and when the processor executes the program, the wake-up device identification method described in the embodiment of the first aspect is implemented.
[0051] According to the non-transitory computer-readable storage medium of the fourth aspect embodiment of the present invention, a computer program is stored thereon, and when the computer program is executed by a processor, the wake-up device identification method described in the first aspect embodiment is implemented.
[0052] A computer program product according to an embodiment of the fifth aspect of the present invention includes a computer program, which, when executed by a processor, implements the wake-up device identification method described in the embodiment of the first aspect.
[0053] The above one or more technical solutions in the embodiments of the present invention have at least one of the following technical effects:
[0054] The present invention is based on different life cycle stages of voice devices, and adopts different calculation methods adapted to each stage to calculate the target gain difference between each target voice device and the reference voice device in the same network. By covering the life cycle stages, various factors that cause the target gain difference between each target voice device and the reference voice device can be covered as much as possible, and accurate target gain differences under the interference of various factors can be obtained through targeted calculation methods. The target gain differences of the target voice devices in different stages are then fused to obtain the fused gain differences under the interference of various factors. Based on the fused gain difference, the voice energy of the user wake-up voice received by the target voice device is gain calibrated, and the gap between the voice energy gain of each target voice device and the voice energy gain of the reference voice device under the interference of various factors is eliminated to the maximum extent, so that the voice energy gain of each target voice device is kept at the same level as the voice energy gain of the reference voice device, and then the voice energy after gain of each target voice device can accurately measure the distance between the corresponding voice device and the user, thereby improving the accuracy of wake-up device recognition and simultaneously improving user experience.
[0055] Additional aspects and advantages of the present invention will be given in part in the following description and in part will be obvious from the following description, or will be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or related technologies, the drawings required for use in the embodiments or related technical descriptions are briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0057] Figure 1 This is one of the flow charts of the wake-up device identification method provided by an embodiment of the present invention;
[0058] Figure 2 This is a second flow chart of the wake-up device identification method provided by an embodiment of the present invention;
[0059] Figure 3 This is a third flow chart of the wake-up device identification method provided by an embodiment of the present invention;
[0060] Figure 4 This is a fourth flow chart of the wake-up device identification method provided by an embodiment of the present invention;
[0061] Figure 5 This is a fifth flow chart of the wake-up device identification method provided by an embodiment of the present invention;
[0062] Figure 6 is a structural diagram of a wake-up device identification apparatus provided by an embodiment of the present invention;
[0063] Figure 7 It is a schematic diagram of the physical structure of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0064] The following embodiments of the present invention are described in further detail in conjunction with the accompanying drawings and examples. The following examples are used to illustrate the present invention, but are not intended to limit the scope of the present invention.
[0065] In the description of the embodiments of the present invention, it should be noted that the terms "center", "longitudinal", "lateral", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside" and the like indicate positions or positional relationships based on the positions or positional relationships shown in the accompanying drawings, and are only for the convenience of describing the embodiments of the present invention and simplifying the description, and do not indicate or imply that the devices or elements referred to must have a specific orientation, be constructed and operate in a specific orientation, and therefore cannot be understood as limitations on the embodiments of the present invention. In addition, the terms "first", "second", and "third" are used for descriptive purposes only and cannot be understood as indicating or implying relative importance.
[0066] In the description of the embodiments of the present invention, it should be noted that, unless otherwise clearly specified and limited, the terms "connected" and "connection" should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection or an indirect connection through an intermediate medium. For ordinary technicians in this field, the specific meanings of the above terms in the embodiments of the present invention can be understood according to specific circumstances.
[0067] In the embodiments of the present invention, unless otherwise clearly specified and limited, the first feature being "above" or "below" the second feature may mean that the first and second features are in direct contact, or the first and second features are in indirect contact through an intermediate medium. Moreover, the first feature being "above", "above" and "above" the second feature may mean that the first feature is directly above or obliquely above the second feature, or simply means that the first feature is higher in level than the second feature. The first feature being "below", "below" and "below" the second feature may mean that the first feature is directly below or obliquely below the second feature, or simply means that the first feature is lower in level than the second feature.
[0068] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the embodiment of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples, without contradiction.
[0069] Figure 1 This is one of the flowcharts of the wake-up device identification method provided by an embodiment of the present invention. Figure 1 , an embodiment of the present invention provides a wake-up device identification method, which may include:
[0070] 101. Based on the life cycle stage of the voice device, a target gain difference between each target voice device and a reference voice device in the target network is calculated by using a stage adaptation method;
[0071] The reference voice device is any voice device in the target network, and the target voice device is any voice device in the target network except the reference voice device;
[0072] 102. Fusion the target gain differences corresponding to each target voice device at multiple life cycle stages to obtain a fusion gain difference corresponding to each target voice device;
[0073] 103. Based on the fusion gain difference corresponding to each target voice device, gain calibration is performed on the voice energy of the user wake-up voice received by each target voice device to obtain the calibrated voice energy of each target voice device;
[0074] 104. Determine a maximum target speech energy from each target speech energy, and identify a speech device corresponding to the maximum target speech energy as a wake-up device.
[0075] The target speech energy includes the calibrated speech energy of each target speech device.
[0076] In step 101, at any life cycle stage, the target gain difference between the target voice device and the reference voice device in the target network is calculated based on the gains of the two for the same sound.
[0077] Assuming that the target network includes voice device A and voice device B, where voice device A is the reference voice device and voice device B is the target voice device, the corresponding target gain difference can be calculated for voice device B in each life cycle stage, thereby obtaining multiple target gain differences that correspond one-to-one between voice device B and multiple life cycle stages.
[0078] In step 102, for example, multiple target gain differences corresponding to the voice device B and multiple life cycle stages are fused to obtain a fused gain difference corresponding to the voice device B.
[0079] Any fusion algorithm may be used to fuse the target gain differences, which is not limited here.
[0080] In step 103, for example, after voice device B receives the user wake-up voice, the voice energy of the user wake-up voice received by it is gain calibrated using its fusion gain difference, so that its voice energy gain is maintained at the same level as the voice energy gain of the same user wake-up voice received by the reference voice device, and the calibrated voice energy of voice device B is obtained.
[0081] In step 104, assuming that the target network also includes voice device C and voice device D, the same steps as above are performed on these two voice devices to obtain the calibrated voice energy of voice device C and voice device D, and the voice device corresponding to the maximum calibrated voice energy among the calibrated voice energies corresponding to voice device B, voice device C and voice device D is identified as the wake-up device.
[0082] Furthermore, the target speech energy may also include the speech energy of the user wake-up speech received by the reference speech device, that is, comparing the speech energy of the reference speech device with the calibrated speech energy of each target speech device, and identifying the speech device corresponding to the maximum speech energy as the wake-up device.
[0083] It should be noted that, since the microphone of the voice device is usually used to gain the received user wake-up voice, the processing of each gain in this embodiment can also be the processing of the microphone gain of each voice device. In addition, the gain calibration of the voice energy of each target voice device is performed on the voice energy of the same user wake-up voice received by each target voice device.
[0084] The wake-up device identification method provided in this embodiment calculates the target gain difference between each target voice device and the reference voice device in the same network group based on the different life cycle stages of the voice device, using different calculation methods adapted to each stage. By covering the life cycle stages, various factors that cause the target gain difference between each target voice device and the reference voice device can be covered as much as possible, and the accurate target gain difference under the interference of various factors can be obtained through targeted calculation methods. The target gain difference of the target voice device in different stages is then fused to obtain the fused gain difference under the interference of various factors. Based on the fused gain difference, the voice energy of the user wake-up voice received by the target voice device is gain calibrated, and the gap between the voice energy gain of each target voice device and the voice energy gain of the reference voice device under the interference of various factors is eliminated to the maximum extent, so that the voice energy gain of each target voice device is kept at the same level as the voice energy gain of the reference voice device, and then the voice energy after gain of each target voice device can accurately measure the distance between the corresponding voice device and the user, thereby improving the accuracy of wake-up device identification and simultaneously improving user experience.
[0085] Figure 2 This is a second flow chart of the wake-up device identification method provided by an embodiment of the present invention. Figure 2 In one embodiment, the target gain difference corresponding to any target speech device in the user use stage can be determined based on the following method:
[0086] 201. Calculate gains of multiple first preset duration background noises collected by the voice device during the late night period to obtain multiple gains;
[0087] 202. Obtain a gain variation coefficient of the voice device based on a gain mean of the multiple gains and a gain standard deviation of the multiple gains;
[0088] 203. If the gain variation coefficient of the voice device is less than the coefficient threshold, it is determined that the voice device is in a gain stable state;
[0089] 204. When each voice device is in a gain stable state, calculate the difference between the gain mean of the target voice device and the gain mean of the reference voice device to obtain a current gain difference corresponding to the target voice device;
[0090] 205. Obtain a target gain difference corresponding to the target voice device based on a weighted sum of the current gain difference and the historical gain difference corresponding to the target voice device.
[0091] Assuming that each voice device is in a very quiet environment late at night and the external noise is very low, the circuit noise floor of each voice device is greater than the external noise, and the sound collected by each voice device is its own circuit noise floor.
[0092] In step 201, the first preset duration background noise is the circuit background noise of the first preset duration, and each voice device collects multiple first preset duration background noises of its own circuit. The first preset duration can be set according to actual needs and is not limited here. The number of first preset duration background noises collected by each voice device can also be set according to actual needs and is not limited here.
[0093] In this embodiment, the first preset time length can be set to 10 seconds, and the number of first preset time length background noises collected by each voice device can be set to 10. Then, starting from 2 a.m., each voice device can collect 10 seconds of its own circuit background noise every 10 minutes, for a total of 10 times, so there are 10 10-second self-circuit background noises, and 10 gains are calculated.
[0094] In step 202, the ratio of the gain standard deviation to the gain mean may be used as the gain variation coefficient of the speech device to measure the dispersion of the gain of the speech device.
[0095] In step 203, when the gain variation coefficient of the voice device is less than the coefficient threshold, it indicates that the gain dispersion is low and the gain of the voice device is relatively stable. The coefficient threshold can be set according to actual needs and is not limited here. In this embodiment, the coefficient threshold can be set to 0.03.
[0096] In step 204, the mean gain corresponding to each voice device is taken as its final gain, and the current gain difference corresponding to the target voice device is calculated.
[0097] Microphones of different voice devices age slowly, which can cause changes in device gain. This embodiment calculates the target gain difference between the target voice device and the reference voice device based on the circuit noise floor gain of each voice device in the late night period during the user use stage. This allows the target gain difference to accurately measure the gain gap between each target voice device and the reference voice device when the microphone is aged, which helps to subsequently calibrate the voice energy of each voice device based on the target gain difference, so that the voice energy gain of each voice device remains at the same level.
[0098] In one embodiment, the gain of any first preset duration background noise may be determined based on the following method:
[0099] The power spectra of the plurality of second preset duration background noises in the first preset duration background noise are calculated to obtain a plurality of power spectra, and the gain of the first preset duration background noise is determined based on an average value of the plurality of power spectra.
[0100] The second preset duration background noise is the second preset duration background noise segment of the 10-second duration self-circuit background noise collected by the voice device. The second preset duration can be set according to actual needs and is not limited here. In this embodiment, the second preset duration can be set to 0.1 second. For the 10-second duration self-circuit background noise collected by the voice device, the power spectrum is calculated every 0.1 second to obtain 100 power spectra corresponding to each 10-second duration self-circuit background noise. The average value of the 100 power spectra is calculated to obtain the gain of the 10-second duration self-circuit background noise.
[0101] In this embodiment, for each first preset duration background noise, the average value of the power spectra of multiple background noise segments is used as the gain of the first preset duration background noise. The power spectra of each background noise segment can be smoothed by averaging, thereby reducing the gain error of the first preset duration background noise.
[0102] In one embodiment, the gain mean and the gain standard deviation may be determined based on the following method:
[0103] The maximum value and the minimum value among the multiple gains are eliminated to obtain multiple gains to be processed, and the average value and the standard deviation of the multiple gains to be processed are calculated to obtain the gain mean value and the gain standard deviation.
[0104] As mentioned above, the voice device has 10 10-second-long circuit background noises of its own, so there are 10 gains in total. After removing the maximum and minimum values of these 10 gains, 8 gains to be processed are left. The average and standard deviation of the remaining 8 gains to be processed are calculated to obtain the gain mean and gain standard deviation.
[0105] In this embodiment, by removing the maximum value, outliers in multiple gains can be effectively eliminated, so that the calculation of the mean and the standard deviation is more accurate.
[0106] Figure 3 This is a flowchart of the third embodiment of the wake-up device identification method provided by the present invention. Figure 3 In one embodiment, the target gain difference corresponding to any target speech device in the user use stage can be determined based on the following method:
[0107] 301. Calculate the gain of the target frequency band speech in the wake-up speech collected by the speech device in multiple historical wake-up periods to obtain multiple gains to be processed;
[0108] 302. Determine a historical gain based on an average value of a plurality of gains to be processed;
[0109] 303. Calculate the difference between the historical gain of the target voice device and the historical gain of the reference voice device to obtain a target gain difference corresponding to the target voice device.
[0110] In step 301, the wake-up speech start point and end point detection module can be used to determine the wake-up period, then the wake-up speech in the wake-up period is intercepted, and the target frequency band speech is obtained from the wake-up speech using a bandpass filter, and then the gain of the target frequency band speech is calculated.
[0111] In step 302, each voice device calculates and saves the gain of the corresponding target frequency band voice after each wake-up. Therefore, for each voice device, multiple to-be-processed gains corresponding to multiple historical wake-up periods are stored. The moving average method can be used to calculate the average value of the multiple to-be-processed gains of each voice device, and the average value is determined as the historical gain.
[0112] In step 303, the historical gain corresponding to each voice device is used as its final gain, and the target gain difference corresponding to the target voice device is calculated.
[0113] When the amount of data is small, the distance between the user and each voice device is relatively random, resulting in unstable target gain differences corresponding to each voice device. Therefore, it can be set that the steps of this embodiment are executed to calculate the target gain difference only when the number of wake-up times reaches a threshold. The threshold can be set according to actual needs and is not limited here. In this embodiment, the threshold can be set to 2000.
[0114] There will be certain obstructions in the installation environment of the voice device, which will cause its gain to change. This embodiment calculates the target gain difference between the target voice device and the reference voice device based on the gain of each voice device in the historical wake-up period during the user use stage, so that the target gain difference can accurately measure the gain gap between each target voice device and the reference voice device when there is obstruction in the actual installation environment, which is helpful for the subsequent calibration of the voice energy of each voice device based on the target gain difference, so that the voice energy gain of each voice device remains at the same level.
[0115] Figure 4 This is a fourth flow chart of the wake-up device identification method provided by an embodiment of the present invention. Figure 4 In one embodiment, the target gain difference corresponding to any target speech device at the target stage can be determined based on the following method:
[0116] 401. Calculate the gain of each frequency point in the swept frequency sound signal collected by the voice device in the target phase to obtain a plurality of gains to be processed;
[0117] The swept frequency sound signal is played by the loudspeaker at at least one angular position relative to the speech device and after being balanced;
[0118] 402. Determine a frequency point gain based on an average value of a plurality of gains to be processed;
[0119] 403. Calculate the difference between the frequency gain of the target voice device and the frequency gain of the reference voice device to obtain a target gain difference corresponding to the target voice device.
[0120] Target phases include the design phase and / or the factory phase.
[0121] Before step 401, the speaker can be placed 1 meter in front of the voice device and the speaker can be balanced using a standard microphone so that when the sound signal played by the speaker reaches the voice device, the sound pressure level reaches a preset decibel. The preset decibel can be set according to actual needs and is not limited here. In this embodiment, the preset decibel can be set to 94 decibels.
[0122] After the speaker is equalized, it can be controlled to play the swept frequency sound signal, and the voice device will collect the swept frequency sound signal.
[0123] Furthermore, the speaker can be placed at other angular positions 1 meter away from the voice device for equalization, and then controlled to play the swept-frequency sound signal. The voice device collects the swept-frequency sound signal again, and so on. The voice device can obtain the swept-frequency sound signals played by the equalized speaker at multiple angular positions.
[0124] In step 401, during the design stage, each voice device can collect swept-frequency sound signals from multiple angular positions, so the gain of each frequency point in the multiple swept-frequency sound signals collected by each voice device needs to be calculated separately; during the factory stage, based on the limitation of factory inspection efficiency, each voice device can only collect swept-frequency sound signals from the front, so the gain of each frequency point in the single swept-frequency sound signal collected by each voice device can be calculated separately.
[0125] In step 402, during the design phase, for each voice device, the average value of the gain of each frequency point in the multiple swept-frequency sound signals collected by it is calculated, and the average value is determined as the frequency gain; during the factory phase, for each voice device, the average value of the gain of each frequency point in the single swept-frequency sound signal collected by it is calculated, and the average value is determined as the frequency gain.
[0126] In step 403, no matter in the design stage or the factory stage, for each voice device, the frequency gain of the respective stage is used as its final gain, and the target gain difference corresponding to the target voice device is calculated.
[0127] Different voice devices may have certain hardware errors during hardware design, resulting in different gains. In the design stage, this embodiment calculates the target gain difference between the target voice device and the reference voice device based on the frequency gain of the swept-frequency sound signal of each voice device at multiple angles, so that the target gain difference can accurately measure the gain gap between each target voice device and the reference voice device in the presence of hardware errors, which is helpful for subsequent calibration of the voice energy of each voice device based on the target gain difference, so that the voice energy gain of each voice device remains at the same level; in the factory stage, the target gain difference between the target voice device and the reference voice device is calculated based on the frequency gain of the swept-frequency sound signal of the positive direction angle of each voice device, so that the target gain difference can further accurately measure the gain gap between each target voice device and the reference voice device in the presence of hardware errors, which is helpful for subsequent further calibration of the voice energy of each voice device based on the target gain difference, so that the voice energy gain of each voice device is more stably maintained at the same level.
[0128] In one embodiment, the target gain differences corresponding to each target voice device at multiple life cycle stages are respectively fused to obtain the fused gain differences corresponding to each target voice device, which may include:
[0129] The target gain differences corresponding to each target voice device at multiple life cycle stages are weighted and summed to obtain the fusion gain difference corresponding to each target voice device;
[0130] The determining factors of the weight of the target gain difference include at least one of the stage selection, the networking time of each voice device, or the wake-up times of each voice device.
[0131] The fusion gain difference corresponding to the target voice device can be obtained by the following formula: :
[0132] ;
[0133] in, is the target gain difference of the target speech device during the design phase, is the target gain difference of the target voice device at the factory stage, is the target gain difference of the target voice device during the late night period of the user's use phase, is the target gain difference of the target voice device during the historical wake-up period of the user's use phase, , , and is the corresponding weight.
[0134] Specifically, the calculation of the target gain difference at the factory stage is optional. ; When the networking time of each voice device is less than 100 days, the wake-up data is likely to be insufficient, so ; When the number of wake-up times is greater than 2000, it may cause interference during late night hours, so .
[0135] The target gain differences of the above stages are applied to the fusion calculation at each wake-up, and an overall fusion gain difference is formed according to different weights. The target gain differences of the design stage and the factory stage will not be updated after each voice device leaves the factory. The target gain differences of the late night period will only be updated at midnight. The target gain differences of the historical wake-up period will be updated at each wake-up.
[0136] In the process of fusing the target gain differences of each stage, this embodiment limits the weights of each target gain difference according to the stage selection, the networking time of each voice device and the wake-up times of each voice device, so that the fused gain difference can be more in line with the actual usage of each voice device. When the fused gain difference is subsequently used to calibrate the gain of the voice energy, the influence of the actual usage can be comprehensively considered to obtain more accurate voice energy.
[0137] Figure 5 This is a flowchart of the fifth embodiment of the method for waking up a device for identifying a device provided by an embodiment of the present invention. Figure 5 In one embodiment, the complete steps of waking up device identification are described as follows:
[0138] 1. Receive input signals sent by each voice device;
[0139] 2. Determine whether to enter the awakening state;
[0140] 3. If yes, monitor the starting point and ending point of the voice in each input signal to obtain the user wake-up voice, calculate the gain of the target frequency band voice in the user wake-up voice in the current wake-up period, and save it; at the same time, for each target voice device, merge the previously calculated and saved target gain difference in the design phase, the target gain difference in the factory phase, the target gain difference in the late night period of the user use phase, and the target gain difference in the historical wake-up period of the user use phase to obtain the fused gain difference corresponding to each target voice device;
[0141] 4. Use a bandpass filter to intercept the frequency band of interest in each user's wake-up voice;
[0142] 5. Calculate the energy of the speech in each frequency band of interest to obtain the speech energy after gain;
[0143] 6. Use the fusion gain difference to calibrate the speech energy of each target speech device;
[0144] 7. Identify the voice device corresponding to the maximum voice energy value of each target voice device after voice energy calibration as the wake-up device.
[0145] Furthermore, the speech energy of the reference speech device and the calibrated speech energy of each target speech device may be compared, and the speech device corresponding to the maximum speech energy may be identified as the wake-up device.
[0146] In actual application scenarios, this embodiment intercepts the speech of the frequency band of interest from the user wake-up speech collected by each voice device, and uses the fusion gain difference of multiple life cycle stages to calibrate the speech energy of the speech of the frequency band of interest of each voice device. This can effectively eliminate the difference in speech energy gain between each target voice device and the reference voice device under various interference factors, so that the speech energy gain of each target voice device is maintained at the same level as the speech energy gain of the reference voice device, and then the speech energy of each target voice device after gain can accurately measure the distance between the corresponding voice device and the user, thereby improving the accuracy of wake-up device recognition and enhancing user experience.
[0147] The following is a description of a nearby wake-up device identification device provided in an embodiment of the present invention. The nearby wake-up device identification device described below and the nearby wake-up device identification method described above can be referenced to each other.
[0148] Figure 6 Schematic diagram of the structure of the wake-up device identification device provided by an embodiment of the present invention. Figure 6 , an embodiment of the present invention provides a wake-up device identification device, which may include:
[0149] The target gain difference calculation module 601 is used to calculate the target gain difference between each target voice device and a reference voice device in the target network by adopting a stage adaptation method based on the life cycle stage of the voice device; the reference voice device is any voice device in the target network, and the target voice device is any voice device in the target network except the reference voice device;
[0150] The target gain difference fusion module 602 is used to: fuse the target gain differences corresponding to each target voice device at multiple life cycle stages respectively to obtain the fused gain differences corresponding to each target voice device;
[0151] The gain calibration module 603 is used to: perform gain calibration on the speech energy of the user wake-up speech received by each target speech device based on the fusion gain difference corresponding to each target speech device, so as to obtain the calibrated speech energy of each target speech device;
[0152] The wake-up device identification module 604 is used to: determine the maximum target speech energy from each target speech energy, and identify the speech device corresponding to the maximum target speech energy as the wake-up device;
[0153] The target speech energy includes the calibrated speech energy of each target speech device.
[0154] The wake-up device recognition device provided in this embodiment calculates the target gain difference between each target voice device and the reference voice device in the same network group based on the different life cycle stages of the voice device, using different calculation methods adapted to each stage. By covering the life cycle stages, various factors that cause the target gain difference between each target voice device and the reference voice device can be covered as much as possible, and accurate target gain differences under the interference of various factors can be obtained through targeted calculation methods. The target gain differences of the target voice devices in different stages are then fused to obtain the fused gain differences under the interference of various factors. Based on the fused gain difference, the voice energy of the user wake-up voice received by the target voice device is gain calibrated, and the gap between the voice energy gain of each target voice device and the voice energy gain of the reference voice device under the interference of various factors is eliminated to the maximum extent, so that the voice energy gain of each target voice device is kept at the same level as the voice energy gain of the reference voice device, and then the voice energy after gain of each target voice device can accurately measure the distance between the corresponding voice device and the user, thereby improving the accuracy of wake-up device recognition and simultaneously improving user experience.
[0155] In one embodiment, the target speech energy further includes speech energy of the user wake-up speech received by the reference speech device.
[0156] In one embodiment, the target gain difference calculation module 601 is specifically used to:
[0157] When each voice device is in a gain stable state, a target gain difference corresponding to the target voice device is calculated.
[0158] In one embodiment, the target gain difference calculation module 601 is specifically used to:
[0159] Calculating the gains of multiple first preset duration background noises collected by the voice device during the late night period to obtain multiple gains;
[0160] Obtaining a gain variation coefficient of the speech device based on a gain mean of the multiple gains and a gain standard deviation of the multiple gains;
[0161] If the gain variation coefficient of the voice device is less than the coefficient threshold, it is determined that the voice device is in a gain stable state.
[0162] In one embodiment, the target gain difference calculation module 601 is specifically used to:
[0163] Calculating the difference between the gain mean of the target voice device and the gain mean of the reference voice device to obtain a current gain difference corresponding to the target voice device;
[0164] Based on a weighted sum of the current gain difference and the historical gain differences corresponding to the target voice device, a target gain difference corresponding to the target voice device is obtained.
[0165] In one embodiment, the target gain difference calculation module 601 is specifically used to:
[0166] Calculating power spectra of multiple second preset duration background noises in the first preset duration background noise to obtain multiple power spectra;
[0167] Based on an average value of the multiple power spectra, a gain of the first preset duration background noise is determined.
[0168] In one embodiment, the target gain difference calculation module 601 is specifically used to:
[0169] Eliminating the maximum value and the minimum value among the multiple gains to obtain multiple gains to be processed;
[0170] The average value and the standard deviation of the plurality of gains to be processed are calculated to obtain the gain mean value and the gain standard deviation.
[0171] In one embodiment, the target gain difference calculation module 601 is specifically used to:
[0172] The difference between the historical gain of the target voice device and the historical gain of the reference voice device is calculated to obtain a target gain difference corresponding to the target voice device.
[0173] In one embodiment, the target gain difference calculation module 601 is specifically used to:
[0174] Calculating the gain of the target frequency band speech in the wake-up speech collected by the speech device in multiple historical wake-up time periods to obtain multiple gains to be processed;
[0175] The historical gain is determined based on an average value of the plurality of gains to be processed.
[0176] In one embodiment, the target gain difference calculation module 601 is specifically used to:
[0177] Calculating the difference between the frequency gain of the target voice device and the frequency gain of the reference voice device to obtain a target gain difference corresponding to the target voice device;
[0178] The target stage includes a design stage and / or a factory delivery stage.
[0179] In one embodiment, the target gain difference calculation module 601 is specifically used to:
[0180] Calculating the gain of each frequency point in the swept-frequency sound signal collected by the voice device in the target stage to obtain a plurality of gains to be processed; the swept-frequency sound signal is played by the speaker at at least one angular position relative to the voice device and after reaching equilibrium;
[0181] The frequency point gain is determined based on an average value of the multiple gains to be processed.
[0182] In one embodiment, the target gain difference fusion module 602 is specifically configured to:
[0183] The target gain differences corresponding to each target voice device at multiple life cycle stages are weighted and summed to obtain the fusion gain difference corresponding to each target voice device;
[0184] The determining factors of the weight of the target gain difference include at least one of the stage selection, the networking time of each voice device, or the wake-up times of each voice device.
[0185] Figure 7 is a schematic diagram of the physical structure of an electronic device provided by an embodiment of the present invention, such as Figure 7 As shown, the electronic device may include: a processor 710, a communication interface 720, a memory 730 and a communication bus 740, wherein the processor 710, the communication interface 720 and the memory 730 communicate with each other through the communication bus 740. The processor 710 may call the logic instructions in the memory 730 to execute the following method:
[0186] Based on the life cycle stage of the voice device, a target gain difference between each target voice device and a reference voice device in the target network is calculated by a stage adaptation method; the reference voice device is any voice device in the target network, and the target voice device is any voice device in the target network except the reference voice device;
[0187] The target gain differences corresponding to each target voice device at multiple life cycle stages are respectively fused to obtain the fused gain differences corresponding to each target voice device;
[0188] Based on the fusion gain difference corresponding to each target voice device, the voice energy of the user wake-up voice received by each target voice device is gain calibrated to obtain the calibrated voice energy of each target voice device;
[0189] Determine a maximum target speech energy from each target speech energy, and identify a speech device corresponding to the maximum target speech energy as a wake-up device;
[0190] The target speech energy includes the calibrated speech energy of each target speech device.
[0191] In addition, the logic instructions in the above-mentioned memory 730 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when it is sold or used as an independent product. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the relevant technology or the part of the technical solution, can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk and other media that can store program codes.
[0192] On the other hand, an embodiment of the present invention discloses a computer program product, wherein the computer program product includes a computer program stored on a non-transitory computer-readable storage medium, wherein the computer program includes program instructions. When the program instructions are executed by a computer, the computer can perform the methods provided by the above-mentioned method embodiments, for example, including:
[0193] Based on the life cycle stage of the voice device, a target gain difference between each target voice device and a reference voice device in the target network is calculated by a stage adaptation method; the reference voice device is any voice device in the target network, and the target voice device is any voice device in the target network except the reference voice device;
[0194] The target gain differences corresponding to each target voice device at multiple life cycle stages are respectively fused to obtain the fused gain differences corresponding to each target voice device;
[0195] Based on the fusion gain difference corresponding to each target voice device, the voice energy of the user wake-up voice received by each target voice device is gain calibrated to obtain the calibrated voice energy of each target voice device;
[0196] Determine a maximum target speech energy from each target speech energy, and identify a speech device corresponding to the maximum target speech energy as a wake-up device;
[0197] The target speech energy includes the calibrated speech energy of each target speech device.
[0198] In another aspect, an embodiment of the present invention further provides a non-transitory computer-readable storage medium having a computer program stored thereon, which is implemented when the computer program is executed by a processor to perform the methods provided in the above embodiments, for example, including:
[0199] Based on the life cycle stage of the voice device, a target gain difference between each target voice device and a reference voice device in the target network is calculated by a stage adaptation method; the reference voice device is any voice device in the target network, and the target voice device is any voice device in the target network except the reference voice device;
[0200] The target gain differences corresponding to each target voice device at multiple life cycle stages are respectively fused to obtain the fused gain differences corresponding to each target voice device;
[0201] Based on the fusion gain difference corresponding to each target voice device, the voice energy of the user wake-up voice received by each target voice device is gain calibrated to obtain the calibrated voice energy of each target voice device;
[0202] Determine a maximum target speech energy from each target speech energy, and identify a speech device corresponding to the maximum target speech energy as a wake-up device;
[0203] The target speech energy includes the calibrated speech energy of each target speech device.
[0204] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Those of ordinary skill in the art may understand and implement it without creative work.
[0205] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the relevant technology can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiment.
[0206] Finally, it should be noted that the above embodiments are only used to illustrate the present invention, rather than to limit the present invention. Although the present invention is described in detail with reference to the embodiments, it should be understood by those skilled in the art that various combinations, modifications or equivalent substitutions of the technical solutions of the present invention do not depart from the spirit and scope of the technical solutions of the present invention, and should be included in the scope of the claims of the present invention.
Claims
1. A method for waking up a device, characterized in that: include: Based on the life cycle stage of the voice device, a stage adaptation method is used to calculate the target gain difference between each target voice device and the reference voice device in the target network; The reference voice device is any voice device in the target network, and the target voice device is any voice device in the target network except the reference voice device; The target gain differences corresponding to each target voice device at multiple life cycle stages are respectively fused to obtain the fused gain differences corresponding to each target voice device; Based on the fusion gain difference corresponding to each target voice device, the voice energy of the user wake-up voice received by each target voice device is gain calibrated to obtain the calibrated voice energy of each target voice device; Determine a maximum target speech energy from each target speech energy, and identify a speech device corresponding to the maximum target speech energy as a wake-up device; The target speech energy includes the calibrated speech energy of each target speech device; The life cycle stages include a user use stage and a target stage. The user use stage includes a late night period and a historical wake-up period. The target stage includes a design stage and a factory delivery stage.
2. The wake-up device identification method according to claim 1, characterized in that: The target voice energy also includes the voice energy of the user wake-up voice received by the reference voice device.
3. The wake-up device identification method according to claim 1, characterized in that: The target gain difference corresponding to any target speech device during the user use phase is determined based on the following method: When each voice device is in a gain stable state, a target gain difference corresponding to the target voice device is calculated.
4. The wake-up device identification method according to claim 3, characterized in that: Whether any voice device is in a gain stable state is determined based on the following: Calculating the gains of multiple first preset duration background noises collected by the voice device during the late night period to obtain multiple gains; Obtaining a gain variation coefficient of the speech device based on a gain mean of the multiple gains and a gain standard deviation of the multiple gains; If the gain variation coefficient of the voice device is less than the coefficient threshold, it is determined that the voice device is in a gain stable state.
5. The wake-up device identification method according to claim 4, characterized in that: The calculating the target gain difference corresponding to the target voice device includes: Calculating the difference between the gain mean of the target voice device and the gain mean of the reference voice device to obtain a current gain difference corresponding to the target voice device; Based on a weighted sum of the current gain difference and the historical gain differences corresponding to the target voice device, a target gain difference corresponding to the target voice device is obtained.
6. The wake-up device identification method according to claim 4, characterized in that: The gain of any first preset duration background noise is determined based on the following method: Calculating power spectra of multiple second preset duration background noises in the first preset duration background noise to obtain multiple power spectra; Based on an average value of the multiple power spectra, a gain of the first preset duration background noise is determined.
7. The wake-up device identification method according to claim 4, characterized in that: The gain mean and the gain standard deviation are determined based on the following method: Eliminating the maximum value and the minimum value among the multiple gains to obtain multiple gains to be processed; The average value and the standard deviation of the plurality of gains to be processed are calculated to obtain the gain mean value and the gain standard deviation.
8. The wake-up device identification method according to claim 1, characterized in that: The target gain difference corresponding to any target speech device during the user use phase is determined based on the following method: The difference between the historical gain of the target voice device and the historical gain of the reference voice device is calculated to obtain a target gain difference corresponding to the target voice device.
9. The wake-up device identification method according to claim 8, characterized in that: The historical gain of any voice device is determined based on: Calculating the gain of the target frequency band speech in the wake-up speech collected by the speech device in multiple historical wake-up time periods to obtain multiple gains to be processed; The historical gain is determined based on an average value of the plurality of gains to be processed.
10. The wake-up device identification method according to claim 1, characterized in that: The target gain difference corresponding to any target speech device in the target stage is determined based on the following method: Calculating the difference between the frequency gain of the target voice device and the frequency gain of the reference voice device to obtain a target gain difference corresponding to the target voice device; The target stage includes a design stage and / or a factory delivery stage.
11. The wake-up device identification method according to claim 10, characterized in that: The frequency gain of any voice device is determined based on the following method: Calculating the gain of each frequency point in the swept-frequency sound signal collected by the voice device in the target stage to obtain a plurality of gains to be processed; the swept-frequency sound signal is played by the speaker at at least one angular position relative to the voice device and after reaching equilibrium; The frequency point gain is determined based on an average value of the multiple gains to be processed.
12. The wake-up device identification method according to claim 1, characterized in that: The target gain differences corresponding to each target voice device at multiple life cycle stages are respectively fused to obtain the fused gain differences corresponding to each target voice device, including: The target gain differences corresponding to each target voice device at multiple life cycle stages are weighted and summed to obtain the fusion gain difference corresponding to each target voice device; The determining factors of the weight of the target gain difference include at least one of the stage selection, the networking time of each voice device, or the wake-up times of each voice device.
13. A wake-up device identification device, characterized in that: include: A target gain difference calculation module is used to: calculate the target gain difference between each target voice device and a reference voice device in a target network by adopting a stage adaptation method based on the life cycle stage of the voice device; the reference voice device is any voice device in the target network, and the target voice device is any voice device in the target network except the reference voice device; The target gain difference fusion module is used to: fuse the target gain differences corresponding to each target voice device at multiple life cycle stages respectively to obtain the fused gain differences corresponding to each target voice device; A gain calibration module is used to: perform gain calibration on the speech energy of the user wake-up speech received by each target speech device based on the fusion gain difference corresponding to each target speech device, so as to obtain the calibrated speech energy of each target speech device; A wake-up device identification module is used to: determine the maximum target voice energy from each target voice energy, and identify the voice device corresponding to the maximum target voice energy as the wake-up device; The target speech energy includes the calibrated speech energy of each target speech device; The life cycle stages include a user use stage and a target stage. The user use stage includes a late night period and a historical wake-up period. The target stage includes a design stage and a factory delivery stage.
14. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the wake-up device identification method according to any one of claims 1 to 12 is implemented.
15. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the wake-up device identification method according to any one of claims 1 to 12 is implemented.
16. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the wake-up device identification method according to any one of claims 1 to 12 is implemented.
Citation Information
Patent Citations
Device waking up method and system for acoustic networking
CN110288997A
Audio synchronization method and device of distributed microphone and storage medium
CN115631764A