Audio calibration method, receiving device, playback device, storage medium

CN115734139BActive Publication Date: 2026-05-26QINGDAO HAIER TECH +2
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
QINGDAO HAIER TECH
Filing Date
2022-10-31
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

当唤醒时刻其中某台设备正在播放声音时,播放设备可以通过回声消除将自身播放的声音的影响消除掉,但对其他设备来说都是噪音干扰,会影响其他设备对唤醒指令的能量计算,进而在决策哪个设备该响应用户时出现错误

Benefits of technology

[0017]This invention, when there are N playback devices in a target area, repeatedly executes the following steps until the receiving device determines the target information corresponding to each playback device: Upon receiving a calibration command from the i-th playback device in the target area, a calibration mode is activated, and the i-th audio played by the i-th playback device is acquired; the target information corresponding to the i-th playback device is calculated based on the i-th audio; wherein, the target information is used to indicate parameters for eliminating the influence of the audio played by the i-th playback device on the receiving device. In other words, the receiving device can automatically determine the target information corresponding to each playback device based on the audio played by each playback device, without the need for a server or central device. This solves the problem that the information required for the device to eliminate interference from other devices' audio needs to be determined through a cloud server or central device, thereby reducing hardware costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115734139B_ABST
    Figure CN115734139B_ABST
Patent Text Reader

Abstract

This application discloses an audio calibration method, a receiving device, a playback device, and a storage medium, relating to the field of smart home technology. The audio calibration method includes: when there are N playback devices in a target area, repeatedly performing the following steps until the receiving device determines the target information corresponding to each playback device: upon receiving a calibration command sent by the i-th playback device in the target area, activating the calibration mode and acquiring the i-th audio played by the i-th playback device; calculating the target information corresponding to the i-th playback device based on the i-th audio; wherein i is greater than or equal to 1 and less than or equal to N, and the target information is used to indicate parameters for eliminating the influence of the audio played by the i-th playback device on the receiving device. By adopting the above technical solution, the problem that the information required for the device to eliminate interference from the audio of other devices needs to be determined through a cloud server or central device is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of smart home technology, and more specifically, to an audio calibration method, a receiving device, a playback device, and a storage medium. Background Technology

[0002] When multiple smart voice devices in a user's home share the same wake-up command, existing wake-up methods typically calculate the audio energy received by each device based on the wake-up command. The magnitude of the audio energy represents the distance between the user and the device, and this distance is used to determine which device should uniquely respond to the user's wake-up command. However, this approach only works in quiet scenarios, where only the user is speaking the wake-up command and there is no other sound interference. When one device is playing sound at the wake-up time, the playback device can eliminate the influence of its own sound through echo cancellation, but this is noise interference for other devices. This affects their energy calculation of the wake-up command, leading to errors in deciding which device should respond to the user. For example, if device A is playing music, and device B is closer to device A than device C, when the user speaks the wake-up command, even though the user is closer to device C, the influence of the sound played by device A will cause device B to receive more audio energy than device C, resulting in device B being woken up. This does not match the user's expectations and negatively impacts the user experience. Therefore, echo energy calibration is required to eliminate the interference of the playback device's echo on other devices. However, existing echo energy calibration solutions require cloud servers or central devices to determine calibration information and distribute it to each device so that the devices can eliminate the echo based on the calibration information. However, cloud servers or central devices need to support the calibration work of a large number of devices, which greatly increases hardware and R&D costs and affects calibration efficiency.

[0003] There is currently no effective solution to the problem that the information needed for a device to eliminate audio interference from other devices needs to be determined through a cloud server or central device in related technologies.

[0004] Therefore, it is necessary to improve the relevant technology to overcome the aforementioned defects. Summary of the Invention

[0005] This invention provides an audio calibration method, a receiving device, a playback device, and a storage medium to at least address the problem that the information required for a device to eliminate audio interference from other devices needs to be determined via a cloud server or a central device.

[0006] According to one aspect of the present invention, an audio calibration method is provided, comprising: when there are N playback devices in a target area, repeatedly performing the following steps until a receiving device determines target information corresponding to each playback device: upon receiving a calibration instruction sent by the i-th playback device in the target area, activating a calibration mode and acquiring the i-th audio played by the i-th playback device; calculating target information corresponding to the i-th playback device based on the i-th audio; wherein i is greater than or equal to 1 and less than or equal to N, and the target information is used to indicate parameters for eliminating the influence of the audio played by the i-th playback device on the receiving device.

[0007] In an exemplary embodiment, when the receiving device does not have a target number of microphones, calculating target information corresponding to the i-th playback device based on the i-th audio includes: acquiring a first audio energy of the i-th audio transmitted by the i-th playback device; determining an interference coefficient corresponding to the i-th playback device based on the first audio energy and a second audio energy, wherein the second audio energy is the audio energy of the i-th audio determined during the acquisition of the i-th audio, and the target information includes the interference coefficient.

[0008] In an exemplary embodiment, when the receiving device does not have a target number of microphones, calculating target information corresponding to the i-th playback device based on the i-th audio includes: acquiring a first set of audio energies of the i-th audio sent by the i-th playback device, wherein each audio energy in the first set of audio energies corresponds one-to-one with each volume in a set of volumes; determining a set of interference coefficients corresponding to the i-th playback device based on the first set of audio energies and a second set of audio energies, wherein each audio energy in the second set of audio energies is the audio energy of the i-th audio determined during the acquisition of the i-th audio at each volume in the set of volumes; the target information includes the set of interference coefficients, and each interference coefficient in the set of interference coefficients corresponds one-to-one with each volume in the set of volumes.

[0009] In an exemplary embodiment, when the receiving device has a target number of microphones, calculating target information corresponding to the i-th playback device based on the i-th audio includes: calculating the positional relationship between the receiving device and the i-th playback device based on the i-th audio, wherein the target information includes the positional relationship.

[0010] In an exemplary embodiment, when the receiving device does not have a target number of microphones and target information corresponding to the i-th playback device is determined, the method further includes: during the process of the receiving device acquiring wake-up audio, when the i-th playback device is playing target audio, determining the audio energy of the acquired audio and acquiring the audio energy of the target audio sent by the i-th playback device; subtracting the product of the audio energy of the target audio and the target interference coefficient from the audio energy of the acquired audio to obtain the audio energy of the wake-up audio, wherein the target information includes the target interference coefficient.

[0011] In an exemplary embodiment, when the receiving device has a target number of microphones and target information corresponding to the i-th playback device is determined, the method further includes: when the receiving device is acquiring wake-up audio and the i-th playback device is playing target audio, performing noise reduction processing on the target audio in the target direction, wherein the target information includes a positional relationship indicating that the i-th playback device is located in the target direction of the receiving device.

[0012] According to another aspect of the present invention, an audio calibration method is also provided, comprising: when there are N receiving devices in a target area, repeatedly performing the following steps until each receiving device determines target information corresponding to a playback device: sending a calibration command to the i-th receiving device in the target area, instructing the i-th receiving device to enable a calibration mode; playing the i-th audio, wherein, when the i-th receiving device enables the calibration mode, it acquires the i-th audio played by the playback device and calculates the target information corresponding to the playback device based on the i-th audio; and, upon receiving feedback information from the i-th receiving device, determining that the i-th receiving device has determined the target information corresponding to the playback device; wherein i is greater than or equal to 1 and less than or equal to N, and the target information is used to indicate parameters for eliminating the influence of the audio played by the playback device on the receiving device.

[0013] According to another aspect of the present invention, a receiving device is also provided, comprising: a first calibration module, configured to repeatedly execute the following steps when there are N playback devices in a target area, until the receiving device determines target information corresponding to each playback device: upon receiving a calibration instruction sent by the i-th playback device in the target area, activating a calibration mode and acquiring the i-th audio played by the i-th playback device; calculating target information corresponding to the i-th playback device based on the i-th audio; wherein i is greater than or equal to 1 and less than or equal to N, and the target information is used to indicate parameters for eliminating the influence of the audio played by the i-th playback device on the receiving device.

[0014] According to another aspect of the present invention, a playback device is also provided, comprising: a second calibration module, configured to repeatedly execute the following steps when there are N receiving devices in a target area, until each receiving device determines target information corresponding to the playback device: sending a calibration instruction to the i-th receiving device in the target area, instructing the i-th receiving device to enable a calibration mode; playing the i-th audio, wherein, when the i-th receiving device enables the calibration mode, it acquires the i-th audio played by the playback device and calculates the target information corresponding to the playback device based on the i-th audio; and, upon receiving feedback information from the i-th receiving device, determining that the i-th receiving device has determined the target information corresponding to the playback device; wherein i is greater than or equal to 1 and less than or equal to N, and the target information is used to indicate parameters for eliminating the influence of the audio played by the playback device on the receiving device.

[0015] According to another aspect of the present invention, a computer-readable storage medium is also provided, wherein a computer program is stored therein, wherein the computer program is configured to execute the above-described audio calibration method at runtime.

[0016] According to another aspect of the present invention, an electronic device is also provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the audio calibration method described above through the computer program.

[0017] This invention, when there are N playback devices in a target area, repeatedly executes the following steps until the receiving device determines the target information corresponding to each playback device: Upon receiving a calibration command from the i-th playback device in the target area, a calibration mode is activated, and the i-th audio played by the i-th playback device is acquired; the target information corresponding to the i-th playback device is calculated based on the i-th audio; wherein, the target information is used to indicate parameters for eliminating the influence of the audio played by the i-th playback device on the receiving device. In other words, the receiving device can automatically determine the target information corresponding to each playback device based on the audio played by each playback device, without the need for a server or central device. This solves the problem that the information required for the device to eliminate interference from other devices' audio needs to be determined through a cloud server or central device, thereby reducing hardware costs. Attached Figure Description

[0018] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0019] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0020] Figure 1 This is a schematic diagram of the hardware environment for an audio calibration method according to an embodiment of this application;

[0021] Figure 2 This is a flowchart (I) of an audio calibration method according to an embodiment of the present invention;

[0022] Figure 3 This is a flowchart (II) of an audio calibration method according to an embodiment of the present invention;

[0023] Figure 4 This is a flowchart (III) of an audio calibration method according to an embodiment of the present invention;

[0024] Figure 5 This is a schematic diagram of a scenario for an audio calibration method according to an embodiment of the present invention;

[0025] Figure 6 This is a structural block diagram of a receiving device according to an embodiment of the present invention;

[0026] Figure 7 This is a structural block diagram of a playback device according to an embodiment of the present invention. Detailed Implementation

[0027] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.

[0028] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0029] According to one aspect of the embodiments of this application, an audio calibration method is provided. This audio calibration method is widely used in whole-house intelligent digital control application scenarios such as smart homes, smart home ecosystems, and intelligencehouse ecosystems. Optionally, in this embodiment, the above-mentioned audio calibration method can be applied to, for example... Figure 1 The hardware environment shown consists of terminal device 102 and server 104. For example... Figure 1 As shown, server 104 is connected to terminal device 102 via a network and can be used to provide services (such as application services) to the terminal or clients installed on the terminal. A database can be set up on the server or independently of the server to provide data storage services for server 104. Cloud computing and / or edge computing services can be configured on the server or independently of the server to provide data processing services for server 104.

[0030] The aforementioned network may include, but is not limited to, at least one of the following: wired network, wireless network. The aforementioned wired network may include, but is not limited to, at least one of the following: wide area network, metropolitan area network, local area network. The aforementioned wireless network may include, but is not limited to, at least one of the following: Wi-Fi (Wireless Fidelity), Bluetooth. The terminal device 102 may not be limited to PC, mobile phone, tablet computer, smart air conditioner, smart range hood, smart refrigerator, smart oven, smart stove, smart washing machine, smart water heater, smart washing equipment, smart dishwasher, smart projector, smart TV, smart clothes rack, smart curtains, smart audio-visual equipment, smart socket, smart speaker, smart speaker box, smart fresh air equipment, smart kitchen and bathroom equipment, smart bathroom equipment, smart robot vacuum cleaner, smart window cleaning robot, smart mopping robot, smart air purifier, smart steam oven, smart microwave oven, smart water heater, smart air purifier, smart water dispenser, smart door lock, etc.

[0031] To address the aforementioned issues, this embodiment provides an audio calibration method. Figure 2 This is a flowchart (a) of an audio calibration method according to an embodiment of the present invention, including but not limited to applications in receiving devices. The process includes the following steps:

[0032] Step S202: If there are N playback devices in the target area, repeat steps S204-S206 until the receiving device determines the target information corresponding to each playback device.

[0033] Step S204: Upon receiving a calibration command from the i-th playback device in the target area, activate the calibration mode and acquire the i-th audio played by the i-th playback device;

[0034] Step S206: Calculate the target information corresponding to the i-th playback device based on the i-th audio;

[0035] Wherein, i is greater than or equal to 1 and less than or equal to N, and the target information is used to indicate parameters for eliminating the influence of the audio played by the i-th playback device on the receiving device.

[0036] It should be noted that the receiving device can use the target information to eliminate the influence of the audio played by the i-th playback device on itself.

[0037] For example, the target area can be the user's house, where both the receiving and playback devices have voice functionality and use the same wake word. Here, the playback device is the one that plays audio during audio calibration, and the receiving device is the one that receives audio during audio calibration.

[0038] It should be noted that if the receiving device does not have the target number of microphones, the corresponding target information is the interference coefficient. The interference coefficient at least indicates the degree of interference caused by the audio played by the playback device during audio acquisition. If the receiving device has the target number of microphones, the corresponding target information is the positional relationship between the playback device and the receiving device. The target number is greater than or equal to 4. The receiving device can use the interference coefficient or positional relationship to eliminate the influence of the audio played by the playback device on itself.

[0039] By repeating the above steps when there are N playback devices in the target area, the following steps are executed until the receiving device determines the target information corresponding to each playback device: Upon receiving a calibration command from the i-th playback device in the target area, calibration mode is activated, and the i-th audio played by the i-th playback device is acquired; the target information corresponding to the i-th playback device is calculated based on the i-th audio; wherein, the target information is used to indicate the parameters for eliminating the influence of the audio played by the i-th playback device on the receiving device. In other words, the receiving device can automatically determine the target information corresponding to each playback device based on the audio played by each playback device, without the need for a server or central device. This solves the problem that the information required for the device to eliminate interference from the audio of other devices needs to be determined by a cloud server or central device, thereby reducing hardware costs.

[0040] In an exemplary embodiment, when the receiving device does not have the target number of microphones, step S206 above can be implemented as follows: acquiring the first audio energy of the i-th audio transmitted by the i-th playback device; determining the interference coefficient corresponding to the i-th playback device based on the first audio energy and the second audio energy, wherein the second audio energy is the audio energy of the i-th audio determined during the acquisition of the i-th audio, and the target information includes the interference coefficient.

[0041] It should be noted that, because the receiving device does not have the target number of microphones, it cannot determine the relative position of the playback device playing the i-th audio frequency based on the acquired i-th audio frequency. In other words, the receiving device can only determine the interference coefficient corresponding to the playback device playing the i-th audio frequency based on the audio energy of the i-th audio frequency.

[0042] It should be noted that after playing the i-th audio, the i-th playback device sends the first audio energy actually generated by playing the i-th audio to the receiving device. During the acquisition of the i-th audio, the receiving device determines the second audio energy of the i-th audio based on the acquired audio. It should be noted that due to the distance between the i-th playback device and the receiving device, the audio energy of the i-th audio played by the i-th playback device will be lost as it propagates through the air. Therefore, the second audio energy of the i-th audio determined by the receiving device is less than the first audio energy received by the receiving device from the i-th playback device. Based on the first and second audio energies, the receiving device can then determine the interference coefficient corresponding to the i-th playback device.

[0043] For example, the receiving device can divide the second audio energy by the first audio energy to obtain the corresponding interference coefficient. For instance, after playing the i-th audio, the i-th playback device (e.g., device A) divides the actual first audio energy (E) of the i-th audio by... A-spk The audio energy of the i-th audio signal is sent to the receiving device (e.g., device B), and the audio acquisition device of the receiving device acquires the audio energy of the i-th audio signal as the second audio energy (E). (A,B)-mic Therefore, the receiving device can calculate the interference coefficient α of the i-th playback device relative to itself. A,B =E (A,B)-mic / E A-spk .

[0044] In an exemplary embodiment, if the receiving device does not have the target number of microphones, step S206 can also be implemented as follows: acquiring a first set of audio energy of the i-th audio transmitted by the i-th playback device, wherein each audio energy in the first set of audio energy corresponds one-to-one with each volume in a set of volumes; determining a set of interference coefficients corresponding to the i-th playback device based on the first set of audio energy and the second set of audio energy, wherein each audio energy in the second set of audio energy is the audio energy of the i-th audio determined during the acquisition of the i-th audio at each volume in the set of volumes; the target information includes the set of interference coefficients, wherein each interference coefficient in the set of interference coefficients corresponds one-to-one with each volume in the set of volumes.

[0045] It should be noted that, in this embodiment, since the degree of interference caused by the i-th audio played by the i-th playback device at different volumes to the receiving device during audio acquisition varies, the receiving device needs to determine the interference coefficient of the i-th audio played by the i-th playback device at different volumes. Therefore, the i-th playback device needs to play the i-th audio at different volumes and send the actual audio energy corresponding to the i-th audio played at different volumes to the receiving device. The receiving device can then determine the acquired audio energy of the i-th audio at different volumes and, based on the acquired audio energy of the i-th audio at different volumes and the audio energy of the i-th audio at different volumes sent by the i-th playback device, determine a corresponding set of interference coefficients.

[0046] In an exemplary embodiment, when the receiving device has a target number of microphones, step S206 above can be implemented by calculating the positional relationship between the receiving device and the i-th playback device based on the i-th audio, wherein the target information includes the positional relationship.

[0047] It should be noted that only devices with the target number of microphones can determine the location of the device playing the audio based on the acquired audio. Since the receiving device has the target number of microphones, it can determine the positional relationship between itself and the i-th playback device based on the acquired i-th audio signal.

[0048] It should be noted that when the receiving device is collecting the user's audio, if the i-th playback device is playing audio, noise reduction processing on the audio in the direction of the i-th playback device is more effective in eliminating the influence of the audio played by the i-th playback device than using an interference coefficient. Furthermore, if the receiving device has the target number of microphones, the positional relationship between the receiving device and the i-th playback device is calculated based on the i-th audio signal.

[0049] In an exemplary embodiment, when the receiving device does not have a target number of microphones and target information corresponding to the i-th playback device is determined, the method further includes: during the process of the receiving device acquiring wake-up audio, while the i-th playback device is playing target audio, determining the audio energy of the acquired audio and acquiring the audio energy of the target audio sent by the i-th playback device; subtracting the product of the audio energy of the target audio and the target interference coefficient from the audio energy of the acquired audio to obtain the audio energy of the wake-up audio, wherein the target information includes the target interference coefficient.

[0050] In other words, during the process of the receiving device collecting the user's wake-up audio, if the i-th playback device is playing the target audio at this time, the receiving device can eliminate the interference caused by the target audio played by the i-th playback device to its own collection of the user's wake-up audio based on the pre-determined target information.

[0051] For example, during the process of the receiving device acquiring the wake-up audio, if the i-th playback device is playing the target audio, the i-th playback device can send a message to the receiving device. The message content is: the i-th playback device is playing the target audio, and the audio energy generated by the played target audio is A. The audio energy acquired by the microphone of the receiving device is B. To accurately determine the audio energy of the wake-up audio, it is necessary to eliminate the influence of the target audio played by the i-th playback device. Since the receiving device has already determined the degree of influence 'a' of the audio from the i-th playback device, the receiving device can then determine the audio energy of the wake-up audio as Ba*A.

[0052] By adopting the above technical solution, the audio energy of the wake-up audio can be accurately determined even when the receiving device does not have the target number of microphones.

[0053] In an exemplary embodiment, when the receiving device has a target number of microphones and target information corresponding to the i-th playback device is determined, the method further includes: when the receiving device is acquiring wake-up audio and the i-th playback device is playing target audio, performing noise reduction processing on the target audio in the target direction, wherein the target information includes a positional relationship indicating that the i-th playback device is located in the target direction of the receiving device.

[0054] It should be noted that during the process of the receiving device acquiring the wake-up audio, if the i-th playback device is playing the target audio, the i-th playback device can send a message to the receiving device informing it that it is currently playing audio. Since the receiving device already knows the positional relationship between the i-th playback device and itself (the i-th playback device is located in the target direction of the receiving device), the receiving device can perform noise reduction processing on the target audio in the target direction when acquiring the wake-up audio, thereby accurately determining the audio energy of the wake-up audio. Using the above technical solution, the receiving device can accurately determine the audio energy of the wake-up audio.

[0055] It should be noted that this embodiment also provides an audio calibration method. Figure 3 This is a flowchart (II) of an audio calibration method according to an embodiment of the present invention, including but not limited to applications in playback devices. The process includes the following steps:

[0056] Step S302: If there are N receiving devices in the target area, repeat steps S304-S308 until each receiving device determines the target information corresponding to the playback device.

[0057] Step S304: Send a calibration command to the i-th receiving device in the target area, instructing the i-th receiving device to enable calibration mode;

[0058] Step S306: Play the i-th audio, wherein the i-th receiving device, when the calibration mode is enabled, acquires the i-th audio played by the playback device, and calculates the target information corresponding to the playback device based on the i-th audio;

[0059] It should be noted that after performing step two above, the playback device will also send the actual audio energy of the i-th audio to the i-th receiving device.

[0060] Step S308: If feedback information from the i-th receiving device is obtained, determine that the i-th receiving device has determined the target information corresponding to the playback device;

[0061] Wherein, i is greater than or equal to 1 and less than or equal to N, and the target information is used to indicate parameters for eliminating the influence of the audio played by the playback device on the receiving device.

[0062] Through the above steps, since the playback device allows the receiving device to automatically determine the target information corresponding to each playback device by playing audio, without the need for the involvement of a server or central device, the problem of needing to determine the information required for the device to eliminate audio interference from other devices through a cloud server or central device is solved, thereby reducing hardware costs.

[0063] Obviously, the embodiments described above are merely some embodiments of the present invention, and not all embodiments. To better understand the above method, the following description, in conjunction with embodiments, illustrates the process, but is not intended to limit the technical solutions of the embodiments of the present invention. Specifically:

[0064] In an optional embodiment, Figure 4 This is a flowchart (III) of an audio calibration method according to an embodiment of the present invention, and the specific steps are as follows:

[0065] Step 1: Power on all devices and connect them to the router;

[0066] Once all devices (equivalent to the N playback and receiving devices in the above embodiment) are powered on and connected to the user-configured network, each device will send its device information to other devices in the network. In this way, any device will save the information of all devices under the user's network, thereby establishing a local network containing all devices.

[0067] Step 2: The user selects a device that can play sound and enters echo calibration mode through the smart home APP or by long-pressing a button on the device.

[0068] Users can select a device in their home that can play sound for calibration. By using a smart home app or by long-pressing a predefined button on the device, the device can enter echo calibration mode. At this time, the device will have a corresponding voice broadcast and a predefined color LED indicator will light up to indicate to the user that it has entered echo calibration mode.

[0069] Step 3: The playback device sends the echo calibration command to all other receiving devices of the user;

[0070] The playback device sends the echo calibration command to all other receiving devices via the network and prepares to play the calibration audio. Figure 5 This is a schematic diagram of a scenario illustrating an audio calibration method according to an embodiment of the present invention, such as... Figure 5 As shown, the television is the playback device, and the refrigerator, speaker, smart screen, and security panel are the receiving devices.

[0071] Step 4: The receiving device enters echo calibration mode;

[0072] After receiving the instruction from the playback device to enter echo calibration mode, the receiving device enters echo calibration mode. At this time, an LED of a predefined color is lit, waiting for the playback device to play the calibration audio.

[0073] Step 5: The playback device plays the calibration audio and sends the calculated echo audio energy (equivalent to the second audio energy in the above embodiment) to all other receiving devices;

[0074] After confirming that all receiving devices have entered echo calibration mode, the playback device begins playing the pre-stored calibration audio at various volume levels, periodically calculating the energy of the echo audio, thus generating a set of echo audio energy (E0). A-spk-vol (Equivalent to the second set of audio energy in the above embodiment), and sends the energy of the echo audio and the corresponding volume to other receiving devices. Accordingly, the LED light will flash periodically in sync to indicate to the user that the calibration process is proceeding normally.

[0075] Step 6: If the receiving device is a device with 4 or more microphones (equivalent to the receiving device in the above embodiment having the target number of microphones), calculate and store the position and orientation of the playback device;

[0076] The receiving device performs sound source localization and beamforming processes to calculate and store the position and direction of the playback device. Then, when the device is actually woken up, it performs noise reduction based on the beam from the playback device's position and direction to obtain the user's actual wake-up command audio.

[0077] Step 7: The receiving device is a device with 4 or fewer microphones (equivalent to the receiving device in the above embodiment not having the target number of microphones), receives the echo audio energy emitted by the playback device in step 5, calculates the echo influence coefficient and stores it;

[0078] The receiving device periodically collects and calibrates audio signals, generating a set of audio energies (E) corresponding to the echo energy of the playback device. (A,B)-mic (Equivalent to the second set of audio energy in the above embodiment), after the calibration audio is played, a calibration calculation is performed with the echo energy of the buffered playback device to calculate the influence coefficient α of the echo at each volume (vol) of the playback device on this device. (A,B)-vol =E (A,B)-mic / E A-spk-vol The data is saved to the receiving device's memory. At this time, the receiving device will illuminate an LED of a corresponding predefined color to indicate to the user that the calibration process is proceeding normally. When the receiving device actually wakes up, it will use α... (A,B)-vol The impact of the echo on the receiving device is calculated, and the decision audio energy of the receiving device is obtained according to a preset intelligent algorithm, thereby improving the accuracy of the decision.

[0079] Step 8: Perform echo calibration on the next device capable of playing sound;

[0080] After calibrating one playback device, proceed to calibrate the next playable device, repeating steps 2 through 8 until all playable devices are calibrated.

[0081] In this embodiment, users can perform echo calibration on their installed smart voice devices at home without requiring specialized equipment, personnel, or expertise. It enables echo calibration of various volume levels on playback devices and automatically saves calibration parameters, resolving noise interference issues caused by the playback device on other devices. The calibration calculations and storage are distributed across the various devices, eliminating the need for a central device or cloud server, thus reducing hardware costs and minimizing the development costs associated with central devices and servers.

[0082] Furthermore, this embodiment employs different calibration strategies for different types of devices, which is closer to the actual usage environment. Therefore, the calibration parameters are more accurate, the device response accuracy is higher, hardware, R&D, and after-sales costs are reduced, and the user experience is improved. Simultaneously, when users move existing devices or add new devices, or when a device experiences a hardware problem, users can perform recalibration themselves without on-site technical support, and without the need to move or replace the device.

[0083] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods of the various embodiments of the present invention.

[0084] This embodiment also provides a receiving device and a playback device for implementing the above embodiments and preferred embodiments; details already described will not be repeated. As used below, the term "module" can refer to a combination of software and / or hardware that performs a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, hardware implementations, or a combination of software and hardware, are also possible and contemplated.

[0085] Figure 6 This is a structural block diagram of a receiving device according to an embodiment of the present invention, the device comprising:

[0086] The first calibration module 62 is used to repeatedly execute the following steps when there are N playback devices in the target area, until the receiving device determines the target information corresponding to each playback device:

[0087] Upon receiving a calibration command from the i-th playback device in the target area, the calibration mode is activated, and the i-th audio signal played by the i-th playback device is acquired.

[0088] Calculate the target information corresponding to the i-th playback device based on the i-th audio;

[0089] Wherein, i is greater than or equal to 1 and less than or equal to N, and the target information is used to indicate parameters for eliminating the influence of the audio played by the i-th playback device on the receiving device.

[0090] Using the above-described device, when there are N playback devices in the target area, the following steps are repeated until the receiving device determines the target information corresponding to each playback device: Upon receiving a calibration command from the i-th playback device in the target area, the calibration mode is activated, and the i-th audio played by the i-th playback device is acquired; the target information corresponding to the i-th playback device is calculated based on the i-th audio; wherein, the target information is used to indicate parameters for eliminating the influence of the audio played by the i-th playback device on the receiving device. In other words, the receiving device can automatically determine the target information corresponding to each playback device based on the audio played by each playback device, without the need for a server or central device. This solves the problem that the information required for the device to eliminate interference from the audio of other devices needs to be determined through a cloud server or central device, thereby reducing hardware costs.

[0091] In an exemplary embodiment, the first calibration module 62 is further configured to acquire the first audio energy of the i-th audio transmitted by the i-th playback device when the receiving device does not have a target number of microphones; determine the interference coefficient corresponding to the i-th playback device based on the first audio energy and the second audio energy, wherein the second audio energy is the audio energy of the i-th audio determined during the acquisition of the i-th audio, and the target information includes the interference coefficient.

[0092] In an exemplary embodiment, the first calibration module 62 is further configured to, when the receiving device does not have a target number of microphones, acquire a first set of audio energies of the i-th audio transmitted by the i-th playback device, wherein each audio energy in the first set of audio energies corresponds one-to-one with each volume in a set of volumes; determine a set of interference coefficients corresponding to the i-th playback device based on the first set of audio energies and the second set of audio energies, wherein each audio energy in the second set of audio energies is the audio energy of the i-th audio determined during the acquisition of the i-th audio at each volume in the set of volumes; the target information includes the set of interference coefficients, wherein each interference coefficient in the set of interference coefficients corresponds one-to-one with each volume in the set of volumes.

[0093] In an exemplary embodiment, the first calibration module 62 is further configured to calculate the positional relationship between the receiving device and the i-th playback device based on the i-th audio, provided that the receiving device has a target number of microphones, wherein the target information includes the positional relationship.

[0094] In an exemplary embodiment, the first calibration module 62 is further configured to, when the receiving device does not have a target number of microphones and target information corresponding to the i-th playback device is determined, determine the audio energy of the acquired audio and obtain the audio energy of the target audio sent by the i-th playback device during the process of the receiving device acquiring wake-up audio and when the i-th playback device is playing target audio; subtract the product of the audio energy of the target audio and the target interference coefficient from the audio energy of the acquired audio to obtain the audio energy of the wake-up audio, wherein the target information includes the target interference coefficient.

[0095] In an exemplary embodiment, the first calibration module 62 is further configured to, when the receiving device has a target number of microphones and target information corresponding to the i-th playback device is determined, further include: when the receiving device is acquiring wake-up audio and the i-th playback device is playing target audio, performing noise reduction processing on the target audio in the target direction, wherein the target information includes a positional relationship indicating that the i-th playback device is located in the target direction of the receiving device.

[0096] Figure 7 This is a structural block diagram of a playback device according to an embodiment of the present invention, the device comprising:

[0097] The second calibration module 72 is used to repeatedly execute the following steps when there are N receiving devices in the target area, until each receiving device determines the target information corresponding to the playback device:

[0098] A calibration command is sent to the i-th receiving device in the target area, instructing the i-th receiving device to enable calibration mode;

[0099] Play the i-th audio, wherein the i-th receiving device, when the calibration mode is enabled, acquires the i-th audio played by the playback device, and calculates the target information corresponding to the playback device based on the i-th audio;

[0100] If feedback information from the i-th receiving device is obtained, it is determined that the i-th receiving device has identified target information corresponding to the playback device;

[0101] Wherein, i is greater than or equal to 1 and less than or equal to N, and the target information is used to indicate parameters for eliminating the influence of the audio played by the playback device on the receiving device.

[0102] With the above-mentioned device, the playback device can automatically determine the target information corresponding to each playback device by playing audio, without the need for the participation of a server or central device. This solves the problem that the information required for the device to eliminate audio interference from other devices needs to be determined by a cloud server or central device, thereby reducing hardware costs.

[0103] Embodiments of the present invention also provide a computer-readable storage medium storing a computer program, wherein the computer program is configured to perform the steps in any of the above method embodiments when executed.

[0104] Optionally, in this embodiment, the storage medium may be configured to store a computer program for performing the following steps:

[0105] S1, if there are N playback devices in the target area, repeat steps S2-S3 until the receiving device determines the target information corresponding to each playback device:

[0106] S2, upon receiving a calibration command from the i-th playback device in the target area, activate the calibration mode and acquire the i-th audio played by the i-th playback device;

[0107] S3, calculate target information corresponding to the i-th playback device based on the i-th audio; wherein i is greater than or equal to 1 and less than or equal to N, and the target information is used to indicate parameters for eliminating the influence of the audio played by the i-th playback device on the receiving device.

[0108] Alternatively, the aforementioned storage medium may be configured to store a computer program for performing the following steps:

[0109] S1, If ​​there are N receiving devices in the target area, repeat steps S2-S4 until each receiving device determines the target information corresponding to the playback device:

[0110] S2, a calibration command is sent to the i-th receiving device in the target area, instructing the i-th receiving device to enable calibration mode;

[0111] S3, play the i-th audio, wherein, when the i-th receiving device is in calibration mode, it acquires the i-th audio played by the playback device and calculates the target information corresponding to the playback device based on the i-th audio;

[0112] S4, upon receiving feedback information from the i-th receiving device, it is determined that the i-th receiving device has identified target information corresponding to the playback device; wherein i is greater than or equal to 1 and less than or equal to N, and the target information is used to indicate parameters for eliminating the influence of the audio played by the playback device on the receiving device.

[0113] In one exemplary embodiment, the aforementioned computer-readable storage medium may include, but is not limited to, various media capable of storing computer programs, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard disk, magnetic disk, or optical disk.

[0114] Specific examples in this embodiment can be found in the examples described in the above embodiments and exemplary implementations, and will not be repeated here.

[0115] Embodiments of the present invention also provide an electronic device including a memory and a processor, the memory storing a computer program and the processor being configured to run the computer program to perform the steps in any of the above method embodiments.

[0116] Optionally, in this embodiment, the processor can be configured to perform the following steps via a computer program:

[0117] S1, if there are N playback devices in the target area, repeat steps S2-S3 until the receiving device determines the target information corresponding to each playback device:

[0118] S2, upon receiving a calibration command from the i-th playback device in the target area, activate the calibration mode and acquire the i-th audio played by the i-th playback device;

[0119] S3, calculate target information corresponding to the i-th playback device based on the i-th audio; wherein i is greater than or equal to 1 and less than or equal to N, and the target information is used to indicate parameters for eliminating the influence of the audio played by the i-th playback device on the receiving device.

[0120] Alternatively, the processor described above can be configured to perform the following steps via a computer program:

[0121] S1, If ​​there are N receiving devices in the target area, repeat steps S2-S4 until each receiving device determines the target information corresponding to the playback device:

[0122] S2, a calibration command is sent to the i-th receiving device in the target area, instructing the i-th receiving device to enable calibration mode;

[0123] S3, play the i-th audio, wherein, when the i-th receiving device is in calibration mode, it acquires the i-th audio played by the playback device and calculates the target information corresponding to the playback device based on the i-th audio;

[0124] S4, upon receiving feedback information from the i-th receiving device, it is determined that the i-th receiving device has identified target information corresponding to the playback device; wherein i is greater than or equal to 1 and less than or equal to N, and the target information is used to indicate parameters for eliminating the influence of the audio played by the playback device on the receiving device.

[0125] In one exemplary embodiment, the electronic device may further include a transmission device and an input / output device, wherein the transmission device is connected to the processor and the input / output device is connected to the processor.

[0126] Specific examples in this embodiment can be found in the examples described in the above embodiments and exemplary implementations, and will not be repeated here.

[0127] It is obvious to those skilled in the art that the modules or steps of the present invention described above can be implemented using general-purpose computing devices. They can be centralized on a single computing device or distributed across a network of multiple computing devices. They can be implemented using computer-executable program code, and thus can be stored in a storage device for execution by a computing device. In some cases, the steps shown or described can be performed in a different order than those described herein, or they can be fabricated as separate integrated circuit modules, or multiple modules or steps can be fabricated as a single integrated circuit module. Thus, the present invention is not limited to any particular combination of hardware and software.

[0128] The above description is only a preferred embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this application, and these improvements and modifications should also be considered within the scope of protection of this application.

Claims

1. An audio calibration method, characterized by, include: If there are N playback devices in the target area, repeat the following steps until the receiving device determines the target information corresponding to each playback device: Upon receiving a calibration command from the i-th playback device in the target area, the calibration mode is activated, and the i-th audio signal played by the i-th playback device is acquired. Calculate the target information corresponding to the i-th playback device based on the i-th audio; Wherein, i is greater than or equal to 1 and less than or equal to N, and the target information is used to indicate parameters for eliminating the influence of the audio played by the i-th playback device on the receiving device, wherein the receiving device and the N playback devices all have voice functions and the same wake word.

2. The method of claim 1, wherein, In the case that the receiving device does not have the target number of microphones, the target information corresponding to the i-th playback device is calculated based on the i-th audio, including: Obtain the first audio energy of the i-th audio signal sent by the i-th playback device; An interference coefficient corresponding to the i-th playback device is determined based on the first audio energy and the second audio energy, wherein the second audio energy is the audio energy of the i-th audio determined during the acquisition of the i-th audio, and the target information includes the interference coefficient.

3. The method according to claim 1 or 2, characterized in that, In the case that the receiving device does not have the target number of microphones, the target information corresponding to the i-th playback device is calculated based on the i-th audio, including: Obtain the first set of audio energy of the i-th audio sent by the i-th playback device, wherein each audio energy in the first set of audio energy corresponds one-to-one with each volume in a set of volumes; Based on the first set of audio energy and the second set of audio energy, a set of interference coefficients corresponding to the i-th playback device is determined, wherein each audio energy in the second set of audio energy is the audio energy of the i-th audio determined during the process of collecting the i-th audio at each volume in the set of volume; the target information includes the set of interference coefficients, and each interference coefficient in the set of interference coefficients corresponds one-to-one with each volume in the set of volume.

4. The method of claim 1, wherein, When the receiving device has a target number of microphones, the target information corresponding to the i-th playback device is calculated based on the i-th audio, including: The positional relationship between the receiving device and the i-th playback device is calculated based on the i-th audio, wherein the target information includes the positional relationship.

5. The method according to any of claims 1-2, characterized by, When the receiving device does not have the target number of microphones, and the target information corresponding to the i-th playback device has been determined, the method further includes: During the process of the receiving device acquiring the wake-up audio, while the i-th playback device is playing the target audio, the audio energy of the acquired audio is determined, and the audio energy of the target audio sent by the i-th playback device is obtained; The audio energy of the wake-up audio is obtained by subtracting the product of the audio energy of the target audio and the target interference coefficient from the audio energy of the acquired audio, wherein the target information includes the target interference coefficient.

6. The method according to claim 1 or 4, characterized in that, When the receiving device has a target number of microphones and the target information corresponding to the i-th playback device is determined, the method further includes: During the process of the receiving device acquiring the wake-up audio, while the i-th playback device is playing the target audio, noise reduction processing is performed on the target audio in the target direction. The target information includes a positional relationship indicating that the i-th playback device is located in the target direction of the receiving device.

7. An audio calibration method, characterized by, include: If there are N receiving devices in the target area, repeat the following steps until each receiving device determines the target information corresponding to the playback device: A calibration command is sent to the i-th receiving device in the target area, instructing the i-th receiving device to enable calibration mode; Play the i-th audio, wherein the i-th receiving device, when the calibration mode is enabled, acquires the i-th audio played by the playback device, and calculates the target information corresponding to the playback device based on the i-th audio; If feedback information from the i-th receiving device is obtained, it is determined that the i-th receiving device has identified target information corresponding to the playback device; Wherein, i is greater than or equal to 1 and less than or equal to N, the target information is used to indicate parameters for eliminating the influence of the audio played by the playback device on the receiving device, wherein the playback device and the N receiving devices all have voice functions and the same wake word.

8. A receiving device, characterized by include: The first calibration module is used to repeatedly execute the following steps when there are N playback devices in the target area, until the receiving device determines the target information corresponding to each playback device: Upon receiving a calibration command from the i-th playback device in the target area, the calibration mode is activated, and the i-th audio signal played by the i-th playback device is acquired. Calculate the target information corresponding to the i-th playback device based on the i-th audio; Wherein, i is greater than or equal to 1 and less than or equal to N, and the target information is used to indicate parameters for eliminating the influence of the audio played by the i-th playback device on the receiving device, wherein the receiving device and the N playback devices all have voice functions and the same wake word.

9. A playback device, characterized by comprising: include: The second calibration module is used to repeatedly execute the following steps when there are N receiving devices in the target area, until each receiving device determines the target information corresponding to the playback device: A calibration command is sent to the i-th receiving device in the target area, instructing the i-th receiving device to enable calibration mode; Play the i-th audio, wherein the i-th receiving device, when the calibration mode is enabled, acquires the i-th audio played by the playback device, and calculates the target information corresponding to the playback device based on the i-th audio; If feedback information from the i-th receiving device is obtained, it is determined that the i-th receiving device has identified target information corresponding to the playback device; Wherein, i is greater than or equal to 1 and less than or equal to N, the target information is used to indicate parameters for eliminating the influence of the audio played by the playback device on the receiving device, wherein the playback device and the N receiving devices all have voice functions and the same wake word.

10. A computer readable storage medium, characterized in that, The computer-readable storage medium includes a stored program, wherein the program, when executed, performs the method of any one of claims 1 to 7.

11. An electronic device comprising a memory and a processor, characterized in that, The memory stores a computer program, and the processor is configured to execute the method of any one of claims 1 to 7 through the computer program.