Electronic device and audio signal playing method

By acquiring channel parameters and determining the fusion scheme for gain adjustment, the problem of channel sensitivity mismatch between display devices and audio devices was solved, achieving synchronized playback of audio signals and accuracy of the sound field, thus improving the joint sound effect.

CN120812482BActive Publication Date: 2026-07-21SHENZHEN XINYANG CHUANGZHI TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SHENZHEN XINYANG CHUANGZHI TECHNOLOGY CO LTD
Filing Date
2025-07-11
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

When the display device and the audio device are not specific matching devices or are from different manufacturers, the mismatch in channel sensitivity leads to large differences in the loudness of the audio signal, resulting in sound positioning deviation, sound field imbalance and dynamic range compression, which affects the combined sound production effect.

Method used

By acquiring channel parameters between the first and second devices, the target fusion scheme is determined, and gain adjustment is performed based on sensitivity and preset adjustment strategies to ensure that the audio signal is played synchronously in different channels.

Benefits of technology

It achieves the matching of audio signal loudness between different channels, ensures accurate sound positioning, corrects sound field errors, maintains a reasonable dynamic range, and improves the combined sound effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120812482B_ABST
    Figure CN120812482B_ABST
Patent Text Reader

Abstract

The present disclosure relates to an electronic device and an audio signal playing method. The first device comprises a controller configured to, in response to an operation of playing a target audio signal, acquire a first sound channel parameter in a case where the first device is connected to a second device, send a target request message, receive a target response message corresponding to the target request message, determine a target fusion scheme based on a first sound channel and a second sound channel, acquire a first audio signal and a second audio signal from the target audio signal, perform gain adjustment processing on the loudness of the first audio signal and the second audio signal based on a first sensitivity, a second sensitivity, and a preset adjustment strategy to obtain a third audio signal and a fourth audio signal to match the sensitivity difference of a third sound channel and a fourth sound channel, send target information based on the fourth sound channel and the fourth audio signal, and play the third audio signal through the third sound channel to synchronize the playing of the third audio signal and the fourth audio signal.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to terminal technology. More specifically, it relates to an electronic device and a method for playing audio signals. Background Technology

[0002] Currently, the common solutions for joint sound production by display and audio devices involve the display device playing the corresponding channel signal from the audio signal according to its own channels, and the audio device also playing the corresponding channel signal from the audio signal according to its own channels. However, when the display and audio devices are not specific matching devices, or when they are not from the same manufacturer, the sensitivity mismatch between the display and audio channels will lead to significant differences in the loudness of the audio signals played by different channels. This results in problems such as sound positioning deviation, sound field imbalance, and dynamic range compression (i.e., loss of detail), which seriously affect the joint sound production effect. Summary of the Invention

[0003] In order to solve the above-mentioned technical problems, or at least partially solve the above-mentioned technical problems, the present disclosure provides an electronic device and an audio signal playback method, which can improve the effect of combined sound generation.

[0004] In a first aspect, embodiments of this disclosure provide a first device, including: a communicator; and a controller configured to: in response to an operation of playing a target audio signal, when the first device is connected to a second device, acquire first channel parameters, the first channel parameters including a first channel supported by the first device and a first sensitivity, the first sensitivity indicating the sensitivity corresponding to each channel in the first channel; send a target request message, the target request message being used to request channel parameters of the second device; receive a target response message corresponding to the target request message, the target response message carrying second channel parameters, the second channel parameters being acquired by the second device in response to the target request message, the second channel parameters including a second channel supported by the second device and a second sensitivity, the second sensitivity indicating the sensitivity corresponding to each channel in the second channel; and determine a target fusion scheme based on the first channel and the second channel, the target fusion scheme including a third channel in the first channel used for playing the target audio signal. The system uses the third and second channels to play the target audio signal in a fourth channel. From the target audio signal, it acquires a first audio signal corresponding to the third channel and a second audio signal corresponding to the fourth channel. The first audio signal includes audio signals corresponding to each channel in the third channel, and the second audio signal includes audio signals corresponding to each channel in the fourth channel. Based on a first sensitivity, a second sensitivity, and a preset adjustment strategy, it determines the gain adjustment amount for each channel in the third and fourth channels. Based on the gain adjustment amount for each channel, it performs gain adjustment processing on the loudness of the first and second audio signals respectively to obtain the third and fourth audio signals, matching the sensitivity difference between the third and fourth channels. Based on the fourth channel and the fourth audio signal, it sends target information, which instructs the second device to play the fourth audio signal through the fourth channel. Finally, it plays the third audio signal through the third channel to synchronize the playback of the third and fourth audio signals.

[0005] Secondly, embodiments of this disclosure provide a second device, including: a communicator; and a controller configured to: receive a target request message, the target request message being used to request channel parameters of the second device, the target request message being generated by a first device in response to an operation of playing a target audio signal, and wherein the first device is connected to the second device; the first device also in response to the operation of playing the target audio signal acquires first channel parameters, the first channel parameters including a first channel supported by the first device and a first sensitivity, the first sensitivity being used to indicate the sensitivity corresponding to each channel in the first channel; in response to the target request message, acquire second channel parameters, the second channel parameters including a second channel supported by the second device and a second sensitivity, the second sensitivity being used to indicate the sensitivity corresponding to each channel in the second channel; and send a target response message corresponding to the target request message, the target response message carrying the second channel parameters, so that the first device, based on the first channel and the second channel, determines a channel in the first channel used for playing the target audio signal. The system utilizes a target fusion scheme for the fourth channel in the third and second channels to play the target audio signal. It obtains a first audio signal corresponding to the third channel and a second audio signal corresponding to the fourth channel from the target audio signal. Based on a first sensitivity, a second sensitivity, and a preset adjustment strategy, it determines the gain adjustment amount for each channel in the third and fourth channels. Based on the gain adjustment amount for each channel, it performs gain adjustment processing on the loudness of the first and second audio signals respectively to obtain the third and fourth audio signals, matching the sensitivity difference between the third and fourth channels. The first audio signal includes the audio signals corresponding to each channel in the third channel, and the second audio signal includes the audio signals corresponding to each channel in the fourth channel. It receives target information, which is generated based on the fourth channel and the fourth audio signal. In response to the target information, it plays the fourth audio signal through the fourth channel, so that the third audio signal played by the first device through the third channel is played synchronously with the fourth audio signal.

[0006] Thirdly, embodiments of this disclosure provide an audio signal playback method applied to a first device, comprising: in response to an operation of playing a target audio signal, when the first device is connected to a second device, acquiring first channel parameters, the first channel parameters including a first channel supported by the first device and a first sensitivity, the first sensitivity being used to indicate the sensitivity corresponding to each channel in the first channel; sending a target request message, the target request message being used to request channel parameters of the second device; receiving a target response message corresponding to the target request message, the target response message carrying second channel parameters, the second channel parameters being acquired by the second device in response to the target request message, the second channel parameters including a second channel supported by the second device and a second sensitivity, the second sensitivity being used to indicate the sensitivity corresponding to each channel in the second channel; and determining a target fusion scheme based on the first channel and the second channel, the target fusion scheme including a third channel in the first channel used for playing the target audio signal. The system uses a fourth channel in the second and third channels to play the target audio signal. From the target audio signal, it acquires a first audio signal corresponding to the third channel and a second audio signal corresponding to the fourth channel. The first audio signal includes audio signals corresponding to each channel in the third channel, and the second audio signal includes audio signals corresponding to each channel in the fourth channel. Based on a first sensitivity, a second sensitivity, and a preset adjustment strategy, it determines the gain adjustment amount for each channel in the third and fourth channels. Based on the gain adjustment amount for each channel, it performs gain adjustment processing on the loudness of the first and second audio signals respectively to obtain the third and fourth audio signals, matching the sensitivity difference between the third and fourth channels. Based on the fourth channel and the fourth audio signal, it sends target information, which instructs the second device to play the fourth audio signal through the fourth channel. Finally, it plays the third audio signal through the third channel to synchronize the playback of the third and fourth audio signals.

[0007] Fourthly, embodiments of this disclosure provide an audio signal playback method applied to a second device, comprising: receiving a target request message, the target request message being used to request channel parameters of the second device, the target request message being generated by a first device in response to an operation of playing a target audio signal, and the first device being connected to the second device; the first device further acquiring first channel parameters in response to the operation of playing the target audio signal, the first channel parameters including a first channel supported by the first device and a first sensitivity, the first sensitivity being used to indicate the sensitivity corresponding to each channel in the first channel; acquiring second channel parameters in response to the target request message, the second channel parameters including a second channel supported by the second device and a second sensitivity, the second sensitivity being used to indicate the sensitivity corresponding to each channel in the second channel; and sending a target response message corresponding to the target request message, the target response message carrying the second channel parameters, so that the first device determines, based on the first channel and the second channel, a channel for playing the target audio signal in the first channel. The system utilizes a target fusion scheme for the fourth channel in the third and second channels to play the target audio signal. It obtains a first audio signal corresponding to the third channel and a second audio signal corresponding to the fourth channel from the target audio signal. Based on a first sensitivity, a second sensitivity, and a preset adjustment strategy, it determines the gain adjustment amount for each channel in the third and fourth channels. Based on the gain adjustment amount for each channel, it performs gain adjustment processing on the loudness of the first and second audio signals respectively to obtain the third and fourth audio signals, matching the sensitivity difference between the third and fourth channels. The first audio signal includes the audio signals corresponding to each channel in the third channel, and the second audio signal includes the audio signals corresponding to each channel in the fourth channel. It receives target information, which is generated based on the fourth channel and the fourth audio signal. In response to the target information, it plays the fourth audio signal through the fourth channel, so that the third audio signal played by the first device through the third channel is played synchronously with the fourth audio signal.

[0008] Fifthly, embodiments of this disclosure provide a computer-readable storage medium, including: storing a computer program on the computer-readable storage medium, wherein when the computer program is executed by a processor, it implements the audio signal playback method as shown in the third or fourth aspect.

[0009] In a sixth aspect, embodiments of this disclosure provide a computer program product, including: when the computer program product is run on a computer, causing the computer to implement the audio signal playback method as shown in the third or fourth aspect.

[0010] Compared with the prior art, the technical solution provided in this disclosure has the following advantages: In this disclosure, in response to the operation of playing a target audio signal, when the first device is connected to the second device, a first channel parameter is obtained. The first channel parameter includes a first channel supported by the first device and a first sensitivity, wherein the first sensitivity is used to indicate the sensitivity corresponding to each channel in the first channel; a target request message is sent, which requests the channel parameters of the second device; a target response message corresponding to the target request message is received, which carries a second channel parameter. The second channel parameter is obtained by the second device in response to the target request message. The second channel parameter includes a second channel supported by the second device and a second sensitivity, wherein the second sensitivity is used to indicate the sensitivity corresponding to each channel in the second channel; based on the first channel and the second channel, a target fusion scheme is determined, which includes a third channel in the first channel used for playing the target audio signal. The system uses the third and second channels to play the target audio signal in a fourth channel. From the target audio signal, it acquires a first audio signal corresponding to the third channel and a second audio signal corresponding to the fourth channel. The first audio signal includes audio signals corresponding to each channel in the third channel, and the second audio signal includes audio signals corresponding to each channel in the fourth channel. Based on a first sensitivity, a second sensitivity, and a preset adjustment strategy, it determines the gain adjustment amount for each channel in the third and fourth channels. Based on the gain adjustment amount for each channel, it performs gain adjustment processing on the loudness of the first and second audio signals respectively to obtain the third and fourth audio signals, matching the sensitivity difference between the third and fourth channels. Based on the fourth channel and the fourth audio signal, it sends target information, which instructs the second device to play the fourth audio signal through the fourth channel. Finally, it plays the third audio signal through the third channel to synchronize the playback of the third and fourth audio signals. In this way, by balancing the loudness of the combined sound from different types of channels in the third and fourth channels, the loudness of the audio signals played between the channels in the target fusion scheme can be matched with each other, with a certain regularity and no significant differences. This ensures accurate sound positioning, corrects sound field errors, maintains a reasonable dynamic range, and improves the effect of combined sound production. Attached Figure Description

[0011] To more clearly illustrate the implementation methods in the embodiments of this disclosure or related technologies, the accompanying drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of this disclosure, and those skilled in the art can obtain other drawings based on these drawings.

[0012] Figure 1 An operational scenario between a display device and a control device according to some embodiments is illustrated;

[0013] Figure 2 A hardware configuration block diagram of a control device 100 according to some embodiments is shown;

[0014] Figure 3 A hardware configuration block diagram of a display device 200 according to some embodiments is shown;

[0015] Figure 4 One of the schematic diagrams of a sound combination scheme in which a display device and an audio device jointly produce sound according to some embodiments is shown;

[0016] Figure 5 A second schematic diagram of a sound combination scheme in which a display device and an audio device jointly produce sound according to some embodiments is shown;

[0017] Figure 6 The third schematic diagram illustrates a sound combination scheme in which a display device and an audio device jointly produce sound according to some embodiments;

[0018] Figure 7 The fourth schematic diagram illustrates a sound combination scheme in which a display device and an audio device jointly produce sound according to some embodiments;

[0019] Figure 8 A schematic diagram illustrating that the relative positions of a display device, an audio device, and a remote control according to some embodiments satisfy target conditions;

[0020] Figure 9 A bandwidth diagram of the same type of audio channel is shown for an audio device and a display device according to some embodiments;

[0021] Figure 10 A flowchart illustrating an audio signal playback method according to some embodiments is shown. Detailed Implementation

[0022] To make the objectives and implementation methods of this disclosure clearer, the exemplary embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings of the exemplary embodiments of this disclosure. Obviously, the exemplary embodiments described are only some embodiments of this disclosure, and not all embodiments.

[0023] It should be noted that the brief descriptions of terms in this disclosure are only for the convenience of understanding the embodiments described below, and are not intended to limit the embodiments of this disclosure. Unless otherwise stated, these terms should be understood in their ordinary and common meaning.

[0024] The terms "first," "second," "third," etc., used in this disclosure, in the specification, claims, and accompanying drawings are used to distinguish similar or related objects or entities, and do not necessarily imply a specific order or sequence, unless otherwise specified. It should be understood that such terms are interchangeable where appropriate.

[0025] The terms “comprising” and “having”, and any variations thereof, are intended to cover but not exclude inclusion, for example, a product or device that includes a range of components is not necessarily limited to all of the components that are clearly listed, but may include other components that are not clearly listed or that are inherent to such product or device.

[0026] The display device provided in this disclosure can have various implementation forms, such as a television, a smart television, a laser projection device, a monitor, an electronic bulletin board, an electronic table, a mobile phone, a tablet computer, a laptop computer, a handheld computer, an in-vehicle electronic device, etc.

[0027] Figure 1 This is a schematic diagram illustrating an operational scenario between a display device and a control device according to an embodiment, wherein the control device includes a smart device or a control apparatus. Figure 1 As shown, the user can operate the display device 200 through the smart device 300 or the control device 100.

[0028] In some embodiments, the control device 100 may be a remote control. Communication between the remote control and the display device includes infrared protocol communication, Bluetooth protocol communication, and other short-range communication methods, controlling the display device 200 wirelessly or via wired means. Users can control the display device 200 by inputting user commands through buttons on the remote control, voice input, control panel input, etc.

[0029] In some embodiments, a smart device 300 (such as a mobile terminal, tablet computer, computer, laptop computer, etc.) can also be used to control the display device 200. For example, an application running on the smart device can be used to control the display device 200.

[0030] In some embodiments, the display device may receive instructions not through the aforementioned smart devices or control devices, but through touch or gestures.

[0031] In some embodiments, the display device 200 can also be controlled in ways other than the control device 100 and the smart device 300. For example, it can be controlled by directly receiving the user's voice commands through a module configured inside the display device 200 for acquiring voice commands, or it can be controlled by receiving the user's voice commands through a voice control device set outside the display device 200.

[0032] In some embodiments, the display device 200 also communicates with the server 400. The display device 200 may communicate via a local area network (LAN), wireless local area network (WLAN), and other networks. The server 400 may provide various content and interactive features to the display device 200. The server 400 may be a cluster or multiple clusters, and may include one or more types of servers.

[0033] Figure 2 An exemplary block diagram of the configuration of the control device 100 according to an exemplary embodiment is shown. Figure 2 As shown, the control device 100 includes a controller 110, a communication interface 130, a user input / output interface 140, an external memory, and a power supply. The control device 100 can receive user input operation commands and convert the operation commands into commands that the display device 200 can recognize and respond to, thus acting as an intermediary for interaction between the user and the display device 200.

[0034] like Figure 3 The display device 200 includes at least one of the following: a tuner 210, a communicator 220, a detector 230, an external device interface 240, a controller 250, a display 260, an audio output interface 270, a user interface 280, an external memory, and a power supply.

[0035] In some embodiments, the controller includes a processor, a video processor, an audio processor, a graphics processor, RAM, ROM, and a first interface to an nth interface for input / output.

[0036] The display 260 includes a display screen assembly for presenting images, a driving assembly for driving image display, a component for receiving image signals from the controller output, and a user control UI interface for displaying video content, image content, menu control interface, and user control UI interface.

[0037] The display 260 can be an LCD display, an OLED display, or a projection display, and can also be a projection device and a projection screen.

[0038] The communicator 220 is a component used to communicate with external devices or servers according to various communication protocol types. For example, the communicator may include at least one of the following: a Wi-Fi module, a Bluetooth module, a wired Ethernet module, other network communication protocol chips or near-field communication protocol chips, and an infrared receiver. The display device 200 can establish the transmission and reception of control signals and data signals with the external control device 100 or the server 400 through the communicator 220.

[0039] User interface 280 can be used to receive control signals from control device 100 (such as an infrared remote control). It can also be used to directly receive user input operation commands and convert the operation commands into commands that the display device 200 can recognize and respond to; in this case, it can be called a user input interface.

[0040] Detector 230 is used to collect signals from the external environment or to interact with the external environment. For example, detector 230 includes a light receiver, a sensor for collecting ambient light intensity; or, detector 230 includes an image acquisition device, such as a camera, which can be used to collect external environmental scenes, user attributes, or user interaction gestures; or, detector 230 includes a sound acquisition device, such as a microphone, for receiving external sounds.

[0041] The external device interface 240 may include, but is not limited to, one or more of the following: High Definition Multimedia Interface (HDMI), analog or high-definition component input interface (component), composite video input interface (CVBS), USB input interface (USB), RGB port, etc. It may also be a composite input / output interface formed by multiple interfaces mentioned above.

[0042] The tuner / demodulator 210 receives broadcast television signals via wired or wireless means, and demodulates audio and video signals, such as EPG data signals, from multiple wireless or wired broadcast television signals.

[0043] In some embodiments, the controller 250 and the tuner 210 may be located in different separate devices, that is, the tuner 210 may also be located in an external device of the main device where the controller 250 is located, such as an external set-top box.

[0044] The controller 250 controls the operation of the display device and responds to user operations through various software control programs stored in memory (internal or external memory). The controller 250 controls the overall operation of the display device 200. For example, in response to receiving a user command to select a UI object to display on the monitor 260, the controller 250 can perform operations related to the object selected by the user command.

[0045] In some embodiments, the controller includes at least one of a central processing unit (CPU), a video processor, an audio processor, a graphics processing unit (GPU), and random access memory (RAM), read-only memory (ROM), a first to an nth interface for input / output, a communication bus, etc.

[0046] RAM, also known as main memory, is an internal memory that directly exchanges data with the controller. It can be read and written at any time (except during refresh) and is very fast, typically serving as temporary data storage for the operating system or other running programs. Its biggest difference from ROM is data volatility; data stored in RAM is lost when power is off. RAM is used in computers and digital systems to temporarily store programs, data, and intermediate results. ROM operates in a non-destructive read-only manner; information can only be read, not written. Once information is written, it is fixed and will not be lost even if power is cut off; therefore, it is also called fixed-function memory.

[0047] Users can input commands through a graphical user interface (GUI) displayed on the monitor 260, and the user input interface receives the user input commands through the GUI. Alternatively, users can input commands by entering specific sounds or gestures, and the user input interface receives the user input commands by recognizing the sounds or gestures through sensors.

[0048] A "user interface" is the medium through which an application or operating system interacts and exchanges information with the user. It converts information from its internal form to a form that the user can accept. A common form of user interface is the graphical user interface (GUI), which refers to a user interface related to computer operation displayed graphically. It can be an icon, window, control, or other interface element displayed on the screen of a display device. Controls can include visual interface elements such as icons, buttons, menus, tabs, text boxes, dialog boxes, status bars, navigation bars, and widgets.

[0049] This disclosure provides an electronic device and an audio signal playback method. In one case, the first device is a display device and the second device is an audio device. In another case, the first device is an audio device and the second device is a display device. The specific device can be determined according to the actual situation and is not limited here.

[0050] Taking televisions (TVs) as an example, more and more TV manufacturers are supporting combined sound output from TVs and their own soundbars. Due to differences in hardware capabilities across different TV platforms—for instance, some TVs are 2.0 channels, some 2.1 channels, some 2.1.2 channels, some 4.1 channels, some 4.1.2 channels, and some 5.1.2 channels; some models have separate overhead speakers, with overhead sound output from the TV, while others lack overhead speakers and output from the soundbar—and the soundbar itself varies in support—some models support 2.1 channels, some 3.1 channels, some 4.1.2 channels, and some 5.1.2 channels. The current strategy is to pair a specific TV model with a specific soundbar model for combined sound output. However, with the increasing number of TV and soundbar models supporting combined sound output, it's impossible to rigidly bind users to specific product types. Therefore, if the display device and the audio device that support joint sound production are not specific matching devices, or if the display device and the audio device are not from the same manufacturer, the channel delay of each channel of the display device will be different from that of each channel of the audio device. This will cause the display device and the audio device to be unable to produce sound synchronously, resulting in problems such as chaotic sound field positioning and spatial distortion, which will seriously affect the joint sound production effect.

[0051] In particular, when the display device and the audio device include the same channels, the two channels of the same type of display device and audio device need to play the same audio signal synchronously. However, since the channel delays of the two channels of the same type are different, the playback time of the same audio signal will be different, resulting in problems such as chaotic sound field positioning and spatial distortion, which seriously affects the joint sound effect.

[0052] In some embodiments of this disclosure, the television can be a traditional television, a laser television, etc.; the audio equipment can include an all-in-one audio system, an audio system with an external subwoofer, an audio system with a different number of satellite speakers, etc.

[0053] To address the sensitivity differences between different audio channels, when an audio device is emitting sound independently, only the virtual surround environment created by each speaker needs adjustment. However, when integrated with a display device, the amplitude of the audio output from each speaker needs to be readjusted to match the larger virtual surround environment created by the display device. Otherwise, in the new immersive sound environment, the primary and secondary audio information heard by the user will be confused, and the sense of direction in the sound field will be inaccurate.

[0054] This disclosure provides a first device, including: a communicator; and a controller for implementing the audio signal playback method provided in this disclosure.

[0055] The controller is configured to: in response to the operation of playing a target audio signal, when the first device is connected to the second device, acquire first channel parameters, the first channel parameters including the first channel supported by the first device and the first sensitivity, the first sensitivity being used to indicate the sensitivity corresponding to each channel in the first channel.

[0056] The target audio signal is a multi-channel audio signal. A multi-channel audio signal refers to an audio signal containing multiple independent channels, each of which can carry different audio information, thereby providing listeners with a richer, more spatial, and immersive audio experience.

[0057] Among them, the combined sound function, also known as the synthesis function, refers to the function of two or more electronic devices playing audio signals together through their respective sound channel systems.

[0058] The controller is further configured to: send a target request message, which requests the channel parameters of the second device; receive a target response message corresponding to the target request message, which carries the second channel parameters. The second channel parameters are obtained by the second device in response to the target request message. The second channel parameters include the second channel supported by the second device and the second sensitivity. The second sensitivity is used to indicate the sensitivity of each channel in the second channel.

[0059] It is understood that, in response to the operation of playing the target audio signal, the first device sends a target request message to the second device when it detects a communication connection with the second device and detects that the first device supports the harmony function, and receives a target response message from the second device based on the target request message.

[0060] In some embodiments of this disclosure, the second device may be assumed to support the harmony function, or a request message may be used to request whether the second device supports the harmony function (hereinafter referred to as requesting to obtain the harmony capability of the second device). The specific method can be determined according to the actual situation and is not limited here.

[0061] In some embodiments of this disclosure, the chord capability of the second device and the channel parameters of the second device can be requested through a single request message, or two separate request messages can be used to request the chord capability of the second device and the channel parameters of the second device. The specific method can be determined according to the actual situation and is not limited here.

[0062] In some embodiments of this disclosure, the target request message includes a first request message and a second request message, and the target response message includes a first response message and a second response message. The process of requesting to obtain the chord capability of the second device and the channel parameters of the second device specifically includes: in response to the operation of playing the target audio signal, if the first device supports the chord function, sending a first request message to the second device, the first request message being used to request the chord capability of the second device; receiving a first response message corresponding to the first request message from the second device, the first response message carrying the chord capability information of the second device; if the chord capability information indicates that the second device supports the chord function, sending a second request message to the second device, the second request message being used to request the channel parameters of the second device; and receiving a second response message corresponding to the second request message from the second device, the second response message carrying the second channel parameters.

[0063] The target request message may also include other request messages, and the target response message may also include other response messages; this is not limited here.

[0064] The controller is further configured to: determine a target fusion scheme based on the first channel and the second channel, the target fusion scheme including a third channel in the first channel for playing the target audio signal and a fourth channel in the second channel for playing the target audio signal.

[0065] In this embodiment of the disclosure, information interaction between the first device and the second device can be achieved by extending the wired or wireless protocol between the first device and the second device.

[0066] In some embodiments of this disclosure, the first device can interact with the second device based on the I2C protocol to obtain the second device's chordal capabilities and the second device's channel parameters, thereby determining the target fusion scheme so that the first device and the second device can jointly produce sound through their respective channels.

[0067] In some embodiments of this disclosure, the first device can interact with the second device based on a broadcast protocol to obtain the chordal capability of the second device and the channel parameters of the second device, thereby determining the target fusion scheme so that the first device and the second device can jointly emit sound through their respective channels.

[0068] In some embodiments of this disclosure, based on the extended HDMI CEC communication protocol between the first device and the second device, the first device obtains the synthesis capability of the second device and the channel parameters of the second device through information interaction with the second device, thereby determining the target fusion scheme so that the first device and the second device can jointly emit sound through their respective channels. This can avoid the problem of the first device and the second device not being able to emit sound synchronously due to the different channel delays of different electronic devices, which would lead to problems such as chaotic sound field positioning and spatial distortion.

[0069] In some embodiments of this disclosure, the following functions can be achieved by extending the manufacturer-customizable parts of the HDMI CEC communication protocol: 1. Adding an information interaction function to indicate whether the chorus function is supported. The second device responds after receiving a first request message from the first device requesting the chorus capability of the second device. After the first and second devices verify their identities during the CEC handshake process, the first device can request the chorus capability of the second device, and the second device responds correctly according to its own capabilities. 2. Adding an information interaction function to request the channel parameters of the second device. After the speaker and the first device have communicated that the chorus function is supported, the first device requests the channel parameters of the second device.

[0070] In some embodiments of this disclosure, it is necessary to design information interaction instructions that satisfy the HDMICEC communication protocol, such as setting the format of a first request message, setting the format of a first response message, setting the format of a second request message, setting the format of a second response message, or setting the format of a target request message and setting the format of a target response message.

[0071] In some embodiments of this disclosure, the second device may also request the channel parameters of the first device, which is not limited here.

[0072] For example, the first and second devices complete identity verification through a CEC handshake (which can be done by matching the supplier identification code or vendor ID, or by other methods, which are not limited here). If the first device supports the chord function, it requests the chord capability of the second device (first request message). Upon receiving this instruction, the second device reports its chord capability (first response message). The first device parses whether the second device supports the chord function. If it does, it further requests the channel parameters of the second device (second request message). The second device responds according to its own chord capability. The first device can also proactively report its own channel parameters to the second device. Upon receiving this, the second device stores the channel parameters of the first device for later use.

[0073] The second channel parameter is used to indicate the second channel supported by the second device. It may include the device type of the second device (e.g., 2.1 channel, 5.1 channel, etc.) or the type of channel supported by the second device. The specific details can be determined according to the actual situation and are not limited in this case.

[0074] The second channel parameters may include channel parameters used to determine the target fusion scheme, such as the second channel itself, or detailed parameter information of the channels supported by the second device (e.g., low-frequency extension (Lf) and power (P) of the subwoofer channel), or the relative position information of the channels (corresponding speakers) in the second device; the second channel parameters may also include channel parameters used to adjust the combined sound effect, such as the channel delay, sensitivity, and bandwidth of each channel in the second channel; the specific second channel parameters can be determined according to actual usage requirements, and are not limited here.

[0075] The description of the parameters of the first channel can be found in the description of the parameters of the second channel described above, and will not be repeated here.

[0076] In some embodiments of this disclosure, the third and fourth channels included in the target fusion scheme may or may not have the same type of channels, and this is not limited here.

[0077] In some embodiments of this disclosure, the controller is specifically configured to: determine the first channel as the third channel and the second channel as the fourth channel to determine the target fusion scheme. This allows for rapid determination of the target fusion scheme, and if the first and second channels contain channels of the same type, then the target fusion scheme includes channels of the same type in the third and fourth channels.

[0078] In some embodiments of this disclosure, a first set of coefficients corresponding to mapping the first channel to the preset channel system is obtained. Each element of the first set of coefficients is 1 or 0, and an element of 1 in the first set of coefficients represents a channel included in both the preset channel system and the first channel. A second set of coefficients corresponding to mapping the second channel to the preset channel system is obtained. Each element of the second set of coefficients is 1 or 0, and an element of 1 in the second set of coefficients represents a channel included in both the preset channel system and the second channel. Based on the first set of coefficients and the second set of coefficients, the target fusion scheme is determined. This facilitates the determination of the target fusion scheme and makes it easier to determine the changed fusion scheme when a certain channel is faulty or disconnected.

[0079] For example, the first device is a TV speaker, and the second device is an external speaker. A completely new channel system, such as the 7-1-4 fusion system, is defined, with the external speaker and TV mapped and managed to this system respectively. Based on the existing hardware and referencing the original Dolby architecture, the parameters in the fusion system are defined. The 7-1-4 system is defined as follows: C-FL (left channel), C-FR (right channel), C-Center (center channel), C-SR (right surround channel), C-SL (left surround channel), C-BR (right rear surround channel), C-BL (left rear surround channel), C-SUB (subwoofer channel), C-FTR (front top right channel), C-FTL (front top left channel), C-RTR (rear top right channel), and C-RTL (rear top left channel), where C (Combine) signifies fusion or combination. If the audio equipment follows the Dolby architecture and is a 5.1.2 system (FL, FR, Center, SR, SL, SUB, FTR, FTL), then the mapping in the fusion system is: C-FL = C1*FL, C-FR = C2*FR, C-Center = C3*Center, C-SR = C4*SR, C-SL = C5*SL, C-RSR = C6*RR, C-RSL = C7*RL, C-SUB = C8*SUB, C-FTR = C9*FTR, C-FTL = C10*FTL, C-RTR = C11*RTR, C-RTL = C12*RTL. Where C1~C12 are 0 or 1, as in the system above 5.1.2, the coefficients can be: C1~C12=(1,1,1,1,1,0,0,1,1,1,0,0). Before system fusion, the TV system parameters TV Speaker (C1~C12) (i.e., the first set of coefficients) and the external speaker system parameters External Speaker (C1~C12) (i.e., the second set of coefficients) are defined respectively. Based on TV Speaker (C1~C12) and External Speaker (C1~C12), the target fusion scheme is determined.

[0080] In some embodiments of this disclosure, the controller is specifically configured to: when there are channels of the same type in the first and second channels, set one of the two channels of the same type as the target channel or mute it to determine the target fusion scheme; when there are no channels of the same type in the third and fourth channels, the target channel is a channel that does not exist in either the first or second channels. In this way, the first device and the second device can jointly emit sound through channels of different types, thereby avoiding the channel conflict and poor sound quality caused by the first device and the second device using the same type of channel to broadcast audio signals.

[0081] In some embodiments of this disclosure, the first device may determine a matching target fusion scheme from at least one preset fusion scheme based on the first channel and the second channel; the first device may also determine the target fusion scheme by combining the first channel and the second channel, as well as the relative position information of the first device and the second device; the specific determination may be made according to the actual situation and is not limited here.

[0082] Among them, at least one preset fusion scheme may include preset fusion schemes corresponding to the first device and various types of second devices respectively, such as preset fusion schemes corresponding to the first device and 2.1 channel second devices, preset fusion schemes corresponding to the first device and 5.1 channel second devices, etc., which are not limited here.

[0083] In some embodiments of this disclosure, the device type (channel system type) corresponding to the first device may be different, and the at least one preset fusion scheme stored in the first device may be different; alternatively, regardless of whether the device types of the first devices are the same, the at least one preset fusion scheme stored in the first device may be the same, and then the target fusion scheme is determined from the at least one preset fusion scheme according to the number and channel type included in the first channel and the second channel; the specific determination can be made according to the actual situation, and is not limited here.

[0084] In some embodiments of this disclosure, multiple candidate fusion schemes matching the first channel can be determined from at least one preset fusion scheme, and then the candidate fusion scheme that best matches the second channel can be selected from the multiple candidate fusion schemes as the target fusion scheme; alternatively, multiple candidate fusion schemes can be displayed to the user, and the user can select the candidate fusion scheme that meets the requirements as the target fusion scheme; the specifics can be determined according to the actual situation, and are not limited here.

[0085] In some embodiments of this disclosure, when the first device is a display device, the controller is further configured to display a first prompt message after determining the target fusion scheme. The first prompt message is used to prompt the user about the relative positional relationship between the display device and the audio device, and / or, the first prompt message is used to prompt the user about the position of the third channel in the display device and the position of the fourth channel in the audio device.

[0086] The first prompt can be a text prompt, a voice prompt, or an image prompt; there is no limitation here.

[0087] In some embodiments of this disclosure, the first device is a display device. After determining the target fusion scheme, the first device can present the target fusion scheme through image prompts. The image prompts display the third channel of the first device, the fourth channel of the second device, the relative positional relationship between the third and fourth channels, and the relative positional relationship between the first device and the second device.

[0088] Example 1: The first device is a Laser TV, and the second device is a 4.1.2 speaker system. Due to the hardware size limitations of the Laser TV, the left and right channels are very close together. Therefore, the left and right channels of the Laser TV are suitable as the center channel. The satellite speakers of the 4.1.2 speaker system can increase the width, surround sound, height, and sound field. Therefore, when the Laser TV is connected to the 4.1.2 speaker system, it defaults to the following mode: Figure 4 As shown, in Laser Home Theater Mode: Laser TV's third channel includes the center channel, and the audio equipment's fourth channel includes left and right channels, subwoofer channel, left surround channel, right surround channel, left sky channel, and right sky channel. The label "41" indicates the components of the TV, and the label "42" indicates the components of the audio equipment. When this mode is selected, Laser TV interacts with the audio equipment to instruct it to play the corresponding audio signal through the fourth channel. Laser TV also plays the corresponding audio signal through the third channel. For example, Laser TV's left and right channels can be used together as the center channel (third channel) to play the center channel's audio signal. The fourth channel, including left and right channels, subwoofer channel, left surround channel, right surround channel, left sky channel, and right sky channel, is used to play the corresponding audio signal from the target audio signal. The audio equipment's center channel is muted or not processed.

[0089] Example 2: The first device is a TV that supports overhead sound, and the audio system is a 5.1 speaker system. In a home environment, the TV is usually placed at a high position. Therefore, the height of the TV can be used, and the overhead sound speakers are generally located near the top of the screen as overhead sound channels. This can provide the following mode 1 (e.g., Figure 5 As shown, where "51" indicates the components of a television and "52" indicates the components of an audio system: the third channel of the TV includes a ceiling channel, and the fourth channel of the audio system includes left and right channels, a center channel, and left and right surround channels; or it provides the following mode 2 (e.g. Figure 6 (As shown): The third channel of a TV includes the left and right channels, the center channel, and the sky channel, while the fourth channel of a speaker includes the left and right surround channels.

[0090] Example 3: The first device is a 2.1.2 TV, and the audio device is a wireless 2.1 speaker. Taking advantage of the ease of placement of the wireless 2.1 channel, the two channels of the speaker can be used as left and right surround channels, forming a 4.1.2 channel configuration together with the TV. (See image prompt for details.) Figure 7 As shown, the label "71" indicates a component of a television, and the label "72" indicates a component of an audio device.

[0091] The aforementioned preset fusion scheme is a target fusion scheme determined based on the first channel, the second channel, and at least one preset fusion scheme. While it resolves the channel conflict issue, it only improves the user experience to a certain extent. Different relative positions of the first and second devices may necessitate different fusion schemes to achieve the best sound effect. For example, in Example 2 above, should mode 1, mode 2, or another mode be used? The only criterion is the sound effect after combined sound production, and two crucial factors affecting sound effect are sensitivity and channel distance difference.

[0092] In some embodiments of this disclosure, when the first device is a display device, the controller is specifically configured to: display a second prompt message, the second prompt message being used to prompt the user whether the relative positions between the first device, the second device, and the remote control meet target conditions, the target conditions being used to indicate: a first perpendicular plane of the line connecting the center of the left channel speaker and the center of the right channel speaker of the first device is parallel to a second perpendicular plane of the line connecting the center of the left channel speaker and the center of the right channel speaker of the second device; the central axis of the remote control is perpendicular to the first perpendicular plane, and the minimum distance from the central axis to a first plane passing through the first device and perpendicular to the first perpendicular plane is less than or equal to a first distance; the minimum distance from the central axis to a second plane passing through the second device and perpendicular to the second perpendicular plane is less than or equal to a second distance; and the distance from the center of the remote control to the center of the first device is greater than or equal to a third distance and less than or equal to a fourth distance, the third distance being less than the fourth distance; in response to a received confirmation operation for the second prompt message, controlling the left and right channel speakers of the first device to sequentially play audio signals, and using the remote control... The microphone of the remote control sequentially receives a first sound signal emitted by the left channel speaker of the first device and a second sound signal emitted by the right channel speaker of the first device; sends a control message to control the left and right channel speakers of the second device to play audio signals sequentially; the microphone of the remote control sequentially receives a third sound signal emitted by the left channel speaker of the second device and a fourth sound signal emitted by the right channel speaker of the speaker; based on the first, second, third, and fourth sound signals, the remote control determines the first and second sound pressure levels corresponding to the left and right channel speakers of the first device, and the third and fourth sound pressure levels corresponding to the left and right channel speakers of the second device, respectively; if the first difference is greater than the second difference, the remote control determines that the third channel includes the left and right channels, and the fourth channel includes the center channel, wherein the first difference is the difference between the first and second sound pressure levels, and the second difference is the difference between the third and fourth sound pressure levels; if the first difference is less than the second difference, the remote control determines that the third channel includes the center channel, and the fourth channel includes the left and right channels.

[0093] The first and second perpendicular planes can coincide, and the minimum distance between the first and second perpendicular planes can be less than a distance threshold. The distance threshold can be determined according to the actual situation and is not limited here.

[0094] Wherein, the first plane is the plane closest to the central axis among the planes passing through the first device and perpendicular to the first perpendicular plane, and the second plane is the plane closest to the central axis among the planes passing through the second device and perpendicular to the first perpendicular plane. The first distance can be greater than or equal to the second distance, or it can be less than the second distance; no limitation is made here.

[0095] The first, second, third, and fourth distances can be determined based on actual usage and are not limited here.

[0096] To improve measurement accuracy, the optimal distance between the remote control's microphone and the first device is between 0.4 meters and 1 meter. Therefore, the optimal distance range between the center of the remote control and the center of the first device can be determined based on the optimal distance range between the remote control's microphone and the first device.

[0097] It is understood that in this embodiment of the disclosure, sensitivity and channel distance are measured using a microphone-enabled remote control of the first device. When the measurement is initiated, the user is prompted to set the relative positions of the first device, the second device, and the remote control to meet the target conditions. Typically, the first and second devices are centered, and their left and right channels are symmetrical. The distance between the speakers of the left and right channels of the first device compared to the distance between the speakers of the left and right channels of the second device can result in three possibilities: greater than, equal to, or less than. For example... Figure 8 As shown, after the relative positions of the first device, the second device, and the remote control meet the target conditions, the measurement program sequentially controls the left and right channels of the first device, the left and right channels of the second device, and the microphone devices record the acquired sound signals of each channel and output the acquired sound signals to the measurement program. The measurement program analyzes the sound signals to obtain the sound pressure level values ​​of the four channels (i.e., the first and second sound pressure levels corresponding to the left and right channel speakers of the first device, and the third and fourth sound pressure levels corresponding to the left and right channel speakers of the second device, respectively), denoted as SL1, SL2, SL3, and SL4. The distances of the microphone from the left and right channel speakers of the first device and the left and right channel speakers of the second device are L1, L2, L3, and L4, respectively. According to the relationship between the sound pressure level of a point source and distance in a free sound field, the following relationship can be obtained:

[0098] SL1-SL2=20lg(L2 / L1)=20lg((L1+D1) / L1)=20lg(1+D1 / L1),

[0099] SL3-SL4=20lg(L4 / L3)=20lg((L3+D2) / L3)=20lg(1+D2 / L3).

[0100] When the measurement program is started, the placement of the microphone on the remote control is predetermined, meaning L1 + D1 / 2 is a constant, and similarly, L3 + D2 / 2 is a constant. Therefore, the larger D1 is, the smaller L1 is, and the smaller D1 is, the larger L1 is; similarly, D2 and L3 follow the same rule. Therefore, according to the formula, the value of SL1 - SL2 depends entirely on D1, and the value of SL3 - SL4 depends entirely on D2. By comparing the values ​​of (SL1 - SL2) and (SL3 - SL4), the values ​​of D1 and D2 can be determined.

[0101] In summary, when the first difference (the difference between the first and second sound pressure levels) is greater than the second difference (the difference between the third and fourth sound pressure levels), since the distance between the left and right speakers of the first device is greater than that of the left and right speakers of the second device, it is determined that the speakers of the first device are suitable for outputting the left and right channels. In this case, the target fusion scheme selects the left and right channels of the first device to output sound, and the center channel of the second device to output sound (i.e., it is determined that the third channel includes the left and right channels, and the fourth channel includes the center channel). When the first difference is less than the second difference, since the distance between the left and right speakers of the first device is less than that of the left and right speakers of the second device, the speakers of the first device are suitable for outputting the center channel, and the speakers of the second device are suitable for outputting the left and right channels (i.e., it is determined that the third channel includes the center channel, and the fourth channel includes the left and right channels).

[0102] In some embodiments of this disclosure, when it is determined that the third channel includes the left channel and the right channel, and the fourth channel includes the center channel, the third channel may also include other channels, and the fourth channel may also include other channels. The specifics can be determined according to the actual situation, and are not limited here.

[0103] In some embodiments of this disclosure, when it is determined that the third channel includes the center channel and the fourth channel includes the left and right channels, the third channel may also include other channels, and the fourth channel may also include other channels. The specifics can be determined according to the actual situation and are not limited here.

[0104] In this embodiment of the disclosure, by setting the relative positions between the first device, the second device, and the remote controller to meet the target conditions, and then measuring the sensitivity and channel distance of the first device and the second device through the microphone of the remote controller, the target fusion scheme can be determined more accurately, and the sound effect of the first device and the second device jointly producing sound can be improved.

[0105] In some embodiments of this disclosure, the controller is further configured to: determine a target difference when the first difference is equal to the second difference, the target difference being the difference between the first sound pressure level and the third sound pressure level, or the difference between the second sound pressure level and the fourth sound pressure level; determine that the third channel includes the center channel and the fourth channel includes the left and right channels when the target difference is greater than or equal to 0; and determine that the third channel includes the left and right channels and the fourth channel includes the center channel when the target difference is less than 0.

[0106] It is understood that the distance between the left and right speakers of the first device is equal to the distance between the left and right speakers of the second device. Therefore, the relationship between the sound pressure level value of the left speaker of the first device and the sound pressure level value of the left speaker of the second device, or the relationship between the sound pressure level value of the right speaker of the first device and the sound pressure level value of the right speaker of the second device, can be used to determine which of the first and second devices outputs the left and right channels and which outputs the center channel.

[0107] The sound pressure level (SPL) is used to characterize the intensity of the sound produced by the vocal tract.

[0108] It can be understood that the target difference is the difference between the first and third sound pressure levels (i.e., SL1-SL3). If SL1-SL3>0, it means that SL1>SL3, indicating that the loudness of the left speaker of the first device is high. If it is used as the center channel, the human voice will be clear. Therefore, it is determined that the left and right speakers of the first device are suitable as the center channel, and the left and right speakers of the second device are suitable as the left and right channels (i.e., the third channel includes the center channel, and the fourth channel includes the left and right channels). Conversely, if SL1-SL3<0, it is determined that the left and right speakers of the first device are suitable as the left and right channels, and the left and right speakers of the second device are suitable as the center channel (i.e., the third channel includes the left and right channels, and the fourth channel includes the center channel).

[0109] In this embodiment of the disclosure, when the first difference is equal to the second difference, determining which of the third and fourth channels includes the left and right channels and which includes the center channel based on the target difference can determine a more accurate target fusion scheme and improve the sound effect of the first device and the second device jointly producing sound.

[0110] In some embodiments of this disclosure, the first channel parameter includes a first bass parameter, which is used to indicate the first bass parameter value corresponding to the subwoofer channel of the first device. The second channel parameter includes a second bass parameter, which is used to indicate the second bass parameter value corresponding to the subwoofer channel of the second device. The controller is further specifically configured to: determine that the third channel includes the subwoofer channel when the first bass parameter value is greater than the second bass parameter value; and determine that the fourth channel includes the subwoofer channel when the first bass parameter value is less than the second bass parameter value.

[0111] The first bass parameter may include a first bass parameter value, or it may include parameters such as the low-frequency extension (Lf) and power (P) used to determine the first bass parameter value. The specific details can be determined based on the actual situation and are not limited here. The description of the second bass parameter can refer to the above description of the first bass parameter, and will not be repeated here.

[0112] For example, if the first bass parameter includes low-frequency extension (Lf) and power (P), the value of the first bass parameter can be determined according to the following formula:

[0113] Lp = a*40 / Lf + (1-a)*P / 200;

[0114] Where 'a' is a weighting coefficient, a < 1; 'Lf' is the low-frequency extension value; 'P' is the power of the bass component; and 'Lp' is the bass parameter value. Table 1 below shows an example of how 'Lp' is calculated.

[0115] Table 1

[0116]

[0117] Since a higher bass parameter value results in better sound quality for the subwoofer channel, the third channel is determined to include the subwoofer channel when the first bass parameter value is greater than the second bass parameter value; conversely, the fourth channel is determined to include the subwoofer channel when the first bass parameter value is less than the second bass parameter value. This allows for a more accurate determination of the subwoofer channel in the target fusion scheme, improving the combined sound quality of the first and second devices.

[0118] It is understandable that if only one of the first and second devices includes a subwoofer channel, then the subwoofer channel included in that device is used as the subwoofer channel of the target fusion scheme. If both the first and second devices include subwoofer channels, the subwoofer channel in the target fusion scheme can be determined as either the subwoofer channel of the first device or the subwoofer channel of the second device based on the relationship between the first and second bass parameter values.

[0119] The controller is further configured to: acquire a first audio signal corresponding to the third channel and a second audio signal corresponding to the fourth channel from the target audio signal, wherein the first audio signal includes audio signals corresponding to each channel in the third channel and the second audio signal includes audio signals corresponding to each channel in the fourth channel.

[0120] Specifically, obtaining the first audio signal corresponding to the third channel and the second audio signal corresponding to the fourth channel from the target audio signal can include directly separating the audio signal corresponding to at least one of the first and second channels from the target audio signal; it can also include downmixing the target audio signal to obtain the audio signal corresponding to at least one of the first and second channels; or it can include upmixing the target audio signal to virtually obtain the audio signal corresponding to at least one of the first and second channels. The specific steps can be determined according to the actual situation and are not limited here.

[0121] In some embodiments of this disclosure, the controller is specifically configured to: render the target audio signal into a multi-channel audio signal matching a preset channel system, the preset channel system including multiple channels, the multi-channel audio signal including audio signals corresponding to each of the multiple channels; when the target fusion scheme includes all the channels of the multiple channels, separate a first audio signal and a second audio signal from the multi-channel audio signal; when the target fusion scheme includes some of the channels of the multiple channels, perform mixing processing on the multi-channel audio signal to obtain the first audio signal and the second audio signal.

[0122] The preset channel system can be a standard channel system, which can be determined based on the device type, usage scenario, and desired immersive sound effect of the first device. For example, it could be a 5.1.2 channel system, a 7.1.4 channel system, etc. The specific system can be determined based on the actual situation and is not limited here.

[0123] It is understandable that if the target audio signal itself is a multi-channel audio signal that matches the preset channel system, then the first audio signal and the second audio signal can be directly obtained from the target audio signal. If the target audio signal itself is not a multi-channel audio signal that matches the preset channel system, then the target audio signal can be rendered into a multi-channel audio signal that matches the preset channel system through mixing processing, virtual surround processing, etc.

[0124] It is understandable that when the first and second channels include the multiple channels, the first audio signal and the second audio signal can be directly separated from the multi-channel audio signal; when the first and second channels include some channels among the multiple channels (i.e., the types of channels included in the first and second channels are less than the types of channels included in the multiple channels), the first audio signal and the second audio signal are obtained by downmixing the multi-channel audio signal.

[0125] In this embodiment, the target audio signal is first rendered into a multi-channel audio signal that matches the preset channel system. Then, based on the channel types of the first and second channels determined by the target fusion scheme and the relationship between them and the channel types of the preset channel system, the first audio signal and the second audio signal are obtained. Based on the first and second audio signals, the first device and the second device are used to jointly emit sound to achieve the sound effect of playing a multi-channel audio signal that matches the preset channel system, thus achieving a panoramic sound effect.

[0126] The controller is further configured to: determine the gain adjustment amount of each channel in the third and fourth channels based on the first sensitivity, the second sensitivity and the preset adjustment strategy; and perform gain adjustment processing on the loudness of the first audio signal and the second audio signal respectively based on the gain adjustment amount of each channel to obtain the third audio signal and the fourth audio signal, so as to match the sensitivity difference between the third and fourth channels.

[0127] Sensitivity is a key indicator in audio systems that measures the response of a channel (such as the left channel, right channel, surround channels, etc.) to an input signal. It is usually expressed as "sound pressure level (SPL, measured in dB)," reflecting the volume (i.e., loudness) that the channel can produce under a specific power input. It not only affects the volume balance of each channel but also directly relates to sound field positioning, detail reproduction, and the overall listening experience.

[0128] In this embodiment of the disclosure, the gain adjustment amount of each channel in the third and fourth channels is determined based on a first sensitivity, a second sensitivity, and a preset adjustment strategy. The preset adjustment strategy includes: First, channels with the same function and symmetrical left and right sides should have consistent sensitivities; for example, the left channel (also called the left front channel) and the right channel (also called the right front channel) should have consistent sensitivities, as should the left surround and right surround channels. Second, for channels with different functions, based on the front (left and right front) channels, the sensitivity of the center channel should be higher than that of the front channels by a first preset value (characterizing the characteristics of the audio system) (e.g., 3dB); the front channels should be higher than the surround channels by a second preset value (e.g., 3dB); and the sky channel, due to its longer path, should be higher than the front channels by a third preset value (e.g., a value between 3 and 5dB).

[0129] For example, the sensitivity of the channels (speakers) of the second device may differ from that of the channels of the first device, requiring matching through an algorithm. The first device pre-stores the sensitivity of each channel, and the second device similarly stores the sensitivity of each channel. This sensitivity matching is not simply matching the sensitivity of the same channels between the first and second devices; for example, it doesn't mean the sensitivity of the left channel of the first device must be identical to that of the left channel of the second device. Instead, after the first and second devices determine the target fusion scheme, the sensitivity of the different functional channels between them is matched to achieve a harmonious value. If the first and second devices in the target fusion scheme have channels of the same type, these channels are considered equivalent to one channel (the sensitivity of the equivalent channel is calculated), and sensitivity adjustments are made between them and other types of channels.

[0130] For example, taking a TV speaker as the first device and an external speaker as the second device, the sensitivity adjustment between the overhead channel of the TV and the left channel of the external speaker in the target fusion scheme is as follows: Calculate the sensitivity difference between the overhead channel of the TV and the left channel of the external speaker: ΔS = S1 - S2, (S1 is the sensitivity of the overhead channel of the TV, and S2 is the sensitivity of the left front channel of the external speaker); convert the sensitivity difference into a gain difference ΔG = 10^{(Gain(ΔS) / 20)}; adjust the amplitude, taking TV Speaker_C'1 (the adjusted gain of the overhead channel of the TV) and External Speaker C”2 (the adjusted gain of the left front channel of the external speaker) as an example, if S1 > S2, then Gain_TV Speaker_C'1_MF = Gain_TV Speaker_C'1_MF = Original - ΔG + ΔG_C; if S1 < S2, then Gain_External Speaker_C”2_MF=Gain_External Speaker_C”2_MFOriginal-ΔG+ΔG_C'. This can be understood as follows: if the TV has high sensitivity, reduce the TV's gain; if the TV has low sensitivity, reduce the audio equipment's gain. In other words, the principle is to reduce the gain of devices with high sensitivity.

[0131] In this algorithm, Gain_TV Speaker_C'1_MFOriginal and Gain_External Speaker_C”2_MFOriginal are the original gain values ​​of the TV and audio equipment, respectively, while ΔG_C and ΔG_C' are compensation values ​​or empirical values. For sensitivity matching between different functional channels, uniform sensitivity is not ideal; a certain degree of difference is more suitable. Therefore, the algorithm reserves ΔG_C and ΔG_C' for adjustment by the setup personnel.

[0132] The controller is further configured to: send target information based on the fourth channel and the fourth audio signal, the target information being used to instruct the second device to play the fourth audio signal through the fourth channel; and play the third audio signal through the third channel so that the third audio signal and the fourth audio signal are played synchronously.

[0133] In this embodiment, the gain adjustment amount of each channel in the third and fourth channels is determined based on the first sensitivity, the second sensitivity, and a preset adjustment strategy. Based on the gain adjustment amount of each channel, the loudness of the first and second audio signals is adjusted to obtain the third and fourth audio signals, so as to match the sensitivity difference between the third and fourth channels. This ensures that the loudness of the audio signals played between different channels in the target fusion scheme is matched with each other, exhibits a certain regularity, and does not have a large difference, thereby ensuring accurate sound positioning, correcting sound field errors, maintaining dynamic range, etc., which can improve the effect of joint sound production.

[0134] In some embodiments of this disclosure, the target information includes a first audio data packet and control information; the controller is specifically configured to: perform low-latency compression processing on the first audio data frame in the fourth audio signal to obtain a first audio data packet, the first audio data packet carrying timestamp information; send the first audio data packet; send the control information; wherein the timestamp information is used by the second device to sort the received audio data packets so that the audio signals played by the first device and the second device are synchronized.

[0135] Among them, low-latency compression processing, namely low-latency audio compression algorithm, uses "audio transmission redundancy coding" to filter out overclocking, empty packets, noise and other data in audio information, which greatly reduces the amount of data transmitted in audio.

[0136] It is understandable that the audio data packets received by the receiving end (second device) are often misaligned due to reasons such as receiving location and interference. Therefore, by extracting the timestamp information from the received audio original data packets, rearranging the audio data frames in the audio data packets according to the timestamp indicated by the timestamp information, and then transmitting them to the sound channel for playback, the audio signals played by the first device and the second device can be synchronized.

[0137] In some embodiments of this disclosure, the first device may send a first audio data packet based on the IIS protocol and send control information based on the I2C protocol.

[0138] In some embodiments of this disclosure, the first device may send first audio data packets and control information based on a broadcast protocol.

[0139] In some embodiments of this disclosure, the controller is specifically configured to: perform low-latency compression processing on a first audio data frame based on a first compression parameter to obtain a first audio data packet, wherein the first compression parameter includes a first sampling rate and a first bit precision; the controller is further configured to: receive first feedback information, wherein the first feedback information is used to indicate the real-time audio detection results and / or network congestion analysis results obtained by the second device; when the first feedback information indicates that the data quality meets a first condition, perform low-latency compression processing on a second audio data frame in a fourth audio signal based on a second compression parameter to obtain a second audio data packet, wherein the second compression parameter includes a second sampling rate and a second bit precision, wherein the second sampling rate is greater than the first sampling rate and the second bit precision is greater than the first bit precision; when the first feedback information indicates that the data quality meets a second condition, perform low-latency compression processing on a second audio data frame in a fourth audio signal based on a third compression parameter to obtain a third audio data packet, wherein the third compression parameter includes a third sampling rate and a third bit precision, wherein the third sampling rate is less than the first sampling rate and the third bit precision is less than the first bit precision.

[0140] Adaptive bitrate control: During audio transmission, the audio bitrate is adjusted in real time according to changes in network bandwidth and latency to ensure optimal transmission performance.

[0141] Among them, the sampling rate and bit precision can determine the transmission bit rate of audio data packets. Therefore, if it is determined that the transmission bit rate needs to be adjusted, the transmission bit rate can be adjusted by adjusting the sampling rate and bit precision.

[0142] Real-time audio detection includes the analysis of digital audio signals, primarily in the time and frequency domains. Key metrics include the relationship between amplitude envelope, zero-crossing rate, and instantaneous energy change, and preset indicators. The reference indicator is that within a certain detection period (e.g., 100 milliseconds), a significant increase or decrease in instantaneous energy is considered a large amplitude (masked by a threshold, assuming the threshold is set at 0.1 times the maximum detection amplitude). Assuming a zero-crossing rate threshold of 0.5, the presence of signals exceeding this threshold indicates a high concentration of high-frequency components, requiring higher transmission bandwidth.

[0143] By analyzing real-time data, the network status for the current period and even the future (adjusted according to the buffer) can be determined and predicted, allowing for real-time adjustments to the transmission strategy. Network congestion analysis is primarily achieved through an algorithm that combines packet transmission and reception latency, packet loss rate, and latency jitter. The packet transmission and reception latency threshold is assumed to be 20 milliseconds, the packet loss rate threshold is assumed to be 0.5%, and the latency jitter rate threshold is assumed to be 5%. Under normal circumstances, the packet loss rate should be below 0.5%, latency jitter should be less than 20 milliseconds, and latency jitter rate should be below 5%. When several parameters show a significant deterioration (i.e., network congestion prediction), a frequency hopping mechanism (a preset frequency hopping algorithm that does not affect audio quality) is first activated. If the frequency hopping mechanism also fails to find a suitable channel for audio transmission, the audio transmission bit rate and audio transmission buffer are dynamically adjusted to ensure smooth audio transmission.

[0144] Based on the above data, the bit rate is dynamically adjusted to compress audio data in frames from 16-bit to 24-bit. When the transmission bit rate is high, the capacity of the audio buffer needs to be appropriately increased to avoid audio dropouts due to insufficient instantaneous bandwidth. The buffer adjustment amount is 10% of the maximum buffer size; for example, if the original audio buffer is 25KB, it will be adjusted to 28KB to accommodate more high-bit audio data.

[0145] The first condition indicates normal signal transmission, including real-time audio detection results indicating normal real-time audio changes and / or network congestion analysis results indicating no network congestion. The second condition indicates abnormal signal transmission, including real-time audio detection results indicating abnormal real-time audio changes and / or network congestion analysis results indicating network congestion.

[0146] In this embodiment of the disclosure, when the audio signal is transmitted normally, a high sampling rate and high bit precision are used to provide high-quality audio data (the higher the sampling rate, the more audio details; the higher the bit precision, the larger the audio data volume and the better the sound quality), and more buffer space is released to provide more computing space for the main controller; when the transmission quality deteriorates, the sampling rate and bit precision in the audio data packet are reduced, and the buffer is increased; after the transmission stabilizes, the sampling rate and bit precision are gradually increased, and the buffer is released to achieve the best transmission quality and the best sound quality output.

[0147] In some embodiments of this disclosure, the audio data packet may also carry packet length, asynchronous sampling rate, etc. Bit precision is directly proportional to packet length; higher bit precision results in a larger packet length.

[0148] In some embodiments of this disclosure, the first channel parameter further includes a first delay, which is the transmission delay of the hardware corresponding to each channel in the first channel of the first device; the second channel parameter further includes a second delay, which is the transmission delay of the hardware corresponding to each channel in the second channel of the second device; the controller is specifically configured to: perform gain adjustment processing on the loudness of the first audio signal and the second audio signal respectively based on the gain adjustment amount of each channel to obtain a fifth audio signal and a sixth audio signal, so as to match the sensitivity difference between the third channel and the fourth channel; and perform delay processing on the fifth audio signal and the sixth audio signal respectively based on the first delay and the second delay to obtain a third audio signal and a fourth audio signal, so as to eliminate the difference in transmission delay between each channel in the third channel and the fourth channel.

[0149] In some embodiments of this disclosure, the controller is specifically configured to: determine the maximum delay between the first delay and the second delay; determine the difference between the maximum delay and each delay in the first and second delays as the delay offset corresponding to each channel; and based on the delay offset corresponding to each channel, add delay offsets to the fifth audio signal and the sixth audio signal respectively to obtain the third audio signal and the fourth audio signal. Thus, by adding corresponding delay offsets to the audio signals corresponding to each channel, the differences in transmission delay between the channels are offset, which to a certain extent ensures that the first device and the second device can synchronously play the third audio signal and the fourth audio signal.

[0150] In some embodiments of this disclosure, if the channel delays between each channel in the first device are the same, then the first delay is a single value; if the channel delays between each channel in the second device are the same, then the second delay is a single value. Specifically, the controller is configured to: when the first delay is greater than or equal to the second delay, add a target delay offset to the sixth audio signal to obtain a fourth audio signal, and the third audio signal becomes the fifth audio signal; when the first delay is less than the second delay, add a target delay offset to the fifth audio signal to obtain a third audio signal, and the fourth audio signal becomes the sixth audio signal; the target delay offset is the absolute value of the difference between the first delay and the second delay. In this way, the difference in transmission delay between each channel in the third and fourth channels can be quickly eliminated, obtaining the third and fourth audio signals.

[0151] In some embodiments of this disclosure, the controller is specifically configured to: play a third audio signal through a third channel based on a third delay, so that the third audio signal and the fourth audio signal are played synchronously, wherein the third delay is used to indicate the difference between the time when the fourth channel receives the fourth audio signal and the time when the third channel receives the third audio signal.

[0152] The third delay includes encoding delay, network transmission delay (related to the network transmission method of the audio signal between the first and second devices), decoding delay, playback buffering delay, etc., which can be calculated by specific algorithms and are not limited here. The third delay is the signal transmission delay between the first and second devices.

[0153] In some embodiments of this disclosure, taking a display device as the first device and an audio device as the second device as an example, if the audio device is a multi-channel integrated unit (such as a soundbar), then the third delay is the difference between the time when each channel in the fourth channel of the audio device receives the audio signal of the corresponding channel and the time when each channel in the third channel of the display device receives the audio signal of the corresponding channel; if the audio device is a combination of an audio host and at least one sub-speaker, then the third delay is the difference between the time when each sub-speaker corresponding to the fourth channel of the audio device receives the audio signal of the corresponding channel and the time when each channel in the third channel of the display device receives the audio signal of the corresponding channel (in this case, the third delay includes the time when the audio signal is transmitted from the display device to the audio host and the time when it is transmitted from the audio host to each sub-speaker).

[0154] In this embodiment of the disclosure, based on the channel delay of the first device and the second device, the fifth audio signal played by the first device and the sixth audio signal played by the second device are delayed to eliminate the difference in transmission delay between each channel in the third and fourth channels. Based on the third delay, the signal transmission delay between the first device and the second device is eliminated, so that the channels of the first device and the second device can play the corresponding audio signals synchronously, which can improve the accuracy of sound field positioning and enhance the sense of space, and improve the effect of joint sound production.

[0155] Regarding channel bandwidth, when an audio device emits sound independently, the allocation of high, mid, and low frequencies only considers the sound field range and number of channels formed by each speaker. After integrating with a display device, the sound field range increases and the number of channels increases. Therefore, the allocation of high, mid, and low frequencies is readjusted based on a larger sound field and more speakers, and redistributed to each speaker to create a larger sound field and more channels in the immersive sound environment. Otherwise, in the new immersive sound environment, users will hear incomplete audio information, and the sound field's spatial orientation and frequency response will be chaotic.

[0156] In some embodiments of this disclosure, the first channel parameter further includes a first bandwidth, which is used to indicate the corresponding sound bandwidth of each channel in the first channel; the second channel parameter further includes a second bandwidth, which is used to indicate the corresponding sound bandwidth of each channel in the second channel; the controller is specifically configured to: perform delay processing on the fifth audio signal and the sixth audio signal based on the first delay and the second delay respectively to obtain the seventh audio signal and the eighth audio signal, so as to eliminate the difference in transmission delay between each channel in the third channel and the fourth channel; and perform gain adjustment processing on the loudness of the seventh audio signal and the eighth audio signal based on the first bandwidth and the second bandwidth respectively to obtain the third audio signal and the fourth audio signal, so as to balance the loudness of the joint sound of the same type of channels in the third channel and the fourth channel.

[0157] Bandwidth refers to the frequency range of sound that a single channel can emit. Gain adjustment processing refers to changing the strength of the sound signal by adjusting the gain parameter of the audio system, thereby controlling the volume / loudness of the audio signal to meet auditory needs or system matching requirements, in order to balance the volume differences of different frequency bands, avoid distortion, and adapt to playback devices or environments.

[0158] like Figure 9 The diagram shows a comparison of the bandwidth of three channels of the same type, with the first device being a display device and the second device being an audio device. The thicker lines represent the bandwidth of the display device, and the thinner lines represent the bandwidth of the audio device. It can be seen that there is a difference in the bandwidth of the same type of audio channel between the audio device and the display device. Therefore, the audio signal is enhanced when the channels of both the audio device and the display device can cover the bandwidth. However, the bandwidth covered by the audio device channel alone and the bandwidth covered by the display device channel alone cannot be enhanced. As a result, the loudness of the played audio signal will be very high in some frequency bands and very low in other frequency bands, leading to problems such as sound imbalance, loss of detail, or auditory fatigue, which seriously affects the combined sound effect.

[0159] Therefore, in this embodiment of the present disclosure, by performing gain adjustment processing on the loudness of the seventh audio signal and the eighth audio signal based on the first bandwidth and the second bandwidth respectively, a third audio signal and a fourth audio signal are obtained, so as to balance the loudness of the combined sound of different frequency bands of the same type of channel in the third channel and the fourth channel, thereby balancing the volume difference of different frequency bands, avoiding distortion, and adapting to playback devices or environments.

[0160] This disclosure provides a second device, including: a communicator; and a controller configured to: receive a target request message, the target request message being used to request channel parameters of the second device, the target request message being generated by a first device in response to an operation of playing a target audio signal, and wherein the first device is connected to the second device; the first device also in response to the operation of playing the target audio signal acquires first channel parameters, the first channel parameters including a first channel supported by the first device and a first sensitivity, the first sensitivity being used to indicate the sensitivity corresponding to each channel in the first channel; in response to the target request message, acquire second channel parameters, the second channel parameters including a second channel supported by the second device and a second sensitivity, the second sensitivity being used to indicate the sensitivity corresponding to each channel in the second channel; and send a target response message corresponding to the target request message, the target response message carrying the second channel parameters, so that the first device, based on the first channel and the second channel, determines a first channel including a channel in the first channel used for playing the target audio signal. A target fusion scheme for the fourth channel in the three-channel and second-channel audio used to play the target audio signal is used. From the target audio signal, a first audio signal corresponding to the third channel and a second audio signal corresponding to the fourth channel are obtained. Based on a first sensitivity, a second sensitivity, and a preset adjustment strategy, the gain adjustment amount for each channel in the third and fourth channels is determined. Based on the gain adjustment amount for each channel, the loudness of the first and second audio signals is adjusted to obtain the third and fourth audio signals to match the sensitivity difference between the third and fourth channels. The first audio signal includes the audio signals corresponding to each channel in the third channel, and the second audio signal includes the audio signals corresponding to each channel in the fourth channel. Target information is received, which is generated based on the fourth channel and the fourth audio signal. In response to the target information, the fourth audio signal is played through the fourth channel so that the third audio signal played by the first device through the third channel is played synchronously with the fourth audio signal.

[0161] In some embodiments of this disclosure, the target information includes a first audio data packet and control information; the controller is specifically configured to: receive the first audio data packet, which is generated by low-latency encoding of a first audio data frame in a fourth audio signal based on a first compression parameter, the first audio data packet carrying timestamp information, the first compression parameter including a first sampling rate and a first bit precision; receive the control information; sort the received first audio data packet based on the timestamp information to synchronize the audio signals played by the first device and the second device; acquire first feedback information, which is used to indicate the real-time audio detection results and / or network congestion analysis results acquired by the second device; and send the first feedback. The information enables the first device to perform low-latency compression processing on the second audio data frame in the fourth audio signal based on the second compression parameters when the first feedback information indicates that the data quality meets the first condition, thereby obtaining a second audio data packet. Then, when the first feedback information indicates that the data quality meets the second condition, the first device performs low-latency compression processing on the second audio data frame in the fourth audio signal based on the third compression parameters, thereby obtaining a third audio data packet. The second compression parameters include a second sampling rate and a second bit precision, where the second sampling rate is greater than the first sampling rate and the second bit precision is greater than the first bit precision. The third compression parameters include a third sampling rate and a third bit precision, where the third sampling rate is less than the first sampling rate and the third bit precision is less than the first bit precision.

[0162] In some embodiments of this disclosure, the controller is further configured to: adjust the receiving buffer from the first buffer to the second buffer, where the second buffer is smaller than the first buffer, when the first feedback information indicates that the data quality meets the first condition; and adjust the receiving buffer from the first buffer to the third buffer, where the third buffer is larger than the first buffer, when the first feedback information indicates that the data quality meets the second condition.

[0163] In this embodiment of the disclosure, when the first feedback information indicates that the data quality meets the first condition, the receiving buffer is adjusted first, and the transmission bit rate is adjusted only when the receiving buffer has been adjusted to the critical value.

[0164] In some embodiments of this disclosure, a detailed description of the second device can be found in the description of the first device in the above embodiments, and will not be repeated here.

[0165] To illustrate this solution in more detail, the following will be described in an exemplary manner. It is understood that the steps involved below may include more or fewer steps in actual implementation, and the order of these steps may also be different, as long as the audio playback method provided in the embodiments of this disclosure can be achieved.

[0166] Figure 10The present invention provides a flowchart of steps for implementing an audio signal playback method according to one or more embodiments of the present disclosure. The audio signal playback method may include the following steps S1001 to S1013.

[0167] S1001. In response to the operation of playing the target audio signal, the first device acquires the first channel parameters when the first device is connected to the second device. The first channel parameters include the first channel and the first sensitivity supported by the first device.

[0168] The first sensitivity is used to indicate the sensitivity of each channel in the first channel.

[0169] S1002, the first device sends a target request message, which is used to request the channel parameters of the second device.

[0170] S1003, The second device receives the target request message.

[0171] The target request message is used to request the channel parameters of the second device. The target request message is generated by the first device in response to the operation of playing the target audio signal and when the first device is connected to the second device. The first device also obtains the first channel parameters in response to the operation of playing the target audio signal. The first channel parameters include the first channel supported by the first device and the first sensitivity. The first sensitivity is used to indicate the sensitivity corresponding to each channel in the first channel.

[0172] S1004. In response to the target request message, the second device obtains the second channel parameters, which include the second channel and the second sensitivity supported by the second device.

[0173] The second sensitivity is used to indicate the sensitivity of each channel in the second channel.

[0174] S1005, The second device sends the target response message corresponding to the target request message.

[0175] The target response message carries second channel parameters, enabling the first device to determine a target fusion scheme based on the first and second channels, including a third channel in the first channel used to play the target audio signal and a fourth channel in the second channel used to play the target audio signal. The device then obtains a first audio signal corresponding to the third channel and a second audio signal corresponding to the fourth channel from the target audio signal. Based on a first sensitivity, a second sensitivity, and a preset adjustment strategy, the device determines the gain adjustment amount for each channel in the third and fourth channels. Based on the gain adjustment amount for each channel, the device performs gain adjustment processing on the loudness of the first and second audio signals respectively to obtain the third and fourth audio signals, thus matching the sensitivity difference between the third and fourth channels. The first audio signal includes the audio signals corresponding to each channel in the third channel, and the second audio signal includes the audio signals corresponding to each channel in the fourth channel.

[0176] S1006, The first device receives a target response message corresponding to the target request message, the target response message carrying the second channel parameter.

[0177] The second channel parameter is obtained by the second device in response to the target request message. The second channel parameter includes the second channel supported by the second device and the second sensitivity. The second sensitivity is used to indicate the sensitivity corresponding to each channel in the second channel.

[0178] S1007. The first device determines a target fusion scheme based on the first channel and the second channel. The target fusion scheme includes a third channel in the first channel for playing the target audio signal and a fourth channel in the second channel for playing the target audio signal.

[0179] S1008. The first device obtains the first audio signal corresponding to the third channel and the second audio signal corresponding to the fourth channel from the target audio signal.

[0180] The first audio signal includes the audio signals corresponding to each channel in the third channel, and the second audio signal includes the audio signals corresponding to each channel in the fourth channel.

[0181] S1009. The first device determines the gain adjustment amount of each channel in the third and fourth channels based on the first sensitivity, the second sensitivity and the preset adjustment strategy; based on the gain adjustment amount of each channel, the loudness of the first audio signal and the second audio signal are respectively subjected to gain adjustment processing to obtain the third audio signal and the fourth audio signal to match the sensitivity difference between the third and fourth channels.

[0182] S1010, the first device sends target information based on the fourth channel and the fourth audio signal.

[0183] The target information is used to instruct the second device to play the fourth audio signal through the fourth channel.

[0184] S1011, The second device receives the target information.

[0185] The target information is generated based on the fourth channel and the fourth audio signal.

[0186] S1012, the second device responds to the target information and plays the fourth audio signal through the fourth channel.

[0187] S1013. The first device plays the third audio signal through the third channel so that the third audio signal and the fourth audio signal are played synchronously.

[0188] In some embodiments of this disclosure, the above-mentioned S1008 can be specifically implemented by the following S1008a to S1008c.

[0189] S1008a, The first device renders the target audio signal into a multi-channel audio signal that matches the preset channel system.

[0190] The preset channel system includes multiple channels, and the multi-channel audio signal includes the audio signal corresponding to each of the multiple channels.

[0191] S1008b: When the target fusion scheme includes all the channels in the multiple channels, the first device separates the first audio signal and the second audio signal from the multi-channel audio signal.

[0192] S1008c, When the target fusion scheme includes some of the multiple channels, the first device performs mixing processing on the multi-channel audio signal to obtain a first audio signal and a second audio signal.

[0193] In some embodiments of this disclosure, the above-mentioned S1007 can be specifically implemented by the following S1007a.

[0194] S1007a. The first device determines the first channel as the third channel and the second channel as the fourth channel to determine the target fusion scheme.

[0195] In some embodiments of this disclosure, the above-mentioned S1007 can be specifically implemented by the following S1007b.

[0196] S1007b: If the first device has channels of the same type in the first channel and the second channel, it sets one of the two channels of the same type as the target channel or performs a mute operation to determine the target fusion scheme.

[0197] Among them, there are no channels of the same type in the third and fourth channels, and the target channel is a channel that does not exist in the first and second channels.

[0198] In some embodiments of this disclosure, the target information includes a first audio data packet and control information; S1010 can be specifically implemented by S1010a to S1010c, and S1011 can be specifically implemented by S1011a to S1011c.

[0199] S1010a, The first device performs low-latency compression processing on the first audio data frame in the fourth audio signal to obtain a first audio data packet, the first audio data packet carrying timestamp information.

[0200] The timestamp information is used by the second device to sort the received audio data packets so that the audio signals played by the first and second devices are synchronized.

[0201] S1010b, The first device sends the first audio data packet.

[0202] S1011a, The second device receives the first audio data packet.

[0203] S1010c, The first device sends the control information.

[0204] S1011b, The second device receives the control information.

[0205] S1011c, the second device sorts the received first audio data packet based on the timestamp information so that the audio signals played by the first device and the second device are synchronized.

[0206] In some embodiments of this disclosure, S1010a can be specifically implemented by S1010a1. After S1011a, the audio signal playback method provided by the embodiments of this disclosure may also include S1014 to S1018.

[0207] S1010a1, The first device performs low-latency compression processing on the first audio data frame based on the first compression parameters to obtain the first audio data packet. The first compression parameters include the first sampling rate and the first bit precision.

[0208] S1014. The second device acquires first feedback information, which is used to indicate the real-time audio detection results and / or network congestion analysis results acquired by the second device.

[0209] S1015, The second device sends the first feedback information.

[0210] S1016, The first device receives the first feedback information.

[0211] S1017. When the first feedback information indicates that the data quality meets the first condition, the first device performs low-latency compression processing on the second audio data frame in the fourth audio signal based on the second compression parameters to obtain the second audio data packet.

[0212] The second compression parameter includes a second sampling rate and a second bit precision, where the second sampling rate is greater than the first sampling rate and the second bit precision is greater than the first bit precision.

[0213] S1018. When the first feedback information indicates that the data quality meets the second condition, the first device performs low-latency compression processing on the second audio data frame in the fourth audio signal based on the third compression parameters to obtain the third audio data packet.

[0214] The third compression parameter includes the third sampling rate and the third bit precision. The third sampling rate is less than the first sampling rate, and the third bit precision is less than the first bit precision.

[0215] In some embodiments of this disclosure, after S1014 described above, the audio signal playback method provided in the embodiments of this disclosure may further include S1019 to S1020 as described below.

[0216] S1019. If the first feedback information indicates that the data quality meets the first condition, the second device adjusts the receiving buffer from the first buffer to the second buffer, where the second buffer is smaller than the first buffer.

[0217] S1020, if the first feedback information indicates that the data quality meets the second condition, the second device will adjust the receiving buffer from the first buffer to the third buffer, and the third buffer is larger than the first buffer.

[0218] In this embodiment of the disclosure, S1016 to S1018 and S1019 to S1020 can be executed in two ways: only S1016 to S1018 can be executed, only S1019 to S1020 can be executed, or both S1016 to S1018 and S1019 to S1020 can be executed. The specific execution method can be determined according to the actual situation, and is not limited here. When both S1016 to S1018 and S1019 to S1020 are executed, S1019 to S1020 can be executed first, followed by S1016 to S1018, or they can be executed simultaneously. The specific execution method can be determined according to the actual situation, and is not limited here.

[0219] When the audio equipment is a combination of a main speaker and at least one sub-speaker, the main speaker also needs to perform low-latency compression processing on the audio signal during audio signal transmission between the main speaker and the sub-speaker. The main speaker also needs to determine whether the data instruction indicated by the feedback information returned by the sub-speaker meets the first condition or the second condition, and adjust the compression parameters of the low-latency compression processing according to the determination result. For details, please refer to the audio signal transmission process between the first device and the second device described above, which will not be repeated here.

[0220] In some embodiments of this disclosure, the above-mentioned S1009 can be specifically implemented by the following S1009a to S1009c.

[0221] S1009a, The first device determines the maximum delay between the first delay and the second delay.

[0222] S1009b, The first device determines the difference between the maximum delay and each delay in the first delay and the second delay as the delay offset corresponding to each channel.

[0223] S1009c, the first device adds delay offsets to the first audio signal and the second audio signal respectively based on the delay offsets corresponding to each channel to obtain the third audio signal and the fourth audio signal.

[0224] In some embodiments of this disclosure, the first channel parameter further includes a first delay, which is the transmission delay of the hardware corresponding to each channel in the first channel of the first device; the second channel parameter further includes a second delay, which is the transmission delay of the hardware corresponding to each channel in the second channel of the second device; the above S1009 can be specifically implemented by the following S1009d to S1009f, and the above S1013 can be specifically implemented by the following S1013a.

[0225] S1009d, the first device determines the gain adjustment amount of each channel in the third and fourth channels based on the first sensitivity, the second sensitivity and the preset adjustment strategy.

[0226] S1009e, the first device performs gain adjustment processing on the loudness of the first audio signal and the second audio signal based on the gain adjustment amount of each channel, to obtain the fifth audio signal and the sixth audio signal, so as to match the sensitivity difference between the third channel and the fourth channel.

[0227] S1009f: The first device performs delay processing on the fifth audio signal and the sixth audio signal based on the first delay and the second delay to obtain the third audio signal and the fourth audio signal, so as to eliminate the difference in transmission delay between the channels in the third channel and the fourth channel.

[0228] S1013a, the first device plays a third audio signal through the third channel based on a third delay, so that the third audio signal and the fourth audio signal are played synchronously. The third delay is used to indicate the difference between the time when the fourth channel receives the fourth audio signal and the time when the third channel receives the third audio signal.

[0229] In some embodiments of this disclosure, the first channel parameter further includes a first bandwidth, which is used to indicate the sound bandwidth corresponding to each channel in the first channel; the second channel parameter further includes a second bandwidth, which is used to indicate the sound bandwidth corresponding to each channel in the second channel; the above S1009f can be specifically implemented by the following S1009f1 to S1009f2.

[0230] S1009f, the first device performs delay processing on the fifth audio signal and the sixth audio signal based on the first delay and the second delay to obtain the seventh audio signal and the eighth audio signal, so as to eliminate the difference in transmission delay between the channels in the third channel and the fourth channel.

[0231] S1009f2, the first device performs gain adjustment processing on the loudness of the seventh audio signal and the eighth audio signal based on the first bandwidth and the second bandwidth, respectively, to obtain the third audio signal and the fourth audio signal, so as to balance the loudness of the same type of channel in the third channel and the fourth channel.

[0232] In this embodiment, the order in which the first and second audio signals are adjusted for bandwidth, sensitivity, and delay is not limited. For example, the first and second audio signals may be adjusted for bandwidth first, then further adjusted for sensitivity, and then further adjusted for delay; or the first and second audio signals may be adjusted for sensitivity first, then further adjusted for bandwidth, and then further adjusted for delay; or the first and second audio signals may be adjusted for delay first, then further adjusted for bandwidth, and then further adjusted for sensitivity; or other orders may be followed, which can be determined according to the actual usage and are not limited here.

[0233] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this disclosure, and are not intended to limit them. Although this disclosure has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this disclosure.

[0234] For ease of explanation, the above description has been provided in conjunction with specific embodiments. However, the above exemplary discussion is not intended to be exhaustive or to limit the embodiments to the specific forms disclosed above. Various modifications and variations can be obtained based on the above teachings. The selection and description of the above embodiments are for the purpose of better explaining the principles and practical applications, thereby enabling those skilled in the art to better utilize the described embodiments and various different variations of embodiments suitable for specific use considerations.

Claims

1. A first device, characterized in that, include: communicator; The controller is configured to: in response to an operation of playing a target audio signal, when the first device is connected to the second device, acquire first channel parameters, the first channel parameters including a first channel supported by the first device and a first sensitivity, the first sensitivity being used to indicate the sensitivity corresponding to each channel in the first channel; the first channel parameters also include a first bandwidth, the first bandwidth being used to indicate the sound bandwidth corresponding to each channel in the first channel. Send a target request message, the target request message being used to request the channel parameters of the second device; The device receives a target response message corresponding to the target request message. The target response message carries a second channel parameter, which is obtained by the second device in response to the target request message. The second channel parameter includes a second channel supported by the second device and a second sensitivity. The second sensitivity is used to indicate the sensitivity corresponding to each channel in the second channel. The second channel parameter also includes a second bandwidth, which is used to indicate the sound bandwidth corresponding to each channel in the second channel. Based on the first channel and the second channel, a target fusion scheme is determined, the target fusion scheme including a third channel in the first channel for playing the target audio signal and a fourth channel in the second channel for playing the target audio signal; From the target audio signal, a first audio signal corresponding to the third channel and a second audio signal corresponding to the fourth channel are obtained. The first audio signal includes audio signals corresponding to each channel in the third channel, and the second audio signal includes audio signals corresponding to each channel in the fourth channel. Based on the first sensitivity, the second sensitivity, and the preset adjustment strategy, the gain adjustment amount of each channel in the third and fourth channels is determined. Based on the gain adjustment amount of each channel, the loudness of the first audio signal and the second audio signal are respectively subjected to gain adjustment processing to obtain the seventh audio signal and the eighth audio signal, so as to match the sensitivity difference between the third channel and the fourth channel. Based on the first bandwidth and the second bandwidth, the loudness of the seventh audio signal and the eighth audio signal are respectively subjected to gain adjustment processing to obtain the third audio signal and the fourth audio signal, so as to balance the loudness of the same type of channels in the third channel and the fourth channel. Based on the fourth channel and the fourth audio signal, target information is sent, which instructs the second device to play the fourth audio signal through the fourth channel. The third audio signal is played through the third channel so that the third audio signal and the fourth audio signal are played synchronously.

2. The first device according to claim 1, characterized in that, The controller is specifically configured as follows: The target audio signal is rendered into a multi-channel audio signal that matches a preset channel system, wherein the preset channel system includes multiple channels and the multi-channel audio signal includes the audio signal corresponding to each of the multiple channels. When the target fusion scheme includes all of the multiple channels, the first audio signal and the second audio signal are separated from the multi-channel audio signal; When the target fusion scheme includes some of the multiple channels, the multi-channel audio signal is mixed to obtain the first audio signal and the second audio signal.

3. The first device according to claim 1, characterized in that, The controller is specifically configured as follows: The first channel is determined as the third channel, and the second channel is determined as the fourth channel, in order to determine the target fusion scheme; or, If there are channels of the same type in the first and second channels, one of the two channels of the same type is set as the target channel or muted to determine the target fusion scheme. If there are no channels of the same type in the third and fourth channels, the target channel is a channel that does not exist in either the first or second channel.

4. The first device according to any one of claims 1 to 3, characterized in that, The target information includes a first audio data packet and control information; the controller is specifically configured as follows: The first audio data frame in the fourth audio signal is subjected to low-latency compression processing to obtain a first audio data packet, which carries timestamp information. Send the first audio data packet; Send the control information, which instructs the second device to play the fourth audio signal through the fourth channel; The timestamp information is used by the second device to sort the received audio data packets so that the audio signals played by the first device and the second device are synchronized.

5. The first device according to claim 4, characterized in that, The controller is specifically configured as follows: Based on the first compression parameters, the first audio data frame is subjected to low-latency compression processing to obtain the first audio data packet. The first compression parameters include the first sampling rate and the first bit precision. The controller is also configured to: Receive first feedback information, which is used to indicate the real-time audio detection results and / or network congestion analysis results obtained by the second device; When the first feedback information indicates that the data quality meets the first condition, the second audio data frame in the fourth audio signal is subjected to low-latency compression processing based on the second compression parameters to obtain the second audio data packet. The second compression parameters include a second sampling rate and a second bit precision, wherein the second sampling rate is greater than the first sampling rate and the second bit precision is greater than the first bit precision. When the first feedback information indicates that the data quality meets the second condition, the second audio data frame in the fourth audio signal is subjected to low-latency compression processing based on the third compression parameters to obtain the third audio data packet. The third compression parameters include a third sampling rate and a third bit precision. The third sampling rate is less than the first sampling rate, and the third bit precision is less than the first bit precision.

6. A second device, characterized in that, include: communicator; The controller is configured to: receive a target request message, the target request message being used to request channel parameters of the second device, the target request message being generated by the first device in response to an operation of playing a target audio signal, and wherein the first device is connected to the second device; the first device also acquires first channel parameters in response to the operation of playing the target audio signal, the first channel parameters including a first channel supported by the first device and a first sensitivity, the first sensitivity being used to indicate the sensitivity corresponding to each channel in the first channel; the first channel parameters also include a first bandwidth, the first bandwidth being used to indicate the sound bandwidth corresponding to each channel in the first channel; In response to the target request message, a second channel parameter is obtained. The second channel parameter includes a second channel supported by the second device and a second sensitivity. The second sensitivity is used to indicate the sensitivity corresponding to each channel in the second channel. The second channel parameter also includes a second bandwidth, which is used to indicate the sound bandwidth corresponding to each channel in the second channel. Send a target response message corresponding to the target request message, the target response message carrying the second channel parameter, so that the first device, based on the first channel and the second channel, determines a target fusion scheme including a third channel in the first channel for playing the target audio signal and a fourth channel in the second channel for playing the target audio signal, and obtains a first audio signal corresponding to the third channel and a second audio signal corresponding to the fourth channel from the target audio signal, and determines the gain adjustment amount of each channel in the third channel and the fourth channel based on the first sensitivity, the second sensitivity and a preset adjustment strategy, based on the parameters of each channel. Gain adjustment is performed on the loudness of the first audio signal and the second audio signal respectively to obtain a seventh audio signal and an eighth audio signal, in order to match the sensitivity difference between the third channel and the fourth channel; based on the first bandwidth and the second bandwidth, the loudness of the seventh audio signal and the eighth audio signal are respectively adjusted to obtain a third audio signal and a fourth audio signal, in order to balance the loudness of the same type of channels jointly emitting sound in the third channel and the fourth channel; the first audio signal includes the audio signals corresponding to each channel in the third channel, and the second audio signal includes the audio signals corresponding to each channel in the fourth channel; Receive target information, which is generated based on the fourth channel and the fourth audio signal; In response to the target information, the fourth audio signal is played through the fourth channel so that the third audio signal played by the first device through the third channel is played synchronously with the fourth audio signal.

7. The second device according to claim 6, characterized in that, The target information includes a first audio data packet and control information; the controller is specifically configured as follows: The first audio data packet is received. The first audio data packet is generated by low-latency encoding of the first audio data frame in the fourth audio signal based on the first compression parameters. The first audio data packet carries timestamp information. The first compression parameters include a first sampling rate and a first bit precision. Receive the control information; The received first audio data packets are sorted based on the timestamp information to synchronize the audio signals played by the first device and the second device. Obtain first feedback information, which is used to indicate the real-time audio detection results and / or network congestion analysis results obtained by the second device; The first feedback information is sent so that, when the first feedback information indicates that the data quality meets the first condition, the first device performs low-latency compression processing on the second audio data frame in the fourth audio signal based on the second compression parameters to obtain a second audio data packet. When the first feedback information indicates that the data quality meets the second condition, the first device performs low-latency compression processing on the second audio data frame in the fourth audio signal based on the third compression parameters to obtain a third audio data packet. The second compression parameters include a second sampling rate and a second bit precision, wherein the second sampling rate is greater than the first sampling rate and the second bit precision is greater than the first bit precision. The third compression parameters include a third sampling rate and a third bit precision, wherein the third sampling rate is less than the first sampling rate and the third bit precision is less than the first bit precision.

8. The second device according to claim 7, characterized in that, The controller is also configured to: If the first feedback information indicates that the data quality meets the first condition, the receiving buffer will be adjusted from the first buffer to the second buffer, where the second buffer is smaller than the first buffer. If the first feedback information indicates that the data quality meets the second condition, the receiving buffer is adjusted from the first buffer to the third buffer, and the third buffer is larger than the first buffer.

9. A method for playing audio signals, characterized in that, Applied to the first device, including: In response to the operation of playing the target audio signal, when the first device is connected to the second device, the first channel parameters are obtained. The first channel parameters include the first channel supported by the first device and the first sensitivity. The first sensitivity is used to indicate the sensitivity corresponding to each channel in the first channel. The first channel parameters also include the first bandwidth, which is used to indicate the sound bandwidth corresponding to each channel in the first channel. Send a target request message, the target request message being used to request the channel parameters of the second device; The device receives a target response message corresponding to the target request message. The target response message carries a second channel parameter, which is obtained by the second device in response to the target request message. The second channel parameter includes a second channel supported by the second device and a second sensitivity. The second sensitivity is used to indicate the sensitivity corresponding to each channel in the second channel. The second channel parameter also includes a second bandwidth, which is used to indicate the sound bandwidth corresponding to each channel in the second channel. Based on the first channel and the second channel, a target fusion scheme is determined, the target fusion scheme including a third channel in the first channel for playing the target audio signal and a fourth channel in the second channel for playing the target audio signal; From the target audio signal, a first audio signal corresponding to the third channel and a second audio signal corresponding to the fourth channel are obtained. The first audio signal includes audio signals corresponding to each channel in the third channel, and the second audio signal includes audio signals corresponding to each channel in the fourth channel. Based on the first sensitivity, the second sensitivity, and the preset adjustment strategy, the gain adjustment amount of each channel in the third and fourth channels is determined. Based on the gain adjustment amount of each channel, the loudness of the first audio signal and the second audio signal are respectively subjected to gain adjustment processing to obtain the seventh audio signal and the eighth audio signal, so as to match the sensitivity difference between the third channel and the fourth channel. Based on the first bandwidth and the second bandwidth, the loudness of the seventh audio signal and the eighth audio signal are respectively subjected to gain adjustment processing to obtain the third audio signal and the fourth audio signal, so as to balance the loudness of the same type of channels in the third channel and the fourth channel. Based on the fourth channel and the fourth audio signal, target information is sent, which instructs the second device to play the fourth audio signal through the fourth channel. The third audio signal is played through the third channel so that the third audio signal and the fourth audio signal are played synchronously.

10. A method for playing audio signals, characterized in that, Applied to a second device, including: The system receives a target request message, which requests the channel parameters of the second device. This target request message is generated by the first device in response to an operation to play a target audio signal, and the first device is connected to the second device. The first device also acquires first channel parameters in response to the operation to play the target audio signal. The first channel parameters include a first channel supported by the first device and a first sensitivity, whereby the first sensitivity indicates the sensitivity corresponding to each channel in the first channel. The first channel parameters also include a first bandwidth, which indicates the sound bandwidth corresponding to each channel in the first channel. In response to the target request message, a second channel parameter is obtained. The second channel parameter includes a second channel supported by the second device and a second sensitivity. The second sensitivity is used to indicate the sensitivity corresponding to each channel in the second channel. The second channel parameter also includes a second bandwidth, which is used to indicate the sound bandwidth corresponding to each channel in the second channel. Send a target response message corresponding to the target request message, the target response message carrying the second channel parameter, so that the first device, based on the first channel and the second channel, determines a target fusion scheme including a third channel in the first channel for playing the target audio signal and a fourth channel in the second channel for playing the target audio signal, and obtains a first audio signal corresponding to the third channel and a second audio signal corresponding to the fourth channel from the target audio signal, and determines the gain adjustment amount of each channel in the third channel and the fourth channel based on the first sensitivity, the second sensitivity and a preset adjustment strategy, based on the parameters of each channel. Gain adjustment is performed on the loudness of the first audio signal and the second audio signal respectively to obtain a seventh audio signal and an eighth audio signal, in order to match the sensitivity difference between the third channel and the fourth channel; based on the first bandwidth and the second bandwidth, the loudness of the seventh audio signal and the eighth audio signal are respectively adjusted to obtain a third audio signal and a fourth audio signal, in order to balance the loudness of the same type of channels jointly emitting sound in the third channel and the fourth channel; the first audio signal includes the audio signals corresponding to each channel in the third channel, and the second audio signal includes the audio signals corresponding to each channel in the fourth channel; Receive target information, which is generated based on the fourth channel and the fourth audio signal; In response to the target information, the fourth audio signal is played through the fourth channel so that the third audio signal played by the first device through the third channel is played synchronously with the fourth audio signal.