Electronic equipment and audio signal playing method
By obtaining channel parameters and delay information, determining the target fusion plan and performing delay processing, the problem of synchronous playback caused by different channel delays of different devices is solved, the accuracy of sound field positioning and spatial sense are improved, and the joint sound effect is enhanced.
Patent Information
- Application Number
- CN202510969685.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-11
- Publication Date
- 2025-10-17
- Estimated Expiration
- 2045-07-11
AI Technical Summary
When the display device and the audio equipment are not specific matching equipment or are from different manufacturers, the different channel delays make it impossible to synchronize the sound, resulting in confusion in the sound field positioning and distortion of the spatial sense, affecting the joint sound effect.
By obtaining channel parameters and delay information between the first device and the second device, determining a target fusion solution, and performing delay processing, the channels of the first device and the second device play audio signals synchronously, eliminating the transmission delay difference.
It improves the accuracy of sound field positioning and spatial sense, enhances the joint sound effect, and ensures synchronous playback of channels between different devices.
Smart Images

Figure CN120812484A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] Embodiments of the present disclosure relate to terminal technology. More particularly, to an electronic device and an audio signal playing method. BACKGROUND
[0002] At present, the schemes of the display device and the sound equipment jointly sounding are that the display device plays the corresponding channel signal in the audio signal according to the channel included by itself, and the sound equipment also plays the corresponding channel signal in the audio signal according to the channel included by itself. In this way, in the case that the display device and the sound equipment are not specific matching devices, or in the case that the display device and the sound equipment are not devices of the same manufacturer, since the channel delay of each channel of the display device is different from the channel delay of each channel of the sound equipment, the display device and the sound equipment cannot synchronously sound, thereby causing problems such as sound field positioning confusion, spatial distortion and the like, and seriously affecting the effect of jointly sounding. SUMMARY
[0003] In order to solve the above technical problems or at least partially solve the above technical problems, embodiments of the present disclosure provide an electronic device and an audio signal playing method, which can improve the effect of jointly sounding.
[0004] In a first aspect, the embodiments of the present disclosure provide a first device, comprising: a communicator; a controller configured to: in response to an operation of playing a target audio signal, acquire a first channel parameter in a case that the first device is connected to a second device, the first channel parameter comprising a first channel supported by the first device and a first delay, the first delay being used to indicate a transmission delay corresponding to each channel in the first channel; send a target request message, the target request message being used to request a channel parameter of the second device; receive a target response message corresponding to the target request message, the target response message carrying a second channel parameter, the second channel parameter being acquired by the second device in response to the target request message, the second channel parameter comprising a second channel supported by the second device and a second delay, the second delay being used to indicate a transmission delay corresponding to each channel in the second channel; determine a target fusion scheme based on the first channel and the second channel, the target fusion scheme comprising a third channel in the first channel for playing the target audio signal and a fourth channel in the second channel for playing the target audio signal; acquire a first audio signal corresponding to the third channel and a second audio signal corresponding to the fourth channel from the target audio signal, the first audio signal comprising an audio signal corresponding to each channel in the third channel, and the second audio signal comprising an audio signal corresponding to each channel in the fourth channel; perform delay processing on the first audio signal and the second audio signal based on the first delay and the second delay respectively to obtain a third audio signal and a fourth audio signal, so as to eliminate a difference in transmission delay between each channel in the third channel and the fourth channel; send target information based on the fourth channel and the fourth audio signal, the target information being used to instruct the second device to play the fourth audio signal through the fourth channel; and play the third audio signal through the third channel based on a third delay, so as to synchronize the third audio signal and the fourth audio signal, the third delay being used to indicate a difference between a time when the fourth channel receives the fourth audio signal and a time when the third channel receives the third audio signal.
[0005] In a second aspect, the embodiments of the present disclosure provide a second device, comprising: a communicator; a controller configured to: receive a target request message, the target request message being used to request a channel parameter of the second device, the target request message being generated by a first device in a case that the first device connects the second device and in response to an operation of playing a target audio signal, the first device further obtaining a first channel parameter in response to the operation of playing the target audio signal, the first channel parameter comprising a first channel supported by the first device and a first delay, the first delay being used to indicate a transmission delay corresponding to each channel in the first channel; obtain a second channel parameter in response to the target request message, the second channel parameter comprising a second channel supported by the second device and a second delay, the second delay being used to indicate a transmission delay corresponding to each channel in the second channel; send a target response message corresponding to the target request message, the target response message carrying the second channel parameter, so that the first device determines a target fusion scheme comprising a third channel in the first channel used to play the target audio signal and a fourth channel in the second channel used to play the target audio signal based on the first channel and the second channel, and obtains a first audio signal corresponding to the third channel and a second audio signal corresponding to the fourth channel from the target audio signal, and performs delay processing on the first audio signal and the second audio signal based on the first delay and the second delay to obtain a third audio signal and a fourth audio signal, so as to eliminate the difference between the transmission delays of each channel in the third channel and the fourth channel; the first audio signal comprises audio signals corresponding to each channel in the third channel, and the second audio signal comprises audio signals corresponding to each channel in the fourth channel; receive target information, the target information being generated based on the fourth channel and the fourth audio signal; play the fourth audio signal through the fourth channel in response to the target information, so that the first device plays the third audio signal and the fourth audio signal synchronously based on a third delay through the third channel, the third delay being used to indicate a difference between a time when the fourth channel receives the fourth audio signal and a time when the third channel receives the third audio signal.
[0006] In a third aspect, the embodiments of the present disclosure provide an audio signal playing method applied to a first device, comprising: in response to an operation of playing a target audio signal, obtaining a first channel parameter in a case that the first device is connected to a second device, the first channel parameter comprising a first channel supported by the first device and a first delay, the first delay being used to indicate a transmission delay corresponding to each channel in the first channel; sending a target request message, the target request message being used to request a channel parameter of the second device; receiving a target response message corresponding to the target request message, the target response message carrying a second channel parameter, the second channel parameter being obtained by the second device in response to the target request message, the second channel parameter comprising a second channel supported by the second device and a second delay, the second delay being used to indicate a transmission delay corresponding to each channel in the second channel; determining a target fusion scheme based on the first channel and the second channel, the target fusion scheme comprising a third channel in the first channel used to play the target audio signal and a fourth channel in the second channel used to play the target audio signal; obtaining a first audio signal corresponding to the third channel and a second audio signal corresponding to the fourth channel from the target audio signal, the first audio signal comprising an audio signal corresponding to each channel in the third channel, and the second audio signal comprising an audio signal corresponding to each channel in the fourth channel; performing delay processing on the first audio signal and the second audio signal based on the first delay and the second delay respectively to obtain a third audio signal and a fourth audio signal, so as to eliminate the difference between the transmission delays of each channel in the third channel and the fourth channel; sending target information based on the fourth channel and the fourth audio signal, the target information being used to instruct the second device to play the fourth audio signal through the fourth channel; and playing the third audio signal through the third channel based on a third delay, so that the third audio signal and the fourth audio signal are played synchronously, the third delay being used to indicate the difference between the time when the fourth channel receives the fourth audio signal and the time when the third channel receives the third audio signal.
[0007] In a fourth aspect, the embodiments of the present disclosure provide an audio signal playing method applied to a second device, including: receiving a target request message, the target request message being used to request a channel parameter of the second device, the target request message being generated by a first device in response to an operation of playing a target audio signal and in a case that the first device is connected to the second device, the first device further acquiring a first channel parameter in response to the operation of playing the target audio signal, the first channel parameter including a first channel supported by the first device and a first delay, the first delay being used to indicate a transmission delay corresponding to each channel in the first channel; acquiring a second channel parameter in response to the target request message, the second channel parameter including a second channel supported by the second device and a second delay, the second delay being used to indicate a transmission delay corresponding to each channel in the second channel; sending a target response message corresponding to the target request message, the target response message carrying the second channel parameter, so that the first device determines a target fusion scheme including a third channel in the first channel used to play the target audio signal and a fourth channel in the second channel used to play the target audio signal based on the first channel and the second channel, and acquires a first audio signal corresponding to the third channel and a second audio signal corresponding to the fourth channel from the target audio signal, and performs delay processing on the first audio signal and the second audio signal respectively based on the first delay and the second delay to obtain a third audio signal and a fourth audio signal, so as to eliminate the difference between the transmission delays of each channel in the third channel and the fourth channel; the first audio signal includes audio signals corresponding to each channel in the third channel, and the second audio signal includes audio signals corresponding to each channel in the fourth channel; receiving target information, the target information being generated based on the fourth channel and the fourth audio signal; playing the fourth audio signal through the fourth channel in response to the target information, so that the first device plays the third audio signal and the fourth audio signal synchronously based on a third delay through the third channel, the third delay being used to indicate a difference between a time when the fourth channel receives the fourth audio signal and a time when the third channel receives the third audio signal.
[0008] In a fifth aspect, the embodiments of the present disclosure provide a computer readable storage medium, including: a computer program stored on the computer readable storage medium, the computer program being executed by a processor to implement the audio signal playing method shown in the third aspect or the fourth aspect.
[0009] In a sixth aspect, the embodiments of the present disclosure provide a computer program product, including: when the computer program product runs on a computer, causing the computer to implement the audio signal playing method shown in the second aspect.
[0010] Compared with the prior art, the technical scheme provided by the embodiments of the present disclosure has the following advantages: in the embodiments of the present disclosure, in the case that the first device is connected to the second device, the first channel parameter is obtained in response to the operation of playing the target audio signal, the first channel parameter includes the first channel supported by the first device and the first delay, the first delay is used to indicate the transmission delay corresponding to each channel in the first channel; a target request message is sent, the target request message is used to request the channel parameter of the second device; a target response message corresponding to the target request message is received, the target response message carries the second channel parameter, the second channel parameter is obtained by the second device in response to the target request message, the second channel parameter includes the second channel supported by the second device and the second delay, the second delay is used to indicate the transmission delay corresponding to each channel in the second channel; based on the first channel and the second channel, a target fusion scheme is determined, the target fusion scheme includes a third channel in the first channel used to play the target audio signal and a fourth channel in the second channel used to play the target audio signal; from the target audio signal, a first audio signal corresponding to the third channel and a second audio signal corresponding to the fourth channel are obtained, the first audio signal includes the audio signal corresponding to each channel in the third channel, and the second audio signal includes the audio signal corresponding to each channel in the fourth channel; based on the first delay and the second delay, the first audio signal and the second audio signal are respectively processed by delay to obtain a third audio signal and a fourth audio signal, so as to eliminate the difference in transmission delay between each channel in the third channel and the fourth channel; based on the fourth channel and the fourth audio signal, a target information is sent, the target information is used to instruct the second device to play the fourth audio signal through the fourth channel; based on the third delay, the third audio signal is played through the third channel, so that the third audio signal and the fourth audio signal are played synchronously, and the third delay is used to indicate the difference between the time when the fourth channel receives the fourth audio signal and the time when the third channel receives the third audio signal. In this way, based on the channel delay of the first device and the second device, the first audio signal played by the first device and the second audio signal played by the second device are processed by delay, so as to eliminate the difference in transmission delay between each channel in the third channel and the fourth channel, so that the channels of the first device and the second device can play the corresponding audio signals synchronously, the sound field positioning accuracy and the sense of space can be improved, and the effect of joint sound production can be improved. BRIEF DESCRIPTION OF DRAWINGS
[0011] In order to more clearly illustrate the embodiments of the present disclosure or the implementation manners in the related art, the drawings needed to be used in the embodiments or related art description will be briefly introduced as follows. Obviously, the drawings in the following description are some embodiments of the present disclosure, and other drawings can also be obtained by those skilled in the art based on these drawings.
[0012] Figure 1An operating scenario between a display device and a control device according to some embodiments is shown;
[0013] Figure 2 A hardware configuration block diagram of the control device 100 according to some embodiments is shown;
[0014] Figure 3 A hardware configuration block diagram of the display device 200 according to some embodiments is shown;
[0015] Figure 4 One of the schematic diagrams of the sound combination scheme of the display device and the sound device according to some embodiments is shown;
[0016] Figure 5 One of the schematic diagrams of the sound combination scheme of the display device and the sound device according to some embodiments is shown;
[0017] Figure 6 One of the schematic diagrams of the sound combination scheme of the display device and the sound device according to some embodiments is shown;
[0018] Figure 7 One of the schematic diagrams of the sound combination scheme of the display device and the sound device according to some embodiments is shown;
[0019] Figure 8 A schematic diagram of the relative positions between the display device, the sound device and the remote controller satisfying the target condition according to some embodiments is shown;
[0020] Figure 9 A schematic diagram of the frequency width of the same type of sound channel of the sound device and the display device according to some embodiments is shown;
[0021] Figure 10 A flowchart of the audio signal playing method according to some embodiments is shown. DETAILED DESCRIPTION
[0022] For the purpose of making the objects and embodiments of the present disclosure clearer, the following will clearly and completely describe the exemplary embodiments of the present disclosure with reference to the accompanying drawings in the exemplary embodiments of the present disclosure. Obviously, the described exemplary embodiments are only a part of the embodiments of the present disclosure, but not all the embodiments.
[0023] It should be noted that the brief description of the terms in the present disclosure is only for the convenience of understanding the following described embodiments, and is not intended to limit the embodiments of the present disclosure. Unless otherwise specified, these terms should be understood according to their ordinary and general meanings.
[0024] The terms "first", "second", "third", and the like in the description and in the claims of the present disclosure and above-described drawings are used for distinguishing between similar or identical objects and entities, and do not necessarily mean a specified order or sequence, unless otherwise specified. It will be understood that the terms so used are interchangeable under appropriate circumstances.
[0025] The terms "comprise", "comprising", "have", "having", "include", "including", "contain", "containing", and any variations thereof, are intended to cover a non-exclusive inclusion, such that a product or device that comprises a list of components does not include only those components but can include other components not expressly listed or inherent to such product or device.
[0026] The display device provided by the embodiments of the present disclosure can have various implementation forms, for example, can be a television, a smart television, a laser projection device, a monitor, an electronic bulletin board, an electronic table, a mobile phone, a tablet computer, a notebook computer, a palm computer, a vehicle-mounted electronic device, etc.
[0027] Figure 1 The operation scenarios between the display device and the control device according to the embodiments are illustrated, wherein the control device includes a smart device or a control apparatus. As shown in Figure 1 The user can operate the display device 200 through the smart device 300 or the control apparatus 100.
[0028] In some embodiments, the control apparatus 100 can be a remote controller, and the communication between the remote controller and the display device includes infrared protocol communication or Bluetooth protocol communication, and other short-distance communication manners, to control the display device 200 through wireless or wired manner. The user can input user instructions through the keys on the remote controller, voice input, control panel input, etc., to control the display device 200.
[0029] In some embodiments, the smart device 300 (such as a mobile terminal, a tablet computer, a computer, a notebook computer, etc.) can also be used to control the display device 200. For example, the display device 200 is controlled by using an application program running on the smart device.
[0030] In some embodiments, the display device can not receive instructions by using the above-described smart device or control apparatus, but receives the control of the user through touch or gesture, etc.
[0031] In some embodiments, the display device 200 can also be controlled in ways other than the control device 100 and the smart device 300, for example, the display device 200 can directly receive the user's voice instruction control through the voice instruction receiving module configured inside the display device 200, or the user's voice instruction control can be received through the voice control device set outside the display device 200.
[0032] In some embodiments, the display device 200 also communicates data with the server 400. The display device 200 can be allowed to be connected in communication through a local area network (LAN), a wireless local area network (WLAN), and other networks. The server 400 can provide various content and interaction to the display device 200. The server 400 can be one cluster or multiple clusters, and can include one or more types of servers.
[0033] Figure 2 An exemplary configuration block diagram of the control device 100 according to an exemplary embodiment is shown. As shown, the control device 100 includes a controller 110, a communication interface 130, a user input / output interface 140, an external storage, a power supply. The control device 100 can receive the user's input operation instruction, and convert the operation instruction into an instruction that the display device 200 can recognize and respond to, and play a mediating role between the user and the display device 200. Figure 2
[0034] As shown, the display device 200 includes at least one of a tuner and demodulator 210, a communicator 220, a detector 230, an external device interface 240, a controller 250, a display 260, an audio output interface 270, a user interface 280, an external storage, a power supply. Figure 3 In some embodiments, the controller includes a processor, a video processor, an audio processor, a graphics processor, a RAM, a ROM, a first interface to an n-th interface for input / output.
[0035] The display 260 includes a display screen component for presenting a picture, and a driving component for driving image display, a component for receiving an image signal originating from the controller output, and displaying video content, image content, and a menu control interface, and a user control UI interface.
[0036] The display 260 can be a liquid crystal display, an OLED display, and a projection display, and can also be a projection device and a projection screen.
[0037]
[0038] The communicator 220 is a component for communicating with external devices or servers according to various communication protocol types. For example, the communicator can include at least one of a Wifi module, a Bluetooth module, a wired Ethernet module, other network communication protocol chips or near field communication protocol chips, and an infrared receiver. The display device 200 can establish transmission and reception of control signals and data signals with the external control apparatus 100 or the server 400 through the communicator 220.
[0039] The user interface 280 can be used to receive control signals of the control apparatus 100 (e.g., an infrared remote controller, etc.). It can also be used to directly receive input operation instructions of a user and convert the operation instructions into instructions that the display device 200 can recognize and respond to. In this case, it can be referred to as a user input interface.
[0040] The detector 230 is used to collect signals of an external environment or interaction with the outside. For example, the detector 230 includes a light receiver for collecting ambient light intensity, or an image collector such as a camera for collecting an external environment scene, user attributes, or user interaction gestures, or a sound collector such as a microphone for receiving external sounds.
[0041] The external device interface 240 can include, but is not limited to, any one or more of the following: a high-definition multimedia interface (HDMI), an analog or digital high-definition component input interface (component), a composite video input interface (CVBS), a USB input interface (USB), an RGB port, etc. It can also be a composite input / output interface formed by a plurality of the above interfaces.
[0042] The tuner demodulator 210 receives broadcast television signals through wired or wireless reception and demodulates audio / video signals and EPG data signals from a plurality of wireless or wired broadcast television signals.
[0043] In some embodiments, the controller 250 and the tuner demodulator 210 can be located in different split devices, i.e., the tuner demodulator 210 can also be in an external device of the main body device where the controller 250 is located, such as an external set-top box, etc.
[0044] The controller 250 controls the operation of the display device and responds to user operations through various software control programs stored on the storage (internal storage or external storage). The controller 250 controls the overall operation of the display device 200. For example, in response to receiving a user command for selecting a UI object displayed on the display 260, the controller 250 can perform an operation related to the object selected by the user command.
[0045] In some embodiments, the controller includes at least one of a central processing unit (CPU), a video processor, an audio processor, a graphics processing unit (GPU), and a random access memory (RAM), a read-only memory (ROM), a first interface to an n-th interface for input / output, a communication bus, and the like.
[0046] The RAM is also called main memory, and is an internal memory that directly exchanges data with the controller. It can be read and written at any time (except during refresh), and is fast, and is usually used as a temporary data storage medium for an operating system or other programs that are running. The biggest difference between it and ROM is the volatility of data, that is, once the power is off, the stored data will be lost. RAM is used in computers and digital systems to temporarily store programs, data, and intermediate results. ROM works in a non-destructive readout mode and can only read out information. Once the information is written, it is fixed and will not be lost even if the power is cut off, so it is also called fixed memory.
[0047] The user can input a user command through a graphical user interface (GUI) displayed on the display 260, and the user input interface receives the user input command through the graphical user interface (GUI). Alternatively, the user can input a user command by inputting a specific sound or gesture, and the user input interface receives the user input command by recognizing the sound or gesture through a sensor.
[0048] The "user interface" is a medium interface for interaction and information exchange between an application program or an operating system and a user, which realizes the conversion between the internal form of information and the form that the user can accept. The commonly used form of user interface is a graphical user interface (GUI), which refers to a user interface related to computer operation displayed in a graphical manner. It can be an icon, window, control, and other interface elements displayed on the display screen of a display device, and the control can include visible interface elements such as icons, buttons, menus, tabs, text boxes, dialog boxes, status bars, navigation bars, and widgets.
[0049] The embodiments of the present disclosure provide an electronic device and an audio signal playing method, wherein in one case, the first device is a display device, and the second device is a sound equipment, and in another case, the first device is a sound equipment, and the second device is a display device, which can be determined according to actual conditions, and is not limited herein.
[0050] Taking a TV as an example, currently, more and more TV manufacturers support joint sound of a TV and a soundbar. Due to differences in hardware functions of different TV platforms, for example, some TV models are 2.0 channels, some are 2.1 channels, some are 2.1.2 channels, some are 4.1 channels, some are 4.1.2 channels, some are 5.1.2 channels, some models have a separate sky sound speaker, the sky sound channel is output from the TV end, some models do not have a sky sound speaker, and the sky sound is output from the soundbar end; the soundbar end also supports 2.1 channels, 3.1 channels, 4.1.2 channels, 5.1.2 channels, etc. The current strategy is to match a specific model of a TV with a specific model of a soundbar for joint sound, however, as the number of TV models and soundbar models supporting joint sound increases, we cannot strongly bind the user's product type. Therefore, in the case that the display device and the soundbar are not specific matching devices, or in the case that the display device and the soundbar are not devices of the same manufacturer, because the channel delay of each channel of the display device is different from the channel delay of each channel of the soundbar, the display device and the soundbar cannot be synchronized to sound, which causes problems such as sound field positioning confusion, spatial distortion, and seriously affects the joint sound effect.
[0051] Especially, in the case that the display device and the soundbar include the same channels, the two same types of channels of the display device and the soundbar need to play the same audio signal synchronously, but because the channel delay of the two same types of channels is different, the playing time of the same audio signal is different, which causes problems such as sound field positioning confusion, spatial distortion, and seriously affects the joint sound effect.
[0052] In some embodiments of the present disclosure, the TV can be a traditional TV, a laser TV, etc.; the soundbar can include an all-in-one soundbar, a soundbar with an external subwoofer, a soundbar with different numbers of satellite boxes, etc.
[0053] The embodiments of the present disclosure provide a first device, including a communicator; a controller configured to implement the audio signal playing method provided by the embodiments of the present disclosure.
[0054] The controller is configured to: in response to an operation of playing a target audio signal, in a case that the first device is connected to a second device, obtain a first channel parameter, the first channel parameter including a first channel supported by the first device and a first delay, the first delay being used to indicate a transmission delay corresponding to each channel in the first channel.
[0055] The target audio signal is a multi-channel audio signal. The multi-channel audio signal refers to an audio signal containing multiple independent channels, each of which can carry different audio information, thereby providing a richer, more spatial and immersive audio experience for the listener.
[0056] The joint sounding function, also referred to as the chorus function, refers to a function in which two or more electronic devices jointly play an audio signal through their respective channel systems.
[0057] The controller is further configured to send a target request message, the target request message being used to request a channel parameter of the second device; receive a target response message corresponding to the target request message, the target response message carrying a second channel parameter, the second channel parameter being obtained by the second device in response to the target request message, and the second channel parameter including a second channel supported by the second device and a second delay, the second delay being used to indicate a transmission delay corresponding to each channel in the second channel.
[0058] It can be understood that, in response to the operation of playing the target audio signal, the first device sends a target request message to the second device, receives a target response message replied by the second device based on the target request message, in a case where it is detected that there is a communication connection with the second device and it is detected that the first device supports the chorus function.
[0059] In some embodiments of the present disclosure, the second device can be defaulted to support the chorus function, or the second device can be requested to obtain whether it supports the chorus function (referred to as the chorus capability of the second device) through a request message. The specific determination can be made according to actual conditions, which is not limited here.
[0060] In some embodiments of the present disclosure, the chorus capability of the second device and the channel parameter of the second device can be requested through one request message, or the chorus capability of the second device and the channel parameter of the second device can be requested through two request messages respectively. The specific determination can be made according to actual conditions, which is not limited here.
[0061] In some embodiments of the present disclosure, the target request message includes a first request message and a second request message, the target response message includes a first response message and a second response message, and the process of obtaining the sound field capability of the second device and the channel parameter of the second device specifically includes: in response to the operation of playing the target audio signal, in a case where the first device supports the sound field function, sending the first request message to the second device, the first request message being used to request the sound field capability of the second device; receiving the first response message corresponding to the first request message from the second device, the first response message carrying the sound field capability information of the second device; in a case where the sound field capability information indicates that the second device supports the sound field function, sending the second request message to the second device, the second request message being used to request the channel parameter of the second device; and receiving the second response message corresponding to the second request message from the second device, the second response message carrying the second channel parameter.
[0062] It should be noted that the target request message can further include other request messages, and the target response message can further include other response messages, which are not limited herein.
[0063] The controller is further configured to determine a target fusion scheme based on the first channel and the second channel, the target fusion scheme including a third channel in the first channel for playing the target audio signal and a fourth channel in the second channel for playing the target audio signal.
[0064] In some embodiments of the present disclosure, on the basis of extending the HDMI CEC communication protocol between the first device and the second device, the sound field capability of the second device and the channel parameter of the second device are obtained through information interaction between the first device and the second device, and then the target fusion scheme is determined, so that the first device and the second device can jointly sound through their respective channels, which can avoid the problem that the first device and the second device cannot sound synchronously due to different sound channel delays of different electronic devices, thereby avoiding the problems of sound field positioning confusion and spatial distortion.
[0065] In the embodiments of the present disclosure, the information interaction between the first device and the second device can be realized by extending the wired protocol or the wireless protocol between the first device and the second device.
[0066] In some embodiments of the present disclosure, the following functions can be realized by extending the part that can be customized by the manufacturer in the HDMI CEC communication protocol: 1. adding the information interaction function of whether to support the sound synthesis function, the second device should respond after receiving the first request message sent by the first device for requesting the sound synthesis capability of the second device; after the identity of the first device and the second device is confirmed in the CEC handshake process, the first device can request the sound synthesis capability of the second device, and the second device responds correctly according to its own capability; 2. adding the information interaction function of requesting the channel parameter of the second device, after the sound and the first device communicate the support of the sound synthesis function, the first device requests the channel parameter of the second device.
[0067] In some embodiments of the present disclosure, the information interaction instruction that meets the HDMI CEC communication protocol needs to be designed, such as setting the format of the first request message, setting the format of the first response message, setting the format of the second request message, setting the format of the second response message, or setting the format of the target request message, and setting the format of the target response message.
[0068] In some embodiments of the present disclosure, the second device can also request the channel parameter of the first device, which is not limited here.
[0069] Exemplarily, the first device and the second device perform CEC handshake to complete identity verification (which can be performed by matching the vendor identification code or the vendor id, or by other ways, which is not limited here), in the case that the first device supports the sound synthesis function, the sound synthesis capability of the second device is requested (the first request message), the second device reports the sound synthesis capability after receiving the instruction (the first response message), the first device analyzes whether the second device supports the sound synthesis function, if the sound synthesis function is supported, the channel parameter of the second device is further requested (the second request message), and the second device responds according to its own sound synthesis capability. The first device can also actively report the channel parameter of itself to the second device, and the second device stores the channel parameter of the first device after receiving it, for subsequent use.
[0070] The second channel parameter is used to indicate the second channel supported by the second device, which can include the device type of the second device (such as 2.1 channel, 5.1 channel, etc.), or the type of the channel supported by the second device, which can be determined according to the actual situation, which is not limited here.
[0071] The second sound channel parameter can include a sound channel parameter used to determine the target fusion scheme, such as a second sound channel, and can further include detailed parameter information of a sound channel supported by the second device (for example, low frequency diving value (Lf) and power (P) of a subwoofer sound channel, and the like), or relative position information of a sound channel (corresponding speaker) in the second device. The second sound channel parameter can include a sound channel parameter used to adjust the joint sound effect, such as sound channel delay, sensitivity, frequency width, and the like of each sound channel in the second sound channel. The second sound channel parameter can be determined according to actual use requirements, and is not limited herein.
[0072] The description of the first sound channel parameter can refer to the related description of the second sound channel parameter, which is not repeated here.
[0073] In some embodiments of the present disclosure, the target fusion scheme can include a third sound channel and a fourth sound channel, and the third sound channel and the fourth sound channel can or can not include the same type of sound channel, which is not limited herein.
[0074] In some embodiments of the present disclosure, the controller is specifically configured to determine the first sound channel as the third sound channel and the second sound channel as the fourth sound channel to determine the target fusion scheme. In this way, the target fusion scheme can be quickly determined, and if the first sound channel and the second sound channel include the same type of sound channel, the third sound channel and the fourth sound channel included in the target fusion scheme include the same type of sound channel.
[0075] In some embodiments of the present disclosure, a first coefficient set corresponding to mapping of the first sound channel to the preset sound channel system is obtained, each element of the first coefficient set is 1 or 0, and an element of the first coefficient set being 1 represents that a sound channel included in the preset sound channel system and the first sound channel. A second coefficient set corresponding to mapping of the second sound channel to the preset sound channel system is obtained, each element of the second coefficient set is 1 or 0, and an element of the second coefficient set being 1 represents that a sound channel included in the preset sound channel system and the second sound channel. The target fusion scheme is determined based on the first coefficient set and the second coefficient set. In this way, the target fusion scheme can be determined, and when a sound channel fails or is disconnected, a changed fusion scheme can be determined.
[0076] Exemplarily, the first device is a TV speaker, the second device is an external speaker, a new sound channel system, such as the fusion system 7-1-4, is defined, and the TV speaker and the external speaker are mapped to and managed by the system. Based on the original Dolby architecture, parameters in the fusion system are defined. The 7-1-4 system is defined, including: C-FL (left channel), C-FR (right channel), C-Center (center channel), C-SR (right surround channel), C-SL (left surround channel), C-BR (right back surround channel), C-BL (left back surround channel), C-SUB (subwoofer channel), C-FTR (front top right channel), C-FTL (front top left channel), C-RTR (rear top right channel), and C-RTL (rear top left channel), where C (Combine) represents the meaning of fusion and combination. If the external speaker is a 5.1.2 system (FL, FR, Center, SR, SL, SUB, FTR, FTL) according to the Dolby architecture, it is mapped to the fusion system as follows: C-FL = C1*FL, C-FR = C2*FR, C-Center = C3*Center, C-SR = C4*SR, C-SL = C5*SL, C-RSR = C6*RR, C-RSL = C7*RL, C-SUB = C8*SUB, C-FTR = C9*FTR, C-FTL = C10*FTL, C-RTR = C11*RTR, and C-RTL = C12*RTL. Wherein, C1-C12 are 0 or 1, such as the 5.1.2 system above, the coefficients can be: C1-C12 = (1, 1, 1, 1, 1, 0, 0, 1, 1, 1, 0, 0). Before system fusion, the TV speaker system parameters TV Speaker (C1-C12) (i.e., the first coefficient set) and the external speaker system parameters External Speaker (C1-C12) (i.e., the second coefficient set) are defined, and the target fusion scheme is determined according to the TV Speaker (C1-C12) and the External Speaker (C1-C12).
[0077] In some embodiments of the present disclosure, the controller is specifically configured to: in the case that there are sound channels of the same type in the first sound channel and the second sound channel, setting one of the two sound channels of the same type as a target sound channel or performing mute processing to determine the target fusion scheme, and there is no sound channel of the same type in the third sound channel and the fourth sound channel, and the target sound channel is a sound channel that does not exist in the first sound channel and the second sound channel. In this way, the first device and the second device can jointly sound through different types of sound channels, thereby avoiding the problems of sound channel conflict and poor sounding effect caused by the first device and the second device playing audio signals through the same type of sound channel.
[0078] In some embodiments of the present disclosure, the first device can determine a target fusion scheme matched from at least one preset fusion scheme based on the first sound channel and the second sound channel; the first device can also determine the target fusion scheme in combination with the first sound channel and the second sound channel, and the relative position information of the first device and the second device, etc. The target fusion scheme can be determined according to actual conditions, which is not limited here.
[0079] The at least one preset fusion scheme can include a preset fusion scheme corresponding to the first device and a second device of each device type, such as a preset fusion scheme corresponding to the first device and a second device of 2.1 sound channel, a preset fusion scheme corresponding to the first device and a second device of 5.1 sound channel, etc., which is not limited here.
[0080] In some embodiments of the present disclosure, the at least one preset fusion scheme stored in the first device can be different due to different device types (sound channel system types) of the first device. The at least one preset fusion scheme stored in the first device can be the same regardless of whether the device types of the first device are the same, and then the target fusion scheme is determined from the at least one preset fusion scheme according to the number and type of sound channels included in the first sound channel and the second sound channel. The actual conditions can be determined, which is not limited here.
[0081] In some embodiments of the present disclosure, a plurality of candidate fusion schemes matched with the first sound channel can be determined from the at least one preset fusion scheme, and then the candidate fusion scheme most matched with the second sound channel is selected from the plurality of candidate fusion schemes to determine the target fusion scheme. The user can be presented with a plurality of candidate fusion schemes, and the user can select a candidate fusion scheme that meets the demand as the target fusion scheme. The actual conditions can be determined, which is not limited here.
[0082] In some embodiments of the present disclosure, in the case that the first device is a display device, the controller is further configured to present first prompt information after determining the target fusion scheme, the first prompt information is used to prompt the user about the relative position relationship between the display device and the sound device, and / or the first prompt information is used to prompt the user about the position of the third sound channel in the display device and the position of the fourth sound channel in the sound device.
[0083] The first prompt information can be text prompt information, voice prompt information, or picture prompt information, which is not limited here.
[0084] In some embodiments of the present disclosure, the first device is a display device. After determining the target fusion scheme, the first device can present the target fusion scheme through picture prompt information, which shows the third sound channel of the first device, the fourth sound channel of the second device, the relative position relationship between the third sound channel and the fourth sound channel, the relative position relationship between the first device and the second device, and the like.
[0085] In Example 1, the first device is a Laser TV, and the second device is a 4.1.2 sound system. Due to the hardware size limitation of the Laser TV, the left and right sound channels are very close to each other. Therefore, the left and right sound channels of the Laser TV are suitable for being used as the center sound channel. The satellite boxes of the 4.1.2 sound system can be increased in width, surround, height, and sound field. Therefore, when the Laser TV is connected to the 4.1.2 sound system, the Laser TV is switched to the following mode by default, as shown in FIG. 1. Figure 4 As shown in FIG. 1, the Laser home theater mode: the third sound channel of the Laser TV includes the center sound channel, and the fourth sound channel of the sound system includes the left and right sound channels, the subwoofer sound channel, the left surround sound channel, the right surround sound channel, the left sky sound channel, and the right sky sound channel. The mark "41" is used to indicate the components of the TV, and the mark "42" is used to indicate the components of the sound system. After selecting this mode, the Laser TV notifies the sound system to play the audio signals of the corresponding sound channels through the fourth sound channel through interaction with the sound system. The Laser TV plays the audio signals of the corresponding sound channels through the third sound channel. For example, the left and right sound channels of the Laser TV are used as the center sound channel (the third sound channel) to play the audio signals of the center sound channel. The fourth sound channel includes the left and right sound channels, the subwoofer sound channel, the left surround sound channel, the right surround sound channel, the left sky sound channel, and the right sky sound channel, which are used to play the audio signals of the corresponding sound channels in the target audio signals. The center sound channel of the sound system is muted or the data is not processed.
[0086] In Example 2, the first device is a TV supporting sky sound, and the sound system is a 5.1 sound system. In a home environment, the TV is placed in a higher position. Therefore, the height of the TV can be used to provide the sky sound channel, and the sky sound speaker of the TV is generally located near the top of the screen to provide the sky sound channel. The following mode 1 (as shown in FIG. 2) can be provided. Figure 5 As shown in FIG. 2, the mark "51" is used to indicate the components of the TV, and the mark "52" is used to indicate the components of the sound system. In the mode 1, the third sound channel of the TV includes the sky sound channel, and the fourth sound channel of the sound system includes the left and right sound channels, the center sound channel, and the left and right surround sound channels. In the mode 2 (as shown in FIG. 3), the third sound channel of the TV includes the left and right sound channels, the center sound channel, and the sky sound channel, and the fourth sound channel of the sound system includes the left and right surround sound channels. Figure 6
[0087] Example 3, the first device is a TV of 2.1.2, and the sound device is a wireless 2.1 sound device. The wireless 2.1 sound device can be easily placed by using the position of the wireless 2.1 sound channel. The two sound channels of the sound device are used as left and right surround sound channels, and the TV and the sound device form a 4.1.2 sound channel together. The picture prompt information is as shown in Figure 7 , wherein the mark "71" is used to indicate the component of the TV, and the mark "72" is used to indicate the component of the sound device.
[0088] The above-mentioned preset fusion scheme is a target fusion scheme determined based on the first sound channel, the second sound channel, and at least one preset fusion scheme. Although the problem of sound channel conflict is solved, the user experience is only improved to a certain extent. The relative position relationship of the first device and the second device is different, and different sound fusion schemes may need to be adopted to achieve the best sound effect. For example, in the above-mentioned example 2, is mode 1 used or mode 2 used, or is another mode used? In fact, the only criterion for judgment is the sound effect after joint sound emission, and the two relatively important factors affecting the sound effect are sensitivity and sound channel distance difference.
[0089] In some embodiments of the present disclosure, when the first device is a display device, the controller is specifically configured to: display second prompt information, the second prompt information being used to prompt a user whether a relative position among the first device, the second device and the remote controller satisfies a target condition, the target condition being used to indicate that a first median plane of a line connecting a center of a left channel speaker and a center of a right channel speaker of the first device is parallel to a second median plane of a line connecting a center of a left channel speaker and a center of a right channel speaker of the second device, a central axis of the remote controller is perpendicular to the first median plane, a minimum distance from the central axis to a first plane passing through the first device and being perpendicular to the first median plane is less than or equal to a first distance, a minimum distance from the central axis to a second plane passing through the second device and being perpendicular to the second median plane is less than or equal to a second distance, and a distance from a center of the remote controller to a center of the first device is greater than or equal to a third distance and less than or equal to a fourth distance, the third distance being less than the fourth distance; in response to a received confirmation operation for the second prompt information, control the left channel speaker and the right channel speaker of the first device to sequentially play audio signals, and sequentially receive, by a microphone of the remote controller, a first sound signal emitted by the left channel speaker of the first device and a second sound signal emitted by the right channel speaker of the first device; send a control message, the control message being used to control the left channel speaker and the right channel speaker of the second device to sequentially play audio signals; sequentially receive, by the microphone of the remote controller, a third sound signal emitted by the left channel speaker of the second device and a fourth sound signal emitted by the right channel speaker of the second device; based on the first sound signal, the second sound signal, the third sound signal and the fourth sound signal, determine a first sound pressure level and a second sound pressure level corresponding to the left channel speaker and the right channel speaker of the first device respectively, and a third sound pressure level and a fourth sound pressure level corresponding to the left channel speaker and the right channel speaker of the second device respectively; in a case where a first difference value is greater than a second difference value, determine that the third channel includes the left channel and the right channel, and the fourth channel includes a center channel, the first difference value being a difference value between the first sound pressure level and the second sound pressure level, and the second difference value being a difference value between the third sound pressure level and the fourth sound pressure level; in a case where the first difference value is less than the second difference value, determine that the third channel includes the center channel, and the fourth channel includes the left channel and the right channel.
[0090] In some embodiments of the present disclosure, the first median plane and the second median plane can coincide, and a minimum distance between the first median plane and the second median plane can be less than a distance threshold value, the distance threshold value being determined according to actual conditions, which is not limited herein.
[0091] In some embodiments of the present disclosure, the first plane is a plane passing through the first device and being perpendicular to the first median plane, and the second plane is a plane passing through the second device and being perpendicular to the first median plane. The first distance can be greater than or equal to the second distance, or can be less than the second distance, which is not limited herein.
[0092] Among them, the first distance, the second distance, the third distance and the fourth distance can be determined according to actual usage and are not limited here.
[0093] To improve measurement accuracy, the optimal distance between the remote control microphone and the first device is between 0.4 meters and 1 meter. Therefore, the optimal distance range between the center of the remote control and the center of the first device can be determined based on the optimal distance range between the remote control microphone and the first device.
[0094] It can be understood that in the embodiment of the present disclosure, the sensitivity and channel distance are measured by the remote control of the first device that supports a microphone, and when the measurement is started, the user is prompted to set the relative positions between the first device, the second device, and the remote control to meet the target conditions. Usually, the first device and the second device are normally placed in the center, and the left and right channels are symmetrical. There are three situations in total when comparing the distance between the left and right channel speakers of the first device to the distance between the left and right channel speakers of the second device: greater than, equal to, and less than. Figure 8 As shown, when the relative positions between the first device, the second device, and the remote control meet the target conditions, the measurement program sequentially controls the left channel of the first device, the right channel of the first device, the left channel of the second device, and the right channel of the second device to play audio signals. The microphone device records the collected sound signals of each channel respectively and outputs the collected sound signals to the measurement program. The measurement program analyzes the sound signals respectively to obtain the sound pressure level values of the four channels (i.e., the first sound pressure level and the second sound pressure level corresponding to the left channel speaker and the right channel speaker of the first device, respectively, and the third sound pressure level and the fourth sound pressure level corresponding to the left channel speaker and the right channel speaker of the second device, respectively), which are recorded as SL1, SL2, SL3, and SL4 respectively. The distances between the microphone and the left channel speaker and the right channel speaker of the first device, and the left channel speaker and the right channel speaker of the second device are L1, L2, L3, and L4 respectively. According to the relationship between the sound pressure level and distance of a point sound source in a free sound field, the following relationship can be obtained:
[0095] SL1-SL2=20lg(L2 / L1)=20lg((L1+D1) / L1)=20lg(1+D1 / L1),
[0096] SL3-SL4=20lg(L4 / L3)=20lg((L3+D2) / L3)=20lg(1+D2 / L3).
[0097] When starting the measurement procedure, since the placement position of the microphone of the remote controller has been determined in advance, that is, L1+D1 / 2 is a constant, and L3+D2 / 2 is a constant, so the greater D1 is, the smaller L1 is, and the smaller D1 is, the greater L1 is; similarly, D2 and L3 follow the same rule. Therefore, according to the formula, it can be concluded that the value of SL1-SL2 depends entirely on D1, and the value of SL3-SL4 depends entirely on D2; by comparing the values of (SL1-SL2) and (SL3-SL4), the sizes of D1 and D2 can be obtained.
[0098] In summary, in the case where the first difference (the difference between the first sound pressure level and the second sound pressure level) is greater than the second difference (the difference between the third sound pressure level and the fourth sound pressure level), since the distance between the left and right speakers of the first device is greater than the distance between the left and right speakers of the second device, it is determined that the speakers of the first device are suitable for outputting left and right channels, at this time, the target fusion scheme selects the first device to output left and right channels, and the second device to output a center channel (that is, it is determined that the third channel includes a left channel and a right channel, and the fourth channel includes a center channel); in the case where the first difference is less than the second difference, since the distance between the left and right speakers of the first device is less than the distance between the left and right speakers of the second device, the speakers of the first device are suitable for outputting a center channel, and the speakers of the second device are suitable for outputting left and right channels (that is, it is determined that the third channel includes a center channel, and the fourth channel includes a left channel and a right channel).
[0099] In some embodiments of the present disclosure, in the case where it is determined that the third channel includes a left channel and a right channel, and the fourth channel includes a center channel, the third channel can also include other channels, and the fourth channel can also include other channels, which can be determined according to actual conditions, and are not limited here.
[0100] In some embodiments of the present disclosure, in the case where it is determined that the third channel includes a center channel, and the fourth channel includes a left channel and a right channel, the third channel can also include other channels, and the fourth channel can also include other channels, which can be determined according to actual conditions, and are not limited here.
[0101] In the embodiments of the present disclosure, by setting the relative positions between the first device, the second device and the remote controller to satisfy the target condition, and then measuring the sensitivity and channel distance of the first device and the second device through the microphone of the remote controller, the target fusion scheme can be more accurately determined, and the sound effect of the joint sound of the first device and the second device can be improved.
[0102] In some embodiments of the present disclosure, the controller is further configured to: in a case where the first difference value is equal to the second difference value, determine a target difference value, the target difference value being a difference value between the first sound pressure level and the third sound pressure level, or a difference value between the second sound pressure level and the fourth sound pressure level; in a case where the target difference value is greater than or equal to 0, determine that the third sound channel includes a center sound channel, and the fourth sound channel includes a left sound channel and a right sound channel; in a case where the target difference value is less than 0, determine that the third sound channel includes the left sound channel and the right sound channel, and the fourth sound channel includes the center sound channel.
[0103] It can be understood that the distance between the left and right speakers of the first device is equal to the distance between the left and right speakers of the second device, and then it can be determined which of the first device and the second device outputs the left and right sound channels and which outputs the center sound channel according to the size relationship between the sound pressure level value corresponding to the left speaker of the first device and the sound pressure level value of the left speaker of the second device, or the size relationship between the sound pressure level value corresponding to the right speaker of the first device and the sound pressure level value of the right speaker of the second device.
[0104] The sound pressure level value is used to represent the strength of the sound generated by the sound channel.
[0105] It can be understood that, when the target difference value is the difference value between the first sound pressure level and the third sound pressure level (i.e., SL1-SL3), and SL1-SL3>0, it indicates that SL1>SL3, which means that the loudness of the left speaker of the first device is high, and if the left and right speakers of the first device are used as the center sound channel, the human voice will be clear, and therefore it is determined that the left and right speakers of the first device are suitable for the center sound channel, and the left and right speakers of the second device are suitable for the left and right sound channels (i.e., the third sound channel includes the center sound channel, and the fourth sound channel includes the left sound channel and the right sound channel). Conversely, if SL1-SL3<0, it is determined that the left and right speakers of the first device are suitable for the left and right sound channels, and the left and right speakers of the second device are suitable for the center sound channel (i.e., the third sound channel includes the left sound channel and the right sound channel, and the fourth sound channel includes the center sound channel).
[0106] In the embodiments of the present disclosure, in a case where the first difference value is equal to the second difference value, it is determined which of the third sound channel and the fourth sound channel includes the left and right sound channels and which includes the center sound channel according to the target difference value, a more accurate target fusion scheme can be determined, and the sound effect of the joint sound of the first device and the second device can be improved.
[0107] In some embodiments of the present disclosure, the first sound channel parameter includes a first bass parameter, the first bass parameter being used to indicate a first bass parameter value corresponding to a bass cannon sound channel of the first device, the second sound channel parameter includes a second bass parameter, the second bass parameter being used to indicate a second bass parameter value corresponding to a bass cannon sound channel of the second device, and the controller is further specifically configured to: in a case where the first bass parameter value is greater than the second bass parameter value, determine that the third sound channel includes the bass cannon sound channel; in a case where the first bass parameter value is less than the second bass parameter value, determine that the fourth sound channel includes the bass cannon sound channel.
[0108] The first bass parameter can include a first bass parameter value, and can also include low frequency diving value (Lf) and power (P) and other parameter information used to determine the first bass parameter value, which can be determined according to actual conditions and is not limited here.
[0109] The description of the second bass parameter can refer to the above description of the first bass parameter, which is not repeated here.
[0110] For example, the first bass parameter includes low frequency diving value (Lf) and power (P), and the first bass parameter value can be determined according to the following formula:
[0111] Lp=a*40 / Lf+(1-a)*P / 200;
[0112] Wherein, a is a weighting coefficient, a<1; Lf is the low frequency diving value, P is the power of the bass part, and Lp is the bass parameter value. Table 1 below is an example of Lp calculation.
[0113] Table 1
[0114]
[0115] Since the higher the bass parameter value, the better the sound effect of the sound emitted by the bass cannon channel, in the case where the first bass parameter value is greater than the second bass parameter value, it is determined that the third channel includes the bass cannon channel; in the case where the first bass parameter value is less than the second bass parameter value, it is determined that the fourth channel includes the bass cannon channel. In this way, the bass cannon channel in the target fusion scheme can be more accurately determined, and the sound effect of the joint sound of the first device and the second device can be improved.
[0116] It can be understood that in the case where only one of the first device and the second device includes a bass cannon channel, the bass cannon channel included by the one device is the bass cannon channel of the target fusion scheme, and in the case where both the first device and the second device include a bass cannon channel, the bass cannon channel in the target fusion scheme is determined to be the bass cannon channel of the first device or the bass cannon channel of the second device according to the size relationship between the first bass parameter value and the second bass parameter value.
[0117] The controller is further configured to: obtain a first audio signal corresponding to the third channel and a second audio signal corresponding to the fourth channel from the target audio signal, the first audio signal including audio signals corresponding to each channel in the third channel, and the second audio signal including audio signals corresponding to each channel in the fourth channel.
[0118] The first audio signal corresponding to the third channel and the second audio signal corresponding to the fourth channel are obtained from the target audio signal. Specifically, the audio signal corresponding to at least one of the first channel and the second channel can be directly separated from the target audio signal. Alternatively, the audio signal corresponding to at least one of the first channel and the second channel can be down-mixed from the target audio signal. Alternatively, the audio signal corresponding to at least one of the first channel and the second channel can be virtually generated by up-mixing the target audio signal. The actual implementation can be determined according to actual conditions, which is not limited herein.
[0119] In some embodiments of the present disclosure, the controller is specifically configured to render the target audio signal into a multi-channel audio signal matching a preset channel system, the preset channel system including a plurality of channels, the multi-channel audio signal including an audio signal corresponding to each channel of the plurality of channels; in a case where the channels included in the target fusion scheme include all channels of the plurality of channels, separating the first audio signal and the second audio signal from the multi-channel audio signal; in a case where the channels included in the target fusion scheme include part of the plurality of channels, performing audio mixing processing on the multi-channel audio signal to obtain the first audio signal and the second audio signal.
[0120] The preset channel system can be a channel system standard, which can be determined according to the device type, device use scenario, and required surround sound effect of the first device, etc. For example, the preset channel system can be a 5.1.2 channel system or a 7.1.4 channel system, etc. The actual implementation can be determined according to actual conditions, which is not limited herein.
[0121] It can be understood that, if the target audio signal itself is a multi-channel audio signal matching the preset channel system, the first audio signal and the second audio signal can be directly obtained from the target audio signal. If the target audio signal itself is not a multi-channel audio signal matching the preset channel system, the target audio signal can be rendered into a multi-channel audio signal matching the preset channel system through audio mixing processing, virtual surround processing, etc.
[0122] It can be understood that, in a case where the first channel and the second channel include the plurality of channels, the first audio signal and the second audio signal can be directly separated from the multi-channel audio signal. In a case where the first channel and the second channel include part of the plurality of channels (i.e., the types of channels included in the first channel and the second channel are less than the types of channels included in the plurality of channels), the first audio signal and the second audio signal can be obtained by down-mixing the multi-channel audio signal.
[0123] In the embodiments of the present disclosure, the target audio signal is first rendered into a multi-channel audio signal matching the preset channel system, then the first audio signal and the second audio signal are obtained based on the relationship between the channel types included in the first channel and the second channel determined according to the target fusion scheme and the channel types included in the preset channel system, and the first device and the second device jointly sound based on the first audio signal and the second audio signal, so as to achieve the sound effect of playing the multi-channel audio signal matching the preset channel system and realize the surround sound effect.
[0124] The controller is further configured to perform delay processing on the first audio signal and the second audio signal based on the first delay and the second delay respectively to obtain a third audio signal and a fourth audio signal, so as to eliminate the difference in transmission delay between the channels in the third channel and the fourth channel.
[0125] The first delay is the transmission delay of the hardware corresponding to each channel in the first channel of the first device, and the second delay is the transmission delay of the hardware corresponding to each channel in the second channel of the second device, so that the first audio signal and the second audio signal can be processed based on the first delay and the second delay respectively to obtain the third audio signal and the fourth audio signal, so as to eliminate the difference in transmission delay between the channels in the third channel and the fourth channel.
[0126] In some embodiments of the present disclosure, the controller is specifically configured to determine the maximum delay in the first delay and the second delay, determine the delay offset corresponding to each channel as the difference between the maximum delay and each delay in the first delay and the second delay, and increase the delay offset in the first audio signal and the second audio signal respectively based on the delay offset corresponding to each channel to obtain the third audio signal and the fourth audio signal. In this way, by increasing the delay offset corresponding to each channel in the audio signal, the difference in transmission delay between the channels is offset, which can ensure that the first device and the second device play the third audio signal and the fourth audio signal synchronously to a certain extent.
[0127] In some embodiments of the present disclosure, if the channel delays between the channels in the first device are the same, the first delay is a value, and if the channel delays between the channels in the second device are the same, the second delay is a value. The controller is specifically configured to increase the target delay offset in the second audio signal to obtain the fourth audio signal when the first delay is greater than or equal to the second delay, and the third audio signal is the first audio signal; increase the target delay offset in the first audio signal to obtain the third audio signal when the first delay is less than the second delay, and the fourth audio signal is the second audio signal; and the target delay offset is the absolute value of the difference between the first delay and the second delay. In this way, the difference in transmission delay between the channels in the third channel and the fourth channel can be eliminated quickly to obtain the third audio signal and the fourth audio signal.
[0128] The controller is further configured to: based on the fourth sound channel and the fourth audio signal, send target information indicating that the second device plays the fourth audio signal through the fourth sound channel; and based on the third delay, play the third audio signal through the third sound channel so that the third audio signal and the fourth audio signal are played synchronously, the third delay indicating a difference between a time at which the fourth sound channel receives the fourth audio signal and a time at which the third sound channel receives the third audio signal.
[0129] The third delay includes an encoding delay, a network transmission delay (related to a network transmission manner of the audio signal between the first device and the second device), a decoding delay, a playing buffer delay, etc., and can be calculated by a specific algorithm, which is not limited here. The third delay is a signal transmission delay between the first device and the second device.
[0130] In the embodiments of the present disclosure, based on the sound channel delay of the first device and the second device, the first audio signal played by the first device and the second audio signal played by the second device are subjected to delay processing to eliminate the difference in transmission delay between the sound channels in the third sound channel and the fourth sound channel, and based on the third delay, the signal transmission delay between the first device and the second device is eliminated, so that the sound channels of the first device and the second device can play the corresponding audio signals synchronously, the sound field positioning accuracy and the spatial sense can be improved, and the joint sound effect can be improved.
[0131] In some embodiments of the present disclosure, the target information includes first audio data packets and control information; the controller is specifically configured to: perform low-delay compression processing on a first audio data frame in the fourth audio signal to obtain a first audio data packet, the first audio data packet carrying timestamp information; send the first audio data packet; and send the control information; wherein the timestamp information is used for the second device to sort the received audio data packets so that the audio signals played by the first device and the second device are synchronous.
[0132] The low-delay compression processing is a low-delay audio compression algorithm: using "audio transmission redundancy coding", filtering data such as super frequency, empty packet, noise, etc. in the audio information, greatly reducing the data amount of audio transmission.
[0133] It can be understood that the audio data packets received by the receiving end (the second device) are often misaligned due to receiving positions, interference, etc. Therefore, by extracting the timestamp information in the received audio data packet, the audio data frames in the audio data packet are rearranged according to the timestamp indicated by the timestamp information, and then transmitted to the sound channel for playing, so that the audio signals played by the first device and the second device are synchronous.
[0134] In some embodiments of the present disclosure, the controller is specifically configured to: based on the first compression parameter, perform low-delay compression processing on the first audio data frame to obtain the first audio data packet, the first compression parameter including a first sampling rate and a first bit precision; the controller is also configured to: receive first feedback information, the first feedback information being used to indicate real-time audio detection results and / or network congestion analysis results obtained by the second device; in a case where the first feedback information indicates that the data quality meets a first condition, based on a second compression parameter, performing low-delay compression processing on the second audio data frame in the fourth audio signal to obtain a second audio data packet, the second compression parameter including a second sampling rate and a second bit precision, the second sampling rate being greater than the first sampling rate, and the second bit precision being greater than the first bit precision; in a case where the first feedback information indicates that the data quality meets a second condition, based on a third compression parameter, performing low-delay compression processing on the second audio data frame in the fourth audio signal to obtain a third audio data packet, the third compression parameter including a third sampling rate and a third bit precision, the third sampling rate being less than the first sampling rate, and the third bit precision being less than the first bit precision.
[0135] Adaptive bit rate control: during the audio transmission process, the bit rate of the audio is adjusted in real time according to the changes in network bandwidth and delay and other conditions, to ensure the best transmission effect.
[0136] Among them, the sampling rate and bit precision can determine the transmission bit rate of the audio data packet, so in the case where it is determined that the transmission bit rate needs to be adjusted, the transmission bit rate can be adjusted by adjusting the sampling rate and bit precision.
[0137] Real-time audio detection includes analysis of digital audio signals, mainly from the time domain and frequency domain, and several key indicators are the relationship between the amplitude envelope, zero-crossing rate, instantaneous energy change value and the preset indicators. Among them, the reference indicator is that when the instantaneous energy has a significant rise or fall (by establishing a threshold to do masking, assuming that the threshold is set at 0.1 times the maximum detection amplitude) within a certain time length (such as 100 milliseconds) of the detection period of the amplitude envelope, it is determined that a larger amplitude occurs. Assuming that the threshold of the zero-crossing rate is 0.5, when a signal higher than the threshold appears, it is determined that there are more high-frequency components at present, and higher transmission bandwidth is needed.
[0138] The network state in the current period and the future period (according to Buffer adjustment) can be judged and predicted by analyzing real-time data, and the transmission strategy is adjusted in real time. The network congestion analysis includes: mainly through the delay of the transceiving packet, the packet loss rate, the delay jitter superposition algorithm is realized. Among them, the transceiving packet delay threshold is assumed to be 20 milliseconds, the packet loss rate threshold is assumed to be 0.5%, and the delay jitter rate threshold is assumed to be 5%. Under normal circumstances, the packet loss rate should be less than 0.5%, the delay jitter should be less than 20 milliseconds, and the delay jitter rate should be less than 5%. When the comprehensive judgment of several parameters appears to be poor (that is, network congestion prediction), the frequency hopping mechanism (preset frequency hopping algorithm, which does not affect the audio quality) is started first, and when the frequency hopping mechanism cannot find a suitable channel for audio transmission, the audio transmission bit rate and the audio transmission buffer need to be dynamically adjusted to ensure the smoothness of the audio transmission channel.
[0139] Based on the above data, the bit rate is dynamically adjusted, and the audio data is frame compressed at 16bit / 24bit. When the transmission bit rate is high, the capacity of the audio buffer buffer needs to be appropriately relaxed to avoid audio interruption due to insufficient instantaneous bandwidth. The adjustment amount of the buffer is 10% of the maximum cache, for example, the original audio buffer is 25KB, and the adjustment is to 28KB to accommodate more high-bit audio data.
[0140] Among them, the first condition is used to indicate that the signal transmission is normal, including that the real-time audio detection result indicates that the real-time audio change is normal and / or the network congestion analysis result indicates that the network is not congested, and the second condition indicates that the signal transmission is abnormal, including that the real-time audio detection result indicates that the real-time audio change is abnormal and / or the network congestion analysis result indicates that the network is congested.
[0141] In the embodiments of the present disclosure, when the audio signal is normally transmitted, high sampling rate and high bit precision transmission are adopted to provide high-quality audio data (the higher the sampling rate, the more audio details. The higher the bit precision, the larger the audio data amount and the better the audio quality), and more cache is released to provide more operation space for the host; when the transmission quality is poor, the sampling rate and bit precision in the audio data packet are reduced, and the cache is increased; after the transmission is stable, the sampling rate and bit precision are gradually increased, the cache is released, and the best transmission quality and the best audio quality output are achieved.
[0142] In some embodiments of the present disclosure, the audio data packet can also carry the packet length, the asynchronous sampling rate, etc. Among them, the bit precision is proportional to the packet length, and the higher the bit precision, the larger the packet length.
[0143] In view of the sensitivity difference of different sound channels, when the sound equipment sounds alone, only the virtual surround environment formed by each loudspeaker of the sound equipment is adjusted, and after the display equipment is fused, the amplitude of the audio output by each loudspeaker needs to be re-adjusted to form a larger virtual surround environment with the display equipment. Otherwise, in the new ambisonic environment, the user hears the primary and secondary audio information mixed up, and the sound field orientation is inaccurate.
[0144] In some embodiments of the present disclosure, the first sound channel parameter further includes a first sensitivity, the first sensitivity being used to indicate the sensitivity corresponding to each sound channel in the first sound channel, and the second sound channel parameter further includes a second sensitivity, the second sensitivity being used to indicate the sensitivity corresponding to each sound channel in the second sound channel; the controller is specifically configured to: based on the first delay and the second delay, delay processes the first audio signal and the second audio signal respectively to obtain a fifth audio signal and a sixth audio signal, so as to eliminate the difference in transmission delay between each sound channel in the third sound channel and the fourth sound channel; based on the first sensitivity, the second sensitivity and a preset adjustment strategy, determines the gain adjustment amount of each sound channel in the third sound channel and the fourth sound channel; based on the gain adjustment amount of each sound channel, gain adjustment processes the loudness of the fifth audio signal and the sixth audio signal respectively to obtain the third audio signal and the fourth audio signal, so as to match the sensitivity difference of the third sound channel and the fourth sound channel.
[0145] Among them, the sensitivity is a key indicator to measure the response ability of a certain sound channel (such as left channel, right channel, surround channel, etc.) to the input signal in the audio system, usually expressed in "sound pressure level (SPL, unit: dB)", reflecting the volume size (i.e. loudness) that the sound channel can produce under a certain power input. It not only affects the volume balance of each sound channel, but also directly relates to sound field positioning, detail restoration and overall listening experience.
[0146] In the embodiments of the present disclosure, based on the first sensitivity, the second sensitivity and a preset adjustment strategy, the gain adjustment amount of each sound channel in the third sound channel and the fourth sound channel is determined. Among them, the preset adjustment strategy includes: first, the sensitivity of the sound channels with the same function and left-right symmetry should be consistent, for example, the left channel (also known as the left front channel) and the right channel (also known as the right front channel) should be consistent, the left surround and the right surround should be consistent, etc.; second, the sensitivity of the sound channels with different functions should be higher than that of the front channel by a first preset value (representing the characteristics of the sound system) (for example, 3dB); the front channel should be larger than the surround channel by a second preset value (for example, 3dB); the sky sound channel should be larger than the front channel by a third preset value (for example, one value of 3-5dB) due to the longer path.
[0147] Exemplarily, the sensitivity of the sound channels (sound channel speakers) of the second device and the sensitivity of the sound channels of the first device will have some differences, which need to be matched by an algorithm, the first device pre-stores the sensitivity of each sound channel, and the second device also stores the sensitivity of each sound channel. The sensitivity matching is not the sensitivity matching between the same sound channels of the first device and the second device, for example, the sensitivity between the left sound channel of the first device and the left sound channel of the second device is not required to be consistent. Instead, after the first device and the second device determine the target fusion scheme, the sensitivity of different functional sound channels between the first device and the second device is matched to reach a harmonious value. For the case that the first device and the second device have the same type of sound channel in the target fusion scheme, the same type of sound channel in the first device and the second device is equivalent to one sound channel (the sensitivity of the equivalent sound channel is calculated), and the sensitivity adjustment is performed between the equivalent sound channel and other types of sound channels.
[0148] Exemplarily, the first device is a TV (TV Speaker), and the second device is an external speaker (External Speaker). Exemplarily, the sensitivity adjustment between the sky sound channel of the TV and the left sound channel of the external speaker in the target fusion scheme is as follows: the sensitivity difference between the sky sound channel of the TV and the left sound channel of the external speaker is calculated: ΔS = S1-S2, (S1 is the sensitivity of the sky sound channel of the TV, and S2 is the sensitivity of the left front sound channel of the external speaker); the sensitivity difference is converted into a gain difference ΔG = 10^{(Gain(ΔS) / 20)}; the amplitude adjustment is exemplified by TV Speaker_C'1 (the adjusted gain of the sky sound channel of the TV) and External Speaker C”2 (the adjusted gain of the left front sound channel of the external speaker). If S1 > S2, Gain_TV Speaker_C'1_MF = Gain_TV Speaker_C'1_MFOriginal-ΔG+ΔG_C; if S1 < S2, Gain_External Speaker_C”2_MF = Gain_External Speaker_C”2_MFOriginal-ΔG+ΔG_C'. It can be understood that if the sensitivity of the TV is high, the gain of the TV is reduced, and if the sensitivity of the TV is low, the gain of the external speaker is reduced, that is, the principle is to reduce the gain of the device with high sensitivity.
[0149] Wherein, Gain_TV Speaker_C'1_MFOriginal and Gain_External Speaker_C"2_MFOriginal are the original gain values of the TV and the audio equipment respectively, and ΔG_C and ΔG_C' are compensation values or empirical values. The sensitivity matching between different functional channels is not the best effect, and there is a certain difference that is the most suitable, therefore, ΔG_C and ΔG_C' are reserved in the algorithm for the adjustment of the debug personnel.
[0150] In the embodiments of the present disclosure, the gain adjustment amount of each channel in the third channel and the fourth channel is determined based on the first sensitivity, the second sensitivity and a preset adjustment strategy; the loudness of the fifth audio signal and the sixth audio signal is respectively subjected to gain adjustment processing based on the gain adjustment amount of each channel, to obtain the third audio signal and the fourth audio signal, so as to match the sensitivity difference of the third channel and the fourth channel. The loudness of the audio signals played between different channels in the target fusion scheme can be matched with each other, there is a certain rule, and there is no big difference, so as to ensure the accurate sound positioning, correct the sound field error, maintain the reasonable dynamic range, etc., and the effect of joint sound emission can be improved.
[0151] For the frequency width of the channel, when the audio equipment emits sound alone, only the sound field range formed by each loudspeaker of the audio equipment and the number of channels are considered. After the display device is fused, the sound field range is increased and the number of channels is superimposed. When high, medium and low frequencies are distributed, the larger sound field and more loudspeakers are re-adjusted and re-distributed to each loudspeaker to form a larger sound field and more channels of a full-sound environment. Otherwise, the user hears incomplete audio information in the new full-sound environment, and the sound field orientation and frequency response are chaotic.
[0152] In some embodiments of the present disclosure, the first channel parameter further includes a first frequency width, and the first frequency width is used to indicate the sound emission frequency width corresponding to each channel in the first channel; the second channel parameter further includes a second frequency width, and the second frequency width is used to indicate the sound emission frequency width corresponding to each channel in the second channel; and the controller is specifically configured to: based on the gain adjustment amount of each channel, the loudness of the fifth audio signal and the sixth audio signal is respectively subjected to gain adjustment processing to obtain the seventh audio signal and the eighth audio signal, so as to match the sensitivity difference of the third channel and the fourth channel; based on the first frequency width and the second frequency width, the loudness of the seventh audio signal and the eighth audio signal is respectively subjected to gain adjustment processing to obtain the third audio signal and the fourth audio signal, so as to balance the loudness of the joint sound emission of the same type of channels in the third channel and the fourth channel.
[0153] Bandwidth refers to the frequency range within which a channel can produce sound. Gain adjustment is the process of adjusting the audio system's gain parameter to alter the strength of the sound signal, thereby controlling the volume / loudness of the audio signal to meet auditory needs or system matching requirements. This helps balance volume differences across frequency bands, avoid distortion, and adapt to the playback device or environment.
[0154] like Figure 9 As shown, a comparison chart of the bandwidths of three groups of the same type of channels is shown, taking the first device as a display device and the second device as an audio device as an example. The thicker line represents the bandwidth of the display device, and the thinner line represents the bandwidth of the audio device. It can be seen that there is a difference in the bandwidth of the same type of channels between the audio device and the display device. Therefore, the audio signal is enhanced in the bandwidth that can be covered by the channels of both the audio device and the display device, while the bandwidth covered by the channels of the audio device alone and the bandwidth covered by the channels of the display device alone cannot be enhanced. As a result, the played audio signal will be very loud in some frequency bands and suddenly very quiet in other frequency bands, resulting in problems such as sound imbalance, loss of details or auditory fatigue, seriously affecting the joint sound effect.
[0155] Therefore, in the embodiment of the present disclosure, gain adjustment processing is performed on the loudness of the seventh audio signal and the eighth audio signal based on the first bandwidth and the second bandwidth, respectively, to obtain the third audio signal and the fourth audio signal. This balances the loudness of the combined sound of different frequency bands of the same type of channels in the third channel and the fourth channel. This can balance the volume differences in different frequency bands, avoid distortion, and adapt to the playback device or environment.
[0156] The embodiment of the present disclosure provides a second device, comprising: a communicator; a controller configured to: receive a target request message, the target request message being used for requesting a channel parameter of the second device, the target request message being generated in a case that a first device connects the second device and the first device plays a target audio signal, the first device further acquires a first channel parameter in response to the operation of playing the target audio signal, the first channel parameter comprising a first channel supported by the first device and a first delay, the first delay being used for indicating a transmission delay corresponding to each channel in the first channel; acquire a second channel parameter in response to the target request message, the second channel parameter comprising a second channel supported by the second device and a second delay, the second delay being used for indicating a transmission delay corresponding to each channel in the second channel; send a target response message corresponding to the target request message, the target response message carrying the second channel parameter, so that the first device determines a target fusion scheme comprising a third channel in the first channel used for playing the target audio signal and a fourth channel in the second channel used for playing the target audio signal based on the first channel and the second channel, and acquires a first audio signal corresponding to the third channel and a second audio signal corresponding to the fourth channel from the target audio signal, and performs delay processing on the first audio signal and the second audio signal based on the first delay and the second delay respectively to obtain a third audio signal and a fourth audio signal, so as to eliminate the difference between the transmission delays of each channel in the third channel and the fourth channel; the first audio signal comprises audio signals corresponding to each channel in the third channel, and the second audio signal comprises audio signals corresponding to each channel in the fourth channel; receive target information, the target information being generated based on the fourth channel and the fourth audio signal; play the fourth audio signal through the fourth channel in response to the target information, so that the first device plays the third audio signal and the fourth audio signal synchronously based on a third delay through the third channel, the third delay being used for indicating a difference between a time when the fourth channel receives the fourth audio signal and a time when the third channel receives the third audio signal.
[0157] In some embodiments of the present disclosure, the target information includes a first audio data packet and control information; the controller is specifically configured to: receive the first audio data packet, the first audio data packet being generated by performing low-delay encoding processing on a first audio data frame in the fourth audio signal based on first compression parameters, the first audio data packet carrying timestamp information, and the first compression parameters including a first sampling rate and a first bit precision; receive the control information; sort the received first audio data packet based on the timestamp information, so as to synchronize the audio signals played by the first device and the second device; obtain first feedback information, the first feedback information being used to indicate real-time audio detection results and / or network congestion analysis results obtained by the second device; and send the first feedback information, so that the first device performs low-delay compression processing on a second audio data frame in the fourth audio signal based on second compression parameters to obtain a second audio data packet in a case where the data quality indicated by the first feedback information meets a first condition, and performs low-delay compression processing on the second audio data frame in the fourth audio signal based on third compression parameters to obtain a third audio data packet in a case where the data quality indicated by the first feedback information meets a second condition, the second compression parameters including a second sampling rate and a second bit precision, the second sampling rate being greater than the first sampling rate, and the second bit precision being greater than the first bit precision, and the third compression parameters including a third sampling rate and a third bit precision, the third sampling rate being less than the first sampling rate, and the third bit precision being less than the first bit precision.
[0158] In some embodiments of the present disclosure, the controller is further configured to: in a case where the data quality indicated by the first feedback information meets the first condition, adjust the receiving buffer from the first buffer to a second buffer, the second buffer being less than the first buffer; and in a case where the data quality indicated by the first feedback information meets the second condition, adjust the receiving buffer from the first buffer to a third buffer, the third buffer being greater than the first buffer.
[0159] In some embodiments of the present disclosure, in a case where the data quality indicated by the first feedback information meets the first condition, the receiving buffer is preferentially adjusted, and when the receiving buffer has been adjusted to a critical value, the transmission bit rate is adjusted.
[0160] In some embodiments of the present disclosure, the detailed description of the second device can refer to the related description of the first device in the above embodiments, which is not described herein again.
[0161] In order to more specifically describe the present scheme, the following will be described in an exemplary manner. It can be understood that the following steps involved in the actual implementation can include more steps or fewer steps, and the order between the steps can also be different, so as to be able to implement the audio playing method provided in the embodiments of the present disclosure.
[0162] Figure 10To implement a step flowchart of an audio signal playing method according to one or more embodiments of the present disclosure, the audio signal playing method can include S1001 to S1013 described below.
[0163] S1001, in response to an operation of playing a target audio signal, a first device acquires first channel parameters in a case where the first device is connected to a second device, the first channel parameters including first channels supported by the first device and first delays.
[0164] The first delays are used to indicate transmission delays corresponding to respective channels in the first channels.
[0165] S1002, the first device sends a target request message, the target request message being used to request channel parameters of the second device.
[0166] S1003, the second device receives the target request message.
[0167] The target request message is used to request the channel parameters of the second device, the target request message being generated in response to the operation of playing the target audio signal and in a case where the first device is connected to the second device, the first device further acquiring the first channel parameters in response to the operation of playing the target audio signal, the first channel parameters including the first channels supported by the first device and the first delays, the first delays being used to indicate transmission delays corresponding to respective channels in the first channels.
[0168] S1004, the second device acquires second channel parameters in response to the target request message, the second channel parameters including second channels supported by the second device and second delays.
[0169] The second delays are used to indicate transmission delays corresponding to respective channels in the second channels.
[0170] S1005, the second device sends a target response message corresponding to the target request message.
[0171] The target response message carries the second channel parameters, so that the first device determines a target fusion scheme including third channels in the first channels for playing the target audio signal and fourth channels in the second channels for playing the target audio signal based on the first channels and the second channels, and acquires first audio signals corresponding to the third channels and second audio signals corresponding to the fourth channels from the target audio signal, and respectively performs delay processing on the first audio signals and the second audio signals based on the first delays and the second delays to obtain third audio signals and fourth audio signals, so as to eliminate differences in transmission delays between respective channels in the third channels and the fourth channels; the first audio signals include audio signals corresponding to respective channels in the third channels, and the second audio signals include audio signals corresponding to respective channels in the fourth channels.
[0172] S1006, the first device receives a target response message corresponding to the target request message, the target response message carrying a second channel parameter.
[0173] The second channel parameter is acquired by the second device in response to the target request message, and the second channel parameter includes a second channel supported by the second device and a second delay, the second delay being used to indicate a transmission delay corresponding to each channel in the second channel.
[0174] S1007, the first device determines a target fusion scheme based on the first channel and the second channel, the target fusion scheme including a third channel in the first channel for playing the target audio signal and a fourth channel in the second channel for playing the target audio signal.
[0175] S1008, the first device acquires a first audio signal corresponding to the third channel and a second audio signal corresponding to the fourth channel from the target audio signal.
[0176] The first audio signal includes an audio signal corresponding to each channel in the third channel, and the second audio signal includes an audio signal corresponding to each channel in the fourth channel.
[0177] S1009, the first device respectively processes the first audio signal and the second audio signal based on the first delay and the second delay to obtain a third audio signal and a fourth audio signal, so as to eliminate the difference between the transmission delays of each channel in the third channel and the fourth channel.
[0178] S1010, the first device sends target information based on the fourth channel and the fourth audio signal.
[0179] The target information is used to instruct the second device to play the fourth audio signal through the fourth channel.
[0180] S1011, the second device receives the target information.
[0181] The target information is generated based on the fourth channel and the fourth audio signal.
[0182] S1012, the second device plays the fourth audio signal through the fourth channel in response to the target information.
[0183] S1013, the first device plays the third audio signal through the third channel based on a third delay, so that the third audio signal and the fourth audio signal are played synchronously, and the third delay is used to indicate the difference between the time when the fourth channel receives the fourth audio signal and the time when the third channel receives the third audio signal.
[0184] In some embodiments of the present disclosure, the above S1008 can be implemented through the following S1008a-S1008c.
[0185] S1008a, the first device renders the target audio signal into a multi-channel audio signal matching the preset channel system.
[0186] The preset channel system includes a plurality of channels, and the multi-channel audio signal includes an audio signal corresponding to each channel in the plurality of channels.
[0187] S1008b, the first device separates the first audio signal and the second audio signal from the multi-channel audio signal in a case where the channels included in the target fusion scheme include all channels in the plurality of channels.
[0188] S1008c, the first device performs a mixing process on the multi-channel audio signal to obtain the first audio signal and the second audio signal in a case where the channels included in the target fusion scheme include part of the plurality of channels.
[0189] In some embodiments of the present disclosure, the above S1007 can be implemented by the following S1007a.
[0190] S1007a, the first device determines the first channel as the third channel and the second channel as the fourth channel to determine the target fusion scheme.
[0191] In some embodiments of the present disclosure, the above S1007 can be implemented by the following S1007b.
[0192] S1007b, the first device sets one of the two channels of the same type as the target channel or performs mute processing in a case where there are channels of the same type in the first channel and the second channel, to determine the target fusion scheme.
[0193] Wherein, there is no channel of the same type in the third channel and the fourth channel, and the target channel is a channel that does not exist in the first channel and the second channel.
[0194] In some embodiments of the present disclosure, the target information includes a first audio data packet and control information; the above S1010 can be implemented by the following S1010a-S1010c, and the above S1011 can be implemented by the following S1011a-S1011c.
[0195] S1010a, the first device performs low-delay compression processing on the first audio data frame in the fourth audio signal to obtain a first audio data packet, and the first audio data packet carries timestamp information.
[0196] The timestamp information is used by the second device to sort the received audio data packets, so that the audio signals played by the first device and the second device are synchronized.
[0197] S1010b, the first device transmits the first audio data packet.
[0198] S1011a, the second device receives the first audio data packet.
[0199] S1010c, the first device transmits the control information.
[0200] S1011b, the second device receives the control information.
[0201] S1011c, the second device sorts the received first audio data packet based on the timestamp information, so that the audio signals played by the first device and the second device are synchronized.
[0202] In some embodiments of the present disclosure, the above S1010a can be implemented by the following S1010a1. After the above S1011a, the audio signal playing method provided by the embodiments of the present disclosure can further include the following S1014 to S1018.
[0203] S1010a1, the first device performs low-delay compression processing on the first audio data frame based on first compression parameters to obtain the first audio data packet, the first compression parameters including a first sampling rate and a first bit precision.
[0204] S1014, the second device obtains first feedback information, the first feedback information being used to indicate real-time audio detection results and / or network congestion analysis results obtained by the second device.
[0205] S1015, the second device transmits the first feedback information.
[0206] S1016, the first device receives the first feedback information.
[0207] S1017, the first device performs low-delay compression processing on the second audio data frame in the fourth audio signal based on second compression parameters to obtain a second audio data packet, in a case where the first feedback information indicates that the data quality meets a first condition.
[0208] The second compression parameters include a second sampling rate and a second bit precision, the second sampling rate being greater than the first sampling rate, and the second bit precision being greater than the first bit precision.
[0209] S1018, the first device performs low-delay compression processing on the second audio data frame in the fourth audio signal based on third compression parameters to obtain a third audio data packet, in a case where the first feedback information indicates that the data quality meets a second condition.
[0210] The third compression parameters include a third sampling rate and a third bit precision, the third sampling rate being less than the first sampling rate, and the third bit precision being less than the first bit precision.
[0211] In some embodiments of the present disclosure, after S1014, the audio signal playing method provided by the embodiments of the present disclosure can further include S1019-S1020.
[0212] S1019, the second device adjusts the receiving buffer from the first buffer to a second buffer in a case where the first feedback information indicates that the data quality meets the first condition, the second buffer being smaller than the first buffer.
[0213] S1020, the second device adjusts the receiving buffer from the first buffer to a third buffer in a case where the first feedback information indicates that the data quality meets the second condition, the third buffer being larger than the first buffer.
[0214] In the embodiments of the present disclosure, S1016-S1018 and S1019-S1020 can be executed only, or S1016-S1018 and S1019-S1020 can be executed together. Specifically, it can be determined according to actual conditions, which is not limited here. In the case where S1016-S1018 and S1019-S1020 are executed together, S1019-S1020 can be executed first, and then S1016-S1018 can be executed, or S1016-S1018 and S1019-S1020 can be executed simultaneously. Specifically, it can be determined according to actual conditions, which is not limited here.
[0215] When the sound equipment is a combination of a sound host and at least one sub-speaker, the sound host also needs to perform low-delay compression processing on the audio signal when the audio signal is transmitted between the sound host and the sub-speaker. The sound host needs to determine whether the data instruction indicated by the feedback information returned by the sub-speaker meets the first condition or the second condition, and adjust the compression parameter of the low-delay compression processing according to the determination result. For details, please refer to the transmission process of the audio signal between the first device and the second device, which is not repeated here.
[0216] In some embodiments of the present disclosure, S1009 can be implemented by S1009a-S1009c.
[0217] S1009a, the first device determines the maximum delay time from the first delay time and the second delay time.
[0218] S1009b, the first device determines the delay offset corresponding to each channel as the difference between the maximum delay time and each delay time from the first delay time and the second delay time.
[0219] S1009c, the first device adds the delay offset corresponding to each channel to the first audio signal and the second audio signal respectively to obtain the third audio signal and the fourth audio signal.
[0220] In some embodiments of the present disclosure, the first sound channel parameter further comprises a first sensitivity, the first sensitivity being used to indicate the sensitivity corresponding to each sound channel in the first sound channel, and the second sound channel parameter further comprises a second sensitivity, the second sensitivity being used to indicate the sensitivity corresponding to each sound channel in the second sound channel; and the S1009 can be specifically implemented by the following S1009d-S1009f.
[0221] The S1009d can be specifically implemented by the following S1009d1-S1009d2.
[0222] The S1009e can be specifically implemented by the following S1009e1-S1009e2.
[0223] The S1009f can be specifically implemented by the following S1009f1-S1009f2.
[0224] In some embodiments of the present disclosure, the first sound channel parameter further comprises a first frequency width, the first frequency width being used to indicate the sound production frequency width corresponding to each sound channel in the first sound channel; and the second sound channel parameter further comprises a second frequency width, the second frequency width being used to indicate the sound production frequency width corresponding to each sound channel in the second sound channel; and the S1009f can be specifically implemented by the following S1009f1-S1009f2.
[0225] The S1009f can be specifically implemented by the following S1009f1-S1009f2.
[0226] The S1009f2 can be specifically implemented by the following S1009f2a-S1009f2b.
[0227] In the embodiments of the present disclosure, the order of adjusting the first audio signal and the second audio signal in terms of bandwidth, sensitivity and delay is not limited. For example, the first audio signal and the second audio signal can be first adjusted in terms of bandwidth, then further adjusted in terms of sensitivity, and then further adjusted in terms of delay. The first audio signal and the second audio signal can also be first adjusted in terms of sensitivity, then further adjusted in terms of bandwidth, and then further adjusted in terms of delay. The first audio signal and the second audio signal can also be first adjusted in terms of delay, then further adjusted in terms of bandwidth, and then further adjusted in terms of sensitivity. The first audio signal and the second audio signal can also be adjusted in other orders, which can be determined according to actual use, and is not limited herein.
[0228] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present disclosure, and not to limit them; although the present disclosure has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that: it can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement for part or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present disclosure.
[0229] For the convenience of explanation, the above description has been made in combination with specific embodiments. However, the above exemplary discussion is not intended to exhaust or limit the embodiments to the specific forms disclosed above. Various modifications and variations can be derived according to the above teachings. The selection and description of the above embodiments are to better explain the principles and practical applications, so that those skilled in the art can better use the embodiments and various different modified embodiments suitable for specific use considerations.
Claims
1. A first device, characterized in that: include: communicator; The controller is configured to: in response to an operation of playing a target audio signal, when the first device is connected to a second device, obtain first channel parameters, the first channel parameters including first channels supported by the first device and first delays, the first delays being used to indicate a transmission delay corresponding to each channel in the first channel; Sending a target request message, where the target request message is used to request a channel parameter of the second device; receiving a target response message corresponding to the target request message, the target response message carrying second channel parameters, where the second channel parameters are acquired by the second device in response to the target request message, the second channel parameters including second channels supported by the second device and second delays, where the second delays are used to indicate a transmission delay corresponding to each channel in the second channel; determining a target fusion scheme based on the first channel and the second channel, the target fusion scheme including a third channel in the first channel for playing the target audio signal and a fourth channel in the second channel for playing the target audio signal; acquiring, from the target audio signal, a first audio signal corresponding to the third channel and a second audio signal corresponding to the fourth channel, wherein the first audio signal includes audio signals corresponding to respective channels in the third channel, and the second audio signal includes audio signals corresponding to respective channels in the fourth channel; performing delay processing on the first audio signal and the second audio signal based on the first delay and the second delay to obtain a third audio signal and a fourth audio signal, so as to eliminate a difference in transmission delay between the third channel and the fourth channel; sending target information based on the fourth sound channel and the fourth audio signal, where the target information is used to instruct the second device to play the fourth audio signal through the fourth sound channel; The third audio signal is played through the third channel based on a third delay so that the third audio signal and the fourth audio signal are played synchronously, where the third delay indicates a difference between a time when the fourth channel receives the fourth audio signal and a time when the third channel receives the third audio signal.
2. The first device according to claim 1, characterized in that The controller is specifically configured to: Rendering the target audio signal into a multi-channel audio signal matching a preset channel system, where the preset channel system includes a plurality of channels, and the multi-channel audio signal includes an audio signal corresponding to each of the plurality of channels; In a case where the channels included in the target fusion solution include all of the multiple channels, separating the first audio signal and the second audio signal from the multi-channel audio signal; In a case where the channels included in the target fusion solution include some of the multiple channels, mixing processing is performed on the multi-channel audio signal to obtain the first audio signal and the second audio signal.
3. The first device according to claim 1, characterized in that The controller is specifically configured to: Determine the first sound channel as the third sound channel and the second sound channel as the fourth sound channel to determine the target fusion solution; or, If the first channel and the second channel have channels of the same type, one of the two channels of the same type is set as a target channel or muted to determine the target fusion solution. If the third channel and the fourth channel do not have channels of the same type, the target channel is a channel that does not exist in either the first channel or the second channel.
4. The first device according to claim 1, characterized in that The target information includes a first audio data packet and control information; the controller is specifically configured to: performing low-latency compression processing on a first audio data frame in the fourth audio signal to obtain a first audio data packet, where the first audio data packet carries timestamp information; sending the first audio data packet; sending the control information, where the control information is used to instruct the second device to play the fourth audio signal through the fourth channel; The timestamp information is used by the second device to sort the received audio data packets so as to synchronize the audio signals played by the first device and the second device.
5. The first device according to claim 4, characterized in that The controller is specifically configured to: performing low-latency compression processing on the first audio data frame based on first compression parameters to obtain the first audio data packet, wherein the first compression parameters include a first sampling rate and a first bit precision; The controller is further configured to: receiving first feedback information, where the first feedback information is used to indicate a real-time audio detection result and / or a network congestion analysis result obtained by the second device; If the first feedback information indicates that the data quality satisfies a first condition, performing low-latency compression processing on a second audio data frame in the fourth audio signal based on second compression parameters to obtain a second audio data packet, where the second compression parameters include a second sampling rate and a second bit precision, the second sampling rate being greater than the first sampling rate, and the second bit precision being greater than the first bit precision; When the first feedback information indicates that the data quality satisfies the second condition, low-latency compression processing is performed on the second audio data frame in the fourth audio signal based on a third compression parameter to obtain a third audio data packet. The third compression parameter includes a third sampling rate and a third bit precision, the third sampling rate is less than the first sampling rate, and the third bit precision is less than the first bit precision.
6. A second device, characterized in that: include: communicator; The controller is configured to: receive a target request message, the target request message being used to request channel parameters of the second device, the target request message being generated by the first device in response to an operation of playing a target audio signal, and when the first device is connected to the second device, the first device further acquiring first channel parameters in response to the operation of playing the target audio signal, the first channel parameters including a first channel and a first delay supported by the first device, the first delay being used to indicate a transmission delay corresponding to each channel in the first channel; Responding to the target request message, obtaining second channel parameters, where the second channel parameters include second channels and second delays supported by the second device, where the second delays are used to indicate a transmission delay corresponding to each channel in the second channel; sending a target response message corresponding to the target request message, the target response message carrying the second channel parameter, so that the first device determines, based on the first channel and the second channel, a target fusion scheme including a third channel in the first channel for playing the target audio signal and a fourth channel in the second channel for playing the target audio signal, obtains a first audio signal corresponding to the third channel and a second audio signal corresponding to the fourth channel from the target audio signal, and performs delay processing on the first audio signal and the second audio signal, respectively, based on the first delay and the second delay, to obtain a third audio signal and a fourth audio signal, so as to eliminate the difference in transmission delay between each channel in the third channel and the fourth channel; the first audio signal includes the audio signal corresponding to each channel in the third channel, and the second audio signal includes the audio signal corresponding to each channel in the fourth channel; receiving target information, the target information being generated based on the fourth channel and the fourth audio signal; In response to the target information, the fourth audio signal is played through the fourth channel, so that the third audio signal played by the first device through the third channel based on a third delay is played synchronously with the fourth audio signal, and the third delay is used to indicate the difference between the time when the fourth channel receives the fourth audio signal and the time when the third channel receives the third audio signal.
7. The second device according to claim 6, characterized in that The target information includes a first audio data packet and control information; the controller is specifically configured to: receiving a first audio data packet, where the first audio data packet is generated by performing low-latency encoding processing on a first audio data frame in the fourth audio signal based on first compression parameters, the first audio data packet carrying timestamp information, and the first compression parameters including a first sampling rate and a first bit precision; receiving the control information; sorting the received first audio data packets based on the timestamp information to synchronize audio signals played by the first device and the second device; Obtaining first feedback information, where the first feedback information is used to indicate a real-time audio detection result and / or a network congestion analysis result obtained by the second device; The first feedback information is sent so that the first device performs low-latency compression processing on the second audio data frame in the fourth audio signal based on the second compression parameter to obtain a second audio data packet when the first feedback information indicates that the data quality meets the first condition; and performs low-latency compression processing on the second audio data frame in the fourth audio signal based on the third compression parameter to obtain a third audio data packet when the first feedback information indicates that the data quality meets the second condition, where the second compression parameter includes a second sampling rate and a second bit precision, the second sampling rate is greater than the first sampling rate, and the second bit precision is greater than the first bit precision; and the third compression parameter includes a third sampling rate and a third bit precision, the third sampling rate is less than the first sampling rate, and the third bit precision is less than the first bit precision.
8. The second device according to claim 7, characterized in that The controller is further configured to: If the first feedback information indicates that the data quality meets the first condition, adjusting the receiving buffer from the first buffer to a second buffer, where the second buffer is smaller than the first buffer; When the first feedback information indicates that the data quality meets the second condition, the receiving buffer is adjusted from the first buffer to a third buffer, and the third buffer is larger than the first buffer.
9. A method for playing an audio signal, characterized in that: Applied to a first device, comprising: In response to an operation of playing a target audio signal, when the first device is connected to a second device, obtaining first channel parameters, where the first channel parameters include first channels supported by the first device and first delays, where the first delays indicate transmission delays corresponding to respective channels in the first channel; Sending a target request message, where the target request message is used to request a channel parameter of the second device; receiving a target response message corresponding to the target request message, the target response message carrying second channel parameters, where the second channel parameters are acquired by the second device in response to the target request message, the second channel parameters including second channels supported by the second device and second delays, where the second delays are used to indicate a transmission delay corresponding to each channel in the second channel; determining a target fusion scheme based on the first channel and the second channel, the target fusion scheme including a third channel in the first channel for playing the target audio signal and a fourth channel in the second channel for playing the target audio signal; acquiring, from the target audio signal, a first audio signal corresponding to the third channel and a second audio signal corresponding to the fourth channel, wherein the first audio signal includes audio signals corresponding to respective channels in the third channel, and the second audio signal includes audio signals corresponding to respective channels in the fourth channel; performing delay processing on the first audio signal and the second audio signal based on the first delay and the second delay to obtain a third audio signal and a fourth audio signal, so as to eliminate a difference in transmission delay between the third channel and the fourth channel; sending target information based on the fourth sound channel and the fourth audio signal, where the target information is used to instruct the second device to play the fourth audio signal through the fourth sound channel; The third audio signal is played through the third channel based on a third delay so that the third audio signal and the fourth audio signal are played synchronously, where the third delay indicates a difference between a time when the fourth channel receives the fourth audio signal and a time when the third channel receives the third audio signal.
10. A method for playing an audio signal, characterized in that: Applied to the second device, comprising: receiving a target request message for requesting channel parameters of the second device, the target request message being generated by the first device in response to an operation of playing a target audio signal and when the first device is connected to the second device, the first device further acquiring first channel parameters in response to the operation of playing the target audio signal, the first channel parameters including a first channel and a first delay supported by the first device, the first delay being used to indicate a transmission delay corresponding to each channel in the first channel; Responding to the target request message, obtaining second channel parameters, where the second channel parameters include second channels and second delays supported by the second device, where the second delays are used to indicate a transmission delay corresponding to each channel in the second channel; sending a target response message corresponding to the target request message, the target response message carrying the second channel parameter, so that the first device determines, based on the first channel and the second channel, a target fusion scheme including a third channel in the first channel for playing the target audio signal and a fourth channel in the second channel for playing the target audio signal, obtains a first audio signal corresponding to the third channel and a second audio signal corresponding to the fourth channel from the target audio signal, and performs delay processing on the first audio signal and the second audio signal, respectively, based on the first delay and the second delay, to obtain a third audio signal and a fourth audio signal, so as to eliminate the difference in transmission delay between each channel in the third channel and the fourth channel; the first audio signal includes the audio signal corresponding to each channel in the third channel, and the second audio signal includes the audio signal corresponding to each channel in the fourth channel; receiving target information, the target information being generated based on the fourth channel and the fourth audio signal; In response to the target information, the fourth audio signal is played through the fourth channel, so that the third audio signal played by the first device through the third channel based on a third delay is played synchronously with the fourth audio signal, and the third delay is used to indicate the difference between the time when the fourth channel receives the fourth audio signal and the time when the third channel receives the third audio signal.
Citation Information
Patent Citations
TV sound system
CN106792333A
Audio and video playing system and audio data playing method applied thereto
CN108012177A
Sound output apparatus and signal processing method thereof
CN110800319A
Display equipment and audio playing method
CN111757171A
Display device, external device and audio output method
CN115884061A