Audio playing method, electronic equipment, storage medium and computer program product
By setting up at least four real speakers in an electronic device and using algorithms to process audio signals to generate multi-channel audio signals, the problem of poor sound playback in the existing technology is solved, and cinema-level three-dimensional sound field playback is achieved to enhance the user's auditory experience without increasing hardware costs.
Patent Information
- Application Number
- CN202410345045.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-21
- Publication Date
- 2025-10-03
AI Technical Summary
When existing electronic devices play audio, the sound output is poor and they are unable to achieve a cinema-level three-dimensional sound field effect. In particular, it is difficult to improve the user's auditory experience without increasing hardware costs.
By setting up at least four real speakers in an electronic device and combining algorithms to process audio signals, a multi-channel audio signal with three-dimensional spatial characteristics is generated. Signal synthesis is performed based on the speaker layout, so that each speaker independently outputs the corresponding audio synthesis signal, thereby realizing three-dimensional sound field playback.
Without increasing hardware costs, it significantly improves the sound playback effect, enhances the user's auditory experience, and achieves cinema-level three-dimensional sound field playback effect.
Smart Images

Figure CN120751320A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of audio playback technology, and in particular to an audio playback method, electronic device, storage medium, and computer program product. Background Art
[0002] With the popularity of electronic devices such as smart tablets, most users use electronic devices to play audio and video on a daily basis. For example, the sounds in daily audio and video entertainment scenes are usually played through external speakers. Especially when using film and television apps or game apps, the quality of the external speaker experience directly affects the overall audio experience of the electronic device.
[0003] When playing audio, conventional electronic devices can only output the audio signal according to the position and direction of the speakers set in the electronic device, resulting in poor playback quality and affecting the user's listening experience. Specifically, with this playback method, users can usually only hear the sound coming from the position of the electronic device's speakers. For example, if the speakers are set on the left and right sides of the electronic device, the user will hear the sound coming from both sides of the electronic device, resulting in poor sound output and a poor listening experience. Summary of the Invention
[0004] The embodiments of the present application provide an audio playback method, an electronic device, a storage medium, and a computer program product, which can improve the sound playback effect when the electronic device plays audio, thereby enhancing the user's listening experience.
[0005] In a first aspect, a method for playing audio externally is provided. The method is applied to an electronic device equipped with at least four real speakers. The electronic device may mix audio data to generate a first height channel audio signal, a first front channel audio signal, a first rear channel audio signal, and a center channel audio signal. The electronic device processes the mixed audio signals to generate a multi-channel audio signal with three-dimensional spatial characteristics. Specifically, the electronic device generates a second front channel audio signal based on the first height channel audio signal and the first front channel audio signal, and performs sound field widening on the first rear channel audio signal to obtain a second rear channel audio signal. The electronic device may then synthesize the second front channel audio signal, the second rear channel audio signal, and the center channel audio signal based on the layout of the real speakers to generate audio synthesis signals corresponding to the real speakers of the electronic device. That is, the multi-channel audio signals with three-dimensional spatial characteristics are synthesized into a multi-channel audio synthesis signal that matches the layout of the real speakers, and each real speaker independently outputs a corresponding audio synthesis signal to achieve three-dimensional sound field reproduction.
[0006] In this solution, an algorithm processes audio signals to generate a multi-channel audio signal with three-dimensional spatial characteristics. This multi-channel audio signal with three-dimensional spatial characteristics is then synthesized based on the layout of at least four real speakers, so that each of the at least four real speakers independently outputs a corresponding audio synthesis signal. In other words, the algorithm enables the electronic device in this solution to achieve a three-dimensional sound field reproduction effect, improving the sound reproduction effect and enhancing the user's listening experience without incurring expensive hardware costs.
[0007] In one possible implementation of the first aspect, the electronic device first performs height rendering on the first height channel audio signal. The electronic device then matches and fuses the first front channel audio signal with the height-rendered second height channel audio signal to obtain a third front channel audio signal. The electronic device may then perform sound field widening on the third front channel audio signal to obtain a second front channel audio signal.
[0008] In this solution, the electronic device first matches and fuses the first front channel audio signal with the highly rendered second height channel audio signal and then widens it, so that the second height channel audio signal can also be widened, thereby generating a second front channel audio signal with a better sense of space.
[0009] In one possible implementation of the first aspect, the second height channel audio signal includes a second height left channel audio signal and a second height right channel audio signal; and the first front channel audio signal includes a first left channel audio signal and a first right channel audio signal. The electronic device fuses the second height left channel audio signal with the first left channel audio signal to obtain a third left channel audio signal; and the electronic device fuses the second height right channel audio signal with the first right channel audio signal to obtain a third right channel audio signal. It will be understood that in this implementation, the third left channel audio signal and the third right channel audio signal are included in the third front channel audio signal.
[0010] In this implementation, the first front channel audio signal and the first height channel audio signal are both subdivided into left and right channels. Then, based on the principle of same-side fusion, the height channel audio signal and the channel audio signal on the same side are fused, which can make the directionality of the audio signal more accurate and help improve the sense of space.
[0011] In one possible implementation of the first aspect, the first rear channel audio signal includes a first left surround channel audio signal and a first right surround channel audio signal. The electronic device performs sound field widening on the first left surround channel audio signal and the first right surround channel audio signal to obtain a second left surround channel audio signal and a second right surround channel audio signal.
[0012] In this implementation, the first rear channel audio signal is subdivided into two left and right channels to widen the sound field, which can make the directionality of the audio signal more accurate and help improve the sense of space.
[0013] In a possible implementation of the first aspect, the electronic device identifies a channel number of the audio data to obtain a channel number corresponding to the audio data. The electronic device mixes the audio data according to a mixing method that matches the identified channel number to generate a first height channel audio signal, a first front channel audio signal, a first rear channel audio signal, and a center channel audio signal.
[0014] In this implementation, the electronic device can adopt matching mixing methods for audio data with different numbers of channels, which not only improves the accuracy of the mixing processing, but also improves the applicability. That is, the solution proposed in this application can be adopted for audio data with any number of channels.
[0015] In a possible implementation of the first aspect, when the identified number of channels is dual channels, the electronic device can perform stereo upmixing processing on the audio data according to a first mixing processing method matching the dual channels to generate a first height channel audio signal, a first front channel audio signal, a first rear channel audio signal, and a center channel audio signal.
[0016] This implementation is applicable to dual-channel audio data, improving applicability. On the other hand, a matching mixing processing method can be adopted for dual-channel audio data, improving the accuracy of upmixing processing.
[0017] In a possible implementation of the first aspect, the electronic device performs height sound object separation on the audio data to obtain a first height channel audio signal; the electronic device performs human voice separation on the audio data to obtain a center channel audio signal; and the electronic device delays the first front channel audio signal in the audio data to obtain a first rear channel audio signal.
[0018] This implementation effectively upmixes the first height channel audio signal, the center channel audio signal, and the first rear channel audio signal based on height sound object separation, human voice separation, and signal delay processing. Furthermore, the height, center, and rear channel audio signals obtained through this processing are more critical audio information within the overall audio data, providing an effective basis for subsequently creating a three-dimensional sound field playback effect.
[0019] In a possible implementation of the first aspect, when the identified number of channels is multi-channel, the electronic device performs mixing processing on the audio data according to a second mixing processing method matching the multi-channel, and generates a first height channel audio signal, a first front channel audio signal, a first rear channel audio signal, and a center channel audio signal.
[0020] This implementation is applicable to any multi-channel audio data. Specifically, by performing mixing processing on the audio data using a mixing processing method that matches the multi-channels, three-dimensional sound field playback can be achieved for any multi-channel audio data, thus improving its applicability. Furthermore, a matching mixing processing method can be used for any multi-channel audio data, improving the accuracy of the mixing processing.
[0021] In a possible implementation of the first aspect, the second front channel audio signal includes a second left channel audio signal and a second right channel audio signal; the second rear channel audio signal includes a second left surround channel audio signal and a second right surround channel audio signal. The electronic device distributes the center channel audio signal to each real speaker to obtain a sub-center channel audio signal corresponding to each real speaker. For each real speaker, the electronic device can determine a target channel audio signal for output in the real speaker from the second left channel audio signal, the second right channel audio signal, the second left surround channel audio signal and the second right surround channel audio signal. Compared with other audio signals, the target channel audio signal is more suitable for playing in the real speaker. Then, the electronic device can synthesize the target channel audio signal and the sub-center channel audio signal corresponding to the real speaker to obtain an audio synthesis signal corresponding to the real speaker.
[0022] In this implementation, for each real speaker, the electronic device can select a target channel audio signal suitable for playing in the real speaker from the obtained channel audio signals with three-dimensional spatial characteristics. Then, the target channel audio signal and the sub-center channel audio signal corresponding to the real speaker are synthesized to obtain an audio synthesis signal corresponding to the real speaker. On the one hand, the audio signal played by the real speaker can be made more accurate and appropriate. On the other hand, a part of the center channel audio signal (i.e., the sub-center channel audio signal) is synthesized in the audio synthesis signal of each real speaker, so that the center sound can be distributed in the middle of multiple real speakers, achieving the effect of centering the subsequent center sound, for example, achieving the effect of centering the human voice.
[0023] In a possible implementation of the first aspect, the electronic device includes at least a first speaker, a second speaker, a third speaker, and a fourth speaker; the first speaker corresponds to a front left channel, the second speaker corresponds to a front right channel, the third speaker corresponds to a rear left channel, and the fourth speaker corresponds to a rear right channel;
[0024] For the first speaker, the electronic device synthesizes the second left channel audio signal and the sub-center channel audio signal corresponding to the first speaker to obtain an audio synthesis signal corresponding to the first speaker.
[0025] For the second speaker, the electronic device synthesizes the second right channel audio signal and the sub-center channel audio signal corresponding to the second speaker to obtain an audio synthesis signal corresponding to the second speaker.
[0026] For the third speaker, the electronic device synthesizes the second left surround channel audio signal and the sub-center channel audio signal corresponding to the third speaker to obtain an audio synthesis signal corresponding to the third speaker.
[0027] For the fourth speaker, the electronic device synthesizes the second right surround channel audio signal and the sub-center channel audio signal corresponding to the fourth speaker to obtain an audio synthesis signal corresponding to the fourth speaker.
[0028] In this implementation, for each real speaker, from among the obtained three-dimensional spatial characteristics of the various channel audio signals, a channel audio signal that matches the channel position corresponding to the real speaker is selected as the channel audio signal suitable for playback by that real speaker. This ensures that the audio signal played by the real speaker is more accurate and appropriate. Furthermore, the selected channel audio signal for each real speaker is synthesized with the sub-center channel audio signal corresponding to that real speaker, achieving the effect of centered playback of the center channel.
[0029] In a possible implementation of the first aspect, the electronic device may determine a center sound allocation weight corresponding to each real speaker; and based on the center sound allocation weight corresponding to each real speaker, the electronic device determines a sub-center channel audio signal corresponding to each real speaker.
[0030] In this implementation, each real speaker has a corresponding center sound distribution weight. Based on the center sound distribution weight, a sub-center channel audio signal is allocated to each real speaker, which can achieve accurate control of the center sound distribution, thereby facilitating the control of subsequent center sound playback.
[0031] In a possible implementation of the first aspect, the electronic device has an extended channel interface in its hardware abstraction layer; the extended channel interface corresponds one-to-one to a real speaker in the electronic device; and the electronic device transmits each audio synthesis signal to the corresponding real speaker for playback based on the extended channel interface in the hardware abstraction layer.
[0032] In this implementation, through the channel interface extended in the hardware abstraction layer, each of at least four real speakers can independently play one audio signal, breaking the limitation of traditional dual speakers and being more conducive to achieving the effect of three-dimensional sound field playback.
[0033] In one possible implementation of the first aspect, the mixing process, the generation of the second front channel audio signal, and the sound field widening and signal synthesis processing steps are performed in the application framework layer of the electronic device. The electronic device transmits the synthesized audio signals to the hardware abstraction layer through the application framework layer. The electronic device then transmits the synthesized audio signals to the audio driver in the core layer of the electronic device based on the expanded channel interfaces in the hardware abstraction layer. The audio driver then controls the real speakers in the electronic device to play the corresponding synthesized audio signals.
[0034] In this implementation, an improved algorithm is implemented in the application framework layer. Based on this improved algorithm, a synthesized audio signal with three-dimensional spatial characteristics is synthesized, corresponding one-to-one with each of the at least four real speakers. Furthermore, by combining this algorithm with the expanded channel interface in the hardware abstraction layer, each of the at least four real speakers can independently play a single audio signal, achieving a three-dimensional sound field reproduction effect.
[0035] In a second aspect, the present application provides an electronic device, which includes at least: at least four speakers, a memory, and one or more processors; the memory, at least four speakers are coupled to the processor; at least four speakers are used to play audio externally, and the memory stores computer program code, the computer program code includes computer instructions, and when one or more processors execute the computer instructions, the electronic device executes a method as described in any one of the first aspects above.
[0036] In a third aspect, the present application provides a computer-readable storage medium, which includes computer instructions. When the computer instructions are executed on an electronic device, the electronic device executes any one of the methods of the first aspect.
[0037] In a fourth aspect, the present application provides a computer program product, which, when executed on a computer, enables the computer to execute any one of the methods of the first aspect. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figure 1 A schematic diagram of the effect of a cinema-level sound field experience provided by an embodiment of the present application;
[0039] Figure 2 A schematic diagram of a traditional audio playback effect of an electronic device provided in an embodiment of the present application;
[0040] Figure 3 A software structure diagram of a tablet provided in an embodiment of the present application;
[0041] Figure 4 A schematic diagram of an electronic device achieving a three-dimensional sound field experience based on virtual speaker technology according to an embodiment of the present application;
[0042] Figure 5 A flowchart of an audio playback method provided in an embodiment of the present application;
[0043] Figure 6 A schematic diagram of the principle of stereo upmixing processing provided in an embodiment of the present application;
[0044] Figure 7 A schematic diagram of the principle of an audio playback method provided in an embodiment of the present application;
[0045] Figure 8 A schematic diagram of a sound field widening scenario provided in an embodiment of the present application;
[0046] Figure 9 A schematic diagram of the principle of sound field widening and audio signal fusion provided in an embodiment of the present application;
[0047] Figure 10 A hardware structure block diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0048] References to "one embodiment" or "some embodiments" etc. described in this specification mean that a particular feature, structure or characteristic described in conjunction with the embodiment is included in one or more embodiments of the present application. Thus, the phrases "in one embodiment", "in some embodiments", "in some other embodiments", "in some other embodiments", etc. appearing in different places in this specification do not necessarily refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized. The terms "including", "comprising", "having" and their variations all mean "including but not limited to", unless otherwise specifically emphasized. The term "connected" includes direct and indirect connections, unless otherwise stated.
[0049] In the embodiments of this application, words such as "exemplary" or "for example" are used to indicate examples, illustrations, or descriptions. Any embodiment or design described as "exemplary" or "for example" in the embodiments of this application should not be interpreted as being preferred or advantageous over other embodiments or designs. Rather, the use of words such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner.
[0050] Hereinafter, the terms "first," "second," "third," and "fourth" are used for descriptive purposes only and should not be understood to indicate or imply relative importance or implicitly specify the number of the technical features indicated. Thus, a feature identified as "first," "second," "third," and "fourth" may explicitly or implicitly include one or more of the features. In the description of this embodiment, unless otherwise specified, "plurality" means two or more.
[0051] An embodiment of the present application provides an audio playback method, which is applied to an electronic device, in which a plurality of real speakers are provided. The real speakers are real, non-virtual speakers in the electronic device. In an embodiment of the present application, even if the electronic device does not have cinema-level audio playback capabilities, in response to the user's operation of playing audio on the electronic device, the electronic device can also execute the solution provided by the embodiment of the present application, and replay the audio data in a three-dimensional sound field to achieve a cinema-level sound field playback effect, and provide the user with a cinema-level sound field experience without the need for other auxiliary equipment. The audio data can be a separate audio file, or audio data in a video file, or audio-visual data in an application or web page, etc., without limitation.
[0052] For example, taking the electronic device as a tablet (i.e., a tablet computer), in response to the operation of playing audio, the tablet can play audio data in a video platform or a music platform, or play audio and video files in a local area, thereby providing Figure 1 The cinema-level sound field experience shown in the figure does not require the purchase of expensive cinema-level audio playback equipment.
[0053] Typically, general electronic devices (i.e., those without cinema-grade audio playback capabilities) cannot achieve cinema-grade sound field effects when playing audio, because electronic devices can generally only output audio signals based on the position and direction of real speakers.
[0054] For example, electronic equipment Figure 2Taking the tablet 200 shown as an example, the tablet 200 does not have cinema-level audio playback capabilities. The tablet 200 is equipped with four real speakers (i.e., the first speaker S1, the second speaker S2, the third speaker S3, and the fourth speaker S4), which are distributed in pairs on the left and right sides of the tablet. Then, when the electronic device needs to play audio externally, the real speakers on the left and right sides will generally only output the sound from the left and right sides of the tablet 200. In this way, the sound will appear to be output from both sides of the tablet, and the sense of layering and spatiality of the sound output will be relatively poor, resulting in a poor listening experience.
[0055] For traditional methods, if you want to experience Figure 1 To achieve the cinema-level three-dimensional sound field effect shown above, you usually need to watch it in a dedicated cinema, or purchase professional home theater equipment to play the audio, which means the hardware cost is relatively high.
[0056] In order to enable general electronic devices to achieve better audio playback effects, some solutions have proposed improvements at the algorithm level for electronic devices with dual speakers, in an attempt to improve the sound playback effect of electronic devices. However, the audio link of the software system in such electronic devices usually does not support multi-channel audio signal output and is limited to outputting two-channel audio signals. In this way, even if an algorithm is used to generate a multi-channel audio signal in the intermediate process, in the end, due to the output limitations at the system level, only two-channel audio signals (i.e., left-channel audio signals and right-channel audio signals) can be provided to the dual speakers for playback. Obviously, the improvement in the sound playback effect is limited.
[0057] It is understood that electronic device software systems can typically adopt a layered architecture, event-driven architecture, microkernel architecture, microservice architecture, or cloud architecture. To facilitate understanding, the Android system with a layered architecture is used as an example to explain why the traditional audio link of the software system does not support the output of multi-channel audio signals.
[0058] In the HAL layer (hardware abstraction layer) of the traditional Android system, there are only two channel interfaces, the left channel interface and the right channel interface, so only left and right audio signals are supported. Therefore, regardless of whether the upper layer (for example, the application framework layer, i.e., the Framework layer) is a multi-channel audio signal, it will still be restricted by the two channel interfaces of the HAL layer. Therefore, before the application framework layer (Framework layer) inputs the audio signal to the HAL layer, it must comply with the interface regulations of the HAL layer, and ultimately only left and right audio signals can be input to the HAL layer. Then, it is transmitted to the audio driver through the left channel interface and the right channel interface of the HAL layer, and the audio driver controls the dual speakers (including the left speaker and the right speaker) to play the left and right audio signals. As a result, the sound playback effect is poor, and the sense of space and layering of the sound are still relatively poor.
[0059] At present, more and more electronic equipment products in the industry are equipped with 4 or even more real speakers, such as Figure 2 The tablet 200 (i.e., a tablet computer) shown typically has four real speakers, and some electronic devices even use eight real speakers. Obviously, if the audio link of the software system only supports outputting two-channel audio signals, it cannot meet the requirements of achieving three-dimensional sound field reproduction for electronic devices with four or more speakers.
[0060] In response to the above problems, the audio link supported by the software system is modified in the embodiments of the present application so that the audio link of the software system supports multi-channel audio signal output. That is to say, at the system level, for the audio link from the application framework layer (Framwork layer) to the bottom layer driving the real speaker, the audio link is modified from supporting dual-channel audio signal output to supporting multi-channel audio signal output. In this way, combined with the algorithms provided in the various embodiments of the present application, multi-channel audio signals can be output, rather than being limited to outputting only left and right dual-channel audio signals, thereby improving the spatial sense and layering of the sound field and achieving the effect of three-dimensional sound field playback.
[0061] Specifically, in an embodiment of the present application, channel expansion is performed based on the actual speaker layout of the electronic device, thereby modifying the audio link so that the audio link of the software system supports multi-channel audio signal output (that is, the audio link supports multi-channel paths), that is, the multi-channels supported by the modified audio link match the real speakers set in the electronic device, so that each real speaker supports an independent channel.
[0062] Based on the modified audio link supporting multi-channel paths described above, using the algorithms provided in the various embodiments of this application, the electronic device can process audio data to generate multi-channel audio signals for forming a three-dimensional sound field (i.e., multi-channel audio signals for playing on real speakers to form a three-dimensional sound field). It will be understood that the multi-channel audio signals for forming a three-dimensional sound field correspond one-to-one with the real speakers, and each real speaker independently outputs a corresponding audio signal to achieve the effect of three-dimensional sound field playback.
[0063] In some embodiments, using the algorithms provided in the embodiments of the present application, the electronic device can render the audio data into a multi-channel audio signal with three-dimensional spatial characteristics. For example, the electronic device can combine the technology of virtual speakers to implement processing such as sound field widening and height rendering of height sound objects to generate a multi-channel audio signal with three-dimensional spatial characteristics. Then, the electronic device can synthesize the rendered multi-channel audio signal into an audio synthesis signal corresponding to the real speaker one by one based on the layout of the real speakers in the electronic device. It can be understood that these audio synthesis signals are multi-channel audio signals used to form a three-dimensional sound field, so that each audio synthesis signal is input into the corresponding real speaker, and each real speaker independently outputs a corresponding audio synthesis signal to achieve three-dimensional sound field playback.
[0064] It should be noted that, since the audio signal input to the real speaker is synthesized from multi-channel audio signals with three-dimensional spatial characteristics, it is named an audio synthesis signal.
[0065] It can be understood that the audio synthesis signal also has three-dimensional spatial characteristics. Compared with the multi-channel audio signal with three-dimensional spatial characteristics generated by rendering, it realizes the virtual speaker of the center channel, which can make the human voice in the audio data emit from the center of the screen of the electronic device, achieving the effect of the center of the human voice.
[0066] In some embodiments, the modification to the audio link is reflected in the expansion of the channel interface in the hardware abstraction layer of the electronic device, and the expanded channel interface corresponds one-to-one with the real speakers in the electronic device. Based on the expanded channel interface in the hardware abstraction layer, the electronic device transmits each audio synthesis signal to the corresponding real speaker for playback. This makes it possible to output multiple audio signals without being limited to the original two channel interfaces, and each of the four real speakers independently outputs one audio signal (i.e., outputs the corresponding audio synthesis signal), which is conducive to forming a three-dimensional sound field.
[0067] In some embodiments, a series of processes for correlating audio data to generate a multi-channel audio synthesis signal are primarily performed in the application framework layer of the electronic device, i.e., an algorithm provided by the application framework layer generates a multi-channel audio signal (i.e., a multi-channel audio synthesis signal) for forming a three-dimensional sound field. The electronic device transmits the synthesized audio synthesis signals to the hardware abstraction layer through the application framework layer; based on the expanded channel interfaces in the hardware abstraction layer, the electronic device transmits the synthesized audio synthesis signals to the audio driver in the kernel layer of the electronic device, and controls the real speakers in the electronic device to play the corresponding audio synthesis signals through the audio driver.
[0068] For ease of understanding, let's take the electronic device as a tablet and the software system as the Android system as an example. Figure 3 The modified audio link in the embodiment of the present application is explained. Figure 3 This is a software structure diagram of the tablet provided in the embodiment of the present application.
[0069] It is understandable that the layered architecture of the Android system can divide the software into several layers, each with a clear role and division of labor. The layers communicate with each other through software interfaces. Figure 3 As shown, the Android system may include four layers, from top to bottom, namely, an application layer 310, an application framework layer (Framework layer) 320, a hardware abstract layer (HAL) layer 330, and a kernel layer (Kernel, also known as a driver layer) 340.
[0070] The application layer (Application) 310 may include a series of application packages. The application layer may include multiple application packages. The multiple application packages may be applications such as camera, gallery, calendar, call, map, navigation, WLAN, Bluetooth, music, video, short message and desktop launcher. For example, Figure 3 As shown, the application layer 310 may include an audio application 311 , which may play audio data.
[0071] The application framework layer 320 provides an application programming interface (API) and a programming framework for the applications in the application layer 310. The application framework layer 320 includes some predefined functions and algorithms.
[0072] In an embodiment of the present application, the audio application 311 can transmit audio data that needs to be played back to the application framework layer 320. The application framework layer 320 may include an algorithm library or algorithm package, which includes algorithm modules such as a multi-channel discrimination module 321, a stereo upmixing module 322, a multi-channel downmixing module 323, and a three-dimensional sound field playback module 324. By combining these algorithm modules to process the audio data input by the application layer 310, a multi-channel audio signal (i.e., a multi-channel audio synthesis signal) for forming a three-dimensional sound field can be generated.
[0073] It is understood that the number of channels of the multi-channel audio signal used to form the three-dimensional sound field matches the number of real speakers of the electronic device. For example, if the electronic device has four real speakers, the number of channels of the multi-channel audio signal used to form the three-dimensional sound field can be four. In other words, the application framework layer 320 can generate a four-channel audio synthesis signal through the three-dimensional sound field playback module 324.
[0074] To facilitate understanding of the algorithm processing flow involved in the application framework layer 320, Figure 3 A more detailed explanation is given.
[0075] like Figure 3 As shown, after the application layer 310 inputs audio data to the application framework layer 320 , the multi-channel identification module 321 in the application framework layer 320 can identify the number of channels corresponding to the audio data.
[0076] If it is a dual-channel audio signal, the multi-channel identification module 321 inputs the audio data into the stereo upmixing module 322 for upmixing processing to obtain 7 audio signals (i.e., a first left channel audio signal, a first right channel audio signal, a center channel audio signal, a first left surround channel audio signal, a first right surround channel audio signal, a first height left channel audio signal, and a first height right channel audio signal).
[0077] If the audio data is multi-channel, the multi-channel identification module 321 inputs the audio data to the multi-channel downmixing module 323 for downmixing, thereby obtaining the aforementioned 7-channel audio signals. It should be noted that the multi-channel audio data requiring downmixing is used as an example for illustration only, and is not limited to downmixing of multi-channel audio data. The specific mixing processing method can be determined based on the characteristics of the channels.
[0078] Then, the stereo upmixing module 322 or the multi-channel downmixing module 323 inputs the above 7-channel audio signals to the 3D sound field playback module 324. In the 3D sound field playback module 324, the sound field is widened and highly rendered to generate a multi-channel audio signal with 3D spatial characteristics, and the multi-channel audio signal with 3D spatial characteristics is merged to obtain the same as the 4 real speakers (i.e., Figure 3 The four-channel audio synthesis signal (ie, the first audio synthesis signal, the second audio synthesis signal, the third audio synthesis signal and the fourth audio synthesis signal) corresponding to the first speaker, the second speaker, the third speaker and the fourth speaker in the embodiment of the present invention.
[0079] The HAL layer 330 is used to connect the application framework layer 320 and the kernel layer 340. For example, the HAL layer 330 can transmit data (e.g., transparently) between the application framework layer 320 and the kernel layer 340. Of course, the HAL layer 330 can also process data from the bottom layer (i.e., the kernel layer 340) before transmitting it to the application framework layer 320. For example, the HAL layer 330 can convert hardware device parameters from the kernel layer 340 into a software programming language that can be recognized by the application framework layer 320 and the application layer 310.
[0080] In the audio playback scenario involved in the embodiments of the present application, the HAL layer 330 can manage the output of audio signals by real speakers based on notifications from upper layers (such as the application framework layer 320 and the application layer 310). For example, based on the audio data input by the application framework layer 320, the audio driver in the kernel layer is called to control the real speakers to play audio.
[0081] In some embodiments, a channel interface is added to the Audio HAL of the HAL layer 330 so that the number of the channel interfaces of the HAL layer 330 matches the number of real speakers provided in the electronic device.
[0082] like Figure 3 As shown, since the HAL layer 330 has two additional channel interfaces, together with the HAL layer 330's existing left and right channel interfaces, there are a total of four channel interfaces. Therefore, the number of channel interfaces in the HAL layer 330 matches the number of the tablet's four real speakers. Therefore, after the 3D sound field playback module 324 in the application framework layer 320 generates the four-channel audio composite signal used to form the 3D sound field, it no longer needs to be combined and compressed. In other words, instead of compressing and combining the four-channel audio signals into two channels as in traditional methods, the four-channel audio composite signal can be directly input into the HAL layer 330.
[0083] The kernel layer 340 is located below the HAL layer 330 and is a layer between hardware and software. The kernel layer 340 at least includes audio drivers, etc., and the present application embodiment does not impose any restrictions on this. For the audio playback scenario involved in the present application embodiment, the hardware can be a real speaker, for example, Figure 3 The first to fourth speakers are shown.
[0084] In this embodiment of the present application, the four-channel audio interface of the HAL layer 330 can transmit the four-channel audio synthesis signal to the audio driver in the kernel layer 340, which controls the four real speakers to play one of the audio synthesis signals. In other words, in this embodiment of the present application, an independent channel interface is implemented in the HAL layer 330 for each real speaker in the tablet, so that each real speaker supports an independent channel. In this way, the four real speakers can play the four-channel audio synthesis signal generated in the application framework layer 320 in a relatively complete and effective manner, greatly enhancing the layering and spatial sense of the sound playback and achieving a three-dimensional sound field playback effect.
[0085] It should be noted that the embodiments of the present application are not limited to electronic devices having only four real speakers; they may also have other numbers of speakers. Furthermore, the electronic device is not limited to being a tablet; it may also be other electronic devices. The embodiments of the present application are merely illustrative of a tablet with four speakers.
[0086] For example, it can also be a tablet with 8 speakers. Compared with 4 speakers, each of the 4 speakers is replaced by 2 speakers. These 2 speakers are responsible for outputting high-frequency signals and low-frequency signals respectively. Therefore, the audio synthesis signal corresponding to the original speaker (i.e., one speaker in the 4-speaker scenario) can be split into high-frequency signals and low-frequency signals, and the 2 replaced speakers output the high-frequency signals and low-frequency signals respectively. For example, taking the first speaker as an example, in the 8-speaker scenario, the first speaker can be replaced by 2 speakers. Then, the first audio synthesis signal input to the first speaker can be split into high-frequency signals and low-frequency signals, and then the high-frequency signals and low-frequency signals are respectively input to the 2 speakers that replace the first speaker. As mentioned above, the speakers involved in this application can also be other numbers or other layouts, which are no longer listed one by one.
[0087] As can be seen from the above example, on the basis of modifying the audio link to support the output of multi-channel audio signals, at the algorithm level, the embodiment of the present application mainly renders the audio data based on virtual speaker technology to generate a multi-channel audio signal with three-dimensional spatial characteristics that conforms to the preset virtual speaker layout. Then, based on the layout of the real speakers in the electronic device, the multi-channel audio signal that conforms to the virtual speaker layout and has three-dimensional spatial characteristics is synthesized into a multi-channel audio synthesis signal that matches the real speaker layout. Then, the multi-channel audio synthesis signal is output by the corresponding real speakers to achieve three-dimensional sound field playback.
[0088] For example, if the preset virtual speaker layout is that of a 5.1-channel system, five virtual speakers are required: the left channel, the right channel, the left surround channel, the right surround channel, and the center channel. Therefore, the electronic device can render and generate channel audio signals corresponding to these five virtual speakers. For example, if there are four real speakers, the electronic device can then synthesize the channel audio signals corresponding to these five virtual speakers into a four-channel audio composite signal that matches the four real speaker layout, with each real speaker outputting one of the audio composite signals.
[0089] In some embodiments, taking the speaker layout of a 5.1-channel system as an example, the electronic device uses virtual speaker technology to achieve the widening of the horizontal sound field and the height rendering of height sound objects for the audio data, so that the sound source position of the human voice part in the audio data is in the center of the screen of the electronic device, and the audio signals in the original left and right channels in the audio data are directly widened from the two sides of the screen of the original electronic device to the position of the virtual speaker, so that the user's experience of sound forms an enclosing nearly hemispherical surface, realizing the rendering of a three-dimensional sound field.
[0090] like Figure 4 The following is a schematic diagram of the speaker layout of a 5.1-channel system. Figure 4 S1 to S4 are real speakers set in the tablet 200. After the audio data is reproduced in the three-dimensional sound field such as sound field widening and height rendering, the listener 100 feels the sound field effect of the five virtual speakers S11, S22, S33, S44 and Sc. Figure 4 It can be seen that the virtual speaker forms a semi-enclosed nearly hemispherical surface, providing users with a three-dimensional sound field sound experience.
[0091] The specific method provided in the embodiments of the present application is described below with reference to the accompanying drawings.
[0092] The embodiment of the present application provides an audio playback method, which is applied to an electronic device equipped with at least 4 real speakers. The method is now described by taking the application of the method to a tablet equipped with 4 real speakers as an example. The 4 real speakers in the tablet are located on both sides of the tablet in pairs, and the real speakers on both sides are symmetrically distributed about the axis center line of the tablet. Figure 5 As shown, the method may include S501-S505.
[0093] S501: The tablet performs mixing processing on audio data to generate a first height channel audio signal, a first front channel audio signal, a first rear channel audio signal, and a center channel audio signal.
[0094] It can be understood that channels refer to independent audio signal paths that are collected or played back at different spatial locations during recording or playback. Channel audio signals refer to audio signals transmitted in the channels. Therefore, the first front channel audio signal refers to the audio signal output in the front channel. The first rear channel audio signal refers to the audio signal output in the rear channel. The first height channel audio signal refers to the audio signal output in the height channel. The center channel audio signal refers to the audio signal transmitted in the center channel. The first height channel audio signal refers to the audio signal output in the height channel.
[0095] Among them, the front channel is the sound output of the electronic device that simulates the user's auditory range.
[0096] The first front channel audio signal may include a first left channel audio signal and a first right channel audio signal. The first left channel audio signal refers to an audio signal transmitted in the left channel. The first right channel audio signal refers to an audio signal transmitted in the right channel.
[0097] The left channel, commonly referred to as the front left channel, is the sound output in electronic devices that simulates the hearing range of the user's left ear. The right channel, commonly referred to as the front right channel, is the sound output in electronic devices that simulates the hearing range of the user's right ear. In a multi-channel audio system, the left and right channels together constitute the main channels, primarily playing the main part of the audio signal, including background music, sound effects, and some vocals. The left and right channels can broadcast the same or different sounds, creating a stereo sound effect that shifts from left to right or right to left.
[0098] The center channel is primarily used to play delicate user conversations, typically located in the mid-frequency portion of the audio spectrum. Its primary purpose is to enhance the sense of positioning of sounds, making the listener perceive them as coming from the center of the screen.
[0099] Rear channels are usually used to simulate the effect of sound coming from behind the listener.
[0100] The first rear channel audio signal may include a first left surround channel audio signal and a first right surround channel audio signal. The first left surround channel audio signal refers to an audio signal transmitted in the left surround channel. The first right surround channel audio signal refers to an audio signal transmitted in the right surround channel.
[0101] The left and right surround channels are surround sound channels. The left surround channel, also known as the rear left channel, is typically used to simulate the effect of sound coming from behind the left side of the listener. The right surround channel, also known as the rear right channel, is typically used to simulate the effect of sound coming from behind the right side of the listener. Together, the left and right surround channels help the listener experience a surround sound effect.
[0102] In the embodiments of the present application, the audio data includes audio signals from high-altitude sound sources. High-altitude sound sources refer to sound sources that typically emit sound in the sky or vertically above the listener. For example, birds and airplanes can be considered high-altitude sound sources. "Above vertically above" refers to sound sources located vertically above the listener's sound perception site relative to the listener. It is understood that the sound perception site can be the ear. If the listener is a robot or other inanimate object, the sound perception site can also be a component with sound perception capabilities.
[0103] The first height channel audio signal is an audio signal of a height sound source in the audio data.
[0104] In some embodiments, the first height channel audio signal may include a first height left channel audio signal and a first height right channel audio signal. The first height left channel audio signal refers to an audio signal transmitted in the height left channel. The first height right channel audio signal refers to an audio signal transmitted in the height right channel.
[0105] The height left and right channels each form a height channel configuration. The height left channel is located above and to the left of the listener, simulating sounds coming from the upper left. The height right channel is located above and to the right of the listener, simulating sounds coming from the upper right. Together, they provide vertical coverage, creating a sense of spatial distribution and layering, enhancing the three-dimensionality and immersiveness of the sound.
[0106] In some embodiments, to generate a three-dimensional sound field that complies with a 5.1.2 virtual speaker layout, the tablet may mix audio data to generate seven audio signals: a first left channel audio signal, a first right channel audio signal, a center channel audio signal, a first left surround channel audio signal, a first right surround channel audio signal, a first height left channel audio signal, and a first height right channel audio signal. The first left channel audio signal, the first right channel audio signal, the center channel audio signal, the first left surround channel audio signal, and the first right surround channel audio signal correspond to the five virtual speakers, while the first height left channel audio signal and the first height right channel audio signal are used to generate overhead sound.
[0107] It can be understood that the above-mentioned seven-channel audio signals are the basic channel audio signals used to form a three-dimensional sound field. They are equivalent to having basic sound information, but they do not meet the spatial characteristic information of a three-dimensional sound field. Horizontal and vertical rendering are required based on the above-mentioned seven-channel audio signals to give them the corresponding spatial characteristic information, thereby forming a three-dimensional sound field with horizontal and vertical spatial characteristics. For example, the first height left channel audio signal and the first height right channel audio signal simply split the audio signal of a high-pitched sound object such as a bird's chirp into two channels. They also need to be given vertical spatial characteristics through height rendering to simulate and enhance the height perception of the high-pitched sound source. In other words, height rendering can simulate the effect of sound sources coming from different height positions, allowing the listener to more accurately determine the vertical position of the sound source. For example, the listener can hear the bird's chirp coming from above. Through height rendering, the sound emitted by the high-pitched sound source can be given a richer sense of space and three-dimensionality in the listener's hearing.
[0108] It should be emphasized that the embodiment of the present application is not limited to generating the above-mentioned 7-channel audio signals, and is merely illustrated by taking the generation of a three-dimensional sound field that complies with the 5.1.2 virtual speaker layout as an example.
[0109] It is understood that for audio data with different numbers of channels, the electronic device can adopt different mixing processing methods. Therefore, the electronic device can determine the mixing processing method that matches the audio data and perform mixing processing on the audio data according to the matching mixing processing method.
[0110] In some embodiments, step S501, i.e., the mixing process step, may include steps (A) to (B), specifically as follows:
[0111] (A) The tablet identifies the number of channels of the audio data and obtains the number of channels corresponding to the audio data.
[0112] Audio data is an audio stream, which can be 2.0 dual-channel or multi-channel. Multi-channel refers to 5.1 or higher channels, for example, 5.1, 5.1.2, 5.1.4, 7.1, and 7.1.2 are all multi-channel.
[0113] After decoding the audio data, the tablet can identify the number of channels in the audio data to determine whether it is dual-channel or multi-channel. For example, the tablet can use the multi-channel identification module in the application framework layer to identify the number of channels. Based on the different channel identification results, the audio data can be mixed using different mixing methods.
[0114] In some embodiments, the audio data carries a channel identifier that directly or indirectly indicates the number of channels in the audio data. Thus, the tablet can determine the number of channels corresponding to the audio data based on the channel identifier through a multi-channel identification module. For example, for dual-channel audio data, the channel identifier can be 2, indicating that the audio data is dual-channel. Alternatively, the channel identifier can be another string corresponding to dual channels, which can be used to identify that the audio data is dual-channel. The embodiments of this application do not limit the specific form of the channel identifier.
[0115] In other embodiments, the tablet may also perform channel separation and identification on the audio data to determine the number of channels corresponding to the audio data.
[0116] (B) The tablet mixes the audio data in a mixing method that matches the identified number of channels to generate a first height channel audio signal, a first front channel audio signal, a first rear channel audio signal, and a center channel audio signal.
[0117] For example, if the identified number of channels is two, the tablet determines that the matching mixing method is the first mixing processing method. The first mixing processing method is upmixing. Upmixing refers to mixing a small number of channel audio signals into multiple channels.
[0118] Next, we will introduce a solution for mixing two-channel audio data using the first mixing processing method:
[0119] In some embodiments, a stereo upmixing module is provided in the application framework layer of the Android system of the tablet. For two-channel audio data, the tablet can use the stereo upmixing module to mix the two-channel audio data into a first height channel audio signal, a first front channel audio signal, a first rear channel audio signal, and a center channel audio signal.
[0120] In some embodiments, the tablet can mix the two-channel audio data into a first left channel audio signal L, a first right channel audio signal R, a center channel audio signal C, a first left surround channel audio signal Ls, a first right surround channel audio signal Rs, a first height left channel audio signal TFL, and a first height right channel audio signal TFR through a stereo upmixing module.
[0121] It will be appreciated that the first left-channel audio signal L and the first right-channel audio signal R can be the original left and right-channel audio signals in the audio data, or they can be left and right-channel audio signals that have undergone pre-processing such as signal filtering or enhancement. The upmixing process also includes the steps of generating a center channel audio signal, generating a surround channel audio signal, and generating a height channel audio signal, which will be described below.
[0122] In the center channel audio signal generation process, the tablet can separate the human voice from the two-channel audio data to obtain the center channel audio signal C. Specifically, the tablet can separate the human voice from the two-channel audio data and enhance the separated human voice signal to obtain the center channel audio signal. In other words, the center channel audio signal is a human voice signal. Optionally, the tablet is provided with a stereo upmixing module for performing stereo upmixing on the audio data. Separating the center channel audio signal is one of the steps in the stereo upmixing process.
[0123] The tablet's stereo upmix module can use a signal separation method or a pre-trained vocal separation model to separate the vocals from the audio data. Furthermore, the tablet's stereo upmix module can also enhance the separated center vocals according to a preset ratio. It is understood that the vocal separation model can be trained based on a DNN (Deep Neural Networks) architecture or a CNN (Convolutional Neural Networks) architecture.
[0124] In some embodiments, the stereo upmixing module in the tablet can obtain the center channel audio signal by the following formula: C=gain*V(L+R).
[0125] Among them, C is the center channel audio signal, L and R are the first left channel audio signal and the first right channel audio signal in the dual channels of the audio data, V(L+R) represents the human voice signal separated based on the first left channel audio signal and the first right channel audio signal, and gain() represents the enhancement function.
[0126] During surround channel audio signal generation, the tablet can delay the first left channel audio signal in the two-channel audio data to generate a first left surround channel audio signal. Similarly, the tablet can delay the first right channel audio signal in the two-channel audio data to generate a first right surround channel audio signal. The delay processing can include filtering delay processing or phase difference control delay processing.
[0127] In some embodiments, regarding filtering and delay processing, the tablet can use a high-pass filter to filter out audio signals that fall within a preset low-frequency range, while retaining audio signals in a high-frequency range. For example, the preset low-frequency range can be below 2000 Hz, so audio signals below 2000 Hz can be filtered out. The tablet can then delay the retained audio signals, thereby generating a first left surround channel audio signal Ls and a first right surround channel audio signal Rs through filtering and delay processing.
[0128] Optionally, the tablet can determine a delay time based on the reflection time from the bottom real speaker to the tablet placement surface. The first left channel audio signal and the first right channel audio signal of the retained audio signal are delayed based on the delay time to generate a first left surround channel audio signal Ls and a first right surround channel audio signal Rs. The tablet placement surface is the support surface for the tablet. For example, if the tablet is placed on a table, the tabletop constitutes the tablet placement surface. In some embodiments, the delay time delay = 7ms - the reflection time t from the bottom real speaker to the tablet placement surface.
[0129] In some embodiments, regarding the phase difference control delay processing, the tablet can perform phase difference control on the first left channel audio signal and the first right channel audio signal to generate a first left surround channel audio signal and a first right surround channel audio signal.
[0130] During the height channel audio signal generation process, the tablet can separate the height sound object (i.e., the audio signal of the height sound source) from the audio data to obtain a first height channel audio signal, such as a first height left channel audio signal and a first height right channel audio signal. Alternatively, the tablet can separate the first height channel audio signal from the first left channel audio signal and the first right channel audio signal in the audio data. It is understood that audio signals such as the sound of an airplane flying or birds singing are all height sound objects.
[0131] Optionally, the tablet can input audio data into a pre-trained height object separation model to separate the height objects. Specifically, the height object separation model can be trained using audio data labeled with height information as sample data, thereby enabling the height object separation model to identify height objects from the audio data.
[0132] In some embodiments, visual information in the video can also be combined to assist in identifying high-pitched sound objects. For example, by detecting object motion or scene changes in the video, high-pitched sound sources can be identified and the sound height information of the high-pitched sound sources can be inferred to separate high-pitched sound objects.
[0133] Now combined Figure 6 The principle of stereo upmixing processing is briefly illustrated.
[0134] from Figure 6It can be seen that the first left channel audio signal L and the first right channel audio signal R in the audio data can be respectively input into the human voice separation module, the filter delay module, and the height sound separation module. The human voice separation module separates and extracts the center channel audio signal C based on the first left channel audio signal L and the first right channel audio signal R. The filter delay module performs filter delay processing on the first left channel audio signal L and the first right channel audio signal R, respectively, to obtain a first left surround channel audio signal Ls and a first right surround channel audio signal Rs. The height sound separation module performs height sound separation processing on the first left channel audio signal L and the first right channel audio signal R, to obtain a first height left channel audio signal TEL and a first height right channel audio signal TFR.
[0135] For example, if the identified number of channels is multi-channel, the tablet determines a second mixing processing method that matches the multi-channel. Specifically, if the number of channels is dual-channel, the mixing processing method can be directly determined as upmixing processing - that is, the first mixing processing method. If the number of channels is multi-channel, the second mixing method that matches the multi-channel can be determined based on the specific channel type and channel characteristics of the multi-channel. Different multi-channels can correspond to different second mixing processing methods. Therefore, the second mixing processing method can be downmixing processing, joint mixing processing, no mixing processing, or only upmixing processing.
[0136] Down-mixing refers to synthesizing a multi-channel audio signal into a small number of channels. Joint mixing processing includes both up-mixing and down-mixing, that is, combining up-mixing and down-mixing.
[0137] To facilitate understanding, the following describes solutions for mixing multi-channel audio data using the second mixing processing method in different situations: In the first situation, if the channel audio signals in the multi-channel audio data happen to match the channel audio signals to be generated, then the second mixing processing method corresponding to the audio data may not perform mixing processing.
[0138] Specifically, assume that the audio data needs to be mixed into 7-channel audio signals, namely, a first left channel audio signal, a first right channel audio signal, a center channel audio signal, a first left surround channel audio signal, a first right surround channel audio signal, a first height left channel audio signal, and a first height right channel audio signal. As mentioned above, multi-channel refers to 5.1 and above channels. For some multi-channel audio data, the channel audio signals it itself has may happen to be the above 7-channel audio signals. For example, 5.1.2 audio data itself has the above 7-channel audio signals.
[0139] Therefore, if the channel audio signals in the multi-channel audio data happen to be the seven-channel audio signals of the first left channel audio signal, the first right channel audio signal, the center channel audio signal, the first left surround channel audio signal, the first right surround channel audio signal, the first height left channel audio signal and the first height right channel audio signal, then the second mixing processing method corresponding to the audio data may be to not perform mixing processing, that is, the audio data may not be mixed, and the seven-channel audio signals may be directly used as input for subsequent processing, for example, as input signals for subsequent three-dimensional sound field playback processing.
[0140] In the second case, if the number of channels in the multi-channel audio data is greater than the number of channels of the channel audio signal to be generated, and the multi-channel audio data contains various channel audio signals to be generated, then the second mixing processing method corresponding to the audio data can be downmixing processing.
[0141] Similarly, taking the need to generate the above-mentioned 7-channel audio signals (first left channel audio signal, first right channel audio signal, center channel audio signal, first left surround channel audio signal, first right surround channel audio signal, first height left channel audio signal and first height right channel audio signal) as an example, if the number of channels in the multi-channel audio data is greater than the number of channels of the above-mentioned 7-channel audio signals, and the multi-channel audio data contains the above-mentioned 7-channel audio signals, then the second mixing processing method corresponding to the audio data can be downmixing processing, that is, the electronic device can use downmixing processing to merge or synthesize channel audio signals of the same type or on the same side to form 7-channel audio signals, namely, the first left channel audio signal, the first right channel audio signal, the center channel audio signal, the first left surround channel audio signal, the first right surround channel audio signal, the first height left channel audio signal and the first height right channel audio signal.
[0142] For example, if it is 7.1.2 multi-channel audio data, then the audio data includes a center left channel audio signal and a center right channel audio signal. The center left channel audio signal and the center right channel audio signal can be merged to generate a center channel audio signal C.
[0143] In a third case, if the multi-channel audio data lacks a height channel audio signal, then the second mixing processing method corresponding to the audio data needs to include an upmixing process.
[0144] Specifically, in some scenes, height channel audio signals will be lacking in multi-channel audio data. For example, if it is 5.1 or 7.1 multi-channel audio data, height channel audio signals will be lacking. So, the second mixing processing method corresponding to the audio data needs to include upmixing, that is, by upmixing, the audio data is processed to separate the height sound objects, so as to separate the height channel audio signals, such as, the first height left channel audio signal and the first height right channel audio signal. It will be understood that the processing of height sound object separation can refer to the relevant steps mentioned in the above-mentioned upmixing, such as, the front channel audio signals (such as the first left channel signal and the first right channel audio signal) in the audio data can be subjected to height sound object extraction to separate and extract the height channel audio signals (such as, the first height left channel audio signal and the first height right channel audio signal).
[0145] It should be noted that if the audio data lacks height channel audio signals and other channel audio signals that need to be generated, there are also redundant channel audio signals, then in addition to upmixing them to generate height channel audio signals, they also need to be downmixed to obtain the required channel audio signals. In this case, the second mixing processing method includes not only upmixing but also downmixing, that is, it belongs to a joint mixing method. If the audio data lacks height channel audio signals and the remaining are all required channel audio signals, then the second mixing processing method is to only perform upmixing.
[0146] It can be understood that after mixing and generating multiple audio signals (a first height channel audio signal, a first front channel audio signal, a first rear channel audio signal and a center channel audio signal) in S501, the above-mentioned multiple audio signals can be input into the three-dimensional sound field playback module for sound field widening, height rendering, signal merging or synthesis, etc., to generate a multi-channel audio synthesis signal that matches the real speaker, and each real speaker plays the corresponding audio synthesis signal.
[0147] For ease of understanding, now combined with Figure 7 The principle of the audio playback method is schematically explained.
[0148] like Figure 7As shown, after the audio data is input into the multi-channel identification module, if the multi-channel identification module identifies it as 2.0 dual-channel, it is input into the stereo upmixing module, which performs upmixing on the audio data to generate 7-channel audio signals, namely, a first left channel audio signal L, a first right channel audio signal R, a center channel audio signal C, a first left surround channel audio signal Ls, a first right surround channel audio signal Rs, a first height left channel audio signal TFL, and a first height right channel audio signal TFR. If it is identified as multi-channel, it is input into the multi-channel downmixing module, which performs downmixing on the audio data to generate the aforementioned 7-channel audio signals.
[0149] Then, the electronic device can input the above-mentioned 7-channel audio signals into the three-dimensional sound field playback module for sound field widening, height rendering, signal merging or synthesis, and generate audio synthesis signals corresponding to the four real speakers, namely, the first audio synthesis signal L_mix, the second audio synthesis signal R_mix, the third audio synthesis signal Ls_mix and the fourth audio synthesis signal Rs_mix, and the first speaker S1 plays the first audio synthesis signal L_mix, the second speaker S2 plays the second audio synthesis signal R_mix, the third speaker S3 plays the third audio synthesis signal Ls_mix, and the fourth speaker S4 plays the fourth audio synthesis signal Rs_mix.
[0150] In some embodiments, after the various audio signals obtained in step S501 are input into the three-dimensional sound field playback module, the three-dimensional sound field playback module may render the multi-channel audio signals to generate a multi-channel audio signal with three-dimensional spatial characteristics, and then synthesize or merge the processed multi-channel audio signals with three-dimensional spatial characteristics with reference to the layout of the real speakers in the electronic device to generate a multi-channel audio synthesis signal that matches the real speakers. Among them, the step of rendering the multi-channel audio signals to generate a multi-channel audio signal with three-dimensional spatial characteristics may include steps S502 to S504, and the step of synthesizing the multi-channel audio signals with three-dimensional spatial characteristics includes S505. Next, steps S502 to S504 and the processing involved in step S505 will be described one by one.
[0151] S502: The tablet generates a second front channel audio signal based on the first height channel audio signal and the first front channel audio signal; the horizontal sound field of the second front channel audio signal is wider than the horizontal sound field of the first front channel audio signal, and the height perception of the height sound source in the second front channel audio signal is higher than the height perception of the height sound source in the first height channel audio signal.
[0152] It can be understood that the horizontal sound field of the second front channel audio signal is wider than the horizontal sound field of the first front channel audio signal, and can simulate the sound effects in a larger spatial range in the horizontal direction or horizontal space, making the listener feel as if they are in a more open environment, which helps to enhance the layering and three-dimensional sense of the audio.
[0153] The second front channel audio signal is generated based on the first height channel audio signal and the first front channel audio signal. Therefore, the second front channel audio signal contains audio information related to height sound sources. Moreover, compared with the first height channel audio signal, the second front channel audio signal has a higher height perception of the height sound sources.
[0154] Among them, the height perception of high sound sources is used to characterize the difficulty of judging the vertical position of the sound source or the vertical space through hearing. The higher the height perception, the easier or more accurate it is to judge the vertical position of the sound source through hearing, so that the sound emitted by the high sound source presents a more obvious vertical spatial distribution and layering in the listener's hearing. Conversely, the lower the height perception, the more difficult it is to judge the vertical position of the sound source through hearing, so that the sound emitted by the high sound source presents a poorer sense of space in the listener's hearing, and may be mixed with other sounds, and the sense of layering is poorer.
[0155] In some embodiments, the height perception of height sound sources in the second front channel audio signal is enhanced through height rendering. Specifically, when generating the second front channel audio signal based on the first height channel audio signal and the first front channel audio signal, the tablet can perform height rendering on the first height channel audio signal, thereby enhancing the height perception of height sound sources in the resulting second front channel audio signal.
[0156] Next, we will introduce how to enhance the height perception of height sound sources in the second front channel audio signal through height rendering.
[0157] In some embodiments, the tablet can perform height rendering and sound field widening on the first height channel audio signal separately, and perform sound field widening on the first front channel audio signal separately, and then merge the height channel audio signal after height rendering and sound field widening with the first front channel audio signal after sound field widening to obtain a second front channel audio signal.
[0158] In other embodiments, the tablet may also perform height rendering on the first height channel audio signal to obtain a second height channel audio signal. The tablet may then match and fuse the height-rendered second height channel audio signal with the first front channel audio signal to obtain a third front channel audio signal. The tablet may perform sound field widening on the third front channel audio signal to obtain a second front channel audio signal. Sound field widening herein refers to horizontal sound field widening (also referred to as "horizontal sound field widening").
[0159] It can be understood that matching and fusing the first front channel audio signal with the highly rendered second height channel audio signal and then widening it can also widen the second height channel audio signal, thereby generating a second front channel audio signal with a better sense of space.
[0160] Specifically, the first height channel audio signal may include a first height left channel audio signal and a first height right channel audio signal, and the first front channel audio signal may include a first left channel audio signal and a first right channel audio signal. Then, step S502 may include the following steps (1) to (3), as follows:
[0161] (1) The tablet performs height rendering on the first height left channel audio signal and the first height right channel audio signal to obtain a height rendered second height left channel audio signal and a second height right channel audio signal.
[0162] That is, the height-rendered second height channel audio signal may include a second height left channel audio signal and a second height right channel audio signal.
[0163] It can be understood that the first height channel audio signal generated by mixing in step S501 refers to the audio signal of the height sound source, such as the sound of an airplane flying or the sound of birds singing. However, the first height channel audio signal generated by mixing in step S501 is generally used to represent sounds such as the sound of an airplane flying or the sound of birds singing. If you want the real speaker to output a sound with a greater sense of space and layering in the vertical space or vertical direction, you also need to perform height rendering on the first height channel audio signal. Specifically, the tablet can perform height rendering on the first height left channel audio signal and the first height right channel audio signal, and give the first height left channel audio signal and the first height right channel audio signal some spatial features in the vertical space to simulate the effect of sound coming from above in the vertical space, thereby improving the height perception of the height sound source. In this way, playing through a real speaker will make the user perceive that the audio signal of the height sound source such as the bird singing is transmitted from above.
[0164] In some embodiments, the tablet can convolve the height channel signal matrix using a first transfer function matrix to achieve height rendering, thereby obtaining a second height left channel audio signal and a second height right channel audio signal. The height channel signal matrix includes the first height left channel audio signal and the first height right channel audio signal. The first transfer function matrix is a matrix composed of transfer functions corresponding to the height channel audio signals.
[0165] Specifically, the first transfer function matrix may include a first height channel transfer function h topl and the second height channel transfer function h topr , the first height channel transfer function h topl , represents the transfer function of the first height left channel audio signal transmitted to the listener's left ear, and the second height channel transfer function h topr The transfer function representing the first height right channel audio signal transmitted to the listener's right ear.
[0166] It should be noted that transfer functions are used to characterize the transmission path of audio signals. Therefore, the transfer functions mentioned in each embodiment of this application are used to characterize the transmission path of the corresponding audio signal. For example, the first height channel transfer function is used to represent the transmission path of the first height left channel audio signal to the listener's left ear.
[0167] In some embodiments, the tablet may perform height rendering on the first height left channel audio signal TFL and the first height right channel audio signal TFR according to the following formula:
[0168]
[0169] Among them, [h topl h topr ] is the first transfer function matrix, It is the height channel signal matrix. TFL_render is the second height left channel audio signal obtained after height rendering, and TFR_render is the second height right channel audio signal obtained after height rendering.
[0170] (2) The tablet fuses the second height left channel audio signal with the first left channel audio signal to obtain a third left channel audio signal, and fuses the second height right channel audio signal with the first right channel audio signal to obtain a third right channel audio signal.
[0171] That is, the third front channel audio signal may include a third left channel audio signal and a third right channel audio signal, so the tablet performs sound field widening on the third front channel audio signal to obtain the second front channel audio signal, including the following step (3). That is, the second front channel audio signal includes a second left channel audio signal and a second right channel audio signal.
[0172] (3) Perform sound field expansion on the third left channel audio signal and the third right channel audio signal to generate a second left channel audio signal and a second right channel audio signal.
[0173] In the embodiment of the present application, sound field widening is equivalent to expanding the sound field width of a real speaker in an electronic device in the horizontal direction, which is also equivalent to widening the output position of the channel audio signal in the horizontal direction.
[0174] For ease of understanding, now combined with Figure 8 Provides a scene diagram for sound field expansion. Figure 8 The first speaker S1 to the fourth speaker S4 are four real speakers actually provided in the tablet 200. The first speaker S1 corresponds to the left channel, the second speaker S2 corresponds to the right channel, the third speaker S3 corresponds to the left surround channel, and the fourth speaker S4 corresponds to the right surround channel.
[0175] Before sound field widening is performed on the front channels, the first left channel audio signal and the first right channel audio signal are output only at the positions of the first speaker S1 and the second speaker S2, respectively. After sound field widening is performed on the front channels, the position of the real first speaker S1 is extended to the position of the first virtual speaker S11, and the position of the real second speaker S2 is extended to the position of the second virtual speaker S22. Therefore, after sound field widening, the output positions of the first left channel audio signal and the first right channel audio signal are equivalent to being extended to the positions of the first virtual speaker S11 and the second virtual speaker S22, respectively.
[0176] In some embodiments, the panel can widen the sound field by means of crosstalk cancellation.
[0177] For ease of understanding, now combined Figure 9 The sound field widening processing of the front channel audio signal is schematically explained.
[0178] See also Figure 9After height rendering the first height left channel audio signal TFL and the first height right channel audio signal TFR, a second height left channel audio signal TFL_render and a second height right channel audio signal TFR_render can be generated. They are input into the CTC (CrossTalk Cancellation) front channel processing module. Then, in the CTC front channel processing module, TFL_render is fused with the first left channel audio signal L, and TFR_render is fused with the first right channel audio signal R. Then, the fused left and right channel audio signals (i.e., the third left channel audio signal and the third right channel audio signal) are subjected to CTC crosstalk cancellation processing, and the second left channel audio signal L_CTC and the second right channel audio signal R_CTC after sound field widening are output.
[0179] Combine Figure 8 This means that the output position of the second left-channel audio signal L_CTC is extended to the position of the first virtual speaker S11 (the virtual speaker corresponding to L_CTC is the first virtual speaker S11), and the output position of the second right-channel audio signal R_CTC is extended to the position of the second virtual speaker S22 (the virtual speaker corresponding to R_CTC is the second virtual speaker S22). Therefore, when the second left-channel audio signal L_CTC and the second right-channel audio signal R_CTC are played, the corresponding sound signals are output at the positions of the first virtual speaker S11 and the second virtual speaker S22, respectively, achieving a widened horizontal sound field.
[0180] It is understood that the tablet can fuse the second height left channel audio signal TFL_render with the first left channel audio signal L to obtain a third left channel audio signal L', and fuse the second height right channel audio signal TFR_render with the first right channel audio signal R to obtain a third right channel audio signal R'. The tablet can form a front channel signal matrix based on the third left channel audio signal L' and the third right channel audio signal R'. The tablet can obtain a preset first crosstalk cancellation matrix and convolve the first crosstalk cancellation matrix with the front channel signal matrix to achieve sound field widening, thereby obtaining a second left channel audio signal L_CTC and a second right channel audio signal R_CTC.
[0181] Optionally, the first crosstalk cancellation matrix can be determined based on the inverse matrix of the second transfer function matrix. The second transfer function matrix is a matrix composed of transfer functions of left and right channel audio signals, including the first left channel transfer function h Ll , the second left channel transfer function h Lr , the first right channel transfer function h Rl and the second right channel transfer function h Rr. Among them, the first left channel transfer function h Ll represents the transfer function of the first left channel audio signal to the left ear of the listener; the second left channel transfer function h Lr The first right channel transfer function h represents the transfer function of the first left channel audio signal to the listener's right ear. Rl The transfer function of the first right channel audio signal to the left ear of the listener is represented by the transfer function of the second right channel h Rr The transfer function representing the first right channel audio signal transmitted to the listener's right ear.
[0182] In some embodiments, the sound field widening of the front channel audio signal can be achieved according to the following formula:
[0183]
[0184] in, is the first crosstalk cancellation matrix, is the second transfer function matrix, is a front channel signal matrix, L' is the third left channel audio signal, R' is the third right channel audio signal, L_CTC is the second left channel audio signal, and R_CTC is the second right channel audio signal.
[0185] S503: The tablet performs sound field widening on the first rear channel audio signal to obtain a second rear channel audio signal.
[0186] It can be understood that the width of the horizontal sound field of the second rear channel audio signal is greater than the width of the horizontal sound field of the first rear channel audio signal.
[0187] In some embodiments, the first rear channel audio signal includes a first left surround channel audio signal and a first right surround channel audio signal. Step S503 includes: the tablet performs sound field widening on the first left surround channel audio signal and the first right surround channel audio signal, respectively, to generate a second left surround channel audio signal and a second right surround channel audio signal. In other words, the second rear channel audio signal includes a second left surround channel audio signal and a second right surround channel audio signal.
[0188] For ease of understanding, also combined with Figure 8 Explain. Figure 8 Before the sound field is widened for the rear channel, the first left surround channel audio signal and the first right surround channel audio signal are output only at the positions where the third speaker S3 and the fourth speaker S4 are located.
[0189] After sound field widening is performed on the rear channels, the position of the real third speaker S3 is extended to the position of the third virtual speaker S33, and the position of the real fourth speaker S4 is extended to the position of the fourth virtual speaker S44. Therefore, the output positions of the audio signals of the first left surround channel and the first right surround channel after sound field widening are equivalent to being extended to the positions of the third virtual speaker S33 and the fourth virtual speaker S44, respectively.
[0190] For ease of understanding, also combined Figure 9 The sound field widening processing of the rear channel audio signal is schematically explained.
[0191] See also Figure 9 In the CTC rear channel processing module, the first left surround channel audio signal Ls and the first right surround channel audio signal Rs can be subjected to CTC crosstalk cancellation processing to generate a second left surround channel audio signal Ls_CTC and a second right surround channel audio signal Rs_CTC after the sound field is widened. Figure 8 This means that the output position of the second left surround channel audio signal Ls_CTC is extended to the location of the third virtual speaker S33 (the virtual speaker corresponding to Ls_CTC is the third virtual speaker S33), and the output position of the second right surround channel audio signal Rs_CTC is extended to the location of the fourth virtual speaker S44 (the virtual speaker corresponding to Rs_CTC is the fourth virtual speaker S44). Therefore, when the tablet outputs the second left surround channel audio signal Ls_CTC and the second right surround channel audio signal Rs_CTC, it can achieve the effect of outputting the corresponding sound signals at the locations of the third virtual speaker S33 and the fourth virtual speaker S44, thereby expanding the width of the horizontal sound field.
[0192] In some embodiments, when processing the sound field widening for the rear channels, the tablet can form a rear channel signal matrix based on the first left surround channel audio signal and the first right surround channel audio signal. The tablet can then obtain a preset second crosstalk cancellation matrix and convolve the second crosstalk cancellation matrix with the rear channel signal matrix to achieve sound field widening, thereby obtaining a second left surround channel audio signal and a second right surround channel audio signal.
[0193] Optionally, the second crosstalk cancellation matrix can be determined based on the inverse matrix of the third transfer function matrix. The third transfer function matrix is a matrix composed of transfer functions of left and right surround channel audio signals, including the first left surround channel transfer function h Lsl , the second left surround channel transfer function h Lsr , the first right surround channel transfer function h Rsl and the second right surround channel transfer function h Rsr. Among them, the first left surround channel transfer function h Lsl represents the transfer function of the first left surround channel audio signal transmitted to the listener's left ear; the second left surround channel transfer function h Lsr The transfer function of the first left surround channel audio signal to the listener's right ear is represented by the first right surround channel transfer function h Rsl The transfer function of the first right surround channel audio signal to the listener's left ear is represented by the transfer function of the second right surround channel audio signal h Rsr Represents the transfer function of the first right surround channel audio signal to the listener's right ear.
[0194] In some embodiments, the sound field widening of the rear channel audio signal can be achieved according to the following formula:
[0195]
[0196] in, is the second crosstalk cancellation matrix, is the third transfer function matrix, is the rear channel signal matrix, Ls is the first left surround channel audio signal, Rs is the first right surround channel audio signal, Ls_CTC is the second left surround channel audio signal, and Rs_CTC is the second right surround channel audio signal.
[0197] In other embodiments, the tablet can also achieve sound field widening through signal delay and superposition. Specifically, a delay time and delay strength can be determined, and the audio signal to be widened can be delayed. The audio signal to be widened can then be superimposed with the delayed audio signal to obtain the audio signal after sound field widening.
[0198] It is understood that the audio signal to be sound field widened may include a third left channel audio signal and a third right channel audio signal. That is, the third left channel audio signal and the third right channel audio signal are delayed, and the delayed left and right channel audio signals are respectively superimposed with the third left channel audio signal and the third right channel audio signal to obtain a second left channel audio signal and a second right channel audio signal after sound field widening.
[0199] The audio signals to be widened may also include a first left surround channel audio signal and a first right surround channel audio signal. For example, the first left surround channel audio signal and the first right surround channel audio signal may be delayed, and then the delayed left and right surround channel audio signals are superimposed with the first left surround channel audio signal and the first right surround channel audio signal, respectively, to obtain a second left surround channel audio signal and a second right surround channel audio signal after the sound field is widened.
[0200] It should be noted that, when the sound field is widened, a preset angle constraint condition needs to be satisfied so that a sound field effect with a sense of space in the horizontal direction can be formed after the sound field is widened.
[0201] Specifically, a first angle corresponds to the angle between the second left channel audio signal and the second right channel audio signal. A second angle corresponds to the angle between the second left surround channel audio signal and the second right surround channel audio signal. The preset angle constraint may include the first angle being smaller than the second angle. It will be appreciated that the vertices of the first angle and the second angle are both at the listener's location.
[0202] For ease of understanding, now combined with Figure 8 The first angle and the second angle are schematically explained. Figure 8 , α is the first angle, β is the second angle. The two sides of the first angle α are edge 1 formed from the position of the listener 100 to the position of the first virtual speaker S11, and edge 2 formed from the position of the listener 100 to the position of the second virtual speaker S22. The two sides of the second angle β are edge 3 formed from the position of the listener 100 to the position of the third virtual speaker S33, and edge 4 formed from the position of the listener 100 to the position of the fourth virtual speaker S44. Figure 8 As shown, the first angle α needs to be smaller than the second angle β.
[0203] Optionally, the first angle can be in the range of 60-80 degrees to avoid a hollow feeling in the front sound field, and the second angle can be in the range of 100-120 degrees to increase the enveloping feeling of the surround sound channel. This can make the entire three-dimensional sound field more profound and layered.
[0204] It should be emphasized that, as described above, in the embodiment of the present application, the front channel audio signals (for example, the third left channel audio signal and the third right channel audio signal) and the rear channel audio signals (for example, the first left surround channel audio signal and the first right surround channel audio signal) are separated for sound field widening, rather than being mixed together for sound field widening, which can make the sound and image layering feel stronger.
[0205] For example, assuming that there are two musical instruments, drums and guitar, in the audio signal, if the front and rear channel audio signals are mixed together for sound field widening, the drums and guitar sounds may overlap. By performing front and rear channel separation widening according to the method in the embodiment of the present application, the drums and guitar sounds can be separated, which may produce an auditory effect of the guitar sound in front and the drums in the back.
[0206] It can be understood that the second front channel audio signal and the second rear channel audio signal in the above steps are obtained through horizontal spatial expansion and vertical spatial height rendering, and have three-dimensional spatial characteristics to a certain extent. Therefore, in step S504, the center channel audio signal can be merged or synthesized with the above-mentioned audio signal with three-dimensional spatial characteristics to generate an audio signal for output and playback from a real speaker (i.e., the audio synthesis signal described below).
[0207] S504: The tablet synthesizes the second front channel audio signal, the second rear channel audio signal, and the center channel audio signal based on the layout of the real speakers to generate an audio synthesis signal corresponding to the real speakers of the electronic device.
[0208] In some embodiments, step S504 includes: the tablet generates an audio synthesis signal corresponding one-to-one to the real speakers of the tablet based on the center channel audio signal, the second left channel audio signal, the second right channel audio signal, the second left surround channel audio signal, and the second right surround channel audio signal.
[0209] It can be understood that the audio synthesis signal refers to the synthesized audio signal. If the tablet has N real speakers, then N channels of audio synthesis signals will be generated, so that each real speaker corresponds to an independent channel and is only used to play one audio signal.
[0210] In some embodiments, the center channel audio signal obtained in step S501 needs to be distributed to each real speaker of the tablet. That is, a certain amount of sub-center channel audio signal is allocated to each real speaker. Therefore, when synthesizing the audio synthesis signal corresponding to the real speaker, for each real speaker, a channel audio signal to be output to the real speaker is determined from the second left channel audio signal, the second right channel audio signal, the second left surround channel audio signal, and the second right surround channel audio signal. The determined channel audio signal is then synthesized with the sub-center channel audio signal allocated to the real speaker to obtain the audio synthesis signal corresponding to the real speaker.
[0211] In some embodiments, Figure 2 Taking the tablet 200 shown in FIG as an example, the tablet can determine the distribution weights of the center channel audio signal corresponding to each of the four real speakers (i.e., the first speaker S1 to the fourth speaker S4). Based on the distribution weights, the tablet can determine the sub-center channel audio signals to be distributed to the four real speakers. The volume of the sub-center channel audio signals distributed to the four real speakers is a portion of the total center channel audio signal.
[0212] It can be understood that in order to place the center vocals in the center of the screen to achieve a better external speaker effect, the audio signal of the center channel audio signal can be evenly distributed to the four real speakers according to the vector amplitude translation principle (VBAP), so that the distribution weights of the four real speakers are equal. For example, the distribution weights of the center channel audio signal corresponding to the four real speakers are a, b, c and d, respectively, where a=b=c=d.
[0213] In some embodiments, in order to balance the volume of the center vocals with the other channel audio signals, the tablet can perform energy normalization processing on the center channel audio signal to satisfy the expression: 1 = a 2 +b 2 +c 2 +d 2 , then, a=b=c=d=1 / 2. The center channel audio signal is allocated to the four real speakers respectively according to the allocated weights. In this way, after synthesizing the audio synthesis signals corresponding to the four real speakers based on the sub-center channel audio signals allocated to the four real speakers, the human voice volume expressed when the corresponding audio synthesis signals are played through the four real speakers (that is, the human voice volume after the three-dimensional sound field is reproduced) is the same as the original human voice volume. For example, the volume of the human voice in the original audio data is 10, and the volume of the human voice after the three-dimensional sound field is reproduced according to the above method is also 10, making the three-dimensional sound field reproduction closer to reality and more accurate.
[0214] In some embodiments, for each real speaker, the tablet can determine a target channel audio signal to be output to the real speaker from the second left channel audio signal, the second right channel audio signal, the second left surround channel audio signal, and the second right surround channel audio signal based on the real speaker's position in the electronic device or the relative position of the real speaker to other real speakers in the tablet. The determined target channel audio signal is then synthesized with the sub-center channel audio signal assigned to the real speaker to obtain an audio synthesis signal corresponding to the real speaker.
[0215] For ease of understanding, combined Figure 9 The principle of audio synthesis is schematically explained.
[0216] For the first speaker S1, since the first speaker S1 is located at the upper left of the tablet and corresponds to the left channel (i.e., the front left channel), it can be determined that the second left channel audio signal L_CTC is used to synthesize the channel audio signal output in the first speaker S1. Figure 8The second left channel audio signal L_CTC and the sub-center channel audio signal C_VBAP allocated to the first speaker S1 are synthesized to obtain a first audio synthesis signal L_mix corresponding to the first speaker S1.
[0217] For the second speaker S2, since the second speaker S2 is located at the upper right of the tablet and corresponds to the right channel (i.e., the front right channel), it can be determined that the second right channel audio signal R_CTC is used to synthesize the channel audio signal output in the second speaker S2. Figure 8 The second right channel audio signal R_CTC and the sub-center channel audio signal C_VBAP allocated to the second speaker S2 are synthesized to obtain a second audio synthesis signal R_mix corresponding to the second speaker S2.
[0218] For the third speaker S3, since the third speaker S3 is located at the lower left of the tablet and corresponds to the left surround channel (i.e., the rear left channel), it can be determined that the second left surround channel audio signal Ls_CTC after the sound field is widened (i.e., the rear left channel audio signal after the sound field is widened) is used to synthesize the channel audio signal output from the third speaker S3. Figure 8 The second left surround channel audio signal Ls_CTC and the sub center channel audio signal C_VBAP allocated to the third speaker S3 are synthesized to obtain a third audio synthesis signal Ls_mix corresponding to the third speaker S3.
[0219] For the fourth speaker S4, since the fourth speaker S4 is located at the lower right of the tablet and corresponds to the right surround channel (i.e., the rear right channel), it can be determined that the second right surround channel audio signal Rs_CTC after the sound field is widened (i.e., the rear right channel audio signal after the sound field is widened) is used to synthesize the channel audio signal output from the fourth speaker S4. Figure 8 The second right surround channel audio signal Rs_CTC and the center channel audio signal C_VBAP allocated to the fourth speaker S4 are synthesized to obtain a fourth audio synthesis signal Rs_mix corresponding to the fourth speaker S4.
[0220] S505 , the tablet controls each real speaker to play the corresponding audio synthesis signal.
[0221] Specifically, the first audio synthesis signal is output through the first speaker, the second audio synthesis signal is output through the second speaker, the second fused third audio synthesis signal is output through the third speaker, and the fourth audio synthesis signal is output through the fourth speaker to achieve a three-dimensional sound field playback effect.
[0222] As described above, when the center channel audio signal is evenly distributed to each real speaker, each real speaker plays the human voice at the same volume when playing its corresponding audio synthesis signal. Figure 8 As shown, it is equivalent to forming a virtual speaker Sc in the center of the tablet to play the human voice, realizing the center placement of the human voice.
[0223] For example, the electronic devices in the above embodiments may be tablet computers, televisions (also referred to as smart TVs, smart screens or large-screen devices), laptop computers, ultra-mobile personal computers (UMPCs), handheld computers, wearable electronic devices (for example, smart watches, smart bracelets, smart glasses), vehicle-mounted devices, virtual reality devices, and other electronic devices with audio playback functions. The embodiments of the present application do not impose any restrictions on this.
[0224] For ease of understanding, now combined Figure 10 The internal structure of the electronic device of each embodiment of the present application is described with examples. Figure 10 In this application, the electronic device 1000 is taken as an example of a tablet computer to introduce the hardware structure of the electronic device 1000 provided in this application.
[0225] like Figure 10 As shown, the electronic device 1000 may include: a processor 1010, an external memory interface 1020, an internal memory 1021, a universal serial bus (USB) interface 1030, a charging management module 1040, a power management module 1041, a battery 1042, an antenna 1, an antenna 2, a mobile communication module 1050, a wireless communication module 1060, an audio module 1070, a speaker 1070A, a receiver 1070B, a microphone 1070C, an earphone interface 1070D, a sensor module 1080, a button 1090, a motor 1091, an indicator 1092, a camera 1093, a display screen 1094, and a subscriber identification module (SIM) card interface 1095, etc.
[0226] It is understood that the speakers 1070A are the actual speakers mentioned in the various embodiments of this application. There are at least four speakers 1070A.
[0227] Among them, the above-mentioned sensor module may include sensors such as pressure sensor, gyroscope sensor, air pressure sensor, magnetic sensor, acceleration sensor, distance sensor, proximity light sensor, fingerprint sensor, temperature sensor, touch sensor, ambient light sensor and bone conduction sensor.
[0228] It should be understood that the structure illustrated in this embodiment does not constitute a specific limitation on the electronic device 1000. In other embodiments, the electronic device 1000 may include more or fewer components than shown, or may combine or separate certain components, or arrange the components differently. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.
[0229] The processor 1010 may include one or more processing units. For example, the processor 1010 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU). Different processing units may be independent devices or integrated into one or more processors.
[0230] The controller can be the nerve center and command center of the electronic device 1000. The controller can generate operation control signals according to instruction operation codes and timing signals to complete the control of instruction fetching and execution.
[0231] Processor 1010 may also include a memory for storing instructions and data. In some embodiments, the memory in processor 1010 is a cache memory. This memory can store instructions or data that have just been used or are being recycled by processor 1010. If processor 1010 needs to use the same instruction or data again, it can directly access the memory. This avoids duplicate accesses, reduces processor 1010 latency, and thus improves system efficiency.
[0232] It is understood that the interface connection relationship between the modules illustrated in this embodiment is merely an illustrative illustration and does not limit the structure of the electronic device 1000. In other embodiments, the electronic device 1000 may also adopt different interface connection methods from the above embodiments, or a combination of multiple interface connection methods.
[0233] The charging management module 1040 is configured to receive charging input from a charger. While charging the battery 1042, the charging management module 1040 can also power the electronic device through the power management module 1041. In other embodiments, the power management module 1041 and the charging management module 1040 can also be provided in the same device.
[0234] The wireless communication function of the electronic device 1000 can be implemented through the antenna 1, the antenna 2, the mobile communication module 1050, the wireless communication module 1060, the modem processor and the baseband processor.
[0235] Antenna 1 and Antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in electronic device 1000 can be used to cover a single or multiple communication frequency bands. Different antennas can also be reused to improve antenna utilization. For example, antenna 1 can be reused as a diversity antenna for a wireless local area network. In other embodiments, the antennas can be used in conjunction with a tuning switch.
[0236] The mobile communication module 1050 can provide solutions for wireless communications including 2G / 3G / 4G / 5G, etc. applied to the electronic device 1000. In the embodiment of the present application, the antenna 1 of the electronic device 1000 is coupled to the mobile communication module 1050, so that the electronic device 1000 can use the mobile cellular network for wireless communication.
[0237] The wireless communication module 1060 can provide wireless communication solutions for the electronic device 1000, including wireless local area networks (WLAN) (such as Wi-Fi networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), infrared (IR), etc. In the embodiment of the present application, the antenna 2 of the electronic device 1000 is coupled to the wireless communication module 1060, so that the electronic device 1000 can use the Wi-Fi network for wireless communication.
[0238] Electronic device 1000 implements display functionality through a GPU, display screen 1094, and an application processor. The GPU is a microprocessor for image processing that connects display screen 1094 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. Processor 1010 may include one or more GPUs that execute program instructions to generate or modify display information.
[0239] Electronic device 1000 can implement a camera function through an ISP, camera 1093, video codec, GPU, display 1094, and application processor. External memory interface 1020 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of electronic device 1000. The external memory card communicates with processor 1010 via external memory interface 1020 to implement data storage. For example, files such as music and videos can be stored on the external memory card.
[0240] The internal memory 1021 can be used to store computer executable program code, which includes instructions. The processor 1010 executes various functional applications and data processing of the electronic device 1000 by running the instructions stored in the internal memory 1021. For example, in an embodiment of the present application, the processor 1010 can execute instructions stored in the internal memory 1021, and the internal memory 1021 can include a program storage area and a data storage area.
[0241] The electronic device 1000 can implement audio functions such as music playback and recording through the audio module 1070, the speaker 1070A, the receiver 1070B, the microphone 1070C, the headphone jack 1070D, and the application processor.
[0242] The gyroscope sensor can be used to determine the motion posture of the electronic device 1000. The accelerometer can detect the magnitude of the acceleration of the electronic device 1000 in various directions (generally three axes). When the electronic device 1000 is stationary, it can detect the magnitude and direction of gravity. It can also be used to identify the electronic device's posture, which is applicable to applications such as switching between landscape and portrait modes and pedometers.
[0243] The buttons 1090 include a power button, a volume button, and the like. The buttons 1090 may be mechanical buttons. Alternatively, they may be touch buttons. The electronic device 1000 may receive button inputs and generate key signal inputs related to the user settings and function controls of the electronic device 1000. The motor 1091 may generate vibration prompts. The motor 1091 may be used for incoming call vibration prompts or for touch vibration feedback. The indicator 1092 may be an indicator light that may be used to indicate the charging status, power level changes, messages, missed calls, notifications, and the like. The SIM card interface 1095 is used to connect a SIM card. The SIM card interface 1095 may support NanoSIM cards, Micro SIM cards, SIM cards, and the like.
[0244] The methods in the following embodiments can all be implemented in the electronic device 1000 having the above hardware structure.
[0245] An embodiment of the present application also provides a computer-readable storage medium, which includes computer instructions. When the computer instructions are executed on the above-mentioned terminal device, the terminal device executes each function or step in the above-mentioned method embodiment.
[0246] The embodiment of the present application further provides a computer program product, which, when executed on a computer, enables the computer to execute each function or step in the above method embodiment.
[0247] Through the description of the above implementation methods, technical personnel in the relevant field can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional modules is used as an example. In actual applications, the above-mentioned functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.
[0248] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the modules or units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0249] The units described as separate components may or may not be physically separate, and the components shown as units may be one physical unit or multiple physical units, that is, they may be located in one place or distributed in multiple places. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0250] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0251] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solution of the embodiment of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for enabling a device (which can be a single-chip microcomputer, chip, etc.) or a processor (processor) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0252] The above content is only a specific embodiment of this application, but the scope of protection of this application is not limited to this. Any changes or replacements within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.
Claims
1. An audio playing method, characterized in that: Applied to an electronic device, wherein the electronic device is provided with at least four real speakers; the method comprises: The electronic device performs mixing processing on the audio data to generate a first height channel audio signal, a first front channel audio signal, a first rear channel audio signal, and a center channel audio signal; the first height channel audio signal is an audio signal of a height sound source in the audio data; The electronic device generates a second front channel audio signal based on the first height channel audio signal and the first front channel audio signal; the horizontal sound field of the second front channel audio signal is wider than the horizontal sound field of the first front channel audio signal, and the height perception of the height sound source in the second front channel audio signal is higher than the height perception of the height sound source in the first height channel audio signal; The electronic device performs sound field widening on the first rear channel audio signal to obtain a second rear channel audio signal; The electronic device synthesizes the second front channel audio signal, the second rear channel audio signal, and the center channel audio signal based on the layout of the real speakers to generate audio synthesis signals corresponding one-to-one to the real speakers of the electronic device; The electronic device controls each of the real speakers to play the corresponding audio synthesis signal.
2. The method according to claim 1, characterized in that The electronic device generates a second front channel audio signal according to the first height channel audio signal and the first front channel audio signal, comprising: The electronic device performs height rendering on the first height channel audio signal to obtain a second height channel audio signal; The electronic device matches and fuses the second height channel audio signal and the first front channel audio signal to obtain a third front channel audio signal; The electronic device performs sound field widening on the third front channel audio signal to obtain a second front channel audio signal.
3. The method according to claim 2, characterized in that The second height channel audio signal includes a second height left channel audio signal and a second height right channel audio signal; the first front channel audio signal includes a first left channel audio signal and a first right channel audio signal; the third front channel audio signal includes a third left channel audio signal and a third right channel audio signal; The electronic device matches and fuses the second height channel audio signal and the first front channel audio signal to obtain a third front channel audio signal, including: The electronic device fuses the second height left channel audio signal with the first left channel audio signal to obtain the third left channel audio signal; The electronic device merges the second height right channel audio signal with the first right channel audio signal to obtain the third right channel audio signal.
4. The method according to claim 1, wherein The first rear channel audio signal includes a first left surround channel audio signal and a first right surround channel audio signal; The electronic device performs sound field widening on the first rear channel audio signal to obtain a second rear channel audio signal, including: The electronic device performs sound field widening on the first left surround channel audio signal and the first right surround channel audio signal to obtain a second left surround channel audio signal and a second right surround channel audio signal.
5. The method according to claim 1, wherein The electronic device performs mixing processing on the audio data to generate a first height channel audio signal, a first front channel audio signal, a first rear channel audio signal, and a center channel audio signal, including: The electronic device identifies the number of channels of the audio data to obtain the number of channels corresponding to the audio data; The electronic device performs mixing processing on the audio data according to a mixing method matching the identified number of channels to generate a first height channel audio signal, a first front channel audio signal, a first rear channel audio signal, and a center channel audio signal.
6. The method according to claim 5, characterized in that The electronic device performs mixing processing on the audio data according to a mixing method matching the identified number of channels to generate a first height channel audio signal, a first front channel audio signal, a first rear channel audio signal, and a center channel audio signal, including: When the number of channels identified is dual channels, the electronic device performs stereo upmixing processing on the audio data according to a first mixing processing method matching the dual channels to generate a first height channel audio signal, a first front channel audio signal, a first rear channel audio signal, and a center channel audio signal.
7. The method according to claim 6, characterized in that The electronic device performs stereo upmixing processing on the audio data according to a first mixing processing mode matching the two channels to generate a first height channel audio signal, a first front channel audio signal, a first rear channel audio signal, and a center channel audio signal, including: The electronic device performs height sound object separation on the audio data to obtain a first height channel audio signal; The electronic device performs voice separation on the audio data to obtain a center channel audio signal; The electronic device delays the first front channel audio signal in the audio data to obtain a first rear channel audio signal.
8. The method according to claim 5, characterized in that The electronic device performs mixing processing on the audio data according to a mixing method matching the identified number of channels to generate a first height channel audio signal, a first front channel audio signal, a first rear channel audio signal, and a center channel audio signal, including: When the number of channels identified is multi-channel, the electronic device performs mixing processing on the audio data according to a second mixing processing method matching the multi-channel to generate a first height channel audio signal, a first front channel audio signal, a first rear channel audio signal and a center channel audio signal.
9. The method according to any one of claims 1 to 8, characterized in that The second front channel audio signal includes a second left channel audio signal and a second right channel audio signal; the second rear channel audio signal includes a second left surround channel audio signal and a second right surround channel audio signal; The electronic device synthesizes the second front channel audio signal, the second rear channel audio signal, and the center channel audio signal based on the layout of the real speakers to generate an audio synthesis signal corresponding to the real speakers of the electronic device, including: The electronic device distributes the center channel audio signal to each of the real speakers to obtain sub-center channel audio signals corresponding to each of the real speakers; For each real speaker, the electronic device determines a target channel audio signal for output in the real speaker from the second left channel audio signal, the second right channel audio signal, the second left surround channel audio signal, and the second right surround channel audio signal; and synthesizes the target channel audio signal with the sub-center channel audio signal corresponding to the real speaker to obtain an audio synthesis signal corresponding to the real speaker.
10. The method according to claim 9, characterized in that The electronic device includes at least a first speaker, a second speaker, a third speaker, and a fourth speaker; the first speaker corresponds to a front left channel, the second speaker corresponds to a front right channel, the third speaker corresponds to a rear left channel, and the fourth speaker corresponds to a rear right channel; For each real speaker, the electronic device determines, from the second left channel audio signal, the second right channel audio signal, the second left surround channel audio signal, and the second right surround channel audio signal, a target channel audio signal for outputting to the real speaker; The method further comprises: synthesizing the target channel audio signal and the sub-center channel audio signal corresponding to the real speaker to obtain an audio synthesis signal corresponding to the real speaker, comprising: For the first speaker, synthesize the second left channel audio signal and the sub-center channel audio signal corresponding to the first speaker to obtain an audio synthesis signal corresponding to the first speaker; For the second speaker, synthesize the second right channel audio signal and the sub-center channel audio signal corresponding to the second speaker to obtain an audio synthesized signal corresponding to the second speaker; For the third speaker, synthesize the second left surround channel audio signal and the sub-center channel audio signal corresponding to the third speaker to obtain an audio synthesized signal corresponding to the third speaker; For the fourth speaker, the second right surround channel audio signal and the sub-center channel audio signal corresponding to the fourth speaker are synthesized to obtain an audio synthesis signal corresponding to the fourth speaker.
11. The method according to claim 9, characterized in that The electronic device distributes the center channel audio signal to each of the real speakers to obtain sub-center channel audio signals corresponding to each of the real speakers, including: The electronic device determines the center sound distribution weight corresponding to each of the real speakers; The electronic device determines the sub-center channel audio signals corresponding to the real speakers based on the center sound distribution weights corresponding to the real speakers.
12. The method according to any one of claims 1 to 11, characterized in that The hardware abstraction layer of the electronic device has an extended channel interface; the extended channel interface corresponds one-to-one with the real speaker in the electronic device; The electronic device controls each of the real speakers to play a corresponding audio synthesis signal, including: The electronic device transmits each of the audio synthesis signals to a corresponding real speaker for playback based on the expanded channel interface in the hardware abstraction layer.
13. The method according to claim 12, characterized in that The steps of mixing, generating the second front channel audio signal, and processing the sound field widening and signal synthesis are performed in the application framework layer of the electronic device; The electronic device transmits each of the audio synthesis signals to a corresponding real speaker for playback based on the expanded channel interface in the hardware abstraction layer, including: The electronic device transmits the synthesized audio signals of each channel to the hardware abstraction layer through the application framework layer; The electronic device transmits each channel of the audio synthesis signal to the audio driver in the kernel layer of the electronic device based on the expanded channel interfaces in the hardware abstraction layer, and controls the real speakers in the electronic device to play the corresponding audio synthesis signal through the audio driver.
14. An electronic device, characterized in that: The electronic device includes a memory, at least four speakers, and one or more processors; the memory and the at least four speakers are coupled to the processor; wherein the at least four speakers are used to play audio externally, and the memory stores computer program code, the computer program code including computer instructions, and when the computer instructions are executed by the processor, the electronic device executes the method according to any one of claims 1 to 13.
15. A computer-readable storage medium, characterized in that The method comprises computer instructions, which, when executed on an electronic device, cause the electronic device to execute the method according to any one of claims 1 to 13.
16. A computer program product, characterized in that When the computer program product is run on a computer, the computer is caused to perform the method according to any one of claims 1 to 13.