Audio surround sound playback method, apparatus, and electronic device
The audio surround playback method allows for customized surround effects for different audio tracks by isolating tracks, selecting playback modules, and setting energy coefficients, thereby improving user experience through personalized surround sound playback.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-09-29
- Publication Date
- 2026-04-09
AI Technical Summary
Conventional surround stereo playback systems provide a constant surround effect, making it impossible to individually set different surround playback effects for different audio tracks, which negatively impacts the user experience.
An audio surround playback method that isolates individual audio tracks, determines a surround playback effect for each track, selects corresponding playback modules, sets energy coefficients, and synthesizes target channel data to achieve customized surround effects for each track.
Enables individual customization of surround playback effects for different audio tracks, enhancing the user experience by providing tailored surround sound experiences for various audio content.
Smart Images

Figure 2026062568000001_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of audio playback, and particularly to an audio surround playback method, apparatus, and electronic device.
Background Art
[0002] In the field of audio playback, surround stereo enhances the sense of depth, presence, and spatiality of sound by different channel playback devices at different positions, so that the viewer is surrounded by the spatial sound field generated by these sound sources, creating acoustic effects as if in a karaoke bar or theater. Taking 5.1-channel surround stereo as an example, to play audio signals from the center channel, the left and right front channels, the left and right rear surround channels, and the subwoofer channel (i.e., the 0.1 channel), usually six playback devices are required.
[0003] However, the surround effect realized by the conventional surround stereo playback system is constant, and it is impossible to individually set different surround playback effects for different audio tracks, such as concentrating the drum sound of some audio in the front, which affects the user experience.
Summary of the Invention
Problems to be Solved by the Invention
[0004] This application provides an audio surround playback method, apparatus, and electronic device to solve the technical problem that the surround effect realized by the conventional surround stereo playback system is constant, and different surround playback effects cannot be individually set for different audio tracks, which affects the user experience.
Means for Solving the Problems
[0005] In the first aspect, the audio surround playback method relating to the present application is: The steps include acquiring the audio to be played back and determining a plurality of first playback modules corresponding to the audio to be played back, The steps include determining at least one audio track in the audio to be played back and audio track data corresponding to the audio track, The steps include determining the surround sound playback effect for any audio track and determining at least one second playback module corresponding to the audio track from a plurality of first playback modules, The steps include determining the energy coefficient of the audio track data corresponding to each of the second playback modules based on the surround sound playback effect, A step of synthesizing target channel data for any of the first playback modules based on an energy coefficient corresponding to the first playback module, which includes the energy coefficient of the audio track data corresponding to the second playback module, the audio track data corresponding to the energy coefficient, and the channel data corresponding to the first playback module. The process includes the step of controlling each of the first playback modules to play the corresponding target channel data in order to play the audio to be played back.
[0006] In one possible embodiment, the step of determining the surround playback effect of the audio track is: The steps include determining whether the aforementioned audio track is set to surround sound playback, If it is determined that the audio track is set to surround sound playback, the step is to determine whether or not a surround sound playback position has been set for the audio track. If it is determined that a surround playback position exists for the audio track, the steps include determining the surround playback position as the surround playback effect of the audio track, If it is determined that no surround playback position exists for the audio track, the step of determining the position where the pre-set first playback module exists as the surround playback effect for the audio track is included.
[0007] As one possible embodiment, A method for outputting a scene graph showing the positional relationship of multiple first playback modules via a visualization interface, A method for determining the initial surround playback position of the audio track in the positional relationship scene graph in response to setting operations on the positional relationship scene graph, The surround playback position of the audio track is set by a method that determines the actual position represented by the initial surround playback position as the surround playback position of the audio track.
[0008] In one possible embodiment, the surround playback position is located between planes in which a symmetrical set of playback modules exists, or is located on a first playback module in the set of playback modules, each of which includes a first playback module corresponding to the left channel and a first playback module corresponding to the right channel, and the surround playback position is a position corresponding to the left channel audio track and the right channel audio track, respectively, included in the audio track. The step of determining at least one second playback module corresponding to the audio track from a plurality of first playback modules is: The steps include determining whether the first playback module includes a Sky playback module that corresponds to Sky Channel, If it is determined that the first regeneration module does not include the Sky regeneration module, the first regeneration module included in the regeneration module set is determined to be the second regeneration module. If it is determined that the first regeneration module includes the Sky regeneration module, the process includes the step of determining the first regeneration module and the Sky regeneration module included in the regeneration module set as the second regeneration module.
[0009] In one possible embodiment, the step of determining the energy coefficient of the audio track data corresponding to each of the second playback modules based on the surround playback effect is: A step of determining the total distance between two opposing sets of playback modules where the surround playback position exists, with respect to the second playback module in the playback module set, For each of the second playback modules, the steps include determining a first distance between the surround playback position and the plane on which the second playback module is located, A step of determining a first distance ratio between the surround playback position and the second playback module based on the first distance and the total distance, A step of obtaining a first energy coefficient by subtracting the first distance ratio from a predetermined value, The process includes the step of determining the energy coefficient of the audio track data corresponding to the second playback module based on the first energy coefficient.
[0010] In one possible embodiment, the step of determining the energy coefficient of the audio track data corresponding to the second playback module based on the first energy coefficient is: The steps include determining whether the Sky Regeneration Module is present in the second regeneration module, If it is determined that the Sky playback module does not exist, the first energy coefficient is determined as the energy coefficient of the audio track data corresponding to the second playback module. If it is determined that the sky regeneration module exists, the steps include determining the vertical distance between the plane on which the sky regeneration module exists and a predetermined plane, The steps include determining a second distance between the surround playback position and the plane on which the sky playback module is located, A step of determining a second distance ratio between the surround playback position and the sky playback module based on the second distance and the vertical distance, A step of obtaining a second energy coefficient by subtracting the second distance ratio from a predetermined value, The method includes the step of determining the first energy coefficient and the second energy coefficient as the energy coefficients of the audio track data corresponding to the second playback module.
[0011] In one possible embodiment, the step of synthesizing target channel data based on an energy coefficient corresponding to the first playback module, audio track data corresponding to the energy coefficient, and channel data corresponding to the first playback module is: The steps include determining the channel data corresponding to the first playback module, For each audio track corresponding to the first playback module, the first audio track data is obtained by multiplying the energy coefficient corresponding to the audio track by the audio track data corresponding to the energy coefficient. The steps include inputting the first audio track data into a filter corresponding to the first playback module to obtain the second audio track data, The method includes the step of synthesizing the channel data, the second audio track data, and audio track data corresponding to a pre-set audio track to obtain target channel data.
[0012] In one possible embodiment, the step of determining the channel data corresponding to the first playback module is: The steps include determining the number of channels in the audio to be played and the number of modules in the first playback module, Comparing the number of channels and the number of modules and obtaining a comparison result; converting the audio to be played based on the comparison result and obtaining channel data corresponding to the first playback module.
[0013] In one possible embodiment, the step of converting the audio to be played based on the comparison result and obtaining channel data corresponding to the first playback module includes: when the comparison result indicates that the number of channels is greater than the number of modules, calling a pre-trained downmix model to convert the audio to be played and obtaining channel data corresponding to the first playback module; when the comparison result indicates that the number of channels is equal to the number of modules, determining channel data corresponding to the first playback module from a plurality of channel data included in the audio to be played; when the comparison result indicates that the number of channels is less than the number of modules, calling a pre-trained upmix model to convert the audio to be played and obtaining channel data corresponding to the first playback module.
[0014] In one possible embodiment, the step of controlling each of the first playback modules to play corresponding target channel data includes: determining, from the second playback module, a third playback module whose energy coefficient is smaller than a preset coefficient threshold; controlling the first playback modules other than the third playback module among the first playback modules to play corresponding target channel data; after a preset time has elapsed, controlling the third playback module to play corresponding target channel data.
[0015] In a second aspect, the audio surround playback device according to the present application acquires audio to be played back, and a first determination module that determines a plurality of first playback modules corresponding to the audio to be played back; a second determination module that determines at least one audio track in the audio to be played back and audio track data corresponding to the audio track; for any audio track, a third determination module that determines a surround playback effect of the audio track and determines at least one second playback module corresponding to the audio track from the plurality of first playback modules; a fourth determination module that determines an energy coefficient of the audio track data corresponding to each second playback module based on the surround playback effect; for any of the first playback modules, based on the energy coefficient corresponding to the first playback module including the energy coefficient of the audio track data corresponding to the second playback module, the audio track data corresponding to the energy coefficient, and the channel data corresponding to the first playback module, a synthesis module that synthesizes target channel data; and a control module that controls each first playback module to play back corresponding target channel data in order to play back the audio to be played back.
[0016] In a third aspect, the electronic device according to the present application includes a processor and a memory. When the processor executes an audio surround playback program stored in the memory, the audio surround playback method according to any one of the first aspects is realized.
[0017] The technical means according to the embodiment of the present application is realized by acquiring audio to be played back, determining a plurality of first playback modules corresponding to the audio to be played back, determining at least one audio track in the audio to be played back and audio track data corresponding to the audio track, determining the surround playback effect of any of the audio tracks, determining at least one second playback module corresponding to the audio track from the plurality of first playback modules, determining the energy coefficient of the audio track data corresponding to each second playback module based on the surround playback effect, synthesizing target channel data for any of the first playback modules based on the energy coefficient corresponding to the first playback module, which includes the energy coefficient of the audio track data corresponding to the second playback module, the audio track data corresponding to the energy coefficient, and the channel data corresponding to the first playback module, and controlling each first playback module to play the corresponding target channel data in order to play back the audio to be played back. In this technology, at least one audio track included in the audio to be played back is first separated, then a surround playback effect to be achieved is individually set for that audio track, and a playback module is determined to achieve that surround playback effect. Based on the surround playback effect, different energy coefficients for the corresponding audio track data are set for the playback module. Finally, when playing back the audio to be played back, the audio track data, energy coefficients, and channel data determined by each playback module can be combined as target channel data. In this way, by controlling each playback module to play back the corresponding target channel data, not only is playback of the audio to be played back achieved, but the surround playback effect of at least one audio track is also achieved. During the process of playing back audio, different surround playback effects can be individually customized for different audio tracks included in the audio, thereby improving the user experience.
[0018] The drawings herein are incorporated into the specification and constitute part of this specification, illustrating the principles of the present invention together with the specification, showing embodiments consistent with the invention.
[0019] To more clearly describe the embodiments of the present invention or the technical means in the prior art, the drawings necessary for describing the embodiments or the prior art will be briefly described below, and obviously, those skilled in the art can obtain other drawings based on these drawings without any creative work.
[0020] One or more embodiments are illustrated by the corresponding drawings, and these illustrative descriptions are not limiting to embodiments. Elements having the same reference numerals in the drawings represent similar elements, and unless otherwise specified, the drawings do not constitute a proportional limitation. [Brief explanation of the drawing]
[0021] [Figure 1] This is a schematic diagram illustrating an application scenario for the audio surround playback method according to an embodiment of the present invention. [Figure 2] This is a schematic diagram illustrating an application scenario for another audio surround playback method according to an embodiment of the present invention. [Figure 3] This is a schematic diagram illustrating an application scenario for another audio surround playback method according to an embodiment of the present invention. [Figure 4] This is a flowchart of an embodiment of the audio surround playback method according to an embodiment of the present invention. [Figure 5] This is a flowchart of an embodiment for determining the channel data of the first regeneration module according to the embodiment of the present application. [Figure 6] This is a flowchart for determining the target channel data according to the embodiment of the present invention. [Figure 7] This is a flowchart of another embodiment of the audio surround playback method according to the embodiment of the present invention. [Figure 8] This is a schematic diagram of the two-dimensional planar structure of the regeneration module according to an embodiment of the present invention. [Figure 9A] This is a schematic diagram of the initial surround sound playback position according to an embodiment of the present invention. [Figure 9B] This is a schematic diagram of another initial surround sound playback position according to an embodiment of the present application. [Figure 9C] This is a schematic diagram of another initial surround sound playback position according to an embodiment of the present application. [Figure 9D] This is a schematic diagram of a further initial surround sound playback position according to an embodiment of the present invention. [Figure 10] This is a flowchart of another embodiment of the audio surround playback method according to the embodiment of the present invention. [Figure 11] This is a schematic diagram of the surround sound playback position according to an embodiment of the present application. [Figure 12] This is a block diagram of an embodiment of an audio surround playback device according to an embodiment of the present application. [Figure 13] This is a schematic diagram of the electronic device according to an embodiment of the present invention. [Modes for carrying out the invention]
[0022] To further clarify the purpose, technical means, and advantages of the embodiments of this application, the technical means of the embodiments of this application will be clearly and completely described below with reference to the drawings of the embodiments, and it is clear that the embodiments described are some, but not all, embodiments of this application. All other embodiments that a person skilled in the art could obtain based on the embodiments of this application without any creative work are all included within the scope of protection of this application.
[0023] The following disclosure provides many different embodiments or examples to realize different structures of the present invention. For the sake of simplicity in the disclosure of the present invention, the components and setups of specific examples are described below. Naturally, these are illustrative only and are not intended to limit the present invention. Furthermore, the present invention may use repeated reference numerals and / or reference letters in different examples. Such repetition is for the purpose of simplification and clarity and does not in itself indicate relationships between the various embodiments and / or setups described.
[0024] In order to solve the technical problem that the surround effect achieved by conventional surround stereo playback systems is constant and it is not possible to individually set different surround playback effects for different audio tracks, thus affecting the user experience, the audio surround playback method, apparatus, and electronic device according to the present invention first isolates at least one audio track included in the audio to be played back, then determines the surround playback effect to be achieved individually set for that audio track and the playback module that will achieve that surround playback effect, thereby setting different energy coefficients for the corresponding audio track data for the playback module based on the surround playback effect, and finally, when playing back the audio to be played back, the audio track data, energy coefficients, and channel data determined by each playback module can be combined as target channel data, and in this way, when each playback module is controlled to play back the corresponding target channel data, not only is playback of the audio to be played back achieved, but the surround playback effect of at least one audio track is also achieved, and different surround playback effects can be individually customized for different audio tracks included in the audio during the process of playing back the audio, thereby improving the user experience.
[0025] To facilitate understanding of the audio surround playback method relating to this application, the following will first describe application scenarios related to the method as examples.
[0026] Figure 1 is a schematic diagram of an application scene for the audio surround playback method according to an embodiment of the present invention. The application scene shown in Figure 1 is a 5.1 channel (five channel playback modules and one bass channel playback module) multi-channel surround system. As shown in Figure 1, the application scene may include a user P, a front left channel playback module FL located to the front left of user P, a front right channel playback module FR located to the front right of user P, a center channel playback module C located in front of user P, a rear left channel surround playback module SL located to the rear left of user P, a rear right channel surround playback module SR located to the rear right of user P, and a bass channel playback module SW.
[0027] The playback modules (FL, FR, SL, SR, C, and SW) included in the application scene shown in Figure 1 may be playback speakers, playback enclosures, or other types of audio players, and the embodiments of this application are not limited thereto.
[0028] Figure 2 is a schematic diagram of an application scene for another audio surround playback method according to an embodiment of the present invention. The application scene shown in Figure 2 is a 7.2 channel (seven channel playback modules and two bass channel playback modules) multi-channel surround system. As shown in Figure 2, the application scene may include a user P, a front left channel playback module FL located to the front left of user P, a front right channel playback module FR located to the front right of user P, a center channel playback module C located in front of user P, a rear left channel surround playback module SL located to the rear left of user P, a rear right channel surround playback module SR located to the rear right of user P, two bass channel playback modules SW located on both sides in front of user P, a rear left channel surround playback module SBL located to the left directly behind user P, and a rear right channel surround playback module SBR located to the right directly behind user P.
[0029] The playback modules (FL, FR, SL, SR, C, SBL, SBR, and two SWs) included in the application scene shown in Figure 2 may be playback speakers, playback enclosures, or other types of audio players, and the embodiments of this application are not limited thereto.
[0030] Figure 3 is a schematic diagram of an application scene for another audio surround playback method according to an embodiment of the present invention. The application scene shown in Figure 3 is a 5.1.4 channel (five channel playback modules, one bass channel playback module, and four sky channel playback modules) multi-channel surround system. As shown in Figure 3, the application scene may include a user P, a front left channel playback module FL located to the front left of user P, a front right channel playback module FR located to the front right of user P, a center channel playback module C located in front of user P, a rear left channel surround playback module RL located to the rear left of user P, a rear right channel surround playback module RR located to the rear right of user P, and four sky channel playback modules FHL, FHR, RHL, and RHR located above user P.
[0031] The playback modules (FL, FR, RL, RR, C, SW, FHL, FHR, RHL, and RHR) included in the application scene shown in Figure 3 may be playback speakers, playback enclosures, or other types of audio players, and the embodiments of this application are not limited thereto.
[0032] In the prior art, the audio to be played back may be played back in a multi-channel surround system in any of the application scenes shown in Figures 1 to 3, and the audio to be played back may be stereo including left and right channels, or it may be audio including multiple channels.
[0033] Currently, when playing audio using the multi-channel surround sound system shown in Figures 1-3, surround sound playback of the audio can generally be achieved by determining the channel data to be played corresponding to each playback module of the audio in the multi-channel surround sound system, and then controlling each playback module to play the corresponding channel data. However, the surround effect achieved by the above method is constant, and it is not possible to individually set different surround playback effects for different audio tracks, such as focusing the drum sounds of some audio tracks in front of the user, which seriously impacts the user experience.
[0034] In contrast, the present invention provides an audio surround playback method that can improve the user experience by individually customizing different surround playback effects for different audio tracks contained in the audio during the audio playback process.
[0035] The audio surround playback method according to the present invention will be further described below with reference to the drawings, and these examples are not intended to limit the embodiments of the present invention.
[0036] Figure 4 is a flowchart of an embodiment of the audio surround playback method according to an embodiment of the present invention. As shown in Figure 4, the process may include the following steps 401 to 406.
[0037] In step 401, the audio to be played is acquired, and multiple first playback modules corresponding to the audio to be played are determined.
[0038] The above-mentioned audio to be played is audio to be played, and the audio to be played may be stereo, may include left channel audio and right channel audio, and may further include other channel audio, but the embodiments of this application are not limited thereto.
[0039] The first playback module described above is any playback module that plays the audio to be played in a playback scene. For example, in the scenes shown in Figures 1 to 3 above, if all playback modules included in the multi-channel surround system in any of the scenes are for playing the audio to be played, then all playback modules in the multi-channel surround system can be designated as the first playback module for the audio to be played.
[0040] In one embodiment, the implementing entity of the embodiment of the present application may be a controller of a multi-channel surround sound system. Based on this, when the implementing entity of the embodiment of the present application receives audio via a wireless module, a Bluetooth® module, or an interface, it determines the received audio to be played. When a user needs to play audio, they can transmit the audio to be played to the implementing entity of the embodiment of the present application via a wireless network connection, a Bluetooth® connection, or an interface.
[0041] In another embodiment, when the implementing entity of the present embodiment detects a voice control command issued by the user, it can identify the voice control command and identify the audio identifier of the audio to be played from the voice control command. Subsequently, it can obtain the audio to be played from a pre-configured audio database based on the audio identifier.
[0042] In one embodiment, the implementer of the embodiment of the present invention can acquire the audio to be played back, then determine a playback module (hereinafter referred to as the first playback module for convenience of explanation) that will play the audio to be played back, and thereby realize the playback of the audio to be played back.
[0043] As one selectable embodiment, the implementer of the embodiment of the present invention can determine a plurality of corresponding first playback modules that will play the audio to be played in the current multi-channel surround system based on the playback recording history. For example, if in the recent playback recording history all playback modules included in the multi-channel surround system have played audio, it can be determined that all playback modules included in the multi-channel surround system are first playback modules. Alternatively, for example, if in the playback recording history the bass channel playback module in the multi-channel surround system has not played audio in any of the multiple playback recordings, it can determine that a playback module other than the bass channel playback module is the first playback module for the audio to be played.
[0044] In another optional embodiment, the implementer of the embodiment of the present invention may acquire the current operating state of each connected playback module and determine the playback module whose current operating state is normal as the first playback module of the audio to be played.
[0045] In step 402, at least one audio track in the audio to be played back and the corresponding audio track data are determined.
[0046] The above-mentioned audio track refers to an audio track that corresponds to any of the sound sources included in the audio being played, for example, an instrument included in the audio being played (electronic drums, piano, bass, etc.).
[0047] Furthermore, the above audio track may also be an audio track for which surround sound playback effects need to be set individually.
[0048] The above audio track data is audio data corresponding to the above audio track in the audio to be played, for example, drum sound data included in the audio to be played, and said drum sound data is audio track data corresponding to the audio track in which an electronic drum exists.
[0049] In one embodiment, the implementer of the embodiment of the present invention can input the audio to be played back into a pre-trained sound source separation model, and perform sound source separation on the audio to be played back using the sound source separation model. As a result, it can obtain all the audio tracks included in the audio to be played back output from the sound source separation model, as well as the audio track data corresponding to each audio track.
[0050] In another embodiment, the implementer of the embodiment of the present invention can first determine the audio tracks in the audio to be played back that need to achieve a user-defined surround sound playback effect, and then determine the audio tracks and the corresponding audio track data by performing sound source separation on the audio to be played back using a pre-trained sound source separation model.
[0051] In step 403, the surround sound playback effect for any of the audio tracks is determined, and at least one second playback module corresponding to the audio track is selected from a plurality of first playback modules.
[0052] The surround sound effect described above refers to the surround sound effect achieved when an audio track plays audio track data, and may also refer to the surround sound position, that is, the position of the audio track that the user hears after the audio track data of the audio track has been played. For example, if the audio track is a drum sound audio track, the surround sound effect set for the drum sound audio track may be that the drum sound perceived by the user is near their ears, in front of them, or behind them.
[0053] The above-mentioned second playback module refers to a playback module that realizes the surround playback effect of the audio track. That is, the audio track data of the audio track is played back by the second playback module, and the playback effect is the surround playback effect. To make it easier to understand, the second playback module may be one of several first playback modules corresponding to the audio to be played back, and it realizes the surround playback effect. Since the surround effect cannot be realized by only one playback module, there are two or more of the above-mentioned second playback modules, and they are relative to each other. For example, in the application scene shown in Figure 1, the front left channel playback module FL is located to the front left of user P, the front right channel playback module FR is located to the front right of user P, the rear left channel surround playback module SL is located to the rear left of user P, and the rear right channel surround playback module SR is located to the rear right of user P.
[0054] In one embodiment, the implementer of the embodiment of the present invention can determine the surround playback effect for any audio track and determine at least one second playback module from a plurality of first playback modules to realize the surround playback effect for the audio track, in order to set different surround playback effects for the audio track before playing the audio to be played back.
[0055] Specifically, how the surround sound playback effect for any given audio track is determined, and how at least one second playback module corresponding to the audio track is selected from multiple first playback modules, can be explained below and will not be explained in detail here.
[0056] In step 404, the energy coefficient of the audio track data corresponding to each second playback module is determined based on the surround sound playback effect described above.
[0057] In step 405, target channel data is synthesized for any of the first playback modules based on the energy coefficient corresponding to the first playback module, which includes the energy coefficient of the audio track data corresponding to the second playback module, the audio track data corresponding to the energy coefficient, and the channel data corresponding to the first playback module.
[0058] In step 406, each first playback module is controlled to play the corresponding target channel data in order to play the audio to be played.
[0059] Steps 404 through 406 are explained below.
[0060] The energy coefficient mentioned above refers to the proportionality coefficient between the energy of the audio track data reproduced by the second playback module and the total energy of the audio track data.
[0061] The channel data mentioned above refers to the channel data determined for each first playback module and to be played back by that first playback module. This channel data consists of several audio data points in the audio to be played back, and multi-channel surround playback of the audio to be played back can be achieved by each first playback module playing back the corresponding channel data.
[0062] The target channel data described above is obtained by combining the channel data initially set for each first playback module, the corresponding audio track data to realize the surround playback effect of the audio track, and the energy coefficient. By playing the corresponding target channel data with each first playback module, multi-channel surround playback of the audio to be played can be realized, and the surround playback effect pre-set for the audio track can be achieved.
[0063] In the embodiments of the present invention, after determining a second playback module that realizes the surround playback effect of an audio track based on the surround playback effect of the audio track, the energy coefficient of the audio track data corresponding to each second playback module can be determined. As a result, the second playback module plays the audio track data corresponding to the audio track based on the energy coefficient, thereby allowing multiple second playback modules to play the audio track data corresponding to the energy coefficient and realize the surround playback effect of the audio track.
[0064] Specifically, how the energy coefficient of the audio track data corresponding to each second playback module is determined can be explained below, but will not be explained in detail here.
[0065] Based on this, the second playback module for surround sound playback effect corresponding to each audio track, and the energy coefficient of the audio track data corresponding to each second playback module are determined. Since the second playback module is determined from the first playback module, the first playback module may include the second playback module, and furthermore, the energy coefficient of the audio track data corresponding to the determined second playback module is the energy coefficient of the first playback module corresponding to the second playback module.
[0066] For example, suppose the first playback module includes first playback module A, first playback module B, first playback module C, and first playback module D. Furthermore, suppose that for a given audio track, first playback module A and first playback module C in the first playback module are determined to be second playback module A and second playback module C, respectively. In this case, the energy coefficient corresponding to second playback module A is the same as the energy coefficient corresponding to first playback module A, and the energy coefficient corresponding to second playback module C is the same as the energy coefficient corresponding to first playback module C.
[0067] Based on this, when playing audio to be played, target channel data can be synthesized for any of the first playback modules based on the energy coefficient corresponding to the first playback module, the audio track data corresponding to the energy coefficient, and the channel data corresponding to the first playback module. The energy coefficient corresponding to the first playback module includes the energy coefficient of the audio track data corresponding to the second playback module.
[0068] In one selectable embodiment, the channel data corresponding to the first playback module can be determined first.
[0069] As one exemplary embodiment, the channel data of the first regeneration module can be determined by the process shown in Figure 5. Figure 5 is a flowchart of an embodiment for determining the channel data of the first regeneration module according to an embodiment of the present application. As shown in Figure 5, the process may include the following steps 501 to 503.
[0070] Step 501 determines the number of channels in the audio to be played and the number of modules in the first playback module.
[0071] In step 502, the number of channels and the number of modules are compared, and the comparison result is obtained.
[0072] Steps 501 and 502 are explained below in a combined manner.
[0073] The number of channels mentioned above refers to the total number of channels included in the audio being played back.
[0074] The number of modules mentioned above refers to the total number of first playback modules corresponding to the audio to be played.
[0075] In one embodiment, the implementing entity of the embodiment of the present invention can determine the number of channels of the audio to be played back and the number of modules of the first playback module, compare the number of channels and the number of modules, and obtain the comparison result.
[0076] In step 503, the audio to be played is converted based on the comparison results, and channel data corresponding to the first playback module is obtained.
[0077] In one embodiment, when determining the channel data for the first playback module, it is necessary to determine the corresponding channel data for each first playback module. Therefore, the implementing entity of the embodiment of the present invention can compare the number of channels of the audio to be played with the number of modules in the first playback module, convert the audio to be played based on the comparison result, and obtain the channel data for each first playback module.
[0078] Preferably, if the above comparison result indicates that the number of channels is greater than the number of modules, then the channel data included in the audio to be played back is greater than the number of first playback modules and cannot be corresponded one-to-one. In this case, a pre-trained downmix model may be called to convert the audio to be played back into channel data corresponding to each first playback module.
[0079] Preferably, if the comparison result indicates that the number of channels is equal to the number of modules, the channel data for each first playback module can be determined by associating the channel data included in the audio to be played with the first playback modules, in order to show that the channel data included in the audio to be played with can correspond one-to-one with the first playback modules.
[0080] Preferably, if the comparison result indicates that the number of channels is smaller than the number of modules, a pre-trained upmix model can be called to convert the audio to be played back, in order to indicate that the number of channels in the audio to be played back needs to be increased, and channel data corresponding to each first playback module can be obtained.
[0081] Furthermore, the upmix model described above can first separate pre-configured audio track data (e.g., audio track data corresponding to human voices) from unconfigured audio track data (e.g., audio track data corresponding to non-human voices) in the audio to be played back. Then, it can use the unconfigured audio track data to synthesize channel data corresponding to each channel playback module in a multi-channel surround system (e.g., the front left channel playback module, front right channel playback module, rear left channel playback module, and rear right channel playback module in a 5.1 channel system), and use the pre-configured audio track data to synthesize the center channel playback module. By separating human voices and background music in the audio to be played back using this upmix model, human voices are played in front of the user, and background music is played in surround sound by the other channel playback modules, thereby improving the multi-channel surround playback effect of the audio to be played back.
[0082] This concludes the explanation of the process shown in Figure 5.
[0083] As can be seen from steps 403 and 404, for each audio track, at least one second playback module corresponding to the audio track is determined from the first playback module to realize a surround playback effect for the audio track. In concrete implementation, the energy coefficient corresponding to the audio track data corresponding to the audio track of each second playback module can be determined. When the second playback module is determined from the first playback module, the energy coefficient of the audio track data corresponding to the second playback module belongs to the energy coefficient of the second playback module corresponding to the first playback module; that is, the energy coefficient corresponding to the first playback module includes the energy coefficient of the audio track data corresponding to the second playback module.
[0084] Specifically, how the energy coefficient of the audio track data corresponding to each second playback module is determined and how the energy coefficient of the corresponding first playback module is obtained can be explained below, but will not be explained in detail here.
[0085] Furthermore, since the audio to be played back may include multiple audio tracks, each first playback module can correspond to one energy coefficient and the audio track data of the audio track corresponding to that energy coefficient for each audio track.
[0086] Based on this, the first audio track data can be obtained by multiplying each audio track corresponding to the first playback module by the energy coefficient corresponding to the audio track and the audio track data corresponding to the energy coefficient.
[0087] Next, the first audio track data can be input to the filter corresponding to the first playback module to obtain the second audio track data. Different first playback modules can correspond to different filters; for example, the bass channel playback module can correspond to a low-pass filter, and other channel playback modules, in particular the second playback module that realizes surround sound playback effects for audio tracks, can correspond to a high-pass filter.
[0088] Finally, the target channel data can be obtained by combining the above channel data, the second audio track data, and the audio track data corresponding to the pre-set audio track. The audio track data corresponding to the pre-set audio track may be a pre-set audio track that does not need to achieve surround sound playback effects but needs to be enhanced, such as an audio track corresponding to a human voice.
[0089] For example, Figure 6 is a flowchart for determining target channel data according to an embodiment of the present invention. Figure 6 shows an example where the audio to be played back is stereo, including left and right channels. As shown in Figure 6, the method includes the following steps.
[0090] First, source separation (SS) and upmix model conversion (Upmix) are performed on the stereo signal. Source separation yields four audio tracks: a human voice audio track (including the left channel human voice audio track (Vocal L) and the right channel audio track (Vocal R)), a bass audio track (including the left channel bass audio track (Bass L) and the right channel bass audio track (Bass R)), a drum audio track (including the left channel drum audio track (Drum L) and the right channel drum audio track (Drum R)), and other audio tracks (including the other audio track on the left channel (Others L) and the other audio track on the right channel (Others R)).
[0091] The upmix model allows for conversion to obtain five channel data sets: front left channel data FL1, front right channel data FR1, center channel data C, rear left channel data SL, and rear right channel data SR.
[0092] Based on this, we assume that pre-set surround playback effects are applied to the drum audio track and other audio tracks, and that the energy coefficients of both the front left channel playback module and the front right channel playback module are 1. In this case, the target channel data FL for the front left channel playback module is obtained by combining four types of audio track data: FL1, the second audio track data obtained by multiplying Vocal L (left channel human voice audio track data, i.e., the first audio track data) and DrumL (left channel drum audio track data) by the energy coefficient and then inputting it into a high-pass filter (HPF), and the second audio track data obtained by multiplying others L (other audio track data for the left channel, i.e., the first audio track data) by the energy coefficient (1) and then inputting it into a high-pass filter.
[0093] At the same time, the target channel data FR of the front right channel playback module is obtained by combining four types of audio track data: FR1, Vocal R (right channel human voice audio track data, i.e., first audio track data), DrumR (right channel drum audio track data) multiplied by an energy coefficient and then input into a high-pass filter (HPF) to obtain second audio track data, and others R (other audio track data of the right channel, i.e., first audio track data) multiplied by an energy coefficient (1) and then input into a high-pass filter to obtain second audio track data.
[0094] Furthermore, since the central playback module does not have audio track data that requires surround sound playback effects, it can directly synthesize human voice audio track data (Vocal) and central channel data (C) to obtain the target central channel data (C1).
[0095] Furthermore, the target channel data LFE corresponding to the bass channel playback module may be obtained by combining two second audio track data sets: the second audio track data obtained by inputting drum sound audio track data (first audio track data) into a low-pass filter (HPF), and the second audio track data obtained by inputting bass audio track data (Bass, i.e., first audio track data) into a low-pass filter.
[0096] Furthermore, for the rear left channel playback module RL and the rear right channel playback module RR, the energy coefficient of the corresponding audio track data is 0, so Drum is not included during synthesis. Specifically, the target channel data RL for the rear left channel playback module is obtained by synthesizing others and the channel data SL for the rear left channel playback module acquired by the upmix model, and is played back with a delay during playback. The target channel data RR for the rear right channel playback module is obtained by synthesizing others and the channel data SR for the rear right channel playback module acquired by the upmix model, and is played back with a delay during playback.
[0097] This concludes the explanation of the process shown in Figure 6.
[0098] Based on this, by controlling each first playback module to play the corresponding target channel data, it is possible to achieve multi-channel surround playback of the audio to be played and surround playback effects for audio tracks.
[0099] As one possible embodiment, a third regeneration module can be selected from the second regeneration module in which the energy coefficient is smaller than a preset coefficient threshold (e.g., 0.5).
[0100] Subsequently, because the energy corresponding to the audio track data to be played back by the third playback module is low, the first playback module other than the third playback module can be controlled to play back the corresponding target channel data in order to increase the surround playback effect. After a preset time has elapsed, the third playback module can be controlled to play back the corresponding target channel data.
[0101] The technical means according to the embodiment of the present application is realized by acquiring audio to be played back, determining a plurality of first playback modules corresponding to the audio to be played back, determining at least one audio track in the audio to be played back and audio track data corresponding to the audio track, determining the surround playback effect of any of the audio tracks, determining at least one second playback module corresponding to the audio track from the plurality of first playback modules, determining the energy coefficient of the audio track data corresponding to each second playback module based on the surround playback effect, synthesizing target channel data for any of the first playback modules based on the energy coefficient corresponding to the first playback module, which includes the energy coefficient of the audio track data corresponding to the second playback module, the audio track data corresponding to the energy coefficient, and the channel data corresponding to the first playback module, and controlling each first playback module to play the corresponding target channel data in order to play back the audio to be played back. In this technology, at least one audio track included in the audio to be played back is first separated, then a surround playback effect to be achieved is individually set for that audio track, and a playback module is determined to achieve that surround playback effect. Based on the surround playback effect, different energy coefficients for the corresponding audio track data are set for the playback module. Finally, when playing back the audio to be played back, the audio track data, energy coefficients, and channel data determined by each playback module can be combined as target channel data. In this way, by controlling each playback module to play back the corresponding target channel data, not only is playback of the audio to be played back achieved, but the surround playback effect of at least one audio track is also achieved. During the process of playing back audio, different surround playback effects can be individually customized for different audio tracks included in the audio, thereby improving the user experience.
[0102] Figure 7 is a flowchart of another embodiment of the audio surround playback method according to an embodiment of the present invention. The process shown in Figure 7 specifically explains how the surround playback effect of an audio track is determined based on the process shown in Figure 4. As shown in Figure 7, the process may include the following steps 701 to 704.
[0103] Step 701 determines whether the audio track is set to surround sound playback. If it is, step 702 is executed; otherwise, the process terminates.
[0104] Step 702 determines whether a surround playback position has been set for the audio track. If it has, step 703 is executed; otherwise, step 704 is executed.
[0105] In step 703, the above surround playback position is determined as the surround playback effect for the audio track.
[0106] In step 704, the position where the pre-configured first playback module is located is determined as the surround playback effect for the audio track.
[0107] The following explains steps 701 through 704 together.
[0108] In one embodiment, the implementer of the embodiment of the present invention can determine whether or not each of the audio tracks determined above is set to surround sound playback.
[0109] One possible implementation is to determine whether or not a surround playback identifier exists for each audio track, and if it is determined that a surround playback identifier exists for that audio track, then it is determined that the audio track is set to surround playback.
[0110] In another optional embodiment, the user can pre-configure audio tracks that require surround sound playback and include them in a pre-configured set of audio tracks. Based on this, the implementer of the embodiment of the present invention can determine whether the audio track is in the set of audio tracks, and if it is, it can determine that the audio track is set for surround sound playback.
[0111] Preferably, if it is determined that an audio track is set to surround sound playback, it is possible to determine whether or not a surround sound playback position exists set for that audio track.
[0112] If it is determined that there is a surround playback position set for an audio track as one of the selectable implementation methods, then that surround playback position is determined as the surround playback effect for that audio track.
[0113] As one exemplary embodiment, the implementer of the embodiment of the present application can set the surround playback position of an audio track in the following manner. First, a positional relationship scene graph of multiple first playback modules can be output by a visualization interface. The positional relationship scene graph may be a three-dimensional scene graph in the corresponding scene of a multi-channel surround system, for example, a three-dimensional scene graph of any of the applicable scenes in Figures 1 to 3, or a two-dimensional plan view showing the related positional relationships, for example, the two-dimensional plan view shown in Figure 5. Figure 8 is a schematic diagram of the two-dimensional planar structure of a playback module according to an embodiment of the present application. As shown in Figure 8, the two-dimensional planar structure diagram may include, using a 5.1 multi-channel surround system as an example, a front left channel playback module, a front right channel playback module, a rear left channel playback module, and a rear right channel playback module corresponding to the user. The center channel playback module and the bass channel playback module in the multi-channel surround system cannot realize the surround playback effect of an audio track, so the two-dimensional plan view does not have to include the center channel playback module and the bass channel playback module.
[0114] In response to this, the user can set the surround playback position of the audio track relative to the positional relationship scene graph, and based on this, the implementing entity of the embodiment of the present invention can, in response to the setting operation on the positional relationship scene graph, determine the initial surround playback position of the audio track in the positional relationship scene graph, and determine the actual position represented by the initial surround playback position as the surround playback position of the audio track.
[0115] Taking the two-dimensional planar structure diagram output in Figure 8 as an example, the two-dimensional planar structure diagram may include several setting methods shown in Figure 9, where the surround playback position is the position of the left channel audio track and the right channel audio track, Figure 9A is a schematic diagram of the initial surround playback position according to an embodiment of the present application, where the initial surround playback position is located at the user's front left channel playback module and front right channel playback module. Figure 9B is a schematic diagram of another initial surround playback position according to an embodiment of the present application, where the left channel audio track is located between the front left channel playback module and the rear left channel playback module, and the right channel audio track is located between the front right channel playback module and the rear right channel playback module. Figure 9C is a schematic diagram of another initial surround playback position according to an embodiment of the present application, where the initial surround playback modules are located on both sides of the user, that is, the left channel audio track is located between the front left channel playback module and the rear left channel playback module, and the right channel audio track is located between the front right channel playback module and the rear right channel playback module. Figure 9D is a schematic diagram of a further initial surround playback position according to an embodiment of the present invention, in which the left channel audio track is located at the rear left channel playback module, and the right channel audio track is located at the rear right channel playback module.
[0116] In another exemplary embodiment, the position information of each first regeneration module (for example, the coordinate information of each regeneration module) can be acquired using a pre-configured Bluetooth® module or distance sensor, and the position information of each first regeneration module can be output.
[0117] Based on this, the user can set surround playback position information for an audio track based on the position information of each first playback module (for example, by inputting coordinate information of the surround playback position via a visualization interface), and the implementing entity of the embodiment of the present invention can, in response to the user's setting operation, acquire the surround playback position information and determine the position corresponding to the surround playback position information as the surround playback position of the audio track.
[0118] As another possible configuration, if it is determined that no surround playback position exists for an audio track, the position of a pre-configured first playback module can be determined as the surround playback effect for that audio track. The pre-configured first playback module may be any playback module in a multi-channel surround system. Preferably, the audio tracks may be classified into left channel audio tracks and right channel audio tracks, in which case the pre-configured first playback module may be a pair of first playback modules, for example, a front left channel playback module corresponding to the left channel audio track and a front right channel playback module corresponding to the right channel audio track.
[0119] In the technical means according to the embodiment of the present application, when it is determined that an audio track is set to surround playback, it is determined whether or not a surround playback position exists set for the audio track. If it exists, the surround playback position is determined as the surround playback effect of the audio track. If it does not exist, the position where a preset first playback module exists is determined as the surround playback effect of the audio track. In this technical means, by presetting a surround playback position for the audio track, the surround playback position is determined as the surround playback effect of the audio track, and the surround playback position can be set arbitrarily, thus enabling diversity in surround playback effects for audio tracks.
[0120] Figure 10 is a flowchart of another embodiment of an audio surround playback method according to an embodiment of the present invention. The process shown in Figure 10 specifically describes how, based on the process shown in Figure 7, at least one second playback module corresponding to an audio track is determined from a plurality of first playback modules, and how the energy coefficient of the audio track data corresponding to each second playback module is determined. As shown in Figure 10, the process may include the following steps 1001 to 1006.
[0121] In step 1001, the first playback module included in the playback module set is determined to be the second playback module, the surround playback effect of the audio track is the surround playback position, which is located between the planes in which the symmetrical playback module sets exist, or located in the first playback module in the playback module set, each playback module set includes a first playback module corresponding to the left channel and a first playback module corresponding to the right channel, and the surround playback position is the position corresponding to the left channel audio track and the right channel audio track included in the audio track, respectively.
[0122] In step 1002, the first energy coefficient of each second regeneration module is determined.
[0123] Steps 1001 and 1002 are explained below in a combined manner.
[0124] As can be seen from the process shown in Figure 7, the surround playback effect of an audio track is the surround playback position, and this surround playback position is located between planes in which symmetrical playback module sets exist, or in the first playback module of a playback module set. Each playback module set includes a first playback module corresponding to the left channel and a first playback module corresponding to the right channel, e.g., a front left channel playback module and a front right channel playback module. A symmetrical playback module set may also include a front playback module and a rear playback module. In 7.1 multichannel, a symmetrical playback module set may include a front playback module (front left channel playback module and front right channel playback module) and an intermediate playback module (e.g., SL and SR shown in Figure 2), and an intermediate playback module and a rear playback module (e.g., SBL and SBR shown in Figure 2). The symmetry may include perfect symmetry or symmetry in the plane in which the two playback module sets exist.
[0125] Based on this, the implementing entity of the embodiment of the present invention can determine that the first playback module included in the above playback module set is the second playback module corresponding to the audio track.
[0126] Based on this, the first energy coefficient of the second regeneration module can be determined.
[0127] In one selectable embodiment, the total distance between two opposing sets of playback modules where surround playback positions exist can be determined, which here can refer to the total distance between the planes where the two sets of playback modules exist.
[0128] Subsequently, for each second playback module, the distance between the surround playback position and the plane on which the second playback module is located (hereinafter referred to as the first distance for convenience of explanation) can be determined. To make it clear, if the surround playback position is at the location where the second playback module is located, the first distance is 0.
[0129] Furthermore, based on the above-mentioned first distance and total distance, the distance ratio between the surround playback position and the second playback module (hereinafter referred to as the first distance ratio for convenience of distinction) can be determined. Preferably, the first distance ratio may be obtained by dividing the first distance by the total distance.
[0130] Based on this, since the energy of the audio track data being played decreases as you move further away from the surround playback position, the first energy coefficient can be obtained by subtracting the above first distance ratio from a preset value (for example, 1).
[0131] Subsequently, based on the first energy coefficient described above, the energy coefficient of the audio track data corresponding to the second playback module can be determined.
[0132] In step 1003, it is determined whether the first regeneration module includes the Sky regeneration module. If it does, step 1004 is performed; otherwise, step 1006 is performed.
[0133] In step 1004, the Sky Regeneration Module is determined as the second regeneration module, and the second energy coefficient of the Sky Regeneration Module is determined.
[0134] In step 1005, both the first and second energy coefficients are determined as the energy coefficients of the audio track data corresponding to the second playback module.
[0135] In step 1006, the above first energy coefficient is determined as the energy coefficient of the audio track data corresponding to the second playback module.
[0136] Steps 1003 through 1006 are explained below.
[0137] The above-mentioned sky playback module refers to the playback module located above the user's head, for example, the four sky channel playback modules FHL, FHR, RHL, and RHR in the application scene shown in Figure 3.
[0138] In the embodiment of the present invention, when determining the energy coefficient of audio track data corresponding to the second playback module based on the first energy coefficient, it is possible to first determine whether or not a Sky playback module exists in the second playback module.
[0139] Preferably, if a Sky playback module is not present, the first energy coefficient can be directly determined as the energy coefficient of the audio track data corresponding to the second playback module, in which case the second playback module is the first playback module included in the playback module set.
[0140] Preferably, if a Sky playback module is present, both the first playback module and the Sky playback module included in the playback module set are determined as the second playback module. Furthermore, the second energy coefficient of the Sky playback module is determined, and both the first and second energy coefficients can be determined as the energy coefficients of the audio track data corresponding to the second playback module.
[0141] In one selectable embodiment, when determining the second energy coefficient of the sky playback module, the vertical distance between the plane on which the sky playback module resides and a preset plane can be determined. The preset plane may be determined based on the heights of other playback modules located on the ground, and since users are generally seated when listening to the audio to be played, the lowest value in the height of the surround playback effect may be the plane on which the heights of other playback modules located on the ground reside.
[0142] Subsequently, the second distance ratio between the surround playback position and the sky playback module can be determined based on the second distance and the vertical distance. Preferably, the ratio of the second distance to the vertical distance may be determined as the second distance ratio.
[0143] Finally, the second energy coefficient can be obtained by subtracting the above second distance ratio from a predetermined value.
[0144] Furthermore, in a multi-channel surround sound system with a Sky Playback Module, the surround playback position of an audio track set by the user may have a certain height in space, and therefore the Sky Playback Module may have an energy coefficient corresponding to the audio track data.
[0145] To make it easier to understand, if the second distance is zero, then the surround playback position in this case is located in the plane where the sky playback module exists, and its second energy coefficient is 1.
[0146] For example, taking the multi-channel surround system shown in Figure 3 as an example, Figure 11 is a schematic diagram of the surround playback position according to the embodiment of the present application. Assuming that the two points in Figure 11 are the surround playback positions, and that the first distance ratio is 30% and the second distance ratio is 50%, the energy audio track data of the audio track data corresponding to each second playback module will be as shown in equation (1) below.
number
[0147] In the technical means according to the embodiment of the present application, a first playback module included in a playback module set is determined as a second playback module, the surround playback effect of an audio track is the surround playback position, the surround playback position is located between planes in which symmetrical playback module sets exist, or is located in the first playback module in the playback module set, each playback module set includes a first playback module corresponding to the left channel and a first playback module corresponding to the right channel, the surround playback position is the position corresponding to the left channel audio track and the right channel audio track included in the audio track, respectively, the first energy coefficient of each second playback module is determined, it is determined whether or not a sky playback module is included in the first playback module, if it is included, the sky playback module is determined as the second playback module, the second energy coefficient of the sky playback module is determined, both the first and second energy coefficients are determined as the energy coefficients of the audio track data corresponding to the second playback module, if it is not included, the first energy coefficient is determined as the energy coefficient of the audio track data corresponding to the second playback module. This technical means achieves a more accurate determination of the energy coefficient of each second regenerative module by determining the energy coefficient of the regenerative module for different dimensions, such as the horizontal and vertical directions.
[0148] Figure 12 is a block diagram of an embodiment of an audio surround playback device according to an embodiment of the present application. As shown in Figure 12, the device is A first determination module 121 acquires the audio to be played and determines a plurality of first playback modules corresponding to the audio to be played, A second determination module 122 that determines at least one audio track in the audio to be played back and the audio track data corresponding to the audio track, A third determination module 123 determines the surround sound playback effect for any audio track and determines at least one second playback module corresponding to the audio track from a plurality of first playback modules, A fourth determination module 124 determines the energy coefficient of the audio track data corresponding to each of the second playback modules based on the surround sound playback effect described above, A synthesis module 125 synthesizes target channel data for any of the above-mentioned first playback modules, based on an energy coefficient corresponding to the first playback module, which includes the energy coefficient of the audio track data corresponding to the second playback module, audio track data corresponding to the energy coefficient, and channel data corresponding to the first playback module. The system includes a control module 126 that controls each of the first playback modules to play the corresponding target channel data in order to play the above-mentioned audio to be played.
[0149] Figure 13 is a schematic diagram of an electronic device according to an embodiment of the present application, which includes a processor 131, a communication interface 132, a memory 133, and a communication bus 134, the processor 131, the communication interface 132, and the memory 133 communicating with each other via the communication bus 134.
[0150] Memory 133 stores computer programs, In one embodiment of the present invention, the processor 131 executes a program stored in the memory 133 to realize an audio surround playback method according to an embodiment of any of the methods described above. The steps include obtaining the audio to be played and determining a plurality of first playback modules corresponding to the audio to be played, The steps include determining at least one audio track in the audio to be played back and the audio track data corresponding to the audio track, The steps include determining the surround sound playback effect for any audio track and determining at least one second playback module corresponding to the audio track from a plurality of first playback modules, The steps include determining the energy coefficient of the audio track data corresponding to each of the second playback modules based on the surround sound playback effect described above, A step of synthesizing target channel data for any of the above-mentioned first playback modules, based on an energy coefficient corresponding to the first playback module, which includes the energy coefficient of the audio track data corresponding to the second playback module, audio track data corresponding to the energy coefficient, and channel data corresponding to the first playback module. The process includes the step of controlling each of the first playback modules to play the corresponding target channel data in order to play the above-mentioned audio to be played.
[0151] Embodiments of the present invention further provide a storage medium which stores a computer program that, when executed by a processor, implements steps of an audio surround playback method according to any embodiment of the above-described method.
[0152] The embodiments of the apparatus described above are illustrative only, and the units described as separating members may or may not be physically separated, and the members referred to as units may or may not be physical units, that is, they may be located in one place or distributed across multiple grid units. Some or all of the modules may be selected as practically necessary to achieve the objectives of the technical means of this embodiment.
[0153] Based on the above description of embodiments, those skilled in the art will clearly understand that each embodiment can be implemented by adding a common hardware platform to the software, and of course, it can be implemented in hardware. Based on this understanding, the above technical means may be embodied in the form of a software product, which may be stored on a computer-readable storage medium such as ROM / RAM, magnetic disk, or optical disk, and may include a number of instructions for causing a computer device (which may be a personal computer, server, or grid equipment, etc.) to perform the method described in each embodiment or in some part of the embodiment.
[0154] To ensure clarity, the terminology used in this specification is intended to describe, and not limit, specific exemplary embodiments. Unless otherwise explicitly indicated in the context, the singular forms “one,” “one,” and “above” used in this specification may include the plural form. The terms “include,” “equip,” “contain,” and “have” are inclusive and thereby specify the presence of the described features, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, elements, components, and / or combinations thereof. The steps, processes, and operations of the methods described in this specification should not be construed as necessarily having to be performed in a specific order described or explained unless the order of execution is explicitly specified. It should be understood that different or alternative steps may be used.
[0155] The above description is merely a specific embodiment of the present invention, and those skilled in the art will be able to understand or implement the invention. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the invention. Accordingly, the present invention is not limited to these embodiments described herein, but rather conforms to the broadest scope that is consistent with the principles and novel features filed herein. [Explanation of Symbols]
[0156] 131 processors, 133 memory
Claims
1. The steps include acquiring the audio to be played and determining a plurality of first playback modules corresponding to the audio to be played, The steps include determining at least one audio track in the audio to be played back and audio track data corresponding to the audio track, The steps include determining the surround sound playback effect for any audio track and determining at least one second playback module corresponding to the audio track from a plurality of first playback modules, The steps include determining the energy coefficient of the audio track data corresponding to each of the second playback modules based on the surround sound playback effect, A step of synthesizing target channel data for any of the first playback modules based on an energy coefficient corresponding to the first playback module, which includes the energy coefficient of the audio track data corresponding to the second playback module, the audio track data corresponding to the energy coefficient, and the channel data corresponding to the first playback module. The process includes controlling each of the first playback modules to play the corresponding target channel data in order to play the audio to be played back, An audio surround sound playback method characterized by the following features.
2. The step of determining the surround sound playback effect of the aforementioned audio track is: The steps include determining whether the aforementioned audio track is set to surround sound playback, If it is determined that the audio track is set to surround sound playback, the step is to determine whether or not a surround sound playback position has been set for the audio track. If it is determined that a surround playback position exists for the audio track, the steps include determining the surround playback position as the surround playback effect of the audio track, If it is determined that no surround playback position exists for the audio track, the step of determining the position where the pre-set first playback module exists as the surround playback effect for the audio track is included. The method according to feature 1.
3. A method for outputting a scene graph showing the positional relationship of multiple first playback modules via a visualization interface, A method for determining the initial surround playback position of the audio track in the positional relationship scene graph in response to setting operations on the positional relationship scene graph, A method for determining the surround playback position of an audio track by determining the actual position represented by the initial surround playback position as the surround playback position of the audio track, and setting the surround playback position of an audio track. The method according to feature 2.
4. The surround playback position is located between planes where symmetrical playback module sets exist, or is located in a first playback module in the playback module set, each of which includes a first playback module corresponding to the left channel and a first playback module corresponding to the right channel, and the surround playback position is a position corresponding to the left channel audio track and the right channel audio track included in the audio track, respectively. The step of determining at least one second playback module corresponding to the audio track from a plurality of first playback modules is: The steps include determining whether the first playback module includes a Sky playback module that corresponds to Sky Channel, If it is determined that the first regeneration module does not include the Sky regeneration module, the first regeneration module included in the regeneration module set is determined to be the second regeneration module. If it is determined that the first regeneration module includes the sky regeneration module, the process includes the step of determining the first regeneration module and the sky regeneration module included in the regeneration module set as the second regeneration module. The method according to feature 2.
5. The step of determining the energy coefficient of the audio track data corresponding to each of the second playback modules based on the surround playback effect is as follows: A step of determining the total distance between two opposing sets of playback modules where the surround playback position exists, with respect to the second playback module in the playback module set, For each of the second playback modules, the steps include determining a first distance between the surround playback position and the plane on which the second playback module is located, A step of determining a first distance ratio between the surround playback position and the second playback module based on the first distance and the total distance, A step of obtaining a first energy coefficient by subtracting the first distance ratio from a predetermined value, The step of determining the energy coefficient of the audio track data corresponding to the second playback module based on the first energy coefficient, The method according to feature 4.
6. The step of determining the energy coefficient of the audio track data corresponding to the second playback module based on the first energy coefficient is: The steps include determining whether the Sky Regeneration Module is present in the second regeneration module, If it is determined that the Sky playback module does not exist, the first energy coefficient is determined as the energy coefficient of the audio track data corresponding to the second playback module. If it is determined that the sky regeneration module exists, the steps include determining the vertical distance between the plane on which the sky regeneration module exists and a predetermined plane, The steps include determining a second distance between the surround playback position and the plane on which the sky playback module is located, A step of determining a second distance ratio between the surround playback position and the sky playback module based on the second distance and the vertical distance, A step of obtaining a second energy coefficient by subtracting the second distance ratio from a predetermined value, The step includes determining that both the first energy coefficient and the second energy coefficient are energy coefficients for the audio track data corresponding to the second playback module, The method according to specification 5.
7. The step of synthesizing target channel data based on the energy coefficient corresponding to the first playback module, the audio track data corresponding to the energy coefficient, and the channel data corresponding to the first playback module is: The steps include determining the channel data corresponding to the first playback module, For each audio track corresponding to the first playback module, the first audio track data is obtained by multiplying the energy coefficient corresponding to the audio track by the audio track data corresponding to the energy coefficient. The steps include inputting the first audio track data into a filter corresponding to the first playback module to obtain the second audio track data, The process includes the step of synthesizing the channel data, the second audio track data, and audio track data corresponding to a pre-set audio track to obtain target channel data. The method according to feature 1.
8. The step of determining the channel data corresponding to the first playback module is: The steps include determining the number of channels in the audio to be played and the number of modules in the first playback module, The steps include comparing the number of channels and the number of modules and obtaining the comparison result, The process includes the step of converting the audio to be played based on the comparison result and obtaining channel data corresponding to the first playback module, The method according to feature 7.
9. The step of converting the audio to be played based on the comparison result and obtaining channel data corresponding to the first playback module is: If the comparison result indicates that the number of channels is greater than the number of modules, the steps include calling a pre-trained downmix model to convert the audio to be played back and obtaining channel data corresponding to the first playback module, If the comparison result indicates that the number of channels is equal to the number of modules, the step is to determine the channel data corresponding to the first playback module from the multiple channel data included in the audio to be played back. If the comparison result indicates that the number of channels is smaller than the number of modules, the comparison includes the step of calling a pre-trained upmix model to convert the audio to be played back and obtaining channel data corresponding to the first playback module. The method according to feature 8.
10. The step of controlling each of the first playback modules to play back the corresponding target channel data is: The steps include determining a third regeneration module from the second regeneration module whose energy coefficient is smaller than a preset coefficient threshold, The steps include controlling the first playback modules other than the third playback module among the first playback modules to play back the corresponding target channel data, The steps include controlling the third playback module to play the corresponding target channel data after a predetermined time has elapsed, The method according to feature 1.
11. A first determination module that acquires the audio to be played and determines a plurality of first playback modules corresponding to the audio to be played, A second determination module that determines at least one audio track in the audio to be played back and audio track data corresponding to the audio track, A third determination module that determines the surround playback effect of any audio track and determines at least one second playback module corresponding to the audio track from a plurality of first playback modules, A fourth determination module that determines the energy coefficient of the audio track data corresponding to each of the second playback modules based on the surround sound playback effect, A synthesis module that synthesizes target channel data for any of the first playback modules based on an energy coefficient corresponding to the first playback module, which includes the energy coefficient of the audio track data corresponding to the second playback module, the audio track data corresponding to the energy coefficient, and the channel data corresponding to the first playback module. Includes a control module that controls each of the first playback modules to play the corresponding target channel data in order to play the audio to be played back, An audio surround sound playback device characterized by the following features.
12. The system includes a processor and memory, and the processor, when it executes an audio surround playback program stored in the memory, realizes the audio surround playback method described in any one of claims 1 to 10. An electronic device characterized by the following features.