Audio splicing method and device, equipment and storage medium
Through HOA reconstruction upmixing technology, 2D channel audio is converted into 3D channel audio and spliced according to the time point, solving the sound quality problem when splicing 2D channel audio and 3D channel audio, achieving high-quality audio splicing effect.
Patent Information
- Application Number
- CN202510556310.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-29
- Publication Date
- 2025-07-18
AI Technical Summary
When the prior art splicing audio of 2D sound into the audio of 3D sound, audio quality problems are easily introduced, especially the problem of inconsistent volume and poor sound quality after advertising and main film rendering.
Through HOA reconstruction upmix technology, the original two-dimensional channel audio is upmixed into the target three-dimensional channel audio, and spliced with the original three-dimensional channel audio according to the audio insertion time point to generate spliced three-dimensional channel audio.
It ensures the high quality of spliced three-dimensional channel audio, avoids sound quality problems during subsequent rendering, and is suitable for mapping of different channels, with high sound field restoration, solving the problem of inconsistent volume.
Smart Images

Figure CN120343484A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of audio splicing technology, and in particular, to an audio splicing method, device, equipment, and storage medium. Background Art
[0002] With the development of audio technology, more and more film and television works are presented in the audio format of 3D (three-dimensional), that is, by adding upper channels, a three-dimensional scene with higher restoration degree is constructed; while advertisements mainly for content dissemination are relatively simple in production and still mainly use the audio format of 2D sound. When inserting a 2D sound advertisement into a 3D sound video source, the traditional audio splicing scheme is to directly map the existing channels of the 2D sound into the corresponding first N channels of the 3D sound and directly perform splicing. However, since this scheme ignores the audio characteristics and volume weights of different channels, and there are also significant differences in the subsequent rendering methods of 2D and 3D sounds, this splicing method often introduces new problems, such as audio quality problems after rendering of advertisements and feature films.
[0003] Therefore, how to splice the audio of a 2D sound video source into the audio of a 3D sound video source with high quality is a problem that needs to be solved by those skilled in the art. Summary of the Invention
[0004] This application provides an audio splicing method, device, equipment, and storage medium to splice the audio of a 2D sound video source into the audio of a 3D sound video source with high quality.
[0005] In a first aspect, this application provides an audio splicing method, including:
[0006] Obtain the original two-dimensional channel audio and the original three-dimensional channel audio;
[0007] Through the HOA reconstruction upmixing technology, upmix the original two-dimensional channel audio into the target three-dimensional channel audio; wherein, the channel type of the target three-dimensional channel audio is the same as that of the original three-dimensional channel audio;
[0008] Splice the target three-dimensional channel audio and the original three-dimensional channel audio according to the audio insertion time point to generate the spliced three-dimensional channel audio.
[0009] Optionally, upmixing the original two-dimensional channel audio into the target three-dimensional channel audio through the HOA reconstruction upmixing technology includes:
[0010] Determine the mono audio of the original two-dimensional channel audio;
[0011] Through the HOA reconstruction upmixing technology, upmix each mono audio into three-dimensional channel audio;
[0012] Superimpose the three-dimensional channel audio corresponding to each monophonic audio to generate the target three-dimensional channel audio.
[0013] Optionally, upmix each monophonic audio to three-dimensional channel audio through HOA reconstruction upmixing technology, including:
[0014] Determine the spatial harmonic signals of the monophonic audio;
[0015] Determine the decoding equation coefficients corresponding to the target channel type; the target channel type is the channel type of the original three-dimensional channel audio;
[0016] Determine the reconstructed signals of each speaker in the target channel type according to the spatial harmonic signals and the decoding equation coefficients;
[0017] Determine the three-dimensional channel audio of the monophonic audio according to the reconstructed signals of each speaker.
[0018] Optionally, after obtaining the original two-dimensional channel audio and the original three-dimensional channel audio, it further includes:
[0019] Perform audio quality detection on the original two-dimensional channel audio and the original three-dimensional channel audio;
[0020] If both detections pass, continue to execute the step of upmixing the original two-dimensional channel audio to the target three-dimensional channel audio through HOA reconstruction upmixing technology.
[0021] Optionally, before splicing the target three-dimensional channel audio and the original three-dimensional channel audio according to the audio insertion time point, it further includes:
[0022] Perform volume normalization processing on the target three-dimensional channel audio to generate the first standard three-dimensional channel audio; perform volume normalization processing on the original three-dimensional channel audio to generate the second standard three-dimensional channel audio.
[0023] Optionally, performing volume normalization processing on the target three-dimensional channel audio to generate the first standard three-dimensional channel audio; performing volume normalization processing on the original three-dimensional channel audio to generate the second standard three-dimensional channel audio includes:
[0024] Calculate the first target loudness value of the target three-dimensional channel audio through the audio loudness measurement standard, and adjust the volume of the target three-dimensional channel audio according to the first target loudness value to generate the first standard three-dimensional channel audio;
[0025] Calculate the second target loudness value of the original three-dimensional channel audio through the audio loudness measurement standard, and adjust the volume of the original three-dimensional channel audio according to the second target loudness value to generate the second standard three-dimensional channel audio.
[0026] Optionally, according to the audio insertion time point, splice the target three-dimensional channel audio and the original three-dimensional channel audio to generate spliced three-dimensional channel audio, including:
[0027] Cut the second standard three-dimensional channel audio according to the audio insertion time point to generate a first cut three-dimensional channel audio and a second cut three-dimensional channel audio;
[0028] Splice the first cut three-dimensional channel audio, the first standard three-dimensional channel audio, and the second cut three-dimensional channel audio in the audio playback order to generate spliced three-dimensional channel audio.
[0029] In a second aspect, the present application provides an audio splicing device, including:
[0030] An acquisition module, configured to acquire the original two-dimensional channel audio and the original three-dimensional channel audio;
[0031] An upmixing module, configured to upmix the original two-dimensional channel audio into a target three-dimensional channel audio through HOA reconstruction upmixing technology; wherein, the channel type of the target three-dimensional channel audio is the same as the channel type of the original three-dimensional channel audio;
[0032] A splicing module, configured to splice the target three-dimensional channel audio and the original three-dimensional channel audio according to the audio insertion time point to generate spliced three-dimensional channel audio.
[0033] In a third aspect, the present application provides an electronic device, including a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory complete communication with each other through the communication bus;
[0034] The memory is used to store a computer program;
[0035] The processor is configured to implement the steps of the above audio splicing method when executing the program stored in the memory.
[0036] In a fourth aspect, the present application further provides a computer storage medium, and the computer storage medium stores computer executable instructions, and the computer executable instructions are used to execute the steps of the above audio splicing method of the present application.
[0037] The above technical solutions provided by the embodiments of the present application have the following advantages compared with the prior art: The present application discloses an audio splicing method, device, equipment and storage medium; the present application obtains the original two-dimensional channel audio and the original three-dimensional channel audio, and through the HOA reconstruction upmixing technology, upmixes the original two-dimensional channel audio into the target three-dimensional channel audio; wherein, the channel type of the target three-dimensional channel audio is the same as that of the original three-dimensional channel audio; according to the audio insertion time point, splices the target three-dimensional channel audio and the original three-dimensional channel audio to generate the spliced three-dimensional channel audio. It can be seen that through the HOA reconstruction upmixing technology, the present application can upmix the original two-dimensional channel audio into a high-quality target three-dimensional channel audio, thereby ensuring the audio quality of the spliced three-dimensional channel audio and avoiding sound quality problems in the subsequent rendering process. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] The accompanying drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present invention and, together with the specification, are used to explain the principles of the present invention.
[0039] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0040] One or more embodiments are illustrated by way of example in the accompanying drawings, and these exemplary illustrations do not limit the embodiments. Elements with the same reference numerals in the drawings are represented as similar elements, unless otherwise stated, and the drawings in the figures do not constitute a proportional limitation.
[0041] Figure 1 Schematic flow chart of an audio splicing method provided by an embodiment of the present application;
[0042] Figure 2 Schematic flow chart of another audio splicing method provided by an embodiment of the present application;
[0043] Figure 3 Schematic flow chart of another audio splicing method provided by an embodiment of the present application;
[0044] Figure 4 Schematic flow chart of another audio splicing method provided by an embodiment of the present application;
[0045] Figure 5 Schematic flow chart of the specific audio splicing process provided by an embodiment of the present application;
[0046] Figure 6 Schematic structural diagram of an audio splicing device provided by an embodiment of the present application;
[0047] Figure 7 A schematic structural diagram of an electronic device provided by an embodiment of the present application. Specific implementation manners
[0048] To make the objectives, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Apparently, the described embodiments are some but not all of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts shall fall within the protection scope of the present application.
[0049] The following disclosure provides many different embodiments or examples for implementing different structures of the present invention. To simplify the disclosure of the present invention, components and settings of specific examples are described below. Of course, they are only examples and are not intended to limit the present invention. In addition, the present invention may repeat reference numerals and / or letters in different examples. This repetition is for the purpose of simplification and clarity and does not itself indicate the relationship between the various embodiments and / or settings discussed.
[0050] An embodiment of the present application discloses an audio splicing method, device, equipment and storage medium for splicing the source audio of 2D sound into the source audio of 3D sound with high quality.
[0051] See Figure 1 , a schematic flowchart of an audio splicing method provided by an embodiment of the present application. The method specifically includes:
[0052] S101. Obtain the original two-dimensional channel audio and the original three-dimensional channel audio;
[0053] In the present application, the original two-dimensional channel audio is the source audio of 2D sound, and the original three-dimensional channel audio is the source audio of 3D sound. Common channel types of two-dimensional channel audio include: 2.0 channel, 5.1 channel, etc. Among them, the 2.0 channel is a basic stereo audio configuration composed of two independent channels, including the left channel (Left) and the right channel (Right); the 5.1 channel refers to a surround sound system including six channels: the center channel, the front left channel, the front right channel, the rear left channel, the rear right channel, and a subwoofer channel; common channel types of three-dimensional channel audio include: 5.1.4 channel, 7.1.4 channel, etc.; among them, the 5.1.4 channel is a configuration of a three-dimensional surround sound system, including five speakers in the horizontal direction, four upper speakers, and one or more subwoofer speakers; the 7.1.4 channel is an advanced configuration of a three-dimensional panoramic sound system, composed of 7 main channels, 1 bass channel, and 4 sky channels.
[0054] S102. Upsample the original two - channel audio to the target three - channel audio through the HOA reconstruction upsampling technology, where the channel type of the target three - channel audio is the same as that of the original three - channel audio.
[0055] In this application, the HOA (Higher Order Ambisonics) reconstruction upsampling technology is a three - dimensional sound field processing technology based on Higher Order Ambisonics. According to the theory of spherical harmonic functions, it can map point sound sources to virtual speakers with arbitrary distributions for sound field reconstruction, that is, it can fit the original 2D sound field under the number and distribution of 3D sound speakers, thereby determining the upsampling method from 2D to 3D sound.
[0056] Through the HOA reconstruction upsampling technology in this application, the original two - channel audio can be upsampled to the target three - channel audio with the same channel type as the original three - channel audio. In this way, traditional stereo or multi - channel content can be converted into a sound field signal suitable for a three - dimensional immersive audio system, ensuring the audio quality of the reconstructed target three - channel audio. For example: If the original two - channel audio is a 2.0 - channel audio and the original three - channel audio is a 5.1.4 - channel audio, through the HOA reconstruction upsampling technology in this application, the 2.0 - channel audio can be upsampled to a 5.1.4 - channel audio and then spliced with the original three - channel audio to improve the quality of the spliced audio.
[0057] S103. Splice the target three - channel audio and the original three - channel audio according to the audio insertion time point to generate a spliced three - channel audio.
[0058] In this application, the audio insertion time point is the time point in the original three - channel audio where the original two - channel audio needs to be inserted. For example: If the original three - channel audio is a 60 - second audio and a section of the original two - channel audio needs to be inserted at 52.8 seconds of the original three - channel audio, then 52.8 seconds at this time is the audio insertion time point.
[0059] Since this application converts the original two - channel audio into the corresponding target three - channel audio, when splicing the audio in this application, specifically, the target three - channel audio is spliced to the audio insertion time point of the original three - channel audio, thereby generating a complete spliced three - channel audio, and then the spliced three - channel audio is subjected to corresponding audio encoding and encapsulation for output.
[0060] In summary, the present application discloses an audio splicing method. After obtaining the original two-dimensional channel audio and the original three-dimensional channel audio, the original two-dimensional channel audio is upmixed into the target three-dimensional channel audio through the HOA reconstruction upmixing technology. Among them, the channel type of the target three-dimensional channel audio is the same as that of the original three-dimensional channel audio. According to the audio insertion time point, the target three-dimensional channel audio and the original three-dimensional channel audio are spliced to generate the spliced three-dimensional channel audio. It can be seen that through the HOA reconstruction upmixing technology, the present application can upmix the original two-dimensional channel audio into high-quality target three-dimensional channel audio, thereby ensuring the audio quality of the spliced three-dimensional channel audio and avoiding audio quality problems in the subsequent rendering process.
[0061] See Figure 2 , which is a schematic flowchart of another audio splicing method provided by an embodiment of the present application. The method specifically includes:
[0062] S201. Obtain the original two-dimensional channel audio and the original three-dimensional channel audio;
[0063] S202. Determine the mono audio of the original two-dimensional channel audio;
[0064] S203. Through the HOA reconstruction upmixing technology, upmix each mono audio into three-dimensional channel audio;
[0065] S204. Perform signal superposition on the three-dimensional channel audio corresponding to each mono audio to generate the target three-dimensional channel audio. Among them, the channel type of the target three-dimensional channel audio is the same as that of the original three-dimensional channel audio;
[0066] S205. According to the audio insertion time point, splice the target three-dimensional channel audio and the original three-dimensional channel audio to generate the spliced three-dimensional channel audio.
[0067] In this embodiment, when upmixing the original two-dimensional channel audio into the target three-dimensional channel audio through the HOA reconstruction upmixing technology, it is necessary to determine the mono audio of the original two-dimensional channel audio. For example: in a 2.0-channel two-dimensional channel audio, the mono audio includes the left-channel audio and the right-channel audio, while in a 5.1-channel two-dimensional channel audio, the mono audio includes: the center-channel audio, the front left-channel audio, the front right-channel audio, the rear left-channel audio, the rear right-channel audio, and a subwoofer-channel audio.
[0068] Therefore, when performing the upmixing operation in this application, it is necessary to use the HOA reconstruction upmixing technology to upmix each monophonic audio into the corresponding three-dimensional channel audio, and then superimpose the signals of the three-dimensional channel audio corresponding to each monophonic audio to generate the target three-dimensional channel audio. For example, when upmixing a two-dimensional channel audio with 2.0 channels into a three-dimensional channel audio with 5.1.4 channels, it is necessary to upmix the left-channel audio of the two-dimensional channel audio into a three-dimensional channel audio with 5.1.4 channels, upmix the right-channel audio of the two-dimensional channel audio into a three-dimensional channel audio with 5.1.4 channels, and then superimpose the signals of the three-dimensional channel audio of the left-channel audio and the three-dimensional channel audio of the right-channel audio to generate the target three-dimensional channel audio with 5.1.4 channels after upmixing the original two-dimensional channel audio.
[0069] In another embodiment of this application, the process of upmixing a monophonic audio into a three-dimensional channel audio by using the HOA reconstruction upmixing technology specifically includes: determining the spatial harmonic signal of the monophonic audio; determining the decoding equation coefficients corresponding to the target channel type, where the target channel type is the channel type of the original three-dimensional channel audio; determining the reconstruction signals of each speaker in the target channel type according to the spatial harmonic signal and the decoding equation coefficients; and determining the three-dimensional channel audio of the monophonic audio according to the reconstruction signals of each speaker.
[0070] Specifically, when this application uses the HOA reconstruction upmixing technology to upmix a monophonic audio into a three-dimensional channel audio, it is necessary to decompose and reconstruct the monophonic audio as a point source to obtain the three-dimensional channel audio. In this embodiment, taking the example of mapping the original two-dimensional channel audio with 2.0 channels to the three-dimensional channel audio with 5.1.4 channels for illustration.
[0071] The L (left) channel audio and R (right) channel audio of the original two-dimensional channel audio are respectively used as two point sources. Among them, the L channel audio is an incident sound wave with a spatial coordinate of Ω S =(θ S , φ S ), an amplitude of S0, θ S is the azimuth angle, φ S is the elevation angle. In this application, the spatial coordinate is specifically Ω S =(0°, 30°); the L channel audio signal is broadened by the (L-1) -order spherical harmonic function to obtain a group of L 2 different-order ideal independent coding signals with pointing characteristics. In this application, this independent coding signal is called the spatial harmonic signal The spatial harmonic signal is only related to the sound source, and the spatial harmonic signals of the same signal are the same.
[0072] This spatial harmonic signal is:
[0073]
[0074] wherein, is a spherical harmonic function, and the normalized spherical harmonic function of the m-th order in real form can be obtained by the following formula:
[0075]
[0076] wherein, is the associated Legendre function; is a normalization factor, and the specific normalization factor is:
[0077]
[0078] When performing 5.1.4 channel reconstruction in this application, specifically, each spatial harmonic signal in the spatial harmonic signal is linearly mixed with different proportionality coefficients, and then fed back to each speaker at the corresponding position. This set of proportionality coefficients is the decoding equation coefficient B lm (Ω i ). This decoding equation coefficient is only related to the positions of the reconstructed speakers. For example, the decoding equation coefficients of 5.1.4 channels and the decoding equation coefficients of 7.1.4 channels need to be determined by their respective speaker distributions.
[0079] When determining the decoding equation coefficient, first let [Y 3D represent the L 2 *M spherical harmonic function matrix determined by the speaker distribution, and its matrix element is The rows of the matrix are arranged in the order of the order of the spherical harmonic function, and the columns of the matrix are arranged in the order of the speaker directions Ω1, Ω2,..., Ω M . The decoding matrix [B 3D can be obtained by inverse or pseudo-inverse:
[0080] [B 3D = pinv[Y 3D = [Y 3D T {[Y 3D [Y 3D T} -1
[0081] The matrix element of the decoding matrix is the decoding equation coefficient wherein, Ω i represents the standard spatial coordinates of the i-th speaker in 5.1.4 channels, i = 1, 2, 3,..., 10.
[0082] After determining the spatial harmonic signal of the L-channel audio and the decoding equation coefficient through the above process, the reconstructed signal of each speaker can be determined through the spatial harmonic signal and the decoding equation coefficient. The specific formula is:
[0083]
[0084] In this application, 2nd order reconstruction can be performed according to the HOA theory, so L is 3. After determining the reconstruction signals of each loudspeaker in this application, a three-dimensional channel audio signal of L-channel audio can be generated.
[0085] Then, the R channel is used as a point sound source at the spatial coordinates (0°, -30°), upmixed into a three-dimensional channel audio of 5.1.4 channels, and the three-dimensional channel audio signal after upmixing the L channel is superimposed on the three-dimensional channel audio signal after upmixing the R channel, so as to obtain the 5.1.4 channel audio after upmixing the original stereo, that is, the target three-dimensional channel audio.
[0086] In summary, this application can reconstruct the monophonic audio in the original two-dimensional channel audio into the corresponding three-dimensional channel audio through the spatial harmonic signal and the decoding equation coefficients, and after superimposing the three-dimensional channel audio corresponding to each monophonic audio, the target three-dimensional channel audio can be generated. Through the HOA reconstruction upmixing technology, this application can generate high-quality target three-dimensional channel audio, thus ensuring the audio quality of the spliced three-dimensional channel audio.
[0087] See Figure 3 , which is a schematic flow diagram of another audio splicing method provided by the embodiment of this application. The method specifically includes:
[0088] S301. Obtain the original two-dimensional channel audio and the original three-dimensional channel audio;
[0089] S302. Perform audio quality detection on the original two-dimensional channel audio and the original three-dimensional channel audio;
[0090] S303. If both detections pass, then through the HOA reconstruction upmixing technology, upmix the original two-dimensional channel audio into the target three-dimensional channel audio; wherein, the channel type of the target three-dimensional channel audio is the same as that of the original three-dimensional channel audio;
[0091] S304. Splice the target three-dimensional channel audio and the original three-dimensional channel audio according to the audio insertion time point to generate the spliced three-dimensional channel audio.
[0092] In this embodiment, in order to improve the quality of the spliced three-dimensional channel audio, quality detection can be performed on the original two-dimensional channel audio and the original three-dimensional channel audio after extracting them.
[0093] Specifically, the present application performs audio quality detection on the original two-channel audio and the original three-channel audio, which can detect audio quality problems such as silence, popping sound, mono-channel, volume imbalance, etc., and generates a detection report based on the detection results. If the detection result indicates that the quality of the original two-channel audio is poor, it is determined that the original two-channel audio fails the detection; if the detection result indicates that the quality of the original three-channel audio is poor, it is determined that the original three-channel audio fails the detection; that is: the present application only continues to execute the subsequent upmixing and splicing processes after both the original two-channel audio and the original three-channel audio pass the detection; if any one of the audio fails the detection, operations such as repair or replacement of the source audio need to be performed.
[0094] Among them, when the present application detects the audio quality of the original two-channel audio and the original three-channel audio, it can be through a machine learning model or a software detection method. For example: when performing silence detection, the waveform amplitude of the audio can be detected through audio software. If the sound intensity of the audio continuously falls below a predetermined threshold, it is determined as a silent segment.
[0095] In summary, after the present application extracts the original two-channel audio and the original three-channel audio, it can perform audio quality detection on the original two-channel audio and the original three-channel audio. Only after both audios pass the detection, the subsequent upmixing and splicing processes are executed. Through this method, the basic audio defects can be eliminated, the input reliability can be guaranteed, the ineffective processing of low-quality audio can be reduced, and the resource utilization rate can be optimized; moreover, through this method, the present application can also improve the audio quality from the source and avoid the secondary repair cost caused by input data defects.
[0096] See Figure 4 , which is a schematic flowchart of an audio splicing method provided by an embodiment of the present application. The method specifically includes:
[0097] S401. Obtain the original two-channel audio and the original three-channel audio;
[0098] S402. Through the HOA reconstruction upmixing technology, upmix the original two-channel audio into a target three-channel audio; among them, the channel type of the target three-channel audio is the same as that of the original three-channel audio;
[0099] S403. Perform volume normalization processing on the target three-channel audio to generate a first standard three-channel audio; perform volume normalization processing on the original three-channel audio to generate a second standard three-channel audio;
[0100] S404. Cut the second standard three-channel audio according to the audio insertion time point to generate a first cut three-channel audio and a second cut three-channel audio;
[0101] S405. Splice the first cut three-dimensional channel audio, the first standard three-dimensional channel audio, and the second cut three-dimensional channel audio in the order of audio playback to generate spliced three-dimensional channel audio.
[0102] It should be noted that in traditional audio splicing solutions, there will also be a problem of inconsistent audio volume after splicing. Therefore, in this embodiment, after obtaining the target three-dimensional channel audio of the original two-dimensional channel audio through the HOA reconstruction upmixing technology, it is also necessary to perform volume normalization processing on the target three-dimensional channel audio to generate the first standard three-dimensional channel audio; perform volume normalization processing on the original three-dimensional channel audio to generate the second standard three-dimensional channel audio; and then splice the processed first standard three-dimensional channel audio and the second standard three-dimensional channel audio according to the audio insertion time point to generate spliced three-dimensional channel audio.
[0103] In another embodiment of the present application, the process of performing volume normalization processing on the target three-dimensional channel audio to generate the first standard three-dimensional channel audio and performing volume normalization processing on the original three-dimensional channel audio to generate the second standard three-dimensional channel audio specifically includes: calculating the first target loudness value of the target three-dimensional channel audio through the audio loudness measurement standard, and adjusting the volume of the target three-dimensional channel audio according to the first target loudness value to generate the first standard three-dimensional channel audio; calculating the second target loudness value of the original three-dimensional channel audio through the audio loudness measurement standard, and adjusting the volume of the original three-dimensional channel audio according to the second target loudness value to generate the second standard three-dimensional channel audio.
[0104] In the present application, the audio loudness measurement standard may specifically be EBU R.128 (European Broadcasting Union Recommendation 128). EBU R.128 is an internationally publicized loudness measurement standard. Through EBU R.128, the problem of poor user experience caused by audio loudness differences can be solved. Therefore, the present application can calculate the first target loudness value of the target three-dimensional channel audio through the EBU R.128 audio loudness measurement standard, and adjust the target three-dimensional channel audio according to the first target loudness value, so as to output the first standard three-dimensional channel audio with a specified loudness; similarly, calculate the second target loudness value of the original three-dimensional channel audio through the EBU R.128 audio loudness measurement standard, and adjust the original three-dimensional channel audio according to the second target loudness value, so as to output the second standard three-dimensional channel audio with a specified loudness. When adjusting the loudness of the audio in the present application, the target loudness can be achieved through gain adjustment, and a compressor and a limiter are used to prevent popping.
[0105] Further, when splicing audio in the present application, the second standard three-dimensional channel audio can be cut according to the audio insertion time point to generate the first cut three-dimensional channel audio and the second cut three-dimensional channel audio; then, according to the audio playback order, the first cut three-dimensional channel audio, the first standard three-dimensional channel audio, and the second cut three-dimensional channel audio are spliced to generate the spliced three-dimensional channel audio.
[0106] See Figure 5 , which is a schematic diagram of the specific audio splicing process provided by the embodiment of the present application. In this embodiment, the original two-dimensional channel audio is the advertisement audio, and the original three-dimensional channel audio is the feature film audio; in the solution, first, audio extraction and audio quality detection need to be performed separately from the feature film and the advertisement. The feature film audio that passes the detection is: ZP_514.pcm, and the feature film audio is in 5.1.4 channels. The advertisement audio that passes the detection is GG_20.pcm, and the advertisement audio is in 2.0 channels; then, the advertisement audio GG_20.pcm is subjected to HOA upmixing to generate the 5.1.4 channel advertisement audio GG_514.pcm; after volume normalization of the feature film audio ZP_514.pcm, ZP_514_gain.w64 is generated, and after volume normalization of the advertisement audio GG_514.pcm, GG_514_gain.w64 is generated. Then, according to the insertion time point of the advertisement, the feature film audio ZP_514_gain.w64 is cut. If an advertisement needs to be inserted at 52.8s, the feature film audio ZP_514_gain.w64 is cut at 52.8s. The first cut three-dimensional channel audio generated after cutting is: ZP_514_gain_part1.w64, and the second cut three-dimensional channel audio ZP_514_gain_part2.w64. Then, the cut feature film audio is spliced with the 5.1.4 channel advertisement, and in chronological order, they are: ZP_514_gain_part1.w64, GG_514_gain.w64, ZP_514_gain_part2.w64; finally, the spliced audio merge.w64 is subjected to corresponding audio encoding and encapsulation output.
[0107] In order to accurately and high-quality splice the 2D sound advertisement audio into the 3D sound feature film source in the present application, the 2D sound advertisement audio is upmixed based on the reconstruction principle of HOA to make it a 3D sound audio with the same channel type as the feature film audio, and volume normalization processing is performed on it; then, volume normalization processing is also performed on the feature film audio, and the feature film is cut and the advertisement is inserted according to the time information, and finally spliced into a new 3D sound audio.
[0108] In summary, through the HOA reconstruction upmixing technology, the present application can upmix the original two-dimensional channel audio into high-quality target three-dimensional channel audio, and then splice it with the original three-dimensional channel audio, ensuring the audio quality and avoiding potential sound quality problems in subsequent rendering. Moreover, the upmixing algorithm based on HOA reconstruction in the present application is applicable to different 2D-3D channel mappings and has a high degree of sound field restoration. The two audio signals to be spliced in the present application are volume-normalized in the 3D channel mode, and the loudness consistency is good.
[0109] Furthermore, when applying this solution to the splicing scenario of advertisements and feature films, it can also solve the problem of online fault reporting. For example, after some film and television dramas are spliced with advertisements, there are problems such as inconsistent volumes between the feature film and the advertisement and several resulting sound quality problems. Moreover, with the increasing number of 3D sound sources currently available, this solution can be used as a standardized splicing process for such situations to improve audio quality and reduce user fault reporting.
[0110] See Figure 6 , Figure 6 which is a schematic structural diagram of an audio splicing device provided by an embodiment of the present application. The device specifically includes:
[0111] An acquisition module 11, configured to acquire the original two-dimensional channel audio and the original three-dimensional channel audio;
[0112] An upmixing module 12, configured to upmix the original two-dimensional channel audio into target three-dimensional channel audio through the HOA reconstruction upmixing technology; wherein, the channel type of the target three-dimensional channel audio is the same as that of the original three-dimensional channel audio;
[0113] A splicing module 13, configured to splice the target three-dimensional channel audio and the original three-dimensional channel audio according to the audio insertion time point to generate spliced three-dimensional channel audio.
[0114] As an optional embodiment, the upmixing module includes:
[0115] A determination unit, configured to determine the mono audio of the original two-dimensional channel audio;
[0116] An upmixing unit, configured to upmix each mono audio into three-dimensional channel audio through the HOA reconstruction upmixing technology;
[0117] A superimposing unit, configured to perform signal superimposition on the three-dimensional channel audio corresponding to each mono audio to generate target three-dimensional channel audio.
[0118] As an alternative embodiment, the upmixing unit includes: determining a spatial harmonic signal of the monophonic audio, and determining decoding equation coefficients corresponding to a target channel type; the target channel type being the channel type of the original three-dimensional channel audio; determining reconstructed signals of each speaker in the target channel type according to the spatial harmonic signal and the decoding equation coefficients; and determining three-dimensional channel audio of the monophonic audio according to the reconstructed signals of each speaker.
[0119] As an alternative embodiment, the audio splicing device further includes:
[0120] a detection unit configured to perform audio quality detection on the original two-dimensional channel audio and the original three-dimensional channel audio; and triggering the upmixing module if both detections pass.
[0121] As an alternative embodiment, the audio splicing device further includes:
[0122] a processing module configured to perform volume normalization processing on the target three-dimensional channel audio to generate a first standard three-dimensional channel audio; and perform volume normalization processing on the original three-dimensional channel audio to generate a second standard three-dimensional channel audio.
[0123] As an alternative embodiment, the processing module includes:
[0124] a first processing unit configured to calculate a first target loudness value of the target three-dimensional channel audio according to an audio loudness measurement standard, and perform volume adjustment on the target three-dimensional channel audio according to the first target loudness value to generate a first standard three-dimensional channel audio;
[0125] a second processing unit configured to calculate a second target loudness value of the original three-dimensional channel audio according to the audio loudness measurement standard, and perform volume adjustment on the original three-dimensional channel audio according to the second target loudness value to generate a second standard three-dimensional channel audio.
[0126] As an alternative embodiment, the splicing module includes:
[0127] a third processing unit configured to cut the second standard three-dimensional channel audio according to an audio insertion time point to generate a first cut three-dimensional channel audio and a second cut three-dimensional channel audio;
[0128] a splicing unit configured to splice the first cut three-dimensional channel audio, the first standard three-dimensional channel audio, and the second cut three-dimensional channel audio according to an audio playback order to generate a spliced three-dimensional channel audio.
[0129] Regarding the device in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated herein.
[0130] See Figure 7 , Figure 7 Figure 7 is a schematic structural diagram of an electronic device provided by an embodiment of the present application, including a processor 21, a communication interface 22, a memory 23, and a communication bus 24. Among them, the processor 21, the communication interface 22, and the memory 23 complete mutual communication through the communication bus 24;
[0131] The memory 23 is used to store a computer program;
[0132] When the processor 21 is used to execute the program stored on the memory 23, it implements the steps of the audio splicing method described in any of the above method embodiments, which will not be elaborated here.
[0133] The communication bus mentioned in the above terminal may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience of representation, Figure 7 only a thick line is used to represent it in Figure 7 , but it does not mean that there is only one bus or one type of bus.
[0134] The communication interface is used for communication between the above terminal and other devices.
[0135] The memory may include a Random Access Memory (RAM), or may also include a non-volatile memory, such as at least one disk memory. Optionally, the memory may also be at least one storage device located far from the aforementioned processor.
[0136] The above-mentioned processor may be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it may also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0137] In another exemplary embodiment, a computer storage medium is further provided. When the program instructions are executed by a processor, the steps of the audio splicing method described in any of the above method embodiments are implemented. Among them, the storage medium may include: various media such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs that can store program codes.
[0138] Optionally, specific examples in this embodiment may refer to the examples described in the above embodiments, and details are not repeated herein.
[0139] It should be understood that the terms used herein are for the purpose of describing specific example embodiments only and are not intended to be limiting. Unless the context clearly indicates otherwise, the singular forms "a", "an", and "the" as used herein may also include the plural forms. The terms "include", "comprise", "contain", and "have" are inclusive and thus specify the presence of the stated features, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, steps, operations, elements, components, and / or their combinations. The method steps, processes, and operations described herein are not to be construed as necessarily requiring them to be performed in the particular order described or illustrated, unless the order of performance is explicitly stated. It should also be understood that additional or alternative steps may be used.
[0140] The above description is only the specific implementation manners of the present invention, enabling those skilled in the art to understand or implement the present invention. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but rather will be accorded the widest scope consistent with the principles and novel features claimed herein.
Claims
1. An audio splicing method, characterized in that, Including: Obtaining the original two - dimensional channel audio and the original three - dimensional channel audio; By using the HOA reconstruction up - mixing technology, up - mixing the original two - dimensional channel audio into a target three - dimensional channel audio; wherein, the channel type of the target three - dimensional channel audio is the same as that of the original three - dimensional channel audio; Splicing the target three - dimensional channel audio and the original three - dimensional channel audio according to the audio insertion time point to generate a spliced three - dimensional channel audio.
2. The audio splicing method according to claim 1, wherein By using the HOA reconstruction up - mixing technology, up - mixing the original two - dimensional channel audio into a target three - dimensional channel audio includes: Determining the mono - channel audio of the original two - dimensional channel audio; By using the HOA reconstruction up - mixing technology, up - mixing each mono - channel audio into a three - dimensional channel audio; Performing signal superposition on the three - dimensional channel audio corresponding to each mono - channel audio to generate a target three - dimensional channel audio.
3. The audio splicing method according to claim 2, wherein By using the HOA reconstruction up - mixing technology, up - mixing each mono - channel audio into a three - dimensional channel audio, including: Determining the spatial harmonic signal of the mono - channel audio; Determining the decoding equation coefficients corresponding to the target channel type; the target channel type is the channel type of the original three - dimensional channel audio; According to the spatial harmonic signal and the decoding equation coefficients, determining the reconstruction signals of each speaker in the target channel type; Determining the three - dimensional channel audio of the mono - channel audio according to the reconstruction signals of each speaker.
4. The audio splicing method according to claim 1, characterized in that After obtaining the original two - dimensional channel audio and the original three - dimensional channel audio, it further includes: Performing audio quality detection on the original two - dimensional channel audio and the original three - dimensional channel audio; If both pass the detection, then continue to execute the step of up - mixing the original two - dimensional channel audio into a target three - dimensional channel audio by using the HOA reconstruction up - mixing technology.
5. The audio splicing method according to any one of claims 1 to 4, characterized in that, Before splicing the target three - dimensional channel audio and the original three - dimensional channel audio according to the audio insertion time point, it further includes: Performing volume normalization processing on the target three - dimensional channel audio to generate a first standard three - dimensional channel audio; performing volume normalization processing on the original three - dimensional channel audio to generate a second standard three - dimensional channel audio.
6. The audio splicing method according to claim 5, wherein Performing volume normalization processing on the target three - dimensional channel audio to generate a first standard three - dimensional channel audio; Performing volume normalization processing on the original three - dimensional channel audio to generate a second standard three - dimensional channel audio includes: Calculating the first target loudness value of the target three - dimensional channel audio through the audio loudness measurement standard, and adjusting the volume of the target three - dimensional channel audio according to the first target loudness value to generate a first standard three - dimensional channel audio; Calculating the second target loudness value of the original three - dimensional channel audio through the audio loudness measurement standard, and adjusting the volume of the original three - dimensional channel audio according to the second target loudness value to generate a second standard three - dimensional channel audio.
7. The audio splicing method according to claim 5, characterized in that, Splicing the target three - dimensional channel audio and the original three - dimensional channel audio according to the audio insertion time point to generate a spliced three - dimensional channel audio, including: Performing a cutting process on the second standard three - dimensional channel audio according to the audio insertion time point to generate a first cut three - dimensional channel audio and a second cut three - dimensional channel audio; Splice the first cut three-dimensional channel audio, the first standard three-dimensional channel audio, and the second cut three-dimensional channel audio according to the audio playback order to generate spliced three-dimensional channel audio.
8. An audio splicing device, characterized in that, It includes: An acquisition module, configured to acquire original two-dimensional channel audio and original three-dimensional channel audio; An upmixing module, configured to upmix the original two-dimensional channel audio into target three-dimensional channel audio through HOA reconstruction upmixing technology; wherein, the channel type of the target three-dimensional channel audio is the same as that of the original three-dimensional channel audio; A splicing module, configured to splice the target three-dimensional channel audio and the original three-dimensional channel audio according to the audio insertion time point to generate spliced three-dimensional channel audio.
9. An electronic device, characterized in that, It includes a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete mutual communication through the communication bus; The memory is used to store computer programs; The processor is configured to implement the steps of the audio splicing method described in any one of claims 1 to 7 when executing the program stored on the memory.
10. A computer storage medium, characterized in that, The computer storage medium stores computer-executable instructions, and the computer-executable instructions are used to execute the steps of the audio splicing method described in any one of claims 1 to 7 of the present application.