A method, device, and storage medium for multimedia data synthesis

By synthesizing sound effects signals in three-dimensional videos and generating and synthesizing sound effects signals based on the spatial position of the sound image, the problem of poor sense of sound spatial position and time synchronization in three-dimensional videos is solved, and a more realistic and immersive viewing is achieved.

CN114630145BActive Publication Date: 2025-06-10TENCENT MUSIC ENTERTAINMENT TECH (SHENZHEN) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210264309.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-17
Publication Date
2025-06-10
Estimated Expiration
2042-03-17

AI Technical Summary

Technical Problem

The existing three-dimensional video technology has failed to effectively simulate the spatial orientation and sound and picture synchronization of sound, resulting in a lack of spatial positional sense of sound in video and poor time synchronization.

Method used

By obtaining three-dimensional video, determine the spatial location of the target video frame and its audio image to be synthesized, generate the corresponding sound effect signal, and synthesize it with the target video frame to generate a synthetic video frame, and finally obtain a new three-dimensional video based on the synthetic video frame and the original video frame.

Benefits of technology

The sound in three-dimensional video has a sense of spatial orientation and is kept in time synchronized with the video frame, improving the immersive viewing of three-dimensional video.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114630145B_ABST
    Figure CN114630145B_ABST
Patent Text Reader

Abstract

The present application discloses a method, device, and storage medium for multimedia data synthesis. The multimedia data synthesis method provided by the present application includes: obtaining a three-dimensional video; for the video frames in the three-dimensional video that require sound effects synthesis, determining the spatial position of the sound image included in the video frame, generating a sound effect signal for the spatial position, and synthesizing the sound effect signal with the video frame to obtain a synthesized video frame; obtaining a new three-dimensional video based on the synthesized video frames and the original video frames in the three-dimensional video that do not require sound effects synthesis. The sound effect signal in the finally obtained synthesized video frame in this solution has a sense of spatial orientation and is synchronized with the video frame in terms of time. Correspondingly, the multimedia data synthesis device and storage medium provided by the present application also have the above technical effects.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and particularly to a method, device, and storage medium for multimedia data synthesis. Background Art

[0002] Currently, 3D videos in virtual scenarios only focus on simulating real 3D scenes, without considering features such as the authenticity of the sound in the video, the coordination and synchronization between the sound and the picture, resulting in out-of-sync sound and picture in 3D videos and poor spatial orientation of the sound. Summary of the Invention

[0003] In view of this, the purpose of this application is to provide a method, device, and storage medium for multimedia data synthesis, so that the sound in the 3D video has a sense of spatial orientation and the sound and picture are synchronized. The specific solutions are as follows:

[0004] To achieve the above purpose, on the one hand, this application provides a method for multimedia data synthesis, including:

[0005] Obtain a 3D video;

[0006] Determine the target video frame in the 3D video that needs to synthesize sound effects, and determine the spatial position of the target sound image in the target video frame, and generate a sound effect signal of the target sound image at the spatial position;

[0007] Synthesize the sound effect signal with the target video frame to obtain a synthesized video frame;

[0008] Based on the synthesized video frame and the original video frames in the 3D video that do not need to synthesize sound effects, obtain a new 3D video.

[0009] Optionally, generating the sound effect signal of the target sound image at the spatial position includes:

[0010] Obtain the target audio corresponding to the target sound image, and encode the target audio based on the spatial position to obtain the sound effect signal.

[0011] Optionally, encoding the target audio based on the spatial position to obtain the sound effect signal includes:

[0012] Determine each encoding channel for encoding the target audio;

[0013] Based on the spatial position, determine the signals of the target audio in each encoding channel;

[0014] Summarize the signals of each encoding channel to obtain the sound effect signal.

[0015] Optionally, it further includes:

[0016] If the sound effect signal in the synthesized video frame is reproduced through a spatially distributed loudspeaker array, the sound effect signal is decoded based on the loudspeaker array, and the decoded signal is played using the loudspeaker array.

[0017] Optionally, the decoding of the sound effect signal based on the loudspeaker array includes:

[0018] Constructing a signal matrix based on the number of loudspeakers in the loudspeaker array and the number of encoding channels;

[0019] Taking the pseudo-inverse matrix of the signal matrix as the decoding matrix;

[0020] Decoding the signals of each encoding channel based on the decoding matrix.

[0021] Optionally, the number of loudspeakers in the loudspeaker array is not less than the number of encoding channels, and satisfies H=(N + 1) 2 ; H is the number of encoding channels, and N is the encoding order.

[0022] Optionally, the decoding of the signals of each encoding channel based on the decoding matrix includes:

[0023] Decoding the signals of each encoding channel according to a target formula; the target formula is: D = A×[A 1 ,A 2 ,…,A H T , D is the decoding result, A is the decoding matrix, A 1 ,A 2 ,…,A H represents the signals of H encoding channels, and H is the number of encoding channels.

[0024] Optionally, the determining of the spatial position of the target sound image in the target video frame includes:

[0025] Taking the object that perceives the target sound image in the target video frame as a reference object, and determining the azimuth angle and elevation angle of the target sound image.

[0026] Optionally, it further includes:

[0027] If the sound effect signal in the synthesized video frame is reproduced through headphones, the sound effect signal is decoded based on a spatially distributed loudspeaker array, and the decoded signal is encoded into a left-channel signal and a right-channel signal, and the left-channel signal and the right-channel signal are played using the headphones.

[0028] On the other hand, the present application also provides a multimedia data synthesis method, including:

[0029] Obtaining a three-dimensional image; ​

[0030] Determine the target object in the three-dimensional image for which a sound effect needs to be synthesized, and determine the spatial position of the target object in the three-dimensional image;

[0031] Generate a sound effect signal for the target object at the spatial position based on the spatial position;

[0032] Synthesize the sound effect signal with the three-dimensional image to obtain a three-dimensional synthesized image.

[0033] Optionally, it further includes:

[0034] Obtain a three-dimensional video based on multiple of the three-dimensional synthesized images.

[0035] In another aspect, the present application also provides an electronic device, which includes a processor and a memory; wherein, the memory is used to store a computer program, and the computer program is loaded and executed by the processor to implement the foregoing multimedia data synthesis method.

[0036] In another aspect, the present application also provides a storage medium, in which computer-executable instructions are stored, and when the computer-executable instructions are loaded and executed by a processor, the foregoing multimedia data synthesis method is implemented.

[0037] The multimedia data synthesis method provided by the present application includes: obtaining a three-dimensional video; determining a target video frame in the three-dimensional video for which a sound effect needs to be synthesized, and determining the spatial position of the target sound image in the target video frame, and generating a sound effect signal for the target sound image at the spatial position; synthesizing the sound effect signal with the target video frame to obtain a synthesized video frame; obtaining a new three-dimensional video based on the synthesized video frame and the original video frames in the three-dimensional video that do not need to synthesize sound effects.

[0038] It can be seen that for the video frames in the three-dimensional video that need to synthesize sound effects, the present application can generate a sound effect signal at the corresponding spatial position according to the spatial position of the sound image therein, and synthesize the sound effect signal with the video frame, so that the sound effect signal in the finally obtained synthesized video frame can have a sense of spatial orientation and be synchronized with the video frame in time. Therefore, the sound in the new three-dimensional video obtained based on each synthesized video frame and the original video frames in the three-dimensional video that do not need to synthesize sound effects has a sense of spatial orientation and the sound and picture are synchronized.

[0039] Correspondingly, the multimedia data synthesis device and storage medium provided by the present application also have the above technical effects. Description of the Drawings

[0040] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only the embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on the provided drawings.

[0041] Figure 1 Schematic diagram of a physical architecture applicable to the present application provided by the present application;

[0042] Figure 2 Flowchart of a multimedia data synthesis method provided by the present application;

[0043] Figure 3 Schematic diagram of a spatial position provided by the present application;

[0044] Figure 4 Schematic diagram of the spatial distribution of a speaker array provided by the present application;

[0045] Figure 5 Flowchart of a sound rendering method in a three-dimensional video provided by the present application;

[0046] Figure 6 Projection display diagram of a three-dimensional video provided by the present application;

[0047] Figure 7 Flowchart of a synthesis method of three-dimensional images and sounds provided by the present application;

[0048] Figure 8 Flowchart of a three-dimensional video audio effect synthesis method provided by the present application;

[0049] Figure 9 Server structure diagram provided by the present application;

[0050] Figure 10 Terminal structure diagram provided by the present application. Detailed implementation manners

[0051] The following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments of the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application. In addition, in the embodiments of the present application, "first", "second", etc. are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence.

[0052] Existing 3D videos only focus on simulating real 3D scenes, without considering features such as the authenticity of the sound in the video, the coordination and synchronization between the sound and the picture, resulting in out-of-sync sound and picture in 3D videos and poor spatial orientation of the sound.

[0053] In view of the above existing problems, the present application proposes a multimedia data synthesis solution, which can endow the sound in 3D videos with spatial orientation and keep the sound and picture in sync.

[0054] For the sake of easy understanding, the physical framework applicable to the present application will be introduced first.

[0055] It should be understood that the multimedia data synthesis method provided by the present application can be applied to a system or program with multimedia data synthesis function. Specifically, a system or program with multimedia data synthesis function can run on devices such as servers and personal computers.

[0056] As Figure 1 shown, Figure 1 is a schematic diagram of the physical architecture applicable to the present application. In Figure 1 , a system or program with multimedia data synthesis function can run on a server, which obtains 3D videos from other terminal devices through a network; determines the target video frames in the 3D videos that need to synthesize sound effects, and determines the spatial positions of the target sound images in the target video frames, generating sound effect signals for the target sound images at the spatial positions; synthesizes the sound effect signals with the target video frames to obtain synthesized video frames; and obtains a new 3D video based on the synthesized video frames and the original video frames in the 3D videos that do not need to synthesize sound effects.

[0057] As can be seen from Figure 1 , the server can establish communication connections with multiple devices and obtain 3D videos from these devices. The server synthesizes corresponding sound effect signals for this 3D video to obtain a new 3D video.

[0058] Figure 1 shows a variety of terminal devices. In actual scenarios, there can be more or fewer types of terminal devices participating in the multimedia data synthesis process. The specific quantity and types depend on the actual scenario and are not limited here. In addition, Figure 1 shows one server, but in actual scenarios, multiple servers can also participate. The specific number of servers depends on the actual scenario.

[0059] It should be noted that the multimedia data synthesis method provided in this embodiment can be performed offline, that is, the server locally stores 3D videos and audio to be used for synthesizing sound effect signals, and it can directly use the solution provided by the present application to generate a new 3D video.

[0060] It can be understood that the above system and program with multimedia data synthesis function can be regarded as a kind of cloud service program. Its specific operation mode depends on the actual scenario and will not be limited here.

[0061] Specifically, after the multimedia data synthesis is completed, the obtained new three-dimensional video can be used for 3D game production, VR (Virtual Reality) scene production, film and television drama production, etc. Of course, the new three-dimensional video can be projected in the three-dimensional space to exhibit the synthesized three-dimensional video, truly achieving an immersive playback effect.

[0062] Combining the above commonalities, please refer to Figure 2 , Figure 2 which is a flowchart of a multimedia data synthesis method provided by an embodiment of this application. As Figure 2 shown, the multimedia data synthesis method may include the following steps:

[0063] S201. Obtain a three-dimensional video.

[0064] In this embodiment, the three-dimensional video may be a three-dimensional virtual animation video, a virtual game video, etc. The three-dimensional video may be an audio video or a silent video.

[0065] S202. Determine the target video frames in the three-dimensional video that need to synthesize sound effects, and determine the spatial positions of the target sound images in the target video frames, and generate sound effect signals of the target sound images at the spatial positions.

[0066] Generally speaking, for these virtual videos such as three-dimensional virtual animation videos and virtual game videos, the sound effect signals can only be obtained through post-configuration. It can be seen that there may be video frames that need to synthesize sound effects in various three-dimensional videos, and generally there is more than one video frame that needs to synthesize sound effects. Each video frame that needs to synthesize sound effects in a three-dimensional video can be used as a target video frame. Usually, in order to reduce the video synthesis time, only some video frames with target sound images are selected. Preferably, the start and end video frames where the target sound image is located and some intermediate video frames are selected at intervals or certain specific video frames are selected for sound effect synthesis. It can be understood that the more video frames are selected, the more realistic the sound effects of the finally synthesized video will be, and the corresponding workload will also be greater.

[0067] Since the 3D video can be an audio video or a silent video, for an audio video, the target video frame is a 3D image frame with sound, while for a silent video, the target video frame is a 3D image frame without sound. Correspondingly, for a 3D image frame with sound, when generating a sound effect signal at the corresponding spatial position for the target sound image therein, it can be directly obtained by encoding the sound of the frame corresponding to the 3D image frame. For a 3D image frame without sound, when generating a sound effect signal at the corresponding spatial position for the target sound image therein, it is necessary to first determine the sound of the frame corresponding to the 3D image frame, and then the sound can be encoded.

[0068] Considering that the sound source may appear at any position in the 3D space and it is necessary to maintain lip-sync in the video, in this embodiment, for any video frame in the 3D video that needs to synthesize the sound effect, the spatial position where the sound image included in the video frame is located is determined, and a sound effect signal at this spatial position is generated, so that the sound effect signal can have a sense of spatial orientation, and then the sound effect signal is synthesized with the video frame, so that the resulting synthesized video frame can maintain lip-sync. Among them, the sound image is: the sound source or the perceived sound source, that is, the sound source felt by the listener in the sense of hearing.

[0069] See Figure 3 As shown, for Figure 3 the cube structure shown, the sound source may be in the front, the back, the connection line between the front surface and the top surface where the front is located, etc. Suppose Figure 3 is a 3D image frame, then the target sound image may be in the front, the back, the connection line between the front surface and the top surface where the front is located, etc. It can be seen that there can be multiple target sound images in a 3D image frame.

[0070] If the Figure 3 cube shown is regarded as a house, and it is assumed that the front of the house is a street, then there may need to appear vehicle honking sounds, people talking sounds, street vendor shouting sounds, etc. in the front, and these sounds may need to move from far to near or from near to far. Correspondingly, the sound images of these sounds need to move from far to near or from near to far in space. Taking the vehicle honking sound moving from far to near as an example, there is: the sound of the synthesized sound effect signal in video frame 1 is smaller and the sense of hearing is located far away from the house, the sound of the synthesized sound effect signal in video frame 2 is a little louder and the sense of hearing is located closer to the house, the sound of the synthesized sound effect signal in video frame 3 is even louder and the sense of hearing is located even closer to the house. Playing video frames 1, 2, and 3 continuously like this can produce the feeling of the vehicle honking sound moving from far to near.

[0071] Correspondingly, if a bird flies over the roof, then bird calls, the sound of the bird flapping its wings, etc. may need to appear at the roof position. Therefore, in video frames 1, 2, and 3, the sounds of the bird flying over, bird calls, and the bird flapping its wings can also be synthesized. It can be seen that there is more than one target sound image in a video frame, and there is more than one audio effect signal that needs to be synthesized. Moreover, the audio effect signal with a sense of spatial orientation is more in line with the actual scene. Correspondingly, the time when the above-mentioned sounds appear also needs to be controlled. For this reason, in this embodiment, audio effect signals are synthesized for each video frame with a specific timestamp. While synthesizing the audio effect signals and the video frames, the synchronization of sound and picture is ensured.

[0072] Of course, there may be more than one audio effect signal that needs to be synthesized for each video frame, which corresponds to the characters, scenes, etc. in the video frame. That is: there may be more than one target sound image in a target video frame, then audio effect synthesis needs to be performed for each target sound image in a target video frame based on its spatial position. In a specific implementation manner, determining the spatial position where the sound image included in the video frame is located includes: taking the object that perceives the target sound image in the target video frame as a reference object to determine the azimuth angle and elevation angle of the target sound image. Generally, the spatial position where the target sound image is located can be determined by the coordinate position of the target sound image in the video frame. Of course, it is necessary to first determine the coordinate position of the object (such as a person in three-dimensional space) that perceives the target sound image in the video frame. Taking the coordinate position of this object in the video frame as the origin, the azimuth angle and elevation angle of the target sound image in the video frame can be determined.

[0073] S203. Synthesize the audio effect signal and the target video frame to obtain a synthesized video frame.

[0074] S204. Obtain a new three-dimensional video based on the synthesized video frame and the original video frame in the three-dimensional video that does not require audio effect synthesis.

[0075] Among them, the original video frame in the three-dimensional video that does not require audio effect synthesis can be with sound or without sound. That is: the original video frame that does not require audio effect synthesis includes: a frame of video without sound and a frame of video with sound but without the need for audio effect synthesis.

[0076] In this embodiment, the audio effect signal can be reproduced either by a speaker array or by headphones.

[0077] It can be seen that in this embodiment, for the video frames in the three-dimensional video that need to synthesize audio effects, audio effect signals corresponding to the spatial positions can be generated according to the spatial positions where the sound images are located, and the audio effect signals and the video frames are synthesized. As a result, the audio effect signals in the finally obtained synthesized video frames can have a sense of spatial orientation and are synchronized with the video frames in terms of time. Therefore, the sounds in the new three-dimensional video obtained based on each synthesized video frame and the original video frames in the three-dimensional video that do not require audio effect synthesis have a sense of spatial orientation and the sound and picture are synchronized.

[0078] Based on the above embodiments, it should be noted that in a specific implementation manner, generating a sound effect signal for a spatial position includes: obtaining a target audio corresponding to a target sound image, and encoding the target audio based on the spatial position to obtain a sound effect signal. Among them, encoding the target audio based on the spatial position to obtain a sound effect signal includes: encoding the target audio using the Ambisonics technology to obtain a sound effect signal.

[0079] In a specific implementation manner, encoding the target audio based on the spatial position to obtain a sound effect signal includes: determining each encoding channel for encoding the target audio; determining the signals of the target audio in each encoding channel based on the spatial position; and summarizing the signals of each encoding channel to obtain a sound effect signal. This process is the Ambisonics encoding process. Among them, the "signals of each encoding channel" can be regarded as the signal representation form of the sound effect signal, that is: the sound effect signal is not a single signal, but a set of signals of each encoding channel.

[0080] It should be noted that the encoding stage does not depend on any speakers or their distribution. As long as the sound image position (i.e., the spatial position where the sound image is located) and the encoding complexity (i.e., how many encoding channels are used for encoding) are known, then on the premise of knowing the spatial position where the sound image is located, it is only necessary to clarify each encoding channel currently used for encoding the target audio. Generally, the number of encoding channels can be flexibly selected. When specifically implemented, the speaker array arranged in the real scene for playing the new 3D video can be considered, as long as "the number of speakers in the speaker array is not less than the number of encoding channels". Of course, the speaker array arranged in the real scene can also be adjusted according to the number of encoding channels used for encoding to meet the above requirements.

[0081] Among them, the number of speakers in the speaker array arranged in the real scene for playing the new 3D video is not less than the number of encoding channels, and satisfies H=(N + 1) 2 ; H is the number of encoding channels, and N is the encoding order. Among them, the speaker array can be any spatial distribution. For example: the speakers can be distributed at Figure 3 each vertex of the cube shown. At this time, the speaker array includes a total of 8 speakers, and the spatial positions of these 8 speakers can be expressed as: azimuth angle: [45°, -45°, 135°, -135°, 45°, -45°, 135°, -135°], elevation angle: [35.3°, 35.3°, 35.3°, 35.3°, -35.3°, -35.3°, -35.3°, -35.3°]. Of course, the speakers can be distributed at Figure 4 each vertex of the regular dodecahedron shown. At this time, the speaker array includes a total of 20 speakers.

[0082] In a specific embodiment, if the sound effect signal in the synthesized video frame is reproduced through a spatially distributed loudspeaker array, the sound effect signal is decoded based on the loudspeaker array, and the decoded signal is played using the loudspeaker array. Among them, decoding the sound effect signal based on the loudspeaker array includes: constructing a signal matrix based on the number of loudspeakers in the loudspeaker array and the number of encoding channels; taking the pseudo-inverse matrix of the signal matrix as the decoding matrix; decoding the signals of each encoding channel based on the decoding matrix. Among them, decoding the signals of each encoding channel based on the decoding matrix includes: decoding the signals of each encoding channel according to the target formula; the target formula is: D = A × [A 1 , A 2 , …, A H T , D is the decoding result, A is the decoding matrix, A 1 , A 2 , …, A H represents the signals of H encoding channels, and H is the number of encoding channels.

[0083] Since the headphones reproduce sound through the left and right channels, in a specific embodiment, if the sound effect signal in the synthesized video frame is reproduced through the headphones, the sound effect signal is decoded based on a spatially distributed loudspeaker array, and the decoded signal is encoded into a left-channel signal and a right-channel signal, and the left-channel signal and the right-channel signal are played using the headphones. Among them, the decoded signal can be encoded into a left-channel signal and a right-channel signal by using HRTF (Head Related Transfer Function, a sound effect signal localization algorithm).

[0084] The following embodiments perform sound rendering for three-dimensional videos. This solution can determine the spatial position of the sound source in any three-dimensional video frame in real time, and use Ambisonics technology to encode the sound signal emitted by the sound source into a sound effect signal with a sense of spatial position, and this sense of spatial position changes with the change of the sound source position. This sound effect signal can be reproduced either by a loudspeaker array or by headphones. If reproduced by headphones, the head-related transfer function in HRTF is used to perform channel processing on the sound effect signal obtained by Ambisonics encoding.

[0085] This embodiment uses Ambisonics technology for encoding audio signals. Ambisonics technology is a spherical surround sound technology and also a codec algorithm. Its physical essence is to decompose, expand, and approximate the sound field according to spatial harmonics of different orders. Among them, the higher the order, the more accurate the approximate reproduction of the physical sound field. The relationship between the order N and the number of Ambisonics channels is: the number of Ambisonics channels = (N + 1) 2 . Here, the encoding is not audio compression encoding, but encoding an audio object into an Ambisonics format audio.​

[0086] Taking the first-order Ambisonics B format as an example, there are a total of 4 channels, and the channel order is W, Y, Z, X. Assuming that a sound needs to be emitted from the spatial position (θ, φ), where θ represents the azimuth angle and φ represents the elevation angle, then the sound object S can be encoded as a 4-channel signal: W = S, Y = S * sinθ * cosφ, Z = S * sinφ, X = S * cosθ * cosφ.

[0087] If it is the third order, the sound object S is encoded into a 16-channel signal: W = S, Y = S * sinθ * cosφ, Z = S * sinφ, X = S * cosθ * cosφ,

[0088]

[0089]

[0090]

[0091]

[0092] For the encoded signal, it can be reproduced either using a loudspeaker array or using headphones. Since the number of channels grows exponentially with the order, to avoid the loudspeaker array used in actual reproduction being too complex, generally up to the third-order Ambisonics is used. If a loudspeaker array is used for reproduction, the requirement for the number of loudspeakers in the loudspeaker array is greater than or equal to (N + 1) 2 .

[0093] For first-order Ambisonics, the loudspeaker array can be as Figure 3 shown, with loudspeakers set at each vertex of a regular hexahedron, for a total of 8 loudspeakers. Specifically, the spatial positions of these 8 loudspeakers can be expressed as: azimuth angles: [45°, -45°, 135°, -135°, 45°, -45°, 135°, -135°]; elevation angles: [35.3°, 35.3°, 35.3°, 35.3°, -35.3°, -35.3°, -35.3°, -35.3°].

[0094] For third-order Ambisonics, there are 16 channels after encoding. At this time, a spherical loudspeaker array of a regular dodecahedron can be used, as Figure 4 shown, with a total of 20 loudspeakers.

[0095] 1. Use a loudspeaker array for reproduction.

[0096] After determining the loudspeaker array, taking first-order Ambisonics as an example, if a regular hexahedron spatial loudspeaker array is used for playback, the 4×8 signal matrix composed of the direction functions of each loudspeaker is as follows:

[0097]

[0098] where θ represents the azimuth angle and φ represents the elevation angle. Taking the pseudo-inverse of Y can obtain the 8×4 decoding matrix A, that is: A = pinv(Y) = Y T {YY T} -1 。

[0099] Decoding is to multiply the signals on the encoded 4 channels by the decoding matrix A to obtain 8 loudspeaker signals: D = [d1, d2, …, d8], that is: D = A * [W, Y, Z, X] T 。

[0100] 2. Use headphones for playback.

[0101] Regarding the above loudspeaker array as a virtual loudspeaker array, the same above process is used for encoding. For the encoded D = [d1, d2, …, d8], convolution using the head-related transfer function in HRTF is performed to obtain two-channel signals.

[0102] Specifically, the left-channel signal L = d1 (45°,35.3°) *HRTF_L(45°, 35.3°) + d2 (-45°,35.3°) *HRTF_L(-45°, 35.3°) + … + d8 (-135°,-35.3°) *HRTF_L(-135°, -35.3°); HRTF_L represents the HRTF from a certain spatial position to the left ear.

[0103] The right-channel signal R = d1 (45°,35.3°) *HRTF_R(45°, 35.3°) + d2 (-45°,35.3°) *HRTF_R(-45°, 35.3°) + … + d8 (-135°,-35.3°) *HRTF_R(-135°, -35.3°). HRTF_R represents the HRTF from a certain spatial position to the right ear, thereby virtualizing the spatial position of the sound.

[0104] Please refer to Figure 5 , and the sound rendering steps in 3D videos can include:

[0105] 1. Obtain a 3D video;

[0106] 2. Determine the spatial positions of the sound sources included in each frame of the 3D video;

[0107] 3. Using Ambisonics, encode the sounds emitted by each sound source determined in step 2 based on their spatial positions.

[0108] 4. Synthesize the encoded results into each frame of the image correspondingly to obtain a new 3D video.

[0109] 5. Play and project the new 3D video, and at the same time play the sound effects in it using a speaker array or headphones.

[0110] As Figure 6 shown, in a 3D video projection exhibition hall, there is an animated picture of flowing water directly in front at a certain moment. At this time, the sound direction of the flowing water can be determined as directly in front. There are animated pictures of wind blowing, rain falling, and birds chirping on the right. Then, the occurrence positions and times of the wind blowing, rain falling, and birds chirping can be determined. The sound position is generally represented by azimuth and elevation angles. As Figure 6 shown, playing the synthesized new 3D video in a 3D projection exhibition hall can truly achieve an immersive effect.

[0111] It can be seen that this embodiment can render and play each sound in combination with the real-time positions of each sound in the 3D video, and can be played using a speaker array or headphones, so that the occurrence times, positions, and pictures of each sound in the picture are kept synchronized and coordinated, thereby enabling the 3D video to have an immersive visual experience.

[0112] Please refer to Figure 7 , another method for synthesizing multimedia data, including:

[0113] S701. Obtain a 3D image;

[0114] S702. Determine the target object in the 3D image that needs to synthesize sound effects, and determine the spatial position of the target object in the 3D image;

[0115] S703. Generate a sound effect signal of the target object at the spatial position based on the spatial position;

[0116] S704. Synthesize the sound effect signal with the 3D image to obtain a 3D synthesized image.

[0117] Among them, the target object in the 3D image that needs to synthesize sound effects is: the sound source that emits sound in the image, that is, the target sound image described in the above embodiment.

[0118] In a specific implementation manner, after synthesizing sound effects for multiple 3D images respectively according to this embodiment, a 3D video can be obtained based on the multiple 3D synthesized images. A 3D image in this embodiment can be regarded as a target video frame in the above embodiment.

[0119] In this embodiment, for the target object in the 3D image that requires synthesized sound effects, sound effect signals corresponding to their spatial positions can be generated according to their spatial positions, and the sound effect signals and the 3D image can be synthesized, so as to obtain a 3D synthesized image, where the sound effect signals can have a sense of spatial orientation. Based on this, a 3D video can be synthesized to obtain a 3D video with synchronized sound and picture.

[0120] The following describes through specific application scenario examples to introduce the solution provided by this application. That is: the specific solution for synthesizing sound effects and 3D videos. This solution can synthesize sound effects with a sense of spatial orientation for any 3D video.

[0121] Please refer to Figure 8 , the specific implementation process of the solution includes:

[0122] S801. The terminal requests the server;

[0123] S802. The server sends a response message to the terminal;

[0124] S803. After receiving the response message, the terminal transmits the 3D video to the server;

[0125] S804. The server determines the spatial positions of the sound sources included in each frame of the 3D video; encodes the sounds emitted by the determined sound sources based on their spatial positions using Ambisonics; synthesizes the encoding results into each frame of the image to obtain a new 3D video;

[0126] S805. The server sends the new 3D video to the terminal;

[0127] S806. The terminal stores the new 3D video.

[0128] Among them, the terminal can be the management terminal that controls the server in the computer room.

[0129] Of course, since the data volume of the 3D video is generally large, the 3D video can also be directly stored in the hard disk, and then the hard disk can be plugged into the server so that the server can directly read the 3D video from the hard disk to perform sound effect synthesis on the 3D video. Correspondingly, the new 3D video can also be directly stored from the server to the hard disk.

[0130] If a new 3D video needs to be played, the terminal storing the new 3D video can be connected to the 3D projection device in the projection exhibition hall, or a hard disk storing the new 3D video can be plugged into the 3D projection device in the projection exhibition hall, or the server storing the new 3D video can be directly connected to the 3D projection device in the projection exhibition hall. Of course, the new 3D video can also be stored locally in the 3D projection device so that the 3D projection device can play the new 3D video. Among them, the 3D projection device in the projection exhibition hall includes: a speaker array, an image projection and display device, headphones, etc. The sound effects in the new 3D video can be played either by the speaker array or by the headphones.

[0131] It can be seen that for the video frames in the 3D video that need to synthesize sound effects in this embodiment, corresponding sound effect signals at the spatial positions can be generated according to the spatial positions where the sound images are located, and the sound effect signals and the video frames can be synthesized, so that the sound effect signals in the finally obtained synthesized video frames can have a sense of spatial orientation and be synchronized with the video frames in terms of time. Therefore, the sound in the new 3D video obtained based on each synthesized video frame and the original video frames in the 3D video that do not need to synthesize sound effects has a sense of spatial orientation and the sound and picture are synchronized.

[0132] Next, an electronic device provided in an embodiment of the present application will be introduced. The relevant implementation steps of the electronic device described below can be mutually referred to those of the above embodiment.

[0133] Furthermore, an embodiment of the present application also provides an electronic device. Among them, the above-mentioned electronic device can be either the Figure 9 server 50 shown as in Figure 10 or the terminal 60 shown as in Figure 9 and Figure 10 are both structural diagrams of electronic devices shown according to an exemplary embodiment. The content in the figure should not be considered as any limitation to the scope of use of the present application.

[0134] Figure 9 is a schematic structural diagram of a server provided in an embodiment of the present application. The server 50 may specifically include: at least one processor 51, at least one memory 52, a power supply 53, a communication interface 54, an input / output interface 55, and a communication bus 56. Among them, the memory 52 is used to store a computer program, and the computer program is loaded and executed by the processor 51 to implement the relevant steps in the multimedia data synthesis disclosed in any of the foregoing embodiments.

[0135] In this embodiment, the power supply 53 is used to provide operating voltages for various hardware devices on the server 50; the communication interface 54 can create a data transmission channel between the server 50 and external devices, and the communication protocol it follows can be any communication protocol applicable to the technical solution of this application, and specific limitations are not imposed here; the input / output interface 55 is used to obtain external input data or output data to the outside, and the specific interface type can be selected according to specific application requirements, and no specific limitations are imposed here.

[0136] In addition, the memory 52, as a carrier for resource storage, can be a read-only memory, a random access memory, a magnetic disk, an optical disc, etc. The resources stored thereon include an operating system 521, a computer program 522, data 523, etc., and the storage method can be temporary storage or permanent storage.

[0137] Among them, the operating system 521 is used to manage and control various hardware devices and the computer program 522 on the server 50 to enable the processor 51 to perform operations and processing on the data 523 in the memory 52, and it can be Windows Server, Netware, Unix, Linux, etc. In addition to the computer program that can be used to complete the multimedia data synthesis method disclosed in any of the foregoing embodiments, the computer program 522 can further include computer programs that can be used to complete other specific tasks. In addition to data such as update information of application programs, the data 523 can also include data such as developer information of application programs.

[0138] Figure 10 It is a schematic structural diagram of a terminal provided by an embodiment of this application. The terminal 60 can specifically include, but is not limited to, a smart phone, a tablet computer, a notebook computer, a desktop computer, etc.

[0139] Generally, the terminal 60 in this embodiment includes: a processor 61 and a memory 62.

[0140] Among them, the processor 61 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 61 may be implemented in at least one of the following hardware forms: DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). The processor 61 may also include a main processor and a coprocessor. The main processor is a processor used to process data in the wake state, also known as the CPU (Central Processing Unit); the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, the processor 61 may be integrated with a GPU (Graphics Processing Unit), and the GPU is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 61 may further include an AI (Artificial Intelligence) processor, and the AI processor is used to process computational operations related to machine learning.

[0141] The memory 62 may include one or more computer-readable storage media, and the computer-readable storage media may be non-transitory. The memory 62 may further include high-speed random access memory and non-volatile memory, such as one or more disk storage devices and flash storage devices. In this embodiment, the memory 62 is at least used to store the following computer program 621. After the computer program is loaded and executed by the processor 61, it can implement the relevant steps in the multimedia data synthesis method executed by the terminal side disclosed in any of the foregoing embodiments. In addition, the resources stored in the memory 62 may further include an operating system 622 and data 623, etc., and the storage method may be temporary storage or permanent storage. Among them, the operating system 622 may include Windows, Unix, Linux, etc. The data 623 may include, but is not limited to, update information of application programs.

[0142] In some embodiments, the terminal 60 may further include a display screen 63, an input / output interface 64, a communication interface 65, a sensor 66, a power supply 67, and a communication bus 68.

[0143] Those skilled in the art can understand that Figure 10 the structure shown in

[0144] The following introduces a storage medium provided by the embodiments of the present application. The relevant implementation steps of the storage medium described below can be referred to each other with the above embodiments.

[0145] Furthermore, the embodiments of the present application also disclose a storage medium. Computer-executable instructions are stored in the storage medium. When the computer-executable instructions are loaded and executed by a processor, the multimedia data synthesis method disclosed in any of the foregoing embodiments is implemented. For the specific steps of this method, reference can be made to the corresponding content disclosed in the foregoing embodiments, and details will not be repeated here.

[0146] It should be noted that the above are only the preferred embodiments of the present application, and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

[0147] In this specification, the embodiments are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The same or similar parts among the embodiments can be referred to each other. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple. For the relevant parts, reference can be made to the description in the method part.

[0148] In this article, specific examples are used to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application. At the same time, for those of ordinary skill in the art, according to the idea of the present application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present application.

Claims

1. A method for synthesizing multimedia data, characterized in that, it includes: Obtain a 3D video; Determine the target video frames in the 3D video that need to synthesize sound effects, and determine the spatial position of the target sound image in the target video frames, and generate a sound effect signal of the target sound image at the spatial position; wherein, select the start and end video frames where the target sound image is located and select some intermediate video frames at intervals as the target video frames; Synthesize the sound effect signal and the target video frames to obtain synthesized video frames; Obtain a new 3D video based on the synthesized video frames and the original video frames in the 3D video that do not need to synthesize sound effects; Among them, the determining the spatial position of the target sound image in the target video frames includes: Taking the object that perceives the target sound image in the target video frame as a reference object, determine the azimuth angle and elevation angle of the target sound image; wherein, taking the coordinate position of the object that perceives the target sound image in the target video frame as the origin, determine the azimuth angle and the elevation angle; And, after generating the sound effect signal of the target sound image at the spatial position according to the spatial position, play the sound effect signal in the new 3D video by using a speaker array or headphones.

2. The method according to claim 1, characterized in that, the generating the sound effect signal of the target sound image at the spatial position includes: Obtain the target audio corresponding to the target sound image, and encode the target audio based on the spatial position to obtain the sound effect signal.

3. The method according to claim 2, characterized in that, the encoding the target audio based on the spatial position to obtain the sound effect signal includes: Determine each encoding channel for encoding the target audio; Determine the signals of the target audio in each encoding channel based on the spatial position; Summarize the signals of each encoding channel to obtain the sound effect signal.

4. The method according to claim 3, characterized in that, it further includes: If the sound effect signal in the synthesized video frames is played back through a spatially distributed speaker array, then decode the sound effect signal based on the speaker array, and play the decoded signal by using the speaker array.

5. The method according to claim 4, characterized in that, the decoding the sound effect signal based on the speaker array includes: Construct a signal matrix based on the number of speakers in the speaker array and the number of encoding channels; Take the pseudo-inverse matrix of the signal matrix as the decoding matrix; Decode the signals of each encoding channel based on the decoding matrix.

6. The method according to claim 4, characterized in that, The number of speakers in the speaker array is not less than the number of encoding channels and satisfies H = (N + 1). 2 ; H is the number of encoding channels, and N is the encoding order.

7. The method according to claim 5, characterized in that, the decoding the signals of each encoding channel based on the decoding matrix includes: Decode the signals of each encoding channel according to the target formula; the target formula is: D = A × [A 1 , A 2 , …, A H T , where D is the decoding result, A is the decoding matrix, and A 1 , A 2 , …, A H represent the signals of H encoding channels, and H is the number of the encoding channels.​ 8. The method according to claim 4, characterized in that, it further includes: If the sound effect signal in the synthesized video frames is played back through headphones, then decode the sound effect signal based on a spatially distributed speaker array, and encode the decoded signal into a left channel signal and a right channel signal, and play the left channel signal and the right channel signal by using the headphones.

9. A method for synthesizing multimedia data, characterized in that, Including: Obtain a three-dimensional image; Determine a target object in the three-dimensional image for which a sound effect needs to be synthesized, and determine the spatial position of the target object in the three-dimensional image; Generate a sound effect signal of the target object at the spatial position based on the spatial position; Synthesize the sound effect signal with the three-dimensional image to obtain a three-dimensional synthesized image; Wherein, the determining the spatial position of the target object in the three-dimensional image includes: Taking the object that perceives the target object in the three-dimensional image as a reference object, determine the azimuth angle and elevation angle of the target object; wherein, taking the coordinate position of the object that perceives the target object in the three-dimensional image as the origin, determine the azimuth angle and the elevation angle; And, after generating the sound effect signal of the target object at the spatial position based on the spatial position, play the sound effect signal in the three-dimensional synthesized image by using a speaker array or headphones.

10. An electronic device Characterized in that The electronic device includes a processor and a memory; wherein, the memory is used to store a computer program, and the computer program is loaded and executed by the processor to implement the method according to any one of claims 1 to 9.

11. A storage medium Characterized in that The storage medium stores computer-executable instructions, and when the computer-executable instructions are loaded and executed by a processor, the method according to any one of claims 1 to 9 is implemented.

Citation Information

Patent Citations

  • Audio processing method and device, readable medium and electronic equipment

    CN113467603A