A method for generating audio
By extracting the motion trajectory and sound source of target image elements from the video, a synchronized audio file is generated, which solves the problem of poor correlation between white noise audio and video images, realizes synchronized audio and video playback, and improves the user experience.
Patent Information
- Application Number
- CN202411486610.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-23
- Publication Date
- 2025-11-14
- Estimated Expiration
- 2044-10-23
AI Technical Summary
In existing technologies, the white noise audio and video image have poor correlation during video file playback, resulting in a poor user experience.
Extract the element image frames of the target image elements from the video to be processed, generate motion trajectories based on the position and time of the image elements, obtain the matching target audio source, and generate a synchronized audio file using spatial rendering to enable the audio and video to play synchronously.
It enables synchronized playback of audio files and video images, improving the user experience.
Smart Images

Figure CN119364129B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of audio processing technology, and more particularly to a method for generating audio. Background Technology
[0002] Nowadays, more and more users like to rest in their cars, and audio and video with white noise elements can help users quickly reach a relaxing state.
[0003] However, when playing video files, the same white noise audio is usually played randomly along with the video file. This white noise audio is often unrelated to the video image or has a poor correlation with it. Summary of the Invention
[0004] This invention provides a method for generating audio, which enables synchronized playback of audio files and target images in a video to be processed, thus achieving audio-visual synchronization in the video to be processed.
[0005] The first aspect of this application provides a method for generating audio, including:
[0006] Extract element image frames containing target image elements from the video to be processed; the target image elements are image elements capable of producing sound effects.
[0007] Based on the position of the target image element in each element image frame and the playback time point corresponding to each element image frame in the video to be processed, the motion trajectory of the target image element in the video to be processed is generated.
[0008] Obtain the target sound source that matches the target image element;
[0009] Spatial rendering is performed on the target sound source using a rendering method that matches the motion trajectory of the target image elements, resulting in a spatially rendered target sound source;
[0010] An audio file corresponding to the video to be processed is generated based on the spatially rendered target sound source; the audio file is used to synchronously play the spatially rendered target sound source that matches the target image element and the position of the target image element in the target element image frame when the video to be processed is played to the target element image frame in the element image frame.
[0011] A second aspect of this application provides an apparatus for generating audio, comprising:
[0012] An extraction unit is used to extract element image frames containing target image elements from the video to be processed; the target image elements are image elements capable of producing sound effects.
[0013] The generation unit is configured to generate the motion trajectory of the target image element in the video to be processed based on the position of the target image element in each element image frame and the playback time point corresponding to each element image frame in the video to be processed.
[0014] The acquisition unit is used to acquire a target sound source that matches the target image element;
[0015] The rendering unit is used to perform spatial rendering on the target sound source using a rendering method that matches the motion trajectory of the target image element, resulting in a spatially rendered target sound source.
[0016] The generation unit is further configured to generate an audio file corresponding to the video to be processed based on the spatially rendered target audio source; the audio file is used to synchronously play the target image element and the spatially rendered target audio source that matches the position of the target image element in the target element image frame when the video to be processed is played to the target element image frame in the element image frame.
[0017] A third aspect of this application provides a computer device, including a processor that, when executing a computer program stored in a memory, implements the method for generating audio provided in the first aspect of this application.
[0018] A fourth aspect of this application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is used to implement the method for generating audio provided in the first aspect of this application.
[0019] The fifth aspect of this application provides a computer program product having computer instructions stored thereon, which, when executed by a processor, are used to implement the method for generating audio provided in the first aspect of this application.
[0020] As can be seen from the above technical solutions, the embodiments of the present invention have the following advantages:
[0021] In this embodiment, element image frames containing target image elements are extracted from the video to be processed; the target image elements are image elements capable of producing sound effects; based on the position of the target image elements in each element image frame and the playback time point corresponding to each element image frame in the video to be processed, the motion trajectory of the target image elements in the video to be processed is generated; a target sound source matching the target image elements is obtained; spatial rendering is performed on the target sound source using a rendering method matching the motion trajectory of the target image elements to obtain a spatially rendered target sound source; an audio file corresponding to the video to be processed is generated based on the spatially rendered target sound source; the audio file is used to synchronously play the target image elements and the spatially rendered target sound source matching the position of the target image elements in the target element image frames when the video to be processed plays to the target element image frame in the element image frame.
[0022] Because the audio file generated in this application embodiment, when the video to be processed plays to the target element image frame containing the target image, synchronously plays the target sound source that matches the target image element and the position of the target image element in the target element image frame, thereby realizing the synchronous playback of the audio file in the video to be processed and the target image in the video to be processed, that is, realizing the audio-visual synchronization in the video to be processed. Attached Figure Description
[0023] Figure 1 This is a schematic diagram of the system architecture for generating audio in the embodiments of this application;
[0024] Figure 2 This is a schematic diagram of one embodiment of the method for generating audio in this application.
[0025] Figure 3 This is a schematic diagram of the target image frame and target image elements in the embodiments of this application;
[0026] Figure 4 This is a schematic diagram of the motion trajectory of the target image element in multiple image frames in an embodiment of this application;
[0027] Figure 5 This is a schematic diagram illustrating an embodiment of spatial rendering of a target sound source using a channel-based rendering method as described in this application.
[0028] Figure 6 This is a schematic diagram illustrating the process of obtaining a 5.1 format multi-channel target audio source in an embodiment of this application;
[0029] Figure 7 This is a schematic diagram illustrating an embodiment of spatial rendering of a target sound source using an object-based rendering method as described in this application.
[0030] Figure 8 This is a schematic diagram of one embodiment of generating extended audio files in this application.
[0031] Figure 9 This is a schematic diagram of another embodiment of generating extended audio files in this application;
[0032] Figure 10 This is a schematic diagram of another embodiment of generating extended audio files in this application;
[0033] Figure 11 This is a schematic diagram of one embodiment of the device for generating audio in this application. Detailed Implementation
[0034] This invention provides a method for generating audio, which automatically generates an audio file that plays synchronously with the target image elements in the element image frames of the video to be processed, based on the target image elements in the video to be processed. This achieves synchronous playback of the audio file in the video to be processed and the target image in the video to be processed, that is, it achieves audio-visual synchronization in the video to be processed.
[0035] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0036] The terms "first," "second," "third," "fourth," etc., used in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0037] This application provides a method for generating audio. The general principle of this method is as follows: When a video needs to be played, the video to be processed is downloaded from the local machine or from the network, and element image frames of target image elements are extracted from the video to be processed; the target image elements are image elements capable of producing sound effects; based on the position of the target image elements in each element image frame and the corresponding playback time point of each element image frame in the video to be processed, the motion trajectory of the target image elements in the video to be processed is generated; a target audio source matching the target image elements is obtained; spatial rendering is performed on the target audio source using a rendering method matching the motion trajectory of the target image elements, resulting in a spatially rendered target audio source. The audio source generates an audio file corresponding to the video to be processed based on the spatially rendered target audio source. The audio file is used to synchronously play the target image element and the spatially rendered target audio source that matches the position of the target image element in the target image frame when the video to be processed is played to the target element image frame in the element image frame. Because the audio file generated by the embodiment of this application can synchronously play the target image element and the spatially rendered target audio source that matches the position of the target image element in the target element image frame when the video to be processed is played to the target element image frame in the element image frame, the synchronous playback of audio and target image in the video is realized.
[0038] To better implement the above-described method for generating audio, this application provides a system for generating audio. Please refer to [link to relevant documentation]. Figure 1 , Figure 1 This is a schematic diagram of the architecture of an audio generation system provided in an embodiment of this application. The audio generation system may include at least one terminal device 101 and a server 102. Different types of applications may be installed on the terminal device 101, such as QQ Music, video players, instant messaging applications, live streaming applications, conferencing applications, etc. The terminal device 101 may be a smartphone, tablet, laptop, desktop computer, smart vehicle, etc. The server 102 may be used to store application data, audio data, video data, or image data generated by different types of applications on the terminal device 101. The server 102 may be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms, etc.
[0039] The method for generating audio described above is executed by either terminal device 101 or server 102. When the method is executed by terminal device 101, the videos that terminal device 101 needs to play in different types of applications can be included in the server. When terminal device 101 needs to generate audio, it can obtain the video to be played from server 102. Then, terminal device 101 extracts multiple image frames from the video and extracts element image frames containing target image elements based on the target image frames. Then, it extracts the motion estimation of the target image elements in the video to be processed, obtains the target audio source that matches the target image elements, and performs spatial rendering on the target audio source using a rendering method that matches the motion trajectory of the target image elements to obtain the spatially rendered target audio source. Finally, it generates an audio file corresponding to the video to be processed based on the spatially rendered target audio source. This audio file is used to synchronously play the spatially rendered target audio source that matches the target image elements and the position of the target image elements in the target element image frames when the video to be processed plays the target element image frames containing image elements, thereby realizing the synchronous playback of audio and target images in the video to be processed.
[0040] For ease of understanding, the method for generating audio in the embodiments of this application is described below. Please refer to [link / reference]. Figure 2 One embodiment of the method for generating audio in this application includes:
[0041] 201. Extract the element image frames containing the target image elements from the video to be processed; the target image elements are image elements that can produce sound effects;
[0042] The execution subject in this application embodiment can be a client, a server, or an APP installed on a client or server. There is no specific limitation on the type of execution subject here.
[0043] For ease of description, this application describes the embodiment with the client as the execution subject. If the client needs to play a video, it can obtain the video to be processed from the local machine or the network, and extract the element image frame containing the target image element from the video to be processed. The target image element is an image element that can produce sound effects.
[0044] Specifically, when the video to be processed contains mountains, water, birds, and trees, the target image elements here can be water, birds, or trees, because these elements can have corresponding water sounds, bird calls, or wind sounds. Of course, the examples here are only for explanation and not limitation on the target image elements. For example, the video to be processed can also contain rain, because rain can also correspond to rain sound effects. For ease of understanding, Figure 3 A schematic diagram of the target image elements in the element image frame is given.
[0045] Specifically, since a video is composed of many image frames, the client can divide the video into multiple image frames to obtain multiple image frames in the video, and then extract the element image frame containing the target image element from the multiple image frames.
[0046] Furthermore, when extracting element image frames containing target image elements, the client can receive element image frames uploaded manually, or the client can automatically extract element image frames containing target image elements from the video based on an algorithm. There are no specific restrictions on the method of extracting element image frames here.
[0047] 202. Based on the position of the target image element in each element image frame and the corresponding playback time point of each element image frame in the video to be processed, generate the motion trajectory of the target image element in the video to be processed;
[0048] In this embodiment of the application, after extracting the element image frames containing the target image elements, the motion trajectory of the target image elements in the video to be processed is further generated based on the position of the target image elements in each element image frame and the playback time point of each element image frame in the video to be processed.
[0049] As an optional embodiment, when extracting the motion trajectory of a target image element across multiple image frames, the position coordinates of the target image element in each of the multiple image frames can be determined separately. Then, according to the playback order of the multiple image frames in the video, the position coordinates of the target image element in each image frame are connected to obtain the motion trajectory of the target image element across multiple image frames. For ease of understanding... Figure 4 A schematic diagram of the motion trajectory of the target image element in multiple element image frames is given.
[0050] 203. Obtain the target sound source that matches the target image elements;
[0051] After obtaining the target image element and its motion trajectory in multiple image frames, the target sound source matching the target image element is further obtained from the material library (which stores a large number of sound sources). The material library can be stored locally on the client or on the server. When the execution subject of this application embodiment is an APP, the material library can also be a small plugin embedded in the APP. Here, there are no specific restrictions on the storage location and presentation of the material library.
[0052] Furthermore, the process of obtaining the target sound source that matches the target image element from the material library will be described in the following embodiments, without specific limitations.
[0053] 204. Spatial rendering is performed on the target sound source using a rendering method that matches the motion trajectory of the target image elements, resulting in a spatially rendered target sound source.
[0054] After obtaining the motion trajectory of the target image element in multiple element image frames, and the target sound source that matches the target image element, spatial rendering is performed on the target sound source based on the motion trajectory of the target image element and using a rendering method that matches the motion trajectory of the target image element, thereby obtaining the spatially rendered target sound source.
[0055] Specifically, the rendering methods in this application embodiment include channel-based rendering or object-based rendering. The process of spatially rendering the target sound source by adopting a rendering method that matches the motion trajectory of the target image elements to obtain the spatially rendered target sound source will be described in detail in the following embodiments, and will not be repeated here.
[0056] 205. Generate an audio file corresponding to the video to be processed based on the spatially rendered target audio source; the audio file is used to synchronously play the spatially rendered target audio source that matches the target image element and the position of the target image element in the target image frame when the video to be processed is played to the target element image frame in the element image frame.
[0057] If the spatially rendered target audio source is obtained in step 204, then an audio file corresponding to the video to be processed is generated based on the spatially rendered target audio source, so as to synchronously play the spatially rendered target audio source that matches the target image element and the position of the target image element in the target image frame when the video to be processed is played.
[0058] Specifically, when generating the audio file corresponding to the video to be processed based on the spatially rendered target sound source, the spatially rendered target sound source can be aligned with, superimposed on, or adaptively balanced with the target image element frame where the target image element is located, in order to generate the audio file corresponding to the video to be processed.
[0059] In this embodiment, element image frames containing target image elements are extracted from the video to be processed; the target image elements are image elements capable of producing sound effects; based on the position of the target image elements in each element image frame and the playback time point corresponding to each element image frame in the video to be processed, the motion trajectory of the target image elements in the video to be processed is generated; a target sound source matching the target image elements is obtained; spatial rendering is performed on the target sound source using a rendering method matching the motion trajectory of the target image elements to obtain a spatially rendered target sound source; an audio file corresponding to the video to be processed is generated based on the spatially rendered target sound source; the audio file is used to synchronously play the spatially rendered target sound source that matches the target image elements and the position of the target image elements in the target element image frames when the video to be processed plays to the target element image frame in the element image frame.
[0060] Because the audio file generated in this application embodiment, when the video to be processed plays to the target element image frame containing the target image, synchronously plays the target sound source that matches the target image element and the position of the target image element in the target element image frame, thereby realizing the synchronous playback of the audio file in the video to be processed and the target image in the video to be processed, that is, realizing the audio-visual synchronization in the video to be processed.
[0061] based on Figure 2 Step 204 in the embodiment is described in detail below as follows: How to perform spatial rendering on the target sound source based on the motion trajectory of the target image:
[0062] Specifically, the motion trajectory of the target image element in the video to be processed includes static motion trajectory and dynamic motion trajectory. Static motion trajectory means that the position of the target image element in the video to be processed does not change over time, while dynamic motion trajectory means that the position of the target image element in the video to be processed changes over time.
[0063] Furthermore, in this embodiment, the spatial rendering of the target audio source includes channel-based rendering and object-based rendering. The channel-based rendering method renders the target audio source on channels based on a preset mixing framework, while the object-based rendering method renders the target audio source based on the position of the target image element in the video and a preset multi-channel audio source format.
[0064] As an optional embodiment, if the motion trajectory of the target image element includes a static motion trajectory, then channel-based rendering is used to perform spatial rendering on the target sound source; if the motion trajectory of the target image element includes the dynamic motion trajectory, then object-based rendering is used to perform spatial rendering on the target sound source. The process of performing spatial rendering on the target sound source is described below from two aspects; please refer to the respective descriptions. Figure 5 and Figure 7 :
[0065] 1. For static motion trajectories, a channel-based rendering method is used to spatially render the target sound source. Please refer to [link / reference]. Figure 5 :
[0066] 501. Use an upmixing algorithm to separate the relevant and irrelevant signals from the left and right channels of the target sound source;
[0067] The acquired video is usually in stereo, which includes left and right channels. When the motion trajectory of the target image element is static, an upmixing algorithm is used to separate the relevant and irrelevant signals from the left and right channels of the target audio source.
[0068] Specifically, the upmixing algorithm is a digital signal processing-based algorithm that mainly separates the relevant and irrelevant signals from stereo, and then places the relevant and irrelevant signals in different channels according to a specific mixing framework, thereby generating audio files of different formats, such as multi-channel audio files like 3.1, 5.1, 5.1.4, or 7.1.4.
[0069] As an optional embodiment, when using the upmixing algorithm, this application combines the Mid / Side decomposition algorithm or the NLMS adaptive filter to find the relevant and irrelevant signals in the left and right channels. When the Mid / Side decomposition algorithm is used, the collected relevant signal refers to the center signal of the stereo image (i.e., the Mid signal), while the irrelevant signal refers to the side signals on both sides of the stereo image (i.e., the Side signal).
[0070] Specifically, in Mid / Side stereo, the Mid channel is located at the center of the stereo image, and its position moves closer to the center as the Mid channel signal increases. The Side channels are located on both sides of the stereo image, and their positions expand to the sides as the Side channel signal increases, making the sound wider.
[0071] When using an NLMS adaptive filter, the input stereo signal is... Filtering is performed to extract relevant and error signals from the stereo signal.
[0072] For ease of understanding, the following description uses the generation of a 5.1 format audio file as an example to illustrate the Mid / Side decomposition algorithm and the NLMS adaptive filter:
[0073] A 5.1 audio file includes Left, Center, Right, LFE (Left Exterior), Left Surround, and Right Surround channels. The Mid / Side decomposition formula is as follows:
[0074] ; ;
[0075] in, This refers to the left channel signal in the input stereo signal. The input stereo signal is the right channel signal, while the Mid signal is used to generate the relevant signal in the embodiments of this application, and the Side signal is used to generate the unrelated signal in the embodiments of this application.
[0076] Furthermore, the specific formula for the NLMS adaptive filter is as follows:
[0077] ; ;
[0078] ;
[0079] in, This refers to the input stereo signal. These are weighting coefficients. It is a relevant signal. It is a fixed value. It is an error signal. These are the weight coefficients after iteration.
[0080] NLMS is a variant of LMS filter used to filter the input stereo signal. Filtering is performed to find the correlation signal and error signal from the stereo signal. Among them, the NLMS filter improves the problem of LMS's sensitivity to input dynamics compared to the LMS filter. NLMS finds the correlation signal and error signal by continuously optimizing the minimum mean square error between the desired signal and the input signal.
[0081] 502. Place relevant and irrelevant signals in different channels according to the preset mixing framework to obtain the target sound source after spatial rendering.
[0082] After obtaining the relevant signals in step 501, the relevant and irrelevant signals are processed according to the preset mixing framework to obtain the target sound source after spatial rendering.
[0083] The upmixing algorithm also includes a mixing framework, which comprises the signal, mapping rules for each channel, and post-processing of the signal. For ease of understanding, the following uses a 5.1 format audio file as an example to illustrate the processing steps of the upmixing algorithm using the mixing framework:
[0084] Specifically, a 5.1 audio file includes a center channel signal, a left front channel signal, a right front channel signal, a left surround channel signal, a right surround channel signal, and a subwoofer. When the input is a stereo left and right channel signal, the output signal of the left front channel in the 5.1 audio file is the original stereo left channel signal minus 3dB, and the output signal of the right front channel is the original stereo right channel signal minus 3dB. The output of the center channel signal in the 5.1 audio file is the correlated signal in step 501 minus 2dB, and the LFE is the signal that has passed through a low-pass filter with a cutoff frequency of 120Hz minus 4dB. The output signal of the left surround channel in the 5.1 audio file is the uncorrelated signal minus 2dB, and the output signal of the right surround channel is also the uncorrelated signal minus 2dB.
[0085] Furthermore, in order to enhance the surround sound effect of the left and right surround channel signals, after obtaining the output signals of the left and right surround channels in the 5.1 audio file, the output signals of the left and right surround channels are further subjected to decorrelation processing. The decorrelation processing here includes, but is not limited to, passing the output signals of the left and right surround channels through an all-pass filter or delaying the sampling points in the time domain.
[0086] For ease of understanding, Figure 6 A schematic diagram of the process of obtaining the 5.1 format multi-channel target audio source is given above.
[0087] II. The motion trajectory is dynamic; spatial rendering of the target sound source is performed using an object-based rendering method. Please refer to [link / reference]. Figure 7 :
[0088] 701. Input the target sound source, the motion trajectory of the target image element, and the preset multi-channel music format into the spatial renderer to obtain the spatially rendered target sound source output by the spatial renderer.
[0089] If the motion trajectory of the target image element in multiple frames is dynamic, then the target sound source, the motion trajectory of the target image element, and the preset multi-channel music format are input into the spatial renderer to obtain the spatially rendered target sound source output by the spatial renderer.
[0090] Specifically, the target audio source is first analyzed to obtain the audio channel information of the target audio source, and multiple positional information of the target audio source in the entire video is determined based on the motion trajectory of the target image elements.
[0091] Then, based on multiple location information of the target sound source, the weight of the target sound source in each channel of the multi-channel music system is calculated. The calculation of this weight can be achieved through various acoustic models and algorithms, such as the VBAP (VectorBase Amplitude Panning) algorithm or the Ambisonics algorithm.
[0092] Further defining the multi-channel music format, taking the 7.1.4-channel music format as an example, this multi-channel music format includes 7 horizontal channels, namely the front left channel, front center channel, front right channel, left channel, right channel, rear left channel and rear right channel, 1 low frequency enhancement (LEF) channel and 4 height channels, namely the front left high channel, front right high channel, rear left high channel and rear right high channel.
[0093] Finally, the target audio source of each location information is multiplied by its weight in each channel, and the result of the multiplication is then superimposed on the corresponding channel. In this way, each channel includes a mixed signal of target audio from different location information. Finally, the mixed signals of each channel are mixed to obtain the multi-channel target audio source output by the spatial renderer.
[0094] To make it easier to understand, the following example is provided:
[0095] Assuming the motion trajectory of a target image element in the video contains three positional information points, the motion trajectory of a target audio source in the video also contains three positional information points, let's say (x1, y1), (x2, y2), and (x3, y3). Then, calculate the weight of the target audio source at (x1, y1) in each of the 12 channels of the 7.1.4 music format. Next, multiply the target audio source at (x1, y1) by its respective weight, and sum the results to the 12 channels. Finally, calculate the weight of the target audio source at (x2, y2) in each of the 12 channels of the 7.1.4 music format. The weights of the target sound source at (x2, y2) are calculated, and then the target sound source at (x3, y3) is multiplied by its respective weight. The result of the multiplication is then added to the 12 channels. Finally, the weight of the target sound source at (x3, y3) in the 12 channels of the 7.1.4 music format is calculated, and then the target sound source at (x3, y3) is multiplied by its respective weight. The result of the multiplication is then added to the 12 channels. In this way, each channel contains a mixed signal of the target sound source from different position information. Finally, the mixed signals of each channel are mixed to obtain the spatially rendered target sound source output in the embodiment of this application.
[0096] As an optional embodiment, before performing step 701, the target sound source may be preprocessed according to the motion trajectory of the target image element, wherein the preprocessing includes at least one of reverberation processing, delay processing, compression processing and volume automation processing.
[0097] Furthermore, if the motion trajectory of the target image element in multiple image frames is along the first direction, then automatic volume processing and compression processing are performed on the target audio source. The first direction is the horizontal direction along the image frame, that is, the direction parallel to the video playback scroll bar. If the motion trajectory of the target image element in multiple image frames is along the second direction, then reverb processing and delay processing are performed on the target audio source. The second direction is perpendicular to the first direction, that is, the second direction is the vertical direction along the image frame.
[0098] Assuming the target sound source in the first direction is a bird call, it's easy to understand that the bird call will become louder and louder as the bird moves in the first direction. Therefore, performing automatic volume processing and compression on the target sound source in the first direction can enhance the scene effect of the target sound source. Assuming the target sound source in the second direction is the sound of flowing water, then performing reverberation and delay processing on the sound of flowing water is because the sound of water heard at different positions along the longitudinal direction is different. Therefore, performing reverberation and delay processing on the target sound source in the second direction can enhance the elemental influence effect of the target image element (i.e., the flowing water).
[0099] based on Figure 5 or Figure 7 In the aforementioned embodiment, if the target image frame includes at least two target image elements, after obtaining at least two spatially rendered target audio sources, each spatially rendered target audio source corresponding to the video to be processed is aligned with the timestamp of its corresponding target image element according to the timestamp in the motion trajectory of each target image element, and the aligned multiple spatially rendered target audio sources are superimposed to obtain a superimposed audio source.
[0100] Furthermore, in order to match the volume of each target audio source in at least two target audio sources, adaptive volume balancing processing can be performed on the superimposed audio sources to obtain the audio file corresponding to the video to be processed.
[0101] To make it easier to understand, the following example is provided:
[0102] Assuming the target image frame contains water and birds as target image elements, after obtaining the spatially rendered water sound and bird call, the spatially rendered water sound and the water image element in the video are aligned on the timestamp based on the timestamp of the water appearing in the video to be processed. Then, the spatially rendered bird call and the bird image element in the video are aligned on the timestamp based on the timestamp of the bird appearing in the video. Finally, the aligned bird call and the aligned water sound are superimposed to obtain the superimposed sound source.
[0103] Furthermore, in order to match the volume of bird calls and flowing water sounds and prevent the volume disharmony that occurs when one sound is too loud and the other is too soft, adaptive volume balancing can be performed on the superimposed sound sources to obtain the audio file corresponding to the video to be processed.
[0104] based on Figure 5 or Figure 7 In the aforementioned embodiment, after obtaining the audio file corresponding to the video to be processed, if the client detects that the video is in loop playback mode, in order to reduce the fatigue caused by the audio file corresponding to the video to be processed in the image elements of the video during loop playback, an extended audio file can be generated based on the audio file of the video to be processed. The audio file is used to play when the video to be processed is played for the first time, while the extended audio file is used to play when the video to be processed is looped.
[0105] As an optional embodiment, when generating an extended audio file based on the audio file of the video to be processed, it may be to obtain audio adjustment elements for the video to be processed, the audio adjustment elements including at least one of the target light music, the associative element sound source of the target image element, and the sampled segment of the target sound source;
[0106] The audio adjustment elements are overlaid on the audio file to obtain an extended audio file. The original audio file is used for the first playback of the video to be processed, while the extended audio file is used for the loop playback of the video to be processed.
[0107] Specifically, when audio adjustment elements include different elements (such as target light music, associative element sound sources of target image elements, and sampled segments of target sound sources), the process of overlaying specific audio adjustment elements with audio files to obtain expanded audio files will be described in detail from the following three aspects:
[0108] 1. If the audio adjustment elements include the target light music
[0109] Please see Figure 8 An example of creating an expanded audio file by overlaying audio adjustment elements onto the audio file includes:
[0110] 801. Select target light music from a preset light music library based on a similarity strategy;
[0111] To avoid fatigue caused by repeated playback of the adaptive target sound source during the playback of multiple image frames, embodiments of this application can also select target light music from a preset light music library based on a similarity strategy and overlay the target light music after the adaptive target sound source.
[0112] As an optional embodiment, the process of selecting target light music from a preset light music library based on a similarity strategy is as follows:
[0113] S1. Extract the scene semantic vector of the video to be processed;
[0114] Specifically, when extracting scene semantic vectors from the video to be processed, word vector models are used as a way to convert natural language into word vectors so that the word vectors can be recognized by computers. Specifically, the word vector models in the embodiments of this application can be common models such as Word2Vec, GloVe, or Bert.
[0115] As an optional embodiment, when extracting the scene semantic vector of the video to be processed, all target image elements in multiple image frames can be extracted separately, and then the identifier of each target image element can be obtained. This identifier can be a natural language identifier such as the name or identification code of the target image element. Then, a word vector model is used to convert the identifier of each target image element into a semantic vector, thereby obtaining multiple semantic vectors for each target image element, such as in... Figure 3 The target image elements in the video include birds, trees, and flowing water. The natural language identifiers of these target image elements are converted into their respective semantic vectors. Then, the average value is calculated based on the semantic vectors of all target image elements. This average value is the scene semantic vector of the video.
[0116] S2. Extract the semantic vector of each light music scene tag in the light music library to obtain scene semantic vectors for multiple light music tracks.
[0117] Similar to step S1, since each piece of light music in the light music library has a scene label, and these scene labels are all natural language labels, the semantic vector of the scene of each piece of light music can be extracted to obtain the scene semantic vectors of multiple pieces of light music. As an optional embodiment, the scene label of each piece of light music can be input into a pre-trained word vector model to obtain multiple scene semantic vectors corresponding to multiple pieces of light music.
[0118] The process of converting the scene tags of light music into scene semantic vectors is similar to the process of converting the identifier of each target image element into a semantic vector in step S1. This process of converting natural language into semantic vectors using word vector models has been described in existing technologies and will not be repeated here.
[0119] S3. Calculate the similarity between the scene semantic vector of the video and the scene semantic vector of each piece of light music to obtain multiple similarity values.
[0120] In order to obtain the target light music that best matches each target image element in the video, this application embodiment can use a similarity algorithm to calculate the similarity between the scene semantic vector of the video and the scene semantic vectors of multiple light music pieces, so as to obtain multiple similarity values. The method for calculating similarity includes, but is not limited to, cosine similarity method and distance similarity method (Euclidean distance or Manhattan distance). No specific limitation is made on the algorithm for calculating similarity here.
[0121] As an optional implementation, when calculating the similarity value, the cosine similarity calculation method can be used to calculate the similarity between the scene semantic vector of the video and the scene semantic vectors of multiple pieces of light music, such as by using the following formula:
[0122]
[0123] in, , These represent the components of vectors A and B, respectively.
[0124] S4. The light music corresponding to the similarity value that meets the preset standard among multiple similarity values is determined as the target light music.
[0125] After obtaining multiple similarity values between the scene semantic vector of the video and the scene semantic vectors of multiple pieces of light music in step S3, the light music corresponding to the similarity value that meets the preset standard can be determined as the target light music.
[0126] As an optional embodiment, the light music corresponding to the highest similarity value can be determined as the target light music, or a random selection can be made from multiple light music pieces that meet a preset similarity threshold as the target light music.
[0127] 802. The target light music is processed using a channel-based rendering method to obtain processed multi-channel light music audio;
[0128] After determining the target light music from the light music library in step 801, the target light music is further processed using a channel-based rendering method to obtain the processed multi-channel light music audio. The channel-based rendering process for processing the target light music is similar to... Figure 5 The channel-based rendering method for the target audio source described in the embodiments is similar and will not be repeated here.
[0129] 803. The processed multi-channel light music source is superimposed on the audio source of the audio file to obtain an extended audio file.
[0130] After obtaining the processed multi-channel light music audio, in order to avoid the monotony and fatigue caused by repeated playback, this embodiment of the application can superimpose the processed multi-channel light music audio source with the audio source of the audio file to obtain an extended audio file.
[0131] Specifically, after overlaying the processed multi-channel light music source with the audio source of the audio file, the processed multi-channel light music source is played simultaneously.
[0132] This application embodiment can superimpose the processed multi-channel light music source with the audio source of the audio file to obtain an extended audio file. The extended audio file can enhance the immersiveness and layering of the audio source and reduce the fatigue caused by loop playback.
[0133] As an optional embodiment, in order to improve the connection and integration between multi-channel light music and multi-channel target audio sources, this embodiment of the application may also superimpose the processed multi-channel light music audio source with the audio source of the audio file in a fade-in and fade-out manner to obtain an extended audio file, so as to avoid the abruptness caused by the interruption between multiple music tracks.
[0134] Because the embodiments of this application generate extended audio files by introducing light music, such processed multi-channel light music sources can improve audio quality on the one hand, and light music can also fully relax the mind and body, and enhance the user's mental and physical pleasure on the other hand.
[0135] 2. If the audio adjustment elements include the target light music
[0136] Please see Figure 9 An example of creating an expanded audio file by overlaying audio adjustment elements onto the audio file includes:
[0137] 901. Randomly extract multiple target audio source segments of equal length from the target audio source;
[0138] Different from Figure 8 The method for generating extended audio files described herein allows for the direct random extraction of multiple equal-length target audio source segments from the target audio source, and the generation of extended audio files based on these target audio source segments, thereby improving the convenience of generating extended audio files.
[0139] Specifically, when using a target audio source to generate an extended audio file, multiple target audio source segments of equal length can be randomly extracted from the target audio source. For example, when the target audio source is a 10-second audio source, multiple target audio source segments of equal length can be randomly determined from the 10-second audio source. The target audio source segments can be 2 seconds, 3 seconds, or 4 seconds, etc. There is no specific restriction on the length of the target audio source segments here.
[0140] 902. If the motion trajectory of the target image element corresponding to the target sound source includes a static motion trajectory, then a channel-based rendering method is used to render the multiple equal-length target sound source segments to obtain multiple spatially rendered target sound source segments.
[0141] After obtaining multiple target sound source segments of equal length in step 901, the target sound source segments can be further rendered based on the motion trajectory of the target image elements corresponding to the target sound source. If the motion trajectory of the target image elements corresponding to the target sound source includes a static motion trajectory, then a channel-based rendering method is used to render the multiple target sound source segments of equal length to obtain multiple spatially rendered target sound source segments. The process of rendering the multiple target sound source segments of equal length using a channel-based rendering method is similar to... Figure 5 The descriptions in the embodiments are similar and can be referred to one another; they will not be repeated here.
[0142] 903. If the motion trajectory of the target image element corresponding to the target sound source includes a dynamic motion trajectory, then an object-based rendering method is used to render the multiple equal-length target sound source segments to obtain multiple spatially rendered target sound source segments.
[0143] After obtaining multiple target sound source segments of equal length in step 901, if the motion trajectory of the target image element corresponding to the target sound source includes a dynamic motion trajectory, then an object-based rendering method is used to render the multiple target sound source segments of equal length to obtain multiple spatially rendered target sound source segments. The process of rendering the multiple target sound source segments of equal length using an object-based rendering method is similar to... Figure 7 The descriptions in the embodiments are similar and can be referred to one another; they will not be repeated here.
[0144] 904. Obtain the first timestamp of the start of playback and the second timestamp of the end of playback of the target image element corresponding to the target audio source in the video;
[0145] After obtaining multiple spatially rendered target audio source segments, the first timestamp of the target image element corresponding to the target audio source in the video and the second timestamp of the video when playback begins and ends are further obtained.
[0146] Specifically, when the target image element is a bird, the first timestamp when the bird starts playing in the video and the second timestamp when the bird stops playing are obtained. When obtaining the first and second timestamps, the client can determine the first image frame in which the bird appears and the second image frame in which the bird disappears based on multiple image frames in the video using image segmentation. Then, the first timestamp of the first image frame appearing in the video and the second timestamp of the second image frame appearing in the video are obtained.
[0147] 905. Randomly overlay one or more of the spatially rendered target audio source segments between the first timestamp and the second timestamp.
[0148] Once the first and second timestamps are determined, one or more spatially rendered target audio source segments can be randomly superimposed between the first and second timestamps.
[0149] Because the embodiments of this application generate extended audio directly based on the sampled segments of the target sound source, the process of obtaining the target light music is reduced compared to the method of selecting light music, thereby improving the convenience of generating extended audio files.
[0150] Furthermore, based on Figure 9 In the aforementioned embodiment, when the duration of the target sound source far exceeds the duration of the target image element appearing in the video, such as a bird call sound source of 3 minutes, but the bird flying in the video is only shown for 10 seconds, and only 2 minutes of valid bird calls are present in the 3-minute bird call sound source, in order to improve the effectiveness of obtaining bird calls contained in the target sound source segment, this embodiment of the application may also perform the following steps before step 901:
[0151] If the duration of the target audio source is greater than the time difference between the first and second timestamps, then a valid audio source segment is extracted from the target audio source. Then, when executing step 901, multiple target audio source segments of equal length are randomly extracted from the valid audio source segments.
[0152] Specifically, because the duration of the target audio source is greater than the time difference between the first and second timestamps, that is, when the duration of the target audio source is greater than the duration of the target image element in the video, in order to obtain a valid target audio source segment, a valid audio source segment is further extracted from the target audio source.
[0153] As an optional embodiment, when extracting valid audio source segments from a target audio source, the target audio source may be subjected to audio framing to obtain multiple frames of audio source signals. Then, the energy value of each frame of audio source signal is calculated to obtain multiple energy values of the multiple frames of audio source signals. Then, multiple target energy values that are greater than a preset energy threshold (the preset energy threshold here is used to characterize the appearance of the target image element in the target audio source) are selected from the multiple energy values. Then, the multiple frames of audio source signals corresponding to the multiple target energy values are determined as valid audio source signals. Then, all timestamps of the valid audio source signals are determined from the video, and valid audio source segments are extracted from the target audio source according to all timestamps.
[0154] The goal of audio framing is to divide continuous speech signals into short frames so that each frame can be further analyzed and processed. A common framing method is to use a fixed-length window, which is used to segment the signal by sliding the window over the target audio source. Typically, the framing window length is chosen to be 20-30ms, which can ensure that the statistical characteristics of the speech signal remain relatively stable within a short period of time.
[0155] Specifically, the energy value for each frame can be calculated using the following formula:
[0156]
[0157] In this embodiment of the application, before randomly extracting multiple target audio source segments of equal length from the target audio source, a valid audio source segment is first extracted from the target audio source, thereby ensuring the probability of extracting a valid audio source segment from the target audio source segment and improving the effectiveness of obtaining the target audio source segment.
[0158] 3. If the audio adjustment element includes the associative element sound source of the target image element.
[0159] Please see Figure 10 An example of creating an expanded audio file by overlaying audio adjustment elements onto the audio file includes:
[0160] 1001. Obtain the sound sources of multiple associative elements of the target image element;
[0161] Different from Figure 8 and Figure 9In the embodiments described above, when generating extended audio files, the extended audio files are generated based on the associative element sound sources of the target image elements, thereby avoiding the monotony and fatigue caused by continuous playback of the target sound source. Here, the associative elements are image elements associated with the target image elements. For example, when the target image element is a forest or rain, the associative elements can be thunder, wind, or animal sounds, etc., and the associative element sound source is a sound source containing the associative elements. For example, when the associative element is thunder, the associative element sound source is a sound source containing thunder, and when the associative element is wind, the associative element sound source is a sound source containing wind, etc.
[0162] As an optional embodiment, when obtaining the associative element sound source of the target image element, the semantic vector of the target image element identifier can be calculated using a preset word vector model (such as Word2Vec, GloVe, Bert, etc.). Here, the target image element identifier refers to the natural language identifier of the target image element. The process of calculating the semantic vector of the target image element identifier using a word vector model is similar to that described in the prior art, and will not be repeated here.
[0163] After obtaining the semantic vector of the target image element identifier, the semantic vector of all sound source identifiers in the sound source library is calculated using the same word vector model. Since all sound sources in the sound source library have their own sound source identifiers, the semantic vector of all sound source identifiers in the sound source library can be calculated using the word vector model. Then, based on the similarity value calculation method, the similarity value (such as cosine similarity or distance similarity) between the semantic vector of the target image element identifier and the semantic vector of all sound source identifiers is calculated, thereby obtaining multiple similarity values.
[0164] After obtaining multiple similarity values, multiple target similarities that are greater than the preset similarity threshold can be determined from the multiple similarity values, and the multiple sound sources corresponding to the multiple target similarity values can be further determined as multiple associative element sound sources.
[0165] To make it easier to understand, the following example is provided:
[0166] Assuming there are 100 audio sources in the audio source library, calculate 100 similarity values between the semantic vector of the target audio source identifier and the semantic vectors of the 100 audio source identifiers. If there are 5 similarity values greater than the similarity threshold, then the 5 audio sources corresponding to these 5 similarity values are identified as the associative element audio sources.
[0167] It should be noted that when identifying multiple associative element sound sources, the number of associated sound sources can be changed by adjusting the similarity threshold. For example, if there are 5 associative element sound sources but only 3 are needed, the number of associative elements can be reduced by increasing the similarity threshold. Conversely, if more associative elements (such as 8) are needed, the number of associative elements can be increased by decreasing the similarity threshold.
[0168] 1002. Randomly arrange and combine the multiple associative element sound sources to obtain associative element sound sources combined in the first order;
[0169] To further enhance the richness and diversity of associative elements, multiple associative element sound sources can be randomly arranged and combined to obtain associative element sound sources combined in the first order.
[0170] Because the variety and quantity of associative element sound sources after random combination are greater than those of single associative elements, the content of the sound source is also richer. The random combination method also ensures that the combined associative element sound sources obtained each time in the loop are not exactly the same, thereby further improving the diversity and richness of the target sound source in the loop process.
[0171] 1003. Obtain the first timestamp of the start of playback and the second timestamp of the end of playback for the target image element in the video;
[0172] Similar to step 903, in order to achieve the effect of audio-visual synchronization, that is, to achieve the synchronous playback of the target audio source when the target image element is played in the video image, this application embodiment first obtains the first time stamp and the second time stamp when the target image element starts playing in the video, and further provides a technical basis for the associative element audio source superimposed and combined outside the first time stamp and the second time stamp.
[0173] 1004. Obtain any target timestamp before the first timestamp or after the second timestamp, and combine the associative element sound sources in the first order at any target timestamp of the audio file.
[0174] After obtaining the first and second timestamps of the target image element in the video, since the embodiments of this application generate extended audio files based on the associative element sound sources of the target image element, during the loop playback, the associative element sound sources are superimposed and combined in the first order at any target timestamp before the first timestamp and after the second timestamp.
[0175] Specifically, since the associative element sound source is an element sound source associated with the target image element, the embodiments of this application can superimpose and combine the associative element sound sources in a first order at any target time point other than the time period between the first and second timestamps, thereby improving the diversity and richness of the sound source content during the loop playback of the video, reducing the monotony of the sound source, and reducing the fatigue caused by monotony.
[0176] In this embodiment, a randomly combined associative element sound source is superimposed at any target time point other than the time period between the first and second timestamps. Compared with the target sound source that is played in a loop, this randomly combined associative element sound source improves the diversity and richness of the sound source during image playback, and further reduces the fatigue caused by playing the target sound source in a loop during playback.
[0177] based on Figure 10 In the embodiment described above, in order to improve the layering of the sound of the randomly combined associative element sound sources during step 1002, this embodiment can also perform spatial rendering on the multiple associative element sound sources based on the object rendering method after obtaining multiple associative element sound sources. Specifically, when performing spatial rendering, the motion trajectory of each associative element sound source, each randomly initialized target image element, and the preset music format (such as 5.1.4, 7.1.4, etc.) are input into the spatial renderer to obtain multiple spatially rendered associative element sound sources, and then the multiple spatially rendered associative element sound sources are randomly arranged and combined.
[0178] Specifically, regarding the process of performing spatial rendering of the associative element sound source based on object rendering, and... Figure 7 The process of performing spatial rendering of the target sound source based on object rendering in the embodiment is similar and will not be described again here.
[0179] Furthermore, based on Figure 10 In the aforementioned embodiment, when performing step 1002, if the combined associative image element sound source obtained after random arrangement and combination of multiple associative image element sound sources is longer than the playback duration of the target image element, then multiple associative element sound source segments of equal length can be randomly extracted from the multiple associative element sound sources, and then the multiple associative element sound source segments of equal length can be randomly combined, so that the playback duration of the randomly combined associative element sound source is adapted to the playback duration of the target image element in the video.
[0180] It is understood that, in various embodiments of the present invention, the order of the steps does not imply the order of execution. The execution order of each step should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0181] The method for generating audio in the embodiments of this application has been described in detail above. The apparatus for generating audio in this application will now be described. Please refer to [link to relevant documentation]. Figure 11 One embodiment of the audio generation apparatus in this application includes:
[0182] Extraction unit 1101 is used to extract element image frames containing target image elements from the video to be processed; the target image elements are image elements that can produce sound effects;
[0183] The generation unit 1102 is used to generate the motion trajectory of the target image element in the video to be processed based on the position of the target image element in each element image frame and the playback time point corresponding to each element image frame in the video to be processed.
[0184] Acquisition unit 1103 is used to acquire a target sound source that matches the target image element;
[0185] The rendering unit 1104 is used to perform spatial rendering on the target sound source using a rendering method that matches the motion trajectory of the target image element, to obtain a spatially rendered target sound source.
[0186] The generation unit 1102 is further configured to generate an audio file corresponding to the video to be processed based on the spatially rendered target audio source; the audio file is used to synchronously play the target image element and the spatially rendered target audio source that matches the position of the target image element in the target image frame when the video to be processed is played to the target element image frame in the element image frame.
[0187] Preferably, the motion trajectory includes static motion trajectory and dynamic motion trajectory, and the rendering method includes channel-based rendering method and object-based rendering method;
[0188] Rendering unit 1104 is specifically used for:
[0189] If the motion trajectory of the target image element includes the static motion trajectory, then the channel-based rendering method is used to perform spatial rendering on the target sound source;
[0190] If the motion trajectory of the target image element includes the dynamic motion trajectory, then the object-based rendering method is used to perform spatial rendering on the target sound source.
[0191] Preferably, the rendering unit 1104 is specifically used for:
[0192] An upmixing algorithm is used to separate the relevant and irrelevant signals from the left and right channels of the target sound source;
[0193] Relevant and irrelevant signals are placed in different channels according to a preset mixing framework to obtain the target sound source after spatial rendering.
[0194] Preferably, the rendering unit 1104 is specifically used for:
[0195] The target sound source, the motion trajectory of the target image elements, and the preset multi-channel music format are input into the spatial renderer to obtain the spatially rendered target sound source output by the spatial renderer.
[0196] Preferably, the device further includes:
[0197] The preprocessing unit 1105 is used to perform automatic volume processing and compression processing on the target audio source if the motion trajectory of the target image element in the video to be processed is along a first direction, so as to obtain the processed target audio source.
[0198] The preprocessing unit 1105 is further configured to perform reverberation processing and delay processing on the target sound source if the motion trajectory of the target image element in the plurality of image frames is a motion along a second direction, thereby obtaining a processed target sound source; the second direction is perpendicular to the first direction;
[0199] Rendering unit 1104 is specifically used for:
[0200] The processed target sound source, the motion trajectory of the target image element, and the preset multi-channel music format are input into the spatial renderer to obtain the spatially rendered target sound source output by the spatial renderer.
[0201] Preferably, the number of target image elements is at least two, the number of rendered target sound sources is at least two, and the generation unit 1102 is specifically used for:
[0202] Based on the timestamps in the motion trajectories of each target image element, each rendered target audio source corresponding to the video to be processed is aligned with the timestamps of its respective target image elements.
[0203] The aligned and rendered target sound sources are superimposed to obtain the superimposed sound source;
[0204] Adaptive volume balancing is performed on the superimposed audio source to obtain the audio file corresponding to the video to be processed.
[0205] Preferably, the device further includes:
[0206] Acquisition unit 1103 is used for:
[0207] Obtain audio adjustment elements for the video to be processed, wherein the audio adjustment elements include at least one of the following: target light music, associative element sound source of the target image element, and sampled segment of the target sound source;
[0208] Stacking unit 1106, used for:
[0209] The audio adjustment element is superimposed on the audio file to obtain an extended audio file, wherein the audio file is used for playback when the video to be processed is played for the first time, and the extended audio file is used for playback when the video to be processed is played in a loop.
[0210] Preferably, if the audio adjustment element includes the target light music, the acquisition unit 1106 is specifically used for:
[0211] The target light music is selected from the preset light music library based on the similarity strategy.
[0212] The superposition unit 1106 is specifically used for:
[0213] The target light music is processed using a channel-based rendering method to obtain a processed multi-channel light music sound source;
[0214] The processed multi-channel light music source is superimposed on the source of the audio file to obtain an extended audio file.
[0215] Preferably, the acquisition unit 1103 is specifically used for:
[0216] Extract the scene semantic vector of the video to be processed;
[0217] Extract the semantic vector of each light music scene tag in the light music library to obtain scene semantic vectors for multiple light music tracks;
[0218] The similarity between the scene semantic vector of the video to be processed and the scene semantic vector of each piece of light music is calculated to obtain multiple similarity values;
[0219] The light music corresponding to the similarity value that meets the preset standard among the multiple similarity values is determined as the target light music.
[0220] Preferably, the acquisition unit 1103 is specifically used for:
[0221] Extract multiple element image frames from the video to be processed;
[0222] Extract all target image elements from the multiple element image frames;
[0223] Obtain the identifiers of each target image element;
[0224] Calculate the semantic vector corresponding to the identifier of each target image element to obtain multiple semantic vectors for each target image element;
[0225] The average value of the multiple semantic vectors is calculated to obtain the scene semantic vector of the video to be processed.
[0226] Preferably, the stacking unit 1106 is also used for:
[0227] The processed multi-channel light music audio is overlaid with the audio file using a fade-in / fade-out method to obtain an extended audio file.
[0228] Preferably, if the audio adjustment element includes a sample segment of the target sound source, the acquisition unit 1103 is further configured to:
[0229] Randomly extract multiple target audio source segments of equal length from the target audio source;
[0230] The superposition unit 1106 is also used for:
[0231] If the motion trajectory of the target image element corresponding to the target sound source includes a static motion trajectory, then a channel-based rendering method is used to render the multiple equal-length target sound source segments to obtain multiple spatially rendered target sound source segments.
[0232] If the motion trajectory of the target image element corresponding to the target sound source includes a dynamic motion trajectory, then an object-based rendering method is used to render the multiple equal-length target sound source segments to obtain multiple spatially rendered target sound source segments.
[0233] Obtain the first timestamp of the start of playback and the second timestamp of the end of playback for the target image element corresponding to the target audio source in the video to be processed;
[0234] One or more spatially rendered target audio source segments are randomly superimposed between the first timestamp and the second timestamp of the audio file.
[0235] Preferably, before randomly extracting multiple equal-length target audio source segments from the target audio source, the acquisition unit 1103 is further configured to:
[0236] If the duration of the target audio source is greater than the time difference between the first timestamp and the second timestamp, then a valid audio source segment is extracted from the target audio source.
[0237] Multiple target audio segments of equal length are randomly selected from the effective audio segments.
[0238] Preferably, the extraction unit 1101 is specifically used for:
[0239] Perform audio framing on the target audio source to obtain multiple frames of audio source signal;
[0240] Calculate the energy value of each frame of the audio source signal to obtain multiple energy values for multiple frames of audio source signals;
[0241] Select multiple target energy values that are greater than a preset energy threshold from a range of energy values;
[0242] The multi-frame audio source signals corresponding to multiple target energy values are determined as valid audio source signals;
[0243] Obtain all timestamps of valid audio source signals;
[0244] Extract valid audio segments from the target audio source based on all timestamps.
[0245] Preferably, if the audio adjustment element includes the associative element sound source of the target image element, the acquisition unit 1103 is specifically used for:
[0246] Extract the semantic vectors of the target image element identifiers;
[0247] Extract the semantic vectors of all audio source identifiers in the audio source library;
[0248] Calculate the similarity value between the semantic vector of the target image element identifier and the semantic vector of all music identifiers to obtain multiple similarity values;
[0249] From the plurality of similarity values, determine a plurality of target similarity values that are greater than a preset similarity threshold;
[0250] The multiple sound sources corresponding to the multiple target similarity values are determined as multiple associative element sound sources of the target image elements.
[0251] Preferably, if the audio adjustment element includes the associative element sound source of the target image element, the acquisition unit 1103 is specifically used for:
[0252] Obtain multiple associative element sound sources for the target image element;
[0253] The multiple associative element sound sources are randomly arranged and combined to obtain associative element sound sources combined in the first order;
[0254] Obtain the first timestamp of the target image element in the video to be processed, and the second timestamp of the start and end timestamp of playback.
[0255] Obtain any target timestamp before the first timestamp and after the second timestamp;
[0256] The associative element sound sources are superimposed in the first order at any of the target timestamps of the audio file.
[0257] Preferably, before randomly arranging and combining multiple associative element sound sources, the rendering unit 1104 is also used for:
[0258] The motion trajectory of each of the multiple associative element sound sources, each randomly initialized associative element sound source, and the preset music format are input into the spatial renderer to obtain multiple spatially rendered associative element sound sources.
[0259] The sound sources of the multiple spatially rendered associative elements are randomly arranged and combined.
[0260] Preferably, if the playback duration of the associated element sound source is greater than the time difference between the first and second timestamps, then before randomly sorting and combining multiple associated element sound sources, the extraction unit 1101 is further used for:
[0261] Multiple associative element sound source fragments of equal length are randomly extracted from the associative element sound source;
[0262] Rendering unit 1104 is also used for:
[0263] The multiple associative element sound source fragments of equal length are randomly arranged and combined.
[0264] For ease of description and brevity, the specific working process of the system, device and unit described above can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0265] In this embodiment of the application, the extraction unit 1101 extracts element image frames containing target image elements from the video to be processed; the target image elements are image elements capable of producing sound effects; the generation unit 1102 generates the motion trajectory of the target image elements in the video to be processed based on the position of the target image elements in each element image frame and the playback time point corresponding to each element image frame in the video to be processed; the acquisition unit 1103 acquires a target sound source matching the target image elements; the rendering unit 1104 performs spatial rendering on the target sound source using a rendering method matching the motion trajectory of the target image elements to obtain a spatially rendered target sound source; an audio file corresponding to the video to be processed is generated based on the spatially rendered target sound source; the audio file is used to synchronously play the target image elements and the spatially rendered target sound source matching the position of the target image elements in the target element image frames when the video to be processed plays to the target element image frame in the element image frame.
[0266] Because the audio file generated in this application embodiment, when the video to be processed plays to the target element image frame containing the target image, synchronously plays the target sound source that matches the target image element and the position of the target image element in the target element image frame, thereby realizing the synchronous playback of the audio file in the video to be processed and the target image in the video to be processed, that is, realizing the audio-visual synchronization in the video to be processed.
[0267] The above description of the audio generation apparatus in the embodiments of the present invention is from the perspective of modular functional entities. The following description of the computer device in the embodiments of the present invention is from the perspective of hardware processing:
[0268] One embodiment of the computer device in this invention includes:
[0269] Processor and memory;
[0270] The memory is used to store computer programs, and when the processor executes the computer programs stored in the memory, it can implement the method steps described in the above method embodiments.
[0271] It is understood that when the processor in the computer device described above executes the computer program, it can also implement the functions of each unit in the corresponding device embodiments described above, which will not be repeated here. For example, the computer program can be divided into one or more modules / units, which are stored in the memory and executed by the processor to complete the present invention. The one or more modules / units can be a series of computer program instruction segments capable of performing specific functions, which describe the execution process of the computer program in the audio generation device. For example, the computer program can be divided into units in the audio generation device described above, and each unit can implement the specific functions described in the corresponding audio generation device above.
[0272] The computer device may be a desktop computer, laptop, handheld computer, or cloud server, etc. The computer device may include, but is not limited to, a processor and memory. Those skilled in the art will understand that the processor and memory are merely examples of a computer device and do not constitute a limitation on the computer device. It may include more or fewer components, or a combination of certain components, or different components. For example, the computer device may also include input / output devices, network access devices, buses, etc.
[0273] The processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor. The processor is the control center of the computer device, connecting various parts of the computer device via various interfaces and lines.
[0274] The memory can be used to store the computer programs and / or modules. The processor implements various functions of the computer device by running or executing the computer programs and / or modules stored in the memory and by calling data stored in the memory. The memory may mainly include a program storage area and a data storage area. The program storage area may store the operating system, at least one application program required for a function, etc.; the data storage area may store data created according to the use of the terminal, etc. In addition, the memory may include high-speed random access memory, and may also include non-volatile memory, such as hard disk, RAM, plug-in hard disk, smart media card (SMC), secure digital card (SD card), flash card, at least one disk storage device, flash memory device, or other volatile solid-state storage device.
[0275] The present invention also provides a computer-readable storage medium for implementing the function of an audio generation device, wherein a computer program is stored thereon, and when the computer program is executed by a processor, the processor can be used to perform the method steps described in the above method embodiments.
[0276] It is understood that if the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a corresponding computer-readable storage medium. Based on this understanding, all or part of the processes in the above-described embodiments of the present invention can also be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the above-described method embodiments. The computer program includes computer program code, which can be in the form of source code, object code, executable file, or some intermediate form. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording media, USB flash drive, portable hard drive, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content included in the computer-readable medium can be appropriately added or removed according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer-readable media do not include electrical carrier signals and telecommunication signals.
[0277] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be an indirect coupling or communication connection between apparatuses or units through some interfaces, and may be electrical, mechanical, or other forms.
[0278] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0279] Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0280] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for generating audio, characterized in that, include: Extract element image frames containing target image elements from the video to be processed; The target image element is an image element capable of producing sound effects; Based on the position of the target image element in each element image frame and the playback time point corresponding to each element image frame in the video to be processed, the motion trajectory of the target image element in the video to be processed is generated. Obtain the target sound source that matches the target image element; Spatial rendering is performed on the target sound source using a rendering method that matches the motion trajectory of the target image elements to obtain the spatially rendered target sound source. The motion trajectory includes static motion trajectory and dynamic motion trajectory, and the rendering method includes channel-based rendering method and object-based rendering. An audio file corresponding to the video to be processed is generated based on the spatially rendered target sound source; the audio file is used to synchronously play the spatially rendered target sound source that matches the target image element and the position of the target image element in the target element image frame when the video to be processed is played to the target element image frame in the element image frame.
2. The method according to claim 1, characterized in that, The step of performing spatial rendering on the target sound source using a rendering method that matches the motion trajectory of the target image elements to obtain the spatially rendered target sound source includes: If the motion trajectory of the target image element includes the static motion trajectory, then the channel-based rendering method is used to perform spatial rendering on the target sound source to obtain the spatially rendered target sound source. If the motion trajectory of the target image element includes the dynamic motion trajectory, then the object-based rendering method is used to perform spatial rendering on the target sound source to obtain the spatially rendered target sound source.
3. The method according to claim 2, characterized in that, The step of performing spatial rendering on the target sound source using the channel-based rendering method to obtain the spatially rendered target sound source includes: An upmixing algorithm is used to separate the relevant and irrelevant signals from the left and right channels of the target sound source; The relevant and unrelated signals are placed in different channels according to a preset mixing framework to obtain the spatially rendered target sound source.
4. The method according to claim 2, characterized in that, The step of performing spatial rendering on the target sound source using the object-based rendering method to obtain the spatially rendered target sound source includes: The target sound source, the motion trajectory of the target image element, and the preset multi-channel music format are input into the spatial renderer to obtain the spatially rendered target sound source output by the spatial renderer.
5. The method according to claim 4, characterized in that, The method further includes: If the motion trajectory of the target image element in the video to be processed is along the first direction, then the target audio source is subjected to automatic volume processing and compression processing to obtain the processed target audio source; If the motion trajectory of the target image element in the multiple image frames is along the second direction, then reverberation processing and delay processing are performed on the target sound source to obtain the processed target sound source; the second direction is perpendicular to the first direction; The step of inputting the target sound source, the motion trajectory of the target image element, and the preset multi-channel music format into the spatial renderer includes: The processed target sound source, the motion trajectory of the target image element, and the preset multi-channel music format are input into the spatial renderer.
6. The method according to claim 1, characterized in that, The number of target image elements is at least two, the number of rendered target sound sources is at least two, and the step of generating the audio file corresponding to the video to be processed based on the rendered target sound sources includes: Based on the timestamps in the motion trajectories of each target image element, each rendered target audio source corresponding to the video to be processed is aligned with the timestamps of its respective target image elements. The aligned and rendered target sound sources are superimposed to obtain the superimposed sound source; Adaptive volume balancing is performed on the superimposed audio source to obtain the audio file corresponding to the video to be processed.
7. The method according to claim 6, characterized in that, If the video to be processed is played in a loop, then after obtaining the audio file corresponding to the video to be processed, the method further includes: Obtain audio adjustment elements for the video to be processed, wherein the audio adjustment elements include at least one of the following: target light music, associative element sound source of the target image element, and sampled segment of the target sound source; The audio adjustment element is superimposed on the audio file to obtain an extended audio file, wherein the audio file is used for playback when the video to be processed is played for the first time, and the extended audio file is used for playback when the video to be processed is played in a loop.
8. The method according to claim 7, characterized in that, If the audio adjustment elements include the target light music, obtaining the audio adjustment elements for the video to be processed includes: The target light music is selected from the preset light music library based on the similarity strategy. The step of overlaying the audio adjustment elements onto the audio file to obtain an expanded audio file includes: The target light music is processed using a channel-based rendering method to obtain a processed multi-channel light music sound source; The processed multi-channel light music source is superimposed on the source of the audio file to obtain an extended audio file.
9. The method according to claim 8, characterized in that, Target light music is selected from a pre-defined light music library based on a similarity strategy, including: Extract the scene semantic vector of the video to be processed; Extract the semantic vector of each light music scene tag in the light music library to obtain scene semantic vectors for multiple light music tracks; The similarity between the scene semantic vector of the video to be processed and the scene semantic vector of each piece of light music is calculated to obtain multiple similarity values; The light music corresponding to the similarity value that meets the preset standard among the multiple similarity values is determined as the target light music.
10. The method according to claim 9, characterized in that, The extraction of the scene semantic vector from the video to be processed includes: Extract multiple element image frames from the video to be processed; Extract all target image elements from the multiple element image frames; Obtain the identifiers of each target image element; Calculate the semantic vector corresponding to the identifier of each target image element to obtain multiple semantic vectors for each target image element; The average value of the multiple semantic vectors is calculated to obtain the scene semantic vector of the video to be processed.
11. The method according to claim 8, characterized in that, The processed multi-channel light music audio source is superimposed on the audio source of the audio file to obtain an extended audio file, including: The processed multi-channel light music audio is overlaid with the audio file using a fade-in / fade-out method to obtain an extended audio file.
12. The method according to claim 7, characterized in that, If the audio adjustment element includes a sample segment of the target audio source, the step of obtaining the audio adjustment element for the video to be processed includes: Randomly extract multiple target sound source segments of equal length from the target sound source; The step of overlaying the audio adjustment elements onto the audio file to obtain an expanded audio file includes: If the motion trajectory of the target image element corresponding to the target sound source includes a static motion trajectory, then a channel-based rendering method is used to render the multiple equal-length target sound source segments to obtain multiple spatially rendered target sound source segments. If the motion trajectory of the target image element corresponding to the target sound source includes a dynamic motion trajectory, then an object-based rendering method is used to render the multiple equal-length target sound source segments to obtain multiple spatially rendered target sound source segments. Obtain the first timestamp of the start of playback and the second timestamp of the end of playback for the target image element corresponding to the target audio source in the video to be processed; One or more spatially rendered target audio source segments are randomly superimposed between the first timestamp and the second timestamp of the audio file.
13. The method according to claim 12, characterized in that, Before randomly extracting multiple equal-length target audio source segments from the target audio source, the method further includes: If the duration of the target audio source is greater than the time difference between the first timestamp and the second timestamp, then a valid audio source segment is extracted from the target audio source. The step of randomly extracting multiple equal-length target audio source segments from the target audio source includes: Multiple target audio segments of equal length are randomly selected from the effective audio segments.
14. The method according to claim 13, characterized in that, The step of extracting effective audio source segments from the target audio source includes: The target audio source is subjected to audio framing to obtain multiple frames of audio source signals; Calculate the energy value of each frame of the audio source signal to obtain multiple energy values for multiple frames of audio source signals; Select multiple target energy values that are greater than a preset energy threshold from the plurality of energy values; The multi-frame audio source signals corresponding to the multiple target energy values are determined as valid audio source signals; Obtain all timestamps of the valid audio source signal; Based on all the timestamps, extract valid audio segments from the target audio source.
15. The method according to claim 7, characterized in that, If the audio adjustment element includes an associative element sound source of the target image element, obtaining the audio adjustment element for the video to be processed includes: Extract the semantic vectors of the target image element identifiers; Extract the semantic vectors of all audio source identifiers in the audio source library; Calculate the similarity value between the semantic vector of the target image element identifier and the semantic vector of all sound source identifiers to obtain multiple similarity values; From the plurality of similarity values, determine a plurality of target similarity values that are greater than a preset similarity threshold; The multiple sound sources corresponding to the multiple target similarity values are determined as multiple associative element sound sources of the target image elements.
16. The method according to claim 7, characterized in that, If the audio adjustment element includes the associative element sound source of the target image element, the step of superimposing the audio adjustment element onto the audio file to obtain an extended audio file includes: Obtain multiple associative element sound sources for the target image element; The multiple associative element sound sources are randomly arranged and combined to obtain associative element sound sources combined in the first order; Obtain the first timestamp of the target image element in the video to be processed, and the second timestamp of the start and end timestamp of playback. Obtain any target timestamp before the first timestamp or after the second timestamp; The associative element sound sources are superimposed in the first order at any of the target timestamps of the audio file.
17. The method according to claim 16, characterized in that, Before randomly arranging and combining the multiple associative element sound sources, the method further includes: The motion trajectory of each of the multiple associative element sound sources, each randomly initialized associative element sound source, and the preset music format are input into the spatial renderer to obtain multiple spatially rendered associative element sound sources. Randomly arranging and combining the multiple associative element sound sources includes: The sound sources of the multiple spatially rendered associative elements are randomly arranged and combined.
18. The method according to claim 17, characterized in that, If the playback duration of the associative element sound source is greater than the time difference between the first timestamp and the second timestamp, then before randomly sorting and combining the multiple associative element sound sources, the method further includes: Randomly extract multiple associative element sound source fragments of equal length from the aforementioned associative element sound sources; The step of randomly arranging and combining the multiple associative element sound sources includes: The multiple associative element sound source fragments of equal length are randomly arranged and combined.
19. A computer device comprising a processor, characterized in that, When the processor executes a computer program stored in a memory, it is used to implement the method for generating audio as described in any one of claims 1 to 18.
20. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it is used to implement the method for generating audio as described in any one of claims 1 to 18.
21. A computer program product storing computer instructions, characterized in that, When the computer instructions are executed by a processor, they are used to implement the method for generating audio as described in any one of claims 1 to 18.
Citation Information
Patent Citations
Image editing device, image editing method and image editing program
JP2013114236A
Electronic device for forming content and method for operating same
WO2021153953A1