Video generation method, live broadcast system and related device

By analyzing the rhythm information of the live audio signal in real time and dynamically adjusting the virtual camera position, the problem of audio and video mismatch in the existing technology is solved, and a richer live visual effect and a higher user experience is achieved.

CN120050439APending Publication Date: 2025-05-27BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510185864.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-19
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The existing live broadcast technology is difficult to match the dynamic effects of real-time audio-driven with video images, and cannot effectively improve the audience's audio-visual experience and sense of participation.

Method used

By obtaining the audio signals collected in real time during the live broadcast process, performing spectrum analysis to determine rhythm information, dynamically adjusting the position information of the virtual camera, and synthesizing the virtual scene with the live broadcast screen collected in real time to generate live videos.

Benefits of technology

Real-time synchronization of virtual camera movement and audio rhythm is achieved, real-time and interactiveness of live broadcasts are enhanced, the visual effect of live broadcasts is enriched, and the overall expressiveness of live broadcasts and the user's viewing experience is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120050439A_ABST
    Figure CN120050439A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a video generation method, a live broadcast system and a related device. The technical scheme comprises the following steps: acquiring an audio signal acquired in real time in a live broadcast process; spectral analysis is carried out on the audio signals, rhythm information in the audio signals is determined, and the rhythm information comprises time points corresponding to rhythm points; according to rhythm information in the audio signal, position information of a virtual camera is determined, and the position information at least comprises position information of the virtual camera at a time point corresponding to the rhythm point; rendering a virtual scene according to the position information of the virtual camera; synthesizing the virtual scene with a live broadcast picture collected in real time to generate a live broadcast video; according to the invention, real-time synchronization of the motion of the virtual camera and the audio rhythm is realized, so that the motion of the virtual camera can accurately reflect the rhythm change of the audio, the visual effect of the live video is enhanced, and the overall expressive force of the live and the watching experience of the user are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of Internet application technologies, and particularly to a video generation method, a live broadcast system, and related devices. Background Art

[0002] With the popularization of webcasting and the development of technologies, the requirements of audiences for the quality and viewing experience of live broadcast content are increasing day by day. Especially in specific scenarios such as music live broadcasts and dance performances, how to achieve the matching of dynamic effects driven by real-time audio with video images to create a better audio-visual experience for the audience has become one of the hotspots in current live broadcast research.

[0003] In traditional live programs, camera movements are often static or simple zooms, which cannot effectively match the rhythm changes of music. Therefore, a new method is needed to achieve virtual camera movements driven by real-time audio to enhance the audio-visual experience and sense of participation of the audience. Summary of the Invention

[0004] In view of this, this application provides a video generation method, a live broadcast system, and related devices to achieve the real-time synchronization of virtual camera movements and audio rhythms, enrich the visual effects of live broadcast images, and enhance the overall expressiveness of live broadcasts and the viewing experience of users.

[0005] This application provides the following solutions:

[0006] In a first aspect, a video generation method is provided, and the method includes:

[0007] Obtain an audio signal collected in real time during a live broadcast;

[0008] Perform spectral analysis on the audio signal to determine the rhythm information in the audio signal, where the rhythm information includes the time points corresponding to the rhythm points;

[0009] Determine the position information of a virtual camera according to the rhythm information in the audio signal, where the position information at least includes the position information of the virtual camera at the time points corresponding to the rhythm points;

[0010] Render a virtual scene according to the position information of the virtual camera;

[0011] Synthesize the virtual scene with the live broadcast images collected in real time to generate a live broadcast video.

[0012] Optionally, the performing spectral analysis on the audio signal to determine the rhythm information in the audio signal includes:

[0013] Divide the audio signal into a plurality of time windows according to time sequence, and each time window corresponds to a segment of sub-audio signal in the audio signal;

[0014] Perform short-time Fourier transform on the sub-audio signals within each time window to obtain the spectral information of each sub-audio signal;

[0015] Extract the rhythm information within the target frequency range according to the spectral information of each sub-audio signal.

[0016] Optionally, the rhythm information further includes the intensity value of the rhythm point; the extracting the rhythm information within the target frequency range according to the spectral information of each sub-audio signal includes:

[0017] Calculate the energy envelope value of each time window within the target frequency range according to the spectral information of each sub-audio signal;

[0018] When the energy envelope value exceeds the preset energy threshold, determine the time point corresponding to the rhythm point as the time point corresponding to the rhythm point, and determine the energy envelope value as the intensity value of the rhythm point.

[0019] Optionally, the rhythm information further includes the intensity value of the rhythm point; determining the position information of the virtual camera according to the rhythm information in the audio signal includes:

[0020] Determine the rhythm amplitude of the virtual camera at the time point corresponding to the rhythm point according to the intensity value of the rhythm point;

[0021] Determine the position information of the virtual camera at the time point corresponding to the rhythm point according to the rhythm amplitude of the virtual camera at the time point corresponding to the rhythm point.

[0022] Optionally, the determining the rhythm amplitude of the virtual camera at the time point corresponding to the rhythm point according to the intensity value of the rhythm point includes:

[0023] Determine the attenuation intensity value of the previous rhythm point of the rhythm point; wherein, the attenuation intensity value is the difference between the intensity value of the previous rhythm point and the preset decline speed;

[0024] Update the maximum value of the intensity value of the rhythm point and the attenuation intensity value of the previous rhythm point as the new intensity value of the rhythm point;

[0025] Perform normalization processing on the new intensity value of the rhythm point to obtain the normalized new intensity value;

[0026] Determine the rhythm amplitude of the virtual camera at the time point corresponding to the rhythm point according to the normalized new intensity value.

[0027] Optionally, the determining the rhythm amplitude of the virtual camera at the time point corresponding to the rhythm point according to the intensity value of the rhythm point includes:

[0028] Map the intensity value of the rhythm point according to a preset mapping relationship to obtain the rhythm amplitude of the virtual camera at the time point corresponding to the rhythm point.

[0029] Optionally, determine the position information of the virtual camera at the time point corresponding to the rhythm point according to the rhythm amplitude of the virtual camera at the time point corresponding to the rhythm point, including:

[0030] Determine the position information of the virtual camera at the time point corresponding to the rhythm point according to the normalized new intensity value and the rhythm amplitude of the time point corresponding to the rhythm point, where the position information includes the depth position of the virtual camera in the virtual scene.

[0031] Optionally, determine the position information of the virtual camera at the time point corresponding to the rhythm point according to the normalized new intensity value and the rhythm amplitude of the time point corresponding to the rhythm point, including:

[0032] Use the normalized new intensity value as a scaling factor to perform linear interpolation between the original depth position and the enlarged depth position of the virtual camera to obtain the depth position of the virtual camera at the time point corresponding to the rhythm point; wherein, the enlarged depth position is determined based on the product of the original depth position of the virtual camera and the rhythm amplitude.

[0033] Optionally, the position information of the virtual camera further includes: the position information of the virtual camera at other time points except the time point corresponding to the rhythm point among the time points included in the audio signal;

[0034] The method further includes: using the original position of the virtual camera and the position information of the virtual camera at the time point corresponding to the rhythm point to perform trajectory fitting to obtain the position information of the virtual camera at other time points.

[0035] Optionally, synthesize the virtual scene with a live broadcast picture collected in real time to generate the live broadcast video, including:

[0036] Extract the image area of the live broadcast object from the live broadcast picture;

[0037] Synthesize the extracted image area of the live broadcast object with the virtual scene according to the corresponding time points to generate the live broadcast video.

[0038] According to a second aspect, there is provided a video generation device, the device includes:

[0039] An audio acquisition unit, configured to acquire an audio signal collected in real time during a live broadcast;

[0040] A spectrum analysis unit, configured to perform spectrum analysis on the audio signal to determine rhythm information in the audio signal, where the rhythm information includes time points corresponding to rhythm points;

[0041] A trajectory determination unit, configured to determine position information of a virtual camera according to the rhythm information in the audio signal, where the position information at least includes position information of the virtual camera at time points corresponding to the rhythm points;

[0042] A video synthesis unit, configured to render a virtual scene according to the position information of the virtual camera; synthesize the virtual scene with a live broadcast image collected in real time to generate a live broadcast video.

[0043] According to a third aspect, a live broadcast system is provided. The live broadcast system includes a first terminal, a sound collection device, and a second terminal, where:

[0044] The first terminal is configured to play audio during a live broadcast;

[0045] The sound collection device is configured to collect an audio signal during a live broadcast in real time and send it to the second terminal;

[0046] The second terminal is configured to perform spectrum analysis on the audio signal to determine rhythm information in the audio signal, where the rhythm information includes time points corresponding to rhythm points; determine position information of a virtual camera according to the rhythm information in the audio signal, where the position information at least includes position information of the virtual camera at time points corresponding to the rhythm points; render a virtual scene according to the position information of the virtual camera; synthesize the virtual scene with a live broadcast image collected in real time to generate a live broadcast video; transmit the live broadcast video to the first terminal;

[0047] The first terminal is further configured to receive the live broadcast video and transmit it to a live broadcast server.

[0048] In a fourth aspect, a computer-readable storage medium is provided, on which a computer program is stored. When the program is executed by a processor, the steps of the method described in any one of the first aspects are implemented.

[0049] In a fifth aspect, an electronic device is provided, including:

[0050] One or more processors; and

[0051] A memory associated with the one or more processors, where the memory is used to store program instructions. When the program instructions are read and executed by the one or more processors, the steps of the method described in any one of the first aspects are executed.

[0052] In a sixth aspect, a computer program product is provided, including a computer program which, when executed by a processor, implements the steps of the method described in any one of the above first aspects.

[0053] According to the specific embodiments provided in the present application, the following technical effects are disclosed in the present application:

[0054] 1) In the present application, by acquiring the audio signal collected in real time during the live broadcast, performing spectrum analysis on it to determine the rhythm information, dynamically adjusting the position information of the virtual camera according to the rhythm information, and then rendering the virtual scene according to the position information of the virtual camera and synthesizing the virtual scene with the live broadcast image collected in real time to generate a live video. This process realizes the real-time synchronization of the movement of the virtual camera and the audio rhythm in the live broadcast scenario, enables the movement of the virtual camera to accurately reflect the rhythm changes of the audio, enhances the real-time and interactivity of the live broadcast, enriches the visual effect of the live broadcast image, and improves the overall expressiveness of the live broadcast and the viewing experience of users.

[0055] 2) In the present application, by performing spectrum analysis on the audio signal, dividing the audio signal into multiple time windows according to time sequence, and performing short-time Fourier transform on each sub-audio signal to obtain spectrum information, and then extracting the rhythm information within the target frequency range. This process can capture the rhythm points in the audio signal more sensitively and accurately, ensuring the accurate matching of the movement of the virtual camera and the audio rhythm, and improving the visual effect of the live broadcast image and the viewing experience of users.

[0056] 3) In the present application, according to the spectrum information of each sub-audio signal, the energy envelope value of each time window within the target frequency range is calculated, and when the energy envelope value exceeds the preset energy threshold, the time point and intensity value corresponding to the rhythm point are determined. This process can effectively extract the rhythm information in the audio signal, avoid misjudgment caused by noise and interference, improve the accuracy and reliability of rhythm point detection, provide a reliable data basis for the movement control of the virtual camera, and further improve the visual effect of the live broadcast image and the viewing experience of users.

[0057] 4) In the present application, the rhythm amplitude is determined according to the intensity value of the rhythm point, so that the movement of the virtual camera can more accurately reflect the rhythm changes of the audio. This process enhances the real-time and interactivity of the live broadcast, enriches the visual effect of the live broadcast image, and improves the overall expressiveness of the live broadcast and the viewing experience of users.

[0058] 5) This application takes into account the attenuation intensity and the falling-back speed of the previous rhythm point. By determining the new intensity value, normalizing, and determining the rhythm amplitude, etc., the movement of the virtual camera becomes smoother and more natural, avoiding incoherence caused by sudden changes in the intensity of the rhythm point. This process improves the stability and smoothness of the virtual camera movement, enhances the real-time performance and interactivity of the live broadcast, and improves the visual effect of the live broadcast screen and the viewing experience of users.

[0059] 6) This application maps the intensity value of the rhythm point according to the preset mapping relationship to obtain the rhythm amplitude. This process simplifies the method for determining the rhythm amplitude, improves the calculation efficiency, ensures the precise matching of the virtual camera movement with the audio rhythm, and improves the visual effect of the live broadcast screen and the viewing experience of users.

[0060] 7) This application makes the movement of the virtual camera richer and more vivid by determining the depth position of the virtual camera in the virtual scene, enhancing the sense of hierarchy and three-dimensionality of the live broadcast screen. This process further improves the overall expressiveness of the live broadcast and the viewing experience of users.

[0061] 8) This application uses the normalized new intensity value as a scaling factor to perform linear interpolation between the original depth position and the enlarged depth position of the virtual camera to obtain the depth position of the virtual camera in the virtual scene. This process makes the adjustment of the depth position of the virtual camera smoother and more natural, avoiding sudden changes, and improving the visual effect of the live broadcast screen and the viewing experience of users.

[0062] 9) In addition to considering the position information of the virtual camera at the time point corresponding to the rhythm point, this application uses the original position of the virtual camera and the position information of the virtual camera at the time points corresponding to each rhythm point to perform trajectory fitting to obtain the position information of the virtual camera at other time points included in the collected audio signal. This way makes the position information of the virtual camera smoother and more detailed, improving the visual coherence of the live broadcast video.

[0063] Of course, it is not necessary for any product implementing this application to achieve all the above-mentioned advantages simultaneously. BRIEF DESCRIPTION OF THE DRAWINGS

[0064] In order to more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the following will briefly introduce the drawings required for the embodiments. Obviously, the drawings in the following description are only some embodiments of this application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0065] Figure 1 FIG. is a flowchart of the video generation method provided by the embodiment of this application;

[0066] Figure 2 Schematic diagram of the "dual-machine shunt" architecture provided by an embodiment of the present application;

[0067] Figure 3 Schematic diagram of using STFT to extract the drumbeat information of an audio signal provided by an embodiment of the present application;

[0068] Figure 4 Schematic diagram of the implementation process of spectrum analysis provided by an embodiment of the present application;

[0069] Figure 5 Schematic diagram of determining the position information of a virtual camera provided by an embodiment of the present application;

[0070] Figure 6 Schematic diagram of virtual-real synthesis provided by an embodiment of the present application;

[0071] Figure 7 Schematic block diagram of a video generation device provided by an embodiment of the present application;

[0072] Figure 8 Schematic block diagram of a live broadcast system provided by an embodiment of the present application;

[0073] Figure 9 Schematic block diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0074] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art belong to the scope of protection of the present application.

[0075] The terms used in the embodiments of the present invention are only for the purpose of describing specific embodiments, and are not intended to limit the present invention. The singular forms of "a", "the" and "said" used in the embodiments of the present invention and the appended claims are also intended to include the plural forms, unless the context clearly indicates otherwise.

[0076] It should be understood that the term " / and / " used herein is only a description of the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " herein generally represents an "or" relationship between the associated objects before and after.

[0077] Depending on the context, as used herein, the word "if" can be interpreted as "when" or "while" or "in response to determining" or "in response to detecting". Similarly, depending on the context, the phrase "if determined" or "if detecting (stated condition or event)" can be interpreted as "when determined" or "in response to determining" or "when detecting (stated condition or event)" or "in response to detecting (stated condition or event)".

[0078] Currently, in the live broadcast industry, two technical solutions, namely offline beat synchronization and handheld camera movement, are mainly adopted. The offline beat synchronization mode can establish a rhythmic connection between audio and video, while the handheld camera movement mode can capture more dynamic pictures through professional operations. However, these existing solutions have many problems, such as lack of real-time performance, insufficient adaptability, high operation threshold, high cost, and limited operation, which seriously restrict their practical applications in various live broadcast scenarios and are difficult to meet the current diverse live broadcast needs.

[0079] In view of this, the present application provides a new idea. A video generation method is provided. As Figure 1 shown, it is a flowchart of the video generation method provided by an embodiment of the present application. The execution subject of this method can be a video generation device, which can be set in any computer device with data storage and processing capabilities, such as a computer terminal with strong data processing capabilities, and this computer terminal can be the second terminal involved in subsequent embodiments. This method may include the following steps:

[0080] Step 101: Obtain the audio signal collected in real time during the live broadcast.

[0081] Step 102: Perform spectral analysis on the audio signal to determine the rhythm information in the audio signal, where the rhythm information includes the time points corresponding to the rhythm points.

[0082] Step 103: Determine the position information of the virtual camera according to the rhythm information in the audio signal, and the position information at least includes the position information of the virtual camera at the time points corresponding to the rhythm points.

[0083] Step 104: Render the virtual scene according to the position information of the virtual camera, and synthesize the virtual scene with the live broadcast picture collected in real time to generate a live broadcast video.

[0084] It can be seen that in this application, by acquiring the audio signal collected in real time during the live broadcast, performing spectrum analysis on it to determine the rhythm information, dynamically adjusting the position information of the virtual camera according to the rhythm information, then rendering the virtual scene according to the position information of the virtual camera, and synthesizing the virtual scene with the live broadcast image collected in real time, a live video is generated. This process realizes the real-time synchronization of the movement of the virtual camera and the audio rhythm in the live broadcast scenario, enables the movement of the virtual camera to accurately reflect the rhythm changes of the audio, enhances the real-time and interactivity of the live broadcast, enriches the visual effect of the live broadcast image, and improves the overall expressiveness of the live broadcast and the viewing experience of users.

[0085] The following will describe in detail each step in the above process and the further effects that can be generated in combination with the embodiments. It should be noted that the "first", "second", etc. limitations involved in this disclosure do not have limitations in terms of size, order, quantity, etc., and are only used to distinguish in name. For example, "the first terminal" and "the second terminal" are used to distinguish two terminals in name. And so on.

[0086] First, the above step 101, that is, "acquiring the audio signal collected in real time during the live broadcast", will be described in detail in combination with the embodiments.

[0087] In the embodiments of this application, the acquisition of the audio signal can be achieved through various devices. Common devices include but are not limited to microphones, audio capture cards, etc. The microphone can convert the sound at the live broadcast site into an analog electrical signal, and the audio capture card can convert it into a digital audio signal through analog-to-digital conversion for subsequent analysis and processing.

[0088] In the current live broadcast scenario, especially in the group live broadcast scenario with multiple audio sources input or high-quality requirements (such as dance performances, band live broadcasts, etc.), due to the complexity and diversity of the audio and video sources, the "dual-machine shunt" architecture can be adopted for live broadcast operations in the embodiments of this application.

[0089] As Figure 2 shown, it is a schematic diagram of the "dual-machine shunt" architecture provided by the embodiments of this application. Among them, the "dual-machine shunt" architecture consists of a first terminal (such as computer A in the figure), a sound acquisition device (such as the sound capture card in the figure), and a second terminal (such as computer B in the figure).

[0090] The first terminal is responsible for playing the audio signal during the live broadcast.

[0091] The sound acquisition device is responsible for collecting the audio signal during the live broadcast in real time and sending it to the second terminal.

[0092] The second terminal receives the audio signal and performs subsequent processing using the method provided in the embodiments of the present application to generate live audio. This acquisition method can reduce the delay and distortion in the transmission of the audio signal. At the same time, separating the subsequent audio and video processing to the second terminal can also improve the stability of the performance of live audio and video.

[0093] The above step 102, that is, "performing spectral analysis on the audio signal to determine the rhythm information in the audio signal", will be described in detail in combination with the embodiments.

[0094] During the live broadcast, the real-time acquired audio signal needs to be spectrally analyzed to determine the rhythm information therein. In the embodiments of the present application, the rhythm information may, but is not limited to, include the time points corresponding to the rhythm points, and may also include the intensity values of the rhythm points. In addition, it may also include other rhythm-related information. Here, the rhythm points usually refer to the points in the audio signal with obvious rhythm characteristics, and these points often correspond to the accents, beats, or drum beats in music or sound. In audio signal processing, the rhythm points can be determined by spectral analysis of the audio signal, usually manifested as the energy peaks or significant changes of the audio signal at specific time points.

[0095] Optionally, before performing spectral analysis on the audio signal in the embodiments of the present application, the acquired audio signal can be preprocessed first. The purpose of preprocessing is to improve the quality of the audio signal, remove unnecessary noise and interference, so as to more accurately extract the rhythm information.

[0096] Among them, the preprocessing steps may, but are not limited to, include:

[0097] Denoising, such as using a filter to remove the background noise in the audio signal. Common filters include low-pass filters, high-pass filters, and band-pass filters, etc., and appropriate filter types and parameters can be selected according to actual needs.

[0098] Gain adjustment, adjusting the amplitude of the audio signal to make it within a suitable range. Gain adjustment can avoid the influence of the signal being too strong or too weak on subsequent analysis.

[0099] The preprocessed audio signal can be spectrally analyzed using a spectral analysis algorithm, such as the Short-time Fourier Transform (STFT), wavelet transform, etc.

[0100] The implementation method of extracting the rhythm information in the audio signal will be introduced below taking STFT as an example.

[0101] STFT is a common spectrum analysis method that divides a long audio signal into multiple shorter sub-audio signals of equal length, and then performs Fourier transform on each sub-audio signal to obtain the spectrum information of each sub-audio signal at different time points.

[0102] The specific implementation can be achieved according to the following steps:

[0103] First, divide the audio signal into multiple time windows according to the time sequence, and each time window corresponds to a segment of sub-audio signal in the audio signal. Among them, the length of each time window is usually from dozens of milliseconds to hundreds of milliseconds. The purpose of dividing the time window here is to convert the non-stationary signal into an approximately stationary short segment signal for Fourier transform.

[0104] Then perform STFT on the sub-audio signal within each time window to obtain the spectrum information of each sub-audio signal.

[0105] Finally, according to the spectrum information of each sub-audio signal, extract the rhythm information within the target frequency range. Among them, the rhythm information can be the time point and intensity value corresponding to the rhythm point.

[0106] As one of the achievable ways, when extracting the rhythm information within the target frequency range according to the spectrum information of each sub-audio signal, the energy envelope value of each time window within the target frequency range can be calculated according to the spectrum information of each sub-audio signal. When the energy envelope value exceeds the preset energy threshold, the time point corresponding to the time window is determined as the time point corresponding to the rhythm point, and the energy envelope value is determined as the intensity value of the rhythm point.

[0107] For example, square and sum the frequency amplitudes of each time window within the target frequency range to obtain the energy envelope value of each time window within the target frequency range. Its calculation formula can be as follows:

[0108]

[0109] Among them, E(t) represents the energy envelope value within time window t, X(t,f) represents the amplitude of frequency f within time window t, and f1 and f2 are respectively the lower and upper limits of the target frequency range.

[0110] Optionally, after calculating the energy envelope value, a low-pass filter can also be applied to smooth the noise in the energy envelope value, thereby reducing the interference of short-time energy mutations and improving the robustness of subsequent rhythm point detection.

[0111] When the smoothed energy envelope value exceeds the preset energy threshold, the time point corresponding to the time window is determined as the time point corresponding to the rhythm point, and the energy envelope value is determined as the intensity value of the rhythm point.

[0112] In addition to using the energy envelope value to determine the time points and intensity values corresponding to each rhythm point as described above, methods such as calculating non-linear energy operators and detecting energy mutation points can also be adopted, which will not be listed one by one here.

[0113] It should also be noted that the target frequency range in the embodiments of the present application can be preset according to the characteristics of the rhythm points. Taking the rhythm point as a drumbeat as an example, drumbeats are usually generated by percussion instruments (such as drums, cymbals, etc.), and their frequency range mainly focuses on the low-frequency and mid-frequency bands. Generally speaking, the frequency range of drumbeats is roughly between 5Hz and 150Hz. The frequency components within this range usually contain the main energy and characteristics of the drumbeats.

[0114] As Figure 3 shown, Figure 3 is a schematic diagram of using STFT to extract the drumbeat information of an audio signal. Among them, Figure 3 the left side is the waveform diagram of the audio signal in a certain time window, the abscissa is time, and the ordinate is amplitude; after processing it with STFT, Figure 3 the spectrum diagram on the right side is obtained, the abscissa is frequency, and the ordinate is the frequency amplitude. That is, the collected audio signal can be transformed from the time domain to the frequency domain through STFT, so as to facilitate the analysis of the frequency components and energy distribution of the audio signal, and then extract the drumbeat information.

[0115] As Figure 4 shown, Figure 4 is a schematic diagram of the implementation process of spectrum analysis. Among them, the sound acquisition card receives the microphone signal collected by the microphone as input and transmits it to computer B (i.e., the second terminal in the embodiments of the present application), and computer B performs acquisition initialization on it. Here, the acquisition initialization refers to the process of preprocessing the collected audio signal before spectrum analysis mentioned above. This process is the starting point of audio signal processing and provides basic data for subsequent spectrum analysis and rhythm information extraction.

[0116] During the spectrum analysis process, the key parameters involved are as follows:

[0117] Time window, that is, the time window mentioned above. Usually, one time window corresponds to one time point;

[0118] Smoothing range, which is used to set the smoothing range of the low-pass filter to smooth the noise in the energy envelope value. The size of the smoothing range will affect the smoothing effect and needs to be adjusted according to the characteristics of the actual audio signal;

[0119] Frequency band limit, which is used to set the target frequency range. The frequency band limit can be set according to the frequency characteristics and analysis purposes of the audio signal. For example, for drumbeat detection, it can be set to 5Hz to 150Hz;

[0120] The number of peaks is used to set the number of peaks to be detected in the spectrum, that is, the number of rhythm points. The setting of the number of peaks can be adjusted according to actual needs and the complexity of the audio signal.

[0121] By setting these key parameters above, effective spectral analysis can be performed on the audio signal to extract the rhythm information therein.

[0122] The above step 103, that is, "determine the position information of the virtual camera according to the rhythm information in the audio signal", will be described in detail in combination with the embodiments.

[0123] The position information of the virtual camera determined in this step is essentially the camera movement data of the virtual camera presented by the position information corresponding to multiple time points, where at least the position information of the virtual camera at the time points corresponding to each rhythm point is included.

[0124] When determining the camera movement data of the virtual camera, an important parameter is the rhythm amplitude. The rhythm amplitude refers to the movement range or amplitude of the virtual camera in the virtual scene, such as the up and down movement and the left and right movement amplitude of the virtual camera.

[0125] In the embodiments of the present application, the intensity value of the rhythm point reflects the energy size of the audio signal at this rhythm point. The larger the intensity value, the stronger the energy of the audio signal. Therefore, the rhythm amplitude of the virtual camera can be determined according to the intensity value of the rhythm point; then, according to the rhythm amplitude of the virtual camera at the time points corresponding to each rhythm point, the position information of the virtual camera at the time points corresponding to each rhythm point is determined.

[0126] As an implementable way, there is a mapping relationship between the rhythm amplitude and the intensity value of the rhythm point, and then the intensity value of the rhythm point can be mapped according to the preset mapping relationship to obtain the rhythm amplitude of the virtual camera.

[0127] Taking linear mapping as an example, linearly map the intensity value of the rhythm point to the rhythm amplitude of the virtual camera, and the formula can be as follows:

[0128] A = k × E(t) + b

[0129] Wherein, A is the rhythm amplitude of the virtual camera, E(t) is the intensity value of the time point t corresponding to the rhythm point (hereinafter referred to as the intensity value of the rhythm point), and k and b are the coefficients of the linear mapping function, which can be set or adjusted according to actual needs.

[0130] In addition, the embodiments of the present application can also use non-linear mapping functions, such as logarithmic functions or exponential functions, to better adapt to the characteristics of the audio signal.

[0131] In the embodiments of the present application, the rhythm amplitude of the virtual camera is updated in real time according to the intensity value of each rhythm point. When a new rhythm point appears in the audio signal, a new rhythm amplitude is calculated according to the intensity value of the rhythm point and applied to the motion control of the virtual camera.

[0132] As another implementable way, for each rhythm point, the following steps can be executed respectively: First, determine the attenuation intensity value of the previous rhythm point of this rhythm point. The attenuation intensity value is the difference between the intensity value of the previous rhythm point and the preset decline speed. Then, update the maximum value of the intensity value of this rhythm point and the attenuation intensity value of the previous rhythm point to the new intensity value of this rhythm point. The formula is as follows:

[0133] E new =max(E(t), E prev -V fall )

[0134] Where, E new is the new intensity value of this rhythm point, E(t) is the intensity value of this rhythm point, E prev is the intensity value of the previous rhythm point of this rhythm point, and V fall is the decline speed.

[0135] After normalization, the normalized new intensity value can be obtained. The formula can be as follows:

[0136]

[0137] Where, E′ new is the normalized new intensity value, and E max is the maximum intensity value of all rhythm points in the audio signal.

[0138] Finally, according to the normalized new intensity value, determine the rhythm amplitude of the virtual camera at the time point t corresponding to the rhythm point.

[0139] It should be noted that the attenuation intensity value of the previous rhythm point refers to the value obtained by adjusting the intensity value of the previous rhythm point through the preset decline speed, which reflects the influence degree of the previous rhythm point on the current rhythm point. By considering the attenuation intensity of the previous rhythm point, the incoherence of the virtual camera movement caused by the sudden change of the rhythm point intensity can be avoided. For example, when the intensity value of the previous rhythm point is relatively high, if the intensity value of the current rhythm point is directly used to determine the rhythm amplitude, it may cause a sudden change in the virtual camera movement. By considering the attenuation intensity, this sudden change can be smoothed, making the movement of the virtual camera smoother and more natural.

[0140] The decay speed refers to the time required for the intensity value of the previous rhythm point to decay to the intensity value of the current rhythm point. It reflects the influence speed of the previous rhythm point on the current rhythm point. By presetting the decay speed, the speed at which the intensity value of the previous rhythm point decays to the intensity value of the current rhythm point can be controlled. For example, when the intensity value of the previous rhythm point is relatively high, if the decay speed is slow, the intensity value of the previous rhythm point will decay to the intensity value of the current rhythm point relatively slowly, making the movement of the virtual camera smoother and more natural.

[0141] Therefore, in the embodiments of the present application, considering the decay intensity and decay speed of the previous rhythm point is to make the movement of the virtual camera smoother and more natural, and to avoid the incoherence of the movement of the virtual camera caused by the sudden change of the rhythm point intensity. By steps such as determining the new intensity value, normalization processing, and determining the rhythm amplitude, the movement of the virtual camera can be matched with the rhythm change of the audio signal, enhancing the real-time performance and interactivity of the live broadcast, enriching the visual effect of the live broadcast screen, and improving the overall expressiveness of the live broadcast and the viewing experience of users.

[0142] In addition to the above method of determining the rhythm amplitude based on the intensity value, the embodiments of the present application can also use other methods to determine the rhythm amplitude, such as predicting the rhythm amplitude based on the rhythm information in the audio signal and a pre-trained machine learning model.

[0143] As one of the achievable ways, the position information of the virtual camera can be embodied as the depth position of the virtual camera in the virtual scene. The so-called depth position refers to the distance between the virtual camera and the three-dimensional model when the virtual camera renders the three-dimensional model corresponding to the virtual scene to obtain the virtual scene, which is the depth of field reflected in the rendered scene.

[0144] In the embodiments of the present application, the depth position of the virtual camera in the virtual scene can be updated in real time according to the new intensity value of each rhythm point and the rhythm amplitude corresponding to the determined time point of the rhythm point. Specifically, the normalized new intensity value is used as a scaling factor to perform linear interpolation between the original depth position of the virtual camera and the enlarged depth position to obtain the depth position of the virtual camera in the virtual scene; among them, the enlarged depth position is determined based on the product of the original depth position of the virtual camera and the rhythm amplitude. The formula can be as follows:

[0145] `

[0146] Z(t) = lerp(Z init , Z init ×A, E new )

[0147] Wherein, Z(t) is the depth position of the virtual camera in the virtual scene at time point t, Z init is the original depth position of the virtual camera, Zinit ×A is the depth position after the virtual camera is magnified, and lerp() is a linear interpolation function.

[0148] In addition to using linear interpolation to determine the depth position of the virtual camera in the virtual scene as described above, the embodiments of the present application can also use non-linear interpolation functions, such as quadratic interpolation and cubic interpolation, to replace the linear interpolation function. The non-linear interpolation function can better handle the non-linear relationship of the data, making the change of the depth position of the virtual camera smoother and more natural.

[0149] It should also be noted that in addition to the camera movement data, the camera parameters, etc. can also be set. The camera parameters can include, for example, the lens aperture and exposure time, etc., which can be set or adjusted according to the needs.

[0150] As Figure 5 shown, it is a schematic diagram for determining the position information of the virtual camera provided by the embodiments of the present application.

[0151] First, when the spectrum analysis result (including the rhythm information of the audio signal) is received, the next step of processing is entered; otherwise, continue to wait for a new spectrum analysis result.

[0152] Extract the rhythm information of the audio signal from the spectrum analysis result and perform weighted smoothing processing to reduce the influence of noise and mutations, making the data smoother and more stable. Weighted smoothing can use some smoothing algorithms, such as the moving average method, the exponential smoothing method, etc.

[0153] Judge whether the intensity value of the rhythm point after weighted smoothing is greater than the decay intensity value of the previous rhythm point. When the intensity value of the rhythm point after weighted smoothing is greater than the decay intensity value of the previous rhythm point, directly determine the rhythm amplitude and depth position (i.e., the virtual camera position in the figure) of the virtual camera according to the intensity value of the rhythm point after weighted smoothing, etc.; if the intensity value of the rhythm point after weighted smoothing is not greater than the decay intensity value of the previous rhythm point, update the intensity value of the rhythm point after weighted smoothing to the decay intensity value of the previous rhythm point, and determine the rhythm amplitude and depth position of the virtual camera according to the updated intensity value, etc.

[0154] Finally, control the virtual camera based on the determined camera movement data.

[0155] In addition, as Figure 5 shown, the user can also interact with the system through the user interface to set some control options or parameters. For example, parameters such as the fallback speed and rhythm amplitude involved in the embodiments of the present application can also be set by the user, and the embodiments of the present application do not limit this.

[0156] Further, in order to increase the fineness of the camera movement trajectory of the virtual camera, the position information of the virtual camera may further include: among the time points included in the audio signal, the position information of the virtual camera at other time points except the time points corresponding to the rhythm points. Then, in this step, the original position of the virtual camera and the position information of the virtual camera at the time points corresponding to each rhythm point may also be used to perform trajectory fitting to obtain the position information of the virtual camera at other time points.

[0157] The above step 104, that is, "render the virtual scene according to the position information of the virtual camera, and synthesize the virtual scene with the live broadcast image collected in real time to generate a live broadcast video", will be described in detail in combination with the embodiments.

[0158] In the embodiment of the present application, the three-dimensional model corresponding to the virtual scene may be rendered using the position information of the virtual camera to obtain a virtual scene (presented as an image). The position and attitude of the virtual camera determine the viewing angle and composition of the image corresponding to the rendered virtual scene.

[0159] In the embodiment of the present application, according to the rhythm amplitude of the virtual camera, the movement range or amplitude of the virtual camera in the virtual scene can be controlled. For example, when the rhythm amplitude is large, the virtual camera can move significantly in the virtual scene to enhance the dynamic effect of the live broadcast image; when the rhythm amplitude is small, the movement range of the virtual camera is relatively small, making the live broadcast image more stable and delicate.

[0160] According to the depth position of the virtual camera, the front-back position of the virtual camera in the virtual scene can be controlled. For example, when the depth position is far, the virtual camera can be far away from the objects in the virtual scene, making the live broadcast image present a wider field of view; when the depth position is close, the virtual camera can be close to the objects in the virtual scene, making the live broadcast image more focused and clear.

[0161] It should be noted that the movement mode of the virtual camera in the embodiment of the present application may include, but is not limited to:

[0162] Translation movement: The virtual camera can perform translation movement in the virtual scene, including up-down translation, left-right translation, and front-back translation. The translation movement can change the viewing angle and range of the live broadcast image, enabling the audience to see different parts of the virtual scene.

[0163] Rotation movement: The virtual camera can perform rotation movement in the virtual scene, including horizontal rotation and vertical rotation. The rotation movement can change the direction and angle of the live broadcast image, enabling the audience to view the virtual scene from different angles.

[0164] Zoom motion: The virtual camera can perform zoom motion in the virtual scene, including zooming in and out. The zoom motion can change the size and details of the live broadcast image, enabling the audience to see the objects in the virtual scene more clearly.

[0165] As one of the achievable ways, the image area of the live object can be extracted from the real-time live broadcast image. For example, through keying processing, the image information of the live object is extracted and the background information is removed; the extracted image information of the live object is synthesized with the virtual scene at the corresponding time point to generate a live video. In addition to this implementation method, other methods can also be used, such as an image synthesis model based on a deep learning model to achieve the synthesis of the virtual scene and the live broadcast image.

[0166] Among them, the live object refers to the main target in the live broadcast image, and this main target can be a target of a specified type. For example, the host, items, etc. in the live broadcast image.

[0167] In this step, by controlling the virtual camera to move in the virtual scene according to the position information of the virtual camera and generating a live video containing the virtual scene, the real-time synchronization of the virtual camera movement and the audio rhythm is achieved. By controlling the movement mode and movement parameters of the virtual camera, the movement of the virtual camera can more accurately reflect the rhythm changes of the audio, enhancing the real-time and interactivity of the live broadcast, enriching the visual effect of the live broadcast image, and improving the overall expressiveness of the live broadcast and the viewing experience of users.

[0168] As Figure 6 shown, it is a schematic diagram of virtual-real synthesis provided by the embodiment of the present application. Among them, the rhythm control result, that is, the position information of the virtual camera, determines the movement mode of the virtual camera in the virtual scene. Figure 6 The camera in is used to collect the real-time live broadcast image and convert the real-time live broadcast image into an electrical signal, while the 4K capture card can digitally process the real-time live broadcast image collected by the camera and convert it into a digital signal to obtain a real-scene camera image. This real-scene camera image is the real-time collected live broadcast image in the live broadcast scene. The virtual scene, as an input, can include a pre-designed virtual environment and objects, usually presented as a 3D model.

[0169] Computer B controls the shooting trajectory of the virtual camera in the virtual scene according to the rhythm control result. The virtual scene image generated by the virtual camera is synthesized with the real-scene camera image. When synthesizing, the virtual scene image and the real-scene camera image are synthesized at the corresponding time points, and then a live video containing the virtual scene is generated and the audio-visual signal is output.

[0170] Through the above process, the real-time synchronization of the virtual camera movement and the audio rhythm is achieved, enhancing the real-time performance and interactivity of the live broadcast, enriching the visual effects of the live broadcast screen, and improving the overall expressiveness of the live broadcast and the viewing experience of users.

[0171] Further, referring to Figure 2 the "dual-camera shunt" architecture shown in, after the second terminal generates the live video, the live video can be transmitted to the first terminal, and then sent by the first terminal to the live server for display. That is to say, in the above "dual-camera shunt" architecture, the first terminal belongs to the push stream host, which is responsible for outputting the audio signal and pushing the received live video to the live server; while the second terminal belongs to the rhythm host, which is responsible for processing the real-time collected audio signal, controlling the virtual camera movement, and synthesizing with the real-time live broadcast screen, that is, merging the audio and video, and generating a merged signal to be transmitted back to the first terminal.

[0172] The specific embodiments of this specification have been described above. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than in the embodiments and still achieve the desired results. Additionally, the processes depicted in the figures do not necessarily require the specific order or sequential order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0173] According to an embodiment of another aspect, a video generation device is provided. Figure 7 The schematic block diagram of the video generation device according to an embodiment is shown. As Figure 7 shown, the device 700 includes: an audio acquisition unit 701, a spectrum analysis unit 702, a trajectory determination unit 703, and a video synthesis unit 704. The main functions of each component unit are as follows:

[0174] The audio acquisition unit 701 is configured to acquire the audio signal collected in real time during the live broadcast.

[0175] The spectrum analysis unit 702 is configured to perform spectrum analysis on the audio signal to determine the rhythm information in the audio signal, and the rhythm information includes the time points corresponding to the rhythm points.

[0176] The trajectory determination unit 703 is configured to determine the position information of the virtual camera according to the rhythm information in the audio signal, and the position information at least includes the position information of the virtual camera at the time points corresponding to the rhythm points;

[0177] The video synthesis unit 704 is configured to render the virtual scene according to the position information of the virtual camera; synthesize the virtual scene with the live broadcast screen collected in real time to generate the live video.

[0178] As one of the achievable ways, when the spectrum analysis unit 702 performs spectrum analysis on the audio signal to determine the rhythm information in the audio signal, it can be specifically configured as follows:

[0179] The audio signal is divided into a plurality of time windows according to the time sequence, each time window corresponds to a sub-audio signal in the audio signal;

[0180] Performing short-time Fourier transform on the sub-audio signal in each time window to obtain frequency spectrum information of each sub-audio signal;

[0181] According to the spectrum information of each sub-audio signal, the rhythm information within the target frequency range is extracted.

[0182] As one of the implementable ways, the rhythm information may also include the intensity value of the rhythm point; then when the spectrum analysis unit 702 extracts the rhythm information within the target frequency range according to the spectrum information of each sub-audio signal, it may be specifically configured as follows:

[0183] According to the spectrum information of each sub-audio signal, the energy envelope value of each time window within the target frequency range is calculated;

[0184] When the energy envelope value exceeds a preset energy threshold, the time point corresponding to the time window is determined as the time point corresponding to the rhythm point, and the energy envelope value is determined as the intensity value of the rhythm point.

[0185] As one possible implementation, the trajectory determination unit 703 may be specifically configured as follows:

[0186] According to the intensity value of the rhythm point, the rhythm amplitude of the virtual camera at the time point corresponding to the rhythm point is determined;

[0187] According to the rhythm amplitude of the virtual camera at the time point corresponding to the rhythm point, the position information of the virtual camera at the time point corresponding to the rhythm point is determined.

[0188] As one of the achievable ways, when the trajectory determination unit 703 determines the rhythm amplitude of the virtual camera at the time point corresponding to the rhythm point according to the intensity value of the rhythm point, it can be specifically configured as follows:

[0189] Determine the decay intensity value of the previous rhythm point; wherein the decay intensity value is the difference between the intensity value of the previous rhythm point and the preset fall-back speed;

[0190] The maximum value of the intensity value of the rhythm point and the decay intensity value of the previous rhythm point is updated as the new intensity value of the rhythm point;

[0191] Normalizing the new intensity value of the rhythm point to obtain a normalized new intensity value;

[0192] Determine the rhythm amplitude of the virtual camera at the time point corresponding to the rhythm point according to the normalized new intensity value.

[0193] As one possible implementation, when determining the rhythm amplitude of the virtual camera at the time point corresponding to the rhythm point according to the intensity value of the rhythm point, the trajectory determination unit 703 can be specifically configured as follows:

[0194] Map the intensity value of the rhythm point according to the preset mapping relationship to obtain the rhythm amplitude of the virtual camera at the time point corresponding to the rhythm point.

[0195] As one possible implementation, when determining the position information of the virtual camera at the time point corresponding to the rhythm point according to the rhythm amplitude of the virtual camera at the time point corresponding to the rhythm point, the trajectory determination unit 703 can be specifically configured as follows:

[0196] Determine the position information of the virtual camera at the time point corresponding to the rhythm point according to the normalized new intensity value and the rhythm amplitude of the virtual camera at the time point corresponding to the rhythm point, where the position information includes the depth position of the virtual camera in the virtual scene.

[0197] As one possible implementation, when determining the position information of the virtual camera at the time point corresponding to the rhythm point according to the normalized new intensity value and the rhythm amplitude of the virtual camera at the time point corresponding to the rhythm point, the trajectory determination unit 703 can be specifically configured as follows:

[0198] Use the normalized new intensity value as a scaling factor to perform linear interpolation between the original depth position and the enlarged depth position of the virtual camera to obtain the depth position of the virtual camera at the time point corresponding to the rhythm point; wherein, the enlarged depth position is determined based on the product of the original depth position and the rhythm amplitude of the virtual camera.

[0199] Furthermore, the position information of the virtual camera further includes: the position information of the virtual camera at other time points except the time point corresponding to the rhythm point among the time points included in the audio signal;

[0200] The trajectory determination unit 703 can also be further configured to: perform trajectory fitting using the original position of the virtual camera and the position information of the virtual camera at the time points corresponding to each rhythm point to obtain the position information of the virtual camera at other time points.

[0201] As one possible implementation, when synthesizing the virtual scene with the live video captured in real time to generate a live video, the video synthesis unit 704 can be specifically configured as follows:

[0202] Extract the image area of the live object from the live video;

[0203] The image region of the live object to be extracted is synthesized with the virtual scene at the corresponding time points to generate a live video.

[0204] According to an embodiment of another aspect, a live broadcast system is provided. Figure 8 FIG. shows a schematic block diagram of the live broadcast system according to an embodiment. The live broadcast system 800 includes a first terminal 801, a second terminal 802, and a sound collection device 803, where:

[0205] The first terminal 801 is configured to play audio during the live broadcast.

[0206] The sound collection device 803 is configured to collect audio signals during the live broadcast in real time and send them to the second terminal 802.

[0207] The second terminal 802 is configured to perform spectral analysis on the audio signal to determine the rhythm information in the audio signal. The rhythm information includes the time points corresponding to the rhythm points; according to the rhythm information in the audio signal, determine the position information of the virtual camera. The position information at least includes the position information of the virtual camera at the time points corresponding to the rhythm points; according to the position information of the virtual camera, render the virtual scene; synthesize the virtual scene with the live broadcast images collected in real time to generate a live video; transmit the live video to the first terminal 801;

[0208] The first terminal 801 is further configured to receive the live video and transmit it to the live broadcast server.

[0209] The above-mentioned sound collection device 803 can be an independent device or a device in the second terminal 802, such as a sound collection card in the second terminal 802, etc.

[0210] Each embodiment in this specification is described in a progressive manner. The same or similar parts among the embodiments can be referred to each other. Each embodiment focuses on the differences from other embodiments. In particular, for system or device embodiments, since they are basically similar to method embodiments, the description is relatively simple. For related parts, refer to the description of the method embodiments. The system and device embodiments described above are only illustrative. The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.

[0211] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. Moreover, the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions, and corresponding operation entrances are provided for users to choose to authorize or reject.

[0212] In addition, an embodiment of the present application further provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the steps of the method described in any one of the foregoing method embodiments.

[0213] And an electronic device, including:

[0214] One or more processors; and

[0215] A memory associated with the one or more processors, the memory is used to store program instructions, and when the program instructions are read and executed by the one or more processors, the steps of the method described in any one of the foregoing method embodiments are executed.

[0216] The present application also provides a computer program product, including a computer program, which implements the steps of the method described in any one of the foregoing method embodiments when executed by a processor.

[0217] Wherein, Figure 9 An exemplary architecture of the electronic device is shown, which may specifically include a processor 910, a video display adapter 911, a disk drive 912, an input / output interface 913, a network interface 914, and a memory 920. The above-mentioned processor 910, video display adapter 911, disk drive 912, input / output interface 913, network interface 914 can be communicatively connected to the memory 920 through a communication bus 930.

[0218] Wherein, the processor 910 can be implemented in ways such as a general-purpose CPU, a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided by the present application.

[0219] The memory 920 can be implemented in the form of a ROM (Read Only Memory), a RAM (Random Access Memory), a static storage device, a dynamic storage device, etc. The memory 920 can store an operating system 921 for controlling the operation of the electronic device, and a Basic Input / Output System (BIOS) 922 for controlling the low-level operations of the electronic device. Additionally, a web browser 923, a data storage management system 924, a video generation device 700, etc. can also be stored. The above-mentioned video generation device 700 can be the application program that specifically implements the operations of the foregoing steps in the embodiments of the present application. In summary, when implementing the technical solution provided by the present application through software or firmware, the relevant program codes are stored in the memory 920 and are called and executed by the processor 910.

[0220] The input / output interface 913 is used to connect to the input / output module to implement information input and output. The input / output module can be configured as a component in the device (not shown in the figure) or can be externally connected to the device to provide corresponding functions. Among them, the input device can include a keyboard, a mouse, a touch screen, a microphone, various sensors, etc., and the output device can include a display, a speaker, a vibrator, an indicator light, etc.

[0221] The network interface 914 is used to connect to a communication module (not shown in the figure) to implement communication interaction between this device and other devices. Among them, the communication module can implement communication in a wired manner (such as USB, network cable, etc.) or in a wireless manner (such as mobile network, WIFI, Bluetooth, etc.).

[0222] The bus 930 includes a path for transmitting information between various components of the device (such as the processor 910, the video display adapter 911, the disk drive 912, the input / output interface 913, the network interface 914, and the memory 920).

[0223] It should be noted that although the above device only shows the processor 910, the video display adapter 911, the disk drive 912, the input / output interface 913, the network interface 914, the memory 920, the bus 930, etc., in the specific implementation process, the device may also include other components necessary for normal operation. In addition, those skilled in the art can understand that the above device may also only include the components necessary to implement the solution of the present application and does not necessarily include all the components shown in the figure.

[0224] As can be seen from the description of the above embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solution of this application, in essence or the part that contributes to the prior art, can be embodied in the form of a computer program product, which can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., including several instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods described in various embodiments or some parts of the embodiments of this application.

[0225] The technical solutions provided in this application have been introduced in detail above. Specific examples are used in this article to elaborate on the principles and implementation manners of this application. The description of the above embodiments is only used to help understand the method and its core idea of this application; at the same time, for those of ordinary skill in the art, according to the idea of this application, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to this application.

Claims

1. A video generation method, characterized in that: The method comprises: Get the audio signal collected in real time during the live broadcast; Performing spectrum analysis on the audio signal to determine rhythm information in the audio signal, wherein the rhythm information includes time points corresponding to rhythm points; Determine position information of a virtual camera according to rhythm information in the audio signal, wherein the position information at least includes position information of the virtual camera at a time point corresponding to the rhythm point; Rendering a virtual scene according to the position information of the virtual camera; The virtual scene is synthesized with the live picture collected in real time to generate a live video.

2. The method according to claim 1, characterized in that The performing spectrum analysis on the audio signal to determine the rhythm information in the audio signal includes: Dividing the audio signal into a plurality of time windows according to time sequence, each time window corresponding to a sub-audio signal in the audio signal; Performing short-time Fourier transform on the sub-audio signal in each time window to obtain frequency spectrum information of each sub-audio signal; According to the frequency spectrum information of each sub-audio signal, the rhythm information within the target frequency range is extracted.

3. The method according to claim 2, characterized in that The rhythm information also includes the intensity value of the rhythm point; The step of extracting rhythm information within a target frequency range according to the frequency spectrum information of each sub-audio signal comprises: Calculating the energy envelope value of each time window within the target frequency range according to the frequency spectrum information of each sub-audio signal; When the energy envelope value exceeds a preset energy threshold, the time point corresponding to the time window is determined as the time point corresponding to the rhythm point, and the energy envelope value is determined as the intensity value of the rhythm point.

4. The method according to claim 1, characterized in that The rhythm information also includes the intensity value of the rhythm point; Determining the position information of the virtual camera according to the rhythm information in the audio signal includes: Determining the rhythm amplitude of the virtual camera at the time point corresponding to the rhythm point according to the intensity value of the rhythm point; According to the rhythm amplitude of the virtual camera at the time point corresponding to the rhythm point, the position information of the virtual camera at the time point corresponding to the rhythm point is determined.

5. The method according to claim 4, characterized in that Determining the rhythm amplitude of the virtual camera at the time point corresponding to the rhythm point according to the intensity value of the rhythm point includes: Determine the decay intensity value of the previous rhythm point of the rhythm point; wherein the decay intensity value is the difference between the intensity value of the previous rhythm point and a preset fall-back speed; Updating the maximum value of the intensity value of the rhythm point and the attenuation intensity value of the previous rhythm point as the new intensity value of the rhythm point; Normalizing the new intensity value of the rhythm point to obtain a normalized new intensity value; The rhythm amplitude of the virtual camera at the time point corresponding to the rhythm point is determined according to the normalized new intensity value.

6. The method according to claim 4, characterized in that Determining the rhythm amplitude of the virtual camera at the time point corresponding to the rhythm point according to the intensity value of the rhythm point includes: According to a preset mapping relationship, the intensity value of the rhythm point is mapped to obtain the rhythm amplitude of the virtual camera at the time point corresponding to the rhythm point.

7. The method according to claim 5, characterized in that Determining the position information of the virtual camera at the time point corresponding to the rhythm point according to the rhythm amplitude of the virtual camera at the time point corresponding to the rhythm point includes: According to the normalized new intensity value and the rhythm amplitude at the time point corresponding to the rhythm point, the position information of the virtual camera at the time point corresponding to the rhythm point is determined, and the position information includes the depth position of the virtual camera in the virtual scene.

8. The method according to claim 7, characterized in that Determining the position information of the virtual camera at the time point corresponding to the rhythm point according to the normalized new intensity value and the rhythm amplitude at the time point corresponding to the rhythm point includes: The normalized new intensity value is used as a scaling factor to perform linear interpolation between the original depth position of the virtual camera and the amplified depth position to obtain the depth position of the virtual camera at the time point corresponding to the rhythm point; wherein the amplified depth position is determined based on the product of the original depth position of the virtual camera and the rhythm amplitude.

9. The method according to claim 1, 7 or 8, characterized in that: The position information of the virtual camera also includes: position information of the virtual camera at other time points other than the time points corresponding to the rhythm points among the time points included in the audio signal; The method further includes: performing trajectory fitting using the original position of the virtual camera and the position information of the virtual camera at the time points corresponding to each rhythm point, so as to obtain the position information of the virtual camera at the other time points.

10. The method according to claim 1, characterized in that The virtual scene is synthesized with the live picture collected in real time to generate the live video, including: Extracting an image area of ​​a live broadcast object from the live broadcast picture; The extracted image area of ​​the live object is synthesized with the virtual scene at corresponding time points to generate the live video.

11. A video generating device, characterized in that: The device comprises: An audio acquisition unit, configured to acquire an audio signal collected in real time during a live broadcast; A spectrum analysis unit, configured to perform spectrum analysis on the audio signal to determine rhythm information in the audio signal, wherein the rhythm information includes a time point corresponding to a rhythm point; A trajectory determination unit is configured to determine position information of a virtual camera according to rhythm information in the audio signal, wherein the position information at least includes position information of the virtual camera at a time point corresponding to the rhythm point; The video synthesis unit is configured to render a virtual scene according to the position information of the virtual camera; synthesize the virtual scene with the live picture collected in real time to generate a live video.

12. A live broadcast system, characterized in that: The live broadcast system includes a first terminal, a sound collection device, and a second terminal, wherein: The first terminal is configured to play audio during the live broadcast; The sound collection device is configured to collect audio signals in real time during the live broadcast and send the audio signals to the second terminal; The second terminal is configured to perform spectrum analysis on the audio signal to determine rhythm information in the audio signal, wherein the rhythm information includes a time point corresponding to a rhythm point; determine position information of a virtual camera according to the rhythm information in the audio signal, wherein the position information includes at least position information of the virtual camera at a time point corresponding to the rhythm point; render a virtual scene according to the position information of the virtual camera; synthesize the virtual scene with a live broadcast picture collected in real time to generate a live video; and transmit the live video to the first terminal; The first terminal is also configured to receive the live video and transmit it to the live server.

13. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the method according to any one of claims 1 to 10 are implemented.

14. An electronic device, characterized in that: include: one or more processors; as well as A memory associated with the one or more processors, the memory being used to store program instructions, wherein the program instructions, when read and executed by the one or more processors, execute the steps of the method described in any one of claims 1 to 10.

15. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 10 are implemented.