Video generation method and apparatus, device, medium, and product
By generating special effects trajectories based on audio information in video footage, the problem of insufficient interactive audio content is solved, resulting in richer and more natural video footage and enhancing the user's interactive experience.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- BEIJING ZITIAO NETWORK TECH CO LTD
- Filing Date
- 2025-09-09
- Publication Date
- 2026-06-04
Smart Images

Figure CN2025120131_04062026_PF_FP_ABST
Abstract
Description
Video generation methods, apparatus, equipment, media and products
[0001] Cross-reference to related applications
[0002] This application claims priority to Chinese Patent Application No. 202411269947.2, filed on September 10, 2024, the disclosure of which is incorporated herein by reference in its entirety. Technical Field
[0003] This disclosure relates to a video generation method, apparatus, device, medium, and product. Background Technology
[0004] With the increasing popularity of intelligent devices, users are demanding more and more interaction with smart terminals.
[0005] Currently, users can interact with smart terminals based on audio information. However, this mainly involves analyzing and processing the audio information to determine the corresponding interactive content. In other words, it primarily relies on the content of the audio information itself to achieve interaction, without combining the richer content of the audio information to generate interactive content, which limits the use of audio information. Summary of the Invention
[0006] This disclosure provides a video generation method, apparatus, device, medium, and product.
[0007] In a first aspect, embodiments of this disclosure provide a video generation method, the method comprising:
[0008] During the video display, the first trajectory is determined based on the collected audio information;
[0009] Based on the first trajectory and a pre-set first source image, a first special effect trajectory associated with the audio information is determined; wherein, the display information corresponding to the first special effect trajectory is related to the pixel information of the first source image;
[0010] Based on the target mask image and the second source image associated with the first trajectory, a second special effect trajectory of the first trajectory is determined;
[0011] Based on the first special effects trajectory and the second special effects trajectory, the target trajectory displayed in the video frame is determined and displayed; wherein, the change information of the target trajectory is associated with the audio information.
[0012] Secondly, embodiments of this disclosure also provide a video generation apparatus, the apparatus comprising:
[0013] The first trajectory determination module is used to determine the first trajectory based on the collected audio information during the video display process;
[0014] The first special effects trajectory generation module is used to determine a first special effects trajectory associated with the audio information based on the first trajectory and a pre-set first material image; wherein, the display information corresponding to the first special effects trajectory is related to the pixel information of the first material image;
[0015] The second special effects trajectory generation module is used to determine the second special effects trajectory of the first trajectory based on the target mask image and the second material image associated with the first trajectory;
[0016] The target trajectory generation module is used to determine and display the target trajectory shown in the video frame based on the first special effect trajectory and the second special effect trajectory; wherein the change information of the target trajectory is associated with the audio information.
[0017] Thirdly, embodiments of this disclosure also provide an electronic device, the electronic device comprising:
[0018] One or more processors;
[0019] Storage device for storing one or more programs.
[0020] When the one or more programs are executed by the one or more processors, the one or more processors implement the video generation method as described in any of the embodiments of this disclosure.
[0021] Fourthly, embodiments of this disclosure also provide a storage medium containing computer-executable instructions, which, when executed by a computer processor, are used to perform the video generation method as described in any of the embodiments of this disclosure.
[0022] Fifthly, embodiments of this disclosure also provide a computer program product, including a computer program that, when executed by a processor, implements the video generation method as described in any of the embodiments of this disclosure. Attached Figure Description
[0023] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and elements are not necessarily drawn to scale.
[0024] Figure 1 is a schematic flowchart of a video generation method provided in an embodiment of this disclosure;
[0025] Figure 2 is a flowchart illustrating a video generation method provided in an embodiment of this disclosure;
[0026] Figure 3 is a schematic flowchart of the process for determining trajectory control points provided in an embodiment of this disclosure;
[0027] Figure 4 is a flowchart illustrating the first trajectory provided in an embodiment of this disclosure;
[0028] Figure 5 is a flowchart illustrating a video generation method provided in an embodiment of this disclosure;
[0029] Figure 6 is a first material image provided in an embodiment of this disclosure;
[0030] Figure 7 is a schematic diagram of the first special effect trajectory provided in the embodiments of this disclosure;
[0031] Figure 8 is a schematic diagram of the first special effect trajectory after Gaussian blur processing provided in the embodiments of this disclosure;
[0032] Figure 9 is a schematic diagram of the generated target mask provided in an embodiment of this disclosure;
[0033] Figure 10 is a target mask image after Gaussian blurring provided in an embodiment of this disclosure;
[0034] Figure 11 is a schematic diagram of the second material diagram provided in the embodiments of this disclosure;
[0035] Figure 12 is a schematic diagram of the second special effect trajectory provided in the embodiments of this disclosure;
[0036] Figure 13 is a schematic diagram of the trajectory to be fused provided in an embodiment of this disclosure;
[0037] Figure 14 is a schematic diagram of the target trajectory provided in the embodiments of this disclosure;
[0038] Figure 15 is a schematic diagram of a video generation device provided in an embodiment of this disclosure; and
[0039] Figure 16 is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. Detailed Implementation
[0040] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.
[0041] It should be understood that the steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect.
[0042] The term "comprising" and its variations as used herein are open-ended inclusions, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the description below.
[0043] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.
[0044] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".
[0045] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.
[0046] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.
[0047] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as the electronic device, application, server, or storage medium performing the operations of this disclosed technical solution, based on the prompt message.
[0048] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0049] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.
[0050] It is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and related provisions.
[0051] Before introducing the technical solutions provided by the embodiments of this disclosure, the application scenarios can be illustrated by example. Users can use application software to shoot videos, make video calls with other users, or integrate the solutions provided by the embodiments of this disclosure as a game function into applications. In these cases, the solutions can be used to generate target trajectories, which are the final special effects trajectories displayed.
[0052] In practical applications, if you want to generate a target trajectory corresponding to the audio information, you can refer to the solution provided in the embodiments of this disclosure.
[0053] [Corrected according to Rule 91, October 29, 2025] This can be understood as follows: The aforementioned application software can be a type of software for image / video processing. Specific application software will not be detailed here, but it mainly needs to achieve image / video processing. Of course, forwarding applications can also be developed to add and display special effects, or integrated into corresponding pages, allowing users to process special effects videos through integrated pages on PCs.
[0054] For example, if the technical solution provided in this disclosure is integrated into an application software capable of recording video or making video calls, the solution can be integrated into a special effects prop and installed in the corresponding application software. When a user triggers the special effects prop, audio information can be collected and the target trajectory displayed in the video can be generated based on the audio information. If the solution provided in this disclosure is integrated as a function into any supported game application, it can be displayed as a functional control on the main interface or a sub-interface of the application. When a trigger is detected, the game map corresponding to the game can be retrieved, and the target trajectory can be generated based on the game map and the collected audio information. Alternatively, when certain conditions are met during gameplay, the solution provided in this disclosure can be automatically invoked to form the target trajectory.
[0055] The specific implementation method for generating the target trajectory can be found in the detailed description of the following embodiments.
[0056] Figure 1 is a schematic flowchart of a video generation method provided in an embodiment of this disclosure. This embodiment is applicable to scenarios where a target trajectory is generated in the currently displayed video frame. The video generation method provided in this embodiment can be executed by a client, executed by a server, or executed jointly by the client and the server. Of course, a video generation device corresponding to the video generation method can be integrated into the client and / or the server.
[0057] As shown in Figure 1, the method of this embodiment may specifically include:
[0058] S110. During the video display process, the first trajectory is determined based on the collected audio information.
[0059] The video footage can be captured after a special effects item is triggered, or it can be displayed after a game function control is triggered. In other words, the video footage is the footage displayed after the special effects item or the game function provided in this embodiment of the disclosure is triggered.
[0060] It should be noted that the video footage can include a variety of background content. The background content can be pre-set content, which is optional. The background content can be flowers, trees, or display elements corresponding to a certain festival. Optional festivals can be the Mid-Autumn Festival, and display elements can be rabbit elements, moon elements, etc.
[0061] During the video display, audio information can be captured in real time or periodically. Alternatively, it can be captured as soon as audio is detected. This audio information primarily originates from objects outside the video frame. However, upon detection, the lip movements of objects within the video frame can be matched to the captured audio lip movements to simulate the effect of the audio being emitted by objects within the video frame, thereby enhancing the realism and visual appeal of the video.
[0062] After acquiring the audio information, it can be analyzed and processed to obtain the first trajectory corresponding to the audio information. At this stage, the first trajectory has not yet undergone further special effects rendering. The length of the first trajectory and its display height in the video frame are related to the duration and pitch of the audio information.
[0063] It should be noted that the first trajectory is not displayed in the video at this time. The first trajectory can be processed again to obtain a cool target trajectory and display it in the video.
[0064] Specifically, after detecting a triggered special effects item, a video screen can be displayed. During the video screen display, the microphone array of the terminal device to which the application belongs can be activated to collect audio information based on the microphone array. Alternatively, if a game function control integrated into the application corresponding to the embodiments of this disclosure is triggered, a pre-set game screen (video screen) can be displayed. Users can play the game based on the video screen and obtain audio information of objects outside the video screen in real time during the game (displaying the video screen). Of course, it can also be during live video streaming or video calls, if a function control corresponding to the solution provided in the embodiments of this disclosure is triggered, audio information of objects on the terminal to which the application belongs can also be collected. After the audio information is collected, it can be analyzed and processed to obtain a first trajectory.
[0065] In this embodiment, during the display of the video screen, the method further includes: when the trajectory generation conditions are detected, collecting the audio information of the first object, wherein the trajectory generation conditions include: determining that the second object moves to a preset position based on a preset map, wherein the second object is a movable object displayed in the video screen; and triggering the trajectory generation control in the video screen or the application software to which the video screen belongs.
[0066] The content displayed in the video can be content designed during the design phase, content captured by the camera, or a composite of the designed content and the captured content. The user corresponding to the terminal device of the application software displaying the video can be considered the first object. For example, if "I" trigger the function control in the application to generate the target trajectory, "I" am the first object. In other words, the first object is an object in the real environment emitting audio information that can be captured by the terminal device of the application. The second object is a movable object displayed in the video. The second object can be a pre-defined object or a virtual object based on the image of the first object.
[0067] In this embodiment, if it is a game scene, a game map can be created. The game map can be used as a preset map. The game map can include visual elements such as mountains, streams, and ravines that the second object needs to traverse. When these elements appear, a target trajectory may need to be generated to traverse the elements in the visual scene based on the target trajectory, achieving the effect of moving from one point to another. Simultaneously, the coordinate information of these elements can be marked on the game map. When the second object moves to the preset position corresponding to the aforementioned element at a preset movement speed, the audio information of the first object can be collected. That is, the preset position is the starting position where the aforementioned element appears.
[0068] In non-game scenarios, such as live streaming, generating special effects videos, or video calls, a trajectory generation control can be displayed in the video frame, or the application software containing the video can display the trajectory generation control. When the aforementioned trajectory generation control is triggered, audio information of the first object can be collected. Alternatively, audio information can be continuously collected, and when there is a need for trajectory generation, the audio information collected at the current moment and after the current moment can be analyzed and processed to obtain the first trajectory.
[0069] The above can be understood as follows: if the trajectory generation conditions are met, the audio information of the first object can be collected.
[0070] S120. Based on the first trajectory and the pre-set first material image, determine the first special effects trajectory associated with the audio information.
[0071] To improve the image quality, the obtained first trajectory can be further processed to obtain the first special effects trajectory. That is, the first special effects trajectory is the trajectory obtained after rendering the special effects of the first trajectory.
[0072] The first source image is used to adjust the display information of each pixel in the first trajectory. A background image with a color and / or transparency gradient can be created as the first source image. The size of this first source image matches the size of the video frame. After obtaining the first trajectory, the first source image can be traversed according to the x-coordinate of each point on the first trajectory to obtain the display information of each trajectory point. The first trajectory with updated display information is then used as the first special effects trajectory.
[0073] This can be understood as follows: A first source image with varying colors and / or transparency can be created based on actual needs. The pixel information of this first source image is traversed based on the x-coordinate of each point in the first trajectory, determining and updating the pixel information of each point in the first trajectory to obtain the first special effects trajectory.
[0074] S130. Based on the target mask image and the second source image associated with the first trajectory, determine the second special effect trajectory of the first trajectory.
[0075] This process involves creating an initial trajectory corresponding to the first trajectory, and then dilating the initial mask image to obtain a target mask image. This can be understood as the target mask image having a shape related to the first trajectory, but a width greater than the first trajectory. The second source image is a pre-defined image matching the target effect. This second source image can be an image simulating a certain phenomenon; optionally, the phenomenon could be an aurora, meaning the second source image is an image simulating an aurora. The second effect trajectory is the trajectory obtained by processing the mask region in the target mask image associated with the first trajectory.
[0076] Specifically, after obtaining the target mask image, the display information of the pixels corresponding to the masked areas in the target mask image can be adjusted based on the pixel values of each pixel in the second source image. After filling the masked areas of the target mask image with pixels, the masked areas can be extracted to obtain the second effect trajectory.
[0077] S140. Based on the first special effects trajectory and the second special effects trajectory, determine and display the target trajectory displayed in the video frame; wherein, the trajectory length of the target trajectory is associated with the audio information.
[0078] This can be understood as follows: after obtaining the first and second special effects trajectories, the first and second special effects trajectories can be merged together to obtain the target trajectory. The target trajectory is then displayed in the video frame.
[0079] It should also be noted that the starting point for displaying the target trajectory can be pre-set or located at a position related to an element in the video frame. Optionally, if the element is a stream, the starting point can be the bank of the stream to which the second object currently resides; if the element is a ravine, the starting point can be any position on one side of the ravine to which the second object currently resides. The final endpoint for displaying the target trajectory (i.e., the pre-set constraint position) can also be pre-set, optionally at the right edge of the video frame, or on the ground on the other side of a mountain peak, stream, or ravine. However, the real-time display position of the trajectory points in the video frame is dynamic, and its change is primarily determined by the pitch of the audio information. This can be understood as the change information of the target trajectory being related to the audio information.
[0080] The technical solution provided in this disclosure analyzes and processes the collected audio information during video display to determine a first trajectory. Then, based on the first trajectory and a first source image, effects are rendered onto the first trajectory to obtain a first special effects trajectory. To further improve the display effect, the target mask image and the second source image corresponding to the first trajectory can be processed again to obtain a second special effects trajectory. Based on the first and second special effects trajectories, a target trajectory corresponding to the audio information can be determined and displayed in the video, achieving a richer and more natural visual effect.
[0081] Figure 2 is a schematic flowchart of a video generation method provided in an embodiment of this disclosure. Based on the foregoing embodiments, the step of "determining the first trajectory based on the collected audio information" can be further refined. For specific implementation methods, please refer to the detailed description of this technical solution. Technical terms that are the same as or corresponding to those in the above embodiments will not be repeated in this embodiment.
[0082] As shown in Figure 2, the method includes:
[0083] S210. Determine at least one trajectory control point based on the pitch information of the valid audio frames in the audio information.
[0084] The audio information consists of multiple audio frames, optionally 25 audio frames per second. Some audio frames may be invalid; therefore, a validity analysis can be performed on the audio frames to determine the first trajectory based on the valid audio frames. Trajectory control points are used to adjust the direction of the first trajectory; optionally, the direction can be understood as the shape of the trajectory.
[0085] Specifically, audio information can be analyzed and processed to determine valid audio frames, and then trajectory control points can be determined based on the valid audio frames.
[0086] In this embodiment, the method for determining a valid audio frame in the audio information can be as follows: after extracting the audio frame from the audio information, the volume and pitch of the audio frame can be obtained. If the volume is greater than the volume in the first volume condition and the pitch meets the pitch in the first pitch condition, the audio frame can be regarded as a valid audio frame.
[0087] For example, the first volume condition is a volume greater than 0.002, and the first pitch condition is a pitch greater than 0. If the volume in the audio frame is greater than 0.002 and the pitch is greater than 0, it means that the audio frame is a valid audio frame.
[0088] In this embodiment, determining at least one trajectory control point based on the pitch information of the valid audio frames in the audio information can be as follows: determining the pitch offset based on the pitch information and pitch control information of the valid audio frames in the audio information, and adjusting the pitch information of the valid audio frames based on the pitch offset to obtain the pitch information to be used; determining the target screen coordinates of the pitch information to be used in the video frame, and determining the trajectory control point based on the target screen coordinates and the historical screen coordinates in the coordinate set.
[0089] Typically, when the first object is of type one (optionally, a male with a lower pitch), the first trajectory's display height in the video frame will be close to the bottom of the screen. Conversely, when the first object is of type two (optionally, a female with a higher pitch), the first trajectory's display height in the video frame will be close to the top of the screen. In other words, the first type is a lower-pitched object type, and the second type is a higher-pitched object type. Before use, it's possible to pre-define which objects are type one and which are type two, and determine the pitch control information accordingly. To avoid poor first trajectory performance due to the aforementioned object types, pitch control information can be generated based on the first pitch corresponding to the first type and the second pitch corresponding to the second type. This allows for adjustment of the pitch of any object type, thereby improving the harmony of the target trajectory within the video frame.
[0090] In this embodiment, the tone control information is determined based on a first tone of a first type and a second tone of a second type. Optionally, the average of the first tone and the second tone can be calculated and used as the tone control information. The advantage of this setting is that it can neutralize the influence of high and low tones on the first trajectory. That is, the tone control information is used to control the tone information to fluctuate within a preset range.
[0091] The pitch offset is the offset information obtained after constraining the sound effect information of valid audio frames based on pitch control information. The pitch information of valid audio frames can be adjusted based on this pitch offset; essentially, all valid audio frames can be normalized based on the pitch offset. The pitch information adjusted based on the pitch offset is then used as the pitch information to be used.
[0092] The target screen coordinates are the coordinates corresponding to the tonal information to be used after mapping it onto the video screen. Based on the target screen coordinates and the historical screen coordinates determined in the coordinate set, trajectory control points are determined to adjust the direction of the Bézier curve. These trajectory control points are mainly used to adjust the trajectory direction of the first trajectory. Optionally, the trajectory direction may include the extension angle of the first trajectory, which can be an angle corresponding to the ground in the video screen.
[0093] The target screen coordinates of the valid audio processed before the current valid audio frame are used as the historical screen coordinates. That is, the coordinate set mainly stores the target screen coordinates of the valid audio frames before the current valid audio frame.
[0094] It should be noted that the solution provided in this embodiment can be processed with the first valid audio frame. Of course, in order to improve the processing accuracy and the transition effect of the picture, the analysis and processing of the determined valid audio frames can only begin when the number of valid audio frames reaches the first threshold, so as to start generating the trajectory.
[0095] Specifically, the pitch offset is first determined based on the pitch information and pitch control information of the valid audio frames in the audio information. Then, the pitch information of the valid audio is adjusted according to this pitch offset to obtain the pitch information to be used. Finally, the trajectory control points can be determined based on the target screen coordinates corresponding to the pitch information to be used in the video frame.
[0096] It should be noted that the processing method is the same for each valid audio frame; this explanation mainly focuses on analyzing and processing a single valid audio frame. In other words, the pitch information to be used for each valid audio frame is dynamically determined.
[0097] For example, the first type is the male type, and the first pitch is the lowest male pitch (lowPitch). The second type is the female type, and the second pitch is the highest female pitch (highPitch). The pitch control information can be calculated by averaging the first and second pitches, midPitch = (lowPitch + highPitch) / 2, where midPitch is the pitch control information.
[0098] It should also be noted that the tone control information can also be determined based on the preset tone weights of the first type and the first tone, the tone weights of the second type and the second tone.
[0099] S220. Determine the first trajectory based on at least one trajectory control point and a preset Bezier function.
[0100] The preset Bézier function can be a function that generates Bézier curves, and correspondingly, the first trajectory corresponds to the generated Bézier curve.
[0101] Specifically, the target screen coordinates corresponding to at least one trajectory control point can be substituted into the Bézier function to obtain a Bézier function with determined coefficients. Next, the independent variable can be adjusted to generate multiple trajectory points; connecting all these trajectory points forms the Bézier curve, which is the final first trajectory.
[0102] The technical solution provided in this disclosure can determine valid audio frames in audio information and determine at least one trajectory control point based on the pitch information of the valid audio frames. Then, based on the trajectory control points and a preset Bezier function, a first trajectory is determined, improving the effectiveness of determining the first trajectory.
[0103] Figure 3 is a flowchart illustrating the determination of trajectory control points provided in the embodiments of this disclosure. Based on the foregoing embodiments, the method of determining trajectory control points in the above embodiments can be further refined. For specific implementation methods, please refer to the detailed description of this technical solution. Technical terms that are the same as or corresponding to those in the above embodiments will not be repeated in this embodiment.
[0104] As shown in Figure 3, the method includes:
[0105] S310. When the number of valid audio frames reaches the first threshold, determine the initial tone information based on the tone information of the valid audio frames.
[0106] The first threshold is a pre-set natural number. The first threshold is primarily used to determine how many valid audio frames need to be used as a reference to determine the initial pitch information. The first threshold is also used to determine from which valid audio frames to begin determining the trajectory control points based on those frames. Optionally, the first threshold can be 10; that is, if the number of valid audio frames reaches 10, the initial pitch information can be determined based on these 10 valid audio frames.
[0107] Specifically, after identifying the valid audio frames in the audio information, the first trajectory can be left undefined and used as a reference data to determine the initial pitch information.
[0108] For example, the first threshold is 10, and the pitch information is the pitch value. After determining 10 valid audio frames based on the audio information, the average pitch value of these 10 valid audio frames can be processed to obtain the initial pitch information (initial pitch value) initPitch.
[0109] S320. Determine the pitch offset based on the initial pitch information and pitch control information.
[0110] The initial pitch information includes the initial pitch value, and the pitch control information is the pitch control value.
[0111] Specifically, by calculating the difference between the initial pitch value and the pitch control value, the pitch offset corresponding to each valid audio frame can be obtained. Determining the pitch offset facilitates subsequent pitch control processing for both the first and second types, ensuring the stability of the first trajectory relative to the video frame.
[0112] S330. Starting from the valid audio frames after the first threshold, the pitch information of the valid audio frames and the preset number of historical valid audio frames before the valid audio frames is averaged to determine the pitch information to be adjusted for the valid audio frames.
[0113] It should be noted that for each scenario, the pitch offset is determined based on the pitch value of the valid audio frames corresponding to the first threshold when the number of valid audio frames reaches the first threshold. That is, the first trajectory is generated starting from the valid audio frames determined after the first threshold. The processing method for valid audio frames generated after the first threshold can be found in the detailed explanation below.
[0114] It should also be noted that the valid audio frames generated after the first threshold are determined in the following way. Here, we will take one of the valid audio frames as an example to illustrate the determination. This valid audio frame can be taken as the current valid audio frame.
[0115] The preset quantity is a value pre-set according to actual needs. This value is mainly used to constrain how many historical valid audio frames need to be combined to constrain the current valid audio frame. Optionally, the preset quantity is 5 or 10 historical valid audio frames. Valid audio frames determined before the current valid audio frame can be called candidate historical valid audio frames. Each historical valid audio frame has a corresponding historical timestamp. The interval duration can be determined based on the historical timestamp and the current timestamp of the current valid audio frame. A preset number of candidate historical valid audio frames can be obtained based on the interval duration as the target historical valid audio frames for the current audio frame. The tone information to be used is the tone information based on the target historical valid audio frames.
[0116] For example, the first threshold is 10 valid audio frames, the preset value is 9 historical valid audio frames, and the current valid audio frame is the 11th valid audio frame. The pitch information of the 11th valid audio frame and the pitch information of the 2nd to 10th target historical valid audio frames are obtained. By averaging the pitch information of the 2nd to 11th valid audio frames, the avgPitch information to be adjusted for the 11th valid audio frame is obtained.
[0117] In other words, in order to improve the smoothness of the first trajectory, the pitch information to be adjusted is obtained by averaging the pitch information of the current valid audio frame and the previous nine target historical valid audio frames.
[0118] S340. Determine the tone information to be used for the valid audio frame based on the tone information to be adjusted and the tone offset.
[0119] This can be understood as follows: For valid audio frames after the first threshold, the pitch information to be used for the valid audio frame can be obtained based on the pitch information to be adjusted and the pitch offset of the valid audio frame.
[0120] For example, after obtaining the pitch information avgPitch of the 11th valid audio frame, the sum of the pitch information avgPitch and the pitch offset pitchOffset can be calculated to obtain the pitch information curPitch of the 11th valid audio frame.
[0121] S350: Based on the first tone, the second tone, and the tone information to be used, determine the tone information of the valid audio frame and map it to the target screen coordinates in the video frame.
[0122] The target screen coordinates are the coordinates obtained after mapping the valid audio frames onto the video screen.
[0123] For example, a video frame includes multiple pixels, each corresponding to a screen coordinate. Typically, the vertical screen coordinate ranges from 0 to 1 from top to bottom, and the pitch range is primarily from the first pitch (lowPitch) to the second pitch (highPitch). Based on this, the pitch information to be used is mapped to the target screen coordinates as follows: curHight = (curPitch - lowPitch) / (highPitch - lowPitch). curHight is the vertical coordinate of the target screen coordinates. The horizontal coordinate in the target screen coordinates is determined based on a pre-set trajectory movement speed.
[0124] S360. Take the newly added historical screen coordinates in the coordinate set as the screen coordinates to be used, and determine the difference between the ordinate of the screen coordinates to be used and the ordinate of the target screen coordinates.
[0125] The coordinate set includes multiple historical screen coordinates. The most recently added historical screen coordinates to the coordinate set can be used as the screen coordinates to be used.
[0126] Specifically, after obtaining the target screen coordinates of a valid audio frame, the most recently added historical screen coordinates can be retrieved from the coordinate set. The difference between the ordinates of the target screen coordinates and the historical screen coordinates can then be calculated.
[0127] S370. Update the coordinate set according to the difference, so as to determine the at least one trajectory control point based on the updated coordinate set.
[0128] The number of at least one trajectory control point can be one or more, and the specific number is related to the order of the Bessel function.
[0129] Specifically, based on the difference, it can be determined how to process the target screen coordinates to update the coordinate set. After updating the coordinate set, at least one trajectory control point can be determined.
[0130] In this embodiment, updating the coordinate set based on the difference to determine at least one trajectory control point includes: processing the target screen coordinates according to a preset condition satisfied by the difference using a coordinate processing method corresponding to the preset condition, and updating the processed target screen coordinates as historical screen coordinates to the coordinate set; when the number of historical screen coordinates in the coordinate set reaches a preset threshold, obtaining multiple screen coordinates to be applied corresponding to the preset threshold based on the time information of adding historical screen coordinates to the coordinate set; and determining at least one trajectory control point based on the multiple screen coordinates to be applied.
[0131] This can be understood as follows: based on the preset conditions satisfied by the difference, the coordinate processing method corresponding to the preset conditions can be retrieved to process the target screen coordinates, and the processed target screen coordinates are updated to the coordinate set as historical screen coordinates.
[0132] In this embodiment, there are at least three preset conditions, and correspondingly, there are at least three coordinate processing methods. The following describes each preset condition and the coordinate processing method corresponding to that preset condition in detail.
[0133] The first method: The preset condition is that the difference is less than the first difference threshold. The coordinate processing method is to keep the target screen coordinates unchanged and delete the screen coordinates to be used from the coordinate set.
[0134] This can be understood as follows: if the coordinate difference is less than the first difference threshold, the target screen coordinates can be kept unchanged, and the screen coordinates to be used can be deleted from the coordinate set.
[0135] The second method: The preset condition is that the difference is greater than the second difference threshold, and the coordinate processing method is to correct the target screen coordinates based on the second difference threshold.
[0136] This can be understood as follows: if the difference is greater than the second difference threshold, the second difference threshold can be superimposed on the vertical coordinate of the target screen coordinates to obtain the updated target screen coordinates.
[0137] The third method: The preset condition is that the difference is greater than the first difference threshold and less than the second difference threshold, and the coordinate processing method is to keep the target screen coordinates unchanged.
[0138] This can be understood as follows: if the difference is greater than the first difference threshold and less than the second difference threshold, the target screen coordinates can remain unchanged.
[0139] After the coordinate set is updated, multiple screen coordinates to be applied, matching a preset threshold, can be determined based on the number of historical screen coordinates in the coordinate set and the time they were added. Through analysis and processing of these screen coordinates, at least one trajectory point can be identified. Optionally, the difference between the time the historical screen coordinates were added to the coordinate set and the current time can be calculated. Based on the ascending order of the difference, a preset number of historical screen coordinates are selected as the screen coordinates to be applied.
[0140] The technical solution provided in this disclosure can process the pitch information and pitch offset of the effective audio frame to obtain the pitch information to be used of the effective audio frame. Then, based on the target screen coordinates corresponding to the pitch information to be used and the historical screen coordinates in the coordinate set, the trajectory control point is determined, which improves the accuracy of determining the trajectory control point. Furthermore, when the accuracy of the trajectory control point is high, the smoothness of the first trajectory can be improved when determining the first trajectory based on the trajectory control point, thereby improving the display effect.
[0141] Figure 4 is a flowchart illustrating the generation of trajectory control points provided in this embodiment. Based on the foregoing embodiments, the following can be further refined: "determining at least one trajectory control point based on the plurality of screen coordinates to be applied" and "determining the first trajectory based on the at least one trajectory control point and a preset Bezier function". For specific implementation details, please refer to the detailed description of this technical solution. Technical terms that are the same as or corresponding to those in the above embodiments will not be repeated in this embodiment.
[0142] As shown in Figure 4, the method includes:
[0143] S410. Based on the time order in which the multiple screen coordinates to be applied are added to the coordinate set, determine the slope value between two adjacent screen coordinates to be applied, starting from the second screen coordinate to be applied.
[0144] It should be noted that the default Bezier function is a second-order Bezier function. Therefore, the number of screen coordinates to be applied can be four, and correspondingly, the number of control points determined in the end is three.
[0145] In this system, each screen coordinate to be applied is added to the coordinate set at a different time. Based on the order in which they were added to the coordinate set, the slope between adjacent screen coordinates can be determined, starting from the second screen coordinate to be applied. For example, the points corresponding to the screen coordinates to be applied are labeled P1, P2, P3, and P4. According to the order in which they were added to the coordinate set from earliest to latest, P1 is earlier than P2, P2 is earlier than P3, and P3 is earlier than P4. Starting from P2, the slope K1 between P2 and P3, and the slope K2 between P3 and P4, can be calculated.
[0146] S420. Based on the slope value, determine the trajectory control point from the plurality of screen coordinates to be applied.
[0147] Specifically, after obtaining the two slope values mentioned above, the relationship between the two slope values can be analyzed and processed so that the trajectory control points can be determined from multiple screen coordinates to be applied based on the analysis results.
[0148] In this embodiment, the method for determining the trajectory control point based on the slope value can be as follows: if the difference between the slope values is greater than the slope difference threshold, then the application screen coordinates other than the last application screen coordinates in the time sequence are used as the trajectory control point; if the difference between the slope values is less than the slope difference threshold, then the application screen coordinates other than the second application screen coordinates are used as the trajectory control point.
[0149] This can be understood as follows: if the difference between slope values is greater than the slope difference threshold, then the screen coordinates of P1, P2, and P3 can be used as trajectory control points. If the difference between slope values is less than or equal to the slope difference threshold, then P1, P2, and P4 can be used as trajectory control points. These trajectory control points are mainly used for subsequent processing of the preset Bezier function to generate the first trajectory.
[0150] S430. Substitute the target screen coordinates of the trajectory control points into the preset Bezier function, and obtain the trajectory points of the effective audio frames in the video frame by adjusting the independent variables of the preset Bezier function, and generate the first trajectory based on the trajectory points.
[0151] The preset Bessel function can be S t =(1-t) 2 S0+2t(1-t)S1+t 2 S2 t∈(0,1); S0 is the first trajectory control point in time sequence, S1 is the second trajectory control point in time sequence, S2 is the third trajectory control point in time sequence, and t is the independent variable.
[0152] Taking trajectory control points P1, P2, and P3 as an example, P1 corresponds to S0, P2 corresponds to S1, and P3 corresponds to S2. The target screen coordinates corresponding to these trajectory control points are substituted into a preset Bezier function. Next, the value of the independent variable t is adjusted to obtain multiple trajectory points. Connecting all trajectory points yields the first trajectory.
[0153] In other words, the independent variables t≥0 and t<0.5 can be substituted into the above-mentioned preset Bezier function to obtain the trajectory points, and these trajectory points can be displayed in the video screen. The splicing of all trajectory points is the first trajectory.
[0154] It should also be noted that the coordinates of the trajectory points obtained by substituting the independent variable t=0.5 into the preset Bezier function can be used to update the coordinates of the second screen to be used. Optionally, the screen coordinates of the P2 trajectory control point can be updated to the coordinates obtained at t=0.5.
[0155] In other words, the coordinate information of the trajectory point corresponding to the target independent variable is used as the second screen coordinate to be applied and updated in the coordinate set.
[0156] In this embodiment, valid audio frames after the first threshold can be processed to determine trajectory points. Valid audio frames can be continuously analyzed and processed to update the length of the first trajectory.
[0157] The technical solution provided in this disclosure can further combine the determined trajectory control points and the preset Bezier function to determine the first trajectory, improve the smoothness of the first trajectory, and thus improve the display effect.
[0158] Figure 5 is a schematic flowchart of a video generation method provided in an embodiment of this disclosure. Based on the foregoing embodiments, the step of "determining a first special effects trajectory associated with the audio information based on the first trajectory and a pre-set first material image" can be further refined. For specific implementation details, please refer to the detailed description of this technical solution. Technical terms that are the same as or corresponding to those in the above embodiments will not be repeated in this embodiment.
[0159] As shown in Figure 5, the method includes:
[0160] S510. During the video display process, the first trajectory is determined based on the collected audio information.
[0161] S520. Based on the coordinate information of each trajectory point in the first trajectory and the pixel information of each pixel in the first material image, update the display information of each trajectory point in the first trajectory to obtain the first special effect trajectory.
[0162] [Corrected according to Rule 91, October 2025] The first source image is a source image set according to a preset effect. Optionally, the first source image, as shown in Figure 6, can be a background image with a color / transparency gradient. For example, a first color is set at the first end of the first source image, and a second color is set at the second end. By interpolating the first and second colors, the color information of each pixel in the first source image is obtained. It should be noted that Figure 6 only schematically illustrates the color / transparency gradient; for example, the first end can be set to the first color, such as blue; the second end can be set to the second color, such as pink; from the first end to the second end, for example, interpolation can be used to make the color gradually change from the first color to the second color, thereby obtaining the color information of each pixel in the first source image. Furthermore, transparency information can also be set. By adjusting the transparency of each pixel in the first source image, an updated first source image is obtained.
[0163] After obtaining the first trajectory, the first source image can be traversed based on the horizontal coordinates of the trajectory points to determine the corresponding pixel in the first source image. The display information of the trajectory point is then updated according to the pixel value of that pixel, and the first trajectory with updated display information is used as the first special effect trajectory. A schematic diagram of the first special effect trajectory after processing the first trajectory based on Figure 6 can be found in Figure 7.
[0164] [Correction based on Rule 91, 29.10.2025] It should be noted that the colors and transparency in Figures 6 and 7 are merely illustrative examples. The image content of the first material image can be set according to actual needs during the design phase. The specific content is not limited in this embodiment.
[0165] The rendered map is queried based on the position information of each trajectory point in the first trajectory, and the display information of the trajectory points is updated according to the transparency and actual color in the rendered map. The size of the first source image is consistent with the size of the video frame, and the display information of any two pixels in the first source image differs.
[0166] In this embodiment, in order to further improve the display effect of the first special effect trajectory, the obtained first special effect trajectory can be further processed. Optionally, the method further includes: performing Gaussian blur processing on the first special effect trajectory to obtain the special effect trajectory to be mixed; and performing fusion processing on the special effect trajectory line to be mixed and the first special effect trajectory before Gaussian blur processing to obtain the first special effect trajectory after the display information is updated.
[0167] This can be understood as follows: the first effect trajectory can be Gaussian blurred to obtain the effect trajectory to be blended. Next, the effect trajectory to be blended is mixed with the first effect trajectory before the Gaussian blur, resulting in the first effect trajectory after the display information is updated.
[0168] For example, after obtaining the first special effect trajectory as shown in Figure 7, a Gaussian blur algorithm is used to blur the first special effect trajectory, resulting in the effect diagram shown in Figure 8. Pixel overlay processing is then performed on the special effect trajectory to be mixed and the first special effect trajectory before Gaussian blur processing to update the display information of each pixel in the first special effect trajectory.
[0169] S530. Determine the target mask image based on the first trajectory.
[0170] The shape of the trajectory in the target mask image is related to the shape of the first trajectory, and there is an image after dilation processing.
[0171] In this embodiment, the method for determining the target mask image may be as follows: obtaining an initial trajectory corresponding to the first trajectory; offsetting the pixels in the initial trajectory by a preset number of pixels along a first direction to obtain a trajectory to be disturbed; perturbing the trajectory to be disturbed according to a preset range of disturbed pixels to obtain the target disturbed trajectory; and using the area formed by the initial trajectory and the target disturbed trajectory as the mask area in the target mask image corresponding to the first trajectory.
[0172] Figure 9 can be used as a reference to understand how to determine the target mask area. The trajectory corresponding to the first trajectory can be copied as the initial trajectory. The initial trajectory is moved along the first direction by a preset number of pixels (optionally 60px). The pixels on the initial trajectory are moved 60px along the first direction perpendicular to the horizontal plane, resulting in the effect shown in Figure 9(a). This obtained trajectory is used as the trajectory to be disturbed. The disturbance pixel range can be a pixel fluctuation range set according to actual needs; optionally, the pixel fluctuation range can be a random fluctuation of 0px-25px along the first or second direction. Each pixel in the trajectory to be disturbed can be randomly disturbed, resulting in the effect shown by the red curve in Figure 9(b). The trajectory corresponding to the red curve is used as the disturbed target disturbance trajectory. The area formed by the initial trajectory and the target disturbance trajectory is used as the mask area in the target mask image corresponding to the first trajectory, as shown in Figure 9(c).
[0173] Based on the above technical solution, a second special effect trajectory can be generated based on the obtained target mask image. Alternatively, the target mask image can be further processed to improve the display effect of the generated second special effect trajectory.
[0174] Optionally, Gaussian blurring is performed on Figure 9(c) to obtain the effect shown in Figure 10. This can be understood as Gaussian blurring the target mask image to obtain an updated target mask image.
[0175] S540: Based on the target mask image and the second source image, determine the second special effect trajectory of the first trajectory.
[0176] The second material image is also a material image created based on the pre-set effect. Optionally, the second material can be a material image simulating an aurora scene, as shown in Figure 11.
[0177] It should be noted that the first and second source images can be changed according to actual needs. This embodiment only uses the source images shown in the accompanying drawings for illustrative purposes and does not limit the solution. The core of the technical solution provided by this disclosure is that this technology can be used to generate special effects trajectories.
[0178] Specifically, by processing the target mask image and the second source image, a second special effects trajectory corresponding to the first trajectory can be obtained.
[0179] In this embodiment, the method for generating the second special effect trajectory based on the target mask image and the second source image can be: updating the pixel information of at least one pixel in the target mask image based on the pixel information of at least one pixel in the second source image to obtain the second special effect trajectory.
[0180] Specifically, the corresponding pixel in the target mask image can be identified as the pixel to be used in the second source image, and the pixel information of the pixel to be used can be updated to the pixel in the target mask image. After processing each pixel in the target mask image, the second effect trajectory can be obtained, as shown in Figure 12.
[0181] S550. Based on the first special effects trajectory and the second special effects trajectory, determine and display the target trajectory displayed in the video frame.
[0182] In this embodiment, the endpoints of the first special effects trajectory and the second special effects trajectory can be kept consistent and then merged to obtain the target trajectory shown in the video.
[0183] Based on the above technical solution, if the endpoint of the target trajectory is a preset screen coordinate, the trajectory length of the target trajectory will no longer be updated; or, if no audio information is collected within a preset time period, the trajectory length of the target trajectory will no longer be updated.
[0184] This can be understood as follows: if the target trajectory's endpoint is a preset screen coordinate (optionally, the preset screen coordinate could be the screen coordinate corresponding to the ground on the other side of a hill, or the preset screen coordinate corresponding to the ground on the other side of a ravine), then the estimated trajectory length of the target trajectory no longer needs to be updated; that is, the collected audio does not need to be further processed. Alternatively, the audio information can be continuously analyzed and processed to generate the target trajectory. If, optionally, no audio information is collected within a preset time period of 20 seconds, the trajectory length of the target trajectory no longer needs to be updated. Of course, if audio information is collected within 20 seconds, the trajectory length of the target trajectory can continue to be updated. Based on the above technical solution, the target trajectory can also include a secondary line. Multiple copies of the first effect trajectory can be copied to obtain the trajectory to be merged, and the thickness value of the trajectory to be merged can be adjusted to obtain the trajectory to be merged as shown in Figure 13.
[0185] The video footage obtained by merging the fusion trajectory, the first special effects trajectory, and the second special effects trajectory together and displaying them in the video frame can be seen in Figure 14.
[0186] The technical solution provided in this disclosure can obtain a first special effect trajectory by processing a first special effect on a first source image, and obtain a second special effect trajectory by processing a mask image corresponding to the first trajectory with a second source image. Based on the first and second special effect trajectories, the target trajectory displayed in the video frame can be determined, achieving not only the display of the target trajectory corresponding to the audio information, which is obtained after multiple special effect processing, but also improving the richness of the video frame content and the interactivity between the video frame and the user.
[0187] Figure 15 is a schematic diagram of a video generation device provided in an embodiment of the present disclosure. As shown in Figure 15, the device includes: a first trajectory determination module 610, a first special effects trajectory generation module 620, a second special effects trajectory generation module 630, and a target trajectory generation module 640.
[0188] The first trajectory determination module 610 is used to determine a first trajectory based on the collected audio information during the display of the video screen; the first special effects trajectory generation module 620 is used to determine a first special effects trajectory associated with the audio information based on the first trajectory and a pre-set first material image; wherein the display information corresponding to the first special effects trajectory is related to the pixel information of the first material image; the second special effects trajectory generation module 630 is used to determine a second special effects trajectory of the first trajectory based on a target mask image and a second material image associated with the first trajectory; the target trajectory generation module 640 is used to determine and display the target trajectory displayed in the video screen based on the first special effects trajectory and the second special effects trajectory; wherein the change information of the target trajectory is associated with the audio information.
[0189] The technical solution of this disclosure embodiment allows for the analysis and processing of collected audio information during video display to determine a first trajectory. Further, a first source image can be used to apply special effects to the first trajectory, resulting in a first special effects trajectory, which enriches the visual effect of the first trajectory. To further enhance the richness and appeal of the video content, a second source image can be used to render the mask image corresponding to the first trajectory, resulting in a second special effects trajectory. Based on the first and second special effects trajectories, the target trajectory displayed in the video can be obtained. This solves the problem that simply converting audio information to text results in a narrow scope of audio information usage and low interactivity with terminal devices. It achieves the generation of a corresponding first trajectory based on audio information, and the fusion display of the first trajectory after multi-dimensional special effects processing, thereby improving the richness of the video content and the interactivity with the user.
[0190] Based on the above technical solution, the device includes: an audio acquisition module, used to acquire audio information of a first object when the trajectory generation conditions are met;
[0191] The trajectory generation conditions include at least one of the following:
[0192] The second object is moved to a preset position based on a preset map, wherein the second object is a movable object shown in the video frame;
[0193] Trigger the video frame or the trajectory generation control in the application software to which the video frame belongs.
[0194] Based on the above technical solutions, the first trajectory determination module includes:
[0195] The trajectory control point determination unit is used to determine at least one trajectory control point based on the pitch information of the valid audio frames in the audio information, wherein the valid audio frame is an audio frame whose volume meets a first volume condition and whose pitch meets a first pitch condition, and the trajectory control point is used to adjust the trajectory shape of the first trajectory.
[0196] The first trajectory determination unit is used to determine the first trajectory based on the at least one trajectory control point and a preset Bezier function.
[0197] Based on the above technical solutions, the trajectory control point determination unit includes:
[0198] The pitch offset determination subunit is used to determine the pitch offset based on the pitch information and pitch control information of the valid audio frames in the audio information; wherein, the pitch control information is determined based on the first pitch of the first type and the second pitch of the second type, and the pitch control information is used to control the pitch information to fluctuate within a preset range;
[0199] The tone information to be used determination subunit is used to adjust the tone information of the effective audio frame according to the tone offset to obtain the tone information to be used;
[0200] The trajectory control point determination subunit is used to determine the target screen coordinates of the tone information to be used in the video frame, and to determine the trajectory control point based on the target screen coordinates and the historical screen coordinates in the coordinate set.
[0201] Based on the above technical solutions, the pitch offset determination subunit includes:
[0202] An initial pitch determination subunit is used to determine initial pitch information based on the pitch information of the valid audio frames when the number of valid audio frames reaches a first threshold.
[0203] The offset determination subunit is used to determine the pitch offset based on the initial pitch information and the pitch control information;
[0204] The subunit for determining the tone information to be used includes:
[0205] The first information determining unit is configured to determine the pitch information to be adjusted for the valid audio frames determined after the first threshold, based on the valid audio frames and the target historical valid audio frames; wherein the target historical valid audio frames are determined based on all selectable historical valid audio frames of the audio information, the interval duration of the valid audio frames, and the preset number of historical valid video frames;
[0206] The second pitch determination unit is used to determine the pitch information to be used in the effective audio frame based on the pitch information to be adjusted and the pitch offset.
[0207] Based on the above technical solutions, the trajectory control point determination subunit includes:
[0208] The first coordinate determination subunit is used to determine the target screen coordinates in the video frame by mapping the tone information of the effective audio frame to the tone information of the video frame, based on the first tone, the second tone and the tone information to be used.
[0209] The difference determination subunit is used to take the newly added historical screen coordinates in the coordinate set as the screen coordinates to be used, and determine the difference between the ordinate of the screen coordinates to be used and the ordinate of the target screen coordinate; wherein, the screen coordinates to be used are the coordinates in the historical screen coordinates;
[0210] The trajectory point determination subunit is used to update the coordinate set according to the difference, so as to determine the at least one trajectory control point based on the updated coordinate set.
[0211] Based on the above technical solutions, the trajectory point determination subunit includes:
[0212] The coordinate set update subunit is used to retrieve the coordinate processing method corresponding to the preset condition according to the difference, process the target screen coordinates, and update the processed target screen coordinates as historical screen coordinates to the coordinate set.
[0213] The screen coordinate determination subunit is used to obtain multiple screen coordinates to be applied corresponding to the preset number threshold by adding time information to the coordinate set based on the historical screen coordinates when the number of historical screen coordinates in the coordinate set reaches a preset number threshold.
[0214] The trajectory point determination subunit is used to determine at least one trajectory control point based on the multiple screen coordinates to be applied.
[0215] Based on the above technical solutions, the coordinate processing methods corresponding to the preset conditions include at least one of the following:
[0216] The preset condition is that the difference is less than a first difference threshold, and the coordinate processing method is to keep the target screen coordinates unchanged and delete the screen coordinates to be used from the coordinate set.
[0217] The preset condition is that the difference is greater than the second difference threshold, and the coordinate processing method is to correct the target screen coordinates based on the second difference threshold;
[0218] The preset condition is that the difference is greater than a first difference threshold and less than a second difference threshold, and the coordinate processing method is to keep the target screen coordinates unchanged.
[0219] Based on the above technical solutions, the trajectory point determination subunit is used for
[0220] Based on the time order in which the multiple screen coordinates to be applied are added to the coordinate set, the slope value between two adjacent screen coordinates to be applied is determined starting from the second screen coordinate to be applied.
[0221] Based on the slope value, the trajectory control point is determined from the plurality of screen coordinates to be applied.
[0222] Based on the above technical solutions, the trajectory point determination subunit is used for
[0223] If the difference between the slope values is greater than the slope difference threshold, then the application screen coordinates other than the application screen coordinates with the latest time sequence will be used as the trajectory control points.
[0224] If the difference between the slope values is less than the slope difference threshold, then the screen coordinates to be applied other than the second screen coordinates to be applied are used as the trajectory control points.
[0225] Based on the above technical solutions, the first trajectory determination unit is also used for
[0226] Substitute the target screen coordinates of the trajectory control points into the preset Bezier function, and obtain the trajectory points of the effective audio frames in the video frame by adjusting the independent variables of the preset Bezier function; determine the first trajectory based on the trajectory points.
[0227] Based on the above technical solutions, the first special effects trajectory generation module is also used for
[0228] Based on the coordinate information of each trajectory point in the first trajectory and the pixel information of each pixel in the first source image, the display information of each trajectory point in the first trajectory is updated to obtain the first special effect trajectory; wherein, the size of the first source image is consistent with the size of the video screen, and the display information of any two pixels in the first source image is different.
[0229] Based on the above technical solutions, the device further includes: an effects update module, used to perform Gaussian blur processing on the first effects trajectory to obtain a effects trajectory to be mixed; and to obtain the first effects trajectory after display information update by merging the effects trajectory line to be mixed with the first effects trajectory before Gaussian blur processing.
[0230] Based on the above technical solutions, the device further includes:
[0231] Obtain an initial trajectory corresponding to the first trajectory; offset the pixels in the initial trajectory by a preset number of pixels along a first direction to obtain a trajectory to be disturbed; perturb the trajectory to be disturbed according to a preset range of disturbed pixels to obtain the target disturbed trajectory; use the area formed by the initial trajectory and the target disturbed trajectory as the mask area in the target mask image corresponding to the first trajectory.
[0232] Based on the above technical solutions, the second special effects trajectory generation module is further used to update the pixel information of at least one pixel in the target mask image according to the pixel information of at least one pixel in the target mask image in the second material image, so as to obtain the second special effects trajectory.
[0233] The video generation apparatus provided in this disclosure can execute the video generation method provided in any embodiment of this disclosure, and has the corresponding functional modules and beneficial effects of executing the method.
[0234] It is worth noting that the various units and modules included in the above-mentioned device are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be realized; in addition, the specific names of each functional unit are only for easy differentiation and are not used to limit the protection scope of the embodiments of this disclosure.
[0235] Figure 16 is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. Referring now to Figure 16, a schematic diagram of the structure of an electronic device (e.g., the terminal device or server in Figure 16) 700 suitable for implementing embodiments of this disclosure is shown. The terminal device in embodiments of this disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. The electronic device shown in Figure 16 is merely an example and should not impose any limitations on the functionality and scope of use of embodiments of this disclosure.
[0236] As shown in Figure 16, the electronic device 700 may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 701, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 702 or a program loaded from a storage device 708 into a random access memory (RAM) 703. The RAM 703 also stores various programs and data required for the operation of the electronic device 700. The processing unit 701, ROM 702, and RAM 703 are interconnected via a bus 704. An edit / output (I / O) interface 705 is also connected to the bus 704.
[0237] Typically, the following devices can be connected to I / O interface 705: input devices 706 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 707 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 708 including, for example, magnetic tapes, hard disks, etc.; and communication devices 709. Communication device 709 allows electronic device 700 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 16 shows electronic device 700 with various devices, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.
[0238] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication device 709, or installed from storage device 708, or installed from ROM 702. When the computer program is executed by processing device 701, it performs the functions defined in the methods of embodiments of this disclosure.
[0239] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.
[0240] The electronic device provided in this disclosure and the video generation method provided in the above embodiments belong to the same inventive concept. Technical details not described in detail in this disclosure can be found in the above embodiments, and this embodiment has the same beneficial effects as the above embodiments.
[0241] This disclosure provides a computer storage medium storing a computer program that, when executed by a processor, implements the video generation method provided in the above embodiments.
[0242] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in connection with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
[0243] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol), and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and end-to-end networks (e.g., ad hoc end-to-end networks), as well as any currently known or future-developed networks.
[0244] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.
[0245] The aforementioned computer-readable medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to:
[0246] During the video display, the first trajectory is determined based on the collected audio information;
[0247] Based on the first trajectory and a pre-set first source image, a first special effect trajectory associated with the audio information is determined; wherein, the display information corresponding to the first special effect trajectory is related to the pixel information of the first source image;
[0248] Based on the target mask image and the second source image associated with the first trajectory, a second special effect trajectory of the first trajectory is determined;
[0249] Based on the first special effects trajectory and the second special effects trajectory, the target trajectory displayed in the video frame is determined and displayed; wherein, the change information of the target trajectory is associated with the audio information.
[0250] Computer program code for performing the operations of this disclosure can be written in one or more programming languages or a combination thereof, including but not limited to object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0251] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0252] The units described in the embodiments of this disclosure can be implemented in software or in hardware. The name of a unit does not necessarily limit the unit itself; for example, the first acquisition unit can also be described as "a unit that acquires at least two Internet Protocol addresses".
[0253] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.
[0254] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0255] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.
[0256] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.
[0257] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative examples of implementing the claims.
Claims
1. A video generation method, comprising: During the video display, the first trajectory is determined based on the collected audio information; Based on the first trajectory and a pre-set first source image, a first special effect trajectory associated with the audio information is determined; wherein, the display information corresponding to the first special effect trajectory is related to the pixel information of the first source image; Based on the target mask image and the second source image associated with the first trajectory, a second special effect trajectory of the first trajectory is determined; Based on the first special effects trajectory and the second special effects trajectory, the target trajectory displayed in the video frame is determined and displayed; wherein, the change information of the target trajectory is associated with the audio information.
2. The method according to claim 1, wherein, During the display of the video footage, the method further includes: When the trajectory generation conditions are met, the audio information of the first object is collected; The trajectory generation conditions include at least one of the following: The second object is moved to a preset position based on a preset map, wherein the second object is a movable object shown in the video frame; Trigger the video frame or the trajectory generation control in the application software to which the video frame belongs.
3. The method according to claim 1 or 2, wherein, Determining the first trajectory based on the collected audio information includes: Based on the pitch information of the valid audio frames in the audio information, at least one trajectory control point is determined, wherein the valid audio frame is an audio frame whose volume meets a first volume condition and whose pitch meets a first pitch condition, and the trajectory control point is used to adjust the trajectory shape of the first trajectory. The first trajectory is determined based on the at least one trajectory control point and a preset Bezier function.
4. The method according to claim 3, wherein, The step of determining at least one trajectory control point based on the pitch information of valid audio frames in the audio information includes: The pitch offset is determined based on the pitch information and pitch control information of the valid audio frames in the audio information; wherein the pitch control information is determined based on the first pitch of the first type and the second pitch of the second type, and the pitch control information is used to control the pitch information to fluctuate within a preset range; Adjust the pitch information of the effective audio frame according to the pitch offset to obtain the pitch information to be used; The target screen coordinates of the tone information to be used in the video frame are determined, and the trajectory control points are determined based on the target screen coordinates and the historical screen coordinates in the coordinate set.
5. The method according to claim 4, wherein, The step of determining the pitch offset based on the pitch information and pitch control information of the valid audio frames in the audio information includes: When the number of valid audio frames reaches a first threshold, initial pitch information is determined based on the pitch information of the valid audio frames; The pitch offset is determined based on the initial pitch information and the pitch control information; The step of adjusting the pitch information of the effective audio frame according to the pitch offset to obtain the pitch information to be used includes: For valid audio frames determined after the first threshold, the pitch information to be adjusted for the valid audio frames is determined based on the valid audio frames and the target historical valid audio frames; wherein, the target historical valid audio frames are determined based on all selectable historical valid audio frames of the audio information, the interval duration of the valid audio frames, and the preset number of historical valid video frames; Based on the pitch information to be adjusted and the pitch offset, the pitch information to be used for the valid audio frame is determined.
6. The method according to claim 4 or 5, characterized in that, The step of determining the target screen coordinates of the tone information to be used in the video frame, and determining the trajectory control points based on the target screen coordinates and historical screen coordinates in the coordinate set, includes: Based on the first tone, the second tone, and the tone information to be used, the tone information of the effective audio frame is mapped to the target screen coordinates in the video frame; The newly added historical screen coordinates in the coordinate set are used as the screen coordinates to be used, and the difference between the ordinate of the screen coordinates to be used and the ordinate of the target screen coordinates is determined; wherein, the screen coordinates to be used are the coordinates in the historical screen coordinates. The coordinate set is updated based on the difference to determine the at least one trajectory control point based on the updated coordinate set.
7. The method according to claim 6, wherein, The step of updating the coordinate set based on the difference, and determining the at least one trajectory control point based on the updated coordinate set, includes: Based on the preset conditions satisfied by the difference, the target screen coordinates are processed by the coordinate processing method corresponding to the preset conditions, and the processed target screen coordinates are updated to the coordinate set as historical screen coordinates. When the number of historical screen coordinates in the coordinate set reaches a preset threshold, time information is added to the coordinate set based on the historical screen coordinates to obtain multiple screen coordinates to be applied corresponding to the preset threshold. Based on the multiple screen coordinates to be applied, at least one trajectory control point is determined.
8. The method according to claim 7, wherein, The coordinate processing methods corresponding to the preset conditions include at least one of the following: The preset condition is that the difference is less than a first difference threshold, and the coordinate processing method is to keep the target screen coordinates unchanged and delete the screen coordinates to be used from the coordinate set. The preset condition is that the difference is greater than the second difference threshold, and the coordinate processing method is to correct the target screen coordinates based on the second difference threshold; The preset condition is that the difference is greater than a first difference threshold and less than a second difference threshold, and the coordinate processing method is to keep the target screen coordinates unchanged.
9. The method according to claim 7 or 8, wherein, Determining the at least one trajectory control point based on the plurality of screen coordinates to be applied includes: Based on the time order in which the multiple screen coordinates to be applied are added to the coordinate set, the slope value between two adjacent screen coordinates to be applied is determined starting from the second screen coordinate to be applied. Based on the slope value, the trajectory control point is determined from the plurality of screen coordinates to be applied.
10. The method according to claim 9, wherein, The step of determining the trajectory control point from the plurality of screen coordinates to be applied based on the slope value includes: If the difference between the slope values is greater than the slope difference threshold, then the application screen coordinates other than the application screen coordinates with the latest time sequence will be used as the trajectory control points. If the difference between the slope values is less than the slope difference threshold, then the screen coordinates to be applied other than the second screen coordinates to be applied are used as the trajectory control points.
11. The method according to any one of claims 3-10, wherein, Determining the first trajectory based on the at least one trajectory control point and a preset Bezier function includes: Substitute the target screen coordinates of the trajectory control points into the preset Bezier function, and by adjusting the independent variable of the preset Bezier function, obtain the trajectory points of the effective audio frames in the video frame; Based on the trajectory points, the first trajectory is determined.
12. The method according to any one of claims 1-11, wherein, The step of determining the first special effects trajectory associated with the audio information based on the first trajectory and a pre-set first material image includes: Based on the coordinate information of each trajectory point in the first trajectory and the pixel information of each pixel in the first material image, update the display information of each trajectory point in the first trajectory to obtain the first special effect trajectory; The size of the first source image is the same as the size of the video frame, and the display information of any two pixels in the first source image is different.
13. The method according to any one of claims 1-12, wherein, After obtaining the first special effect trajectory, the method further includes: The first special effect trajectory is Gaussian blurred to obtain the trajectory to be mixed with the special effect; By fusing the trajectory line to be mixed with the first trajectory line before Gaussian blur processing, the first trajectory line after the display information is updated is obtained.
14. The method according to any one of claims 1-13, further comprising: Obtain the initial trajectory corresponding to the first trajectory; The pixels in the initial trajectory are shifted along the first direction by a preset number of pixels to obtain the trajectory to be disturbed. The trajectory to be disturbed is perturbed according to a pre-set range of perturbed pixels to obtain the target perturbed trajectory; The region formed by the initial trajectory and the target disturbance trajectory is used as the mask region in the target mask map corresponding to the first trajectory.
15. A video generation apparatus, comprising: The first trajectory determination module is configured to determine the first trajectory based on the collected audio information during the video display process; The first special effects trajectory generation module is configured to determine a first special effects trajectory associated with the audio information based on the first trajectory and a pre-set first material image; wherein the display information corresponding to the first special effects trajectory is related to the pixel information of the first material image; The second special effects trajectory generation module is configured to determine the second special effects trajectory of the first trajectory based on the target mask image and the second material image associated with the first trajectory. The target trajectory generation module is configured to determine and display the target trajectory shown in the video frame based on the first special effect trajectory and the second special effect trajectory; wherein the change information of the target trajectory is associated with the audio information.
16. An electronic device comprising: One or more processors; A storage device is configured to store one or more programs, wherein, When the one or more programs are executed by the one or more processors, the one or more processors implement the video generation method as described in any one of claims 1-14.
17. A storage medium containing computer-executable instructions, wherein, The computer-executable instructions, when executed by a computer processor, are used to perform the video generation method as described in any one of claims 1-14.
18. A computer program product comprising a computer program, wherein, When the computer program is executed by a processor, it implements the video generation method as described in any one of claims 1-14.