Video generation method and apparatus, device, medium, and product
By generating special effects trajectories based on audio information in video footage, the problem of insufficient interactive audio information is solved, resulting in rich and natural video effects and enhancing the user's interactive experience.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-09-09
- Publication Date
- 2026-03-19
AI Technical Summary
In existing technologies, the limitations of audio information content analysis and processing during user interaction with smart terminals result in insufficient interactive content and fail to fully utilize the rich potential of audio information.
By determining the first trajectory based on the collected audio information during the video display process, and combining it with pre-set material images and target mask images, a special effects trajectory associated with the audio information is generated, and finally the target trajectory is displayed in the video, enhancing the richness and naturalness of the interactive content.
It enables the generation of rich and natural video effects based on audio information, enhancing the interactive experience between users and smart terminals.
Smart Images

Figure CN2025120131_19032026_PF_FP_ABST
Abstract
Description
Video generation method, device, apparatus, medium and product
[0001] Cross-reference to Related Applications
[0002] This application claims priority to Chinese Patent Application No. 202411269947.2, filed on September 10, 2024, the disclosure of which is incorporated herein in its entirety as part of the present application. TECHNICAL FIELD
[0003] Embodiments of the present disclosure relate to a video generation method, device, apparatus, medium and product. BACKGROUND
[0004] With the popularization of intelligent devices, users have more and more interactive needs with intelligent terminals.
[0005] Currently, users can interact with intelligent terminals based on audio information, but at this time, the content analysis and processing of the audio information are mainly used to determine the corresponding interactive content, that is, the interaction is mainly realized by relying on the content of the audio information itself, without combining the more abundant content of the audio information to generate interactive content, which has the problem of limited use of audio information. SUMMARY
[0006] The present disclosure provides a video generation method, device, apparatus, medium and product.
[0007] In a first aspect, the embodiments of the present disclosure provide a video generation method, which comprises:
[0008] In the process of video picture display, a first track is determined based on collected audio information;
[0009] A first special effect track associated with the audio information is determined according to the first track and a first material image set in advance, wherein the display information corresponding to the first special effect track is related to the pixel information of the first material image;
[0010] A second special effect track of the first track is determined based on a target mask image associated with the first track and a second material image;
[0011] A target track displayed in the video picture is determined and displayed based on the first special effect track and the second special effect track, wherein the change information of the target track is associated with the audio information.
[0012] In a second aspect, the embodiments of the present disclosure also provide a video generation device, which comprises:
[0013] The first trajectory determination module is configured to determine a first trajectory based on the collected audio information during the process of displaying the video picture.
[0014] The first special effect trajectory generation module is configured to determine a first special effect trajectory associated with the audio information according to the first trajectory and a first material image set in advance, wherein display information corresponding to the first special effect trajectory is related to pixel information of the first material image.
[0015] The second special effect trajectory generation module is configured to determine a second special effect trajectory of the first trajectory based on a target mask image associated with the first trajectory and a second material image.
[0016] The target trajectory generation module is configured to determine a target trajectory displayed in the video picture based on the first special effect trajectory and the second special effect trajectory and display the target trajectory, wherein change information of the target trajectory is associated with the audio information.
[0017] In a third aspect, an electronic device is provided, and the electronic device includes:
[0018] One or more processors;
[0019] A storage device configured to store one or more programs,
[0020] When the one or more programs are executed by the one or more processors, the one or more processors implement the video generation method according to any of the embodiments of the present disclosure.
[0021] In a fourth aspect, a storage medium containing computer executable instructions is provided, and the computer executable instructions are used to execute the video generation method according to any of the embodiments of the present disclosure when executed by a computer processor.
[0022] In a fifth aspect, a computer program product is provided, and the computer program product includes a computer program which, when executed by a processor, implements the video generation method according to any of the embodiments of the present disclosure. BRIEF DESCRIPTION OF DRAWINGS
[0023] The above and other features, advantages, and aspects of the embodiments of the present disclosure will become more apparent with reference to the following detailed description when taken in conjunction with the accompanying drawings. Throughout the drawings, the same or similar reference numerals are used to represent the same or similar elements. It should be understood that the drawings are schematic and elements and features are not necessarily drawn to scale.
[0024] FIG. 1 is a flow diagram of a video generation method according to an embodiment of the present disclosure;
[0025] FIG. 2 is a flowchart of a video generation method according to an embodiment of the present disclosure;
[0026] FIG. 3 is a flowchart of determining a track control point according to an embodiment of the present disclosure;
[0027] FIG. 4 is a flowchart of a first track according to an embodiment of the present disclosure;
[0028] FIG. 5 is a flowchart of a video generation method according to an embodiment of the present disclosure;
[0029] FIG. 6 is a first material image according to an embodiment of the present disclosure;
[0030] FIG. 7 is a first special effect track according to an embodiment of the present disclosure;
[0031] FIG. 8 is a first special effect track after Gaussian blur processing according to an embodiment of the present disclosure;
[0032] FIG. 9 is a schematic diagram of generating a target mask according to an embodiment of the present disclosure;
[0033] FIG. 10 is a target mask after Gaussian blur processing according to an embodiment of the present disclosure;
[0034] FIG. 11 is a second material image according to an embodiment of the present disclosure;
[0035] FIG. 12 is a second special effect track according to an embodiment of the present disclosure;
[0036] FIG. 13 is a schematic diagram of a track to be fused according to an embodiment of the present disclosure;
[0037] FIG. 14 is a schematic diagram of a target track according to an embodiment of the present disclosure;
[0038] FIG. 15 is a schematic diagram of a video generation device according to an embodiment of the present disclosure; and
[0039] FIG. 16 is a schematic diagram of an electronic device according to an embodiment of the present disclosure. DETAILED DESCRIPTION
[0040] Embodiments of the present disclosure will be described in more detail with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the drawings, it is understood that the present disclosure can be embodied in various forms and should not be interpreted as being limited to the embodiments set forth herein, but rather, these embodiments are provided so that the present disclosure can be more thoroughly and completely understood. It should be understood that the drawings and embodiments of the present disclosure are only for illustrative purposes and are not intended to limit the scope of protection of the present disclosure.
[0041] It should be understood that each step recited in the method embodiments of the present disclosure can be performed in different orders and / or in parallel. In addition, the method embodiments can include additional steps and / or omit performing the steps shown. The scope of the present disclosure is not limited in this respect.
[0042] The term "comprising" and variations thereof as used herein are used inclusively, i.e., "comprising, but not limited to." The term "based on" means "based, at least in part, on." The term "one embodiment" means "at least one embodiment." The term "another embodiment" means "at least one additional embodiment." The term "some embodiments" means "at least some embodiments." Related definitions are given below.
[0043] It should be noted that the terms "first", "second", and the like in the present disclosure are merely used to distinguish different devices, modules or units, and do not imply the order or interdependence of the functions performed by these devices, modules or units.
[0044] It should be noted that the terms "one", "multiple" in the present disclosure are illustrative and not restrictive, and those skilled in the art should understand that unless otherwise explicitly stated in the context, it should be understood as "one or more".
[0045] The names of the messages or information exchanged between the devices in the embodiments of the present disclosure are only for illustrative purposes, and are not intended to limit the scope of the messages or information.
[0046] It can be understood that before using the technical solutions disclosed in the embodiments of the present disclosure, the type, use range, use scenario, etc. of the personal information involved in the present disclosure should be informed to the user and the authorization of the user should be obtained through appropriate means according to relevant laws and regulations.
[0047] For example, in response to receiving the active request of the user, the user is sent prompt information to explicitly prompt the user that the operation requested to be performed will require obtaining and using the personal information of the user. Thus, the user can voluntarily choose whether to provide personal information to the electronic device, application program, server or storage medium, etc. software or hardware performing the operation of the technical solutions of the present disclosure according to the prompt information.
[0048] As an optional but not limited implementation manner, in response to receiving the active request of the user, the manner of sending prompt information to the user may, for example, be a pop-up window manner, and the prompt information can be presented in the form of text in the pop-up window. In addition, the pop-up window can also carry selection controls for the user to select "agree" or "disagree" to provide personal information to the electronic device.
[0049] It can be understood that the above notification and user authorization process is only illustrative and does not limit the implementation of the present disclosure, and other methods that meet relevant laws and regulations can also be applied to the implementation of the present disclosure.
[0050] It can be understood that the data involved in the technical solution (including but not limited to the data itself, the acquisition or use of the data) should comply with the requirements of relevant laws and regulations and relevant provisions.
[0051] Before introducing the technical solution provided by the embodiment of the present disclosure, the application scenario can be exemplarily described. The user can generate a target trajectory by using the scheme provided by the embodiment of the present disclosure when shooting a video through an application software, having a video call with other users, in some application software supporting games, or integrating the scheme provided by the embodiment of the present disclosure as a game function in an application program. The target trajectory is the final special effect trajectory to be displayed.
[0052] In actual application, if a target trajectory corresponding to audio information is to be generated, the scheme provided by the embodiment of the present disclosure can be referred to for implementation.
[0053] [According to Rule 91, correct 29.10.2025] It can be understood that the above application software can be a type of software for image / video processing, and specific application software is not described one by one here, and it can be implemented only for image / video processing. Of course, a forwarding application program can also be developed to implement the software for adding special effects and displaying special effects, or integrated in the corresponding page, and the user can process the special effect video through the PC integrated page.
[0054] For example, if the technical solution provided by the embodiment of the present disclosure is integrated in an application software that can shoot a video or have a video call, the scheme provided by the embodiment of the present disclosure can be integrated as a special effect prop and installed in the corresponding application software. When detecting a trigger operation of the user on the special effect prop, audio information can be collected and a target trajectory displayed in a video picture can be generated based on the audio information. If the scheme provided by the embodiment of the present disclosure is integrated as a function in any game application program that can be supported, it can be displayed as a function control in the main interface or a certain sub-interface of the application program. When detecting a trigger operation on the function control, the game map corresponding to the game can be called, and the target trajectory can be generated based on the game map and the collected audio information. Alternatively, when a certain condition is detected during the game process, the scheme provided by the embodiment of the present disclosure can be automatically called to form a target trajectory.
[0055] The specific implementation of generating a target trajectory can be referred to the detailed description of the following embodiments.
[0056] FIG. 1 is a flowchart of a video generation method according to an embodiment of the present disclosure. The video generation method according to the embodiment of the present disclosure can be executed by a client, a server, or a combination of the client and the server. Of course, the client and / or the server can be integrated with a video generation apparatus corresponding to the video generation method.
[0057] As shown in FIG. 1, the method according to the embodiment can include the following steps.
[0058] In S110, a first trajectory is determined based on collected audio information during display of a video screen.
[0059] The video screen can be a video screen collected after a special effect prop is triggered, or a game screen displayed after a game function control is triggered. That is, the video screen is a screen displayed after a triggering operation of a special effect prop or a game function provided by the embodiment of the present disclosure.
[0060] It should be noted that the video screen can include a lot of background content. The background content can be pre-set content. Alternatively, the background content can be flowers, trees, or display elements corresponding to a festival. Alternatively, the festival can be the Mid-Autumn Festival, and the display elements can be rabbit elements and moon elements, etc.
[0061] During display of the video screen, audio information can be collected in real time or periodically. Alternatively, the audio information can be collected when it is detected that the audio information exists. The audio information at this time is mainly audio emitted by an object in a non-video screen. However, when it is detected that the audio information exists, the mouth shape of the object in the video screen can be made consistent with the mouth shape of the collected audio information, so as to simulate an effect that the audio information is emitted by the object in the video screen, so as to achieve a realistic and appreciable screen effect.
[0062] After the audio information is collected, the audio information can be analyzed and processed to obtain a first trajectory corresponding to the audio information. The first trajectory at this time is a trajectory that has not been further rendered with special effects. The length of the first trajectory and the display height in the video screen are related to the time length and the pitch of the audio information.
[0063] It should be noted that the first trajectory is not displayed in the video screen at this time. The first trajectory can be processed again to obtain a target trajectory with an eye-catching effect and displayed in the video screen.
[0064] Specifically, after detecting the trigger special effect prop, a video picture can be displayed. During the display of the video picture, a microphone array of a terminal device to which the application belongs can be started to collect audio information based on the microphone array. Alternatively, if a game function control corresponding to the embodiment of the present disclosure is triggered in the application, a pre-set game picture (video picture) can be displayed. The user can play the game based on the video picture, and the audio information of the object in the non-video picture is acquired in real time during the game (display of the video picture). Of course, during a video live broadcast or a video call, if a function control corresponding to the scheme provided by the embodiment of the present disclosure is triggered, the audio information of the object of the terminal to which the application belongs can also be collected. After the audio information is collected, the audio information can be analyzed and processed to obtain the first track.
[0065] In the embodiment, during the display of the video picture, the method further includes: collecting the audio information of the first object when it is detected that a track generation condition is met, wherein the track generation condition includes: determining that the second object moves to a preset position based on a preset map, wherein the second object is a movable object displayed in the video picture; and triggering a track generation control in the video picture or an application software to which the video picture belongs.
[0066] The picture content displayed in the video picture can be content designed in a design stage, can be picture content collected based on a camera device, or can be a picture obtained by superimposing the content designed in the design stage and the picture collected by the camera device. A user of a terminal device to which an application software to which the video picture belongs corresponds can be taken as the first object. For example, if the function control in the application is triggered by “me” to generate the target track, “me” is the first object. That is, the first object is an object that emits audio information in a real environment and can be collected by the audio information of the terminal to which the application belongs. The second object is a movable object displayed in the video picture. The second object can be a pre-set object or an object virtually created based on the image of the first object.
[0067] In the embodiment, if it is a game scene, a game map can be created. The game map can be taken as the preset map. The game map can include picture elements such as mountains, streams, and gullies that need to be crossed by the second object. When the elements appear, the target track can be generated to cross the elements in the picture based on the target track to achieve the effect of moving from one point to another. Meanwhile, the coordinate information of the elements can be marked in the game map. When the second object moves to a preset position corresponding to the elements at a preset moving speed, the audio information of the first object can be collected. That is, the preset position is the starting position of the appearance of the elements.
[0068] If it is a non-game scene, such as in a live scene, a special effect video generation scene, or a video call scene. The video picture can display a track generation control, or the application software to which the video picture belongs can display a track generation control. When the track generation control is detected, the audio information of the first object can be collected. It can also be that the audio information is collected all the time, and when there is a track generation requirement, the audio information collected at the current time and after the current time is analyzed and processed to obtain the first track.
[0069] The above can be understood as: if the track generation condition is met, the audio information of the first object can be collected.
[0070] S120, determining a first special effect track associated with the audio information according to the first track and a first material image set in advance.
[0071] In order to improve the picture quality, the obtained first track can be further processed to obtain the first special effect track. That is, the first special effect track is the track obtained after the first track is rendered.
[0072] The first material image is used to adjust the display information of each pixel point in the first track. A bottom picture with gradually changing color and / or transparency can be made as the first material image. The size of the first material image is consistent with the size of the video picture. After obtaining the first track, the first material image can be traversed according to the horizontal coordinates of each point on the first track to obtain the display information of each track point on the first track. The first track with updated display information is taken as the first special effect track.
[0073] It can be understood that a first material image with gradually changing color and / or transparency can be made according to actual needs. The pixel information in the first material image is traversed according to the horizontal coordinates of each point in the first track to determine and update the pixel information of each point in the first track to obtain the first special effect track.
[0074] S130, determining a second special effect track of the first track based on a target mask image associated with the first track and a second material image.
[0075] The initial track corresponding to the first track can be created, and the initial mask image can be dilated to obtain the target mask image. It can be understood that the shape of the target mask image is related to the first track, and the width is greater than the width of the first track. The second material image is a material image set in advance to match the target effect. The second material image can be an image simulating a certain phenomenon, and optionally, the certain phenomenon can be an aurora phenomenon, that is, the second material image is an image simulating an aurora. The second special effect track is the track obtained after processing the mask region associated with the first track in the target mask image.
[0076] Specifically, after obtaining the target mask image, the display information of the pixel points corresponding to the mask region in the target mask image can be adjusted according to the pixel values of the pixel points in the second material image. After filling the mask region of the target mask image with the completed pixel points, the second special effect track can be obtained by extracting the mask region.
[0077] S140, based on the first special effect track and the second special effect track, determining and displaying a target track displayed in the video picture; wherein the track length of the target track is associated with the audio information.
[0078] It can be understood that, after obtaining the first special effect track and the second special effect track, the first special effect track and the second special effect track can be fused together to obtain the target track. The target track is displayed in the video picture.
[0079] It should be noted that the starting point of the display of the target track can be pre-set, and can also be a position related to an element in the video picture. Optionally, if the element is a stream, the starting point of the display can be one bank of the stream to which the second object currently belongs; if the element is a gully, the starting point of the display can be any position on one side of the gully to which the second object currently belongs. The final display end point (i.e., the pre-set constraint position) of the target track can be pre-set, and can be the right edge of the video picture, or the ground on the other side of the mountain peak, the ground on the other side of the stream, or the ground on the other side of the gully. However, the real-time display position of the track point in the target track in the video picture is dynamically changed, and the main factor of the change is determined according to the pitch of the audio information. It can be understood that the change information of the target track is related to the audio information.
[0080] The technical solution provided by the embodiments of the present disclosure can determine the first track by analyzing and processing the collected audio information during the display of the video picture. Next, the first track is rendered according to the first track and the first material image to obtain the first special effect track. In order to further improve the display effect of the picture, the target mask image and the second material image corresponding to the first track can be processed again to obtain the second special effect track. Based on the first special effect track and the second special effect track, the target track corresponding to the audio information can be determined and displayed in the video picture, and the richness and naturalness of the picture content are realized.
[0081] FIG. 2 is a flowchart of a video generation method provided by an embodiment of the present disclosure. On the basis of the foregoing embodiment, the “determining the first track based on the collected audio information” can be further refined, and the specific implementation can be referred to the detailed description of the present technical solution. Among them, the same or corresponding technical terms as the above embodiments are not described again in the present embodiment.
[0082] As shown in FIG. 2, the method comprises:
[0083] S210, determining at least one trajectory control point according to the pitch information of the valid audio frame in the audio information.
[0084] The audio information is composed of a plurality of audio frames, and there are 25 audio frames per second. For some audio frames, they can be invalid. Based on this, the audio frames can be analyzed for validity to determine the first trajectory based on the valid audio frames. The trajectory control point is used to adjust the direction of the first trajectory. The direction can be understood as the shape of the trajectory.
[0085] Specifically, the audio information can be analyzed and processed to determine the valid audio frames in the audio information, and then the trajectory control point is determined according to the valid audio frames.
[0086] In this embodiment, the way to determine the valid audio frames in the audio information can be: after extracting the audio frames in the audio information, the volume and the pitch of the audio frames can be obtained. If the volume is greater than the volume in the first volume condition, and the pitch meets the pitch in the first pitch condition, the audio frame can be regarded as a valid audio frame.
[0087] For example, the first volume condition is that the volume is greater than 0.002, and the first pitch condition is that the pitch is greater than 0. If the volume in the audio frame is greater than 0.002 and the pitch is greater than 0, it means that the audio frame is a valid audio frame.
[0088] In this embodiment, the way to determine the at least one trajectory control point according to the pitch information of the valid audio frame in the audio information can be: determining the pitch offset according to the pitch information of the valid audio frame in the audio information and the pitch control information, adjusting the pitch information of the valid audio frame according to the pitch offset to obtain the to-be-used pitch information; determining the target screen coordinates of the to-be-used pitch information in the video picture, and determining the trajectory control point according to the target screen coordinates and the historical screen coordinates in the coordinate set.
[0089] Generally, in a case that the first object is of a first type, optionally, the first type is a male type with a lower tone, then the display height of the first track in the video picture approaches the bottom of the picture; in a case that the first object is of a second type, optionally, the second type is a female with a higher tone, then the display height of the first track in the video picture approaches the top of the picture. That is, the first type is an object type with a lower tone, and the second type is an object type with a higher tone, before use, it can be pre-set which objects are of the first type and which objects are of the second type, and the tone control information is determined based on this. In order to avoid the problem that the first track effect is not good due to the above object type, the tone control information can be generated based on the first tone corresponding to the first type and the second tone corresponding to the second type, so as to control the tone of any object type based on the tone control information, thereby improving the harmony of the target track in the video picture.
[0090] In the embodiment, the tone control information is determined according to the first tone of the first type and the second tone of the second type. Optionally, the mean value of the first tone and the second tone can be calculated, and the mean value is taken as the tone control information. The advantage of this setting is that the influence of high and low tones on the first track can be neutralized. That is, the tone control information is used to control the tone information to fluctuate within a preset range.
[0091] The tone offset is offset information obtained after the tone control information is used to constrain the sound effect information of the effective audio frame. The tone information of the effective audio frame can be adjusted based on the tone offset, which is equivalent to normalizing all effective audio frames based on the tone offset. The tone information adjusted based on the tone offset is taken as the tone information to be used.
[0092] The target screen coordinates are coordinate information corresponding to the video picture after the tone information to be used is mapped to the video picture. According to the target screen coordinates and the historical screen coordinates determined in the coordinate set, the track control point for adjusting the track of the Bezier curve is determined. The track control point is mainly used to adjust the track of the first track. Optionally, the track can include the extension angle of the first track, and the extension angle can be an angle corresponding to the ground of the video picture.
[0093] The target screen coordinates of the effective audio processed before the current effective audio are taken as the historical screen coordinates, that is, the target screen coordinates of the effective audio frames before the current effective audio frame are mainly stored in the coordinate set.
[0094] It should be noted that the scheme provided in the embodiment can be processed with the first effective audio frame. Of course, in order to improve the processing accuracy and the transition effect of the picture, the effective audio frames are analyzed and processed only when the number of effective audio frames reaches a first threshold, so as to start generating the track.
[0095] Specifically, the pitch offset is determined according to the pitch information and the pitch control information of the valid audio frame in the audio information. Then, the pitch information of the valid audio is adjusted according to the pitch offset to obtain the to-be-used pitch information. Finally, the trajectory control point can be determined according to the target screen coordinates corresponding to the to-be-used pitch information in the video picture.
[0096] It should be noted that the processing manner of each valid audio frame is the same, and here, the analysis and processing of one valid audio frame is mainly taken as an example for illustration. That is, for each valid audio frame, the to-be-used pitch information corresponding to the valid audio frame is dynamically determined.
[0097] For example, the first type is a male type, and the first pitch is a male lowest pitch lowPitch. The second type is a female type, and the second pitch is a female high pitch highPitch. The pitch control information can be the average of the first pitch and the second pitch, midPitch = (lowPitch + highPitch) / 2, where midPitch is the pitch control information.
[0098] It should be further noted that the pitch control information can also be determined according to the pre-set pitch weight of the first type and the first pitch, and the pitch weight of the second type and the second pitch.
[0099] S220, determining a first trajectory according to at least one trajectory control point and a preset Bezier function.
[0100] The preset Bezier function can be a function for generating a Bezier curve, and correspondingly, the first trajectory corresponds to the generated Bezier curve.
[0101] Specifically, the target screen coordinates corresponding to the at least one trajectory control point can be substituted into the Bezier function to obtain a Bezier function with determined coefficients. Next, the independent variable can be adjusted to generate a plurality of trajectory points, and all the trajectory points are connected to form a Bezier curve, that is, the final first trajectory.
[0102] The technical scheme provided by the embodiments of the present disclosure can determine the valid audio frame in the audio information, and determine at least one trajectory control point according to the pitch information of the valid audio frame. Then, the first trajectory is determined according to the trajectory control point and the preset Bezier function, which improves the effectiveness of determining the first trajectory.
[0103] FIG. 3 is a flowchart of determining a trajectory control point according to an embodiment of the present disclosure. The embodiment can be further refined based on the foregoing embodiment, and the specific implementation can be referred to the detailed description of the technical solution. The same or corresponding technical terms as the foregoing embodiment are not described herein.
[0104] As shown in FIG. 3, the method comprises:
[0105] S310, when the number of valid audio frames reaches a first threshold, determining initial pitch information according to the pitch information of the valid audio frames.
[0106] The first threshold is a natural number set in advance. The first threshold is mainly used to determine how many valid audio frames are needed as a reference to determine the initial pitch information. The first threshold is also used to determine from which valid audio frame to start determining the trajectory control point based on the valid audio frames. Optionally, the number of the first threshold can be 10, that is, if the number of valid audio frames reaches 10, the initial pitch information can be determined based on the 10 valid audio frames.
[0107] Specifically, after determining the valid audio frames in the audio information, the first trajectory can not be determined, but it can be used as a reference data to determine the initial pitch information.
[0108] For example, the first threshold is 10, and the pitch information is the pitch value. After determining 10 valid audio frames based on the audio information, the pitch values of the 10 valid audio frames can be averaged to obtain the initial pitch information (initial pitch value) initPitch.
[0109] S320, determining a pitch offset according to the initial pitch information and the pitch control information.
[0110] The initial pitch information includes an initial pitch value, and the pitch control information is a pitch control value.
[0111] Specifically, by calculating the difference between the initial pitch value and the pitch control value, the pitch offset corresponding to each valid audio frame can be obtained. The advantage of determining the pitch offset is that it is convenient for subsequent pitch control processing of the first type and the second type, so as to ensure the trajectory stability of the first trajectory relative to the video picture.
[0112] S330, starting from the valid audio frame after the first threshold, the pitch information of the valid audio frame and the pitch information of a preset number of historical valid audio frames before the valid audio frame are averaged to determine the adjusted pitch information of the valid audio frame.
[0113] It should be noted that for each scene, the tonal offset is determined according to the tonal value of the valid audio frame corresponding to the first threshold when the number of valid audio frames reaches the first threshold. That is, the first track is generated from the valid audio frame determined after the first threshold. For the processing mode of the valid audio frame generated after the first threshold, please refer to the detailed description below.
[0114] It should also be noted that for the valid audio frame generated after the first threshold, the following method is used to determine it. Herein, one of the valid audio frames is taken as an example to illustrate, which can be taken as the current valid audio frame.
[0115] The preset number is a value set in advance according to actual needs, which is mainly used to constrain how many historical valid audio frames need to be combined to constrain the current valid audio frame. Optionally, the preset number is 5 or 10 historical valid audio frames. The valid audio frames determined before the current valid audio frame can be called as the historical valid audio frames to be selected. Each historical valid audio frame has a corresponding historical timestamp, and the interval duration can be determined according to the historical timestamp and the current timestamp of the current valid audio frame. The preset number of historical valid audio frames to be selected can be obtained as the target historical valid audio frames of the current audio frame according to the interval duration. The tonal information to be used is based on the tonal information of the target historical valid audio frames.
[0116] For example, the first threshold is 10 valid audio frames, the preset value is 9 historical valid audio frames, and the current valid audio frame is the 11th valid audio frame. The tonal information of the 11th valid audio frame and the tonal information of the 2nd to 10th target historical valid audio frame are obtained. The tonal information of the 2nd to 11th valid audio frame is processed by averaging to obtain the adjusted tonal information avgPitch of the 11th valid audio frame.
[0117] That is, in order to improve the smoothness of the first track, the tonal information of the current valid audio frame and the nine target historical valid audio frames before it is processed by averaging to obtain the adjusted tonal information.
[0118] S340, according to the adjusted tonal information and the tonal offset, determine the tonal information to be used of the valid audio frame.
[0119] It can be understood that for the valid audio frame after the first threshold, the tonal information to be used of the valid audio frame can be obtained according to the adjusted tonal information and the tonal offset of the valid audio frame.
[0120] Exemplarily, after obtaining the to-be-adjusted pitch information avgPitch of the 11th valid audio frame, a sum of the to-be-adjusted pitch information avgPitch and the pitch offset pitchOffset can be calculated to obtain the to-be-used pitch information curPitch of the 11th valid audio frame.
[0121] S350, determining, according to the first pitch, the second pitch, and the to-be-used pitch information, a target screen coordinate in the video picture to which the pitch information of the valid audio frame is mapped.
[0122] The target screen coordinate is a coordinate obtained after the valid audio frame is mapped into the video picture.
[0123] Exemplarily, the video picture includes a plurality of pixel points, and each pixel point corresponds to a screen coordinate. Generally, a screen vertical coordinate is 0 to 1 from top to bottom of the video picture, and the pitch range is mainly from the first pitch lowPitch to the second pitch highPitch. Based on this, a mapping manner of the to-be-used pitch information to the target screen coordinate is: curHight=(curPitch-lowPitch) / (highPitch-lowPitch). curHight is a vertical coordinate of the target screen coordinate. A horizontal coordinate in the target screen coordinate is determined according to a pre-set trajectory moving speed.
[0124] S360, taking a newly-added historical screen coordinate in the coordinate set as a to-be-used screen coordinate, and determining a difference between a vertical coordinate of the to-be-used screen coordinate and a vertical coordinate of the target screen coordinate.
[0125] The coordinate set includes a plurality of historical screen coordinates. The historical screen coordinate newly added to the coordinate set can be taken as the to-be-used screen coordinate.
[0126] Specifically, after obtaining the target screen coordinate of the valid audio frame, the historical screen coordinate newly added to the coordinate set can be obtained. The difference between the vertical coordinates of the target screen coordinate and the historical screen coordinate can be calculated.
[0127] S370, updating the coordinate set according to the difference, to determine the at least one trajectory control point based on the updated coordinate set.
[0128] The number of the at least one trajectory control point can be one or more, and the specific number is related to the order of the Bezier function.
[0129] Specifically, according to the difference, it can be determined how to process the target screen coordinate to update the coordinate set based on the processed target screen coordinate. After the coordinate set is updated, the at least one trajectory control point can be determined.
[0130] In the embodiment, the coordinate set is updated according to the difference value, to determine the at least one trajectory control point based on the updated coordinate set, including: according to a preset condition met by the difference value, calling a coordinate processing mode corresponding to the preset condition to process the target screen coordinate, and updating the processed target screen coordinate to the coordinate set as a historical screen coordinate; when the number of historical screen coordinates in the coordinate set reaches a preset number threshold, adding time information of the historical screen coordinates into the coordinate set to obtain a plurality of to-be-applied screen coordinates corresponding to the preset number threshold; and determining the at least one trajectory control point based on the plurality of to-be-applied screen coordinates.
[0131] It can be understood that, according to the preset condition met by the difference value, the coordinate processing mode corresponding to the preset condition can be called to process the target screen coordinate, and the processed target screen coordinate can be updated to the coordinate set as a historical screen coordinate.
[0132] In the embodiment, the preset condition includes at least three kinds, and correspondingly, the coordinate processing mode also includes at least three kinds. Next, each preset condition and the coordinate processing mode corresponding to the preset condition will be described in detail.
[0133] The first kind: the preset condition is that the difference value is less than a first difference threshold, and the coordinate processing mode is to keep the target screen coordinate unchanged and delete the to-be-used screen coordinate from the coordinate set.
[0134] It can be understood that, if the coordinate difference value is less than the first difference threshold, the target screen coordinate can be kept unchanged, and the to-be-used screen coordinate can be deleted from the coordinate set.
[0135] The second kind: the preset condition is that the difference value is greater than a second difference threshold, and the coordinate processing mode is to correct the target screen coordinate according to the second difference threshold.
[0136] It can be understood that, if the difference value is greater than the second difference threshold, the second difference threshold can be superimposed on the ordinate of the target screen coordinate to obtain the updated target screen coordinate.
[0137] The third kind: the preset condition is that the difference value is greater than the first difference threshold and less than the second difference threshold, and the coordinate processing mode is to keep the target screen coordinate unchanged.
[0138] It can be understood that, if the difference value is greater than the first difference threshold and less than the second difference threshold, the target screen coordinate can be kept unchanged.
[0139] After the coordinate set is updated, a plurality of to-be-applied screen coordinates consistent with the preset quantity threshold can be determined according to the quantity of historical screen coordinates in the coordinate set and the time information of the historical screen coordinates added to the coordinate set. Through analysis and processing of the to-be-applied screen coordinates, at least one trajectory point can be determined. Optionally, a difference between the time information of the historical screen coordinates added to the coordinate set and the current time is calculated. According to the difference from small to large, a preset quantity of historical screen coordinates is selected as the to-be-applied screen coordinates.
[0140] The technical solution provided by the embodiments of the present disclosure can process the pitch information and the pitch offset of the effective audio frame to obtain the to-be-used pitch information of the effective audio frame, and then determine the trajectory control point according to the target screen coordinates corresponding to the to-be-used pitch information and the historical screen coordinates in the coordinate set, thereby improving the accuracy of determining the trajectory control point. Furthermore, in the case that the accuracy of the trajectory control point is high, the smoothness of the first trajectory can be improved when the first trajectory is determined according to the trajectory control point, thereby improving the picture display effect.
[0141] FIG. 4 is a flowchart of generating a trajectory control point provided by the embodiments of the present disclosure. On the basis of the foregoing embodiments, the “determining the at least one trajectory control point based on the plurality of to-be-applied screen coordinates” and “determining the first trajectory according to the at least one trajectory control point and a preset Bezier function” can be further refined. For specific implementation, reference can be made to the detailed description of the present technical solution. Among them, the same or corresponding technical terms as the above embodiments are not described again in the present embodiment.
[0142] As shown in FIG. 4, the method comprises:
[0143] S410, according to the time order of the plurality of to-be-applied screen coordinates added to the coordinate set, starting from the second to-be-applied screen coordinate, determining the slope value between the adjacent two to-be-applied screen coordinates.
[0144] It should be noted that the preset Bezier function adopts a second-order Bezier function, and the quantity of the plurality of to-be-applied screen coordinates can be four, and correspondingly, the quantity of the finally determined control points is three.
[0145] Wherein, the time of each to-be-applied screen coordinate added to the coordinate set is different, and the slope value between the adjacent two to-be-applied screen coordinates can be determined starting from the second to-be-applied screen coordinate according to the time order of adding to the coordinate set. For example, the points corresponding to the to-be-applied screen coordinates are marked as P1, P2, P3, and P4, and according to the order of adding to the coordinate set from early to late, P1 is earlier than P2, P2 is earlier than P3, and P3 is earlier than P4. Starting from P2, the slope K1 between P2 and P3 and the slope K2 between P3 and P4 can be calculated.
[0146] S420, determining a trajectory control point from the plurality of to-be-applied screen coordinates according to the slope values.
[0147] Specifically, after obtaining the two slope values, the relationship between the two slope values can be analyzed, and the trajectory control point can be determined from the plurality of to-be-applied screen coordinates according to the analysis result.
[0148] In this embodiment, the manner of determining the trajectory control point according to the slope values can be: if the difference between the slope values is greater than a slope difference threshold, taking the to-be-applied screen coordinates other than the to-be-applied screen coordinates at the last time sequence as the trajectory control point; if the difference between the slope values is less than the slope difference threshold, taking the to-be-applied screen coordinates other than the second to-be-applied screen coordinates as the trajectory control point.
[0149] It can be understood that if the difference between the slope values is greater than the slope difference threshold, the to-be-applied screen coordinates of P1, P2 and P3 can be taken as the trajectory control point. If the difference between the slope values is less than or equal to the slope difference threshold, P1, P2 and P4 can be taken as the trajectory control point. The trajectory control point is mainly used for subsequent processing of the preset Bezier function to generate the first trajectory.
[0150] S430, substituting the target screen coordinates of the trajectory control point into the preset Bezier function, adjusting the independent variable of the preset Bezier function to obtain a trajectory point of the effective audio frame in the video picture, and generating the first trajectory according to the trajectory point.
[0151] The preset Bezier function can be S t =(1-t) 2 S0+2t(1-t)S1+t 2 S2 t∈(0,1); S0 is the first trajectory control point in time sequence, S1 is the second trajectory control point in time sequence, S2 is the third trajectory control point in time sequence, and t is the independent variable.
[0152] Taking P1, P2 and P3 as the trajectory control point as an example, P1 corresponds to S0, P2 corresponds to S1, and P3 corresponds to S2. The target screen coordinates corresponding to the trajectory control point are substituted into the preset Bezier function. Next, the value of the independent variable t is adjusted to obtain a plurality of trajectory points. All the trajectory points are connected to obtain the first trajectory.
[0153] That is, the independent variable t≥0 and t<0.5 can be substituted into the preset Bezier function to obtain the trajectory point, and the trajectory point is displayed in the video picture. The splicing of all the trajectory points is the first trajectory.
[0154] It should be noted that the coordinate information of the trajectory point obtained by substituting t=0.5 into the preset Bezier function can be used to update the second to-be-used screen coordinate. Alternatively, the to-be-used screen coordinate of the P2 trajectory control point is updated to the coordinate information obtained when t=0.5.
[0155] That is, the coordinate information of the trajectory point corresponding to the target independent variable is used as the second to-be-applied screen coordinate and is updated to the coordinate set.
[0156] In this embodiment, for the valid audio frames after the first threshold, the trajectory points can be determined by using the above-mentioned manner. The valid audio frames can be continuously analyzed and processed to update the length of the first trajectory.
[0157] The technical solution provided by the embodiments of the present disclosure can further determine the first trajectory in combination with the determined trajectory control point and the preset Bezier function, improve the smoothness of the first trajectory, and thus improve the display effect of the picture.
[0158] FIG. 5 is a flowchart of a video generation method provided by an embodiment of the present disclosure. On the basis of the foregoing embodiments, the step of “determining a first special effect trajectory associated with the audio information according to the first trajectory and a first material image set in advance” can be further refined. The specific implementation can be referred to the detailed description of the technical solution. The same or corresponding technical terms as those in the foregoing embodiments are not described herein.
[0159] As shown in FIG. 5, the method comprises the following steps.
[0160] S510, determining a first trajectory based on the collected audio information during the display of the video picture.
[0161] S520, updating the display information of each trajectory point in the first trajectory according to the coordinate information of each trajectory point in the first trajectory and the pixel information of each pixel point in the first material image, to obtain a first special effect trajectory.
[0162] [According to the rules 91 correction 29.10.2025] Wherein, the first material image is a material image according to preset effect setting, optionally, the first material image refers to FIG. 6, a bottom picture with color / transparency gradient can be made. For example, a first color is set at a first end of the first material image, and a second color is set at a second end, and the color information of each pixel point in the first material image is obtained by interpolating the first color and the second color. It should be noted that FIG. 6 is only a schematic illustration of the color / transparency gradient; for example, the first end can be set to correspond to the first color, which can be blue, for example; the second end can be set to correspond to the second color, which can be pink, for example; from the first end to the second end, the color can be interpolated to gradually change from the first color to the second color, thereby obtaining the color information of each pixel point in the first material image. Further, transparency information can also be set. By adjusting the transparency of each pixel point in the first material image, an updated first material image is obtained.
[0163] After obtaining the first trajectory, the first material image can be traversed based on the abscissa of the trajectory points in the first trajectory to determine the pixel points in the first material image corresponding to the trajectory points, and the display information of the trajectory points is updated according to the pixel values of the pixel points, and the first trajectory with updated display information is taken as the first special effect trajectory. The schematic diagram of the first special effect trajectory processed based on FIG. 6 can be seen in FIG. 7.
[0164] [According to the rules 91 correction 29.10.2025] It should be noted that the colors and transparencies in FIGS. 6 and 7 are only exemplary illustrations, and the image content of the first material image can be set according to actual needs during the design phase, and the specific content is not limited in this embodiment.
[0165] The rendering map is queried based on the position information of each trajectory point in the first trajectory to update the display information of the trajectory points according to the transparency and real color in the rendering map. Wherein, the size of the first material image is consistent with the size of the video picture, and the display information of any two pixel points of the first material image is different.
[0166] In this embodiment, in order to further improve the display effect of the first special effect trajectory, the obtained first special effect trajectory can be further processed, and optionally, the method further comprises: Gaussian blur processing the first special effect trajectory to obtain a to-be-mixed special effect trajectory; and fusing the to-be-mixed special effect trajectory and the first special effect trajectory before Gaussian blur processing to obtain the first special effect trajectory with updated display information.
[0167] It can be understood that the first special effect track can be Gaussian blurred to obtain a to-be-mixed special effect track. Next, the to-be-mixed special effect track and the first special effect track before Gaussian blur processing are mixed to obtain the first special effect track with updated display information.
[0168] For example, after obtaining the first special effect track as shown in FIG. 7, the first special effect track is blurred by using a Gaussian blur algorithm to obtain an effect diagram as shown in FIG. 8. The to-be-mixed special effect track and the first special effect track before Gaussian blur processing are pixel superimposed to update the display information of each pixel point in the first special effect track.
[0169] In S530, a target mask diagram is determined according to the first track.
[0170] The shape of the track in the target mask diagram is related to the shape of the first track, and there is an image after inflation processing.
[0171] In this embodiment, the manner of determining the target mask diagram can be: obtaining an initial track corresponding to the first track; offsetting pixel points in the initial track by a preset number of pixels along a first direction to obtain a to-be-disturbed track; performing disturbance processing on the to-be-disturbed track according to a pre-set disturbance pixel range to obtain the target disturbance track; and taking a region formed by the initial track and the target disturbance track as a mask region corresponding to the first track in the target mask diagram.
[0172] How to determine the target mask region can be understood in combination with FIG. 9. A track corresponding to the first track can be copied as an initial track. The initial track is offset by a preset number of pixels along a first direction, which is optional. The preset number of pixels is 60px. The pixel points on the initial track are moved by 60px along the first direction perpendicular to the horizontal plane to obtain an effect diagram as shown in (a) of FIG. 9. The track obtained at this time is taken as the to-be-disturbed track. The disturbance pixel range can be a pixel fluctuation range set according to actual needs. Optionally, the pixel fluctuation range can be randomly fluctuated by 0px-25px along the first direction or the second direction. Each pixel point in the to-be-disturbed track can be randomly disturbed to obtain an effect diagram as shown in (b) of FIG. 9. The track corresponding to the red curve is taken as the target disturbance track after disturbance. The region formed by the initial track and the target disturbance track is taken as a mask region corresponding to the first track in the target mask diagram, as shown in (c) of FIG. 9.
[0173] On the basis of the above technical solutions, the second special effect track can be generated based on the obtained target mask diagram. Further, the target mask diagram can be further processed to improve the display effect of the generated second special effect track.
[0174] Optionally, Gaussian blur processing is performed on (c) of FIG. 9 to obtain an effect schematic diagram as shown in FIG. 10. It can be understood that the target mask diagram is subjected to Gaussian blur processing to obtain an updated target mask diagram.
[0175] In S540, a second special effect track of the first track is determined based on the target mask diagram and the second material image.
[0176] The second material image is also a material image made according to a preset effect. Optionally, the second material can be a material image simulating an aurora scene. The material image shown in FIG. 11 can be the second material image.
[0177] It should be noted that the first material image and the second material image can be changed according to actual needs. In the present embodiment, only the material images shown in the drawings are used for exemplary description, and the present scheme is not limited. The core of the technical scheme provided in the present embodiment is that the special effect track can be generated by using the technology.
[0178] Specifically, the second special effect track corresponding to the first track can be obtained by processing the target mask diagram and the second material image.
[0179] In the present embodiment, the way of generating the second special effect track according to the target mask diagram and the second material image can be: updating pixel information of at least one pixel point in the target mask diagram in pixel information of the at least one pixel point in the second material image to obtain the second special effect track.
[0180] Specifically, the pixel point to be used corresponding to each pixel point in the target mask diagram in the second material image can be determined, and the pixel information of the pixel point to be used is updated to the pixel point of the target mask diagram. After processing each pixel point in the target mask diagram, the second special effect track can be obtained, as shown in FIG. 12.
[0181] In S550, a target track displayed in the video picture is determined based on the first special effect track and the second special effect track and displayed.
[0182] In the present embodiment, the endpoints of the first special effect track and the second special effect track can be kept consistent and fused to obtain the target track displayed in the video picture.
[0183] On the basis of the above technical scheme, if the endpoint of the target track is a preset screen coordinate, the track length of the target track is no longer updated; or, if no audio information is collected within a preset time length, the track length of the target track is no longer updated.
[0184] It can be understood that if the terminal of the target trajectory is a preset screen coordinate, optionally, the preset screen coordinate can be a screen coordinate corresponding to the ground on the other side of the hill, or a preset screen coordinate corresponding to the ground at the other end of the ditch, the length of the target estimated trajectory can not be updated, that is, the collected audio can not be continuously processed. It can also be that the audio information is continuously analyzed and processed to generate the target trajectory. If no audio information is collected within a preset time, optionally, 20S, the length of the target trajectory can not be updated. Of course, if 20S audio information is collected, the length of the target trajectory can be continuously updated. On the basis of the above technical solutions, the target trajectory can further include a sub-line. The first special effect trajectory can be copied multiple times to obtain a to-be-fused trajectory, and the thickness value of the to-be-fused trajectory can be adjusted to obtain the to-be-fused trajectory as shown in FIG. 13.
[0185] By fusing the to-be-fused trajectory, the first special effect trajectory and the second special effect trajectory together, and displaying them in the video picture, the obtained video picture can be seen from FIG. 14.
[0186] The technical solutions provided by the embodiments of the present disclosure can rely on the first material image to obtain the first special effect trajectory by first special effect processing, and rely on the second material image to obtain the second special effect trajectory by special effect processing on the mask image corresponding to the first trajectory. Based on the first special effect trajectory and the second special effect trajectory, the target trajectory displayed in the video picture can be determined, which not only displays the target trajectory corresponding to the audio information, the target trajectory being obtained after multiple special effect processing, but also improves the richness of the content of the video picture and the interactivity between the video picture and the user.
[0187] FIG. 15 is a structural schematic diagram of a video generation device provided by an embodiment of the present disclosure, as shown in FIG. 15, the device includes a first trajectory determination module 610, a first special effect trajectory generation module 620, a second special effect trajectory generation module 630 and a target trajectory generation module 640.
[0188] The first trajectory determination module 610 is configured to determine a first trajectory based on collected audio information during display of a video picture; the first special effect trajectory generation module 620 is configured to determine a first special effect trajectory associated with the audio information according to the first trajectory and a pre-set first material image; wherein display information corresponding to the first special effect trajectory is related to pixel information of the first material image; the second special effect trajectory generation module 630 is configured to determine a second special effect trajectory of the first trajectory based on a target mask image associated with the first trajectory and a second material image; and the target trajectory generation module 640 is configured to determine a target trajectory displayed in the video picture based on the first special effect trajectory and the second special effect trajectory and display the target trajectory; wherein change information of the target trajectory is associated with the audio information.
[0189] The technical scheme of the embodiments of the present disclosure can analyze and process the collected audio information to determine the first track in the process of video picture display. Further, the first track can be processed with special effects by using the first material image to obtain the first special effect track, which enriches the picture effect of the first track. In order to further improve the richness and appreciability of the video picture content, the mask image corresponding to the first track can be rendered based on the second material image to obtain the second special effect track. The target track displayed in the video picture can be obtained based on the first special effect track and the second special effect track, which solves the problem that the use range of audio information is narrow and the interactivity with the terminal device is low when only the audio information is converted into text, and realizes the generation of the corresponding first track based on the audio information, and the fusion display of the first track after special effect processing from multiple dimensions, thereby improving the richness of the video picture content and the interactivity with the user.
[0190] On the basis of the above technical scheme, the device comprises an audio acquisition module configured to acquire audio information of a first object when it is detected that a track generation condition is met.
[0191] The track generation condition comprises at least one of the following:
[0192] The second object is a movable object displayed in the video picture.
[0193] The track generation control is triggered in the video picture or the application software to which the video picture belongs.
[0194] On the basis of the above technical scheme, the first track determination module comprises:
[0195] A track control point determination unit configured to determine at least one track control point according to the pitch information of the valid audio frame in the audio information, wherein the valid audio frame is an audio frame whose volume meets a first volume condition and whose pitch meets a first pitch condition, and the track control point is used to adjust the track shape of the first track.
[0196] A first track determination unit configured to determine the first track according to the at least one track control point and a preset Bezier function.
[0197] On the basis of the above technical scheme, the track control point determination unit comprises:
[0198] The tone offset amount determining sub-unit is configured to determine a tone offset amount according to tone information of valid audio frames in the audio information and tone control information, wherein the tone control information is determined according to a first tone of a first type and a second tone of a second type, and the tone control information is used to control the tone information to fluctuate within a preset range.
[0199] The to-be-used tone information determining sub-unit is configured to adjust the tone information of the valid audio frames according to the tone offset amount to obtain to-be-used tone information.
[0200] The trajectory control point determining sub-unit is configured to determine target screen coordinates of the to-be-used tone information in the video picture, and determine a trajectory control point according to the target screen coordinates and historical screen coordinates in the coordinate set.
[0201] On the basis of the above technical solutions, the tone offset amount determining sub-unit comprises:
[0202] The initial tone determining sub-unit is configured to determine initial tone information according to tone information of the valid audio frames when a number of the valid audio frames reaches a first threshold.
[0203] The offset amount determining sub-unit is configured to determine a tone offset amount according to the initial tone information and the tone control information.
[0204] The to-be-used tone information determining sub-unit comprises:
[0205] The first information determining unit is configured to determine to-be-adjusted tone information of the valid audio frames according to the valid audio frames and target historical valid audio frames for the valid audio frames determined after the first threshold, wherein the target historical valid audio frames are determined according to all to-be-selected historical valid audio frames of the audio information, an interval duration of the valid audio frames and a preset number of the historical valid video frames.
[0206] The second tone determining unit is configured to determine to-be-used tone information of the valid audio frames according to the to-be-adjusted tone information and the tone offset amount.
[0207] On the basis of the above technical solutions, the trajectory control point determining sub-unit comprises:
[0208] The first coordinate determining sub-unit is configured to determine target screen coordinates of tone information of the valid audio frames in the video picture according to the first tone, the second tone and the to-be-used tone information.
[0209] The difference determination sub-unit is configured to determine a difference between a vertical coordinate of a to-be-used screen coordinate and a vertical coordinate of the target screen coordinate, where the to-be-used screen coordinate is a coordinate in the historical screen coordinates.
[0210] The trajectory point determination sub-unit is configured to update the coordinate set according to the difference, and determine the at least one trajectory control point based on the updated coordinate set.
[0211] On the basis of the above technical solutions, the trajectory point determination sub-unit comprises:
[0212] The coordinate set updating sub-unit is configured to process the target screen coordinate by using a coordinate processing mode corresponding to a preset condition according to the difference satisfying the preset condition, and update the processed target screen coordinate to the coordinate set as a historical screen coordinate.
[0213] The to-be-applied screen coordinate determination sub-unit is configured to obtain a plurality of to-be-applied screen coordinates corresponding to the preset number threshold according to time information of the historical screen coordinates in the coordinate set when a number of the historical screen coordinates in the coordinate set reaches a preset number threshold.
[0214] The trajectory point determination sub-unit is configured to determine the at least one trajectory control point based on the plurality of to-be-applied screen coordinates.
[0215] On the basis of the above technical solutions, the coordinate processing mode corresponding to the preset condition comprises at least one of the following:
[0216] The preset condition is that the difference is less than a first difference threshold, and the coordinate processing mode is to keep the target screen coordinate unchanged and delete the to-be-used screen coordinate from the coordinate set.
[0217] The preset condition is that the difference is greater than a second difference threshold, and the coordinate processing mode is to correct the target screen coordinate according to the second difference threshold.
[0218] The preset condition is that the difference is greater than a first difference threshold and less than a second difference threshold, and the coordinate processing mode is to keep the target screen coordinate unchanged.
[0219] On the basis of the above technical solutions, the trajectory point determination sub-unit is configured to
[0220] According to a time sequence of the plurality of to-be-applied screen coordinates added to the coordinate set, the slope value between two adjacent to-be-applied screen coordinates is determined from the second to-be-applied screen coordinate.
[0221] According to the slope values, a trajectory control point is determined from the plurality of to-be-applied screen coordinates.
[0222] On the basis of the above technical solutions, the trajectory point determination subunit is configured to
[0223] If the difference between the slope values is greater than a slope difference threshold value, a to-be-applied screen coordinate other than the to-be-applied screen coordinate that is chronologically last is taken as the trajectory control point.
[0224] If the difference between the slope values is less than the slope difference threshold value, a to-be-applied screen coordinate other than the second to-be-applied screen coordinate is taken as the trajectory control point.
[0225] On the basis of the above technical solutions, the first trajectory determination unit is further configured to
[0226] The target screen coordinate of the trajectory control point is substituted into the preset Bezier function, and the independent variable of the preset Bezier function is adjusted to obtain a trajectory point of the effective audio frame in the video picture; and the first trajectory is determined based on the trajectory point.
[0227] On the basis of the above technical solutions, the first special effect trajectory generation module is further configured to
[0228] According to the coordinate information of each trajectory point in the first trajectory and the pixel information of each pixel point in the first material image, display information of each trajectory point in the first trajectory is updated to obtain the first special effect trajectory; wherein the size of the first material image is consistent with the size of the video picture, and the display information of any two pixel points of the first material image is different.
[0229] On the basis of the above technical solutions, the device further comprises a special effect updating module configured to perform Gaussian blur processing on the first special effect trajectory to obtain a to-be-mixed special effect trajectory; and perform fusion processing on the to-be-mixed special effect trajectory and the first special effect trajectory before Gaussian blur processing to obtain the first special effect trajectory with updated display information.
[0230] On the basis of the above technical solutions, the device further comprises:
[0231] An initial trajectory corresponding to the first trajectory is obtained; pixel points in the initial trajectory are offset by a preset number of pixels along a first direction to obtain a to-be-disturbed trajectory; the to-be-disturbed trajectory is disturbed according to a pre-set disturbance pixel range to obtain the target disturbance trajectory; and a region formed by the initial trajectory and the target disturbance trajectory is taken as a mask region corresponding to the first trajectory in a target mask image.
[0232] On the basis of each of the technical solutions above, the second special effect track generation module is further configured to update pixel information of at least one pixel point in the target mask image according to pixel information of the at least one pixel point in the second material image, and obtain the second special effect track.
[0233] The video generation apparatus provided by the embodiments of the present disclosure can perform the video generation method provided by any of the embodiments of the present disclosure, and has the corresponding function modules and beneficial effects of performing the method.
[0234] It should be noted that each unit and module included in the apparatus is only divided according to the function logic, but is not limited to the above division, as long as the corresponding function can be implemented; in addition, the specific name of each functional unit is only for convenient distinction, and does not limit the protection scope of the embodiments of the present disclosure.
[0235] FIG. 16 is a structural schematic diagram of an electronic device provided by an embodiment of the present disclosure. Referring to FIG. 16, a structural schematic diagram of an electronic device (for example, a terminal device or a server in FIG. 16) 700 suitable for implementing the embodiments of the present disclosure is shown. The terminal device in the embodiments of the present disclosure can include, but is not limited to, a mobile terminal such as a mobile phone, a notebook computer, a digital broadcast receiver, a PDA (Personal Digital Assistant), a PAD (Tablet Personal Computer), a PMP (Portable Multimedia Player), a vehicle-mounted terminal (for example, a vehicle-mounted navigation terminal), and the like, and a fixed terminal such as a digital TV, a desktop computer, and the like. The electronic device shown in FIG. 16 is only an example, and should not bring any limitation to the functions and use range of the embodiments of the present disclosure.
[0236] As shown in FIG. 16, the electronic device 700 can include a processing device (for example, a central processing unit, a graphics processing unit, and the like) 701, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 702 or a program loaded from a storage device 708 into a random access memory (RAM) 703. In the RAM 703, various programs and data required for the operation of the electronic device 700 are also stored. The processing device 701, the ROM 702, and the RAM 703 are connected to each other through a bus 704. An edit / output (I / O) interface 705 is also connected to the bus 704.
[0237] In general, the following devices can be connected to the I / O interface 705: input devices 706 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, and the like; output devices 707 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, and the like; storage devices 708 including, for example, a magnetic tape, a hard disk, and the like; and communication devices 709. The communication devices 709 can allow the electronic device 700 to communicate wirelessly or wiredly with other devices to exchange data. Although FIG. 16 illustrates the electronic device 700 with various devices, it is understood that all of the illustrated devices are not required to be implemented or possessed. More or fewer devices can be alternatively implemented or possessed.
[0238] In particular, according to embodiments of the present disclosure, the processes described above with reference to the flowcharts can be implemented as a computer software program. For example, embodiments of the present disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for executing the methods illustrated by the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network through the communication devices 709, or installed from the storage devices 708, or installed from the ROM 702. When the computer program is executed by the processing devices 701, the above-described functions defined in the methods of the embodiments of the present disclosure are performed.
[0239] The names of the messages or information exchanged between the plurality of devices in the embodiments of the present disclosure are only for illustrative purposes, and are not intended to limit the scope of the messages or information.
[0240] The electronic device provided by the embodiments of the present disclosure and the video generation method provided by the above-described embodiments belong to the same inventive concept, and the technical details not described in detail in the embodiments of the present disclosure can be referred to the above-described embodiments, and the present embodiments have the same beneficial effects as the above-described embodiments.
[0241] The embodiments of the present disclosure provide a computer storage medium having a computer program stored thereon, which, when executed by a processor, implements the video generation method provided by the above-described embodiments.
[0242] It should be noted that the computer-readable medium described above can be a computer-readable signal medium or a computer-readable storage medium or any combination thereof. The computer-readable storage medium, for example, can be, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus or device, or any suitable combination of the foregoing. More specific examples of the computer-readable storage medium can include, but are not limited to, an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing. In the disclosure, the computer-readable storage medium can be any tangible medium that contains or stores a program used by or in connection with an instruction execution system, apparatus or device. In the disclosure, the computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, in which the computer-readable program code is contained. Such a propagated data signal can take any of a variety of forms, including, but not limited to, an electromagnetic signal, an optical signal, or any suitable combination of the foregoing. The computer-readable signal medium can also be any computer-readable medium that is not a storage medium and that can communicate, propagate or transport a program for use by or in connection with an instruction execution system, apparatus or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including, but not limited to, wire, cable, RF (radio frequency), etc., or any suitable combination of the foregoing.
[0243] In some embodiments, the client, server, or both can communicate using any current known or future developed network protocol, such as HTTP (HyperText Transfer Protocol), and can be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include local area networks ("LAN"), wide area networks ("WAN"), the Internet, and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any current known or future developed networks.
[0244] The computer-readable medium described above can be included in the electronic device; or exist separately from the electronic device, and not be assembled into the electronic device.
[0245] The computer-readable medium described above carries one or more programs, when the one or more programs are executed by the electronic device, cause the electronic device to:
[0246] In the process of displaying the video picture, based on the collected audio information, a first track is determined;
[0247] According to the first track and a first material image preset, a first special effect track associated with the audio information is determined; wherein the display information corresponding to the first special effect track is related to the pixel information of the first material image;
[0248] Based on a target mask image associated with the first track and a second material image, a second special effect track of the first track is determined;
[0249] Based on the first special effect track and the second special effect track, a target track displayed in the video picture is determined and displayed; wherein the change information of the target track is associated with the audio information.
[0250] Computer program code for carrying out operations of the present disclosure can be written in one or more programming languages or combinations of languages including object oriented programming languages such as Java, Smalltalk, C++ as well as conventional procedural programming languages such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider).
[0251] The flow diagrams and the block diagrams in the drawings are illustrations of architectures, functionalities, and operations of possible implementations of systems, methods, and computer program products according to various embodiments of present disclosure. In this regard, each block in the flow diagrams or block diagrams can represent a module, a procedure, or a portion of code, which comprises one or more executable instructions for implementing the specified logical functions. It should also be noted that in some alternative implementations, the functions noted in the blocks can occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently or in the reverse order, depending on the functionality involved. It will also be noted that each block in the block diagrams and / or flow diagrams, and combinations of blocks in the block diagrams and / or flow diagrams, can be implemented by special purpose hardware-based systems that perform the specified functions or operations, or combinations of special purpose hardware and computer instructions.
[0252] The units described in the embodiments of the present disclosure can be implemented in the form of software, or can be implemented in the form of hardware. In some cases, the names of the units do not constitute a limitation on the units themselves. For example, the first obtaining unit can also be described as a unit that obtains at least two Internet protocol addresses.
[0253] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, and without limitation, exemplary types of hardware logic components that can be used include: Field-programmable Gate Arrays (FPGAs), Application-specific Integrated Circuits (ASICs), Application-specific Standard Products (ASSPs), System-on-a-chip systems (SOCs), Complex Programmable Logic Devices (CPLDs), etc.
[0254] In the context of the present disclosure, a machine-readable medium can be a tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium will include one or more lines of electrical connections, portable computer disks, hard disk drives, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or Flash memory), optical fibers, portable compact disc read-only memories (CD-ROMs), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0255] The above description is merely the preferred embodiments of the present disclosure and the explanation of the principles of the applied technology. It should be understood by those skilled in the art that the disclosed scope of the present disclosure is not limited to the technical solutions formed by the specific combinations of the above technical features, and also covers other technical solutions formed by any combinations of the above technical features or equivalent features without departing from the disclosed concept. For example, the above technical features can be replaced with the technical features disclosed in the present disclosure (but not limited to) that have similar functions to form technical solutions.
[0256] Moreover, while operations are depicted in a particular order, this should not be understood as requiring such an order nor infringing on the scope of the disclosure. Certain of the operations described in the discussion are combinable into a single operation, and certain operations can be separated into several operations. In some embodiments, the operations described in the discussion can be performed in an order different than presented in the discussion. In some embodiments, the operations described in the discussion can be performed concurrently. Also, while several specific implementation details are discussed in the discussion, these should not be interpreted as limiting the scope of the disclosure. Rather, certain features described in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment can also be implemented in multiple embodiments separately or in any suitable sub-combination.
[0257] Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are disclosed as example forms of implementing the claims.
Claims
1. A video generation method, comprising: determining a first track based on collected audio information during a process of video screen display; determining a first special effect track associated with the audio information according to the first track and a first material image preset; wherein display information corresponding to the first special effect track is related to pixel information of the first material image; determining a second special effect track of the first track based on a target mask image associated with the first track and a second material image; determining and displaying a target track displayed in the video screen based on the first special effect track and the second special effect track; wherein change information of the target track is associated with the audio information.
2. The method of claim 1, wherein, During the process of video screen display, the method further comprises: collecting audio information of a first object when it is detected that a track generation condition is met; wherein the track generation condition comprises at least one of the following: determining that a second object moves to a preset position based on a preset map, wherein the second object is a movable object displayed in the video screen; triggering a track generation control in the video screen or an application software to which the video screen belongs.
3. The method of claim 1 or 2, wherein, The determining of the first track based on the collected audio information comprises: determining at least one track control point according to tone information of an effective audio frame in the audio information, wherein the effective audio frame is an audio frame whose volume meets a first volume condition and whose tone meets a first tone condition, and the track control point is used to adjust a track shape of the first track; determining the first track according to the at least one track control point and a preset Bezier function.
4. The method of claim 3, wherein, The determining of the at least one track control point according to the tone information of the effective audio frame in the audio information comprises: determining a tone offset according to the tone information of the effective audio frame in the audio information and tone control information, wherein the tone control information is determined according to a first tone of a first type and a second tone of a second type, and the tone control information is used to control the tone information to fluctuate within a preset range; adjusting the tone information of the effective audio frame according to the tone offset to obtain to-be-used tone information; determining a target screen coordinate of the to-be-used tone information in the video screen, and determining a track control point according to the target screen coordinate and historical screen coordinates in a coordinate set.
5. The method of claim 4, wherein, The determining of the tone offset according to the tone information of the effective audio frame in the audio information and the tone control information comprises: determining initial tone information according to the tone information of the effective audio frame when a number of the effective audio frames reaches a first threshold; determining a tone offset according to the initial tone information and the tone control information; The adjusting of the tone information of the effective audio frame according to the tone offset to obtain to-be-used tone information comprises: For the effective audio frame determined after the first threshold, according to the effective audio frame and a target historical effective audio frame, determine to-be-adjusted pitch information of the effective audio frame; wherein, the target historical effective audio frame is determined according to all to-be-selected historical effective audio frames of the audio information, an interval duration of the effective audio frame and a preset number of the historical effective video frame; According to the to-be-adjusted pitch information and the pitch offset, determine to-be-used pitch information of the effective audio frame.
6. The method according to claim 4 or 5, characterized in that, The determination of the to-be-used pitch information in the target screen coordinates in the video picture, and the determination of the track control point according to the target screen coordinates and the historical screen coordinates in the coordinate set, include: According to the first pitch, the second pitch and the to-be-used pitch information, determine the target screen coordinates in the video picture where the pitch information of the effective audio frame is mapped to; Add the newly added historical screen coordinates in the coordinate set as to-be-used screen coordinates, and determine the difference between the vertical coordinates of the to-be-used screen coordinates and the target screen coordinates; wherein, the to-be-used screen coordinates are coordinates in the historical screen coordinates; According to the difference, update the coordinate set to determine the at least one track control point based on the updated coordinate set.
7. The method of claim 6, wherein, The updating of the coordinate set according to the difference to determine the at least one track control point based on the updated coordinate set, includes: According to a preset condition met by the difference, call a coordinate processing mode corresponding to the preset condition to process the target screen coordinates, and update the processed target screen coordinates to the coordinate set as historical screen coordinates; When the number of historical screen coordinates in the coordinate set reaches a preset number threshold, add time information of the historical screen coordinates in the coordinate set to obtain a plurality of to-be-applied screen coordinates corresponding to the preset number threshold; Determine the at least one track control point based on the plurality of to-be-applied screen coordinates.
8. The method of claim 7, wherein, The coordinate processing mode corresponding to the preset condition includes at least one of the following: The preset condition is that the difference is less than a first difference threshold, and the coordinate processing mode is to keep the target screen coordinates unchanged and delete the to-be-used screen coordinates from the coordinate set; The preset condition is that the difference is greater than a second difference threshold, and the coordinate processing mode is to correct the target screen coordinates according to the second difference threshold; The preset condition is that the difference is greater than a first difference threshold and less than a second difference threshold, and the coordinate processing mode is to keep the target screen coordinates unchanged.
9. The method of claim 7 or 8, wherein, The determination of the at least one track control point based on the plurality of to-be-applied screen coordinates, includes: According to the time sequence of adding the plurality of to-be-applied screen coordinates to the coordinate set, start from the second to-be-applied screen coordinates to determine the slope value between adjacent two to-be-applied screen coordinates; According to the slope value, determine a track control point from the plurality of to-be-applied screen coordinates.
10. The method of claim 9, wherein, The determination of the track control point from the plurality of to-be-applied screen coordinates according to the slope value, includes: if the difference between the slope values is greater than a slope difference threshold, then a to-be-applied screen coordinate outside the to-be-applied screen coordinates that are chronologically last is taken as the track control point; if the difference between the slope values is less than the slope difference threshold, then a to-be-applied screen coordinate outside the second to-be-applied screen coordinate is taken as the track control point.
11. The method according to any one of claims 3-10, wherein, The determining the first track according to the at least one track control point and a preset Bezier function comprises: substituting target screen coordinates of the track control point into the preset Bezier function, and obtaining track points of the effective audio frame in the video picture by adjusting an independent variable of the preset Bezier function; determining the first track based on the track points.
12. The method of any one of claims 1-11, wherein, The determining the first special effect track associated with the audio information according to the first track and a first material image pre-set comprises: updating display information of each track point in the first track according to coordinate information of each track point in the first track and pixel information of each pixel point in the first material image, to obtain the first special effect track; wherein a size of the first material image is consistent with a size of the video picture, and display information of any two pixel points of the first material image is different.
13. The method of any one of claims 1-12, wherein, After obtaining the first special effect track, the method further comprises: performing Gaussian blur processing on the first special effect track to obtain a to-be-mixed special effect track; performing fusion processing on the to-be-mixed special effect track and the first special effect track before Gaussian blur processing to obtain the first special effect track with updated display information.
14. The method of any one of claims 1-13, further comprising: obtaining an initial track corresponding to the first track; offsetting pixel points in the initial track by a preset number of pixels in a first direction to obtain a to-be-disturbed track; performing disturbance processing on the to-be-disturbed track according to a pre-set disturbance pixel range to obtain the target disturbed track; taking a region formed by the initial track and the target disturbed track as a mask region corresponding to the first track in a target mask image.
15. A video generation apparatus, comprising: a first track determination module configured to determine a first track based on collected audio information during display of a video picture; a first special effect track generation module configured to determine a first special effect track associated with the audio information according to the first track and a first material image pre-set; wherein display information corresponding to the first special effect track is related to pixel information of the first material image; a second special effect track generation module configured to determine a second special effect track of the first track based on a target mask image associated with the first track and a second material image; a target track generation module configured to determine and display a target track displayed in the video picture based on the first special effect track and the second special effect track; wherein change information of the target track is associated with the audio information.
16. An electronic device, comprising: one or more processors; a storage apparatus configured to store one or more programs, wherein When the one or more programs are executed by the one or more processors, the one or more processors implement a video generation method as claimed in any of claims 1-14.
17. A storage medium containing computer-executable instructions, wherein, The computer executable instructions, when executed by a computer processor, perform a video generation method as claimed in any of claims 1-14.
18. A computer program product comprising a computer program, wherein, The computer program, when executed by a processor, implements a video generation method as claimed in any of claims 1-14.
Citation Information
Patent Citations
Audio visualization method and terminal
CN112667828A
Method and device for determining special effect video, electronic equipment and storage medium
CN114630057A
Method and device for generating sticking point special effect video and storage medium
CN115174823A
Special effect display method and device, equipment and medium
CN118301261A
Video generation method and device, equipment, medium and product
CN119211663A