Video generation method and device, electronic equipment and storage medium
By identifying keyword categories in video dialogue text, determining the style of visual effects elements, and defining the display area in the video frame, the problem of low video generation efficiency in existing technologies is solved. This enables the efficient generation of videos with visual effects elements, improving the information focus and user experience of the video.
Patent Information
- Application Number
- CN202511990170.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-26
- Publication Date
- 2026-03-03
AI Technical Summary
In existing technologies, the efficiency of users manually generating videos with visual effects elements is relatively low.
By identifying keyword categories in the dialogue text of the video to be processed, the element styles of visual effects elements are determined, and the display area is determined in the video frame to generate a video displaying visual effects elements.
It improves the efficiency of generating and displaying videos with visual effects, enhances the information focus and content diversity of the videos, and improves the user viewing experience.
Smart Images

Figure CN121603746A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing technology, and in particular to a video generation method, apparatus, electronic device, and storage medium. Background Technology
[0002] As major online video platforms release an increasing number of video works, users use these works as source material for secondary creation, generating derivative videos. Currently, when creating derivative videos, users manually add visual effects elements, resulting in videos with these effects. However, manually generating videos with visual effects is inefficient. Summary of the Invention
[0003] The purpose of this invention is to provide a video generation method, apparatus, electronic device, and storage medium to improve the efficiency of generating and displaying videos with visual effects elements. The specific technical solution is as follows:
[0004] According to one aspect of the present invention, a video generation method is provided, the method comprising:
[0005] Obtain the word categories of keywords in each line of the dialogue text of the video to be processed, as well as the playback time of each line in the video to be processed;
[0006] Based on the word categories of the obtained keywords, determine the element styles of the visual effects elements corresponding to each line of dialogue;
[0007] Within the video frame playing during the time slot of each line of dialogue, determine the display area for the visual effects elements corresponding to each line of dialogue;
[0008] Based on the element style of the visual effects elements corresponding to each line of dialogue, a video is generated in which the display area corresponding to each line of dialogue in the video to be processed displays the visual effects elements corresponding to each line of dialogue.
[0009] In one embodiment of the present invention, the display area for displaying the visual effects element corresponding to each line of dialogue is determined in the video frame played during the dialogue playback period in the following manner:
[0010] In the video footage played during the time period of the dialogue, identify the area to avoid elements that contain the main subject of the video footage;
[0011] In the video frame, outside the area where the element is avoided, a display area is determined for the visual effects element corresponding to the line of dialogue.
[0012] In one embodiment of the present invention, determining the display area for the visual effects element corresponding to the line of dialogue in the area outside the element avoidance area in the video frame includes:
[0013] Identify the key target points of the main subject in the video frame within the element avoidance area;
[0014] Based on the identified target key points, the display area of the visual effects element corresponding to the line of dialogue is determined in the area outside the element avoidance area in the video frame.
[0015] In one embodiment of the present invention, determining the display area for the visual effects element corresponding to the line of dialogue in the video frame, excluding the element avoidance area, includes:
[0016] If there are multiple element avoidance areas, then the overall connected region is determined, wherein the overall connected region is: a region of a preset shape that can cover all element avoidance areas;
[0017] Determine whether the remaining region is smaller than the region size threshold, wherein the remaining region is: the region in the overall connected region excluding the avoidance regions of each element;
[0018] If the remaining area is smaller than the area size threshold, the display area of the line is determined based on the area in the video frame excluding the overall connected area.
[0019] If the remaining area is greater than or equal to the area size threshold, the display area of the line is determined based on the area in the video frame excluding the avoidance areas of each element.
[0020] In one embodiment of the present invention, the main body of the video frame includes at least one of the following elements:
[0021] Facial elements in the video footage;
[0022] Key items and elements in the video footage;
[0023] Fixed elements in a fixed position within a video frame.
[0024] In one embodiment of the present invention, the element style includes at least one of the following styles:
[0025] Captions with text effects that correspond to the keyword's word category;
[0026] Display images of the people described by the keywords;
[0027] Display images of the objects described by the keywords;
[0028] Display images representing the time periods described by the keywords;
[0029] Display images that represent the emotions described by the keywords;
[0030] Display animation effects corresponding to the word categories of the keywords.
[0031] In one embodiment of the present invention, determining the element style of the visual effect element corresponding to each line of dialogue based on the word category of the obtained keywords includes: for each line of dialogue, determining the element style of the visual effect element corresponding to the word category of the keywords in the line of dialogue based on the correspondence between word category and element style of visual effect element; if there are multiple determined element styles of visual effect element, then the determined element styles of visual effect element are combined, and the combined element style is used as the element style of the visual effect element corresponding to the line of dialogue.
[0032] According to another aspect of the present invention, a video generation apparatus is provided, the apparatus comprising:
[0033] The word category acquisition module is used to obtain the word category of keywords in each line of dialogue in the video to be processed, as well as the playback time of each line of dialogue in the video to be processed;
[0034] The element style determination module is used to determine the element style of the visual effects elements corresponding to each line of dialogue based on the word category of the obtained keywords.
[0035] The display area determination module is used to determine the display area of the visual effects elements corresponding to each line of dialogue in the video frame played during the time period of each line of dialogue.
[0036] The video generation module is used to generate a video in which the display area of each line of dialogue in the video to be processed displays the visual effects elements corresponding to each line of dialogue, according to the element style of the visual effects elements corresponding to each line of dialogue.
[0037] In one embodiment of the present invention, the display area determination module determines the display area for displaying the visual effects element corresponding to the line in the video frame played during the line playback time period of each line in the following manner: in the video frame played during the line playback time period of the line, an element avoidance area containing the main body of the video frame is determined; in the area of the video frame other than the element avoidance area, a display area for displaying the visual effects element corresponding to the line is determined.
[0038] In one embodiment of the present invention, the display area determination module is specifically used to determine the target key points of the main body of the video screen in the element avoidance area; based on the determined target key points, the display area of the visual effect element corresponding to the line of dialogue is determined in the area outside the element avoidance area in the video screen.
[0039] In one embodiment of the present invention, the display area determination module is specifically used to determine an overall connected region if there are multiple element avoidance regions, wherein the overall connected region is a region of a preset shape that can cover all element avoidance regions; determine whether the remaining region is less than a region size threshold, wherein the remaining region is the region in the overall connected region excluding each element avoidance region; if the remaining region is less than the region size threshold, determine the display area of the line based on the region in the video frame excluding the overall connected region; if the remaining region is greater than or equal to the region size threshold, determine the display area of the line based on the region in the video frame excluding each element avoidance region.
[0040] In one embodiment of the present invention, the main body of the video frame includes at least one of the following elements: facial elements in the video frame; key object elements in the video frame; and fixed elements with fixed positions in the video frame.
[0041] In one embodiment of the present invention, the element style includes at least one of the following styles: a caption for the keyword with text effects corresponding to the word category of the keyword; an image of the person described by the keyword; an image of the object described by the keyword; an image of the time described by the keyword; an image of the emotion described by the keyword; and an animation effect corresponding to the word category of the keyword.
[0042] In one embodiment of the present invention, the element style determination module is specifically used to determine, for each line of dialogue, the element style of the visual effect element corresponding to the word category of the keyword in the line of dialogue, according to the correspondence between word category and element style of visual effect element; if there are multiple determined element styles of visual effect element, the determined element styles of visual effect element are combined, and the combined element style is used as the element style of visual effect element corresponding to the line of dialogue.
[0043] According to another aspect of the present invention, an electronic device is provided, including a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus;
[0044] Memory, used to store computer programs;
[0045] The processor, when executing a program stored in memory, implements any of the video generation methods described above.
[0046] According to another aspect of the present invention, a computer-readable storage medium is provided, wherein a computer program is stored therein, and the computer program, when executed by a processor, implements any of the video generation methods described above.
[0047] According to another aspect of the present invention, a computer program product containing instructions is provided, which, when run on a computer, causes the computer to perform any of the video generation methods described above.
[0048] As can be seen from the above, the video generation method provided by the embodiments of the present invention can determine the element style of the visual effect elements of each line based on the word category of the keywords in each line of the dialogue text of the video to be processed, and thus efficiently generate a video in which the display area corresponding to each line of the video to be processed displays the visual effect elements corresponding to each line of the dialogue, thereby improving the efficiency of generating videos displaying visual effect elements. Attached Figure Description
[0049] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below.
[0050] Figure 1 This is a flowchart illustrating a video generation method provided in an embodiment of the present invention;
[0051] Figure 2 This is a flowchart illustrating a display area confirmation method provided in an embodiment of the present invention;
[0052] Figure 3 This is a schematic diagram of an overall connected region provided in an embodiment of the present invention;
[0053] Figure 4 This is a schematic diagram of the structure of a video generation device provided in an embodiment of the present invention;
[0054] Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation
[0055] The technical solutions of the present invention will now be described with reference to the accompanying drawings in the embodiments of the present invention.
[0056] The execution subject of the embodiments of the present invention will be described below.
[0057] The solutions provided in this invention can be applied to electronic devices such as desktop computers, laptops, and servers. They can also be applied to video generation platforms that generate videos containing visual effects elements; however, this application does not limit the scope of the application. For ease of description, the executing entity of this invention will be referred to as a video generation platform.
[0058] In one embodiment of the present invention, a video generation method is provided. The method includes: obtaining the word categories of keywords in each line of dialogue in the text of the video to be processed, and the playback time period of each line in the video to be processed; determining the element style of the visual effect element corresponding to each line according to the obtained word categories of keywords; determining the display area of the visual effect element corresponding to each line in the video frame played during the playback time period of each line; and generating a video in the video to be processed where the display area of each line displays the visual effect element corresponding to each line, according to the element style of the visual effect element corresponding to each line.
[0059] As can be seen from the above, the video generation method provided by the embodiments of the present invention can determine the element style of the visual effect elements of each line based on the word category of the keywords in each line of the dialogue text of the video to be processed, and thus efficiently generate a video in which the display area corresponding to each line of the video to be processed displays the visual effect elements corresponding to each line of the dialogue, thereby improving the efficiency of generating videos displaying visual effect elements.
[0060] In addition, visual effects elements with styles corresponding to the word categories of keywords in each line of dialogue are used to make the style of the visual effects elements more consistent with the dialogue text. This enhances the information focus of the video to be processed through visual effects elements. Furthermore, adding visual effects elements appropriately to the display area of the video to be processed can improve the content diversity of the video and enhance the user experience of watching the generated video with visual effects elements.
[0061] In one embodiment of the present invention, see Figure 1 A flowchart of a video generation method is provided, which includes the following steps S101-S104.
[0062] Step S101: Obtain the word category of each line in the dialogue text of the video to be processed, as well as the playback time of each line in the video to be processed.
[0063] The videos to be processed can be movies, TV series, short dramas, variety shows, sports events, and other videos. The videos to be processed can also be videos created by users based on the above videos as materials, without the addition of visual effects.
[0064] Keywords can be categorized into categories such as names, locations, times, objects, and emotions. For example, assuming a line of dialogue reads, "Zhang XX buried a key in Liufeng Mountain three years ago," the resulting keywords would include "Zhang XX," "Liufeng Mountain," "three years ago," and "key." "Zhang XX" would be categorized as a name, "Liufeng Mountain" as a location, "three years ago" as a time, and "key" as an object. This keyword categorization can be achieved by using an LLM (Large Language Model) model to identify and refine the structured word categories within the pre-defined dialogue text.
[0065] Video generation platforms can use video processing tools to preprocess the video to be processed, obtaining the decoded video frames and audio tracks of the video. For example, the video processing tool could be FFmpeg (an open-source audio and video processing tool).
[0066] The following explains how to obtain the dialogue text from the video to be processed.
[0067] In one implementation, the video generation platform can perform speech-to-text recognition on the audio of the video to be processed, obtain the audio text of the video to be processed, and use the audio text as dialogue text.
[0068] In another implementation, the video generation platform can perform optical character recognition (OCR) on the screen of the video to be processed to obtain the on-screen text. Alternatively, if the video generation platform can extract the subtitle file of the video to be processed using video processing tools, the text in the subtitle file can be used as the on-screen text of the video to be processed. Then, the video generation platform can use the above-mentioned on-screen text as dialogue text.
[0069] In another implementation, the video generation platform can obtain the video text and audio text of the video to be processed according to the two implementation methods mentioned above, and use the intersection text between the audio text and the video text as the dialogue text.
[0070] Among the three implementation methods mentioned above, the video generation platform can use ASR (Automatic Speech Recognition) technology to perform speech-to-text recognition on the audio of the video to be processed, and can use OCR (Optical Character Recognition) technology to perform optical character recognition on the screen of the video to be processed.
[0071] The following explains how to obtain the word categories of keywords in each line of dialogue text.
[0072] In one implementation, the video generation platform can extract semantic features from each line of the dialogue text, and input these semantic features into a word classification model to obtain the keywords in the dialogue lines labeled by the word classification model and the word categories of the keywords. For example, the video generation platform can extract feature vectors from the dialogue text based on a feature extraction model, which can be used as semantic features of the dialogue text. The feature extraction model can be an LLM model, a CLIP (Contrastive Language-Image Pre-training, a cross-modal neural network model), or a combination of CLIP and LLM. Specifically, when using a combination of CLIP and LLM to extract features from the dialogue text, the CLIP model can be used to encode the dialogue text, obtaining encoded features. These encoded features are then input into the LLM model for feature extraction to obtain the feature vectors of the dialogue text. This enhances the feature representation capability of the LLM model's input features.
[0073] In another implementation, the video generation platform can segment each line of the dialogue text to obtain the segmentation results, and then perform keyword matching on the segmentation results to obtain the keywords in each line and the word category of the keywords.
[0074] The following explains how to obtain the playback time period of each line in the video to be processed.
[0075] In one implementation, during the process of obtaining the dialogue text of the video to be processed, the video generation platform can record the start and end timestamps of the playback of each line in the video to be processed. For example, the recorded start timestamp can be 00:05 and the end timestamp can be 00:08. Then, the playback time period of the line in the video to be processed is 00:05-00:08.
[0076] Step S102: Determine the element style of the visual effects elements corresponding to each line of dialogue based on the word category of the obtained keywords.
[0077] In one implementation, for each line of dialogue, the video generation platform can determine the element style of the visual effect element corresponding to the word category of the keyword in the line of dialogue based on the correspondence between word category and element style of visual effect element.
[0078] The correspondence between word categories and visual effect elements, as well as the description of the element styles, will be explained in the examples below, and will not be detailed here.
[0079] Step S103: In the video frame played during the time period of each line of dialogue, determine the display area of the visual effects element corresponding to each line of dialogue.
[0080] The method for determining the display area in step S103 will be described in the following embodiments, and will not be detailed here.
[0081] Step S104: Based on the element style of the visual effect elements corresponding to each line of dialogue, generate a video in which the display area corresponding to each line of dialogue in the video to be processed displays the visual effect elements corresponding to each line of dialogue.
[0082] In one implementation, the video generation platform can generate visual effects elements corresponding to each line of dialogue in the display area corresponding to each line of dialogue in the video to be processed, according to the element style of the visual effects elements corresponding to each line of dialogue, and obtain a video displaying the visual effects elements corresponding to each line of dialogue.
[0083] In one approach, for each line of dialogue, the video generation platform can determine the playback time period of that line, obtain the number of original video frames in the video to be processed that play during the dialogue playback time period, generate several original video frames corresponding to the line of dialogue, set the transparency of the area containing the visual effect elements in the visual effect video frames to 0, and the transparency of other areas to 100%, align the original video frames and the visual effect video frames from the starting frame to obtain several pairs of original video frames containing the original video frames and the visual effect video frames, and superimpose the original video frames and the visual effect video frames in each pair of video frames according to the transparency settings, to obtain a video containing the visual effect elements corresponding to the line of dialogue.
[0084] The resolution of the special effects video frame is the same as that of the original video frame. The area in the special effects video frame that is the same as the display area corresponding to the line contains the visual effects element corresponding to the line.
[0085] When the visual effects element is a static element, the element style of the visual effects element is a static style. In this way, the original video frames and several special effects video frames are all the same video frames, and the style of the visual effects element in each special effects video frame is the above-mentioned static style.
[0086] When the visual effect element is a dynamic element, the element style of the visual effect element is a dynamic style. In this way, the visual effect element in several effect video frames of the original video frame changes according to the above dynamic style. The relative change of the visual effect element in adjacent effect video frames is determined based on the number of original video frames and the total change corresponding to the above dynamic style. For example, the above relative change is equal to the ratio of the total change to the number of original video frames.
[0087] In another implementation, the video generation platform can generate visual effects elements corresponding to each line of dialogue in the video to be processed, based on the resolution of the video and the position of the display area corresponding to each line of dialogue, according to the element style of the visual effects elements corresponding to each line of dialogue in the display area corresponding to each line of dialogue in the video to be processed. This results in a video displaying visual effects elements corresponding to each line of dialogue. The video tracks of the video effects element video, the video track of the video to be processed, and the audio track of the video to be processed are synchronized in time. For example, the timestamps of the video tracks of the video effects element video, the video track of the video to be processed, and the audio track of the video to be processed can be aligned. For example, starting from the beginning timestamp, video frames and audio frames with the same timestamp in the video tracks of the video effects element video, the video track of the video to be processed, and the audio track of the video to be processed can be aligned to achieve time synchronization. This ensures that the video effects element video, the video to be processed, and the audio of the video to be processed play synchronously during playback, resulting in a video where the display area corresponding to each line of dialogue in the video to be processed displays the visual effects elements corresponding to each line of dialogue.
[0088] Furthermore, after generating a video displaying visual effects elements corresponding to each line of dialogue, the video generation platform can use a video quality assessment model to detect quality loss in the generated video. If the detected quality loss is less than or equal to a preset loss threshold, the generated video displaying visual effects elements corresponding to each line of dialogue is retained. If the detected quality loss is greater than the preset loss threshold, the generated video is not retained, and video generation is performed again. For example, the aforementioned video quality assessment model can be a VMAF (Video Multimethod Assessment Fusion) model, and the resulting video can have a resolution of 1080P and a bitrate of 8Mbps. In this way, the video generation platform can input the generated video and the video to be processed into the VMAF model. The VMAF model can evaluate the quality loss of the generated video relative to the video to be processed based on indicators such as pixel-level error, preservation of texture and edge information in the image, preservation of image structural information, and inter-frame motion coherence. Among them, the preservation of image structural information can assess the loss of contour, texture, and structural information of the main subject in the video frame, while inter-frame motion coherence can assess the motion continuity of the main subject in the video frame between video frames. The human eye has a high degree of perception of such quality loss, and the above indicators can also assess the degree of human perception of the quality loss of the generated video. By performing quality loss detection in the above way, even when the generated visual effects elements occlude the main subject of the video frame, the detected quality loss is usually high, exceeding the preset loss threshold. If the detected quality loss is low, less than or equal to the preset loss threshold, it indicates that the main structure, contour, texture, and motion continuity of the main subject in the video frame are well preserved. By removing videos with high quality loss, the occlusion of the main subject in the generated video can be reduced.
[0089] As can be seen from the above, the video generation method provided by the embodiments of the present invention can determine the element style of the visual effect elements of each line based on the word category of the keywords in each line of the dialogue text of the video to be processed, and thus efficiently generate a video in which the display area corresponding to each line of the video to be processed displays the visual effect elements corresponding to each line of the dialogue, thereby improving the efficiency of generating videos displaying visual effect elements.
[0090] In addition, visual effects elements with styles corresponding to the word categories of keywords in each line of dialogue are used to make the style of the visual effects elements more consistent with the dialogue text. This enhances the information focus of the video to be processed through visual effects elements. Furthermore, adding visual effects elements appropriately to the display area of the video to be processed can improve the content diversity of the video and enhance the user experience of watching the generated video with visual effects elements.
[0091] The following explains how to determine the display area in step S103.
[0092] In one embodiment of the present invention, see Figure 2 A flowchart of a method for confirming a display area is provided. According to the following steps S201-S202, the display area for visual effects elements corresponding to each line of dialogue is determined in the video frame played during the dialogue playback time period.
[0093] Step S201: In the video frame played during the time period of the dialogue, determine the element avoidance area that contains the main body of the video frame.
[0094] In one implementation, the video generation platform can perform video frame subject recognition on the video frame to be processed, determine the timestamp of the video frame containing the video frame subject, and the element avoidance area corresponding to the determined timestamp containing the video frame subject. Then, it determines the element avoidance area where the timestamp belongs to the playback time period of the dialogue line.
[0095] In this way, when multiple different videos with visual effects elements corresponding to each line of dialogue are generated using the same video to be processed, it is not necessary to repeatedly perform video subject recognition. Instead, the avoidance area containing the video subject can be determined by using timestamps.
[0096] In another implementation, the video generation platform can perform video subject recognition on the video frame corresponding to the time period of the dialogue in the video to be processed, and determine the element avoidance area containing the main element of the video frame.
[0097] The following explains the method for identifying the subject in a video frame.
[0098] The main body of the video frame includes at least one of the following elements: facial elements in the video frame, key object elements in the video frame, and fixed elements in a fixed position in the video frame.
[0099] For facial elements and / or key object elements in a video frame, the video generation platform can use object recognition to identify the main subject of the video frame, designating the identified object as the main subject. For example, the object can be a face and / or a key object. The platform can use models such as MTCNN (Multi-task Convolutional Neural Network) or YOLO (You Only LookOnce) for object recognition. Furthermore, the platform can use a multimodal classification model to analyze the identified main subject, further determining whether it is a key theme in the video frame. The identified avoidance region containing the main subject can be represented by two coordinates of a rectangle, such as the coordinates of the two corners connected by the rectangle's diagonals.
[0100] For fixed elements in a video frame, the video generation platform can determine the fixed elements and the element avoidance area according to preset positions and sizes. For example, if you want to retain the subtitles in the video to be processed, you can designate an area starting from the bottom edge of the video, 100 pixels high and 1920 pixels wide as the element avoidance area.
[0101] In this way, the main subject of the video frame to be avoided can be set more flexibly according to the scene described by the video frame to be processed, so that the generated visual effects elements can adapt to different video frame scenes and improve the user's viewing experience.
[0102] Step S202: In the area outside the element avoidance area in the video frame, determine the display area for the visual effects element corresponding to the line of dialogue.
[0103] In one implementation, a display area that can fully display the visual effects element corresponding to the line can be randomly determined in the video frame, excluding the element avoidance area.
[0104] In another implementation, the video generation platform can determine the target key points of the main body of the video frame in the element avoidance area, and based on the determined target key points, determine the display area of the visual effect element corresponding to the line in the video frame outside the element avoidance area.
[0105] In cases where the main subject of the video is a face, the target key points can be key points around the mouth or the top of the head. When the video generation platform performs target object recognition, if the target object is a face, it can further identify the coordinates of the mouth or the top of the head to obtain the coordinates of the target key points.
[0106] In one approach, the video generation platform can identify display points in the video frame, excluding the element avoidance area, that are less than a preset distance threshold from the target key point. If multiple display points exist, the closest one can be selected. Then, based on the identified display points, and assuming the overlap area between the display area and the element avoidance area is less than a preset overlap area threshold, an area that can be fully displayed and includes the aforementioned display points is determined as the display area. For example, the preset distance threshold could be 100 pixels, 150 pixels, 200 pixels, etc., and the preset overlap area threshold could be 5%, 10%, etc.
[0107] Furthermore, the dialogue can be classified based on a natural language model to determine whether it is a line of dialogue or a line of psychological description. If it is a line of dialogue, the mouth is taken as the target key point of the main body of the video frame; if it is a line of psychological description, the top of the head is taken as the target key point of the main body of the video frame.
[0108] In this way, by identifying the key points of the main subject of the video frame, the generated visual effects elements can better match the key points of the main subject of the video frame, thereby improving the viewing experience of users watching videos that display visual effects elements corresponding to each line of dialogue.
[0109] The following describes another way to implement step S202, which determines the display area.
[0110] If multiple elements need to be avoided, the video generation platform can determine the display area by following these steps:
[0111] Step A: Determine the overall connected regions.
[0112] The aforementioned overall connected region is a region of a preset shape that can cover all element-avoidance areas. The preset shape can be a rectangle, ellipse, or circle, etc. The overall connected region is a single, preset-shaped region that can cover all element-avoidance areas.
[0113] In one implementation, the video generation platform can determine a region of a preset shape that can cover all the areas avoided by elements, as the overall connected region.
[0114] In one approach, the video generation platform can determine the leftmost left boundary point, the rightmost right boundary point, the topmost top boundary point, and the bottommost bottom boundary point of all element avoidance areas based on each element avoidance area. Then, based on the left boundary point, right boundary point, top boundary point, and bottom boundary point, it can determine a region of a preset shape that can cover all element avoidance areas.
[0115] In one scenario, the preset shape can be a rectangle, and the overall connected region can be the smallest bounding rectangle covering all the areas avoided by elements. For example, if the preset shape is a rectangle, the video generation platform can determine a column of pixels containing the left boundary point as the left boundary line of the overall connected region, a column of pixels containing the right boundary point as the right boundary line of the overall connected region, a row of pixels containing the top boundary point as the top boundary line of the overall connected region, and a row of pixels containing the bottom boundary point as the bottom boundary line of the overall connected region. The area enclosed by the left, right, top, and bottom boundary lines is then considered as the overall connected region.
[0116] In another scenario, the preset shape can be a rectangle, and the overall connected region can be the circumscribed rectangle covering all the areas avoided by elements. For example, after determining the left, right, top, and bottom boundary points, a column of pixels located to the left of the left boundary point and spaced a predetermined number of pixels horizontally from the left boundary point can be determined as the left boundary line of the overall connected region; a column of pixels located to the right of the right boundary point and spaced a predetermined number of pixels horizontally from the right boundary point can be determined as the right boundary line of the overall connected region; a row of pixels located above the top boundary point and spaced a predetermined number of pixels vertically from the top boundary point can be determined as the top boundary line of the overall connected region; and a row of pixels located below the bottom boundary point and spaced a predetermined number of pixels vertically from the bottom boundary point can be determined as the bottom boundary line of the overall connected region. The area enclosed by the left, right, top, and bottom boundary lines is then considered the overall connected region.
[0117] In another approach, the video generation platform can determine the smallest preset shape that can completely cover all the areas avoided by the elements, as the overall connected region.
[0118] The video generation platform can determine discrete boundary points on the boundary contours of all element avoidance areas, and determine the area with the minimum preset shape that covers the element avoidance areas based on the determined discrete boundary points.
[0119] Given a rectangular shape, the video generation platform can use the convex hull algorithm to calculate the minimum convex polygon of the avoidance area of each element based on the determined discrete boundary points, and use the rotating caliper method to determine the minimum enclosing rectangle of the minimum convex polygon, thus obtaining the minimum enclosing rectangle that covers the avoidance area of all elements.
[0120] Given a pre-defined circular shape, the video generation platform can use an efficient random incremental method to determine the minimum enclosing circle that surrounds the discrete boundary points, thereby obtaining the minimum enclosing circle that covers the avoidance area of all elements.
[0121] Given an elliptical shape, the video generation platform can use eigenvalue decomposition to determine the minimum bounding ellipse surrounding the discrete boundary points, thus obtaining the minimum bounding ellipse covering the avoidance area of all elements.
[0122] For example, taking a rectangle as the preset shape, if there are two elements that need to avoid each other, the overall connected region can be referenced. Figure 3 A schematic diagram of a fully connected region is provided. Figure 3 The two smaller rectangles containing the human-shaped image are the two element avoidance regions, and the larger rectangle containing the two element avoidance regions is the overall connected region. If the overall connected region continues to shrink, it cannot completely cover the entire area. Figure 3 The two elements in the avoidance area.
[0123] Step B: Determine whether the remaining area is smaller than the area size threshold.
[0124] The remaining region refers to the area within the overall connected region excluding the avoidance regions of individual elements. The region size threshold can include a length threshold and a width threshold; for example, the region size threshold could be 400 pixels long and 300 pixels wide.
[0125] In one implementation, when determining whether the remaining region is smaller than the region size threshold, the video generation platform can determine the comparison region based on the region size threshold, and then determine whether the remaining region can accommodate the comparison region. If there is space in the remaining region that can accommodate the comparison region, then the remaining region is determined to be greater than or equal to the region size threshold. If there is no space in the remaining region that can accommodate the comparison region, then the remaining region is determined to be smaller than the region size threshold.
[0126] The region size threshold can be the dimensions of a circle, such as its diameter. It can also be the dimensions of a rectangle, such as its length and width. Different region size thresholds can be used to determine the size and shape of the comparison region.
[0127] The following explains how to determine whether the remaining area can accommodate the comparison area.
[0128] In one implementation, if the comparison region is circular, the video generation platform can determine the largest inscribed circle of the remaining region. If the radius of the largest inscribed circle is greater than or equal to the radius of the comparison region, the remaining region can accommodate the comparison region; if the radius of the largest inscribed circle is less than the radius of the comparison region, the remaining region cannot accommodate the comparison region. For example, the video generation platform can determine the largest inscribed circle of the remaining region using a binary search method.
[0129] If the comparison area is rectangular, the video generation platform can determine the maximum inscribed rectangle of the remaining area. If the length and width of the maximum inscribed rectangle are greater than or equal to the length and width of the comparison area, the remaining area can accommodate the comparison area. If the length of the maximum inscribed rectangle is less than the length of the comparison area, or the width of the maximum inscribed rectangle is less than the width of the comparison area, the remaining area cannot accommodate the comparison area. For example, the video generation platform can determine the maximum inscribed rectangle of the remaining area using the rotating caliper method.
[0130] In this way, by judging the size of the area outside the avoidance area of each element in the overall connected area of the preset shape covering each element's avoidance area, it is possible to determine whether there is enough space between each element's avoidance area to add visual effects elements. While reducing the obstruction of the main body of the video by visual effects elements, the display space of visual effects elements is also taken into consideration, so that visual effects elements are displayed in the display area as reasonably as possible, thereby improving the user's experience of watching videos that display visual effects elements corresponding to each line of dialogue.
[0131] Step C: If the remaining area is smaller than the area size threshold, then determine the display area of the line based on the area in the video frame other than the overall connected area.
[0132] Step D: If the remaining area is greater than or equal to the area size threshold, then the display area of the line is determined based on the area in the video frame excluding the avoidance areas of each element.
[0133] The methods for determining the display area of the dialogue line in steps C and D are similar to those in step S202, except that the areas on which they are based are different.
[0134] In step C, the display area of the line is determined in the area outside the overall connected region. In steps D and S202, the display area of the line is determined in the area outside the avoidance areas of each element.
[0135] As can be seen from the above, by identifying the avoidance area for elements containing the main body of the video frame that need to be avoided in the video frame, and then determining the display area from the area other than the avoidance area, it is possible to reduce the situation where generated visual effects elements obscure the main body of the video frame to be processed, reduce the loss of picture quality in the generated video displaying visual effects elements corresponding to each line of dialogue, and improve the user's viewing experience of the generated video.
[0136] The following section explains the correspondence between word categories and element styles of visual effect elements, based on the types of element styles.
[0137] The element style of a visual effects element can include at least one of the following styles:
[0138] 1. Captions with text effects that correspond to the keyword's word category.
[0139] In this case, the generated visual effect elements of the above-mentioned element style can be keyword captions with added text effects.
[0140] For example, if the keyword belongs to the category of personal names, then the text effects corresponding to personal names can be gold borders, silver borders, red borders, etc.
[0141] If the keyword's category is emotion-related, the video generation platform can analyze the keyword's sentiment and determine the corresponding text effects within the emotion-related text effects. For example, the keyword's sentiment could include anger, joy, sweetness, etc. The text effect corresponding to anger could be a red flame border, joy could be a blooming flower border, and sweetness could be a red heart border, etc.
[0142] 2. Display images of the people described by the keywords.
[0143] For example, if the keyword category is "person's name", the corresponding image of the person can be determined based on the correspondence between the person's name and the person's portrait image.
[0144] 3. Display images of the objects described by the keywords.
[0145] For example, if the keyword is categorized as an object or location, the video generation platform can determine the object described by the keyword and then generate an image of that object. For instance, if the keyword is "Liufeng Mountain," the platform can render an image of the mountain based on a 3D model of the word "mountain." Similarly, if the keyword is "XXX Diary," the platform can render an image of the diary based on a 3D model of the word "diary."
[0146] IV. Display images representing the time periods described by the keywords.
[0147] For example, if the keyword category is time-related, the video generation platform can obtain images of clocks and calendars.
[0148] 5. Display images describing the emotions described by the keywords.
[0149] When the keyword category is emotion-related, the video generation platform can analyze the emotion of the keyword and determine the corresponding facial expression image from the facial expression images associated with that emotion category.
[0150] 6. Display animation effects corresponding to the word categories of the keywords.
[0151] Different animation effects can be displayed for different keyword categories. If a keyword can generate visual effects elements of other element styles, the animation effect can be added to the visual effects elements of other element styles generated for that keyword. If it is not possible to generate visual effects elements of other element styles, a caption with animated effects can be generated for the keyword.
[0152] The following example illustrates the correspondence between word categories and element styles of visual effects elements.
[0153] I. Personal Names
[0154] The style of elements corresponding to the name category can include: text effects for the name category, keyword captions, animation effects for the name category, and images displaying the person described by the keywords. For example, you can set a gold-bordered text effect as a text effect for the name category, and a fade-in / fade-out animation effect as an animation effect for the name category.
[0155] II. Emotional Category
[0156] The style of elements corresponding to emotions can include: text effects with keywords related to the emotion, animation effects for the emotion, and images displaying the emotion described by the keywords. For example, a text effect with a red flame border can be set as the text effect for the keyword of the emotion category of anger, and a pulse amplification animation effect can be set as the animation effect for the emotion category.
[0157] III. Time-related
[0158] The style of elements corresponding to the emotion category can include: animation effects corresponding to the time category and images displaying the time described by the keyword. For example, you can set an image of a clock or a calendar to display the time described by the keyword, and you can set a clock countdown or a calendar page turning as animation effects corresponding to the time category.
[0159] IV. Items
[0160] The element styles corresponding to the item category can include: animation effects corresponding to the item category and images of the objects described by the keywords. For example, you can set an image based on the object described by the item category keywords as the image of the object described by the keywords, and you can set a rotation display as an animation effect corresponding to the time category.
[0161] The following explains the case where there are multiple element styles for visual effects elements corresponding to the word categories of keywords in a sentence.
[0162] Suppose a sentence contains multiple keywords, such as "Zhang XX buried a key in Liufeng Mountain three years ago." The video generation platform can randomly select one keyword from these keywords and use the element style of the visual effect element corresponding to the selected keyword's word category as the element style of the visual effect element corresponding to that sentence. In this case, the visual effect element corresponding to that sentence in the generated video's display area includes: the visual effect element corresponding to the randomly selected keyword.
[0163] The video generation platform can also determine the element style of the visual effect element corresponding to the word category of each keyword in the dialogue, combine the determined element styles of the visual effect element, and use the combined element style as the element style of the visual effect element corresponding to the dialogue.
[0164] This allows for the generation of diverse visual effects elements, increasing the variety of generated visual effects. Furthermore, a wider range of visual effects elements can adapt to various video scenes, enhancing the user experience when watching videos with visual effects.
[0165] The following explains how to combine element styles.
[0166] For example, a video generation platform can determine the arrangement order of visual effect elements corresponding to the keyword categories according to the order of keywords in the dialogue. The determined visual effect elements are then combined according to the above arrangement order to obtain the combined element style. The combined element style is: the visual effect elements under each element style are arranged and displayed in the order of arrangement.
[0167] In this case, the generated video will display the visual effects elements corresponding to multiple keywords in the order they appear in the corresponding display area.
[0168] For example, video generation platforms can combine the element styles of visual effects elements corresponding to keyword categories based on their stacking priority. For instance, during combination, the element style with the highest stacking priority is used as the top-level element style in the layer, and the element style with the lowest stacking priority is used as the bottom-level element style. The visual effects elements of each element style are then arranged sequentially in the layer from lowest to highest stacking priority to obtain the combined element style. The combined element style displays the visual effects elements under each element style in ascending order of stacking priority.
[0169] In this scenario, the generated video displays multiple visual effects elements corresponding to the keywords in the designated display area, arranged in ascending order of priority. For example, the order of priority for element styles from lowest to highest could be: name, item, time, and emotion. If the line is "Zhang XX buried a key in Liufeng Mountain three years ago," then "Zhang XX" is a name, "key" is an item, and "three years ago" is a time. The resulting element styles would be: visual effects elements corresponding to "Zhang XX," "key," and "three years ago" would be displayed in ascending order of priority within the layers.
[0170] This allows for the combination of various visual effects elements, increasing the diversity of visual effects elements corresponding to each line of dialogue and improving the user experience when watching videos with visual effects elements.
[0171] Corresponding to the video generation method described above, this embodiment of the invention also provides a video generation apparatus.
[0172] In one embodiment of the present invention, see Figure 4 A schematic diagram of a video generation device is provided, the device comprising:
[0173] The word category acquisition module 401 is used to obtain the word category of keywords in each line of dialogue in the video to be processed, as well as the playback time of each line of dialogue in the video to be processed.
[0174] The element style determination module 402 is used to determine the element style of the visual effect elements corresponding to each line of dialogue based on the word category of the obtained keywords.
[0175] The display area determination module 403 is used to determine the display area of the visual effects elements corresponding to each line of dialogue in the video frame played during the time period of each line of dialogue.
[0176] The video generation module 404 is used to generate a video in which the display area of each line of dialogue in the video to be processed displays the visual effects elements corresponding to each line of dialogue, according to the element style of the visual effects elements corresponding to each line of dialogue.
[0177] As can be seen from the above, the video generation method provided by the embodiments of the present invention can determine the element style of the visual effect elements of each line based on the word category of the keywords in each line of the dialogue text of the video to be processed, and thus efficiently generate a video in which the display area corresponding to each line of the video to be processed displays the visual effect elements corresponding to each line of the dialogue, thereby improving the efficiency of generating videos displaying visual effect elements.
[0178] In addition, visual effects elements with styles corresponding to the word categories of keywords in each line of dialogue are used to make the style of the visual effects elements more consistent with the dialogue text. This enhances the information focus of the video to be processed through visual effects elements. Furthermore, adding visual effects elements appropriately to the display area of the video to be processed can improve the content diversity of the video and enhance the user experience of watching the generated video with visual effects elements.
[0179] In one embodiment of the present invention, the display area determination module determines the display area for displaying the visual effects element corresponding to the line in the video frame played during the line playback time period of each line in the following manner: in the video frame played during the line playback time period of the line, an element avoidance area containing the main body of the video frame is determined; in the area of the video frame other than the element avoidance area, a display area for displaying the visual effects element corresponding to the line is determined.
[0180] As can be seen from the above, by identifying the avoidance area for elements containing the main body of the video frame that need to be avoided in the video frame, and then determining the display area from the area other than the avoidance area, it is possible to reduce the situation where generated visual effects elements obscure the main body of the video frame to be processed, reduce the loss of picture quality in the generated video displaying visual effects elements corresponding to each line of dialogue, and improve the user's viewing experience of the generated video.
[0181] In one embodiment of the present invention, the display area determination module is specifically used to determine key points of the main body of the video frame in the element avoidance area; based on the determined key points, the display area of the visual effect element corresponding to the line of dialogue is determined in the area outside the element avoidance area in the video frame.
[0182] In this way, by identifying the key points of the main subject of the video frame, the generated visual effects elements can better fit the key points of the main subject of the video frame, improving the viewing experience of users watching videos that display visual effects elements corresponding to each line of dialogue.
[0183] In one embodiment of the present invention, the display area determination module is specifically used to determine, if there are multiple element avoidance areas, an overall connected region of a preset shape covering all element avoidance areas; determine whether the remaining area is less than a region size threshold, wherein the remaining area is: the area in the overall connected region excluding each element avoidance area; if the remaining area is less than the region size threshold, then the display area of the line is determined based on the area in the video frame excluding the overall connected region; if the remaining area is greater than or equal to the region size threshold, then the display area of the line is determined based on the area in the video frame excluding each element avoidance area.
[0184] In this way, by judging the size of the area outside the individual element avoidance areas in the overall connected area of the preset shape that covers all element avoidance areas, it is possible to determine whether there is enough space between the element avoidance areas to add visual effects elements. While reducing the obstruction of the main video screen by visual effects elements, the display space of visual effects elements is also taken into consideration, so that visual effects elements are displayed in the display area as reasonably as possible, thereby improving the user's experience of watching videos that display visual effects elements corresponding to each line of dialogue.
[0185] In one embodiment of the present invention, the main body of the video frame includes at least one of the following elements: facial elements in the video frame; key object elements in the video frame; and fixed elements with fixed positions in the video frame.
[0186] In this way, the main subject of the video frame to be avoided can be set more flexibly according to the scene described by the video frame to be processed, so that the generated visual effects elements can adapt to different video frame scenes and improve the user's viewing experience.
[0187] In one embodiment of the present invention, the element style includes at least one of the following styles: a caption for the keyword with text effects corresponding to the word category of the keyword; an image of the person described by the keyword; an image of the object described by the keyword; an image of the time described by the keyword; an image of the emotion described by the keyword; and an animation effect corresponding to the word category of the keyword.
[0188] This allows for the generation of diverse visual effects elements, increasing the variety of generated visual effects. Furthermore, a wider range of visual effects elements can adapt to various video scenes, enhancing the user experience when watching videos with visual effects.
[0189] In one embodiment of the present invention, the element style determination module is specifically used to determine, for each line of dialogue, the element style of the visual effect element corresponding to the word category of the keyword in the line of dialogue, according to the correspondence between word category and element style of visual effect element; if there are multiple determined element styles of visual effect element, the determined element styles of visual effect element are combined, and the combined element style is used as the element style of visual effect element corresponding to the line of dialogue.
[0190] This allows for the combination of various visual effects elements, increasing the diversity of visual effects elements corresponding to each line of dialogue and improving the user experience when watching videos with visual effects elements.
[0191] This invention also provides an electronic device, such as... Figure 5 As shown, it includes a processor 501, a communication interface 502, a memory 503, and a communication bus 504, wherein the processor 501, the communication interface 502, and the memory 503 communicate with each other through the communication bus 504.
[0192] Memory 503 is used to store computer programs;
[0193] The processor 501, when executing the program stored in the memory 503, implements the video generation method described in any of the above embodiments.
[0194] The communication bus mentioned above can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into address bus, data bus, control bus, etc. For ease of illustration, only one thick line is used to represent it in the diagram, but this does not mean that there is only one bus or one type of bus.
[0195] The communication interface is used for communication between the aforementioned terminal and other devices.
[0196] The memory may include random access memory (RAM) or non-volatile memory, such as at least one disk storage device. Optionally, the memory may also be at least one storage device located remotely from the aforementioned processor.
[0197] The processors mentioned above can be general-purpose processors, including central processing units (CPUs), network processors (NPs), etc.; they can also be digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.
[0198] In another embodiment of the present invention, a computer-readable storage medium is also provided, wherein a computer program is stored therein, and when the computer program is executed by a processor, it implements any of the video generation methods described in the above embodiments.
[0199] In another embodiment of the present invention, a computer program product containing instructions is also provided, which, when run on a computer, causes the computer to perform any of the video generation methods described in the above embodiments.
[0200] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of the present invention are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium accessible to a computer or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state disk (SSD)).
[0201] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0202] The various embodiments in this specification are described in a related manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the embodiments of apparatus, electronic devices, storage media, and computer program products are basically similar to the method embodiments, so the descriptions are relatively simple; relevant parts can be referred to the descriptions of the method embodiments.
[0203] The above description is merely a preferred embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention are included within the scope of protection of the present invention.
Claims
1. A video generation method, characterized in that, The method includes: Obtain the word categories of keywords in each line of the dialogue text of the video to be processed, as well as the playback time of each line in the video to be processed; Based on the word categories of the obtained keywords, determine the element styles of the visual effects elements corresponding to each line of dialogue; Within the video frame playing during the time slot of each line of dialogue, determine the display area for the visual effects elements corresponding to each line of dialogue. Based on the element style of the visual effects elements corresponding to each line of dialogue, a video is generated in which the display area corresponding to each line of dialogue in the video to be processed displays the visual effects elements corresponding to each line of dialogue.
2. The method according to claim 1, characterized in that, In the video frame during the playback period of each line of dialogue, determine the display area for the visual effects element corresponding to that line of dialogue in the following manner: In the video footage played during the time period of the dialogue, identify the area to avoid elements that contain the main subject of the video footage; In the video frame, outside the area where the element is avoided, a display area is determined for the visual effects element corresponding to the line of dialogue.
3. The method according to claim 2, characterized in that, The step of determining the display area for the visual effects element corresponding to the line of dialogue in the video frame, excluding the element avoidance area, includes: Identify the key target points of the main subject in the video frame within the element avoidance area; Based on the identified target key points, the display area of the visual effects element corresponding to the line of dialogue is determined in the area outside the element avoidance area in the video frame.
4. The method according to claim 2, characterized in that, The step of determining the display area for the visual effects element corresponding to the line of dialogue in the video frame, excluding the element avoidance area, includes: If there are multiple element avoidance areas, then the overall connected region is determined, wherein the overall connected region is: a region of a preset shape that can cover all element avoidance areas; Determine whether the remaining region is smaller than the region size threshold, wherein the remaining region is: the region in the overall connected region excluding the avoidance regions of each element; If the remaining area is smaller than the area size threshold, the display area of the line is determined based on the area in the video frame excluding the overall connected area. If the remaining area is greater than or equal to the area size threshold, the display area of the line is determined based on the area in the video frame excluding the avoidance areas of each element.
5. The method according to claim 2, characterized in that, The main body of the video frame includes at least one of the following elements: Facial elements in the video footage; Key items and elements in the video footage; Fixed elements in a fixed position within a video frame.
6. The method according to any one of claims 1-5, characterized in that, The element style includes at least one of the following styles: Captions with text effects that correspond to the keyword's word category; Display images of the people described by the keywords; Display images of the objects described by the keywords; Display images representing the time periods described by the keywords; Display images that represent the emotions described by the keywords; Display animation effects corresponding to the word categories of the keywords.
7. The method according to any one of claims 1-5, characterized in that, The step of determining the element style of the visual effects elements corresponding to each line of dialogue based on the word category of the obtained keywords includes: For each line of dialogue, based on the correspondence between word category and element style of visual effect element, determine the element style of visual effect element corresponding to the word category of keyword in the line of dialogue; if there are multiple determined element styles of visual effect element, combine the determined element styles of visual effect element, and use the combined element style as the element style of visual effect element corresponding to the line of dialogue.
8. A video generation apparatus, characterized in that, The device includes: The word category acquisition module is used to obtain the word category of keywords in each line of dialogue in the video to be processed, as well as the playback time of each line of dialogue in the video to be processed; The element style determination module is used to determine the element style of the visual effects elements corresponding to each line of dialogue based on the word category of the obtained keywords. The display area determination module is used to determine the display area of the visual effects elements corresponding to each line of dialogue in the video frame played during the time period of each line of dialogue. The video generation module is used to generate a video in which the display area of each line of dialogue in the video to be processed displays the visual effects elements corresponding to each line of dialogue, according to the element style of the visual effects elements corresponding to each line of dialogue.
9. An electronic device, characterized in that, It includes a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus; Memory, used to store computer programs; A processor, when executing a program stored in memory, implements the steps of the method described in any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the steps of the method described in any one of claims 1-7.