Character dynamic effect data generation method and device, storage medium and electronic equipment
By constructing a text library and a background library, and using a grid partitioning strategy to generate text animation videos, the problem of lacking diverse dynamic text and background pairing data in existing technologies is solved, and the automatic generation of high-quality text animation training data and the improvement of model accuracy are realized.
Patent Information
- Application Number
- CN202511549880.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-28
- Publication Date
- 2026-02-17
AI Technical Summary
Existing technologies lack diverse and real-world dynamic text-background pairing data, making it difficult to automatically generate high-quality text animation training samples, resulting in inaccurate control of text attributes and insufficient model generalization ability.
We construct a text library and a background library, and use a grid partitioning strategy to generate a fusion of black text animation video and background image. We extract the mask image and combine it with the text index information to generate a training dataset.
It achieves high-quality and diverse text animation training data generation, ensuring consistency between text and background, reducing manual costs, and improving model recognition and generation accuracy and generalization ability.
Smart Images

Figure CN121544743A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of multimedia intelligent processing, and in particular to a method and device for generating text motion effect data, a storage medium, and an electronic device. BACKGROUND
[0002] In the field of modern advertising and digital marketing, dynamic visual content (such as short videos) gradually replaces static pictures due to its stronger attraction and propagation effect. Advertisers usually want to convert a single advertisement picture into a dynamic video while maintaining the consistency of core elements (such as the text and the main body of the product). However, the existing technology has significant limitations: manual production relies on designers to design text motion effects frame by frame, which is costly, inefficient, and difficult to apply on a large scale; traditional automated methods generate motion effects by erasing the original text after OCR recognition, but this method loses the design attributes such as font, color, and stroke, thereby destroying the visual consistency; and although the video generation model can generate videos, it is difficult to accurately maintain the stability of the text and the main body, the text is prone to deformation or blurring, and the generated content may introduce irrelevant elements. The root cause of the existing problems is the lack of training data, and there is a lack of paired samples of “dynamic text + static background”, which leads to a lack of fine-grained control of text attributes by the large model. At the same time, the existing synthetic data is mostly single-text motion effect segments, and lacks complex background interference such as multi-text layout, which makes it difficult to meet the needs of real advertising scenarios and limits the generalization ability and practicality of the model. SUMMARY
[0003] The present application provides a method and device for generating text motion effect data, a storage medium, and an electronic device to solve the technical problem of lacking a large amount of diversified and real-scenario paired data of dynamic text and background, and being difficult to automatically generate high-quality text motion effect training samples.
[0004] In a first aspect, the present application provides a method for generating text motion effect data, comprising: constructing a text library and a background library; randomly selecting target text content from the text library according to a text motion effect template to generate a corresponding black-text motion effect video; using a grid division strategy to fuse the black-text motion effect video with a target background image randomly selected from the background library to generate a target motion effect video; extracting a first frame mask image of the target motion effect video to obtain a text mask label image of the target motion effect video, and combining text index information of the target motion effect video to generate a training data set of the target text content, wherein the text index information is the target text content, and the training data set includes the target motion effect video, the text mask label image, and the text index information.
[0005] In a second aspect, the application provides a device for generating text motion effect data, comprising: a construction module configured to construct a script library and a background library; a first generation module configured to generate a corresponding black-text motion effect video by randomly selecting target text content from the script library according to a text motion effect template; a second generation module configured to generate a target motion effect video by fusing the black-text motion effect video and a target background image randomly selected from the background library using a grid division strategy; and a third generation module configured to extract a first frame mask image of the target motion effect video to obtain a text mask label image of the target motion effect video, and generate a training data set of the target text content in combination with script index information of the target motion effect video, wherein the script index information is the target text content, and the training data set comprises the target motion effect video, the text mask label image, and the script index information.
[0006] As an optional example, the construction module comprises a first construction unit configured to generate a plurality of pieces of text content by collecting or using a text generation large model to construct the script library, wherein the text content comprises phrases without punctuation marks and phrases with punctuation marks, and covers different industry scenarios and various emotional tendencies, and the language types of the text content comprise Chinese, English, and numbers.
[0007] As an optional example, the construction module comprises a first generation unit configured to generate a plurality of background images of different scene types by collecting or using a text generation large model, a second construction unit configured to remove the text in each background image using an optical character recognition detection algorithm to construct the background library, and a classification unit configured to extract the theme color of each background image in the background library and classify all the background images in the background library according to the theme color.
[0008] As an optional example, the first generation module comprises a preset unit configured to preset the text motion effect template, wherein the text motion effect template comprises a plurality of types of text motion effects, and a second generation unit configured to randomly select target text content from the script library and apply the target text content to the text motion effect template to generate a corresponding black-text motion effect video.
[0009] As an optional example, the first generation module further comprises a third generation unit configured to randomly generate a text color of the target text content during the generation of the corresponding black-text motion effect video, an adding unit configured to randomly add an outline effect and an outline color to the target text content, and a selection unit configured to randomly select a font type and a font size for the target text content from a system font library.
[0010] As an optional example, the second generation module includes: a partitioning unit, used to divide the canvas into M grids, and assign a unique identifier and coordinates to each grid to form a grid information list; a processing unit, used to randomly extract N videos from the black-background text animation videos, randomly assign them to N grids as video grids, define the remaining grids as interference grids, and randomly mark whether the text in each video grid is a button type, wherein N is less than M; and a fourth generation unit, used to generate random static element information for each interference grid, wherein the static element information includes text content, font type, and font size. The system includes: a text color, stroke color, and button color; a drawing unit, used to randomly select a target background image from the background library as the background base image, and draw the static element information of each interference grid onto the grid position corresponding to the background base image; a filling unit, used to extract the text region mask of each video grid frame by frame, and fill the text region of the video grid marked as button type with a randomly generated button color; and an overlay unit, used to overlay the processed video of each video grid frame by frame onto the grid position corresponding to the background base image to obtain the target motion effect video.
[0011] As an optional example, the third generation module includes: an extraction unit, used to extract the text mask of the first frame image from the target motion video to form the text mask marker image; an establishment unit, used to establish a mapping relationship between the text content in each video frame of the target motion video and the text mask marker image according to the text index information; and a combination unit, used to combine the target motion video, the text mask marker image and the text index information to form a training dataset of the target text content.
[0012] Thirdly, this application provides a storage medium storing a computer program, wherein the computer program is executed by a processor to perform the above-described method for generating text animation data.
[0013] Fourthly, this application also provides an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to execute the above-described method for generating text animation data through the computer program.
[0014] The technical solutions provided in this application have the following advantages compared with the prior art: This application employs the following methods: constructing a text library and a background library; randomly selecting target text content from the text library based on text animation templates to generate corresponding black-background text animation videos; using a grid partitioning strategy, fusing the black-background text animation videos with target background images randomly selected from the background library to generate target animation videos; extracting the first frame mask image of the target animation video to obtain the text mask marker image of the target animation video, and combining it with the text index information of the target animation video to generate a training dataset of the target text content, wherein the text index information is the target text content, and the training dataset includes the target animation video, the text mask marker image, and the text index information. The proposed method, by constructing diverse text and background libraries and generating black-background text animation videos based on preset text animation templates, employs a grid partitioning strategy to fuse the black-background text animation videos with the background images. Finally, it combines text index information to generate text mask markers and training datasets. This achieves automated batch generation of high-quality, diverse text animation training data, ensuring consistency between text animations and backgrounds, improving data diversity and realism, reducing manual costs, and enhancing the accuracy and generalization ability of text animation recognition or generation models. Consequently, it solves the technical problem of lacking a large amount of diverse, real-world dynamic text and background pairing data, making it difficult to automatically generate high-quality text animation training samples. Attached Figure Description
[0015] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0016] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0017] One or more embodiments are illustrated by way of example with reference numerals in the accompanying drawings. These illustrations do not constitute a limitation on the embodiments. Elements with the same reference numerals in the drawings are denoted as similar elements. Unless otherwise stated, the figures in the drawings are not to be limited by scale.
[0018] Figure 1 This is a flowchart of an optional method for generating text animation data according to an embodiment of this application; Figure 2 This is a flowchart illustrating the specific implementation of an optional method for generating text animation data according to an embodiment of this application. Figure 3This is a schematic diagram of an optional text animation data generation device according to an embodiment of this application; Figure 4 This is a schematic diagram of an optional electronic device according to an embodiment of this application. Detailed Implementation
[0019] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0020] The following disclosure provides many different embodiments or examples for implementing different structures of this application. To simplify the disclosure, specific examples of components and arrangements are described below. Of course, these are merely examples and are not intended to limit the scope of this application. Furthermore, reference numerals and / or letters may be repeated in different examples. Such repetition is for simplification and clarity and does not in itself indicate a relationship between the various embodiments and / or arrangements discussed.
[0021] According to a first aspect of the embodiments of this application, a method for generating text animation data is provided, optionally, as follows: Figure 1 As shown, the above method includes: S102, Build a copywriting library and a background library; S104: Based on the text animation template, randomly select target text content from the text library and generate the corresponding black background text animation video; S106, using a grid division strategy, merges the black-background text animation video with a target background image randomly selected from the background library to generate the target animation video; S108, extract the first frame mask image of the target motion effect video to obtain the text mask marker image of the target motion effect video, and combine it with the text index information of the target motion effect video to generate a training dataset of the target text content. The text index information is the target text content. The training dataset includes the target motion effect video, the text mask marker image and the text index information.
[0022] Optionally, this embodiment discloses a method for generating text animation data, the goal of which is to automatically generate high-quality, diverse text animation training data to meet the high consistency and diversity requirements of dynamic text and background in short advertising videos and digital marketing. Figure 2 The flowchart shown illustrates the specific implementation process, which includes the following steps: Data preparation phase: Building a text library and a background library. The text library stores diverse text content, covering different themes, emotional tendencies, and mixed Chinese and English text, ensuring text diversity; the background library includes various text-free background images, such as natural scenery, cityscapes, people, indoor and outdoor backgrounds, and solid color backgrounds. Images containing text are removed using optical character recognition (OCR) algorithms to ensure clean backgrounds.
[0023] Initial animation generation stage: Based on the preset text animation template, target text content is randomly selected from the text library to generate a corresponding black-background text animation video. During the generation process, the text color, font, font size, and outline effect can be randomly set, and button-type text can be generated in some cases to enrich the animation performance.
[0024] Canvas Compositing Stage: A grid-based strategy is used to blend the black-background text animation video with a randomly selected background image from the background library. Specifically, the canvas is divided into several grids, each assigned as either a video grid or a distractor grid. The corresponding text animation video is overlaid in the video grid, while random static elements are drawn in the distractor grid to create a complex and diverse background layout. Text area masks are extracted frame-by-frame from the video grid, and button-type text areas are filled with random colors to ensure accurate and controllable text effects.
[0025] Post-processing stage: Extract the first frame mask image from the target motion effect video to obtain the text mask marker image. Establish a mapping relationship by combining the text index information. Combine the target motion effect video, text mask marker image, and text index information to generate a training dataset of the target text content. This training dataset can be used to train text motion effect recognition or generation models.
[0026] Optionally, in this embodiment, a batch generation of text animation data is achieved through an automated process, ensuring a high degree of consistency between the text animation and the background. It supports rich control over text attributes (such as font, color, stroke, and button effects) and introduces complex background interference to improve the diversity and realism of the training data. Compared to traditional manual production and existing automated methods, this method significantly reduces costs and workload, improves generation efficiency, and provides high-quality training data suitable for real-world advertising scenarios, helping to improve the accuracy and generalization ability of the text animation recognition and generation model.
[0027] As an optional example, building a copywriting library and background library includes: By collecting or generating multiple text content using the Wenshengwen model, a copywriting library is constructed. The text content includes phrases without punctuation and phrases with punctuation, covering different industry scenarios and various emotional tendencies. The language types of the text content include Chinese, English, and numbers.
[0028] Optionally, in this embodiment, multiple text contents are collected manually or generated in batches using a text-to-text model to construct a text library. The text contents include phrases without punctuation and phrases with commonly used punctuation, covering different industry scenarios and various emotional tendencies, such as promotions, warmth, or warnings; the language types of the text contents include Chinese, English, and numbers to ensure text diversity and broad applicability.
[0029] As an optional example, building a copywriting library and background library includes: Multiple background images of different scene types are generated by collecting or creating large-scale models; An optical character recognition (OCR) algorithm is used to remove text from each background image, thus constructing a background library. Extract the theme color of each background image in the background library, and classify all background images in the background library according to the theme color.
[0030] Optionally, in this embodiment, text-free background images are generated by collecting real-scene images, film screenshots, or using a large-scale text-generating model. Background types include natural, urban, portrait, indoor, outdoor, and solid-color backgrounds. An optical character recognition (OCR) algorithm is used to remove images containing text, ensuring a clean background. Simultaneously, theme color detection and classification are performed on each background image to ensure balanced color distribution.
[0031] Optionally, in this embodiment, automated batch generation of text animation data can be achieved, ensuring that the text animation is highly consistent with the background, supporting rich text attribute control, while introducing complex background interference to improve the diversity and authenticity of training data, significantly reducing manual production costs, improving generation efficiency, and enhancing the recognition accuracy and generalization ability of the text animation model.
[0032] As an optional example, based on a text animation template, target text content is randomly selected from the text library to generate a corresponding black-background text animation video, including: Preset text animation templates, which include various types of text animations; Randomly select target text content from the text library and apply it to the text animation template to generate a corresponding black background text animation video.
[0033] As an optional example, the above method also includes the following steps in generating the corresponding black-background text animation video: Randomly generate the text color of the target text content; Randomly add a stroke effect and stroke color to the target text content; Randomly select the font type and size for the target text content from the system font library.
[0034] Optionally, in this embodiment, firstly, multiple types of text animation templates are preset, including but not limited to wave, rotating fly-in, typewriter, heartbeat, and word-by-word zoom-in, to meet diverse dynamic text expression needs. Then, target text content is randomly selected from the text library and applied to the text animation template to generate a corresponding black-background text animation video. During the generation process, the text color (excluding pure black), font, font size, and stroke effect can be randomly set, with a 50% probability of adding a stroke, the stroke color of which is independently randomized. In some cases, the text is marked as a button type to enrich the animation expression. Simultaneously, the generated video can be processed frame-by-frame to ensure the continuity and visual consistency of the text animation in different frames, thereby forming a data foundation suitable for subsequent synthesis and training.
[0035] Optionally, in this embodiment, automated batch generation of black-background text animation videos can be achieved, ensuring the diversity of text animations and controllable fine-grained attributes, providing a high-quality foundation for subsequent fusion with background images and the construction of training datasets. This significantly reduces generation costs and time, improves efficiency, and provides rich and realistic training samples for text animation recognition or generation models, enhancing the model's generalization ability and recognition accuracy.
[0036] As an optional example, a grid partitioning strategy is used to fuse a black-background text animation video with a target background image randomly selected from a background library to generate a target animation video, including: Divide the canvas into M grids and assign a unique identifier and coordinates to each grid to form a grid information list; N videos are randomly selected from the black background text animation videos and randomly assigned to N grids as video grids. The remaining grids are defined as interference grids. Each video grid is randomly marked with whether the text is a button type. The integer N is less than the integer M and is greater than 2. For each interference cell, generate random static element information, including text content, font type, font size, text color, stroke color, and button color. Randomly select a target background image from the background library as the background base image, and draw the static element information of each interference grid onto the corresponding grid position of the background base image; For each video cell, extract the text region mask frame by frame, and for video cells marked as button type, fill the text region determined based on the text region mask of the first frame with a randomly generated button color. The processed video of each video frame is superimposed onto the corresponding grid position of the background image to obtain the target motion effect video.
[0037] Optionally, in this embodiment, the canvas is first divided into M grids, and the division method is randomly selected from the following three: (1) 4 grids: 2×2 uniform grid; (2) 6 grids: height divided into 2 equal parts and width divided into 3 equal parts; (3) 9 grids: 3×3 uniform grid. A unique identifier and start and end coordinates are assigned to each grid to form a grid information list, so as to accurately locate the position of video and static elements in the future.
[0038] Subsequently, N videos are randomly selected from the black-background text animation videos and randomly assigned to N grids as video grids. The remaining grids are defined as interference grids, where N is less than M. Simultaneously, the text content of each video grid is randomly labeled to indicate whether it is a button type, in order to enhance the visual presentation of the text animation.
[0039] Next, random static element information is generated for each interference grid, including text content, font type, font size, text color, stroke color and button color, and the static element information is drawn onto the corresponding grid position of the background image, thus forming a complex and diverse background layout.
[0040] Then, the video is read and its frames are decomposed into images; for each frame image... The image is converted to grayscale, and pixels with values greater than th1 are extracted as text information. This mask image is denoted as... Each video has a 25% chance of being converted to a button type. In this case, based on the first frame's static mask image, the coordinates of the top, bottom, left, and right rectangles of the text mask area need to be determined, and these coordinate areas should be filled with random colors to create the button effect. At this point, the button type... The image calculation formula is as follows:
[0041] In the above formula, The image is a solid color, and its color is the same as the button's color.
[0042] After completing the text mask and button processing, each frame of video image is expanded to the target size and then composited frame by frame with the background image using the following formula:
[0043] in, To expand the video frame image to the target size, This is the text mask image for the corresponding frame (range normalized to 0-1, white area is text area). This serves as the background image. The complete target motion video is obtained by synthesizing frame by frame using the method described above. .
[0044] Optionally, in this embodiment, automated fusion of text animation video and background image is achieved, ensuring controllable text animation attributes and diverse layouts. Simultaneously, complex background interference is introduced to improve the diversity and realism of training data. This significantly reduces production costs and time, improves generation efficiency, and provides high-quality training data for text animation recognition or generation models, enhancing the model's recognition accuracy and generalization ability in complex scenes.
[0045] As an optional example, the first frame mask image of the target motion effect video is extracted to obtain the text mask marker image of the target motion effect video. Combined with the text index information of the target motion effect video, a training dataset for generating the target text content is generated, including: Extract the text mask from the first frame of the target motion video to form a text mask marker image; Based on the text index information, establish a mapping relationship between the text content in each video frame of the target motion effect video and the text mask mark image; The target motion video, text mask labeled image, and text index information are combined to form a training dataset of the target text content.
[0046] Optionally, in this embodiment, firstly, the text mask of the first frame image is extracted from the generated target motion effect video to obtain a text mask marker image. This image is used to accurately identify the pixel regions of the text in the video, thereby providing basic information for the localization of text motion effects. Subsequently, combined with the text index information of the target motion effect video, a precise mapping relationship is established between the text content of each video frame and the text mask marker image, so that each piece of text content corresponds one-to-one with the corresponding pixel region, achieving accurate association between text content and visual information. Finally, the target motion effect video, the text mask marker image, and the text index information are combined to form a complete target text content training dataset. The generated dataset can be used for training subsequent text motion effect recognition or generation models.
[0047] Optionally, in this embodiment, automated and large-scale training data generation can be achieved while ensuring the consistency of the visual attributes of text animation effects. This not only allows for precise annotation of text positions, improving the accuracy and effectiveness of training data, but also enables the mapping between text content and pixel regions based on text indexes. This provides high-quality, controllable, and diverse training samples for the text animation effect model, significantly improving the accuracy and generalization ability of the text animation effect recognition and generation model in complex backgrounds and diverse layout scenarios. Simultaneously, it eliminates the need for manual frame-by-frame annotation or design, reducing data production costs and improving data generation efficiency, providing a feasible solution for short advertising videos, digital marketing, and related automated visual generation technologies.
[0048] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this application.
[0049] According to another aspect of the embodiments of this application, a device for generating text animation data is also provided, such as... Figure 3 As shown, it includes: Module 302 is used to build the copywriting library and background library; The first generation module 304 is used to randomly select target text content from the text library based on the text animation template and generate a corresponding black background text animation video. The second generation module 306 is used to merge the black text animation video with a target background image randomly selected from the background library using a grid division strategy to generate the target animation video. The third generation module 308 is used to extract the first frame mask image of the target motion effect video, obtain the text mask mark image of the target motion effect video, and combine it with the text index information of the target motion effect video to generate a training dataset of the target text content. The text index information is the target text content, and the training dataset includes the target motion effect video, the text mask mark image, and the text index information.
[0050] It should be noted that the construction module 302 in this embodiment can be used to execute step S102 in this application embodiment, the first generation module 304 in this embodiment can be used to execute step S104 in this application embodiment, the second generation module 306 in this embodiment can be used to execute step S106 in this application embodiment, and the third generation module 308 in this embodiment can be used to execute step S108 in this application embodiment.
[0051] As an optional example, the building blocks include: The first building unit is used to collect or generate multiple text content through the text generation model to build a copywriting library. The text content includes phrases without punctuation and phrases with punctuation, and covers different industry scenarios and various emotional tendencies. The language types of the text content include Chinese, English and numbers.
[0052] As an optional example, the building blocks include: The first generation unit is used to generate multiple background images of different scene types by collecting or generating large-scale models; The second construction unit is used to remove text from each background image using an optical character recognition detection algorithm to build a background library; The classification unit is used to extract the theme color of each background image in the background library and classify all background images in the background library according to the theme color.
[0053] As an optional example, the first generation module includes: The preset unit is used to preset text animation templates, which include various types of text animations; The second generation unit is used to randomly select target text content from the text library and apply the target text content to the text animation template to generate the corresponding black background text animation video.
[0054] As an optional example, the first generation module also includes: The third generation unit is used to randomly generate the text color of the target text content during the process of generating the corresponding black-background text animation video. Add a cell to randomly add a stroke effect and stroke color to the target text content; The selection unit is used to randomly select the font type and size for the target text content from the system font library.
[0055] As an optional example, the second generation module includes: The division unit is used to divide the canvas into M grids, and assign a unique identifier and coordinates to each grid to form a grid information list; The processing unit is used to randomly extract N videos from the black text animation video, randomly assign them to N grids as video grids, define the remaining grids as interference grids, and randomly mark whether the text in each video grid is a button type, where N is less than M; The fourth generation unit is used to generate random static element information for each interference cell. The static element information includes text content, font type, font size, text color, stroke color, and button color. The drawing unit is used to randomly select a target background image from the background library as the background base image, and draw the static element information of each interference grid onto the grid position corresponding to the background base image; The filling unit is used to extract the text region mask for each video cell frame by frame, and fill the text region determined according to the text region mask of the first frame with a randomly generated button color for video cells marked as button type. The overlay unit is used to overlay the processed video of each video cell onto the grid position corresponding to the background image frame by frame to obtain the target motion effect video.
[0056] As an optional example, the third generation module includes: The extraction unit is used to extract the text mask of the first frame image from the target motion effect video and form a text mask marker image; Establish a unit to create a mapping relationship between the text content in each video cell of the target motion effect video and the text mask markup image based on the text index information; The combination unit is used to combine the target motion video, text mask marker image and text index information to form a training dataset of the target text content.
[0057] For other examples of this embodiment, please refer to the examples above, which will not be repeated here.
[0058] Figure 4 This is a schematic diagram of an optional electronic device according to an embodiment of this application, such as... Figure 4 As shown, it includes a processor 402, a communication interface 404, a memory 406, and a communication bus 408. The processor 402, communication interface 404, and memory 406 communicate with each other via the communication bus 408. Memory 406 is used to store computer programs; When processor 402 executes a computer program stored in memory 406, it performs the following steps: Build a copywriting library and a background library; Based on the text animation template, randomly select target text content from the text library to generate a corresponding black background text animation video; A grid partitioning strategy is used to fuse a black-background text animation video with a target background image randomly selected from a background library to generate a target animation video. Extract the first frame mask image of the target motion effect video to obtain the text mask marker image of the target motion effect video, and combine it with the text index information of the target motion effect video to generate a training dataset of the target text content.
[0059] Optionally, in this embodiment, the communication bus can be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, etc. This communication bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 4 The symbol is represented by a single thick line, but this does not indicate that there is only one bus or one type of bus. The communication interface is used for communication between the aforementioned electronic devices and other devices.
[0060] The memory may include RAM, or non-volatile memory, such as at least one disk storage device. Optionally, the memory may also be at least one storage device located remotely from the aforementioned processor.
[0061] As an example, the memory 406 described above may include, but is not limited to, the construction module 302, the first generation module 304, the second generation module 306, and the third generation module 308 in the text animation data generation device. Furthermore, it may include, but is not limited to, other module units in the text animation data generation device, which will not be elaborated upon in this example.
[0062] The processor mentioned above can be a general-purpose processor, including but not limited to: CPU (Central Processing Unit), NP (Network Processor), etc.; it can also be DSP (Digital Signal Processor), ASIC (Application Specific Integrated Circuit), FPGA (Field-Programmable Gate Array) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0063] Optionally, specific examples in this embodiment can refer to the examples described in the above embodiments, and will not be repeated here.
[0064] Those skilled in the art will understand that Figure 4 The structure shown is for illustrative purposes only. The device that implements the above method for generating text animation data can be a terminal device, such as a smartphone (e.g., an Android phone, an iOS phone), a tablet computer, a PDA, a mobile internet device (MID), a PAD, or other terminal devices. Figure 4 This does not limit the structure of the aforementioned electronic devices. For example, the electronic device may also include components that are more... Figure 4 The more or fewer components shown (such as network interfaces, display devices, etc.), or having the same Figure 4 The different configurations shown.
[0065] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by a program instructing the hardware related to the terminal device. The program can be stored in a computer-readable storage medium, which may include: flash drive, ROM, RAM, disk or optical disk, etc.
[0066] According to another aspect of the embodiments of this application, a computer-readable storage medium is also provided, wherein a computer program is stored in the computer program, which, when executed by a processor, performs the steps in the above-described method for generating text animation data.
[0067] Optionally, in this embodiment, those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by a program instructing the hardware related to the terminal device. The program can be stored in a computer-readable storage medium, which may include: flash drive, read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.
[0068] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.
[0069] If the integrated units in the above embodiments are implemented as software functional units and sold or used as independent products, they can be stored in the aforementioned computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause one or more computer devices (which may be personal computers, servers, or network devices, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application.
[0070] In the above embodiments of this application, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0071] In the several embodiments provided in this application, it should be understood that the disclosed client can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces, indirect coupling or communication connection between units or modules, and may be electrical or other forms.
[0072] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0073] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0074] The above description is only a preferred embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this application, and these improvements and modifications should also be considered within the scope of protection of this application.
Claims
1. A method for generating text animation data, characterized in that, include: Build a copywriting library and a background library; Based on the text animation template, target text content is randomly selected from the text library to generate a corresponding black background text animation video; A grid partitioning strategy is used to fuse the black-background text animation video with a target background image randomly selected from the background library to generate a target animation video. Extract the first frame mask image of the target motion effect video to obtain the text mask marker image of the target motion effect video, and combine it with the text index information of the target motion effect video to generate a training dataset of the target text content, wherein the text index information is the target text content, and the training dataset includes the target motion effect video, the text mask marker image and the text index information.
2. The method according to claim 1, characterized in that, Building a copywriting library and background library includes: The copywriting library is constructed by collecting or generating multiple text content through the Wenshengwen model. The text content includes phrases without punctuation and phrases with punctuation, covering different industry scenarios and various emotional tendencies. The language types of the text content include Chinese, English, and numbers.
3. The method according to claim 1, characterized in that, Building a copywriting library and background library includes: Multiple background images of different scene types are generated by collecting or creating large-scale models; The background library is constructed by removing text from each background image using an optical character recognition (OCR) detection algorithm. Extract the theme color of each background image in the background library, and classify all background images in the background library according to the theme color.
4. The method according to claim 1, characterized in that, Based on the text animation template, target text content is randomly selected from the text library to generate a corresponding black-background text animation video, including: The text animation template is preset, wherein the text animation template includes various types of text animations; Randomly select target text content from the text library and apply the target text content to the text animation template to generate a corresponding black background text animation video.
5. The method according to claim 4, characterized in that, In the process of generating the corresponding black-background text animation video, the method also includes: Randomly generate the text color of the target text content; Randomly add a stroke effect and stroke color to the target text content; Randomly select a font type and size for the target text content from the system font library.
6. The method according to claim 1, characterized in that, Using a grid partitioning strategy, the black-background text animation video is fused with a target background image randomly selected from the background library to generate the target animation video, including: Divide the canvas into M grids and assign a unique identifier and coordinates to each grid to form a grid information list; N videos are randomly selected from the black-background text animation videos and randomly assigned to N grids as video grids. The remaining grids are defined as interference grids. Each video grid is randomly marked with whether the text is of the button type. Wherein, N is less than M. For each interference cell, generate random static element information, wherein the static element information includes text content, font type, font size, text color, stroke color, and button color; A target background image is randomly selected from the background library as the background base image, and the static element information of each interference grid is drawn onto the grid position corresponding to the background base image. For each video cell, extract the text region mask frame by frame, and for video cells marked as button type, fill the text region determined based on the text region mask of the first frame with a randomly generated button color. The processed video of each video frame is superimposed onto the grid position corresponding to the background image frame by frame to obtain the target motion effect video.
7. The method according to any one of claims 1 to 6, characterized in that, Extract the first frame mask image of the target motion effect video to obtain the text mask marker image of the target motion effect video, and combine it with the text index information of the target motion effect video to generate the training dataset of the target text content, including: Extract the text mask from the first frame of the target motion video to form the text mask marker image; Based on the text index information, a mapping relationship is established between the text content in each video cell of the target motion effect video and the text mask mark image; The target motion video, the text mask image, and the text index information are combined to form a training dataset for the target text content.
8. A device for generating text animation data, characterized in that, include: The building module is used to build the copywriting library and background library; The first generation module is used to randomly select target text content from the text library based on the text animation template and generate a corresponding black background text animation video. The second generation module is used to use a grid division strategy to fuse the black text animation video with a target background image randomly selected from the background library to generate a target animation video. The third generation module is used to extract the first frame mask image of the target motion effect video, obtain the text mask mark image of the target motion effect video, and combine it with the text index information of the target motion effect video to generate a training dataset of the target text content, wherein the text index information is the target text content, and the training dataset includes the target motion effect video, the text mask mark image and the text index information.
9. A computer-readable storage medium storing a computer program, characterized in that, The computer program is executed by the processor to perform the method described in any one of claims 1 to 7.
10. An electronic device comprising a memory and a processor, characterized in that, The memory stores a computer program, and the processor is configured to execute the method described in any one of claims 1 to 7 through the computer program.