Method, apparatus, device and storage medium for generating hand-drawn videos based on text
By identifying feature labels and segmenting the image of text materials, a hand-painted video based on preset hand-painted paths is generated, which solves the problems of low efficiency and high technical requirements of hand-painted video production in the prior art, and achieves low-cost and high-efficiency video generation.
Patent Information
- Application Number
- CN202110736329.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-06-30
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2041-06-30
AI Technical Summary
The existing hand-painted video production methods require a lot of labor and time costs, and the post-editing requirements for high technical requirements, resulting in low production efficiency and high difficulty.
By obtaining the feature labels of text materials, obtaining matching portrait pictures from the preset portrait database, and performing video frame configuration processing on the picture, dividing the picture into several video frames, and generating hand-painted videos in association with the preset hand-painted path.
It reduces labor costs, improves the efficiency of video generation, simplifies the production process, and reduces dependence on technical requirements.
Smart Images

Figure CN113392231B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image processing technology, and in particular, to a method, apparatus, device, and storage medium for generating a hand-drawn video based on text. Background Art
[0002] In the business service industry, in order to improve the service experience of users and the service efficiency of business personnel, business personnel usually introduce business products to users by combining some hand-drawn videos. However, the current existing method for producing hand-drawn videos is generally to shoot the entire process of drawing, and then combine audio and video through post-editing. Such a method often requires a large amount of labor cost and time cost, and the post-editing has high technical requirements for manpower, and requires the producer to have certain professional shooting ability and video editing ability, which results in low production efficiency and high difficulty of hand-drawn videos. Summary of the Invention
[0003] In view of this, the embodiments of the present application provide a method, apparatus, device, and storage medium for generating a hand-drawn video based on text, which can generate a hand-drawn video by performing video frame configuration on portrait pictures, with low labor cost, simple to use, and high efficiency in generating videos.
[0004] The first aspect of the embodiments of the present application provides a method for generating a hand-drawn video based on text, including:
[0005] Obtain text materials corresponding to the video to be generated, and use a text element recognition model to recognize the text materials to obtain element tags of the text materials;
[0006] According to the element tags of the text materials, obtain portrait pictures matching the element tags from a preset portrait database;
[0007] Perform video frame configuration processing on the portrait pictures, and divide the portrait pictures into several video frames;
[0008] Associate the several video frames with coordinate point information corresponding to a preset hand-drawn path respectively to generate a hand-drawn video based on the preset hand-drawn path.
[0009] Combined with the first aspect, in the first possible implementation manner of the first aspect, the step of performing video frame configuration processing on the portrait pictures and dividing the portrait pictures into several video frames includes:
[0010] Obtain audio information of the text materials, and determine the video duration of the portrait pictures according to the audio information of the text materials;
[0011] Calculate the number of video frames of the portrait picture according to the preset playback frame rate and the video duration of the portrait picture, and perform video frame configuration processing on the portrait picture according to the calculated number of video frames.
[0012] Combined with the first possible implementation manner of the first aspect, in the second possible implementation manner of the first aspect, the step of performing video frame configuration processing on the portrait picture according to the calculated number of video frames includes:
[0013] Obtain the number of coloring effect frames of the portrait picture;
[0014] Calculate the number of segmentation frames of the portrait picture according to the number of video frames of the portrait picture and the number of coloring effect frames of the portrait picture;
[0015] Divide the portrait picture into the number of segmentation frames of video frames, and configure the number of coloring effect frames of default video frames behind the last video frame obtained by segmentation. Among them, the number of segmentation frames of video frames is configured to display the portrait picture, and the number of coloring effect frames of default video frames is configured to perform coloring processing on the portrait picture.
[0016] Combined with the second possible implementation manner of the first aspect, in the third possible implementation manner of the first aspect, after the step of dividing the portrait picture into the number of segmentation frames of video frames and configuring the number of coloring effect frames of default video frames behind the last video frame obtained by segmentation, it further includes:
[0017] Perform portrait color value configuration on the number of coloring effect frames of default video frames, where the portrait color value of the previous default video frame is greater than the portrait color value of the next default video frame.
[0018] Combined with the first aspect, in the fourth possible implementation manner of the first aspect, after the step of associating the several video frames with the coordinate point information corresponding to the preset hand-drawn path to generate a hand-drawn video based on the preset hand-drawn path, it further includes:
[0019] Obtain the audio information of the text material, and add audio subtitles to the hand-drawn video according to the audio information.
[0020] Combined with the first aspect, in the fifth possible implementation manner of the first aspect, before the step of obtaining the portrait picture matching the element label from the preset portrait database according to the element label of the text material, it further includes:
[0021] Construct a multi-level material label classification in the preset portrait database to classify and store the portrait pictures in the preset portrait database according to the multi-level material label classification.
[0022] Combined with the fifth possible implementation manner of the first aspect, in the sixth possible implementation manner of the first aspect, the step of obtaining a portrait picture matching the element label from a preset portrait database according to the element label of the text material includes:
[0023] Compare the element label of the text material with the multi-level material label classification in the preset portrait database to identify whether there is a unique material label classification in the multi-level material label classification that matches the element label of the text material;
[0024] If so, obtain a portrait picture matching the element label of the text material from the preset portrait database according to the unique material label classification; otherwise, perform context analysis on the text material based on the element label to determine the unique matching material label classification, and obtain a portrait picture matching the element label of the text material from the preset portrait database according to the uniquely matched material label classification.
[0025] A second aspect of the embodiments of the present application provides a device for generating a hand-drawn video based on text. The device for generating a hand-drawn video based on text includes:
[0026] A text material acquisition module, configured to acquire text material corresponding to a video to be generated, and use a text element recognition model to recognize the text material to obtain the element label of the text material;
[0027] A portrait picture acquisition module, configured to obtain a portrait picture matching the element label from a preset portrait database according to the element label of the text material;
[0028] A portrait picture segmentation module, configured to perform video frame configuration processing on the portrait picture and segment the portrait picture into several video frames;
[0029] A video generation module, configured to associate the several video frames with coordinate point information corresponding to a preset hand-drawn path respectively to generate a hand-drawn video based on the preset hand-drawn path.
[0030] A third aspect of the embodiments of the present application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the electronic device. When the processor executes the computer program, the method for generating a hand-drawn video based on text provided in the first aspect is implemented.
[0031] A fourth aspect of the embodiments of the present application provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method for generating a hand-drawn video based on text provided in the first aspect is implemented.
[0032] The method, device, equipment and storage medium for generating a hand-drawn video based on text provided by the embodiments of the present application have the following beneficial effects:
[0033] The method described in the present application obtains the text material corresponding to the video to be generated, uses a text element recognition model to recognize the text material, and obtains the element tags of the text material; according to the element tags of the text material, obtains the portrait pictures matching the element tags from a preset portrait database; performs video frame configuration processing on the portrait pictures, and divides the portrait pictures into several video frames; associates the several video frames with the coordinate point information corresponding to the preset hand-drawn path respectively to generate a hand-drawn video based on the preset hand-drawn path. This method generates a hand-drawn video by performing video frame configuration on portrait pictures, with low labor cost, being simple and easy to use, having little dependence on materials, and high efficiency in generating videos. Description of the Drawings
[0034] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0035] Figure 1 It is the basic method flow chart of a method for generating a hand-drawn video based on text provided by the embodiments of the present application;
[0036] Figure 2 It is a schematic diagram of the hand-drawn video display process in the method for generating a hand-drawn video based on text provided by the embodiments of the present application;
[0037] Figure 3 It is a schematic diagram of a method flow for video frame configuration in the method for generating a hand-drawn video based on text provided by the embodiments of the present application;
[0038] Figure 4 It is another schematic diagram of a method flow for video frame configuration in the method for generating a hand-drawn video based on text provided by the embodiments of the present application;
[0039] Figure 5 It is another schematic diagram of the hand-drawn video display process in the method for generating a hand-drawn video based on text provided by the embodiments of the present application;
[0040] Figure 6 It is a schematic diagram of a method flow for obtaining portrait pictures in the method for generating a hand-drawn video based on text provided by the embodiments of the present application;
[0041] Figure 7 This is the basic structural block diagram of a device for generating hand-drawn videos based on text provided by an embodiment of the present application;
[0042] Figure 8 This is the basic structural block diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0043] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0044] Please refer to Figure 1 , Figure 1 This is the basic method flow chart of a method for generating hand-drawn videos based on text provided by an embodiment of the present application. Details are as follows:
[0045] Step S11: Obtain the text material corresponding to the video to be generated, and use a text element recognition model to recognize the text material to obtain the element labels of the text material.
[0046] In this embodiment, the text element recognition model is a neural network model trained to recognize entity word elements in text. In this embodiment, the text material for generating a video can be a paragraph or a sentence. By inputting the paragraph or sentence into the text element recognition model, the text element recognition model can extract the entity words that appear in the input paragraph or sentence, and output the entity words extracted from the input text as element labels.
[0047] Exemplarily, assume that the text material input into the text element recognition model is "As long as you are willing to work hard, reaching the peak of life is not a dream", then the text element recognition model can obtain two element labels, namely "you" and "peak", based on this text material. It can be understood that when the text material is a text with a large amount of content, such as the script of a complete story, at this time, a text summary extraction algorithm can also be used to extract the summary of the script, and then the summary is split according to a preset text splitting strategy to obtain several sentences, and then the text element recognition model is used to perform recognition processing on the several sentences respectively, so as to obtain the element labels of each sentence, so that in the subsequent steps, hand-drawn videos can be generated by sentence segmentation, and then the segmented hand-drawn videos can be spliced to obtain a complete hand-drawn video.
[0048] Step S12: According to the element labels of the text material, obtain the portrait pictures matching the element labels from a preset portrait database.
[0049] In this embodiment, a portrait database is pre-constructed. This portrait database is used to store portrait pictures corresponding to various element tags, and a correspondence relationship is established between each portrait picture and its corresponding element tag. Among them, one portrait picture can correspond to multiple element tags. For example, for a portrait of a person, the corresponding element tags can be "me", "you", "her", "him", "person's name", and so on. In this embodiment, the portrait pictures stored in the portrait database can be configured as transparent pictures (such as line drawings without a background) to make the hand-drawn effect of the generated hand-drawn video better. In this embodiment, after obtaining the element tags of the text material through the text element recognition model, the preset portrait database can be queried according to the element tags of the text material, and based on the correspondence relationship between the portrait pictures and the element tags, the portrait pictures matching the element tags of each text material can be obtained from the portrait database.
[0050] Step S13: Perform video frame configuration processing on the portrait picture, and divide the portrait picture into several video frames.
[0051] In this embodiment, the hand-drawn video is a video synthesized by multiple video frames. In this embodiment, the portrait picture is segmented according to the user's requirements to divide the portrait picture into several video frames for synthesizing the hand-drawn video. Exemplarily, assuming that the number of video frames corresponding to the text material is 100 frames, the portrait picture can be first pre-processed by edge extraction, binarization, pixel point dilation, etc. of the image, and then all pixel points of the portrait picture can be obtained by traversing the portrait picture. Then, all pixel points are sorted according to a preset hand-drawn path, and then all pixel points are evenly distributed to the 100 video frames according to the sorting. The corresponding pixel points are drawn in each video frame, so as to divide the portrait picture into 100 video frames.
[0052] Step S14: Associate the several video frames with the coordinate point information corresponding to the preset hand-drawn path respectively to generate a hand-drawn video based on the preset hand-drawn path.
[0053] In this embodiment, please refer to Figure 2 , Figure 2 which is a schematic diagram of the display process of the hand-drawn video in the method for generating a hand-drawn video based on text provided by the embodiment of the present application. As Figure 2As shown, the hand-drawn path is defaulted to a Z-shaped path from the upper left to the lower right. In this embodiment, the length of the hand-drawn path can be set by the user customarily. After setting the preset hand-drawn path, several layout points for displaying video frames can be set on the preset hand-drawn path according to the length of the preset hand-drawn path, and each layout point has corresponding coordinate point information. Exemplarily, the layout points set on the preset hand-drawn path can be evenly spaced. For example, when associating several previously obtained video frames with the coordinate information corresponding to the preset hand-drawn path respectively, first associate the first video frame of the portrait picture with the first layout point of the preset hand-drawn path for displaying video frames, and associate the last video frame of the portrait picture with the last layout point of the preset hand-drawn path for displaying video frames. Then, evenly distribute the remaining video frames of the portrait picture at equal intervals on the layout points of the preset hand-drawn path. Thus, as Figure 2 shown, in the generated hand-drawn video, when the hand-drawn action reaches the layout point of the preset hand-drawn path for displaying video frames, the video frame associated with the layout point is displayed.
[0054] As can be seen from the above, the method for generating a hand-drawn video based on text provided in this embodiment is based on text materials. By using a text element recognition model to identify the element tags of the text materials, and then obtaining portrait pictures matching the element tags from the portrait database according to the element tags. Furthermore, through video frame configuration processing of the portrait pictures, the portrait pictures are segmented into several video frames, and finally the several video frames are respectively associated with the coordinate point information corresponding to the preset hand-drawn path to generate a hand-drawn video based on the preset path. In this way, a hand-drawn video is generated by performing video frame configuration on the portrait pictures, and the hand-drawn video is automatically generated by a computer, with low labor cost, simple to use, less dependent on materials, and high efficiency in generating videos.
[0055] In some embodiments of the present application, please refer to Figure 3 , Figure 3 which is a schematic flowchart of a method for video frame configuration in the method for generating a hand-drawn video based on text provided in the embodiments of the present application. Details are as follows:
[0056] Step S31: Obtain the audio information of the text material, and determine the video duration of the portrait picture according to the audio information of the text material;
[0057] Step S32: Calculate the number of video frames of the portrait picture according to the preset playback frame rate and the video duration of the portrait picture, and perform video frame configuration processing on the portrait picture according to the calculated number of video frames.
[0058] In this embodiment, since the hand-drawn video generated based on the text material is played in combination with audio, and the audio is a voice file reading the text, it is manifested as follows: when the first word in the text material is read out by the audio, the hand-drawn video starts to play, and when the last word in the text material is read by the audio, the hand-drawn video ends. In this embodiment, by obtaining the audio information of the text material, the duration of the audio can be obtained, and the duration of the audio is set as the video duration of the portrait picture corresponding to the text material in the hand-drawn video. By presetting and fixing the playback frame rate of a video (the number of frames displayed per second when the video is played), the video frames of the portrait picture are calculated based on the fixed playback frame rate and the video duration of the portrait picture obtained from the audio information. Specifically, the video frames can be obtained by dividing the video duration by the frame rate. In this embodiment, by determining the duration of the hand-drawn video to be produced according to the audio duration cited in the hand-drawn video, it is possible to generate a hand-drawn video with a 1:1 time ratio between the hand-drawn video and the audio, without the need for post-production such as video editing, audio editing, and audio-video time alignment, solving the problem that the technical requirements for manpower in post-editing are relatively high and saving labor costs.
[0059] In some embodiments of the present application, please refer to Figure 4 , Figure 4 which is another schematic flowchart of the video frame configuration in the method for generating a hand-drawn video based on text provided by the embodiments of the present application. Details are as follows:
[0060] Step S41: Obtain the number of coloring special effect frames of the portrait picture;
[0061] Step S42: Calculate the number of segmented frames of the portrait picture according to the video frames of the portrait picture and the number of coloring special effect frames of the portrait picture;
[0062] Step S43: Divide the portrait picture into the number of video frames of the segmented frames, and configure the number of default video frames of the coloring special effect frames behind the last video frame obtained by segmentation. The number of default video frames of the coloring special effect frames is configured to perform coloring processing on the portrait picture.
[0063] In this embodiment, after the hand-drawn video draws a complete portrait picture, it is also possible to perform coloring special effect processing on the complete portrait picture in the hand-drawn video. Specifically, when configuring the video frames of the portrait picture, the number of segmented frames for displaying the portrait picture can be configured first in the order of segmentation, and then the number of default video frames of the coloring special effect frames can be configured behind the last video frame for displaying the portrait picture, thereby completing the configuration processing of the video frames.
[0064] In some embodiments of the present application, by configuring the portrait color values of the default video frames, the coloring of the portrait pictures can be achieved. In this embodiment, the first default video frame to the last default video frame following the last video frame used to display the portrait picture can be configured with image color values from light to dark in sequence, that is, the portrait color value of the previous default video frame is greater than that of the subsequent default video frame, thereby achieving the effect of gradient coloring. Exemplarily, please refer to Figure 5 , Figure 5 which is another schematic diagram of the process of displaying a hand-drawn video in the method for generating a hand-drawn video based on text provided by an embodiment of the present application. As Figure 5 shown, Figure 5-1 to Figure 5-2 represent the first video frame used to display the portrait picture to the last video frame used to display the portrait picture, which is the drawing process of the portrait picture, Figure 5-3 to Figure 5-4 represent the first default video frame to the last default video frame, which is the coloring process of the portrait picture.
[0065] In some embodiments of the present application, after obtaining the audio information of the text material, the time nodes corresponding to each byte read out in the audio can also be identified, and according to these time nodes, the video frames corresponding to these time nodes can be obtained from the hand-drawn video, and then the subtitles of these bytes can be added to the video frames to synthesize an audio-visual video.
[0066] In some embodiments of the present application, in a preset portrait database, a multi-level material label classification can be constructed to classify different meaning representations represented by the same entity word. For example, for a person, based on the multi-level material label classification, the portrait pictures of a person can be correspondingly divided into portrait pictures of male adults, portrait pictures of female adults, portrait pictures of male children, portrait pictures of female children, and so on. Another example is an apple, which can refer to an apple in fruits or an iPhone, an electronic product. Through the multi-level material label classification constructed in this embodiment, the two can be classified and stored.
[0067] Exemplarily, in some embodiments of the present application, based on the multi-level material label classification constructed in the preset portrait database, please refer to Figure 6 , Figure 6 which is a schematic diagram of a method flow for obtaining a portrait picture in the method for generating a hand-drawn video based on text provided by an embodiment of the present application. Details are as follows:
[0068] Step S61: Compare the element labels of the text material with the multi-level material label classification in the preset portrait database to identify whether there is a unique material label classification in the multi-level material label classification that matches the element labels of the text material;
[0069] Step S62: If so, obtain portrait pictures matching the element tags of the text material from a preset portrait database according to the unique material tag classification; otherwise, perform context analysis on the text material based on the element tags to determine the material tag classification uniquely matching the element tags, and obtain portrait pictures matching the element tags of the text material from the preset portrait database according to the uniquely matching material tag classification.
[0070] In this embodiment, when obtaining portrait pictures matching the element tags from a preset portrait database according to the multi-level material tag classification constructed in the preset portrait database, there may be a situation where the obtained portrait pictures are not unique. In this regard, this embodiment can compare the requirement tags of the text material with the multi-level material tag classification in the preset portrait database to identify whether there is a unique material tag classification in the multi-level material tag classification that matches the element tags of the text material. If there is a unique material tag classification that matches the element tags of the text material, it means that a unique portrait picture can be obtained, and the obtained unique portrait picture is the portrait picture matching the element tags of the text material. For example, if the text material is "There is a little boy standing on the stone steps", the element tag "little boy" can be obtained at this time. The element tag "little boy" can accurately match the classification of the portrait picture of the little boy from the multi-level material tag classification divided by the portrait pictures of people. Thus, the portrait picture of the little boy is obtained as the portrait picture matching the element tag "little boy". If the material tag classification matching the element tags of the text material is not unique, it is necessary to perform context analysis on the text material based on the element tags to determine the material tag classification uniquely matching the element tags, so as to obtain portrait pictures matching the element tags of the text material from the preset portrait database according to the uniquely matching material tag classification. For example, if the text material is "She wore a new red dress to work today", the element tag "She" can be obtained at this time. The element tag "She" can only determine that the material tag classification is female characters, and it is impossible to distinguish between adult women and child women. At this time, context analysis can be performed on the text material based on the element tag to determine the material tag classification uniquely matching the element tag. From the word "work" obtained in the text material, it can be inferred that the element tag "She" refers to an adult woman who has already started working. Thus, the classification of the portrait picture of an adult woman can be matched, and the portrait picture of an adult woman is obtained as the portrait picture matching the element tag "She".
[0071] It should be understood that the sequence numbers of the steps in the above embodiments do not indicate the order of execution. The order of execution of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.
[0072] In some embodiments of the present application, please refer to Figure 7 , Figure 7 , which is a basic structural block diagram of a device for generating a hand-drawn video based on text provided by an embodiment of the present application. In this embodiment, each unit included in the device is used to execute each step in the above method embodiment. For specific details, please refer to the relevant descriptions in the above method embodiment. For the sake of convenience of description, only the part related to this embodiment is shown. As Figure 7 shown, the device for generating a hand-drawn video based on text includes: a text material acquisition module 71, a portrait picture acquisition module 72, a portrait picture segmentation module 73, and a video generation module 74. Among them: the text material acquisition module 71 is used to acquire the text material corresponding to the video to be generated, and use a text element recognition model to recognize the text material to obtain the element tags of the text material; the portrait picture acquisition module 72 is used to acquire portrait pictures matching the element tags from a preset portrait database according to the element tags of the text material; the portrait picture segmentation module 73 is used to perform video frame configuration processing on the portrait pictures, and segment the portrait pictures into several video frames; the video generation module 74 is used to associate the several video frames with the coordinate point information corresponding to a preset hand-drawn path respectively, and generate a hand-drawn video based on the preset hand-drawn path.
[0073] It should be understood that the above device for generating a hand-drawn video based on text corresponds one-to-one with the above method for generating a hand-drawn video based on text, and details are not described herein again.
[0074] In some embodiments of the present application, please refer to Figure 8 , Figure 8 , which is a basic structural block diagram of an electronic device provided by an embodiment of the present application. As Figure 8 shown, the electronic device 8 in this embodiment includes: a processor 81, a memory 82, and a computer program 83 stored in the memory 82 and executable on the processor 81, such as a program for the method of generating a hand-drawn video based on text. When the processor 81 executes the computer program 83, it implements the steps in each of the above embodiments of the method for generating a hand-drawn video based on text. Alternatively, when the processor 81 executes the computer program 83, it implements the functions of each module in the corresponding embodiment of the above device for generating a hand-drawn video based on text. For specific details, please refer to the relevant descriptions in the embodiments, and details are not described herein.
[0075] Exemplarily, the computer program 83 may be divided into one or more modules (units). The one or more modules are stored in the memory 82 and executed by the processor 81 to complete this application. The one or more modules may be a series of computer program instruction segments capable of performing specific functions, and these instruction segments are used to describe the execution process of the computer program 83 in the electronic device 8. For example, the computer program 83 may be divided into a text material acquisition module, an image picture acquisition module, an image picture segmentation module, and a video generation module, and the specific functions of each module are as described above.
[0076] The turntable device may include, but is not limited to, the processor 81 and the memory 82. Those skilled in the art can understand that Figure 8 merely examples of the electronic device 8, which do not constitute a limitation on the electronic device 8. It may include more or fewer components than shown in the figure, or combine certain components, or different components. For example, the turntable device may also include an input / output device, a network access device, a bus, etc.
[0077] The so-called processor 81 may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0078] The memory 82 may be an internal storage unit of the electronic device 8, such as the hard disk or memory of the electronic device 8. The memory 82 may also be an external storage device of the electronic device 8, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the electronic device 8. Further, the memory 82 may also include both the internal storage unit and the external storage device of the electronic device 8. The memory 82 is used to store the computer program and other programs and data required by the turntable device. The memory 82 may also be used to temporarily store the data that has been output or will be output.
[0079] It should be noted that for the information interaction, execution process, etc. between the above-mentioned devices / units, since they are based on the same concept as the method embodiments of the present application, for their specific functions and the technical effects brought, reference can be specifically made to the method embodiment part, and details are not described herein again.
[0080] The embodiments of the present application further provide a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in the above-mentioned various method embodiments can be implemented. In this embodiment, the computer-readable storage medium can be non-volatile or volatile.
[0081] The embodiments of the present application provide a computer program product. When the computer program product runs on a mobile terminal, the mobile terminal can execute the steps in the above-mentioned various method embodiments.
[0082] Those skilled in the art can clearly understand that for the convenience and simplicity of description, only the above-mentioned division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiment can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of the present application. The specific working process of the units and modules in the above system can refer to the corresponding process in the foregoing method embodiments, and details are not described herein again.
[0083] If the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above-described embodiment methods of the present application, it can also be completed by instructing relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-described various method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.
[0084] In the above embodiments, the descriptions of the various embodiments have their own emphases. For the parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0085] The above-described embodiments are only used to illustrate the technical solutions of the present application, and are not intended to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included in the protection scope of the present application.
Claims
1. A method for generating a hand-drawn video based on text, characterized in that, it includes: Obtain the text material corresponding to the video to be generated, and use a text element recognition model to recognize the text material to obtain the element tags of the text material; According to the element tags of the text material, obtain portrait pictures matching the element tags from a preset portrait database; Perform video frame configuration processing on the portrait pictures, and divide the portrait pictures into several video frames; Associate the several video frames with the coordinate point information corresponding to a preset hand-drawn path respectively to generate a hand-drawn video based on the preset hand-drawn path; Among them, the step of dividing the portrait pictures into several video frames includes: first performing image preprocessing on the portrait pictures, and the preprocessing includes edge extraction, binarization, and pixel point dilation. By traversing the portrait pictures, all pixel points of the portrait pictures are obtained, and all pixel points are sorted according to a preset hand-drawn path, and all pixel points are evenly divided into the several video frames according to the sorting, and the corresponding pixel points are drawn in each video frame, so as to divide the portrait pictures into the several video frames.
2. The method for generating a hand-drawn video based on text according to claim 1, characterized in that, The step of performing video frame configuration processing on the portrait pictures and dividing the portrait pictures into several video frames includes: Obtain the audio information of the text material, and determine the video duration of the portrait pictures according to the audio information of the text material; Calculate the number of video frames of the portrait pictures according to a preset playback frame rate and the video duration of the portrait pictures, and perform video frame configuration processing on the portrait pictures according to the calculated number of video frames.
3. The method for generating a hand-drawn video based on text according to claim 2, characterized in that, The step of performing video frame configuration processing on the portrait pictures according to the calculated number of video frames includes: Obtain the number of coloring special effect frames of the portrait pictures; Calculate the number of divided frames of the portrait pictures according to the number of video frames of the portrait pictures and the number of coloring special effect frames of the portrait pictures; Divide the portrait pictures into the number of divided frames of video frames, and configure the number of coloring special effect frames of default video frames behind the last video frame obtained by division. Among them, the number of divided frames of video frames is configured to display the portrait pictures, and the number of coloring special effect frames of default video frames is configured to perform coloring processing on the portrait pictures.
4. The method for generating a hand-drawn video based on text according to claim 3, characterized in that, After the step of dividing the portrait pictures into the number of divided frames of video frames and configuring the number of coloring special effect frames of default video frames behind the last video frame obtained by division, it further includes: Configure the portrait color values of the number of coloring special effect frames of default video frames, where the portrait color value of the previous default video frame is greater than the portrait color value of the next default video frame.
5. The method for generating a hand-drawn video based on text according to claim 1, characterized in that, After the step of associating the several video frames with the coordinate point information corresponding to the preset hand-drawn path respectively to generate a hand-drawn video based on the preset hand-drawn path, the method further includes: Obtaining the audio information of the text material, and adding audio subtitles to the hand-drawn video according to the audio information.
6. The method for generating a hand-drawn video based on text according to claim 1, wherein, Before the step of obtaining a portrait picture matching the element label from a preset portrait database according to the element label of the text material, the method further includes: Constructing a multi-level material label classification in the preset portrait database to classify and store the portrait pictures in the preset portrait database according to the multi-level material label classification.
7. The method for generating a hand-drawn video based on text according to claim 5, wherein, The step of obtaining a portrait picture matching the element label from a preset portrait database according to the element label of the text material includes: Comparing the element label of the text material with the multi-level material label classification in the preset portrait database to identify whether there is a unique material label classification in the multi-level material label classification that matches the element label of the text material; If so, obtaining a portrait picture matching the element label of the text material from the preset portrait database according to the unique material label classification, otherwise, performing context analysis on the text material based on the element label to determine a uniquely matching material label classification, and obtaining a portrait picture matching the element label of the text material from the preset portrait database according to the uniquely matching material label classification.
8. An apparatus for generating a hand-drawn video based on text, wherein, it includes: A text material acquisition module, configured to acquire the text material corresponding to the video to be generated, and use a text element recognition model to recognize the text material to obtain the element label of the text material; A portrait picture acquisition module, configured to obtain a portrait picture matching the element label from a preset portrait database according to the element label of the text material; A portrait picture segmentation module, configured to perform video frame configuration processing on the portrait picture, and segment the portrait picture into several video frames; A video generation module, configured to associate the several video frames with the coordinate point information corresponding to a preset hand-drawn path respectively to generate a hand-drawn video based on the preset hand-drawn path; wherein, the portrait picture segmentation module is specifically configured to: first perform image preprocessing on the portrait picture, the preprocessing includes edge extraction, binarization, and pixel point dilation, traverse all pixel points of the portrait picture to obtain all pixel points of the portrait picture, sort all pixel points according to a preset hand-drawn path, evenly distribute all pixel points to the several video frames according to the sorting, and draw the corresponding distributed pixel points in each video frame, thereby segmenting the portrait picture into the several video frames.
9. An electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein, when the processor executes the computer program, the method according to any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium storing a computer program, wherein, when the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Method for fast combining hand-painted video elements in video
CN107592565A
Animation draft generation method and device based on character paragraphs
CN112270197A