Method, apparatus, computing device and storage medium for generating a video using text

By building a video template library and automatically generating short videos by parsing dynamic resource parameters, the problems of high cost and uncontrollable quality in existing technologies have been solved, achieving efficient generation of diverse and high-quality short videos and improving advertising effectiveness.

CN115294491BActive Publication Date: 2026-02-10BEIJING CHESHANGHUI SOFTWARE
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210793156.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-05
Publication Date
2026-02-10
Estimated Expiration
2042-07-05

AI Technical Summary

Technical Problem

Existing short video generation solutions are costly and have uncontrollable quality, making it difficult to save manpower and resources while ensuring the quality of the generated videos.

Method used

By building a video template library, selecting video templates based on business scenarios, parsing dynamic resource parameters, obtaining dynamic resources from text, and writing them into video templates to generate short videos.

Benefits of technology

This approach saves manpower and resources while improving the diversity and quality of short video generation, thereby increasing ad clicks and conversion rates.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115294491B_ABST
    Figure CN115294491B_ABST
Patent Text Reader

Abstract

The application discloses a method and device for generating a video from a text, a computing device and a storage medium. The method for generating a video from a text comprises the steps of: selecting any video template from a video template library based on a business scenario, wherein the video template library comprises at least one video template under different business scenarios; determining dynamic resource parameters in the video template, wherein the dynamic resource parameters comprise a display position, a display time and a display sequence of a dynamic resource; obtaining corresponding dynamic resources from the text based on the dynamic resource parameters; and writing the obtained dynamic resources into the video template to generate a video.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to the technical field of computer networks, and particularly relates to a method and device for generating a video from text, a computing device, and a storage medium. BACKGROUND

[0002] With the popularity of mobile terminals such as mobile phones and tablets and the increasing speed of the network, short videos have gradually become popular. Compared with static text and pictures, short videos are more popular among people. In addition, short videos also play an important role in online advertising and other fields. Compared with an article form of advertising, a short video form of advertising can bring more traffic and improve the conversion rate of advertising.

[0003] Existing short video production schemes are mostly based on manual production of short videos, which requires a large investment in manpower, material resources, and time costs. Another scheme is to automatically generate short videos using a deep learning GAN network. However, the video quality of the short videos generated based on deep learning is extremely dependent on the training data, and cannot guarantee 100% reliability. It is possible to produce short videos with poor overall effects. In other words, the generation quality of short videos is uncontrollable.

[0004] Therefore, there is a need for a short video generation scheme that can save costs and control the quality of short videos. SUMMARY

[0005] The present disclosure provides a scheme for generating a video from text in an attempt to solve or at least alleviate at least one of the problems existing above.

[0006] According to one aspect of the present disclosure, a method for generating a video from text is provided, comprising the steps of: selecting any video template from a video template library based on a business scenario, wherein the video template library contains at least one video template under different business scenarios; determining dynamic resource parameters in the video template, the dynamic resource parameters including the display position, display time, and display order of the dynamic resource; obtaining the corresponding dynamic resource from the text based on the dynamic resource parameters; and writing the obtained dynamic resource into the video template to generate a video.

[0007] Optionally, the method according to the present disclosure further comprises the step of constructing the video template library: parsing each article corresponding to each business scenario to obtain corresponding key information, the key information including text and pictures; generating at least one video template for the key information of each article; and constructing the video template library based on the generated video module.

[0008] Optionally, in the method according to the present disclosure, the step of generating at least one video template for the key information of each article comprises: determining static resources and dynamic resources based on the key information of each article, wherein the static resources comprise at least one of pictures, texts, backgrounds, and animation effects associated with the video template, and the dynamic resources comprise at least one of texts, pictures, and audio to be written; defining dynamic resource parameters; and generating one video template corresponding to the article based on the static resources and the dynamic resource parameters.

[0009] Optionally, in the method according to the present disclosure, the step of defining the dynamic resource parameters comprises: generating a first sequence for sequentially indicating positions of each text to be written; generating a second sequence for sequentially indicating positions of each picture to be written; and generating a third sequence comprising at least a plurality of background musics.

[0010] Optionally, in the method according to the present disclosure, the third sequence further comprises whether to turn on speech synthesis, and if the speech synthesis is turned on, text related to speech needs to be obtained.

[0011] Optionally, in the method according to the present disclosure, the step of writing the obtained dynamic resources into the video template to generate a video comprises: writing the obtained dynamic resources except the audio to be written into the video template to generate an initial video template; dividing the initial video template into a plurality of segments and realizing video effects of each segment to obtain corresponding video segments; connecting the video segments in sequence to form a video segment; and fusing the obtained audio to be written with the video segment to obtain a complete video.

[0012] Optionally, the method according to the present disclosure further comprises: when the speech synthesis is turned on, obtaining text related to speech; converting the text related to speech into speech based on speech synthesis technology; and fusing the speech with the obtained background music as the audio to be written.

[0013] According to another aspect of the present disclosure, there is provided an apparatus for generating a video by using a text, comprising: a video template matching unit adapted to select any video template from a video template library based on a business scenario, wherein the video template library comprises at least one video template under different business scenarios; an analysis unit adapted to determine dynamic resource parameters in the video template, wherein the dynamic resource parameters comprise display positions, display times, and display sequences of the dynamic resources; a resource obtaining unit adapted to obtain corresponding dynamic resources from the text based on the dynamic resource parameters; and a video synthesizing unit adapted to write the obtained dynamic resources into the video template to generate a video.

[0014] According to yet another aspect of the present disclosure, a computing device is provided, comprising: one or more processors memory; one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the one or more programs comprising instructions for performing any of the methods above.

[0015] According to yet another aspect of the present disclosure, a computer-readable storage medium storing one or more programs is provided, the one or more programs comprising instructions, which when executed by a computing device, cause the computing device to perform any of the methods described above.

[0016] In summary, according to the scheme of the present disclosure, a video template is selected according to a business scenario, and a dynamic resource required is extracted from the text and written into the video template by analyzing the dynamic resource parameters of the video template, so as to generate a short video. The video is automatically generated from the text, which can greatly save the labor cost and time cost.

[0017] In addition, according to the present scheme, more than one video template is provided under each business scenario, so that multiple videos can be generated for one text, improving the diversity of generated videos. At the same time, the dynamic resource of the generated video is provided from the text to ensure the effect of the generated video.

[0018] When the present scheme is applied to the advertising business, compared with the traditional article advertising method, the short video advertising form using this method can greatly improve the click rate and improve the advertising conversion rate. BRIEF DESCRIPTION OF DRAWINGS

[0019] To the accomplishment of the foregoing and related ends, certain illustrative aspects are described herein in connection with the following description and the annexed drawings. These aspects are indicative of various ways in which the principles disclosed herein can be practiced and all aspects and equivalents thereof are intended to be within the scope of the claimed subject matter. The above- and other advantages of the present disclosure, as defined solely by the claims, will become more fully apparent from the detailed description given herein below and the accompanying drawings, wherein like elements are referred to by like reference numerals. Such description is given for the sake of

[0020] Figure 1 A schematic diagram of a computing device 100 according to some embodiments of the present disclosure is shown;

[0021] Figure 2 A flowchart schematic diagram of a method 200 of generating a short video using text according to some embodiments of the present disclosure is shown;

[0022] Figure 3 A schematic diagram of an apparatus 300 of generating a short video using text according to some embodiments of the present disclosure is shown. DETAILED DESCRIPTION

[0023] Exemplary embodiments of the present disclosure will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the disclosure to those skilled in the art.

[0024] As mentioned above, to address the problems existing in existing solutions, this disclosure proposes a method for generating videos (especially short videos) using text. Short videos generally refer to videos with a duration of no more than 5 minutes. According to this disclosure, the text can be a complete text paragraph, an article, or several sentences, keywords, images, etc., containing short video ideas; this disclosure does not impose many restrictions on this.

[0025] Figure 1 A structural block diagram of a computing device 100 according to an embodiment of the present disclosure is shown.

[0026] like Figure 1 As shown, in the basic configuration 102, the computing device 100 typically includes a system memory 106 and one or more processors 104. A memory bus 108 can be used for communication between the processors 104 and the system memory 106.

[0027] Depending on the desired configuration, processor 104 can be any type of processor, including but not limited to: microprocessors (μP), microcontrollers (μC), digital information processors (DSPs), or any combination thereof. Processor 104 may include one or more levels of cache such as L1 cache 110 and L2 cache 112, processor core 114, and registers 116. Example processor core 114 may include an arithmetic logic unit (ALU), a floating-point unit (FPU), a digital signal processing (DSP) core, or any combination thereof. Example memory controller 118 may be used with processor 104, or in some implementations, memory controller 118 may be an internal part of processor 104.

[0028] Depending on the desired configuration, the system memory 106 can be of any type including but not limited to volatile memory (such as RAM), non-volatile memory (such as ROM, flash memory, etc.) or any combination thereof. Physical memory in a computing device typically refers to volatile memory, such as RAM, and data is typically transferred between data storage and physical memory before the data is processed by the processor 104. The system memory 106 can include an operating system 120, one or more applications 122, and program data 124. In some embodiments, the applications 122 are arranged to execute instructions from the operating system 120 using the program data 124 to perform tasks. The operating system 120, for example, can be a Linux, Windows, etc., that includes program instructions for handling basic system services and for performing hardware dependent tasks. The applications 122 include program instructions for implementing user desired functions, such as a browser, instant messaging software, software development tools (e.g., an integrated development environment (IDE), a compiler, etc.), etc., but are not limited thereto.

[0029] When the computing device 100 is in operation, the processor 104 is configured to read and execute instructions from the memory 106, such as the operating system 120. The applications 122 are executed on the operating system 120 to implement various user desired functions using the interfaces provided by the operating system 120 and underlying hardware. When a user launches an application 122, the application 122 is loaded into the memory 106 and the processor 104 reads and executes the program instructions of the application 122 from the memory 106.

[0030] The computing device 100 also includes storage devices 132, which include a removable storage 136 (e.g., a CD, DVD, USB flash drive, a removable hard disk drive, etc.) and a non-removable storage 138 (e.g., a hard disk drive (HDD), etc.), which are connected to the storage interface bus 134.

[0031] The computing device 100 can also include a storage interface bus 134. The storage interface bus 134 enables communication between the storage devices 132 (e.g., the removable storage 136 and the non-removable storage 138) and the basic configuration 102 via the bus / interface controller 130. The operating system 120, the applications 122, and the program data 124 are at least partially stored either in the removable storage 136 and / or the non-removable storage 138 during the manufacturing of the computing device 100, and are loaded in the system memory 106 and executed by the one or more processors 104 at the time of booting of the computing device 100 or when the applications 122 are required to be executed.

[0032] The computing device 100 may also include an interface bus 140 that facilitates communication from various interface devices (e.g., output devices 142, peripheral interfaces 144, and communication devices 146) to the basic configuration 102 via a bus / interface controller 130. Example output devices 142 include a graphics processing unit 148 and an audio processing unit 150. They may be configured to facilitate communication with various external devices such as displays or speakers via one or more A / V ports 152. Example peripheral interfaces 144 may include a serial interface controller 154 and a parallel interface controller 156, which may be configured to facilitate communication with external devices such as input devices (e.g., keyboards, mice, pens, voice input devices, touch input devices) or other peripherals (e.g., printers, scanners, etc.) via one or more I / O ports 158. Example communication devices 146 may include a network controller 160, which may be arranged to facilitate communication with one or more other computing devices 162 via a network communication link through one or more communication ports 164.

[0033] A network communication link can be an example of a communication medium. A communication medium can typically be embodied in a modulated data signal, such as a carrier wave or other transmission mechanism, and can include any information delivery medium. A “modulated data signal” can be a signal whose data set, or its modifications, can be encoded as information within the signal. As a non-limiting example, a communication medium can include wired media such as wired networks or leased lines, and various wireless media including sound, radio frequency (RF), microwave, infrared (IR), or other wireless media. The term “computer-readable medium” as used herein can include both storage media and communication media.

[0034] The computing device 100 can be implemented as a personal computer, including desktop and laptop computer configurations. Of course, the computing device 100 can also be implemented as part of a small-sized portable (or mobile) electronic device, such as a cellular phone, digital camera, personal digital assistant (PDA), personal media player device, wireless network browsing device, personal head-mounted device, application-specific device, or a hybrid device that may include any of the above functions. It can even be implemented as a server, such as a file server, database server, application server, and web server. The embodiments of the present invention do not limit this.

[0035] In an embodiment according to the present disclosure, computing device 100 is configured to execute method 200 for generating video from text according to the present disclosure. Application 122, arranged on an operating system, includes multiple program instructions for executing method 200 for generating video from text according to the present disclosure. These program instructions can instruct processor 104 to execute method 200 of the present disclosure to generate short videos more quickly.

[0036] Figure 2 A flowchart illustrating a method 200 for generating video from text according to an embodiment of the present disclosure is shown.

[0037] like Figure 2 As shown, method 200 begins with step S210. In step S210, based on the business scenario, any video template is selected from the video template library. The video template library contains at least one video template for different business scenarios.

[0038] Business scenarios can include promotional scenarios, holiday greeting scenarios, new product display scenarios, etc., and can be applied to various business scenarios related to short video production. According to the embodiments of this disclosure, considering that the article structure under the same business scenario will be similar, the video templates in the video template library are classified according to the business scenario. For example, for articles on vehicle promotions, different dealers across the country will release promotional information for different brands and models of vehicles from time to time. The number of promotional articles is large, but many different articles are just different brands or models of vehicles, or prices, so they can be classified into the same category.

[0039] For each business scenario, there can be more than one video template.

[0040] In one embodiment, a template identifier can be assigned to a video template to identify the business scenario to which the video template belongs. This way, when a short video needs to be generated, a suitable video template can be quickly matched from the video template library based on the current business scenario and the template identifier.

[0041] According to this disclosure, before performing step S210, method 200 further includes the step of constructing a video template library.

[0042] In one embodiment, firstly, different articles from various business scenarios are collected, and the articles corresponding to each business scenario are parsed to obtain the corresponding key information. Key information includes text, images, etc. In one embodiment, key information may be, for example, words or phrases with high repetition in the article and located in key positions (such as the article title), which can be obtained through some key information extraction algorithms; this disclosure does not impose many limitations on this. In another embodiment, key information may also be aesthetically pleasing images that are strongly relevant to the business scenario; for example, in the business scenario of holiday greetings, some articles may contain images displaying holiday greetings.

[0043] Next, at least one video template is generated for each article's key information. In one embodiment, the video module can be generated through the following three steps, as detailed below.

[0044] The first step is to determine static and dynamic resources for each article based on its key information. In one embodiment, static resources include at least one of the following associated with the video template: images, text, background (such as background color, background image, etc.), and animation effects. In other words, static resources are text, images, and animation effects of various elements in the short video that are applicable to all short videos generated using the video template. Animation effects include, for example, scaling, rotation, changes in transparency, tilting, and effects that change position over time.

[0045] Dynamic resources include at least one of text, images, and audio to be written. That is, dynamic resources are data that needs to be input from external sources (such as text), and are relevant content that needs to be extracted using the text of the video to be generated.

[0046] The second step is to define dynamic resource parameters.

[0047] Dynamic resource parameters are used to define the type, display position, display time, and display order of the dynamic resources to be passed in. The type of dynamic resource includes at least one of text, image, or audio. The display position refers to the location area within a frame of an image, and the display time refers to the display time within a video segment; these two define the display area of ​​the dynamic resource from the temporal and spatial domains, respectively. When there are multiple dynamic resources, a display order is needed to define the display sequence of each dynamic resource. In one embodiment, the display position, display time, and display order are collectively referred to as display parameters.

[0048] Based on the above description, according to one embodiment, dynamic resource parameters are defined in the following manner.

[0049] 1) Generate a first sequence to sequentially indicate the positions of each text element to be written. The first sequence can be denoted as [T1, T2, ..., T...]. p], where p is the total number of texts that need to be dynamically replaced in the video template, and T i (i∈{1,2,...,p}) represents the string content of the dynamic text at the i-th position in a predefined order. According to this embodiment, each text to be written (T) i It could be a single word, a sentence, or a paragraph; in short, a T i Represents a text region.

[0050] The first sequence is input from the outside during the process of generating short videos using text. It is dynamically obtained according to different texts and written into the corresponding positions in the short video according to the instructions of the first sequence.

[0051] 2) Generate a second sequence to sequentially indicate the positions of each image to be written. The second sequence can be denoted as [I1, I2, ..., I...]. q ], where q is the total number of images that need to be dynamically replaced in the video template, and I i (i∈{1,2,...,q}) represents the storage address of the i-th position of the animated image in a predefined order. According to this disclosure, the second sequence is also input from an external source during the automatic generation of short videos using text. It is dynamically obtained based on different texts and written to the corresponding positions in the short video according to the instructions of the second sequence.

[0052] 3) Generate a third sequence containing at least multiple background music tracks.

[0053] In one embodiment, the third sequence can be denoted as [I1,I2,...,I...]. r ], where r is the total number of background music tracks included in the video template, I i (i∈{1,2,...,r}) represents the i-th background music provided by the video template. When generating a short video, one of these can be selected as the audio content and integrated with the video content.

[0054] In another embodiment, the third sequence further includes whether speech synthesis is enabled, denoted as [I1,I2,...,I...]. r [T], where T represents whether speech synthesis is enabled. If T = 1, speech synthesis is enabled; if T = 0, speech synthesis is disabled. Furthermore, if speech synthesis is enabled, speech-related text needs to be dynamically input from an external source so that it can be converted into audio via TTS (Text To Speech). According to embodiments of this disclosure, from I... i The audio content selected from (i∈{1,2,...,r}) is used as the first audio, and the audio converted by TTS is used as the second audio.

[0055] According to this embodiment, the display order of each dynamic resource has been defined when generating the first sequence and the second sequence. The display position and display time of each dynamic resource can also be defined correspondingly in the first and second sequences. Alternatively, the display position of a dynamic resource can be defined by calling a specific dynamic resource (e.g., I1 in the first sequence) at a corresponding position in the video template. This disclosure does not impose further limitations in this regard.

[0056] The third step is to generate a video template corresponding to the article based on static and dynamic resource parameters.

[0057] According to one embodiment, the predefined static and dynamic resources are coded. As mentioned above, the static resources are pre-arranged in the video template as static parameters. The dynamic resource parameters are arranged in the video template through the first sequence, and / or, the second sequence, and / or the third sequence generated above.

[0058] Finally, based on the generated video modules, a video template library is built.

[0059] For example, a total of M articles from N business scenarios have been collected (assuming M is 100,000 and N is 3). For the i-th business scenario, where i∈{1,2,...,N}, C can be designed and produced. i ∈{C1,C2,...,C N There are} video templates, and the total number of P videos in this business scenario is recorded. i This article satisfies For each article in various business scenarios, C can be automatically generated. i Several short videos. In total, the design requires the production of... One video template, the number of short videos that can be automatically generated is [number]. For example, if C1=3, C2=3, C3=3, P1=30000, P2=40000, P3=30000, then designing a total of 9 video templates can automatically generate 300,000 short videos.

[0060] Generating such a massive number of short videos would require enormous human, material, and time resources if done manually. Therefore, this solution can bring about a significant increase in labor productivity.

[0061] Subsequently, in step S220, the dynamic resource parameters in the video template are determined. As mentioned above, the dynamic resource parameters include at least one of the display position, display time, and display order of the dynamic resource. The dynamic resource parameters to be written can be determined by the first sequence, second sequence, and third sequence defined in the video template.

[0062] Subsequently, in step S230, the corresponding dynamic resources are obtained from the text based on the dynamic resource parameters.

[0063] According to one implementation, text and images can be automatically extracted from the text using a key information extraction algorithm. Of course, as mentioned earlier, when the text used to generate a short video is merely an idea, sentences, keywords, and images representing that idea can be used as the acquired dynamic resources. For example, in a vehicle promotional video sent by a dealer to its customers, the text and images used as dynamic resources could include: vehicle model, color, price, and images of each model, which can be used as the text and images to be written, respectively.

[0064] For a third sequence containing audio content, if speech synthesis is not enabled, the background music selected by the user (i.e., the first audio) can be obtained and passed in as a dynamic resource parameter.

[0065] If speech synthesis is enabled, additional text related to the speech needs to be acquired. Then, based on text-to-speech (TTS) technology, this text is converted into speech (i.e., the second audio). Finally, the converted speech (second audio) is blended with the acquired background music (first audio) to generate a blended audio, which is then used as the audio to be written.

[0066] Subsequently, in step S240, the acquired dynamic resources are written into the video template to generate a video.

[0067] In summary, the acquired dynamic resources may include: text to be written, images to be written, and audio to be written.

[0068] In one embodiment, the acquired dynamic resources, excluding the audio to be written, are first written into the video template to generate an initial video template.

[0069] Next, the initial video template is divided into multiple segments, and the video effects for each segment are implemented to obtain the corresponding video clips. Optionally, the initial video template can be segmented based on animation effects, camera transitions, etc., with the segmentation points defined when setting the video template. Each segment includes the display positions of images and text corresponding to dynamic and static resources, animation effects, and the order and timing of each resource's appearance in the video clip. Optionally, the position, animation effects, and appearance times of each resource in each segment can be implemented using the moviepy library.

[0070] Then, the video clips are connected in sequence to form a video segment.

[0071] Finally, the acquired audio and video segments are merged to obtain a complete video. In one embodiment, the audio and video are merged using the FFmpeg library to form a final, complete short video with sound effects.

[0072] According to the method 200 for generating video from text disclosed herein, a video template is pre-constructed by parsing articles from different business scenarios. After selecting a video template, the required dynamic resources are extracted from the text and written into the video template by parsing the dynamic resource parameters of the video template to generate a short video.

[0073] Applying this solution to advertising can significantly increase click-through rates and improve ad conversion rates compared to traditional article ads.

[0074] It should be noted that this method 200 is not only applicable to advertising, but can also be applied to other scenarios. Any scheme that uses this method 200 to generate video templates by parsing dynamic and static resources to automatically generate short videos falls within the protection scope of this disclosure.

[0075] Accordingly, Figure 3 A schematic diagram of an apparatus 300 for generating video from text according to some embodiments of the present disclosure is shown. The apparatus 300 implements the automatic video generation scheme of the present disclosure by executing method 200. The description of the apparatus 300 complements the description of method 200 above, and repeated details will not be repeated.

[0076] like Figure 3 As shown, the device 300 includes: a video template matching unit 310, a parsing unit 320, a resource acquisition unit 330, and a video synthesis unit 340.

[0077] The video template matching unit 310 selects any video template from the video template library based on the business scenario. The video template library contains at least one video template for different business scenarios. For an explanation of the video template library, please refer to the preceding text.

[0078] Then, the parsing unit 320 determines the dynamic resource parameters in the video template, including the display position, display time, and display order of the dynamic resources.

[0079] Resource acquisition unit 330 acquires corresponding dynamic resources from the text based on dynamic resource parameters.

[0080] The video synthesis unit 340 writes the acquired dynamic resources into a video template to generate a video.

[0081] According to the solution disclosed herein, automatically generating videos from text can significantly save labor and time costs. Furthermore, this solution allows for the generation of multiple videos from a single text, increasing the diversity of generated videos. Simultaneously, the text provides dynamic resources for video generation, ensuring the quality of the generated videos.

[0082] The various techniques described herein can be implemented in combination with hardware or software, or a combination thereof. Thus, the methods and apparatus of this disclosure, or certain aspects or portions thereof, may take the form of program code (i.e., instructions) embedded in a tangible medium, such as a removable hard disk, USB flash drive, floppy disk, CD-ROM, or any other machine-readable storage medium, wherein when the program is loaded into and executed by a machine such as a computer, the machine becomes an apparatus for practicing this disclosure.

[0083] When the program code is executed on a programmable computer, the computing device generally includes a processor, a processor-readable storage medium (including volatile and non-volatile memory and / or storage elements), at least one input device, and at least one output device. The memory is configured to store program code; the processor is configured to execute the methods of this disclosure according to instructions in the program code stored in the memory.

[0084] By way of example, and not limitation, readable media include readable storage media and communication media. Readable storage media stores information such as computer-readable instructions, data structures, program modules, or other data. Communication media generally embodies computer-readable instructions, data structures, program modules, or other data in the form of modulated data signals such as carrier waves or other transmission mechanisms, and includes any information delivery medium. Any combination of the above is also included within the scope of readable media.

[0085] In the specification provided herein, the algorithms and displays are not inherently related to any particular computer, virtual system, or other device. Various general-purpose systems can also be used with the examples of this disclosure. Based on the above description, the required structure for constructing such systems is apparent. Furthermore, this disclosure is not directed to any particular programming language. It should be understood that the contents of this disclosure described herein can be implemented using various programming languages, and the above description of specific languages ​​is for the purpose of disclosing preferred embodiments of this disclosure.

[0086] Numerous specific details are set forth in the specification provided herein. However, it will be understood that embodiments of this disclosure may be practiced without these specific details. In some instances, well-known methods, structures, and techniques have not been shown in detail so as not to obscure the understanding of this specification.

[0087] Similarly, it should be understood that, for the sake of brevity and to aid in understanding one or more of the various aspects of the disclosure, in the foregoing description of exemplary embodiments of the disclosure, various features of the disclosure are sometimes grouped together in a single embodiment, figure, or description thereof. However, this approach to disclosure should not be construed as reflecting an intention that the claimed disclosure requires more features than are expressly recited in each claim. Rather, as reflected in the following claims, the aspects of the disclosure consist of fewer than all features of a single foregoing disclosed embodiment. Therefore, the claims following the detailed description are hereby expressly incorporated into that detailed description, wherein each claim itself is a separate embodiment of the disclosure.

[0088] Those skilled in the art will understand that modules, units, or components of the devices disclosed in the examples herein can be arranged in the devices described in this embodiment, or alternatively, can be located in one or more devices different from the devices in this example. The modules in the foregoing examples can be combined into a single module or, in addition, can be divided into multiple sub-modules.

[0089] Those skilled in the art will understand that modules in the device of the embodiments can be adaptively changed and placed in one or more devices different from that embodiment. Modules, units, or components in the embodiments can be combined into a single module, unit, or component, and further, they can be divided into multiple sub-modules, sub-units, or sub-components. Except where at least some of such features and / or processes or units are mutually exclusive, any combination can be used to combine all features disclosed in this specification (including the accompanying claims, abstract, and drawings) and all processes or units of any method or device so disclosed. Unless expressly stated otherwise, each feature disclosed in this specification (including the accompanying claims, abstract, and drawings) may be replaced by an alternative feature that serves the same, equivalent, or similar purpose.

[0090] Furthermore, those skilled in the art will understand that although some embodiments described herein include certain features included in other embodiments but not others, combinations of features from different embodiments are intended to be within the scope of this disclosure and form different embodiments. For example, in the following claims, any of the claimed embodiments can be used in any combination.

[0091] Furthermore, some of the embodiments described herein are methods or combinations of method elements that can be implemented by a processor of a computer system or by other means of performing the functions. Therefore, a processor having the necessary instructions for implementing the method or method elements forms means for implementing the method or method elements. Furthermore, the elements described herein in the apparatus embodiments are examples of means for implementing the functions performed by the elements for the purposes of this disclosure.

[0092] As used herein, unless otherwise specified, the use of ordinal numbers such as “first,” “second,” “third,” etc., to describe ordinary objects merely indicates different instances of similar objects and is not intended to imply that the objects being described must have a given order in time, space, ordering, or any other manner.

[0093] Although this disclosure has been described with reference to a limited number of embodiments, those skilled in the art will understand from the foregoing description that other embodiments are conceivable within the scope of this disclosure. Furthermore, it should be noted that the language used in this specification has been chosen primarily for readability and edibility purposes, and not for interpreting or limiting the subject matter of this disclosure. Therefore, many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the appended claims. Regarding the scope of this disclosure, the disclosure is illustrative and not restrictive, and the scope of this disclosure is defined by the appended claims.

Claims

1. A method for generating video from text, comprising the following steps: Based on the business scenario, select any video template from the video template library, where... The video template library contains at least one video template for different business scenarios; Determine the dynamic resource parameters in the video template, including the display position, display time, and display order of the dynamic resources; Based on the dynamic resource parameters, the corresponding dynamic resources are obtained from the text; as well as The acquired dynamic resources are written into the video template to generate a video; The steps for constructing the video template library include: The articles corresponding to each business scenario are analyzed to obtain the corresponding key information, including text and images. Generate at least one video template for each article's key information; Based on the generated video modules, a video template library is built; The step of generating at least one video template for each article's key information includes: For each article, static and dynamic resources are determined based on the key information of the article. The static resources include images, text, backgrounds, and animation effects associated with the video template. The dynamic resources are relevant content extracted from the text of the video to be generated, including text, images, and audio to be written. Define dynamic resource parameters; Based on the static resources and the dynamic resource parameters, a video template corresponding to the article is generated; The steps for defining dynamic resource parameters include: Generate a first sequence to sequentially indicate the positions of each text to be written; Generate a second sequence to sequentially indicate the positions of each image to be written; Generate a third sequence containing at least multiple background music tracks.

2. The method as described in claim 1, wherein, The third sequence also includes whether to enable speech synthesis, and if speech synthesis is enabled, then text related to the speech needs to be acquired.

3. The method as described in claim 1 or 2, wherein, The step of writing the acquired dynamic resources into a video template to generate a video includes: Write the acquired dynamic resources, excluding the audio to be written, into the video template to generate the initial video template; The initial video template is divided into multiple segments, and the video effects of each segment are implemented to obtain the corresponding video segments; Connect the video clips in sequence to form a video segment; and The acquired audio to be written is merged with the video segment to obtain a complete video.

4. The method of claim 3, wherein, Before writing the acquired dynamic resources into the video template, the following steps are also included: When speech synthesis is enabled, acquire the text related to the speech. Based on speech synthesis technology, the speech-related text is converted into speech; The spoken audio is then blended with the acquired background music to form the audio to be written.

5. An apparatus for generating video from text, comprising: The video template matching unit is adapted to select any video template from the video template library based on the business scenario, wherein the video template library contains at least one video template under different business scenarios; The parsing unit is adapted to determine the dynamic resource parameters in the video template, wherein the dynamic resource parameters include the display position, display time, and display order of the dynamic resources; The resource acquisition unit is adapted to acquire corresponding dynamic resources from the text based on the dynamic resource parameters; as well as A video synthesis unit is adapted to write the acquired dynamic resources into the video template to generate a video; The device is also adapted to construct the video template library, including: The articles corresponding to each business scenario are analyzed to obtain the corresponding key information, including text and images. Generate at least one video template for each article's key information; Based on the generated video modules, a video template library is built; The step of generating at least one video template for each article's key information includes: For each article, static and dynamic resources are determined based on the key information of the article. The static resources include images, text, backgrounds, and animation effects associated with the video template. The dynamic resources are relevant content extracted from the text of the video to be generated, including text, images, and audio to be written. Define dynamic resource parameters; Based on the static resources and the dynamic resource parameters, a video template corresponding to the article is generated; The steps for defining dynamic resource parameters include: Generate a first sequence to sequentially indicate the positions of each text to be written; Generate a second sequence to sequentially indicate the positions of each image to be written; Generate a third sequence containing at least multiple background music tracks.

6. A computing device, comprising: One or more processors; Memory; One or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions for performing the method as described in any one of claims 1-4.

7. A computer-readable storage medium storing one or more programs, said one or more programs including instructions that, when executed by a computing device, cause the computing device to perform the method as described in any one of claims 1-4.

Citation Information

Patent Citations

  • Method and device for converting article into video, storage medium and equipment

    CN110807126A