Video generation method and apparatus, electronic device, and storage medium
By obtaining text information to generate and display target pictures and videos, the problems of single display process and complex operation in the prior art are solved, and richer interaction methods and higher user experience are achieved.
Patent Information
- Application Number
- PCT/CN2025/075774
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-02-05
- Filing Date
- 2025-02-05
- Publication Date
- 2025-08-14
AI Technical Summary
When generating videos, the presentation process and interaction methods are single, and the user operation requirements are high, making it difficult to meet the diverse needs of users.
By obtaining the text information entered by the user, generating and displaying the target image, and then generating and displaying the target video based on the target image, various interactive methods such as regeneration, reference templates and historical video display are provided to lower the user's operation threshold.
It enriches the display process and interaction methods, reduces user operation requirements, improves user experience, and meets users' diverse needs.
Smart Images

Figure CN2025075774_14082025_PF_FP_ABST
Abstract
Description
Video generation method and device, electronic device and storage medium
[0001] This application claims priority to Chinese Patent Application No. 202410166960.9 filed on February 5, 2024, and the contents of the above-mentioned Chinese patent application disclosure are hereby incorporated by reference in their entirety as a part of this application. Technical Field
[0002] Embodiments of the present disclosure relate to a video generation method and apparatus, an electronic device, and a storage medium. Background Art
[0003] Artificial intelligence has developed rapidly in recent years, driven by advances in deep learning algorithms. AI has numerous applications in modern society, including but not limited to intelligent question-answering, natural language processing, autonomous driving, and intelligent transportation. Intelligent question-answering systems are software systems that can engage in conversations with users, based on large amounts of corpus data, using mathematical models and relevant programming languages. Some software systems incorporate image generation capabilities, enabling them to generate and display images based on user needs. Summary of the Invention
[0004] When generating and displaying images based on user needs, the display process and interaction methods are simple and require high user operation. To address the above issues, at least one embodiment of the present disclosure provides a video generation method and device, an electronic device, and a storage medium that can reduce user operation requirements, enrich the display process and interaction methods, and enhance the user experience.
[0005] At least one embodiment of the present disclosure provides a video generation method, including: obtaining first text information, where the first text information is used to describe first video features; generating a first target picture based on the first text information and displaying the first target picture, where the first target picture matches first image features, and the first video features include first image features; generating a first target video based on the first target picture and displaying the first target video, where the first target video matches the first video features.
[0006] At least one embodiment of the present disclosure also provides a video generation device, including an acquisition unit, a first display unit, and a second display unit. The acquisition unit is configured to acquire first text information, where the first text information is used to describe a first video feature; the first display unit is configured to generate a first target picture based on the first text information and display the first target picture, where the first target picture matches a first image feature, and the first video feature includes the first image feature; the second display unit is configured to generate a first target video based on the first target picture and display the first target video, where the first target video matches the first video feature.
[0007] At least one embodiment of the present disclosure also provides an electronic device, comprising: a processor; a memory, comprising one or more computer program modules; wherein the one or more computer program modules are stored in the memory and configured to be executed by the processor, and the one or more computer program modules include instructions for implementing the video generation method described in any embodiment of the present disclosure.
[0008] At least one embodiment of the present disclosure further provides a storage medium for storing non-transitory computer-readable instructions. When the non-transitory computer-readable instructions are executed by a computer, the video generation method described in any embodiment of the present disclosure can be implemented.
[0009] At least one embodiment of the present disclosure further provides a computer program product, including a computer program carried on a non-transitory computer-readable medium, wherein the computer program includes program code for executing the video generation method described in any embodiment of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] The above and other features, advantages, and aspects of the various embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the drawings, like reference numerals represent like elements. It should be understood that the drawings are schematic and that the originals and elements are not necessarily drawn to scale.
[0011] FIG1 is a flow chart of a video generation method provided by some embodiments of the present disclosure;
[0012] FIG2A is a schematic diagram of a first page provided in some embodiments of the present disclosure;
[0013] FIG2B is a schematic diagram of a first target image provided by some embodiments of the present disclosure;
[0014] FIG2C is a schematic diagram of a first target video provided by some embodiments of the present disclosure;
[0015] FIG3 is a schematic diagram of another first page provided by some embodiments of the present disclosure;
[0016] FIG4A is a schematic diagram of a task card provided in some embodiments of the present disclosure;
[0017] FIG4B is a schematic diagram of a first target image provided by some embodiments of the present disclosure;
[0018] FIG4C is a schematic diagram of a first target video provided by some embodiments of the present disclosure;
[0019] FIG5 is a schematic diagram of a gradual animation effect provided by some embodiments of the present disclosure;
[0020] FIG6 is a schematic diagram of an alternative animation effect provided by some embodiments of the present disclosure;
[0021] FIG7 is a schematic diagram of another first page provided by some embodiments of the present disclosure;
[0022] FIG8 is a system that can be used to implement the video generation method provided by an embodiment of the present disclosure;
[0023] FIG9 is a schematic block diagram of a video generating apparatus provided by some embodiments of the present disclosure;
[0024] FIG10 is a schematic block diagram of an electronic device provided by some embodiments of the present disclosure;
[0025] FIG11 is a schematic block diagram of another electronic device provided by some embodiments of the present disclosure; and
[0026] FIG12 is a schematic diagram of a storage medium provided in some embodiments of the present disclosure. DETAILED DESCRIPTION
[0027] The following describes embodiments of the present disclosure in more detail with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments described herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.
[0028] It should be understood that the various steps described in the method embodiments of the present disclosure may be performed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this respect.
[0029] As used herein, the term "including" and its variations are open-ended, i.e., "including but not limited to." The term "based on" means "based, at least in part, on." The term "one embodiment" means "at least one embodiment," the term "another embodiment" means "at least one additional embodiment," and the term "some embodiments" means "at least some embodiments." Other terms are defined in the following description.
[0030] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.
[0031] It should be noted that the modifications of "one" and "plurality" mentioned in this disclosure are illustrative rather than restrictive, and those skilled in the art will understand that unless the context clearly indicates otherwise, they should be understood as "one or more." "Plurality" should be understood as two or more.
[0032] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.
[0033] In some embodiments, videos can be generated based on user input. For example, a video can be generated and displayed based on various control elements, such as a text description, a first frame image, and motion control parameters. However, this approach requires a lot of user input and requires a high level of user experience, creating a certain barrier to entry. Furthermore, after a user submits a request and before a video is generated, the user can only wait for the video to be generated. After the video is generated, the video is displayed directly to the user, resulting in a simple display process and interactive method.
[0034] At least one embodiment of the present disclosure provides a video generation method. The video generation method includes: obtaining first text information, the first text information being used to describe first video features; generating a first target image based on the first text information and displaying the first target image, the first target image matching first image features, the first video features including the first image features; generating a first target video based on the first target image and displaying the first target video, the first target video matching the first video features.
[0035] At least one embodiment of the present disclosure provides a video generation method and device, an electronic device, and a storage medium. After a user inputs text information and submits a generation request, a first target image can be generated based on the text information and displayed. Then, a first target video can be generated based on the first target image and displayed. On the one hand, the corresponding video can be generated and displayed based on the text information input by the user, which requires less user operation and lowers the threshold. On the other hand, after the user submits the request and before the video is generated, interaction with the user can be achieved by first displaying the image, so that the user can have a general understanding of the video image in advance, and feedback from the user can be received during the image display process. Therefore, the display process and interaction method are enriched, and the user experience is improved.
[0036] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the accompanying drawings.
[0037] Figure 1 is a flow chart of a video generation method provided by some embodiments of the present disclosure. As shown in Figure 1, in at least one embodiment, the method includes the following operations.
[0038] Step S110: Acquire first text information, where the first text information is used to describe a first video feature.
[0039] Step S120: generating a first target image based on the first text information and displaying the first target image, wherein the first target image matches the first image feature, and the first video feature includes the first image feature.
[0040] Step S130: Generate a first target video based on the first target image and display the first target video, where the first target video matches the first video feature.
[0041] For example, before step S110 , the method may further include: displaying the first page.
[0042] FIG2A is a schematic diagram of a first page provided in some embodiments of the present disclosure.
[0043] As shown in FIG2A , first page 200 includes an information input area 210 and an information display area 220. Information input area 210 is used to input text information (e.g., a first text message), and information display area 220 is used to display a first target image, a first target video, etc. In response to the operation of submitting a generation request executed on first page 200, a generation request is obtained, and the text in information input area 210 is used as the text information in the generation request.
[0044] For example, the text information may include Chinese, foreign languages, numbers, symbols, etc. The user may input text information using a keyboard, a touch screen, voice, or the like, and the text information input by the user may be displayed in the information input area 210. The first page 200 may further include a generation control 211. When the generation control 211 is triggered (such as being clicked), it is considered that the user has performed an operation of submitting a generation request on the first page 200, and the text information in the information input area 210 may be used as the text information included in the generation request. In addition, in other embodiments, the operation of submitting a generation request may also be performed in other forms. For example, when it is detected that the user has performed an enter operation in the information input area 210 or when it is detected that the user has entered an end symbol such as a period in the information input area 210, it may also be considered that the user has performed an operation of submitting a generation request on the first page 200.
[0045] For example, the first text information is used to describe the first video feature, that is, the first text information is a description of the characteristics of the video to be generated. For example, the first video feature may include information such as the type of object (scenery, person, animal), the state of the object, and the action of the object. In addition, the first video feature may also include a description of information such as color tone and atmosphere.
[0046] For example, in step S120, a first target image is generated based on the first text information, and in response to the generation of the first target image, the first target image is displayed. After the first target image is generated, the first target image can be immediately displayed to the user.
[0047] For example, the first target image matches the first image feature, which can be understood as the first target image conforming to the description of the first image feature. The first image feature can be all or part of the first video feature. For example, the operation of generating the first target image based on text information can be implemented using a pre-trained first deep learning model. The text information is input into the first deep learning model, and the model outputs the first target image. The first deep learning model can be pre-trained using sample text information and sample images.
[0048] FIG2B is a schematic diagram of a first target image provided by some embodiments of the present disclosure.
[0049] As shown in FIG2B , after the first target image 221 is generated, the first target image is displayed in the information display area 220. Before the first target video is generated, the first target image 221 may always be displayed in the information display area 220.
[0050] For example, in step S130, after obtaining the first target image, a first target video may be generated based on the first target image, and in response to generating the first target video, the first target video may be displayed. For example, the first target image may be replaced with the first target video. After generating the first target video, the first target image may be immediately replaced with the first target video.
[0051] For example, if the first target video matches the first video feature, it can be understood that the first target video meets the description of the first video feature. For example, the task of generating the first target video based on the first target image can be implemented using a pre-trained second deep learning model. The first target image is input into the second deep learning model, and the model outputs the first target video. The second deep learning model can be pre-trained using sample images and sample videos.
[0052] For example, in some other embodiments, a first target video can be generated based on the first text information and the first target image. The first text information and the first target image are input into a second deep learning model, which then outputs the first target video. The second deep learning model can be pre-trained using sample text information, sample images, and sample videos.
[0053] FIG2C is a schematic diagram of a first target video provided by some embodiments of the present disclosure.
[0054] As shown in FIG. 2C , after the first target video 222 is generated, the first target video 222 is displayed in the information display area 220 . For example, the first target picture 221 may be replaced by the first target video 222 .
[0055] For example, in some embodiments, the picture of the first target image may be similar to but not identical to the picture of the first target video, and the first target image may not be any frame in the first target video. The first target image can be provided to the user as a reference schematic before the first target video is generated, so that the user can have a general understanding of the composition, structure and outline of the picture of the first target video in advance.
[0056] According to the video generation method of the embodiment of the present disclosure, after the user enters text information and submits a generation request, a first target image can be generated based on the text information and displayed. Then, a first target video can be generated based on the first target image and displayed. On the one hand, the corresponding video can be generated and displayed based on the text information entered by the user, which requires less user operation and lowers the threshold. On the other hand, after the user submits the request and before the video is generated, interaction with the user can be achieved by first displaying the image, so that the user can have a general understanding of the video image in advance. In addition, user feedback can be received during the image display process. Therefore, the display process and interaction method are enriched, and the user experience is improved.
[0057] For example, the video generation method may also include: receiving second text information, the second text information is used to describe the second video feature; generating a second target picture based on the second text information and displaying the second target picture; in response to receiving a cancel operation on the second target picture, generating a third target picture different from the second target picture based on the second text information, and displaying the third target picture; in response to receiving a confirmation operation on the third target picture, generating a second target video based on the third target picture and displaying the second target video, the second target video matching the second video feature.
[0058] For example, after displaying the second target image, if the user is not satisfied with the second target image, a third target image can be regenerated based on the second text information. If the user is satisfied with the third target image, a corresponding video can be generated based on the third target image. If the user is not satisfied with the image and continues to generate a video based on the image, the final generated video will not meet the user's expectations. The embodiment of the present disclosure can avoid the above problem. The embodiment of the present disclosure generates a target video based on a user-approved image, which can meet the user's expectations and further enhance the user experience.
[0059] For example, after displaying the first target video, the video generation method may further include: in response to a regeneration request for the first target video, generating a fourth target picture different from the first target picture based on the first text information, and displaying the fourth target picture, wherein the fourth target picture matches the first image feature; generating a fourth target video different from the first target video based on the fourth target picture, and displaying the fourth target video, wherein the fourth target video matches the first video feature.
[0060] For example, as shown in Figure 2C, after the first target video is shown, a regeneration control 223 can be shown around the first target video. When the regeneration control 223 is triggered, different pictures and videos can be generated based on the text information identical to the first target video. There is no need for the user to re-enter the same text information. The regeneration operation can be achieved by one key through the regeneration control. For example, a random value (seed) can be set in the algorithm. When the regeneration task is triggered, sampling can be performed according to the random value so that the regenerated pictures and videos are different from the pictures and videos generated last time. Based on this approach, a regeneration function can be provided to the user, and different pictures and videos can be generated based on the same text information to meet the needs of the user.
[0061] For example, the video generation method may further include: displaying multiple reference templates, each of which includes a reference video image and reference text information; in response to a first operation being performed on one of the multiple reference templates, obtaining reference text information of the operated target reference template; and obtaining first text information based on the reference text information of the operated target reference template. For example, the reference text information of the operated target reference template may be directly used as the first text information, or edited text obtained by editing the reference text information may be used as the first text information.
[0062] FIG3 is a schematic diagram of another first page provided by some embodiments of the present disclosure.
[0063] As shown in Figure 3, in addition to the information input area 310 and the information display area 320, the first page 300 may also include a reference template area 330. The reference template area 330 can be used to display multiple reference templates, each of which includes a reference video image 331 and reference text information 332. Each reference template may also include, for example, a usage control 333. When a user clicks the usage control 333 of a reference template, it can be considered that a first operation has been performed on the reference template, and the reference text information 332 of the selected reference template has been entered into the information input area 310. In addition, the first operation can also be performed in other ways, for example, the first operation can also be performed by clicking on the reference text information 332 area.
[0064] For example, in some embodiments, after the reference text information 332 of the selected reference template is displayed in the information input area 310, the user can directly submit a generation request based on the reference text information 332 to generate corresponding pictures and videos based on the reference text information 332. The video screen generated based on the reference text information 332 may be different from the reference video screen 331 corresponding to the reference text information 332. The reference video screen 331 corresponding to the reference text information 332 refers to the reference video screen 331 in the reference template to which the reference text information 332 belongs.
[0065] For example, in some other embodiments, after the reference text information 332 of the selected reference template is displayed in the information input area 310 , the user can edit the reference text information 332 and submit a generation request based on the edited text information.
[0066] For example, by providing multiple reference templates, users can use the text information of the reference templates to quickly generate the videos they want, and the reference templates can provide users with creative inspiration, further enhancing the user experience.
[0067] For example, the video generation method may also include: in response to a second operation being performed on one of multiple reference templates, loopingly playing the reference video picture of the operated reference module; in response to a third operation being performed on one of the multiple reference templates, enlarging and displaying the reference video picture of the operated reference module.
[0068] For example, when the cursor is placed on any of the reference video screen 331, reference text information 332, and usage control 333 of a reference template, it can be considered that the second operation has been performed on the reference template, and the video corresponding to the reference video screen 331 can be automatically played in a loop so that the user can watch the entire reference video. In addition, the second operation can also be performed in other ways.
[0069] For example, when the cursor is placed on any of the reference video screen 331, reference text information 332, and usage control 333 of a reference template, a zoom control can be displayed within the reference template area (e.g., a zoom control can be displayed within the area where the reference video screen 331 is located). If the user clicks the zoom control, it can be considered that the third operation has been performed on the reference template. In response to the third operation, a large-screen preview mode can be entered, and the reference video screen 331 can be magnified and played in a loop. In addition, the third operation can also be performed in other ways.
[0070] For example, due to page limitations, the reference text information of some reference templates may not be fully displayed. When the user wants to understand the complete reference text information, the user can perform a fourth operation on the reference template. In response to the fourth operation, the reference text information of the selected reference template is fully displayed. For example, if the cursor stays on the reference text information 332 of a reference template for a period of time greater than or equal to a predetermined time (e.g., 600ms), it can be considered that the fourth operation has been performed on the reference template, and the complete reference text information is magnified and displayed.
[0071] For example, the video generation method may further include: when a first operation is performed on a reference template among the multiple reference templates, if the first operation is the first operation of the user on the multiple reference templates, displaying a tutorial that teaches the user how to use the reference template.
[0072] For example, the information display area is also used to display at least one historical video, and the at least one historical video is generated in response to at least one history generation request.
[0073] For example, as shown in FIG3 , multiple historical videos can be displayed in chronological order in information display area 320. A historical video refers to a target video generated before the current moment based on a user's historical generation request. The area surrounding each historical video can display corresponding historical text information, the time the video was generated (or the time the request was submitted), and other information.
[0074] For example, in some embodiments, the video generation method may further include: displaying multiple functional controls. In response to a first control among the multiple functional controls being triggered, filling in the information input area with the frame setting parameters and text information corresponding to the first target video; in response to a second control among the multiple functional controls being triggered, storing the first target video to a cloud space; in response to a third control among the multiple functional controls being triggered, importing the first target video into a video editor; in response to a fourth control among the multiple functional controls being triggered, saving the first target video locally; and in response to a fifth control among the multiple functional controls being triggered, removing the first target video.
[0075] For example, as shown in FIG3 , in the information display area 320 , multiple functional controls may be displayed in the surrounding area of each video, such as display controls R, C1 - C5 , where control R may represent the aforementioned regeneration control.
[0076] For example, control C1 represents the first control mentioned above. When control C1 of a certain video is triggered, the text information and frame setting parameters corresponding to the video can be filled in the information input area 310, so that the user can view or use the text information and frame setting parameters at that time, wherein the frame setting parameters include parameters such as the frame ratio of the video.
[0077] For example, the control C2 represents the second control mentioned above, which can be used as a cloud storage control. When the control C2 of a certain video is triggered, the operation of saving the video to the cloud space can be executed.
[0078] For example, control C3 represents the third control mentioned above, which can be used as a video editing control. When control C3 of a video is triggered, it can jump to the video editor page and import the video into the video editor. In the video editor, you can add subtitles, add audio, and other operations to the video.
[0079] For example, control C4 represents the fourth control mentioned above, which can be used as a local storage control. When control C4 of a certain video is triggered, the operation of saving the video locally can be executed.
[0080] For example, control C5 represents the fifth control mentioned above, which can be used as a removal control. When control C5 of a certain video is triggered, the video can be removed from the image display area 320.
[0081] For example, when the area of the control display area is not enough to display all the controls in control R, C1-C5, some controls can be hidden. Based on the above multiple controls, various operations on the video can be implemented to meet the user's video operation needs.
[0082] For example, as shown in Figure 3, in some embodiments, when the information display area displays a first historical video and a first target video, the video generation method may further include: displaying two positioning auxiliary controls 324 corresponding to the first historical video and the first target video, respectively, in the information display area 320; in response to one of the two positioning auxiliary controls 324 being triggered, positioning the video corresponding to the triggered positioning auxiliary control 324 to a predetermined area of the information display area 320 (such as the center area of the window), where N is a positive integer.
[0083] For example, when the number of messages (number of videos) in the information display area is greater than two, thumbnails of each video can be displayed on one side of the information display area as positioning auxiliary controls 324. Multiple positioning auxiliary controls 324 can be arranged along the vertical direction of the page, where the vertical direction of the page refers to the direction from the top of the page to the bottom of the page (or from the bottom to the top). The arrangement order of the multiple positioning auxiliary controls 324 is consistent with the arrangement order of the multiple videos. When a positioning auxiliary control 324 is clicked, the corresponding historical video can be positioned in the center area of the window of the information display area 320. When multiple positioning auxiliary controls 324 exceed the window height, they can be scrolled up and down. When multiple videos are scrolled in the historical results area, the multiple positioning auxiliary controls in the positioning auxiliary area can follow the scrolling; when the multiple positioning auxiliary controls in the positioning auxiliary area are scrolled, the historical results area may not follow the scrolling. Based on this method, multiple videos can be quickly viewed in the form of thumbnails, and the video of interest can be quickly positioned in the center area of the window.
[0084] For example, in some embodiments, before generating the first target image, a task card may be displayed, wherein the task card includes text information and / or time information. In response to generating the first target image, the task card is replaced with the first target image. In response to generating the first target video, the first target image is replaced with the first target video. The task card, the first target image, and the first target video all have the same frame setting parameters.
[0085] Figure 4A is a schematic diagram of a task card provided by some embodiments of the present disclosure, Figure 4B is a schematic diagram of a first target picture provided by some embodiments of the present disclosure, and Figure 4C is a schematic diagram of a first target video provided by some embodiments of the present disclosure.
[0086] As shown in Figure 4A, after the user submits the generation request, a task card 421 can be displayed first. The task card can display information such as the text information corresponding to the generation request, the submission time of the generation request, and the frame setting parameters. This information can be displayed in the card area or around the card area. The frame size and frame ratio of the task card, the frame size and frame ratio of the first target image, and the frame size and frame ratio of the first target video can be the same. Based on the task card, the user can be informed of the frame size and frame ratio of the images and videos to be generated before the first target image is generated. After the first target image is generated, the task card can be replaced with the first target image. After the first target video is generated, the first target video can be used to replace the first target image.
[0087] For example, in some embodiments, as shown in Figures 4A to 4C, the video generation method may further include: displaying progress information 426, wherein, while the generation task of the first target image is waiting for processing, the progress information is first progress information, and the first progress information indicates that the task is in queue; in the process of generating the first target image, the progress information is second progress information, and the second progress information indicates that the image is being generated; after displaying the first target image and before generating the first target video, the progress information is third progress information, and the third progress information indicates that the video is being generated; in the process of displaying the first target video, the progress information is fourth progress information, and the fourth progress information indicates that the task has been completed; in the case where the generation of the first target image fails or the generation of the first target video fails, the progress information is fifth progress information, and the fifth progress information indicates that the generation failed; in the case of receiving negative feedback input by the user, the progress information is sixth progress information, and the sixth progress information indicates that the video has received negative feedback.
[0088] For example, if there are currently multiple generation tasks that need to be processed and the current generation task is in the queue, the task card 425 can be displayed first and the progress information "queued" can be displayed at the same time, prompting the user to wait. As shown in Figure 4A, after the current generation task begins to be processed and before the first target image 421 is generated, the progress information "image is being generated" can be displayed. As shown in Figure 4B, after the first target image 421 is displayed and before the first target video 422 is generated, the progress information "video is being generated" can be displayed. As shown in Figure 4C, after the first target video 422 is displayed, the progress information "completed" can be displayed. If the user submits negative feedback to the first target video 422 (such as reporting and other operations), the progress information indicating that the video has been negatively fed back can be displayed. By displaying progress information, users can understand real-time progress and further enhance user experience.
[0089] For example, when negative user feedback on a video is detected, a pop-up window can be displayed, allowing the user to enter a reason, and then enter the review process based on the reason entered by the user.
[0090] For example, in some embodiments, the video generation method may further include: in the process of replacing the task card with the first target picture, superimposing a gradient animation effect on the first target picture; in the process of replacing the first target picture with the first target video, displaying the animation effect of changing from the first target picture to the first target video.
[0091] FIG5 is a schematic diagram of a gradual animation effect provided by some embodiments of the present disclosure.
[0092] As shown in FIG5 , when switching from the task card to the first target image, an animation with a gradient effect (or scanning effect) can be superimposed on the screen to indicate that the task card is currently being switched to the first target image, serving as a prompt.
[0093] FIG6 is a schematic diagram of an alternative animation effect provided by some embodiments of the present disclosure.
[0094] As shown in Figure 6, in the process of replacing the first target image with the first target video, the first target video can be gradually changed starting from one side of the first target image, presenting a gradual change effect, such as the dynamic effect of sweeping light and disappearance of the black mask, which increases the fun of switching from the image to the video, and serves as a prompt, thereby improving the user experience.
[0095] For example, after receiving the generation request, in response to the generation request, the text information may be firstly subjected to a security check, and if the check passes, the first target image may be generated based on the text information. For example, the security check of the text information may include determining whether the text information contains any illegal words.
[0096] For example, after receiving a build request and before executing a build task, the build request text is checked to see if it contains any offending words. If so, the build task is not executed, and the user is prompted that the text contains offending words. If not, the build task is executed.
[0097] For example, in some embodiments, the video generation method may further include: obtaining an aspect ratio setting parameter, wherein the aspect ratio setting parameter is at least used to determine an aspect ratio of the first target image and the first target video.
[0098] For example, the user can set parameters such as the aspect ratio of the image and / or video to be generated, and can pre-set multiple candidate setting parameters for selection, such as pre-setting five ratios of 1:1, 3:4, 16:9, 4:3, and 9:16.
[0099] For example, obtaining the frame setting parameter may include: when the text information includes the first setting parameter, if the first setting parameter belongs to a predetermined plurality of candidate setting parameters, then using the first setting parameter as the frame setting parameter; when the text information does not include the first setting parameter or the first setting parameter does not belong to a plurality of candidate setting parameters, if a second setting parameter selected by the user from a plurality of candidate setting parameters is received, then using the second setting parameter as the frame setting parameter; when the text information does not include the first setting parameter or the first setting parameter does not belong to a plurality of candidate setting parameters, and the second setting parameter selected by the user from a plurality of candidate setting parameters is not received, then obtaining the historical setting parameter from the historical data as the frame setting parameter; when the text information does not include the first setting parameter or the first setting parameter does not belong to a plurality of candidate setting parameters, the second setting parameter selected by the user from a plurality of candidate setting parameters is not received, and there is no historical setting parameter in the historical data, then using the default setting parameter as the frame setting parameter.
[0100] For example, as shown in FIG3 , a user can edit desired setting parameters in information input area 310. For example, if the user wants to generate a video with an aspect ratio of 3:4, they can enter 3:4 in information input area 310. In the process of determining the aspect setting parameters, it can first be determined whether the text information of the generation request includes the setting parameters. If the text information of the generation request includes the setting parameters, the setting parameters in the text information are preferentially used. Before use, it is first determined whether the setting parameters belong to multiple preset candidate setting parameters. If so, they can be used directly; if not, they are discarded.
[0101] For example, as shown in FIG3 , the information input area may display a setting control 312. When the setting control 312 is triggered, multiple parameters may be displayed for selection, such as five preset ratios: 1:1, 3:4, 16:9, 4:3, and 9:16. If the text information does not include a first setting parameter or the first setting parameter does not belong to the multiple candidate setting parameters, if the user selects a second setting parameter using the setting control 312, the second setting parameter will be used as the frame setting parameter.
[0102] For example, the first setting parameter and the second setting parameter are only used to distinguish different parameter sources, and are not used to display specific values of the parameters.
[0103] FIG7 is a schematic diagram of another first page provided in some embodiments of the present disclosure.
[0104] As shown in FIG7 , for example, in some embodiments, the video generation method may further include: if the historical data does not include a historical video and no generation request has been received, displaying a preset video 701 and first preset text information 702 in the information display area; wherein the frame setting parameters of the preset video are different from those of the first target video. If the historical data does not include a historical video and no generation request has been received, displaying the preset video 701 and first preset text information 702 can enhance the visual effect of the image.
[0105] For example, in some embodiments, the video generation method may further include: displaying second preset text information in the information input area. For example, a default text is displayed in the information input area and provided to the user as a reference text.
[0106] It should be noted that, in the embodiments of the present disclosure, the execution order of the various steps of the video generation method is not limited. Although the execution process of the various steps is described above in a specific order, this does not constitute a limitation on the embodiments of the present disclosure. The various steps in the video generation method can be executed serially or in parallel, which can be determined according to actual needs. The video generation method can also include more or fewer steps, for example, to achieve a better preview effect and add some preprocessing steps, or to store some intermediate process data and use it for subsequent processing and calculation, so as to omit some similar steps.
[0107] Figure 8 illustrates a system that can be used to implement the video generation method provided by an embodiment of the present disclosure. As shown in Figure 8 , the system 10 may include a user terminal 11, a network 12, a server 13, and a database 14. For example, the system 10 may be used to implement the video generation method provided by any embodiment of the present disclosure.
[0108] The user terminal 11 is, for example, a computer 11-1. It is understood that the user terminal 11 can be any other type of electronic device capable of performing data processing, including but not limited to a desktop computer, a laptop computer, a tablet computer, a workstation, etc. The user terminal 11 can also be any device equipped with an electronic device. The embodiments of the present disclosure do not limit the hardware configuration or software configuration of the user terminal (for example, the type (e.g., Windows, MacOS, etc.) or version of the operating system).
[0109] The user can operate the application installed on the user terminal 11 or the website logged in on the user terminal 11. The application or website transmits the user behavior data to the server 13 through the network 12. The user terminal 11 can also receive the data transmitted by the server 13 through the network 12.
[0110] For example, the user terminal 11 is installed with video generation software, and the user uses the video generation software to achieve the purpose of video generation on the user terminal 11. The user terminal 11 can execute the video generation method provided by the embodiment of the present disclosure by running code.
[0111] The network 12 may be a single network or a combination of at least two different networks. For example, the network 12 may include, but is not limited to, a local area network, a wide area network, a public network, a private network, or a combination of several thereof.
[0112] The server 13 can be a single server or a server group, wherein each server in the group is connected via a wired or wireless network. A server group can be centralized, such as a data center, or distributed. The server 13 can be local or remote.
[0113] Database 14 generally refers to a device with storage functionality. Database 14 is primarily used to store various data used, generated, and output by user terminal 11 and server 13 during operation. Database 14 can be local or remote. Database 14 can include various types of memory, such as random access memory (RAM) and read-only memory (ROM). The storage devices mentioned above are merely examples, and the storage devices that system 10 can use are not limited to these.
[0114] The database 14 may be connected or communicated with the server 13 or a part thereof via the network 12, or directly connected or communicated with the server 13, or a combination of the two.
[0115] In some examples, database 14 may be a standalone device. In other examples, database 14 may be integrated into at least one of user terminal 11 and server 13. For example, database 14 may be located on user terminal 11 or on server 13. In another example, database 14 may be distributed, with a portion located on user terminal 11 and another portion located on server 13.
[0116] For example, the embodiments of the present disclosure do not limit the type of database, for example, it can be a relational database or a non-relational database.
[0117] Figure 9 is a schematic block diagram of a video generation device provided in some embodiments of the present disclosure. As shown in Figure 9, the video generation device 900 includes an acquisition unit 910, a first display unit 920, and a second display unit 930. For example, the video generation device 900 can be applied to a user terminal or any device or system that needs to generate and display videos, and the embodiments of the present disclosure are not limited thereto.
[0118] The acquisition unit 910 is configured to acquire first text information, where the first text information is used to describe a first video feature. For example, the acquisition unit 910 may execute step S110 of the video generation method shown in FIG1 .
[0119] The first display unit 920 is configured to generate a first target image based on the first text information and display the first target image, wherein the first target image matches the first image feature, and the first video feature includes the first image feature. For example, the first display unit 920 can perform step S120 of the video generation method shown in Figure 1.
[0120] The second display unit 930 is configured to generate a first target video based on the first target image and display the first target video, wherein the first target video matches the first video feature. For example, the second display unit 930 may execute step S130 of the video generation method shown in FIG1 .
[0121] For example, the acquisition unit 910, the first display unit 920, and the second display unit 930 can be hardware, software, firmware, or any feasible combination thereof. For example, the acquisition unit 910, the first display unit 920, and the second display unit 930 can be dedicated or general-purpose circuits, chips, or devices, or can be a combination of a processor and memory. The embodiments of the present disclosure do not limit the specific implementation of the acquisition unit 910, the first display unit 920, and the second display unit 930.
[0122] It should be noted that in the embodiments of the present disclosure, the various units of the video generation device 900 correspond to the various steps of the aforementioned video generation method. For the specific functions of the video generation device 900, reference can be made to the relevant description of the video generation method above and will not be repeated here. The components and structure of the video generation device 900 shown in Figure 9 are merely exemplary and non-restrictive. The video generation device 900 may also include other components and structures as needed.
[0123] For example, in some examples, the first display unit is further configured to: generate the first target image based on the first text information, and in response to generating the first target image, display the first target image. The second display unit is further configured to: generate a first target video based on the first target image, and in response to generating the first target video, replace the first target image with the first target video.
[0124] For example, in some examples, the acquisition unit 910 is further configured to receive second text information, wherein the second text information is used to describe the second video feature. The first display unit is further configured to: generate a second target image based on the second text information and display the second target image; in response to receiving a cancel operation on the second target image, generate a third target image different from the second target image based on the second text information and display the third target image. The second display unit is further configured to: in response to receiving a confirmation operation on the third target image, generate a second target video based on the third target image and display the second target video, wherein the second target video matches the second video feature.
[0125] For example, in some examples, the video generating device 900 may also include a regeneration unit, configured to, after displaying the first target video, generate a fourth target image different from the first target image based on the first text information in response to a regeneration request for the first target video, and display the fourth target image, wherein the fourth target image matches the first image feature; generate a fourth target video different from the first target video based on the fourth target image, and display the fourth target video, wherein the fourth target video matches the first video feature.
[0126] For example, in some examples, the video generating device 900 can also refer to the template unit and be configured to: display multiple reference templates, wherein each of the reference templates includes a reference video picture and reference text information; in response to a first operation being performed on one of the multiple reference templates, obtain the reference text information of the operated target reference template; and obtain the first text information based on the reference text information of the operated target reference template.
[0127] For example, in some examples, the video generating device 900 may also include a page display unit configured to display a first page, wherein the first page includes an information input area, an information display area, and a reference template area; wherein the information input area is used to input the first text information, the information display area is used to display the first target picture and the first target video, and the reference template area is used to display the multiple reference templates.
[0128] For example, in some examples, the reference template unit is further configured to: in response to a second operation being performed on one of the multiple reference templates, loop the reference video frame of the operated reference template; and in response to a third operation being performed on one of the multiple reference templates, zoom in and display the reference video frame of the operated reference template.
[0129] For example, in some examples, the video generating device 900 may also include a page display unit, configured to display an information input area and an information display area, wherein the information input area is used to display the first text information, and the information display area is used to display the first target picture and the first target video; wherein the information display area is also used to display at least one historical video, and the at least one historical video is generated in response to at least one historical generation request.
[0130] For example, in some examples, the at least one historical video includes a first historical video. The video generation device 900 may further include a positioning assistance unit configured to: display two positioning assistance controls corresponding to the first historical video and the first target video, respectively, in the information display area; and in response to one of the two positioning assistance controls being triggered, position the video corresponding to the triggered positioning assistance control to a predetermined area of the information display area.
[0131] For example, in some examples, the video generation device 900 may further include a card display unit configured to display a task card before generating the first target image, wherein the task card includes the first text information and / or time information. The first display unit is further configured to: in response to generating the first target image, replace the task card with the first target image. The second display unit is further configured to: in response to generating the first target video, replace the first target image with the first target video. The frame setting parameters of the task card, the first target image, and the first target video are consistent.
[0132] For example, in some examples, the video generating device 900 may also include an animation unit configured to: superimpose a gradient animation effect on the first target picture during the process of replacing the task card with the first target picture; and display an animation effect of changing from the first target picture to the first target video during the process of replacing the first target picture with the first target video.
[0133] For example, in some examples, the video generating device 900 may further include a progress unit configured to: display progress information, wherein, while the generation task of the first target image is waiting for processing, the progress information is first progress information, and the first progress information indicates that the task is in queue; during the generation of the first target image, the progress information is second progress information, and the second progress information indicates that the image is being generated; after displaying the first target image and before generating the first target video, the progress information is third progress information, and the third progress information indicates that the video is being generated; during the display of the first target video, the progress information is fourth progress information, and the fourth progress information indicates that the task has been completed; in the case that the generation of the first target image fails or the generation of the first target video fails, the progress information is fifth progress information, and the fifth progress information indicates that the generation failed.
[0134] For example, in some examples, the first display unit is further configured to: perform a security check on the text information; if the check passes, generate a first target image based on the first text information and display the first target image.
[0135] For example, in some examples, the video generating device 900 may further include a parameter unit configured to obtain an aspect ratio setting parameter, wherein the aspect ratio setting parameter is at least used to determine an aspect ratio of the first target image and the first target video.
[0136] For example, in some examples, the parameter unit is further configured to: when the first text information includes a first setting parameter, if the first setting parameter belongs to a predetermined plurality of candidate setting parameters, use the first setting parameter as the frame setting parameter; when the first text information does not include the first setting parameter or the first setting parameter does not belong to the plurality of candidate setting parameters, if a second setting parameter selected by the user from the plurality of candidate setting parameters is received, use the second setting parameter as the frame setting parameter; when the first text information does not include the first setting parameter or the first setting parameter does not belong to the plurality of candidate setting parameters, and the second setting parameter selected by the user from the plurality of candidate setting parameters is not received, obtain a historical setting parameter from historical data as the frame setting parameter; when the first text information does not include the first setting parameter or the first setting parameter does not belong to the plurality of candidate setting parameters, the second setting parameter selected by the user from the plurality of candidate setting parameters is not received, and there is no historical setting parameter in the historical data, use a default setting parameter as the frame setting parameter.
[0137] For example, in some examples, the video generating device 900 may also include a functional unit configured to: display multiple functional controls; in response to the first control among the multiple functional controls being triggered, fill in the frame setting parameters and the first text information corresponding to the first target video into the information input area; in response to the second control among the multiple functional controls being triggered, store the first target video to the cloud space; in response to the third control among the multiple functional controls being triggered, import the first target video into the video editor; in response to the fourth control among the multiple functional controls being triggered, save the first target video locally; in response to the fifth control among the multiple functional controls being triggered, remove the first target video.
[0138] For example, in some examples, the video generating device 900 may further include a third display unit, configured to: display a preset video and a first preset text message in the information display area when the historical data does not include a historical video and the generation request is not received; wherein the frame setting parameters of the preset video are different from those of the first target video.
[0139] Figure 10 is a schematic block diagram of an electronic device provided in some embodiments of the present disclosure. As shown in Figure 10, the electronic device 1000 includes a processor 1010 and a memory 1020. The memory 1020 is used to store non-transitory computer-readable instructions (e.g., one or more computer program modules). The processor 1010 is used to execute non-transitory computer-readable instructions. When the non-transitory computer-readable instructions are executed by the processor 1010, one or more steps in the video generation method described above can be executed. The memory 1020 and the processor 1010 can be interconnected via a bus system and / or other forms of connection mechanisms (not shown).
[0140] For example, the processor 1010 may be a central processing unit (CPU), a digital signal processor (DSP), or other processing units with data processing capabilities and / or program execution capabilities, such as a field programmable gate array (FPGA). For example, the central processing unit (CPU) may be an X86 or ARM architecture. The processor 1010 may be a general-purpose processor or a dedicated processor, and may control other components in the electronic device 1000 to perform desired functions.
[0141] For example, the memory 1020 may include any combination of one or more computer program products, which may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) and / or cache memory. Non-volatile memory may include, for example, read-only memory (ROM), a hard disk, an erasable programmable read-only memory (EPROM), a portable compact disk read-only memory (CD-ROM), a USB memory, a flash memory, etc. One or more computer program modules may be stored on the computer-readable storage medium, and the processor 1010 may execute one or more computer program modules to implement various functions of the electronic device 1000. Various applications and various data, as well as various data used and / or generated by the applications, may also be stored in the computer-readable storage medium.
[0142] It should be noted that, in the embodiment of the present disclosure, the specific functions and technical effects of the electronic device 1000 can be referred to the above description of the video generation method, which will not be repeated here.
[0143] Figure 11 is a schematic block diagram of another electronic device provided in some embodiments of the present disclosure. This electronic device 1100 is, for example, suitable for implementing the video generation method provided in embodiments of the present disclosure. The electronic device 1100 may be a user terminal, etc. It should be noted that the electronic device 1100 shown in Figure 11 is merely an example and does not impose any limitations on the functionality and scope of use of the embodiments of the present disclosure.
[0144] As shown in FIG11 , the electronic device 1100 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 1110, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1120 or a program loaded from a storage device 1180 into a random access memory (RAM) 1130. Various programs and data required for the operation of the electronic device 1100 are also stored in the RAM 1130. The processing device 1110, the ROM 1120, and the RAM 1130 are connected to each other via a bus 1140. An input / output (I / O) interface 1150 is also connected to the bus 1140.
[0145] Typically, the following devices may be connected to the I / O interface 1150: an input device 1160 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 1170 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 1180 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1190. The communication device 1190 may allow the electronic device 1100 to communicate with other electronic devices wirelessly or by wire to exchange data. Although FIG11 shows the electronic device 1100 with various devices, it should be understood that it is not required to implement or have all of the devices shown, and the electronic device 1100 may alternatively implement or have more or fewer devices.
[0146] For example, according to an embodiment of the present disclosure, the video generation method can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes program code for executing the above-mentioned video generation method. In such an embodiment, the computer program can be downloaded and installed from the network through the communication device 1190, or installed from the storage device 1180, or installed from the ROM 1120. When the computer program is executed by the processing device 1110, the functions defined in the video generation method provided by the embodiment of the present disclosure can be executed.
[0147] At least one embodiment of the present disclosure further provides a storage medium for storing non-transitory computer-readable instructions. When the non-transitory computer-readable instructions are executed by a computer, the video generation method described in any embodiment of the present disclosure can be implemented.
[0148] Figure 12 is a schematic diagram of a storage medium provided by some embodiments of the present disclosure. As shown in Figure 12, storage medium 1200 is used to store non-transitory computer-readable instructions 1210. For example, when non-transitory computer-readable instructions 1210 are executed by a computer, one or more steps of the video generation method described above can be performed.
[0149] For example, the storage medium 1200 can be applied to the electronic device 1000. For example, the storage medium 1200 can be the memory 1020 in the electronic device 1000 shown in FIG12. For example, the relevant description of the storage medium 1200 can refer to the corresponding description of the memory 1020 in the electronic device 1000 shown in FIG12, and will not be repeated here.
[0150] It should be noted that the storage medium (computer-readable medium) mentioned above in the present disclosure may be a computer-readable signal medium or a non-transitory computer-readable storage medium or any combination of the two. Non-transitory computer-readable storage media may be, for example, but not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices or devices, or any combination of the above. More specific examples of non-transitory computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a non-transitory computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, device or device. In the present disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries a computer-readable program code. Such propagated data signals may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium may also be any computer-readable medium other than a non-transitory computer-readable storage medium that can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to wires, optical cables, RF (radio frequency), etc., or any suitable combination thereof.
[0151] In some embodiments, the client and server can communicate using any currently known or future developed network protocol, such as the Hyper Text Transfer Protocol (HTTP), and can be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), internetworks (e.g., the Internet), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future developed networks.
[0152] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.
[0153] The computer-readable medium carries one or more programs. When the one or more programs are executed by the electronic device, the electronic device: obtains at least two Internet Protocol addresses; sends a node evaluation request including the at least two Internet Protocol addresses to a node evaluation device, wherein the node evaluation device selects an Internet Protocol address from the at least two Internet Protocol addresses and returns it; receives the Internet Protocol address returned by the node evaluation device; wherein the obtained Internet Protocol address indicates an edge node in a content distribution network.
[0154] Alternatively, the computer-readable medium carries one or more programs, which, when executed by the electronic device, causes the electronic device to: receive a node evaluation request including at least two Internet Protocol addresses; select an Internet Protocol address from the at least two Internet Protocol addresses; and return the selected Internet Protocol address; wherein the received Internet Protocol address indicates an edge node in a content distribution network.
[0155] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages, or a combination thereof, including, but not limited to, object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, such as a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0156] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the module, program segment, or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or can be implemented using a combination of dedicated hardware and computer instructions.
[0157] The units involved in the embodiments described in this disclosure may be implemented in software or hardware, wherein the name of a unit does not necessarily limit the unit itself.
[0158] The functions described above herein may be performed, at least in part, by one or more hardware logic components. For example, and without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chip (SOCs), complex programmable logic devices (CPLDs), and the like.
[0159] In the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or apparatus. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or apparatus, or any suitable combination of the foregoing. More specific examples of machine-readable storage media may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0160] The above description is only a partial embodiment of the present disclosure and an illustration of the technical principles used. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above-mentioned technical features, but should also cover other technical solutions formed by any combination of the above-mentioned technical features or their equivalents without departing from the above-mentioned disclosed concepts. For example, the above-mentioned features can be replaced with (but not limited to) technical features with similar functions disclosed in this disclosure.
[0161] In addition, although each operation is described in a specific order, this should not be understood as requiring these operations to be performed in the specific order shown or in a sequential order. Under certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although some specific implementation details have been included in the above discussion, these should not be interpreted as limiting the scope of the present disclosure. Some features described in the context of a separate embodiment can also be implemented in a single embodiment in combination. On the contrary, the various features described in the context of a single embodiment can also be implemented in multiple embodiments individually or in any suitable sub-combination mode.
[0162] Although the subject matter has been described in language specific to structural features and / or methodological logical acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms of implementing the claims.
Claims
1. A video generation method, comprising: Acquire first text information, where the first text information is used to describe a first video feature; generating a first target picture based on the first text information and displaying the first target picture, wherein the first target picture matches a first image feature, and the first video feature includes the first image feature; A first target video is generated based on the first target image and displayed, where the first target video matches the first video feature.
2. The video generation method according to claim 1, wherein: Generating a first target picture based on the first text information and displaying the first target picture includes: generating the first target picture based on the first text information, and displaying the first target picture in response to generating the first target picture; Generating a first target video based on the first target picture and displaying the first target video includes: generating the first target video based on the first target picture, and replacing the first target picture with the first target video in response to generating the first target video.
3. The video generation method according to claim 1 or 2, further comprising: receiving second text information, wherein the second text information is used to describe a second video feature; generating a second target image based on the second text information and displaying the second target image; In response to receiving a cancel operation on the second target image, generating a third target image different from the second target image based on the second text information, and displaying the third target image; In response to receiving a confirmation operation on the third target image, a second target video is generated based on the third target image and displayed, wherein the second target video matches the second video feature.
4. The video generation method according to any one of claims 1 to 3, further comprising: after displaying the first target video; In response to a regeneration request for the first target video, generating a fourth target picture different from the first target picture based on the first text information, and displaying the fourth target picture, wherein the fourth target picture matches the first image feature; A fourth target video different from the first target video is generated based on the fourth target image, and the fourth target video is displayed, wherein the fourth target video matches the first video feature.
5. The video generation method according to any one of claims 1 to 4, further comprising: Displaying a plurality of reference templates, wherein each of the reference templates includes a reference video frame and reference text information; In response to a first operation being performed on a reference template among the plurality of reference templates, obtaining reference text information of the operated target reference template; The first text information is obtained based on the reference text information of the operated target reference template. The video generation method according to claim 5 , wherein: Before receiving the generation request, the video generation method further includes: Displaying a first page, wherein the first page includes an information input area, an information display area, and a reference template area; Among them, the information input area is used to input the first text information, the information display area is used to display the first target picture and the first target video, and the reference template area is used to display the multiple reference templates.
7. The video generation method according to claim 5 or 6, further comprising: In response to a second operation being performed on one of the plurality of reference templates, loopingly playing a reference video frame of the operated reference template; In response to a third operation being performed on one of the plurality of reference templates, a reference video picture of the operated reference template is magnified and displayed.
8. The video generation method according to any one of claims 1 to 7, further comprising: Displaying an information input area and an information display area, wherein the information input area is used to display the first text information, and the information display area is used to display the first target picture and the first target video; The information display area is further used to display at least one historical video, and the at least one historical video is generated in response to at least one historical generation request.
9. The video generation method according to claim 8, wherein: The at least one historical video includes a first historical video; The video generation method further includes: displaying two positioning auxiliary controls corresponding to the first historical video and the first target video respectively in the information display area; In response to one of the two positioning auxiliary controls being triggered, the video corresponding to the triggered positioning auxiliary control is positioned to a predetermined area of the information display area.
10. The video generation method according to any one of claims 1 to 9, further comprising: Before generating the first target image, displaying a task card, wherein the task card includes the first text information and / or time information; The step of generating a first target image based on the first text information and displaying the first target image includes: replacing the task card with the first target image in response to generating the first target image; Generating a first target video based on the first target image and displaying the first target video includes: in response to generating the first target video, replacing the first target image with the first target video; Among them, the frame setting parameters of the task card, the first target picture and the first target video are consistent.
11. The video generation method according to claim 10, further comprising: In the process of replacing the task card with the first target image, a gradient animation effect is superimposed on the first target image; During the process of replacing the first target image with the first target video, an animation effect of changing from the first target image to the first target video is displayed.
12. The video generation method according to any one of claims 1 to 11, further comprising: Display progress information, Wherein, during the process of the generation task of the first target image waiting to be processed, the progress information is first progress information, and the first progress information indicates that the task is in queue; During the process of generating the first target image, the progress information is second progress information, and the second progress information indicates that the image is being generated; After displaying the first target image and before generating the first target video, the progress information is third progress information, and the third progress information indicates that the video is being generated; During the display of the first target video, the progress information is fourth progress information, and the fourth progress information indicates that the task has been completed; In a case where the first target image fails to be generated or the first target video fails to be generated, the progress information is fifth progress information, and the fifth progress information indicates generation failure.
13. The video generation method according to any one of claims 1 to 12, wherein: Generating a first target image based on the first text information and displaying the first target image includes: Performing a security check on the first text information; If the verification is passed, a first target image is generated based on the first text information and the first target image is displayed.
14. The video generation method according to any one of claims 1 to 13, further comprising: An aspect ratio setting parameter is obtained, wherein the aspect ratio setting parameter is at least used to determine an aspect ratio of the first target image and the first target video.
15. The video generation method according to claim 14, wherein: The obtaining of the frame setting parameters includes: In a case where the first text information includes a first setting parameter, if the first setting parameter belongs to a plurality of predetermined candidate setting parameters, using the first setting parameter as the frame setting parameter; In a case where the first text information does not include the first setting parameter or the first setting parameter does not belong to the multiple candidate setting parameters, if a second setting parameter selected by the user from the multiple candidate setting parameters is received, the second setting parameter is used as the frame setting parameter; If the first setting parameter is not included in the first text information or the first setting parameter does not belong to the multiple candidate setting parameters, and a second setting parameter selected by the user from the multiple candidate setting parameters is not received, obtaining a historical setting parameter from historical data as the frame setting parameter; If the first setting parameter is not included in the first text information or the first setting parameter does not belong to the multiple candidate setting parameters, the second setting parameter selected by the user from the multiple candidate setting parameters is not received, and there is no historical setting parameter in the historical data, a default setting parameter is used as the frame setting parameter.
16. The video generation method according to claim 6, further comprising: Display multiple functional controls; In response to a first control among the plurality of function controls being triggered, filling the frame setting parameter corresponding to the first target video and the first text information into the information input area; In response to a second control among the plurality of function controls being triggered, storing the first target video in a cloud space; In response to a third control among the plurality of function controls being triggered, importing the first target video into a video editor; In response to a fourth control among the plurality of function controls being triggered, saving the first target video locally; In response to a fifth control among the plurality of function controls being triggered, the first target video is removed.
17. The video generation method according to claim 8 or 9, further comprising: In the case where the historical data does not include the historical video and the generation request is not received, displaying the preset video and the first preset text information in the information display area; The frame setting parameters of the preset video are different from those of the first target video.
18. A video generating device, comprising: an acquiring unit, configured to acquire first text information, where the first text information is used to describe a first video feature; a first display unit configured to generate a first target image based on the first text information and display the first target image, wherein the first target image matches a first image feature, and the first video feature includes the first image feature; The second display unit is configured to generate a first target video based on the first target image and display the first target video, where the first target video matches the first video feature.
19. An electronic device comprising: processor; a memory comprising one or more computer program modules; The one or more computer program modules are stored in the memory and configured to be executed by the processor, and the one or more computer program modules include instructions for implementing the video generation method according to any one of claims 1 to 17.
20. A storage medium configured to store non-transitory computer-readable instructions, which, when executed by a computer, can implement the video generation method according to any one of claims 1 to 17.
21. A computer program product comprising a computer program carried on a non-transitory computer readable medium, wherein: The computer program includes program code for executing the video generation method according to any one of claims 1 to 17.
Citation Information
Patent Citations
Video generation method and device, electronic equipment and storage medium
CN120434454A
Video generation method and device, medium and equipment
CN109889849A
Video synthesis method and device, computer equipment and storage medium
CN114390217A
Video generation method and server
CN116233491A
Video generation method and device, medium and computer equipment
CN117095085A
Cited By
Video generation method and device, electronic equipment and medium
CN121711542A