A method, device, medium, and equipment for generating a storyboard script

Generating storyboard scripts through multimodal large models solves the problem of ordinary users conceiving outlines and formulating storyboard scripts in video creation, achieving an efficient and automated video creation process, and improving user experience.

CN118764694BActive Publication Date: 2025-07-22GUANGZHOU YIHU NETWORK TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410934151.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-11
Publication Date
2025-07-22
Estimated Expiration
2044-07-11

AI Technical Summary

Technical Problem

When ordinary users create videos, it is difficult to conceive outlines and formulate storyboard scripts, resulting in inefficient and unstable video creation.

Method used

Through user input setting framework, use the trained multimodal mockup model to generate script prompts, combine lens text and images to automatically generate storyboard scripts, and allow users to evaluate and adjust to optimize the generation of results.

Benefits of technology

It greatly improves the user's video creation experience, and automatically generates storyboard scripts, reducing the complexity of the creative process and improving the video quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118764694B_ABST
    Figure CN118764694B_ABST
Patent Text Reader

Abstract

In the method, apparatus, medium, and device for generating a storyboard script provided in this specification, first, the set framework input by the user is input into the trained multi-modal large model to obtain a script prompt. Secondly, the script prompt and the set framework are input again to obtain the shot texts of each shot and multiple shot images of each shot. Thirdly, the multiple shot images corresponding to each shot text are respectively input into the multi-modal large model to obtain the description texts of the multiple shot images, and from the description texts of the multiple shot images, the description text that matches the shot text is determined, and the shot image corresponding to the matching description text is used as the shot image that matches the shot text. Finally, a storyboard script is generated and displayed based on each shot text and the shot images that match each shot text. The storyboard script generated quickly greatly improves the user's video creation experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification relates to the field of computer technology, and in particular, to a method, apparatus, medium, and device for generating a storyboard script. Background Art

[0002] In recent years, with the development of computer technology, more and more users have become video creators. For each user, how to create high-quality videos that stand out from the vast number of videos has become a key research direction.

[0003] In the prior art, users need to first collect inspiration for ideas, then formulate an outline, and then create to obtain a storyboard script, so as to refer to the storyboard script for video creation to obtain high-quality videos.

[0004] It is not difficult to see that the storyboard script determines the quality of the created video. However, for ordinary users who have not studied professional shooting techniques, the process of idea generation and outline formulation is difficult, which leads to low efficiency in obtaining the storyboard script and seriously affects the user's video creation experience.

[0005] Therefore, this specification provides a method for generating a storyboard script. Summary of the Invention

[0006] This application provides a method to partially solve the above problems existing in the prior art.

[0007] This application provides a method for generating a storyboard script, including:

[0008] Responding to a set framework input by a user, and inputting the set framework into a trained multi-modal large model to obtain a script prompt;

[0009] Inputting the script prompt and the set framework into the multi-modal large model to obtain the shot text of each shot and multiple shot images corresponding to each shot text;

[0010] For the shot text of each shot, inputting the multiple shot images corresponding to the shot text into the multi-modal large model respectively to obtain description texts of the multiple shot images;

[0011] Determining, from the description texts of the multiple shot images, a description text that matches the shot text, and using the shot image corresponding to the matching description text as the shot image that matches the shot text;

[0012] Generating a storyboard script according to each shot text and the shot images that match each shot text, and presenting it to the user.

[0013] Optionally, the method further includes:

[0014] In response to the evaluation text input by the user, input the evaluation text into the multi-modal large model to obtain an evaluation prompt;

[0015] Input the evaluation prompt and the set framework into the multi-modal large model to obtain the shot text of each shot and multiple shot images corresponding to each shot text;

[0016] For the shot text of each shot, input the multiple shot images corresponding to the shot text into the multi-modal large model respectively to obtain the description text of the multiple shot images;

[0017] Determine the description text that matches the shot text from the description texts of the multiple shot images, and use the shot image corresponding to the matching description text as the shot image that matches the shot text;

[0018] Generate a shot script according to the shot texts and the shot images matching the shot texts, and display it to the user.

[0019] Optionally, the method further includes:

[0020] In response to the rewrite instruction input by the user according to the shot script, determine the temperature parameter of the multi-modal large model that generates the shot script, and adjust the temperature parameter of the multi-modal large model;

[0021] Input the set framework and the rewrite instruction into the adjusted multi-modal large model to obtain a script prompt different from the script prompt;

[0022] Input the different script prompt and the set framework into the multi-modal large model to re-obtain the shot text of each shot and multiple shot images corresponding to each re-obtained shot text;

[0023] For each re-obtained shot text, determine the shot image that matches the re-obtained shot text from the multiple shot images corresponding to the shot text;

[0024] Generate a new shot script according to the re-obtained shot texts and the shot images matching the re-obtained shot texts, and display the newly generated shot script to the user.

[0025] Optionally, inputting the set framework into the trained multi-modal large model to obtain a script prompt specifically includes:

[0026] Determine each keyword from the set framework, and determine the script element corresponding to each keyword according to a preset script element comparison table;

[0027] For each keyword, determine the position of the keyword in the script outline according to the position of the keyword in the set framework.

[0028] According to the positions of the keywords, the script elements corresponding to the keywords, and the preset number of shots, determine the script outline of the shots with the number of shots, and input the script outline into the trained multi-modal large model to obtain a script prompt.

[0029] Optionally, determine the description text that matches the shot text from the description texts of the multiple shot images, specifically including:

[0030] For each description text, determine the keywords in the description text and the positions of the keywords in the description text.

[0031] Determine the keywords in the shot text and the positions of the keywords in the shot text.

[0032] According to the positions of the keywords in each description text and the positions of the keywords in the shot text, determine the description text that matches the shot text.

[0033] Optionally, the method further includes:

[0034] In response to the shot text selected by the user from the storyboard script, determine the keywords of the selected shot text.

[0035] According to the preset script element comparison table, determine the script elements corresponding to the keywords respectively.

[0036] For each keyword, determine the position of the keyword in the script outline according to the position of the keyword in the shot text.

[0037] According to the positions of the keywords, the script elements corresponding to the keywords, and the preset number of shots, determine the script outline of the shots with the number of shots, and input the script outline into the trained multi-modal large model to obtain a script prompt;

[0038] Input the script prompt and the set framework into the multi-modal large model to obtain the shot text of each storyboard and multiple shot images corresponding to each shot text.

[0039] For the shot text of each storyboard, input the multiple shot images corresponding to the shot text into the multi-modal large model respectively to obtain the description texts of the multiple shot images.

[0040] Determine the descriptive text that matches the shot text from the descriptive texts of the multiple shot images, and use the shot image corresponding to the matching descriptive text as the shot image that matches the shot text;

[0041] Generate a storyboard script based on each shot text and the shot images that match the shot texts, and display it to the user.

[0042] Optionally, the set framework includes at least: background plot and world view.

[0043] This specification provides a computer storage medium storing a computer program, and when the computer program is executed by a processor, it implements a method for generating a storyboard script.

[0044] This application provides a device for generating a storyboard script, including:

[0045] A prompt module that responds to a set framework input by a user, inputs the set framework into a trained multi-modal large model, and obtains a script prompt;

[0046] A determination module that inputs the script prompt and the set framework into the multi-modal large model to obtain the shot texts of each shot and multiple shot images corresponding to each shot text;

[0047] A description module that, for the shot text of each shot, inputs the multiple shot images corresponding to the shot text into the multi-modal large model respectively to obtain the descriptive texts of the multiple shot images;

[0048] A matching module that determines the descriptive text that matches the shot text from the descriptive texts of the multiple shot images, and uses the shot image corresponding to the matching descriptive text as the shot image that matches the shot text;

[0049] A display module that generates a storyboard script based on each shot text and the shot images that match the shot texts, and displays it to the user.

[0050] This specification provides an electronic device including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements a method for generating a storyboard script.

[0051] In a method for generating a storyboard script provided in this specification, first, the set framework input by the user is input into the trained multi-modal large model to obtain a script prompt. Secondly, the script prompt and the set framework are input again to obtain the shot text of each shot and multiple shot images of each shot. Thirdly, the multiple shot images corresponding to each shot text are respectively input into the multi-modal large model to obtain the description text of the multiple shot images, and from the description texts of the multiple shot images, the description text matching the shot text is determined, and the shot image corresponding to the matching description text is used as the shot image matching the shot text. Finally, a storyboard script is generated and displayed based on each shot text and the shot images matching each shot text.

[0052] As can be seen from the above method, by quickly generating a storyboard script according to the set framework, the video creation experience of users is greatly improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] The drawings described herein are used to provide a further understanding of this specification and form a part of this specification. The schematic embodiments of this specification and their descriptions are used to explain this specification and do not constitute an improper limitation to this specification. In the drawings:

[0054] Figure 1 is a schematic flow chart of a method for generating a storyboard script provided in this specification;

[0055] Figure 2 is a schematic diagram of a storyboard script of a method for generating a storyboard script provided in this specification;

[0056] Figure 3 is another schematic diagram of a storyboard script of a method for generating a storyboard script provided in this specification;

[0057] Figure 4 is a schematic diagram of a device of a method for generating a storyboard script provided in this specification;

[0058] Figure 5 corresponding to Figure 1 is a schematic diagram of an electronic device provided in this specification. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0059] To make the objectives, technical solutions, and advantages of this specification clearer, the technical solutions of this specification will be clearly and completely described below in conjunction with the specific embodiments of this specification and the corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this specification, rather than all the embodiments. Based on the embodiments in this specification, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of this application.

[0060] The following will, in conjunction with the accompanying drawings, elaborate on the technical solutions provided in each embodiment of this specification in detail.

[0061] Figure 1 It is a schematic flowchart of a storyboard script generation method provided in this specification, which specifically includes the following steps:

[0062] S101: Respond to the set framework input by the user, and input the set framework into the trained multi-modal large model to obtain a script prompt.

[0063] In the current video creation process, the user first conceives an outline based on inspiration, then creates a storyboard script according to the outline, and finally creates a video with reference to the storyboard script. However, for ordinary users without professional training, the process of conceiving an outline is quite difficult, and even with an outline, creating a storyboard script is not easy. It can be seen that for ordinary users, the efficiency of creating a storyboard script is relatively low, and the quality of the created videos is unstable.

[0064] Therefore, this specification provides a storyboard script generation method to generate a complete storyboard script based on the user's set framework. This storyboard script generation process can be executed by devices with computing capabilities such as personal computers, mobile terminals, servers, etc. Specifically, this specification does not limit which device is used to execute this storyboard script generation process. For unified description, this specification takes the server executing this storyboard script generation method as an example for illustration.

[0065] In one or more embodiments of this specification, the server does not generate a storyboard script out of thin air. The server needs to generate a storyboard script based on at least the basic information given by the user. Specifically, the server can respond to the set framework input by the user, input the set framework into the trained multi-modal large model, and obtain a script prompt.

[0066] Among them, the set framework refers to the text input by the user, which has a worldview and a background plot. Both the worldview and the background plot are used to guide the generation of the to-be-generated storyboard script. And the script prompt is used to prompt the multi-modal large model in the subsequent steps, that is, to prompt the large model what kind of shot text and what kind of shot image to generate. It can be seen that the script prompt can be understood as the story outline of the storyboard script generated based on the worldview specified by the user and the background plot specified by the user. That is, based on the set framework, through the multi-modal large model, a framework of a video plot is given. And the subsequent steps can continue to further refine and determine the storyboard script based on this script prompt.

[0067] In addition, in one or more embodiments of this specification, the content input by the user does not only include the set framework, but can also input some prompt words at the same time, as a reference for the multi-modal large model to generate script prompts. For example, the prompt words are "Please refer to the background plot () under the world view () given by me and randomly arrange a story outline for me, with no restrictions on the theme, plot, or length." The content input in the first "()" is the world view specified by the user, and the content input in the second "()" is the background plot specified by the user. By using the prompt words, the user inputs the corresponding set framework content respectively, and enables the multi-modal large model that receives the set framework and the prompt words to more quickly and accurately identify the world view or the background plot, and determine the script prompt corresponding to the "story outline". In the above embodiment, the types of variables input by the user are the world view and the background plot. However, in addition to the world view and the background plot, other variable types can also be selected. For example, other variable types such as character setting or background music can be additionally selected.

[0068] In the embodiments of this specification, the trained multi-modal large model can determine the world view corresponding to the set framework and the background plot corresponding to the set framework according to the text in the set framework, and then determine the main plot of the storyboard script to be generated within the background plot under the world view according to the world view and the background plot. Then, according to the preset number of storyboard scripts, the determined main plot is divided into the number of storyboard plots of the storyboard script, and each storyboard plot is input into the multi-modal large model to generate a script prompt in the form of a script outline.

[0069] It should be additionally noted here that in the embodiments of this specification, the script prompt obtained by analyzing the set framework through the world view and analyzing the set framework through the background plot in the multi-modal large model means that at least the set framework needs to be analyzed from the perspective of the world view and the background plot to obtain the script prompt. For example, after analyzing the set framework from the perspective of the world view and the background plot, the set framework can also be analyzed from the perspective of the character setting, and then the script prompt is jointly generated according to the obtained world view, background plot, and character setting. Generally speaking, in order to obtain a more accurate script prompt, more angles for analyzing the set framework are often selected.

[0070] S103: Input the script prompt and the set framework into the multi-modal large model to obtain the lens text of each storyboard and multiple lens images corresponding to each lens text.

[0071] After the server determines the script prompt, the multi-modal large model generates the shot text for each shot according to the script prompt and the set framework. For each shot, multiple shot images are also generated. Among them, to help users preview the narrative flow and rhythm, check the smoothness of the transitions between scenes, and obtain the degree of compliance with the background plot, shot images are generated for each shot. Also, because the quality of the generated shot images is unstable, multiple shot images are generated for each shot for subsequent steps to select. Each shot in the shot script includes at least one paragraph of shot text and one generated shot image. Through the shot text and the shot image, users can accurately understand the specific content of the shot in the shot script, so as to display the evaluation of the shot or carry out corresponding shooting activities.

[0072] Specifically, the server inputs the set framework and the script prompt into the multi-modal large model, enabling the multi-modal large model to supplement the script prompt according to the set framework and the script prompt. In one or more embodiments of this specification, it is the multi-modal large model that supplements the script prompt in the form of a script outline. Based on the shot plot of each shot included in the script outline, for each shot plot, according to the set framework, the shot plot is expanded to obtain the shot text of each shot. For example, for the input shot plot of "a battle", the multi-modal large model expands "a battle" according to "the other world" in the set framework to obtain "Under the sky of the other world, on top of the ancient trees, two forces collide violently, with rays of light shining everywhere, sword shadows crisscrossing, and a battle that decides fate is quietly staged, and the air is filled with the atmosphere of magic and courage."

[0073] It should be additionally noted here that to ensure the correct narrative logic, the shot script usually contains multiple shot texts arranged in a fixed order. And each shot text has text that echoes the adjacent shot text, and the echoing text reflects the coherence and overall sense of the content of the shot script. In one or more embodiments of this specification, whether the text has an echo is judged by the multi-modal large model through visual elements, dialogue, actions, or emotional clues. Similarly, each shot text also has text that reflects the narrative rhythm and emotional ups and downs of the content of the shot script. That is to say, for the overall content of the shot script, each shot text has one or more unique functions, such as functions like promoting the plot, showing the character's psychology, creating an atmosphere, or providing more key information, etc. Through these different functions, each shot text helps users accurately understand the shot script generated in the subsequent steps.

[0074] S105: For the shot text of each shot, input the multiple shot images corresponding to the shot text into the multi-modal large model respectively to obtain the description text of the multiple shot images.

[0075] By determining the descriptive texts of each shot image, different descriptive texts represent different images, so that when selecting images in subsequent steps, the comparison between the shot text and the shot image is transformed into the comparison between the shot image and the descriptive text, that is, the comparison between text and image is transformed into the comparison between text and text, greatly reducing the difficulty of selecting shot images in subsequent steps.

[0076] Specifically, the server inputs each shot image into a multi-modal large model to determine the descriptive text of each shot image. In one or more of this specification, the server enables the multi-modal large model to extract multiple features of each shot image, and screens the extracted features through the attention mechanism in the multi-modal large model, and then determines the vocabulary that can represent the screened features, and forms the determined vocabulary into a descriptive text and outputs it.

[0077] S107: Determine the descriptive text that matches the shot text from the descriptive texts of the multiple shot images, and use the shot image corresponding to the matching descriptive text as the shot image that matches the shot text.

[0078] By comparing the shot text of each sub-shot with the descriptive texts of the multiple shot images of this sub-shot, determine the shot image that matches the shot text of this sub-shot.

[0079] Specifically, in one or more embodiments of this specification, the server first determines the descriptive texts of multiple shot images of each sub-shot for each sub-shot. Secondly, for each descriptive text of this sub-shot, compare the content difference between the shot text of this sub-shot and this descriptive text. Finally, according to the differences of the descriptive texts of each sub-shot, select the matching descriptive texts for each sub-shot, and use the shot images represented by the matching descriptive texts of each sub-shot as the shot images that match the shot text of this sub-shot. For example, for each sub-shot, the descriptive texts of this sub-shot can be sorted from the smallest proportion to the largest proportion according to the proportion of differences in the descriptive text, and the first descriptive text in the sorting can be selected as the descriptive text that matches the shot text of this sub-shot.

[0080] S109: Generate a storyboard script based on the shot texts and the shot images that match the shot texts, and display it to the user.

[0081] In one or more embodiments of this specification, after the server determines the shot texts of each sub-shot generated based on the script prompt and the shot images that match the shot texts, it can combine the shot texts and the shot images in the order of the shot texts to generate a storyboard script. And return the storyboard script to the user for display, so that the user can preview the generated storyboard script.

[0082] Specifically, the server first determines the order of the shot texts for each shot according to the shot prompts. By numbering each shot text, the numerical values of the shot numbers from small to large represent the sequence in which the shot texts occur in the video. Secondly, for each shot, the shot image matching the shot text of this shot is placed after the shot number, and the position of the entire shot text is shifted backward, thereby obtaining a shot script and presenting the generated shot script to the user.

[0083] Further, to facilitate understanding of the shot script containing each shot image and each shot text, this specification provides a partial schematic diagram of the shot script, as Figure 2 shown.

[0084] Figure 2 is the schematic diagram of the shot script provided by the embodiment of this specification. Similar to a table in Figure 2 , the shot script also has a header. The header in the figure includes the shot number, the shot image, and the shot text. Each cell corresponding to the header represents the shot number, the shot image, and the shot text of the corresponding shot. The specific content of the shot image is omitted in the figure, and the differences between the shot texts are shown through different thumbnail representations.

[0085] Based on Figure 1 The shot script generation method shown can achieve: First, input the set framework input by the user into the trained multi-modal large model to obtain a script prompt. Secondly, input the script prompt and the set framework again to obtain the shot texts of each shot and multiple shot images of each shot. Thirdly, input the multiple shot images corresponding to each shot text into the multi-modal large model respectively to obtain the description texts of the multiple shot images, and determine the description text matching the shot text from the description texts of the multiple shot images, and use the shot image corresponding to the matching description text as the shot image matching the shot text. Then, generate and present a shot script based on each shot text and the shot image matching the shot text. By quickly generating a shot script according to the set framework, the video creation experience of the user is greatly improved.

[0086] This specification also provides a method for training a multi-modal large model to generate a trained multi-modal large model for Figure 1The storyboard generation method shown in the figure is used, specifically including: for step S101, through the high-score film reviews of classic movies, the setting framework of the movie is used as the input sample, and the background story of the movie and the world view of the movie are used as the corresponding samples, and the movie outline summarized for the key storyboards of the movie is used as a reference. The setting framework in the input sample is input into the multi-modal large model to be trained, and the generated script prompt and the background story and world view according to which are obtained. If there are differences between the script prompt and the movie outline, by comparing the background story in the corresponding sample and the background story according to which the generated script prompt is obtained, and comparing the world view in the corresponding sample and the world view according to which the generated script prompt is obtained, the parameters of the subsystem that generates the script prompt in the multi-modal large model to be trained are adjusted until the generated script outline is the same as the movie outline in the corresponding sample, and the training is stopped.

[0087] Further, in one or more embodiments of this specification, after the server executes the above training method, for step S103, the movie outline referred to in the training of the S101 model is used as the input sample, and the shot text of each shot and the shot image of each shot are obtained from the official website or the storyboard summarized by professional photographers as references. The movie outline in the input sample is input into the multi-modal large model to be trained, and the generated shot text and shot image are obtained. If there are differences between the shot text and the reference shot text, by comparing the shot text with the reference shot text, the parameters of the subsystem that generates the shot text and shot image in the multi-modal large model to be trained are adjusted until the generated shot text is the same as the shot text in the corresponding sample, and the training is stopped.

[0088] Furthermore, in one or more embodiments of this specification, after the server executes the above training method, for step S105, when multiple multi-modal large models can already implement the content of generating descriptive text according to the shot image in S105, the corresponding subsystem of the existing multi-modal large model that can implement this function can be directly selected and incorporated into the multi-modal large model completed in S103 training, or the existing multi-modal large model can be directly bridged outside the multi-modal large model completed in S103 training.

[0089] It should be additionally noted here that the above method for training the multi-modal large model is trained with the movie storyboard as a sample. However, for game videos, since the official game website contains more complete relevant setting frameworks than movies, and the script prompts and storyboards used are easier to obtain, the above training method is more suitable for game video generation. The corresponding storyboard generation method can also be understood as the game promotion video storyboard generation method or the game publicity video storyboard generation method.

[0090] It should be additionally noted here that the multimodal large model trained above is a single multimodal large model jointly used in step S101, step S103, and step S105, and is a multifunctional multimodal large model with the functions required for each step. However, it does not mean that the multimodal large model implementing this method must be a multifunctional multimodal large model with the functions required for each step. Specifically, the functions required for each step can be split into multiple multimodal large models. This specification does not make specific restrictions on how to split and the number of multimodal large models, and can be selected as needed. For example, when the computing power for training the multimodal large model is insufficient, three different multimodal large models can be trained separately for each step using the multimodal large model.

[0091] The method provided in this specification further includes that after step S109, the user is partially dissatisfied with the generated text. Therefore, the user inputs the dissatisfied opinions into the multimodal large model as evaluation text. The server then, in response to the evaluation text input by the user, first inputs the evaluation text into the multimodal large model to obtain an evaluation prompt. Subsequently, according to the evaluation prompt, it then inputs the evaluation prompt and the set framework into the multimodal large model to obtain the new shot text for each storyboard panel and multiple new shot images corresponding to each shot text. Again, for the shot text of each storyboard panel, the multiple shot images corresponding to the shot text are respectively input into the multimodal large model to obtain the descriptive text of the multiple shot images. Again, from the descriptive text of the multiple shot images, the descriptive text that matches the shot text is determined, and the shot image corresponding to the matching descriptive text is used as the shot image that matches the shot text. Finally, the server generates a new storyboard script based on each shot text and the shot images that match each shot text, and displays it to the user.

[0092] For the server's response to the user's evaluation of the storyboard script, the server re-inputs the evaluation and the previously input large-scale set framework by the user into the multimodal large model, enabling the multimodal large model to generate the adjusted shot text for each step, regenerate the shot images corresponding to each shot text, and generate a storyboard script based on the adjusted shot text and each shot image, and the server displays the newly generated storyboard script to the user.

[0093] The method provided in this specification further includes that after step S109, the user is not satisfied with the generated text. Therefore, the user inputs the rewritten opinions into the multimodal large model as a rewriting instruction. In response to the rewriting instruction input by the user according to the storyboard script, the server determines the temperature parameter of the multimodal large model that generates the storyboard script and adjusts the temperature parameter of the multimodal large model so that the multimodal large model can generate text different from the text that the user is not satisfied with, and then gradually regenerates the storyboard script again. That is, the set framework and the rewriting instruction are input into the adjusted multimodal large model to obtain a script prompt different from the script prompt. Again, the different script prompt and the set framework are input into the multimodal large model to re-obtain the shot text of each shot and multiple shot images corresponding to each re-obtained shot text. Again, for each re-obtained shot text, the shot image matching the re-obtained shot text is determined from the multiple shot images corresponding to the shot text. Finally, according to the re-obtained shot texts and the shot images matching the re-obtained shot texts, the storyboard script is regenerated and the regenerated storyboard script is displayed to the user.

[0094] The server adjusts the temperature parameter used to generate the shot text in the multimodal large model according to the rewriting instruction input by the user, and regenerates the shot text and the shot images of the regenerated shot text by inputting the set framework into the adjusted multimodal large model. And as the method shown in steps S103 - S109, a new storyboard script is generated according to the regenerated shot text and shot images.

[0095] It should be additionally noted here that the adjusted temperature parameter of the multimodal large model is a parameter that controls the randomness and diversity of the output in the multimodal large model. A higher temperature parameter will result in more random and diverse generated content, while a lower temperature parameter will make the generated result more certain and conservative. For example, in a model that generates text and images simultaneously, the creativity of the generated text and images can be controlled by adjusting the temperature parameter. Correspondingly, adjusting other functions with the same effect as the temperature parameter can also achieve the same effect. For example, through different loss functions, regularization techniques, or innovations in the model structure.

[0096] The method provided in this specification further includes that, to ensure a more accurate understanding of the user's setting framework and improve the quality of the generated script, in step S101, the server first determines each keyword from the setting framework, and determines the script elements corresponding to each keyword according to a preset script element comparison table. Secondly, for each keyword, according to the position of the keyword in the setting framework, the position of the keyword in the script outline is determined. By adding the script elements and positions, the user's setting framework can be understood more accurately. That is, finally, the server determines the script outline of the shots with the number of shots according to the positions of the keywords, the script elements corresponding to the keywords, and the preset number of shots, and inputs the script outline into the trained multi-modal large model to obtain a script prompt.

[0097] The server analyzes the keywords for script elements through the script element comparison table to determine the script elements to which each keyword belongs. The script elements may specifically include: venue, picture content, narration, props, shot type, subtitles, and remarks, etc., which can all be set in the script outline. For example, Figure 3 as shown in the storyboard schematic diagram provided in the embodiment of this specification, the rest is the same as Figure 2 the structure. Combining the positions of the keywords in the setting framework, the script outline can be quickly generated, enabling the multi-modal large model to generate a more accurate script prompt to improve the quality of the subsequent generated storyboard.

[0098] This specification also provides a method for training a multi-modal large model for use in the above generation method, which specifically includes: generating a script element comparison table by summarizing the storyboard elements of classic movies. Through the high-scoring movie reviews of classic movies, taking the setting framework of the movie as an input sample, determining and taking the keywords of the movie as corresponding samples, and taking the movie outline summarized for the key storyboards of the movie as a reference, inputting the setting framework in the input sample into the multi-modal large model to be trained, so that the multi-modal large model obtains the generated script prompt and the keywords according to which the script prompt is generated according to the script element comparison table. If there are differences between the script prompt and the movie outline, by comparing the keywords in the corresponding sample and the keywords according to which the generated script prompt is generated, the parameters of the subsystem for generating the script prompt in the multi-modal large model to be trained are adjusted until the generated script outline is the same as the movie outline in the corresponding sample, and the training is stopped. As for the methods for other steps, the same training methods as steps S103 and S105 can be adopted.

[0099] Further, to obtain more accurate script prompts, in one or more embodiments of this specification, the server may first determine the keywords in each description text and determine the positions of the keywords in the description text. Secondly, determine the keywords in the shot text and the positions of the keywords in the shot text. Finally, determine the description text that matches the shot text according to the positions of the keywords in each description text and the positions of the keywords in the shot text.

[0100] The method provided in this specification further includes, after determining the keywords in step S101, providing a method for more accurately matching texts, that is, in step S107, the server may, for each description text, first determine the keywords in the description text and determine the positions of the keywords in the description text, secondly determine the keywords in the shot text and the positions of the keywords in the shot text, and finally determine the description text that matches the shot text according to the positions of the keywords in each description text and the positions of the keywords in the shot text.

[0101] The server may use the keywords determined in S101 as the basis for matching the shot text and the description text.

[0102] After the user obtains the storyboard script and finds that the shot text has correct content but a slightly simple plot, it is necessary to supplement the plot of the shot text. Therefore, the user inputs the shot text into the server. Since text generation based on the text is a process of expanding from existing "lines" to "surfaces". Here, the "lines" represent text fragments that contain rich semantic information and context relationships. The multimodal large model generates new content based on understanding these existing "lines" to make the overall text more rich and coherent. In this specification, it means that the multimodal large model relies on learning the existing text structure and style, as well as in-depth understanding of the context, to generate text that can naturally integrate into the original text and maintain consistency and fluency, focusing on generating more new details rather than more new plots. For example, given a description of a cat on a rainy day: "The cat likes to play on the grass after the rain." The text generation algorithm based on this may expand this sentence by adding more details: "The cat likes to play on the grass after the rain. It jumps lightly, chasing the grass blades made wet by the rain, and occasionally stops to sniff the earthy smell in the air." In this example, the original sentence is a "line", and the algorithm expands it into a more rich and concrete "surface" by adding details. In contrast, keywords are like "points", which are the basic elements that make up the entire image, but each point itself does not have complete picture information. Starting from these independent points, a complete and coherent "surface" is constructed. In this specification, it means that the multimodal large model understands the semantics of each keyword and then generates a new text based on the understanding of these keywords and inference of the context. Through semantic understanding and creativity, without the guidance of specific context, more new plots are created just based on the keywords. For example, assume the keywords are "cat", "rainy day", and "window". The algorithm needs to understand that "cat" may refer to a pet cat, "rainy day" is a weather condition, and "window" is part of a building. Based on these understandings, the algorithm may generate such a sentence: "A wet kitten curls up under the window, watching the raindrops gently tapping on the glass." In this example, the keywords are like three isolated points, and the generated text constitutes a new plot. Therefore, the method of determining text based on keywords will be more in line with the user's expectations. Accordingly, the server re-determines the keywords from the determined shot text, generates a script prompt based on the re-determined keywords, and generates an expanded storyboard script according to the script prompt.

[0103] This specification provides an expansion method. After step S109, after the user obtains the storyboard script and determines that the content is correct but the shot text is slightly simple. To obtain more new content related to the shot text, the user inputs the shot text into the server. In response to the shot text selected by the user from the storyboard script. First, determine each keyword of the selected shot text. Secondly, according to the preset storyboard element comparison table, determine the storyboard elements corresponding to each keyword. Thirdly, for each keyword, determine the position of the keyword in the script outline according to the position of the keyword in the shot text. Fourthly, according to the positions of the keywords, the storyboard elements corresponding to the keywords, and the preset number of shots, determine the script outline of the shots with the number of shots. Input the script outline into the trained multi-modal large model to obtain a script prompt. Fifthly, input the script prompt and the set framework into the multi-modal large model to obtain the shot text of each storyboard and multiple shot images corresponding to each shot text. Sixthly, for the shot text of each storyboard, input the multiple shot images corresponding to the shot text into the multi-modal large model respectively to obtain the description text of the multiple shot images. Seventhly, from the description texts of the multiple shot images, determine the description text that matches the shot text, and use the shot image corresponding to the matching description text as the shot image that matches the shot text. Finally, generate a storyboard script according to each shot text and the shot images that match each shot text, and display it to the user. Furthermore, execute as shown in steps S103 - S109 to expand the determined shot text, obtain the expanded storyboard script, and display it to the user.

[0104] This specification also provides Figure 1 a device corresponding to the flowchart of the storyboard script generation method, as Figure 4 shown:

[0105] A prompt module 201, in response to the set framework input by the user, and input the set framework into the trained multi-modal large model to obtain a script prompt;

[0106] A determination module 203, input the script prompt and the set framework into the multi-modal large model to obtain the shot text of each storyboard and multiple shot images corresponding to each shot text;

[0107] A description module 205, for the shot text of each storyboard, input the multiple shot images corresponding to the shot text into the multi-modal large model respectively to obtain the description text of the multiple shot images;

[0108] A matching module 207, from the description texts of the multiple shot images, determine the description text that matches the shot text, and use the shot image corresponding to the matching description text as the shot image that matches the shot text;

[0109] The display module 209 generates a storyboard script based on each shot text and the shot images matched with the shot texts, and displays it to the user.

[0110] Optionally, the device further includes a feedback module 211. The feedback module 211 is used to respond to the evaluation text input by the user. First, the evaluation text is input into the multi-modal large model to obtain an evaluation prompt. Second, the evaluation prompt and the set framework are input into the multi-modal large model to obtain the shot texts of each scene and multiple shot images corresponding to each shot text. Third, for the shot text of each scene, the multiple shot images corresponding to the shot text are respectively input into the multi-modal large model to obtain the description texts of the multiple shot images. Fourth, from the description texts of the multiple shot images, the description text matching the shot text is determined, and the shot image corresponding to the matching description text is used as the shot image matched with the shot text. Finally, the server generates a storyboard script based on each shot text and the shot images matched with the shot texts, and displays it to the user.

[0111] Optionally, the device further includes a feedback module 211. The feedback module 211 is used to respond to the rewrite instruction input by the user according to the storyboard script. First, the temperature parameter of the multi-modal large model for generating the storyboard script is determined and the temperature parameter of the multi-modal large model is adjusted. Second, the set framework and the rewrite instruction are input into the adjusted multi-modal large model to obtain a script prompt different from the original one. Third, the different script prompt and the set framework are input into the multi-modal large model to re-obtain the shot texts of each scene and multiple shot images corresponding to each re-obtained shot text. Fourth, for each re-obtained shot text, the shot image matching the re-obtained shot text is determined from the multiple shot images corresponding to the shot text. Finally, a new storyboard script is generated based on the re-obtained shot texts and the shot images matched with the re-obtained shot texts, and the newly generated storyboard script is displayed to the user.

[0112] Optionally, the prompt module 201 is used to first determine each keyword from the set framework, and determine the script element corresponding to each keyword according to a preset script element comparison table. Second, for each keyword, determine the position of the keyword in the script outline according to the position of the keyword in the set framework. Finally, according to the positions of the keywords, the script elements corresponding to the keywords, and the preset number of shots, determine the script outline of the shots with the number of shots, and input the script outline into the trained multi-modal large model to obtain a script prompt.

[0113] Optionally, the matching module 207 is configured to, for each description text, first determine the keywords in the description text and the positions of the keywords in the description text, second determine the keywords in the shot text and the positions of the keywords in the shot text, and finally determine the description text that matches the shot text according to the positions of the keywords in each description text and the positions of the keywords in the shot text.

[0114] Optionally, the device further includes a feedback module 211. The feedback module 211 is configured to, in response to the shot text selected by the user from the storyboard script, first determine the keywords of the selected shot text, second determine the script elements corresponding to the keywords according to a preset script element comparison table, third, for each keyword, determine the position of the keyword in the script outline according to the position of the keyword in the shot text, third, determine the script outline of the shot with the number of shots according to the positions of the keywords, the script elements corresponding to the keywords, and the preset number of shots, input the script outline into the trained multi-modal large model to obtain a script prompt, then input the script prompt and a set framework into the multi-modal large model to obtain the shot text of each storyboard and multiple shot images corresponding to each shot text, then, for the shot text of each storyboard, input the multiple shot images corresponding to the shot text into the multi-modal large model respectively to obtain description texts of the multiple shot images, then, determine the description text that matches the shot text from the description texts of the multiple shot images, and use the shot image corresponding to the matching description text as the shot image that matches the shot text, and finally, generate a storyboard script according to each shot text and the shot images that match each shot text and display it to the user.

[0115] This specification also provides a computer-readable storage medium storing a computer program that can be used to execute the above method.

[0116] This specification also provides Figure 5 a schematic structural diagram of an electronic device corresponding to Figure 1 the storyboard script generation method. As Figure 5 , at the hardware level, the electronic device includes a processor, an internal bus, a network interface, a memory, and a non-volatile memory. Of course, it may also include other hardware required for other services. The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it to implement the above Figure 1 storyboard script generation method. Of course, in addition to the software implementation, this specification does not exclude other implementation methods, such as logic devices or a combination of software and hardware. That is to say, the execution subject of the following processing flow is not limited to each logic unit, and can also be hardware or logic devices.

[0117] In the 1990s, improvements to a technology could be clearly distinguished as either hardware improvements (e.g., improvements to circuit structures such as diodes, transistors, switches, etc.) or software improvements (improvements to method flows). However, with the development of technology, many method flow improvements today can be regarded as direct improvements to hardware circuit structures. Almost all designers obtain the corresponding hardware circuit structure by programming the improved method flow into the hardware circuit. Therefore, it cannot be said that an improvement to a method flow cannot be implemented using a hardware entity module. For example, a Programmable Logic Device (PLD) (such as a Field Programmable Gate Array (FPGA)) is an integrated circuit whose logical function is determined by the user programming the device. Designers can program themselves to "integrate" a digital system onto a single PLD, without having to ask a chip manufacturer to design and fabricate a dedicated integrated circuit chip. Moreover, today, instead of manually fabricating integrated circuit chips, this programming is mostly implemented using "logic compiler" software, which is similar to the software compilers used in program development and writing. The original code before compilation also has to be written in a specific programming language, which is called a Hardware Description Language (HDL). There is not just one type of HDL, but many, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc. The most commonly used ones currently are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should also be aware that by simply performing a little logical programming on the method flow using the above-mentioned several hardware description languages and programming it into an integrated circuit, it is easy to obtain the hardware circuit that implements the logical method flow.

[0118] The controller can be implemented in any suitable manner. For example, the controller can take the form of, for example, a microprocessor or a processor and a computer-readable medium storing computer-readable program code (such as software or firmware) executable by the (micro)processor, logic gates, switches, an application specific integrated circuit (ASIC), a programmable logic controller, and an embedded microcontroller. Examples of the controller include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320. The memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art also know that in addition to implementing the controller in the form of pure computer-readable program code, it is entirely possible to implement the same function by logically programming method steps so that the controller is in the form of logic gates, switches, application specific integrated circuits, programmable logic controllers, and embedded microcontrollers, etc. Therefore, such a controller can be considered a hardware component, and the devices included therein for implementing various functions can also be regarded as structures within the hardware component. Or even, the devices for implementing various functions can be regarded as either software modules for implementing the method or structures within the hardware component.

[0119] The systems, devices, modules, or units illustrated in the above embodiments can be specifically implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer can be, for example, a personal computer, a laptop computer, a cellular phone, a sensor phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or any combination of these devices.

[0120] For the convenience of description, the above devices are described by dividing them into various units according to functions. Of course, when implementing this specification, the functions of each unit can be implemented in the same or multiple software and / or hardware.

[0121] Those skilled in the art should understand that the embodiments of this specification can be provided as a method, a system, or a computer program product. Therefore, this specification can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, this specification can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) containing computer-usable program code.

[0122] This specification is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the specification. It should be understood that each flow and / or block in the flowchart and / or block diagram, and combinations of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general purpose computer, special purpose computer, embedded processor, or other programmable data processing device to produce a machine, such that the instructions executed by the processor of the computer or other programmable data processing device generate means for implementing the functions specified in one or more of the flows Figure 1 one or more of the flows and / or blocks Figure 1 means for implementing the functions specified in one or more of the blocks.

[0123] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to operate in a particular manner, such that the instructions stored in the computer-readable memory produce a manufacture including instruction means that implement the functions specified in one or more of the flows Figure 1 one or more of the flows and / or blocks Figure 1 means for implementing the functions specified in one or more of the blocks.

[0124] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to produce a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more of the flows Figure 1 one or more of the flows and / or blocks Figure 1 means for implementing the functions specified in one or more of the blocks.

[0125] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.

[0126] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM), and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. The memory is an example of computer-readable media.

[0127] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.

[0128] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.

[0129] Those skilled in the art will appreciate that the embodiments of this specification may be provided as methods, systems or computer program products. Therefore, this specification may take the form of a complete hardware embodiment, a complete software embodiment or an embodiment combining software and hardware. Moreover, this specification may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0130] This specification may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. This specification may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules may be located in local and remote computer storage media, including storage devices.

[0131] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and for the relevant parts, reference can be made to the partial description of the method embodiment.

[0132] The above description is only for the embodiments of this specification and is not intended to limit this specification. For those skilled in the art, various changes and modifications can be made to this specification. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of this specification shall be included within the scope of the claims of this specification.

Claims

1. A storyboard generation method, characterized in that, Including: Responding to a set framework input by a user, inputting the set framework into a trained multi-modal large model to obtain a script prompt, where the script prompt represents the story outline of a storyboard script; Inputting the script prompt and the set framework into the multi-modal large model to obtain the shot text of each shot and multiple shot images corresponding to each shot text; For the shot text of each shot, inputting the multiple shot images corresponding to the shot text into the multi-modal large model respectively to obtain description texts of the multiple shot images; Determining each keyword from the set framework, and for each description text, determining the keyword in the description text and the position of the keyword in the description text; Determining the keyword in the shot text and the position of the keyword in the shot text; Determining a description text that matches the shot text according to the position of the keyword in each description text and the position of the keyword in the shot text, and using the shot image corresponding to the matching description text as the shot image that matches the shot text; Generating a storyboard script according to each shot text and the shot images that match each shot text, and presenting it to the user.

2. The method according to claim 1, characterized in that The method further includes: Responding to an evaluation text input by a user, inputting the evaluation text into the multi-modal large model to obtain an evaluation prompt; Inputting the evaluation prompt and the set framework into the multi-modal large model to obtain the shot text of each shot and multiple shot images corresponding to each shot text; For the shot text of each shot, inputting the multiple shot images corresponding to the shot text into the multi-modal large model respectively to obtain description texts of the multiple shot images; Determining a description text that matches the shot text from the description texts of the multiple shot images, and using the shot image corresponding to the matching description text as the shot image that matches the shot text; Generating a storyboard script according to each shot text and the shot images that match each shot text, and presenting it to the user.

3. The method according to claim 1, characterized in that The method further includes: Responding to a rewrite instruction input by a user according to the storyboard script, determining the temperature parameter of the multi-modal large model that generates the storyboard script, and adjusting the temperature parameter of the multi-modal large model; Inputting the set framework and the rewrite instruction into the adjusted multi-modal large model to obtain a script prompt different from the script prompt; Inputting the different script prompt and the set framework into the multi-modal large model to obtain the shot text of each shot again and multiple shot images corresponding to each shot text obtained again; For each shot text obtained again, determining the shot image that matches the shot text obtained again from the multiple shot images corresponding to the shot text; Regenerating a storyboard script according to the shot texts obtained again and the shot images that match the shot texts obtained again, and presenting the regenerated storyboard script to the user.

4. The method according to claim 1, wherein Inputting the set framework into a trained multi-modal large model to obtain a script prompt, specifically including: Determine the script elements corresponding to each of the keywords determined from the set framework according to a preset script element comparison table; For each keyword, determine the position of the keyword in the script outline according to the position of the keyword in the set framework; According to the positions of the keywords, the script elements corresponding to the keywords, and a preset number of shots, determine the script outline of the shots with the number of shots, and input the script outline into the trained multi-modal large model to obtain a script prompt; 5. The method according to claim 1, characterized in that, The method further includes: In response to the user selecting shot text from the storyboard script, determine the keywords of the selected shot text; Determine the script elements corresponding to the keywords according to a preset script element comparison table; For each keyword, determine the position of the keyword in the script outline according to the position of the keyword in the shot text; According to the positions of the keywords, the script elements corresponding to the keywords, and a preset number of shots, determine the script outline of the shots with the number of shots, and input the script outline into the trained multi-modal large model to obtain a script prompt; Input the script prompt and the set framework into the multi-modal large model to obtain the shot text of each storyboard and multiple shot images corresponding to each shot text; For the shot text of each storyboard, input the multiple shot images corresponding to the shot text into the multi-modal large model respectively to obtain the description text of the multiple shot images; Determine the description text that matches the shot text from the description texts of the multiple shot images, and use the shot image corresponding to the matching description text as the shot image that matches the shot text; Generate a storyboard script based on the shot texts and the shot images that match the shot texts, and display it to the user; 6. The method according to claim 1, characterized in that The set framework at least includes: background plot and world view; 7. A storyboard generation device, characterized in that, including: A prompt module, which responds to the set framework input by the user and inputs the set framework into the trained multi-modal large model to obtain a script prompt, and the script prompt represents the story outline of the storyboard script; A determination module, which inputs the script prompt and the set framework into the multi-modal large model to obtain the shot text of each storyboard and multiple shot images corresponding to each shot text; A description module, which inputs the multiple shot images corresponding to the shot text of each storyboard into the multi-modal large model respectively to obtain the description text of the multiple shot images; A matching module, which determines the keywords from the set framework, and for each description text, determines the keywords in the description text and the positions of the keywords in the description text; determines the keywords in the shot text and the positions of the keywords in the shot text; according to the positions of the keywords in each description text and the positions of the keywords in the shot text, determines the description text that matches the shot text, and uses the shot image corresponding to the matching description text as the shot image that matches the shot text; A display module generates a storyboard script based on each shot text and the shot image matched with each shot text, and displays it to the user.

8. A computer storage medium, characterized in that, The storage medium stores a computer program, and when the computer program is executed by a processor, the method described in any one of claims 1 to 6 above is implemented.

9. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, the method described in any one of claims 1 to 6 above is implemented.

Citation Information

Patent Citations

  • Data processing method, text and graph generation method and related devices

    CN117671055A

  • Video generation method and system based on large language model

    CN117676195A