Method and system for generating intelligent picture or video of large model interaction

By building a systematic storyboard template and real-time response to user interaction methods, the existing multimodal large model generation interaction problems are solved, and more efficient and richer user interaction experience and generation quality are achieved.

CN120226048APending Publication Date: 2025-06-27余音
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480004085.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-03-01
Filing Date
2024-07-19
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

When generating text to pictures or videos, existing multimodal large models have problems such as single interaction mode, lack of systematic management, low generation efficiency and limited playback interaction mode.

Method used

It provides an intelligent image or video generation method and system for large-model interaction. By building a systematic storyboard template, it realizes systematic, structured editing and management of propt requests, supports batch generation, and responds to user interaction in real time during playback.

Benefits of technology

It improves user interaction experience, improves the efficiency of image and video generation, ensures generation quality, and expands the richness of playback interaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120226048A_ABST
    Figure CN120226048A_ABST
Patent Text Reader

Abstract

The invention provides a method and a system for generating an intelligent picture or video of large model interaction, and the method comprises a picture or video generation process: S11, receiving an original prompt request of a user, extracting description information in the original prompt request, and filling the description information into a sub-mirror template in the system; or directly importing the original prompt request and the sub-mirror template outside the system, and mapping the sub-mirror template outside the system into the sub-mirror template inside the system; s12, serializing and textualizing the content in the sub-mirror template to obtain at least one prompt request segment, each prompt request segment corresponding to a picture or a sub-mirror video; and S13, submitting the prompt request to the large model one by one or in batches in a segmented manner, and returning the picture or video content generated by the large model. According to the method and the system, a systematized sub-mirror template is provided for user construction, so that systematized and structured editing and management are realized, meanwhile, batch generation can be supported, interaction modes are rich and diversified, and the user interaction experience is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of large model interactions, and particularly to a method and system for generating intelligent pictures or videos for large model interactions. Background Art

[0002] Currently, the development of multi-modal large models in artificial intelligence is rapid. For example, models such as stableVideo and openAI sora have realized the function of generating pictures or videos from text.

[0003] The current process of generating pictures or videos from text for existing multi-modal large models in artificial intelligence is as follows: The user makes a prompt request, provides a language description of the image or scene to the large model, and the large model can then return the generated image or video.

[0004] Since the function of the prompt request is mainly based on text / string templates, which include the theme of the picture or video scene, the description of the main body, the description of the picture background and environment, the definition of the painting style, and model parameters, etc.

[0005] As Figure 1 shown, when a large model realizes the function of generating pictures from text according to a prompt request, the user provides a prompt request as shown on the left, and the large model generates a picture as shown on the right. Figure 1 shown on the left, and the large model generates a picture as shown on the right. Figure 1 on the right.

[0006] As Figure 2 shown, for a large model to realize the function of generating videos from text according to a prompt request, the prompt request provided by the user, and the large model generates a video as shown in Figure 3 shown.

[0007] However, when the existing multi-modal large models in artificial intelligence realize the function of generating pictures or videos from text, there are the following several disadvantages:

[0008] 1. The interaction method on the generated page is single. The page interaction is only suitable for single-item and item-by-item text-to-picture and text-to-video interactions; all dimensions of information of the picture are included in the prompt; the newly released stableVideo places the camera state configuration on the secondary generation draft page for selection, resulting in a lack of systematic management of the prompt requests for each picture.

[0009] 2. When the prompt request involves a large amount of text description, the user may lack description dimensions, affecting the generation quality of pictures or videos by the large model.

[0010] 3. When the large model returns the generated content, the playback interaction method is single. Usually, only browsers for pictures and videos are provided, which only support playback interaction methods such as playing, zooming, speed adjustment, and skipping; 4. It does not support batch generation and has low efficiency. Summary of the Invention

[0011] The technical problem to be solved by the present invention is to provide a method and system for generating intelligent pictures or videos for large model interaction, which constructs a systematic storyboard template for users, thereby realizing systematic and structured editing and management. At the same time, it supports batch generation, has a rich variety of interaction methods, and improves the user interaction experience.

[0012] In the first aspect, the present invention provides a method for generating intelligent pictures or videos for large model interaction, including the generation process of pictures or videos. The generation process of pictures or videos includes the following steps:

[0013] S11. Receive the user's original prompt request, extract the description information in the original prompt request and fill it into the storyboard template inside the system; or

[0014] Directly import the original prompt request and the storyboard template outside the system, and map the storyboard template outside the system to the storyboard template inside the system;

[0015] S12. Serialize and literalize the content in the storyboard template to obtain at least one prompt request segment, and each prompt request segment corresponds to a picture or a storyboard video;

[0016] S13. Submit the prompt request segments one by one or in batches to the large model, and the large model returns the generated picture or video content.

[0017] In the second aspect, the present invention provides a system for generating intelligent pictures or videos for large model interaction, including a template management module. The template management module is used to execute the following steps:

[0018] S11. Receive the user's original prompt request, extract the description information in the original prompt request and fill it into the storyboard template inside the system; or

[0019] Directly import the original prompt request and the storyboard template outside the system, and map the storyboard template outside the system to the storyboard template inside the system;

[0020] S12. Serialize and literalize the content in the storyboard template to obtain at least one prompt request segment, and each prompt request segment corresponds to a picture or a storyboard video;

[0021] S13. Submit the prompt requests segment by segment, one by one or in batches to the large model, and let the large model generate image or video content.

[0022] One or more technical solutions provided by the present invention have at least the following technical effects or advantages: constructing a systematic storyboard template for users, automatically extracting the entity configuration of the prompt request, importing and editing the configuration to generate a storyboard template, or directly importing a storyboard template outside the system, so as to realize the systematic and structured editing and management of the prompt request, and improve the interaction experience of the generated page; at the same time, it supports single generation or batch generation of segmented prompt requests, greatly improving the generation efficiency of images and videos; based on the information dynamic editing function of the structured storyboard template, users can supplement the description dimensions at any time to ensure the generation quality of the large model for images or videos; and during the playback process of the generated result, corresponding parameters on the storyboard template can be modified or added in real time according to the user's playback interaction input and submitted to the large model to realize playback dynamic interaction, greatly improving the playback interaction experience.

[0023] The above description is only an overview of the technical solutions of the present invention. In order to be able to understand the technical means of the present invention more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features and advantages of the present invention more obvious and understandable, the specific embodiments of the present invention are specifically given below. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] The present invention will be further described below with reference to the accompanying drawings in conjunction with embodiments.

[0025] Figure 1 It is a schematic diagram of an interactive page for an existing large model to realize the generation function from text to image according to a prompt request.

[0026] Figure 2 It is a schematic diagram of an interactive page for an existing large model to realize the generation function from text to video according to a prompt request.

[0027] Figure 3 It is a diagram of one frame of a video generated by an existing large model according to a prompt request.

[0028] Figure 4 It is a flowchart of the method in Embodiment 1 of the present invention.

[0029] Figure 5 It is a schematic diagram of the structure of a storyboard template inside the system in one embodiment of the present invention.

[0030] Figure 6 It is a schematic diagram of the structure of a storyboard template outside the system in one embodiment of the present invention.

[0031] Figure 7This is the flowchart of the method in the second embodiment of the present invention.

[0032] Figure 8 This is the system block diagram in the third embodiment of the present invention.

[0033] Figure 9 This is the flowchart of the dynamic interaction process playback in an embodiment of the present invention. Detailed implementation manners

[0034] In the embodiments of the present application, by providing a method and system for generating intelligent pictures or videos for large model interaction, a systematic storyboard template is constructed for users, so as to achieve systematic and structured editing and management. At the same time, batch generation is supported, the interaction methods are rich and diverse, and the user interaction experience is improved.

[0035] The overall idea of the technical solution in the embodiments of the present application is as follows: construct a systematic storyboard template for users, automatically extract the entity configuration of the prompt request, import and edit the configuration to generate a storyboard template, or directly import the storyboard template outside the system, so as to achieve systematic and structured editing and management of the prompt request, and improve the interaction experience of the generated page; at the same time, single generation or batch generation of segmented prompt requests is supported, greatly improving the generation efficiency of pictures and videos; based on the information dynamic editing function of the structured storyboard template, users can supplement the description dimensions at any time to ensure the generation quality of pictures or videos by the large model; and during the playback process of the generated result, corresponding parameters on the storyboard template can be modified or added in real time according to the user's playback interaction input and submitted to the large model to achieve playback dynamic interaction, greatly improving the playback interaction experience.

[0036] Embodiment 1

[0037] As Figure 4 shown, this embodiment provides a method for generating intelligent pictures or videos for large model interaction, including the generation process of pictures or videos. The generation process of pictures or videos includes the following steps:

[0038] S11. Receive the user's original prompt request, extract the description information in the original prompt request and fill it into the storyboard template inside the system; or

[0039] Directly import the original prompt request and the storyboard template outside the system, and map the storyboard template outside the system to the storyboard template inside the system; the storyboard template outside the system is a text file, a table file, or a markdown configuration file in json or yaml.

[0040] As Figure 5 and Figure 6As shown, the storyboard template is displayed in the form of a table on the generated interactive page and is in a state of being editable for a second time;

[0041] The content of the storyboard template includes scene number, shot number, picture reference, object, scene type, camera position number, shooting angle, camera movement method, equipment, shot length, sound, picture description, editing and transition method, lighting, generation times number, shooting time; thus covering the tools for creative planning, picture design, shooting production, and post - production editing to understand and communicate about the pictures, which can more accurately convey the needs of visual planning, production, or the director, and speed up the communication and production efficiency.

[0042] S12. Serialize and literalize the content in the storyboard template to obtain at least one prompt request segment, and each prompt request segment corresponds to a picture or a storyboard video;

[0043] Among them, the prompt request segment can be correspondingly filled in the storyboard template.

[0044] S13. Submit the prompt request segments one by one or in batches to the large - model, that is, set a trigger button for single - item execution for each prompt request segment and set a trigger button for batch execution at the entire storyboard template. After the trigger execution, the large - model can generate picture or video content and then return.

[0045] As Figure 4 shown, usually video playback includes conventional playback methods such as normal playback, zoom playback, speed - up playback, skip playback, etc. And the generation method of this embodiment can also include a dynamic interaction process during playback:

[0046] S21. When playing the generated video result, receive the user's playback interaction input in real - time;

[0047] S22. Based on the preset playback interaction rules or playback interaction use cases, according to the user's playback interaction input and based on the time stamp, find the corresponding video reference frames and description texts, modify or add the corresponding parameters on the storyboard template in real - time, and submit them to the large - model;

[0048] S23. The large - model regenerates the picture or video content in real - time and then returns.

[0049] As Figure 9As shown, the dynamic interaction process during playback allows users to interact with the sequential frames in the video. Through playback interaction input, based on the timestamp, the video reference frame corresponding to the index, and the descriptive text, the corresponding parameters on the storyboard template can be modified or added in real time to achieve interactive picture content. For example, video frame actions simulated based on mouse interaction, or the supplementary generation of picture special effects corresponding to keyboard shortcuts, etc. Thus, the video playback interaction is extended from static sequential frame playback interaction to immersive interaction within the generated content.

[0050] Among them, the user's playback interaction input includes but is not limited to keyboard input, mouse input, audio input, motion input, video input, etc.; the playback interaction rules or the playback interaction use cases are provided by the system or defined by the user themselves.

[0051] Taking mouse operation as an example for the playback interaction rules:

[0052] During video playback, when clicking the left mouse button and dragging left / right, the camera movement parameter is modified to push left / push right;

[0053] During video playback, when clicking the right mouse button and dragging left / right, the camera movement parameter is modified to pan left / pan right; when clicking the right mouse button and dragging forward / backward, pan up / pan down;

[0054] During video playback, when scrolling the mouse wheel, it is for zooming in; when clicking the mouse wheel forward / backward, it is for moving forward / backward.

[0055] The playback interaction use case can be, for example, rain or snow falling on the screen.

[0056] Embodiment 2

[0057] As Figure 7 As shown, Embodiment 2 of the present invention provides a method for generating intelligent pictures or videos with large model interaction. Based on Embodiment 1, it also provides the function of generating a suggested prompt. That is, in step S12, after serializing and verbalizing the content in the storyboard template to obtain at least one prompt request segment, each of the prompt request segments can also generate a suggested prompt according to the user's trigger, and replace the prompt request segment with the suggested prompt and submit it to the large model. Thus, it can make up for the deficiencies of the user's prompt request segments and further improve the generation quality of pictures or videos by the large model.

[0058] Among them, the user's trigger can be set at the position corresponding to each of the prompt request segments in the storyboard template on the generation interaction page.

[0059] Embodiment 2 also provides an option to submit prompt requests segment by segment, either one by one or in batches, to the large model. That is, a single - execution trigger button is set for each prompt request segment, and a batch - execution trigger button is set for the entire storyboard template. After the user triggers the execution, the large model can generate image or video content and then return it.

[0060] Embodiment 3

[0061] As Figure 8 shown, in this embodiment, a system for generating intelligent images or videos for large - model interaction is provided, which is a device corresponding to the method in Embodiment 1, including a template management module, and may also include a playback dynamic interaction module.

[0062] As Figure 4 shown, the template management module is used to perform the following steps:

[0063] S11. Receive the user's original prompt request, extract the description information in the original prompt request and fill it into the storyboard template inside the system; or

[0064] Directly import the original prompt request and the storyboard template outside the system, and map the storyboard template outside the system to the storyboard template inside the system; the storyboard template outside the system is a text file, a table file, or a markdown configuration file in json or yaml format.

[0065] S12. Serialize and literalize the content in the storyboard template to obtain at least one prompt request segment, and each prompt request segment corresponds to an image or a storyboard video;

[0066] S13. Submit the prompt request segments to the large model one by one or in batches, and the large model generates image or video content.

[0067] Embodiment 2 also provides an option to submit prompt requests segment by segment, either one by one or in batches, to the large model. That is, a single - execution trigger button is set for each prompt request segment, and a batch - execution trigger button is set for the entire storyboard template. After the user triggers the execution, the large model can generate image or video content and then return it.

[0068] As Figure 5 and Figure 6 shown, the storyboard template is displayed in the form of a table on the generation interaction page and is in a state of being editable twice.

[0069] The content of the storyboard template includes scene number, shot number, picture reference, object, scene type, camera position number, shooting angle, camera movement method, equipment, shot length, sound, picture description, editing and transition method, lighting, generation times number, shooting time;

[0070] The template management module is also used for:

[0071] Filling the prompt requests in segments correspondingly into the storyboard template;

[0072] Generating recommended prompts for the corresponding prompt request segments according to the user's trigger to replace the prompt request segments and submitting them to the large model.

[0073] As Figure 4 shown, the playback dynamic interaction module is used to execute the following process:

[0074] S21. When playing the generated video result, real-time receive the user's playback interaction input;

[0075] S22. Based on the preset playback interaction rules or playback interaction use cases, according to the user's playback interaction input, based on the timestamp, find the corresponding video reference frame and description text, and real-time modify or add the corresponding parameters on the storyboard template and submit them to the large model;

[0076] S23. After the large model regenerates the picture or video content in real time, it returns.

[0077] The playback interaction rules or the playback interaction use cases are provided by the system or defined by the user himself.

[0078] As Figure 9 shown, the playback dynamic interaction process allows the user to interact with the sequence frames in the video. Through the playback interaction input, based on the timestamp and index, the corresponding video reference frame and description text are used to real-time modify or add the corresponding parameters on the storyboard template, realizing the interaction of the picture content, such as the video picture lens movement simulated based on mouse interaction, or the supplementary generation of picture special effects corresponding to keyboard shortcuts, etc., so that the video playback interaction expands from the static sequence frame playback interaction to the immersive interaction within the generated content.

[0079] Among them, the user's playback interaction input includes but is not limited to keyboard input, mouse input, audio input, action input, video input, etc.; the playback interaction rules or the playback interaction use cases are provided by the system or defined by the user himself.

[0080] Taking the mouse operation as an example for the playback interaction rules:

[0081] When the video is played, click the left button and drag it left / right, and the camera movement parameter is modified to push left / push right;

[0082] When playing the video, right-click and drag left / right, and the camera movement parameters are modified to pan left / pan right; right-click and drag forward / backward, and pan up / pan down.

[0083] When playing the video, the operation of scrolling the mouse wheel is to zoom in; click the mouse wheel forward / backward to fast forward / rewind.

[0084] The playback interaction use case can be, for example, rain or snow falling on the screen.

[0085] In summary, the technical solution provided in the embodiment of the present application has at least the following technical effects or advantages: constructing a systematic storyboard template for users, automatically extracting the entity configuration of the prompt request, importing and editing the configuration to generate a storyboard template, or directly importing the storyboard template outside the system, so as to realize the systematic and structured editing and management of the prompt request, and improve the interaction experience of the generated page; at the same time, it supports single generation or batch generation of prompt request segments, greatly improving the generation efficiency of pictures and videos; based on the information dynamic editing function of the structured storyboard template, users can supplement the description dimension at any time to ensure the generation quality of pictures or videos by the large model; and during the playback of the generated result, corresponding parameters on the storyboard template can be modified or added in real time according to the user's playback interaction input and submitted to the large model to realize playback dynamic interaction, greatly improving the playback interaction experience.

[0086] Although the specific implementation manners of the present invention have been described above, those skilled in the art of this technology should understand that the specific embodiments we described are illustrative rather than used to limit the scope of the present invention. Equivalent modifications and changes made by those skilled in the art in accordance with the spirit of the present invention should be covered by the scope protected by the claims of the present invention.

Claims

1. A method for generating intelligent pictures or videos of large model interactions, characterized in that: The process of generating a picture or a video includes the following steps: S11, receiving an original prompt request from a user, extracting description information from the original prompt request and filling it into a storyboard template inside the system; or Directly import the original prompt request and the storyboard template outside the system, and map the storyboard template outside the system to the storyboard template inside the system; S12, serializing and textualizing the content in the storyboard template to obtain at least one prompt request segment, each prompt request segment corresponding to a picture or a storyboard video; S13. Submit the prompt request in segments one by one or in batches to the big model, and the big model generates the image or video content and returns it.

2. The method for generating a large model interactive intelligent picture or video according to claim 1, characterized in that: The storyboard template is displayed in the generated interactive page in the form of a table and is in a secondary editable state; The content of the storyboard template includes scene number, shot number, picture reference, object, scene type, camera position number, shooting angle, camera movement method, equipment, shot length, sound, picture description, editing and transition method, lighting, generation number, and shooting time; The storyboard template outside the system is a text file, a table file, or a markdown configuration file of json or yaml.

3. The method for generating a large model interactive intelligent picture or video according to claim 1, characterized in that: The prompt request segment is filled in the storyboard template accordingly. Each prompt request segment can also generate a suggestion prompt according to the user's trigger. The suggestion prompt replaces the prompt request segment and is submitted to the big model.

4. The method for generating a large model interactive intelligent picture or video according to claim 1, characterized in that: It also includes the playback of dynamic interactions: S21, when generating a video result to play, receiving a user's playback interaction input in real time; S22, based on preset playback interaction rules or playback interaction use cases, according to the playback interaction input of the user, based on the timestamp, searching for the corresponding video reference frame and description text, modifying or adding corresponding parameters on the storyboard template in real time, and submitting the parameters to the big model; S23, the large model regenerates the image or video content in real time and then returns.

5. The method for generating a large model interactive intelligent picture or video according to claim 4, characterized in that: The playback interaction rules or the playback interaction use cases are provided by the system or defined by the user.

6. A system for generating intelligent pictures or videos of large model interactions, characterized in that: It includes a template management module, which is used to perform the following steps: S11, receiving an original prompt request from a user, extracting description information from the original prompt request and filling it into a storyboard template inside the system; or Directly import the original prompt request and the storyboard template outside the system, and map the storyboard template outside the system to the storyboard template inside the system; S12, serializing and textualizing the content in the storyboard template to obtain at least one prompt request segment, each prompt request segment corresponding to a picture or a storyboard video; S13. Submit the prompt request in segments one by one or in batches to the big model, and the big model generates pictures or video content.

7. The system for generating intelligent pictures or videos of large-scale model interactions according to claim 6, characterized in that: The storyboard template is displayed in the generated interactive page in the form of a table and is in a secondary editable state; The content of the storyboard template includes scene number, shot number, picture reference, object, scene type, camera position number, shooting angle, camera movement method, equipment, shot length, sound, picture description, editing and transition method, lighting, generation number, and shooting time; The storyboard template outside the system is a text file, a table file, or a markdown configuration file of json or yaml.

8. The system for generating intelligent pictures or videos of large-scale model interactions according to claim 6, characterized in that: The template management module is also used for: Fill the prompt request segments into the storyboard template accordingly; According to the user's trigger, the corresponding prompt request segment generates a suggestion prompt to replace the prompt request segment and submit it to the big model.

9. The system for generating intelligent pictures or videos of large-scale model interaction according to claim 6, characterized in that: It also includes a dynamic interaction module for playing, and the dynamic interaction module for playing is used to perform the following process: S21, when generating a video result to play, receiving a user's playback interaction input in real time; S22, based on preset playback interaction rules or playback interaction use cases, according to the playback interaction input of the user, based on the timestamp, searching for the corresponding video reference frame and description text, modifying or adding corresponding parameters on the storyboard template in real time, and submitting the parameters to the big model; S23, the large model regenerates the image or video content in real time and then returns.

10. The system for generating intelligent pictures or videos of large-scale model interaction according to claim 9, characterized in that: The playback interaction rules or the playback interaction use cases are provided by the system or defined by the user.