Lens splitting image generation method and device, electronic equipment and storage medium
Through the combination of large language model and literary and artistic drawing model, the storyboard image is automatically generated, which solves the problems of high cost and low efficiency of storyboard drawing in traditional film and television production, and achieves fast and accurate storyboard image generation, reducing costs and improving production efficiency.
Patent Information
- Application Number
- CN202510852833.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-24
- Publication Date
- 2025-08-22
AI Technical Summary
In traditional film and television production, the drawing of storyboard pictures relies on professional storyboarders, which is costly and time-consuming, which affects production efficiency.
By obtaining the target script, using the large language model to generate scene description information, and converting it into the Wensheng Diagram model into a storyboard style, and combining the Wensheng Diagram model to train the sample storyboard to generate a storyboard.
It reduces the cost of generating storyboard pictures, improves the efficiency of film and television production, and improves the accuracy and user experience of storyboard pictures.
Smart Images

Figure CN120529147A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data processing technology, and in particular to a method, device, electronic device and storage medium for generating storyboards. Background Art
[0002] In the traditional film and television production process, storyboards are usually needed to guide the director and creative team to complete the filming of the film.
[0003] The creation of storyboards relies on professional storyboard artists, which is expensive and unaffordable for small-budget crews. Furthermore, storyboard artists often spend a long time creating storyboards, which affects the efficiency of film and television production. Therefore, a method for generating storyboards is needed to reduce the cost of generating storyboards and improve the efficiency of film and television production. Summary of the Invention
[0004] The purpose of the embodiments of the present application is to provide a method, device, electronic device, and storage medium for generating storyboards to reduce the cost of generating storyboards and improve the efficiency of film and television production. The specific technical solutions are as follows:
[0005] In a first aspect of the present application, a method for generating a storyboard is provided, the method comprising:
[0006] Get the target script;
[0007] Inputting the target script into a large language model to obtain first scene description information output by the large language model; wherein the first scene description information is used to describe the scene represented by the target script;
[0008] Inputting the first scene description information into a Wensheng graph model, generating a target scene graph whose displayed content matches the first scene description information, and performing stylization processing on the target scene graph to convert the image style of the target scene graph into a storyboard style, thereby obtaining a first storyboard;
[0009] Obtain a first storyboard output by the Wensheng graph model.
[0010] In a possible embodiment, after obtaining the Wensheng graph model to output the first storyboard, the method further includes:
[0011] Acquire second scene description information, where the second scene description information is obtained after updating the lens angle and / or lens focal length in the first scene description information;
[0012] Inputting the second scene description information and the first storyboard into a Wensheng graph model to obtain a second storyboard output by the Wensheng graph model; wherein the second storyboard displays the same scene as the first storyboard, and the lens angle and / or lens focal length of the second storyboard conform to the second scene description information;
[0013] Obtain a second storyboard output by the Wensheng graph model.
[0014] In a possible embodiment, after obtaining the target script, the method further includes:
[0015] Determining era information of the target script according to the content of the target script; wherein the era information represents the target era described by the target script;
[0016] Inputting the target script into the large language model to obtain first scene description information output by the large language model includes:
[0017] The target script and the era information are input into the large language model, the third scene description information in the target script is extracted, and the third scene description information is modified to the first scene description information that conforms to the target era, and the first scene description information output by the large language model is obtained.
[0018] In a possible embodiment, before inputting the first scene description information into the text graph model to generate a target scene graph whose displayed content matches the first scene description information, the method further includes:
[0019] Adjusting the first scene description information by using a target adjustment method;
[0020] The step of inputting the first scene description information into a text graph model to generate a target scene graph whose displayed content matches the first scene description information includes:
[0021] The adjusted first scene description information is input into the Wensheng graph model to generate a target scene graph whose displayed content matches the adjusted first scene description information.
[0022] In a possible embodiment, the method further includes:
[0023] Acquiring first information, where the first information includes multiple values of information of a preset type in the first scene description information, where the preset type of information is used to describe shooting information, where the shooting information is used to describe a relative angle and / or relative distance between a shooting device and a shooting subject;
[0024] Inputting the target script into the large language model to obtain first scene description information output by the large language model includes:
[0025] The first information and the target script are input into a large language model to select a target value from multiple values of the preset type of information based on the target script as the value of the preset type of information in the first scene description information.
[0026] In a second aspect of the present application, a storyboard generating device is provided, the device comprising:
[0027] The first acquisition module is used to acquire the target script;
[0028] A first input module is configured to input the target script into a large language model to obtain first scene description information output by the large language model; wherein the first scene description information is used to describe the scene represented by the target script;
[0029] A second input module is configured to input the first scene description information into a Wensheng graph model, generate a target scene graph whose displayed content matches the first scene description information, and perform stylization processing on the target scene graph to convert the image style of the target scene graph into a storyboard style to obtain a first storyboard;
[0030] The second acquisition module is used to acquire the first storyboard output by the Wensheng graph model.
[0031] In a possible embodiment, the device further includes:
[0032] a third acquisition module, configured to acquire second scene description information, where the second scene description information is obtained after updating the lens angle and / or lens focal length in the first scene description information;
[0033] a third input module, configured to input the second scene description information and the first storyboard into a Wensheng diagram model, and obtain a second storyboard output by the Wensheng diagram model; wherein the second storyboard displays the same scene as the first storyboard, and the lens angle and / or lens focal length of the second storyboard conform to the second scene description information;
[0034] The fourth acquisition module is used to obtain the second storyboard output by the Wensheng graph model.
[0035] In a possible embodiment, the device further includes:
[0036] A first determining module is configured to determine era information of the target script based on the content of the target script; wherein the era information indicates a target era described by the target script;
[0037] The first input module is specifically used to:
[0038] The target script and the era information are input into the large language model, the third scene description information in the target script is extracted, and the third scene description information is modified to the first scene description information that conforms to the target era, and the first scene description information output by the large language model is obtained.
[0039] In a possible embodiment, the device further includes:
[0040] an information adjustment module, adjusting the first scene description information by a target adjustment method;
[0041] The second input module is specifically used to:
[0042] The adjusted first scene description information is input into the Wensheng graph model to generate a target scene graph whose displayed content matches the adjusted first scene description information, and the target scene graph is stylized to convert the image style of the target scene graph into a storyboard style to obtain a first storyboard.
[0043] In a possible embodiment, the device further includes:
[0044] a fifth acquisition module, configured to acquire first information, the first information including multiple values of information of a preset type in the first scene description information, the preset type of information being used to describe shooting information, the shooting information being used to describe a relative angle and / or relative distance between a shooting device and a subject;
[0045] The first input module is specifically used to:
[0046] The first information and the target script are input into a large language model to select a target value from multiple values of the preset type of information based on the target script as the value of the preset type of information in the first scene description information.
[0047] In a third aspect of the present application, an electronic device is provided, including a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other via the communication bus;
[0048] Memory for storing computer programs;
[0049] The processor is configured to implement any of the method steps described in the first aspect when executing a program stored in the memory.
[0050] In a fourth aspect of the implementation of the present application, a computer-readable storage medium is further provided, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, any of the above-mentioned storyboard generation methods is implemented.
[0051] The storyboard generation method provided by the embodiment of the present application, when the target script is obtained, the target script is input into the large language model, and the first scene description information output by the large language model can be obtained, wherein the first scene description information is used to describe the scene represented by the target script; the first scene description information is input into the Wensheng graph model, a target scene graph whose displayed content matches the first scene description information is generated, and the target scene graph is stylized, and the image style of the target scene graph is converted into the storyboard style, and the first storyboard can be obtained. By applying the embodiment of the present application, the first storyboard corresponding to the target script can be quickly obtained, which saves the cost of professional storyboard artists, reduces the cost of generating storyboards, and improves the efficiency of film and television production. In addition, the embodiment of the present application will first generate a target scene graph that matches the first scene description information, and then convert the image style of the target scene graph into the storyboard style to obtain the first storyboard, so that the picture finally obtained can conform to the storyboard style, improve the accuracy of generating storyboards, and enhance the user experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art.
[0053] Figure 1 A schematic flow chart of a first storyboard generation method provided in an embodiment of the present application;
[0054] Figure 2 A schematic flow chart of a second storyboard generation method provided in an embodiment of the present application;
[0055] Figure 3 A schematic flow chart of a third storyboard generation method provided in an embodiment of the present application;
[0056] Figure 4 A schematic flow chart of a fourth storyboard generation method provided in an embodiment of the present application;
[0057] Figure 5a A schematic diagram of the first storyboard provided in an embodiment of the present application;
[0058] Figure 5b A schematic diagram of a second storyboard provided in an embodiment of the present application;
[0059] Figure 5c A schematic diagram of a third storyboard provided in an embodiment of the present application;
[0060] Figure 5d A schematic diagram of a fourth storyboard provided in an embodiment of the present application;
[0061] Figure 6 A schematic structural diagram of a storyboard generating device provided in an embodiment of the present application;
[0062] Figure 7 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0063] The technical solutions in the embodiments of the present application will be described below in conjunction with the drawings in the embodiments of the present application.
[0064] The present application embodiment provides a method for generating a storyboard. Figure 1 , which is a flow chart of the first storyboard generation method provided in an embodiment of the present application, and the above method includes S101-S104.
[0065] S101, obtain the target script.
[0066] In S101, the target script can be any script used for film and television production. The script can be divided according to different classification standards.
[0067] For example, based on the content of the script, scripts can be divided into film and television scripts, drama scripts, and graphic design scripts. Film and television scripts can include movie scripts, TV scripts, animation scripts, etc. Drama scripts can include drama scripts, opera scripts, dance scripts, etc. Graphic design scripts can include comic scripts, advertising scripts, etc.
[0068] Based on the source of the script story, the script can be divided into original script, adapted script, etc. Original script is a newly created story, and adapted script is a story based on the original novel or historical events.
[0069] The format of the target script can be .doc, or it can be .docx, .txt, .pdf, .ppt and other formats. The embodiment of the present application does not limit the format of the target script.
[0070] S102: Input the target script into the large language model to obtain first scene description information output by the large language model.
[0071] The first scene description information is used to describe the scene represented by the target script. The first scene description information may include at least one of the following information: a brief description of the content, scene information, character information, and shooting information.
[0072] In S102, before inputting the target script into the large language model, a scene description information extraction instruction may be input into the large language model so that the large language model extracts the first scene description information from the target script according to the storyboard description information extraction instruction.
[0073] The scene description information extraction instruction includes the operations that the large language model needs to perform and the format of the output first scene description information.
[0074] For example, the scene description information extraction instruction may be the following:
[0075] You are a director with rich experience in storyboard drawing. Based on the content of the provided script, analyze the text description of the corresponding storyboard (including: brief description of the content, character setting, emotional expression, background structure, picture setting, photography techniques), and use it as prompt words for the text image (with as much details as possible, including detailed descriptions of the surrounding scenes and characters, etc.).
[0076] [Output format]: content description, scene information, character information, shooting information.
[0077] The summary is used to describe the main plot of the story. For example, an ancient man carries a dish and walks towards two other men sitting at a table.
[0078] Character information refers to the relevant information about the character, such as the number of characters, character modeling information, character action information, etc.
[0079] Shooting information refers to the elements that affect the shooting effect, such as shooting angle, camera position, exposure, lighting, etc.
[0080] In a possible embodiment, the above method further includes the following step A.
[0081] Step A: Obtain first information.
[0082] The first information includes multiple numerical values of information of a preset type in the first scene description information. The preset information type is used to describe shooting information, which is used to describe the relative angle and / or relative distance between the camera and the subject. In one example, the shooting information includes at least one of the following: shooting angle information, camera position information, and focal length information. Shooting angle information describes the camera's shooting angle; camera position information describes the relative position of the camera and the subject; and focal length information is used to select an appropriate lens.
[0083] Specifically, the shooting angle represented by the shooting angle information can be: Dutch angle, overhead shot, overhead shot, straight-on shot, straight-on shot, and flat shot; among them, "Dutch angle", "overhead shot", "overhead shot", "straight-on shot", "straight-on shot", and flat shot" are multiple values in the shooting angle information.
[0084] and / or
[0085] The shooting position indicated by the shooting position information can be: high position, low position, and flat position; high position means that the camera is above the subject; low position means that the camera is below the subject; flat position means that the camera and the subject are on the same horizontal line; among them, "high position", "low position", and "flat position" are multiple values in the shooting position information.
[0086] and / or
[0087] The lens represented by the shooting focal length information can be: fisheye lens, wide-angle lens, standard lens, telephoto lens; the focal length of the fisheye lens is smaller than the focal length of the wide-angle lens, the focal length of the wide-angle lens is smaller than the focal length of the standard lens, and the focal length of the standard lens is smaller than the focal length of the telephoto lens; among them, "fisheye lens", "wide-angle lens", "telephoto lens", and "standard lens" are multiple values in the shooting focal length information.
[0088] Among them, the Dutch angle refers to the camera being set to a certain tilt angle during shooting, which is used to increase the visual impact of the picture and create an uneasy, tense or disorienting atmosphere. It is usually used in horror films, film noir and other film and television dramas that need to express an uneasy or weird atmosphere.
[0089] Bird's-eye view refers to shooting from above the subject, and is often used to capture large scenes, such as cityscapes, construction sites, and large gatherings.
[0090] Shooting from above means shooting from below the subject, and is often used to photograph buildings, great people, etc., to show their height and majesty. For example, when photographing skyscrapers.
[0091] A straight-on camera refers to a camera lens facing the subject, with the lens and the subject at a 90-degree angle. This is often used to capture facial features.
[0092] A straight-up shot refers to shooting with the camera lens facing the sky or upwards. This angle makes the subject below appear smaller and is often used to shoot people looking up at the distance or magnificent scenery.
[0093] Flat shooting means that the camera lens is completely parallel to the ground, with no obvious upward or downward angle. It can truly reflect the actual size of the object and the proportional relationship with the surrounding environment. It is often used for shooting still lifes.
[0094] High camera positions are suitable for shooting large scenes, vast landscapes, buildings, etc.
[0095] The low camera position is suitable for shooting people or buildings that need to appear tall.
[0096] The flat camera position is suitable for shooting scenes that need to appear natural and realistic, such as daily conversations, walking, etc.
[0097] The focal length of a fisheye lens is usually less than 16mm, and the imaging angle is close to or even exceeds 180 degrees, which can capture a circular distorted field of view.
[0098] The focal length of a wide-angle lens is relatively short, between 14mm and 40mm, and is capable of capturing a wide range of images.
[0099] The focal length of a standard lens is close to the human eye's viewing angle, between 40mm and 50mm, and the image is basically distortion-free.
[0100] The focal length of a telephoto lens is longer, between 70mm and 300mm. It can compress the sense of space, bring distant objects closer, and create an obvious zoom effect, but the viewing angle is narrow.
[0101] In this embodiment, the above S102 includes the following steps a.
[0102] Step a: Input the first information and the target script into the large language model to select a target value from multiple values of the preset type of information based on the target script as the value of the preset type of information in the first scene description information.
[0103] In step a, after the first information and the target script are input into the large language model, the large language model can select a target value from multiple values of the preset type of information according to the content of the target script as the value of the preset type of information in the first scene description information.
[0104] For example, when the first information is shooting angle information, and the values in the first information are "Dutch angle", "overhead shot", "upward shot", "straight on", "straight up", and "horizontal shot", the large language model can select a most appropriate value from "Dutch angle, overhead shot, upward shot, straight on, straight up, and horizontal shot" according to the content of the target script as the value of the shooting angle information in the first scene description information.
[0105] By selecting the above embodiment, the multiple numerical values in the first information can provide the large language model with a selection instruction for a preset type of information, so that the large language model selects a target numerical value from the multiple numerical values of the preset type of information according to the content of the target script, as the numerical value of the preset type of information in the first scene description information, thereby making the first scene description information more precise and accurate, and further improving the accuracy of generating storyboards.
[0106] In a possible embodiment, the scene information includes at least one of the following information: environmental information, weather information;
[0107] The character information includes at least one of the following information: the number of characters, character modeling information, character position information, character angle information, character action information, and character emotion information;
[0108] The shooting information includes at least one of the following information: shooting angle information, shooting position information, and shooting focal length information.
[0109] Specifically, the environment information describes the time of day, the setting, the objects in the setting, and whether the setting is interior or exterior. Interior refers to indoor shots, while exterior refers to outdoor shots. For example, the environment information could be "woods," "exterior," and "daytime"; or "classroom," "podium," "interior," and "nighttime."
[0110] Weather information is used to describe the climatic conditions in which the story takes place, such as sunny, rainy, or snowy days.
[0111] The number of characters is the number of main characters involved in the story. For example, if the story is briefly about a man in ancient times carrying a dish to two other men sitting at a table, the number of characters involved is three.
[0112] Character modeling information is used to describe the character's clothing, hairstyle, makeup, etc. For example, a man in ancient costume, a modern man in casual clothes, a woman in ancient costume with a snake bun, a woman in ancient costume with a flower ornament, etc.
[0113] Character position information is used to describe the relative positions of characters in the scene. For example, Zhang San is on the left side of the screen, and Li Si and Wang Wu are on the right side of the screen.
[0114] Character angle information describes the character's position and angle relative to the camera. For example, Zhang San's position angle is three-quarters right, Li Si's position angle is frontal, and Wang Wu's position angle is sideways.
[0115] Character action information is used to describe the character's actions. For example, if Zhang San carries a dish and walks from the left toward Li Si and Wang Wu, who are sitting at the table, Zhang San's action is "carrying the dish," Li Si's action is "sitting down," and Wang Wu's action is "sitting down."
[0116] Character emotion information is used to describe the character's emotional state and its changes, including happiness, anger, sadness, fear, calmness, etc. For example, Zhang San's emotional state is calm, Li Si's emotional state is angry, and Wang Wu's emotional state is happy.
[0117] In a possible embodiment, the first scene description information further includes a shot number.
[0118] The shot number, also known as the shot sequence number, is a serial number assigned to each shot, used to identify the sequence of shots. Shot numbers can be one or more of Arabic numerals, English letters, and special characters, for example, 1, 1-1, 1A, or 1B. Shot numbers increase in order of their sequence.
[0119] For example, the script "XX Building" is input into the large language model to obtain the first scene description information output by the large language model. The first scene description information is as follows:
[0120] Storyboard number: 1
[0121] Content: An ancient man carries a dish toward two other men sitting at a table. Three men, woods, exterior, daytime
[0122] Characters: [Zhang San, ancient man, ancient robe, black hair, left, 3 / 4 angle right, serving food, calm], [Li Si, ancient man, ancient robe, black hair, right, profile, sitting, unfriendly], [Wang Wu, ancient man, ancient robe, black hair, right, profile, sitting, calm]
[0123] Shooting angle: overhead
[0124] Shooting position: High position
[0125] Shooting focal length: standard lens;
[0126] Scenario No.: 2
[0127] Content: An ancient man pushes a dish to another ancient man, who picks up his chopsticks and tastes a bite of fish. Two people, small table, woods, exterior, daytime
[0128] Characters: [Li Si, ancient man, ancient robe, black hair, left, side view, pushing food, intentionally], [Wang Wu, ancient man, ancient robe, black hair, right, front view, tasting fish, reluctantly]
[0129] Angle: Flat shot
[0130] Seat: Flat
[0131] Focal length: standard lens.
[0132] In the above embodiment, since the shooting angle is one of Dutch angle, overhead shot, overhead shot, straight-on shot, straight-on shot, and horizontal shot, the shooting position is one of high camera position, low camera position, and horizontal camera position, and the shooting lens is one of fisheye lens, wide-angle lens, standard lens, and telephoto lens, the large language model can select a suitable shooting angle, shooting position, and shooting lens according to the target script, thereby making the first storyboard generated by the text graph model closer to the scene described in the target script.
[0133] In a possible embodiment, the first scene description information further includes at least one of the following information: camera movement information, composition information, color tone information, and viewing angle information.
[0134] The camera movement information indicates the camera movement method. This refers to how the camera moves during filming to achieve a specific visual effect. Common camera movement methods include push, pull, pan, shift, follow, swing, lift, and surround.
[0135] Push means that the lens moves forward slowly, gradually approaching the subject, which is used to focus and highlight the subject.
[0136] Pull means that the camera gradually moves backward and away, which is used to explain the environment and background information.
[0137] Pan means rotating the camera in a fixed position to capture the surrounding environment, and is often used to show multiple parts of a scene.
[0138] Shift means parallel movement of the lens, which can be horizontal or vertical. It is suitable for shooting wide scenes or following moving objects.
[0139] Follow means that the lens moves along with the subject, maintaining a relative distance. It is often used to track and shoot moving objects or people.
[0140] Swinging means quickly switching perspectives to create a strong visual impact.
[0141] Lifting means moving the lens vertically to create unique perspectives and visual effects.
[0142] Surround means wrapping around the subject to show its all-round perspective.
[0143] The composition information indicates a composition method, which may be a three-part composition, a two-part composition, or the like.
[0144] Hue information indicates the hue of the storyboard. Hue refers to factors such as color, brightness, and saturation that affect the audience's visual experience, including cool tones and warm tones.
[0145] Perspective information represents the camera's perspective, which can be categorized as either objective or subjective. Objective lenses present events from the narrator's perspective, impersonal and allowing the audience to understand the events through an unseen observer, unrelated to the character's perspective and emotions. Subjective lenses simulate the first-person perspective of a character in a film or television drama, presenting the audience with the characters or objects they see. This lens carries a strong sense of personal emotion and subjectivity, reflecting the character's impressions and feelings about a specific object.
[0146] By selecting the above embodiment, the first scene description information also includes at least one of camera movement information, composition information, color tone information, and perspective information, which can make the first scene description information more accurate, and the generated target storyboard is closer to the content described in the target script, which helps to improve the user experience.
[0147] S103: Input the first scene description information into the Wensheng graph model to generate a target scene graph whose displayed content matches the first scene description information, and perform stylization processing on the target scene graph to convert the image style of the target scene graph into a storyboard style to obtain a first storyboard.
[0148] In S103, the text-based graph model is a multimodal deep learning model that can generate matching images based on text description information. The core principle of the text-based graph model is to convert natural language text into image space and simultaneously link visual features with language information to achieve mapping between natural language text and images.
[0149] Commonly used cultural graph models include Stable Diffusion and Midjourney. Stable Diffusion is an image generation technology based on LDMs (Latent Diffusion Models), while Midjourney is an image generation tool based on artificial intelligence technology.
[0150] When generating a storyboard, the Wensheng graph model in the embodiment of the present application generates two images: a target scene graph and a first storyboard. The target scene graph merely matches the first scene description in terms of displayed content, but its image style is random and may not necessarily be the storyboard style that meets the user's needs. The embodiment of the present application also performs stylization processing on the target scene graph, converting its image style into a storyboard style, and further obtaining a first storyboard with a storyboard style.
[0151] S104, obtaining a first storyboard output by the Wensheng graph model.
[0152] In S104, after the Chinese raw image model generates the first storyboard in the aforementioned S103, the first storyboard output by the Chinese raw image model can be obtained by directly saving it locally, calling and obtaining it through a cloud API (Application Programming Interface), or outputting it through an integrated plug-in / software.
[0153] By selecting the above embodiment, when the target script is obtained, the target script is input into the large language model, and the first scene description information output by the large language model can be obtained, wherein the first scene description information is used to describe the scene represented by the target script; the first scene description information is input into the Wensheng graph model to generate a target scene graph whose displayed content matches the first scene description information, and the target scene graph is stylized, and the image style of the target scene graph is converted into the storyboard style, so as to obtain the first storyboard. By applying the embodiment of the present application, the first storyboard corresponding to the target script can be quickly obtained, which saves the cost of professional storyboard artists, reduces the cost of generating storyboards, and improves the efficiency of film and television production. In addition, the embodiment of the present application will first generate a target scene graph that matches the first scene description information, and then convert the image style of the target scene graph into the storyboard style to obtain the first storyboard, so that the picture finally obtained can conform to the storyboard style, improve the accuracy of generating storyboards, and enhance the user experience.
[0154] In a possible embodiment, the Wensheng graph model is trained using sample storyboards and sample storyboard description information describing the sample storyboards;
[0155] The sample storyboard description information includes at least one of the following information: content summary, scene information, character information, and shooting information.
[0156] The sample storyboards are a visual representation of the information described in the sample storyboards. They transform the storyline, character dialogue, and scene descriptions in the sample storyboards into concrete visual images, allowing the creators and filming team to intuitively understand every detail of the story and the filming requirements, helping the filming team to better execute the filming task.
[0157] By selecting the above embodiment and using the sample storyboards and the sample storyboard description information describing the sample storyboards to train the Wensheng graph model, the first storyboard generated by the Wensheng graph model can be made more accurate, which helps to improve the user experience.
[0158] In one possible embodiment, see Figure 2 , is a flow chart of the second method for generating a storyboard provided in the embodiment of the present application, which is similar to the aforementioned Figure 1 Compared with the embodiment shown in FIG. 1 , S105 to S107 are further included after S104 described above.
[0159] S105: Obtain second scene description information.
[0160] The second scene description information is obtained after updating the lens angle and / or lens focal length in the first scene description information, and is applicable to the situation where the same scene needs to be shot at different shooting angles / positions.
[0161] S106: Input the second scene description information and the first storyboard into the Wensheng graph model to obtain a second storyboard output by the Wensheng graph model.
[0162] The second storyboard displays the same scene as the first storyboard, and the lens angle and / or lens focal length of the second storyboard conform to the second scene description information. For example, if the first storyboard shows Xiao Wang doing homework in the classroom, shot from the front with a short focal length, then the second storyboard can be shot from the side with a short focal length, from the front with a long focal length, or from the side with a long focal length.
[0163] S107, obtaining a second storyboard output by the Wensheng graph model.
[0164] After obtaining the second storyboard output by the Wensheng graph model, the second storyboard output by the Wensheng graph model is obtained. The method of obtaining the second storyboard is similar to the method of obtaining the first storyboard in the aforementioned S104 and will not be repeated here.
[0165] By selecting the above embodiment, a second storyboard can be obtained based on the second scene description information and the first storyboard. Since the second scene description information is obtained after updating the lens angle and / or lens focal length in the first scene description information, when it is necessary to shoot the same scene but at different shooting angles / positions, it can be ensured that the scene displayed by the generated second storyboard is the same as that of the first storyboard, and the lens angle and / or lens focal length of the second storyboard conform to the second scene description information, which can improve the accuracy of storyboard generation at different shooting angles / positions.
[0166] In one possible embodiment, see Figure 3 , is a flow chart of the third method for generating storyboards provided in the embodiment of the present application, which is similar to the aforementioned Figure 1 Compared with the embodiment shown, after S101, S108 is further included.
[0167] S108, determining the era information of the target script based on the content of the target script.
[0168] The era information indicates the target era described by the target scenario.
[0169] In S108 , the era information of the target script may be determined from multiple dimensions according to the content of the target script.
[0170] In one example, the era information of the target script can be determined by lexical or syntactic features. For example, it can be detected whether the sentence pattern of the target script contains classical Chinese sentence patterns. If it contains, the era information of the script is determined to be ancient; if not, the era information of the script is determined to be modern. Further, more specific era information can be determined according to the sentence pattern or lexical style of a specific era. For example, if the sentence patterns of five-character or seven-character poems appear frequently in the target script, or specific words such as "必先, 轮对" appear, the era information of the script is determined to be the Tang Dynasty.
[0171] In another example, it can also be determined by identifying typical scene elements in the target script. For example, the era information of the script can be determined to be ancient through scene elements such as post stations, inns, and marketplaces, and the era information of the script can be determined to be modern through scene elements such as offices, subway stations, and cafes. Further, more specific era information can be determined according to the scene elements of a specific era. For example, if the scene element of "瓦舍" (or "瓦子", "瓦市") appears frequently in the target script, the era information of the script is determined to be the Song Dynasty.
[0172] In this embodiment, S102 can be implemented by S102A.
[0173] S102A: Input the target script and era information into the large language model, extract the third scene description information in the target script, and modify the third scene description information to the first scene description information that conforms to the target era, so as to obtain the first scene description information output by the large language model.
[0174] The large language model generates the first scene description information according to the target script. However, when there is information in the target script that does not conform to the era described by the script, it will reduce the accuracy of the first scene description information output by the large language model. In S102A, after determining the era information of the target script, inputting the target script and era information into the large language model together can extract the third scene description information in the target script, and modify the third scene description information to the first scene description information that conforms to the target era, so as to obtain the first scene description information.
[0175] Exemplarily, in the case where the era information of the target script is determined to be ancient, if the third scene description information in the target script is "dress", after inputting the target script and the era information of ancient times into the large language model, the third scene description information of "dress" will be extracted and modified to "ruqun" to obtain the first scene description information of "ruqun". It should be noted that this is only an example here, and the specific types of the third scene description information in this application are not limited, and can be character clothing information, scene element information, etc.
[0176] By selecting the above embodiment, the first scene description information output by the large language model can be made completely consistent with the era information of the target script, thereby improving the matching accuracy between the first scene description information output by the large language model and the target script, and further improving the accuracy of generating storyboards.
[0177] In one possible embodiment, see Figure 4 , is a flow chart of the third method for generating storyboards provided in the embodiment of the present application, which is similar to the aforementioned Figure 1 Compared with the embodiment shown, before S103, S109 is also included.
[0178] S109: Adjust the first scene description information through a target adjustment method.
[0179] Among them, the target adjustment method includes adding, deleting or modifying the first scene description information.
[0180] For example, it is assumed that the first scene description information is as follows:
[0181] Storyboard number: 1
[0182] Content: An ancient man carries a dish toward two other men sitting at a table. Three men, woods, exterior, daytime
[0183] Characters: [Zhang San, ancient man, ancient robe, black hair, left, 3 / 4 angle right, serving food, calm], [Li Si, ancient man, ancient robe, black hair, right, profile, sitting, unfriendly], [Wang Wu, ancient man, ancient robe, black hair, right, profile, sitting, calm]
[0184] Angle: Overhead shot
[0185] Camera position: High camera position
[0186] Focal length: standard lens.
[0187] If the target adjustment method is: the camera movement method is to pull, the camera position is deleted, and the day is changed to night, the description information of the first scene after adjustment is as follows:
[0188] Storyboard number: 1
[0189] Content: An ancient man carries a dish toward two other men sitting at a table. Three men, woods, exterior, evening
[0190] Characters: [Zhang San, ancient man, ancient robe, black hair, left, 3 / 4 angle right, serving food, calm], [Li Si, ancient man, ancient robe, black hair, right, profile, sitting, unfriendly], [Wang Wu, ancient man, ancient robe, black hair, right, profile, sitting, calm]
[0191] Angle: Overhead shot
[0192] Focal length: Standard lens
[0193] Camera movement: pull.
[0194] In this embodiment, S103 can be implemented through S103A.
[0195] S103A, input the adjusted first scene description information into the Wensheng graph model, generate a target scene graph whose displayed content matches the adjusted first scene description information, and perform stylization processing on the target scene graph, converting the image style of the target scene graph into a storyboard style to obtain a first storyboard.
[0196] By selecting the above embodiment, before inputting the first scene description information into the Wensheng graph model to obtain the first scene output by the Wensheng graph model, the first scene description information is adjusted. This can make the first storyboard generated based on the adjusted first scene description information more in line with user needs, thereby helping to improve user experience.
[0197] The following combination Figure 5a-5d The above-mentioned storyboard generation method is described.
[0198] For example, assume that the target script is a scene from the TV series "XX Building", which is divided into four storyboards, which are as follows:
[0199] In the woods, there is a building called xx. There is a small table next to it.
[0200] Storyboard 1:
[0201] Li Si looked at Wang Wu with a bad expression.
[0202] Zhang San served the dishes.
[0203] Zhang San: Here are some new dishes I've developed. Try them. This is pickled fish, this is boiled chicken legs, and this is called Willow Sparrow. I stuffed some young sparrows I shot, stuffed them with lime and five-spice, and then roasted them over a scorching fire for a whole day. Try them.
[0204] …
[0205] Storyboard 2:
[0206] Li Si deliberately pushed the two dishes to Wang Wu: You are the guest, you go first.
[0207] Wang Wu didn't suspect anything. He picked up the chopsticks and tasted the fish. It was so fishy that he almost vomited.
[0208] …
[0209] Storyboard 3:
[0210] Li Si and Wang Wu grabbed the chicken leg at the same time, and neither of them gave in or let go of their chopsticks.
[0211] …
[0212] Storyboard 4:
[0213] Zhang San coughed for a while and felt that the pain was coming back.
[0214] Wang Wu suddenly asked: Does he know that you were poisoned?
[0215] Zhang San paused: This matter is only known to you and me, I hope you keep it a secret.
[0216] Wang Wu: I have a feeling...
[0217] The above script content is input into the large language model, and the first scene description information obtained is as follows:
[0218] Storyboard number: 1
[0219] Content: An ancient man carries a dish toward two other men sitting at a table. Three men, woods, exterior, daytime
[0220] Characters: [Zhang San, ancient man, ancient robe, black hair, left, 3 / 4 angle right, serving food, calm], [Li Si, ancient man, ancient robe, black hair, right, profile, sitting, unfriendly], [Wang Wu, ancient man, ancient robe, black hair, right, profile, sitting, calm]
[0221] Angle: Overhead shot
[0222] Camera position: High camera position
[0223] Focal length: standard lens;
[0224] Scenario No.: 2
[0225] Content: An ancient man pushes a dish to another ancient man, who picks up his chopsticks and tastes a bite of fish. Two people, small table, woods, exterior, daytime
[0226] Characters: [Li Si, ancient man, ancient robe, black hair, left, side view, pushing food, intentionally], [Wang Wu, ancient man, ancient robe, black hair, right, front view, tasting fish, reluctantly]
[0227] Angle: Flat shot
[0228] Seat: Flat
[0229] Focal length: standard lens;
[0230] Scenario No.: 3
[0231] Content: Two ancient men sitting at a table, using chopsticks to pick up cooked chicken legs from the same plate. Two people, small table, woods, exterior, daytime
[0232] Characters: [Li Si, ancient man, ancient robe, black hair, left, profile, mid shot, grabbing a chicken leg, nervous], [Wang Wu, ancient man, ancient robe, black hair, right, profile, mid shot, grabbing a chicken leg, nervous and intense]
[0233] Angle: Flat shot
[0234] Seat: Flat
[0235] Focal length: standard lens;
[0236] Scenario No.: 4
[0237] Content: An ancient man covers his mouth with his hand and coughs violently, while another ancient man looks at him worriedly. Two people, a small table, in the woods, exterior, daytime
[0238] Characters: [Li Si, ancient man, ancient robe, black hair, right, front, medium shot, looking, concerned], [Wang Wu, ancient man, ancient robe, black hair, left, profile, medium shot, covering mouth, suffering]
[0239] Angle: Flat shot
[0240] Seat: Flat
[0241] Focal length: standard lens.
[0242] The first scene description information is input into the Wensheng graph model to obtain the first storyboard output by the Wensheng graph model as follows: Figure 5a-5d shown.
[0243] See also Figure 5a , which is a schematic diagram of the first storyboard provided in an embodiment of the present application.
[0244] in, Figure 5a It is the first storyboard generated by the Wensheng graph model according to the first scene description information carrying the storyboard number 1.
[0245] See also Figure 5b , which is a schematic diagram of the second storyboard provided in an embodiment of the present application.
[0246] in, Figure 5b It is the first storyboard generated by the Wensheng graph model according to the first scene description information carrying the storyboard number 2.
[0247] See also Figure 5c , which is a schematic diagram of the third storyboard provided in an embodiment of the present application.
[0248] in, Figure 5cIt is the first storyboard generated by the Wensheng graph model according to the first scene description information carrying the storyboard number 3.
[0249] See also Figure 5d , which is a schematic diagram of the fourth storyboard provided in an embodiment of the present application.
[0250] in, Figure 5d It is the first storyboard generated by the Wensheng graph model according to the first scene description information carrying the storyboard number 4.
[0251] It should be noted that the above Figure 5a-5d It is only for reference and does not mean that the first storyboard generated by the Wensheng graph model based on the first scene description information can only be Figure 5a-5d shown.
[0252] Corresponding to the aforementioned storyboard generation method, the present application embodiment further provides a storyboard generation device, see Figure 6 , is a schematic structural diagram of a storyboard generating device provided in an embodiment of the present application, the device comprising:
[0253] The first acquisition module 601 is used to acquire the target script;
[0254] A first input module 602 is configured to input the target script into a large language model to obtain first scene description information output by the large language model; wherein the first scene description information is used to describe the scene represented by the target script;
[0255] A second input module 603 is configured to input the first scene description information into a Wensheng graph model, generate a target scene graph whose displayed content matches the first scene description information, and perform stylization processing on the target scene graph to convert the image style of the target scene graph into a storyboard style to obtain a first storyboard;
[0256] The second acquisition module 604 is configured to acquire the first storyboard output by the Wensheng graph model.
[0257] By selecting the above embodiment, when the target script is obtained, the target script is input into the large language model, and the first scene description information output by the large language model can be obtained, wherein the first scene description information is used to describe the scene represented by the target script; the first scene description information is input into the Wensheng graph model to generate a target scene graph whose displayed content matches the first scene description information, and the target scene graph is stylized, and the image style of the target scene graph is converted into the storyboard style, so as to obtain the first storyboard. By applying the embodiment of the present application, the first storyboard corresponding to the target script can be quickly obtained, which saves the cost of professional storyboard artists, reduces the cost of generating storyboards, and improves the efficiency of film and television production. In addition, the embodiment of the present application will first generate a target scene graph that matches the first scene description information, and then convert the image style of the target scene graph into the storyboard style to obtain the first storyboard, so that the picture finally obtained can conform to the storyboard style, improve the accuracy of generating storyboards, and enhance the user experience.
[0258] In a possible embodiment, the device further includes:
[0259] a third acquisition module, configured to acquire second scene description information, where the second scene description information is obtained after updating the lens angle and / or lens focal length in the first scene description information;
[0260] a third input module, configured to input the second scene description information and the first storyboard into a Wensheng diagram model, and obtain a second storyboard output by the Wensheng diagram model; wherein the second storyboard displays the same scene as the first storyboard, and the lens angle and / or lens focal length of the second storyboard conform to the second scene description information;
[0261] The fourth acquisition module is used to obtain the second storyboard output by the Wensheng graph model.
[0262] By selecting the above embodiment, a second storyboard can be obtained based on the second scene description information and the first storyboard. Since the second scene description information is obtained after updating the lens angle and / or lens focal length in the first scene description information, when it is necessary to shoot the same scene but at different shooting angles / positions, it can be ensured that the scene displayed by the generated second storyboard is the same as that of the first storyboard, and the lens angle and / or lens focal length of the second storyboard conform to the second scene description information, which can improve the accuracy of storyboard generation at different shooting angles / positions.
[0263] In a possible embodiment, the device further includes:
[0264] A first determining module is configured to determine era information of the target script based on the content of the target script; wherein the era information indicates a target era described by the target script;
[0265] The first input module is specifically used to:
[0266] The target script and the era information are input into the large language model, the third scene description information in the target script is extracted, and the third scene description information is modified to the first scene description information that conforms to the target era, and the first scene description information output by the large language model is obtained.
[0267] By selecting the above embodiment, the first scene description information output by the large language model can be made completely consistent with the era information of the target script, thereby improving the matching accuracy between the first scene description information output by the large language model and the target script, and further improving the accuracy of generating storyboards.
[0268] In a possible embodiment, the device further includes:
[0269] an information adjustment module, adjusting the first scene description information by a target adjustment method;
[0270] The second input module is specifically used to:
[0271] The adjusted first scene description information is input into the Wensheng graph model to generate a target scene graph whose displayed content matches the adjusted first scene description information, and the target scene graph is stylized to convert the image style of the target scene graph into a storyboard style to obtain a first storyboard.
[0272] By selecting the above embodiment, before inputting the first scene description information into the Wensheng graph model to obtain the first scene output by the Wensheng graph model, the first scene description information is adjusted. This can make the first storyboard generated based on the adjusted first scene description information more in line with user needs, thereby helping to improve user experience.
[0273] In a possible embodiment, the device further includes:
[0274] a fifth acquisition module, configured to acquire first information, the first information including multiple values of information of a preset type in the first scene description information, the preset type of information being used to describe shooting information, the shooting information being used to describe a relative angle and / or relative distance between a shooting device and a subject;
[0275] The first input module is specifically used to:
[0276] The first information and the target script are input into a large language model to select a target value from multiple values of the preset type of information based on the target script as the value of the preset type of information in the first scene description information.
[0277] Before inputting the first scene description information into the Wensheng graph model to obtain the first scene output by the Wensheng graph model, adjusting the first scene description information can make the first storyboard generated according to the adjusted first scene description information more in line with user needs, which helps to improve user experience.
[0278] In a possible embodiment, the Wensheng graph model is trained using sample storyboards and sample storyboard description information describing the sample storyboards;
[0279] The sample storyboard description information includes at least one of the following information: storyboard number, content summary, number of characters, environment information, weather information, character modeling information, character position information, character angle information, character action information, character emotion information, shooting angle information, shooting camera position information, and shooting focal length information.
[0280] By selecting the above embodiment and using the sample storyboards and the sample storyboard description information describing the sample storyboards to train the Wensheng graph model, the first storyboard generated by the Wensheng graph model can be made more accurate, which helps to improve the user experience.
[0281] In a possible embodiment, the scene information includes at least one of the following information: environmental information, weather information;
[0282] The character information includes at least one of the following information: the number of characters, character modeling information, character position information, character angle information, character action information, and character emotion information;
[0283] The shooting information includes at least one of the following information: shooting angle information, shooting position information, and shooting focal length information.
[0284] In the above embodiment, the scene information includes at least one of environmental information and weather information, the character information includes at least one of the number of characters, character modeling information, character position information, character angle information, character action information, and character emotional information, and the shooting information includes at least one of the shooting angle information, shooting camera position information, and shooting focal length information. The large language model can output matching scene information, character information, and shooting information based on the target script, thereby making the first storyboard generated by the literary graph model closer to the scene described in the target script.
[0285] In a possible embodiment, the shooting angle represented by the shooting angle information is one of the following angles: Dutch angle, overhead shot, overhead shot, straight shot, straight upward shot, and horizontal shot;
[0286] and / or
[0287] The camera position information indicates one of the following camera positions: high camera position, low camera position, and level camera position. High camera position indicates that the camera is above the subject; low camera position indicates that the camera is below the subject; and level camera position indicates that the camera and the subject are on the same horizontal line.
[0288] and / or
[0289] The lens represented by the shooting focal length information is one of the following lenses: fisheye lens, wide-angle lens, standard lens, and telephoto lens; the focal length of the fisheye lens is smaller than the focal length of the wide-angle lens, the focal length of the wide-angle lens is smaller than the focal length of the standard lens, and the focal length of the standard lens is smaller than the focal length of the telephoto lens.
[0290] In the above embodiment, since the shooting angle is one of Dutch angle, overhead shot, overhead shot, straight-on shot, straight-on shot, and horizontal shot, the shooting position is one of high camera position, low camera position, and horizontal camera position, and the shooting lens is one of fisheye lens, wide-angle lens, standard lens, and telephoto lens, the large language model can select a suitable shooting angle, shooting position, and lens according to the target script, thereby making the first storyboard generated by the literary graph model closer to the scene described in the target script.
[0291] The present application also provides an electronic device, such as Figure 7 As shown, it includes a processor 701, a communication interface 702, a memory 703 and a communication bus 704, wherein the processor 701, the communication interface 702, and the memory 703 communicate with each other through the communication bus 704.
[0292] Memory 703, used for storing computer programs;
[0293] The processor 701 is configured to execute the program stored in the memory 703, and implement the following steps:
[0294] Get the target script;
[0295] Inputting the target script into a large language model to obtain first scene description information output by the large language model; wherein the first scene description information is used to describe the scene represented by the target script;
[0296] Inputting the first scene description information into a Wensheng graph model, generating a target scene graph whose displayed content matches the first scene description information, and performing stylization processing on the target scene graph to convert the image style of the target scene graph into a storyboard style, thereby obtaining a first storyboard;
[0297] Obtain a first storyboard output by the Wensheng graph model.
[0298] By selecting the above embodiment, when the target script is obtained, the target script is input into the large language model, and the first scene description information output by the large language model can be obtained, wherein the first scene description information is used to describe the scene represented by the target script; the first scene description information is input into the Wensheng graph model to generate a target scene graph whose displayed content matches the first scene description information, and the target scene graph is stylized, and the image style of the target scene graph is converted into the storyboard style, so as to obtain the first storyboard. By applying the embodiment of the present application, the first storyboard corresponding to the target script can be quickly obtained, which saves the cost of professional storyboard artists, reduces the cost of generating storyboards, and improves the efficiency of film and television production. In addition, the embodiment of the present application will first generate a target scene graph that matches the first scene description information, and then convert the image style of the target scene graph into the storyboard style to obtain the first storyboard, so that the picture finally obtained can conform to the storyboard style, improve the accuracy of generating storyboards, and enhance the user experience.
[0299] The communication bus mentioned in the electronic device mentioned above may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus. This communication bus can be divided into an address bus, a data bus, a control bus, etc. For ease of illustration, the figure shows only one thick line, but this does not mean that there is only one bus or only one type of bus.
[0300] The communication interface is used for communication between the above electronic device and other devices.
[0301] The memory may include random access memory (RAM) or non-volatile memory, such as at least one disk storage. Alternatively, the memory may be at least one storage device located away from the processor.
[0302] The above-mentioned processor can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, and discrete hardware components.
[0303] In another embodiment provided in the present application, a computer-readable storage medium is further provided, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the storyboard generation method described in any of the above embodiments is implemented.
[0304] By selecting the above embodiment, when the target script is obtained, the target script is input into the large language model, and the first scene description information output by the large language model can be obtained, wherein the first scene description information is used to describe the scene represented by the target script; the first scene description information is input into the Wensheng graph model to generate a target scene graph whose displayed content matches the first scene description information, and the target scene graph is stylized, and the image style of the target scene graph is converted into the storyboard style, so as to obtain the first storyboard. By applying the embodiment of the present application, the first storyboard corresponding to the target script can be quickly obtained, which saves the cost of professional storyboard artists, reduces the cost of generating storyboards, and improves the efficiency of film and television production. In addition, the embodiment of the present application will first generate a target scene graph that matches the first scene description information, and then convert the image style of the target scene graph into the storyboard style to obtain the first storyboard, so that the picture finally obtained can conform to the storyboard style, improve the accuracy of generating storyboards, and enhance the user experience.
[0305] In another embodiment provided by the present application, a computer program product comprising instructions is also provided, which, when executed on a computer, enables the computer to execute the storyboard generation method described in any one of the above embodiments.
[0306] By selecting the above embodiment, when the target script is obtained, the target script is input into the large language model, and the first scene description information output by the large language model can be obtained, wherein the first scene description information is used to describe the scene represented by the target script; the first scene description information is input into the Wensheng graph model to generate a target scene graph whose displayed content matches the first scene description information, and the target scene graph is stylized, and the image style of the target scene graph is converted into the storyboard style, so as to obtain the first storyboard. By applying the embodiment of the present application, the first storyboard corresponding to the target script can be quickly obtained, which saves the cost of professional storyboard artists, reduces the cost of generating storyboards, and improves the efficiency of film and television production. In addition, the embodiment of the present application will first generate a target scene graph that matches the first scene description information, and then convert the image style of the target scene graph into the storyboard style to obtain the first storyboard, so that the picture finally obtained can conform to the storyboard style, improve the accuracy of generating storyboards, and enhance the user experience.
[0307] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When software is used for implementation, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website, computer, server or data center to another website, computer, server or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) method. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more available media integrations. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive (SSD)).
[0308] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply the existence of any such actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or device comprising the element.
[0309] Each embodiment in this specification is described in a related manner. Similar portions between the various embodiments can be referenced to each other. Each embodiment focuses on the differences from other embodiments. In particular, the device, electronic device, and storage medium embodiments are generally similar to the method embodiments, so their descriptions are relatively simple. For related portions, reference can be made to the descriptions of the method embodiments.
[0310] The above description is only a preferred embodiment of the present application and is not intended to limit the scope of protection of the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application are included in the scope of protection of the present application.
Claims
1. A method for generating a storyboard, characterized in that: The method comprises: Get the target script; Inputting the target script into a large language model to obtain first scene description information output by the large language model; wherein the first scene description information is used to describe the scene represented by the target script; Inputting the first scene description information into a Wensheng graph model, generating a target scene graph whose displayed content matches the first scene description information, and performing stylization processing on the target scene graph to convert the image style of the target scene graph into a storyboard style, thereby obtaining a first storyboard; Obtain a first storyboard output by the Wensheng graph model.
2. The method according to claim 1, characterized in that After obtaining the Wensheng graph model to output the first storyboard, the method further includes: Acquire second scene description information, where the second scene description information is obtained after updating the lens angle and / or lens focal length in the first scene description information; Inputting the second scene description information and the first storyboard into a Wensheng graph model to obtain a second storyboard output by the Wensheng graph model; wherein the second storyboard displays the same scene as the first storyboard, and the lens angle and / or lens focal length of the second storyboard conform to the second scene description information; Obtain a second storyboard output by the Wensheng graph model.
3. The method according to claim 1, characterized in that After obtaining the target script, the method further includes: Determining era information of the target script according to the content of the target script; wherein the era information represents the target era described by the target script; Inputting the target script into the large language model to obtain first scene description information output by the large language model includes: The target script and the era information are input into the large language model, the third scene description information in the target script is extracted, and the third scene description information is modified to the first scene description information that conforms to the target era, and the first scene description information output by the large language model is obtained.
4. The method according to any one of claims 1 to 3, characterized in that Before inputting the first scene description information into the text graph model to generate a target scene graph whose displayed content matches the first scene description information, the method further includes: Adjusting the first scene description information by using a target adjustment method; The step of inputting the first scene description information into a text graph model to generate a target scene graph whose displayed content matches the first scene description information includes: The adjusted first scene description information is input into the Wensheng graph model to generate a target scene graph whose displayed content matches the adjusted first scene description information.
5. The method according to claim 1, characterized in that The method further comprises: Acquiring first information, where the first information includes multiple values of information of a preset type in the first scene description information, where the preset type of information is used to describe shooting information, where the shooting information is used to describe a relative angle and / or relative distance between a shooting device and a shooting subject; Inputting the target script into the large language model to obtain first scene description information output by the large language model includes: The first information and the target script are input into a large language model to select a target value from multiple values of the preset type of information based on the target script as the value of the preset type of information in the first scene description information.
6. A storyboard generating device, characterized in that: The device comprises: The first acquisition module is used to acquire the target script; A first input module is configured to input the target script into a large language model to obtain first scene description information output by the large language model; wherein the first scene description information is used to describe the scene represented by the target script; A second input module is configured to input the first scene description information into a Wensheng graph model, generate a target scene graph whose displayed content matches the first scene description information, and perform stylization processing on the target scene graph to convert the image style of the target scene graph into a storyboard style to obtain a first storyboard; The second acquisition module is used to acquire the first storyboard output by the Wensheng graph model.
7. The device according to claim 6, characterized in that The device further comprises: a third acquisition module, configured to acquire second scene description information, where the second scene description information is obtained after updating the lens angle and / or lens focal length in the first scene description information; a third input module, configured to input the second scene description information and the first storyboard into a Wensheng diagram model, and obtain a second storyboard output by the Wensheng diagram model; wherein the second storyboard displays the same scene as the first storyboard, and the lens angle and / or lens focal length of the second storyboard conform to the second scene description information; The fourth acquisition module is used to obtain the second storyboard output by the Wensheng graph model.
8. The device according to claim 6, characterized in that The device further comprises: A first determining module is configured to determine era information of the target script based on the content of the target script; wherein the era information indicates a target era described by the target script; The first input module is specifically used to: The target script and the era information are input into the large language model, the third scene description information in the target script is extracted, and the third scene description information is modified to the first scene description information that conforms to the target era, and the first scene description information output by the large language model is obtained.
9. An electronic device, characterized in that: It includes a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other via the communication bus; Memory for storing computer programs; A processor, configured to implement the method steps described in any one of claims 1 to 5 when executing a program stored in a memory.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method steps described in any one of claims 1 to 5 are implemented.