Video script generation method and device, equipment and medium
By integrating character design information, posture design information, and background keywords into a textual logic, the system generates storyboard description text corresponding to the content description text of each storyboard scene, and identifies the visual and auditory elements associated with the initial video script. This solves the problem of the lack of flexibility and diversity in existing video script generation methods, and achieves more efficient video script generation.
Patent Information
- Application Number
- CN202510809699.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-17
- Publication Date
- 2025-11-11
AI Technical Summary
Existing technologies lack flexibility and diversity in video script generation, making it impossible to personalize video scripts according to the design requirements of different video elements.
By integrating character design information, posture design information, and background keywords into a textual logic, storyboard description text corresponding to each storyboard content description text is generated. Visual and auditory elements associated with the initial video script are identified and aligned to generate the final video script.
It improves the flexibility and diversity of video script generation, and enables element design based on video script reference information.
Smart Images

Figure CN120935427A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of computer technology, and specifically relates to a method, apparatus, device and medium for generating video scripts. Background Technology
[0002] With the rapid development of short video platforms, film and television production, and the advertising industry, the demand for video scripts is increasing daily. A video script is a written plan prepared before video production, containing a detailed outline of the video content. It describes in detail every shot, frame, dialogue, action, background music, special effects, etc., guiding the production team in shooting and editing the video to ensure the final product meets expectations. Efficient video script generation technology has become a hot research area in video creation.
[0003] In existing technologies, video script generation mainly involves predefining a script template, filling the received keywords into the corresponding positions in the script structure, and then organizing the results according to the script format to generate the video script. However, due to the diversity of related elements in a video script and the mutual influence between elements, existing technologies can only generate scripts based on a single generation path. They cannot personalize the generation of video script elements based on the design requirements of different video elements, and their generation methods lack flexibility and diversity. Summary of the Invention
[0004] The purpose of this application is to provide a method, apparatus, device, and medium for generating video scripts, which solves the problem of lack of flexibility and diversity in existing video script generation methods. By integrating character design information, posture design information, and background keywords into text logic, storyboard description text corresponding to each storyboard content description text is generated. Visual and auditory elements associated with the initial video script are determined, and the storyboard content description text, storyboard description text, auditory elements, and visual elements are aligned to obtain the final video script. This achieves the goal of designing elements according to video script reference information, improving the flexibility and diversity of script generation.
[0005] In a first aspect, embodiments of this application provide a method for generating a video script, the method comprising: Obtain multiple scene description texts from the initial video script, and perform semantic parsing on each scene description text to obtain character keywords, action keywords, and background keywords; Determine the character design information associated with character keywords and the posture design information associated with action keywords. Integrate the character design information, posture design information and background keywords into text logic to generate storyboard description text corresponding to the content description text of each storyboard. Identify the visual and auditory elements associated with the initial video script, align the storyboard content description text, storyboard image description text, auditory elements, and visual elements to obtain the final video script.
[0006] Furthermore, identify the visual elements associated with the initial video script, including: The action keywords in the initial video script are grouped into action groups, and the action range corresponding to each group of actions in the action grouping results is determined. Determine the first shot range corresponding to the action range, match the first shot range with the second shot range of the preset shot type, and determine the target shot type for each group of actions based on the matching result; Obtain reference images for the initial video script, perform image analysis on the reference images to determine the visual style, and combine the target shot type and visual style to obtain the visual elements associated with the initial video script.
[0007] Furthermore, the first shot range corresponding to the range of motion is determined, including: Based on the range of motion, determine the limbs associated with each group of motions in the motion grouping results. Based on the mapping relationship between the preset shot range and the range of motion and the maximum extension of the limbs, determine the first shot range corresponding to the range of motion.
[0008] Furthermore, identify the auditory elements associated with the initial video script, including: Obtain the story type and video design duration from the initial video script, determine multiple candidate auditory elements corresponding to the story type in the preset auditory database, and determine the narrative rhythm corresponding to each candidate auditory element based on the music rhythm of each candidate auditory element. Identify the number of action keywords and determine the story rhythm of the initial video script based on the number of action keywords and the video design duration; Identify the target narrative rhythm that is the same as the story rhythm in the narrative rhythm, and determine the auditory elements associated with the initial video script based on the target candidate auditory elements corresponding to the target narrative rhythm.
[0009] Furthermore, identify the character design information associated with the character keywords, including: Identify character skills from character keywords and determine the skill types corresponding to those skills; Obtain character reference images corresponding to skill types from the preset character reference image library, extract features from the character reference images, and obtain character prop design information and character clothing design information; Multiple character personalities corresponding to the clothing design information are identified, and the facial parameters corresponding to each character personality are combined to obtain the character's facial feature design information.
[0010] Furthermore, determine the posture design information associated with the action keywords, including: Identify action adjectives and verbs in action keywords, determine the action speed corresponding to the action adjectives, and the action posture corresponding to the verbs; The instantaneous motion amplitude is determined based on the motion speed, and the motion posture is adjusted according to the instantaneous motion amplitude to obtain the posture design information.
[0011] Furthermore, before obtaining the description text of multiple scene contents from the initial video script, the method also includes: Obtain the video requirement text and narration text from the initial video script, and perform first-scene text splitting on the video requirement text and second-scene text splitting on the narration text according to the preset scene identifier. The first and second segment text splitting results are arranged and combined according to the segment sequence corresponding to the preset segment identifier to obtain multiple segment content description texts of the initial video script.
[0012] Secondly, embodiments of this application provide a video script generation apparatus, the apparatus comprising: The keyword acquisition module is used to acquire multiple scene description texts from the initial video script, and to perform semantic parsing on each scene description text to obtain character keywords, action keywords, and background keywords; The text integration module is used to determine the character design information associated with the character keywords and the posture design information associated with the action keywords. It performs text logic integration of the character design information, posture design information and background keywords to generate the storyboard description text corresponding to the content description text of each storyboard. The script generation module is used to determine the visual and auditory elements associated with the initial video script, and to align the storyboard content description text, storyboard scene description text, auditory elements, and visual elements to obtain the final video script.
[0013] Thirdly, embodiments of this application provide an electronic device including a processor, a memory, and a program or instructions stored in the memory and executable on the processor, wherein the program or instructions, when executed by the processor, implement the steps of the method described in the first aspect.
[0014] Fourthly, embodiments of this application provide a readable storage medium on which a program or instructions are stored, which, when executed by a processor, implement the steps of the method described in the first aspect.
[0015] Fifthly, embodiments of this application also provide a computer program product comprising a computer program stored in a computer-readable storage medium, wherein at least one processor of the device reads from the computer-readable storage medium and executes the computer program, causing the device to perform the method described in the first aspect.
[0016] In this embodiment, multiple scene description texts of the initial video script are obtained. Semantic parsing is performed on each scene description text to obtain character keywords, action keywords, and background keywords. Character design information associated with character keywords and posture design information associated with action keywords are determined. The character design information, posture design information, and background keywords are logically integrated to generate scene description texts corresponding to each scene description text. Visual and auditory elements associated with the initial video script are determined. The scene description texts, scene description texts, auditory elements, and visual elements are aligned to obtain the final video script. The above-described video script generation method solves the problem of lack of flexibility and diversity in existing video script generation methods. By integrating character design information, posture design information, and background keywords into text logic, it generates storyboard description text corresponding to each storyboard content description text, and determines the visual and auditory elements associated with the initial video script. By aligning the storyboard content description text, storyboard description text, auditory elements, and visual elements, the final video script is obtained. This achieves the goal of designing elements based on video script reference information, thus improving the flexibility and diversity of script generation. Attached Figure Description
[0017] Figure 1 This is a flowchart of a video script generation method provided in an embodiment of this application; Figure 2 This is a schematic diagram illustrating the process of aligning script elements provided in this application; Figure 3 This is a flowchart of determining visual elements provided in an embodiment of this application; Figure 4 This is a flowchart of determining auditory elements provided in an embodiment of this application; Figure 5 This is a structural block diagram of a video script generation device provided in an embodiment of this application; Figure 6 This is a structural block diagram of the electronic device provided in the embodiments of this application. Detailed Implementation
[0018] To make the objectives, technical solutions, and advantages of this application clearer, specific embodiments of this application are described in detail below with reference to the accompanying drawings. It is understood that the specific embodiments described herein are merely for explaining this application and not for limiting it. Furthermore, it should be noted that, for ease of description, only the parts relevant to this application are shown in the drawings, not all of them. Before discussing exemplary embodiments in more detail, it should be mentioned that some exemplary embodiments are described as processes or methods depicted as flowcharts. Although the flowcharts describe operations (or steps) as sequential processes, many of these operations can be performed in parallel, concurrently, or simultaneously. Furthermore, the order of the operations can be rearranged. The process can be terminated when its operation is completed, but may also have additional steps not included in the drawings. The process can correspond to a method, function, procedure, subroutine, subprogram, etc.
[0019] The technical solutions of the embodiments of this application will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application are within the scope of protection of this application.
[0020] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such use of data can be interchanged where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class and the number of objects is not limited; for example, a first object can be one or more. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.
[0021] Firstly, this solution can be used for scenarios involving automatic script generation before video production, especially when combining multiple video-related elements. By logically integrating character design information, posture design information, and background keywords, it generates storyboard description text corresponding to each storyboard content description text. It also identifies the visual and auditory elements associated with the initial video script. Aligning the storyboard content description text, storyboard description text, auditory elements, and visual elements yields the final video script. This allows for element design based on the video script reference information, improving the flexibility and diversity of script generation. Based on these usage scenarios, it is understandable that electronic devices can be the primary implementers of this solution.
[0022] The following description, in conjunction with the accompanying drawings, details a method, apparatus, device, and medium for generating video scripts provided in this application, through specific embodiments and application scenarios.
[0023] Figure 1 This is a flowchart illustrating a video script generation method provided in an embodiment of this application. Figure 1 As shown, the specific steps include the following: S101: Obtain multiple scene description texts from the initial video script, and perform semantic parsing on each scene description text to obtain character keywords, action keywords, and background keywords.
[0024] The initial video script can serve as the foundational reference material for generating the video. This includes, for example, the required parameters for video production and an overview of the corresponding story. The storyboard description text includes descriptions of the visual content, story content, and character actions in each shot. Character keywords are words used to describe the characteristics of characters in the video. Action keywords are words used to describe the characteristics and movement of characters in the video. Background keywords are words used to describe the visual features of the environment in which the characters are situated in each shot.
[0025] In one embodiment, an initial video script can be received, and text splitting can be performed based on the storyboard identifiers in the initial video script to obtain multiple storyboard content description texts. A semantic parsing tool is then used to perform semantic parsing on the storyboard story corresponding to each storyboard content description text to obtain character keywords, action keywords, and background keywords for each storyboard story.
[0026] In one embodiment, before obtaining multiple scene description texts of the initial video script, the method further includes: obtaining video requirement text and spoken text in the initial video script; performing a first scene text split on the video requirement text according to a preset scene identifier, and performing a second scene text split on the spoken text; and arranging and combining the first scene text splitting result and the second scene text splitting result according to the scene sequence corresponding to the preset scene identifier to obtain multiple scene description texts of the initial video script.
[0027] The video requirement text can be text describing the video production requirements and a brief summary of the video content. The spoken text can be text used for story narration or character commentary during video playback. The video requirement text and spoken text correspond to the same story content. Preset storyboard markers are pre-set tags that map a portion of text to a video scene. Each preset storyboard marker corresponds to a storyboard sequence, and each storyboard sequence represents the order in which the scenes appear in the video. First storyboard text splitting can be an operation that splits the video requirement text according to the preset storyboard markers. Second storyboard text splitting can be an operation that splits the spoken text according to the preset storyboard markers.
[0028] In one embodiment, the video requirement text and spoken text in the initial video script can be obtained. Preset scene markers in the video requirement text and the text range covered by each preset scene marker are identified. The text range covered by each preset scene marker in the video requirement text is then split according to the preset scene markers to obtain a first scene text splitting result. Similarly, preset scene markers in the spoken text and the spoken text content corresponding to each preset scene marker are identified. The spoken text content corresponding to each preset scene marker is then split according to the preset scene markers to obtain a second scene text splitting result. Since the scene text splitting results correspond to the preset scene markers, the first and second scene text splitting results can be arranged and combined according to the scene sequence corresponding to the preset scene markers. The first and second scene text splitting results of the same scene sequence are combined together as the scene content description text for that scene. The combined results of different scene sequences are then sorted according to the sequence order to obtain multiple scene content description texts for the initial video script.
[0029] This solution splits the video requirement text and spoken text in the initial video script into segmentation texts based on preset segmentation identifiers, and then arranges and combines the segmentation results according to the segmentation sequence to obtain multiple segmentation content description texts of the initial video script. It can split the segmentation content description texts while retaining the necessary segmentation content description texts and the segmentation order, thus improving the efficiency and accuracy of segmentation content description texts.
[0030] S102, determine the character design information associated with the character keywords and the posture design information associated with the action keywords, integrate the character design information, posture design information and background keywords with text logic, and generate the storyboard description text corresponding to the storyboard content description text.
[0031] The character design information can be derived by expanding or filling in character keywords to further refine the character's characteristics. Examples include hair color, clothing, facial features, and limb length. Posture design information can be derived by expanding or filling in action keywords to further describe the specific posture of the character in a particular frame of the storyboard. For example, a character running with their left leg forward and right leg back, feet shoulder-width apart. Text logic can represent the priority of the storyboard description text, the way relationships are expressed between content, and the order of descriptive phrases. For example, text logic might require the storyboard text to be described in the order of background environment, character information, and character posture, with the overall picture described first, then the details. Changes in information between the background environment, character information, and character posture should be described in a cause-and-effect order, reflecting the spatial relationship between character information and the background environment. Following this text logic, character design information, posture design information, and background keywords can be integrated to generate text that meets the narrative requirements of the video storyboard. Storyboard description text can also serve as reference text for generating keyframes in the video storyboard. The storyboard description text includes the content of the storyboard keyframes, the characteristics of the scenes, and the relationships between the elements in the scenes.
[0032] In one embodiment, character information associated with each character keyword can be searched online. This information is then categorized by facial features, clothing, and props. The most frequently occurring information in each category is used as the character's information for that category. All category-related information is combined to generate character design information associated with the character keyword. Posture design information associated with action keywords can be determined based on the relationship between action keywords and postures. Character design information, posture design information, and background keywords are then logically integrated to generate storyboard description text corresponding to each storyboard content description text.
[0033] In one embodiment, determining the character design information associated with the character keywords includes: identifying the character skills in the character keywords and determining the skill type corresponding to the character skills; obtaining the character reference image corresponding to the skill type in the preset character reference image library, extracting features from the character reference image to obtain character prop design information and character clothing design information; determining multiple character personalities corresponding to the clothing design information, and combining the facial parameters corresponding to each character personality to obtain character facial feature design information.
[0034] Character skills can be the functions a character can perform in the video. Skill types can include defensive, offensive, and healing skills, etc. A preset character reference image library is a pre-set database used to provide reference for character design information in the video script. This library includes character images from multiple videos related to the video story background of the initial video script. Character reference images can be character images from the preset library that have the same or similar skill types. Character prop design information can be information describing the props used by the character, generated based on the prop features in the character reference images. Character clothing design information can be information describing the clothing worn by the character, generated based on the clothing features in the character reference images. Facial parameters can be facial feature parameters and proportions. Character facial feature design information can be information describing the proportions of the character's facial features as ultimately presented in the video.
[0035] In one embodiment, character keywords include character skills. Character skills within the keywords can be identified, and the character's skill type can be determined based on the correspondence between character skills and skill types. Reference images of characters with the same skill type are searched in a preset character reference image library. Feature extraction is performed on the reference images using image recognition, including the prop features and clothing features of the props used by the characters in the reference images. The most frequently occurring prop features are used as character prop design information, and the most frequently occurring clothing features are used as character clothing design information. Based on the correspondence between clothing features and character information, multiple character personalities corresponding to the clothing design information are determined. Based on the correspondence between character personalities and facial parameters, the facial parameters corresponding to each character personality are determined. These facial parameters are combined according to the facial feature structure to obtain the character's facial feature design information. The character prop design information, character clothing design information, and character facial feature design information are collectively used as the character design information.
[0036] This solution extracts features from character reference images to obtain character prop design information and character clothing design information. Based on the character personality corresponding to the clothing design information, it determines the character facial feature design information, thereby achieving the goal of character design based on existing character images and improving the rationality of character design.
[0037] In one embodiment, determining the posture design information associated with the action keyword includes: identifying the action adjective and verb in the action keyword, determining the action speed corresponding to the action adjective, and the action posture corresponding to the verb; determining the instantaneous action amplitude based on the action speed, and adjusting the action posture according to the instantaneous action amplitude to obtain the posture design information.
[0038] Among these, action adjectives can be words that describe the characteristics or manner of an action. Examples include words describing speed, intensity, state, and emotional tone. Verbs can be words that describe the meaning of the action itself. For example, if the key action is "walking quickly," then the action adjective is "quickly," and the verb is "walk." Instantaneous action amplitude can be the degree of physical or visual change exhibited by an action within a very short period of time.
[0039] In one embodiment, the action adjectives and verbs in the action keywords can be determined by identifying the part of speech of each word in the action keywords. The corresponding action speed can be determined based on the correspondence between the action adjectives and action speeds. Since a continuous action can be broken down into multiple instantaneous action postures, the action corresponding to a verb can be determined, and the corresponding action posture can be determined based on the correspondence between the action and the instantaneous action posture. Because the faster the action speed, the faster the instantaneous action posture changes, and the larger the corresponding instantaneous action amplitude, the instantaneous action amplitude can be determined based on the action speed. Each action posture corresponding to a verb can then be adjusted according to the instantaneous action amplitude to obtain posture design information.
[0040] This solution identifies action adjectives and verbs in action keywords to determine action speed and posture, determines instantaneous action amplitude based on action speed, and adjusts action posture according to instantaneous action amplitude to obtain posture design information, which can improve the accuracy of posture design.
[0041] S103, determine the visual and auditory elements associated with the initial video script, align the storyboard content description text, storyboard scene description text, auditory elements, and visual elements to obtain the final video script.
[0042] Visual elements can be information that affects the visual effect of the video, such as camera angle, focal length, lighting, color, and composition. Auditory elements can be information that affects the audio quality of the video, such as the timbre and rhythm of the narration and the style of the song. Finally, the video script can be the text that describes the entire content of the video.
[0043] In one embodiment, visual and auditory elements associated with the initial video script can be determined based on visual and auditory requirement information in the initial script. The music duration in the auditory elements can be determined based on the total duration of all spoken parts corresponding to the storyboard content description text. The playback duration of the video storyboard corresponding to the storyboard description text can be adjusted based on the duration of each spoken part. The visual elements are then superimposed with all video storyboards to obtain the final video script.
[0044] Figure 2 This is a schematic diagram illustrating the process of aligning script elements provided in this application.
[0045] like Figure 2 As shown, multiple video storyboard segments, such as Segment 1, Segment 2, and Segment 3, can be generated based on the storyboard narrative, scene description text, and visual elements in the storyboard content description text. Each video storyboard segment has an original storyboard design playback duration in the initial video script. The storyboard content description text includes voice-over content for storyboard segments, such as Voice-over 1, Voice-over 2, and Voice-over 3, with each voice-over content corresponding to a specific voice-over duration. The storyboard design playback duration is adjusted according to the voice-over duration. The video storyboard segments with adjusted playback durations are then spliced together in storyboard order. Voice-over timbres and background music corresponding to the voice-over content are generated based on auditory elements. The background music duration is extracted based on the total duration of the spliced video storyboard segments. The voice-over timbres are used to dub the voice-over content of the video storyboard segments, and the extracted background music is used to add background music to the spliced video storyboard segments, generating the final video script.
[0046] The technical solution provided in this application involves obtaining multiple scene description texts of an initial video script, performing semantic parsing on each scene description text to obtain character keywords, action keywords, and background keywords; determining character design information associated with character keywords and posture design information associated with action keywords; integrating the character design information, posture design information, and background keywords using text logic to generate scene description texts corresponding to each scene description text; determining visual and auditory elements associated with the initial video script, and aligning the scene description texts, scene description texts, auditory elements, and visual elements to obtain the final video script. The above-described video script generation method solves the problem of lack of flexibility and diversity in existing video script generation methods. By integrating character design information, posture design information, and background keywords into text logic, it generates storyboard description text corresponding to each storyboard content description text, and determines the visual and auditory elements associated with the initial video script. By aligning the storyboard content description text, storyboard description text, auditory elements, and visual elements, the final video script is obtained. This achieves the goal of designing elements based on video script reference information, thus improving the flexibility and diversity of script generation.
[0047] Figure 3 This is a flowchart illustrating the determination of visual elements provided in an embodiment of this application. For example... Figure 3 As shown, the specific steps include the following: S301, group the action keywords in the initial video script into action groups, and determine the action range corresponding to each group of actions in the action grouping results.
[0048] The action grouping result can be the action keyword grouping obtained by dividing multiple action keywords associated with the same action into the same group. The action range can be the maximum range of motion corresponding to each group of actions. In this scheme, the action range can be the limb range associated with each group of actions of the character.
[0049] In one embodiment, the action corresponding to each action keyword in the initial video script can be identified, and the action keywords can be grouped according to the same action. The action range corresponding to each group of actions in the action grouping result can be determined according to the limbs associated with the action and the characteristics of limb activities. For example, if a certain group of actions in the action grouping result is running, and the limbs associated with running are the arms and legs, then the action range corresponding to this group of actions is the character's arms and legs.
[0050] S302, determine the first shot range corresponding to the action range, match the first shot range with the second shot range of the preset shot type, and determine the target shot type for each group of actions based on the matching result.
[0051] The first shot range can be the minimum shot range required for the character's movement. For example, if the character's movement is a hand movement, the minimum shot range required is the area that can show the entire hand. Preset shot types can include: close-up, long shot, and extreme close-up. The second shot range can be the frame area corresponding to each preset shot type.
[0052] In one embodiment, a first camera range corresponding to the action range can be determined based on the correspondence between the action range and the camera range. The first camera range is then matched with the second camera ranges corresponding to multiple preset camera types to determine a target second camera range that is greater than or equal to the first camera range. The preset camera type corresponding to the target second camera range is then used as the target camera type for that group of actions. All action groups are then traversed to obtain the target camera type for each group of actions.
[0053] In one embodiment, determining the first shot range corresponding to the action range includes: determining the limbs associated with each group of actions in the action grouping results based on the action range, and determining the first shot range corresponding to the action range based on the mapping relationship between the preset shot range and the action range and the maximum extension of the limbs.
[0054] Among them, the maximum extension of the moving limb can be the maximum stretch length of the moving limb.
[0055] In one embodiment, the associated limbs in the action grouping results can be determined based on the action range, and the maximum action range corresponding to each action group can be determined based on the maximum extension and designed length of the limbs. The designed length of the limbs can be obtained by reading the character design information. The first shot range corresponding to the action range is determined based on the mapping relationship between the preset shot range and the action range and the maximum action range.
[0056] This solution identifies the limbs associated with each group of actions in the action grouping results, and determines the first shot range corresponding to the action range based on the mapping relationship between the preset shot range and the action range and the maximum extension of the limbs. This ensures that the shot range can cover all actions and improves the accuracy of determining the first shot range.
[0057] S303, obtain reference images for the initial video script, perform image analysis on the reference images to determine the visual style, combine the target shot type and visual style to obtain visual elements associated with the initial video script.
[0058] Reference images can be those used in the initial video script to adjust the visual style of video frames. Visual style can represent the overall atmosphere and emotional tone of the reference image. It can be conveyed through parameters such as color, brightness, and lighting of the reference image.
[0059] In one embodiment, reference images of the initial video script can be obtained, and the reference images can be analyzed to determine the visual style by analyzing image parameters such as color, brightness, and light. The visual style can be applied to the target shot type of each group of actions to obtain visual elements associated with the initial video script.
[0060] The technical solution provided in this application determines the first shot range by grouping the action keywords in the initial video script, matching the first shot range with the second shot range of the preset shot type to determine the target shot type, and combining the target shot type with the visual style of the reference image to obtain the visual elements associated with the initial video script. This achieves the goal of determining visual elements based on both shot type and visual style, improving the comprehensiveness and accuracy of visual element determination.
[0061] Figure 4 This is a flowchart illustrating the determination of auditory elements provided in an embodiment of this application. For example... Figure 4 As shown, the specific steps include the following: S401, obtain the story type and video design duration from the initial video script, determine multiple candidate auditory elements corresponding to the story type in the preset auditory database, and determine the narrative rhythm corresponding to each candidate auditory element based on the music rhythm of each candidate auditory element.
[0062] The story type can be the narrative pattern or framework to which the story belongs in terms of structure, theme, or plot. Story types can be obtained by classifying stories according to their historical period, theme, emotional tone, and target audience. The video design duration can be the pre-set video playback duration in the initial video script. The preset auditory database can be a pre-set database containing timbre-related data and music-related data. Candidate auditory elements can be auditory elements in the preset auditory database that correspond to the story type. For example, if the story type is historical science popularization, then the candidate auditory elements can be narrative timbre and soothing background music. Candidate auditory elements can contain both candidate timbre data and candidate music data. The narrative rhythm can be the narration speed when using candidate auditory elements to explain the story.
[0063] In one embodiment, the story type in the initial video script can be obtained based on the story theme and plot, and the video design duration in the initial video script can be read. Multiple candidate auditory elements corresponding to the story type are determined from a preset auditory database based on the correspondence between auditory elements and story types. The narrative rhythm corresponding to each candidate auditory element is determined based on the musical rhythm of the candidate music data in each candidate auditory element and the narrative rhythm corresponding to the musical rhythm.
[0064] S402, Identify the number of action keywords, and determine the story rhythm of the initial video script based on the number of action keywords and the video design duration.
[0065] The story rhythm can refer to the speed at which the story develops and the plot transitions.
[0066] In one embodiment, the number of action keywords in the initial video script can be identified based on part-of-speech tags, and the average number of actions per unit time can be determined based on the number of action keywords and the video design duration. The average number of actions can then be used as the story rhythm of the initial video script.
[0067] S403, determine the target narrative rhythm that is the same as the story rhythm in the narrative rhythm, and determine the auditory elements associated with the initial video script based on the target candidate auditory elements corresponding to the target narrative rhythm.
[0068] In one embodiment, the story rhythm corresponding to the narrative rhythm can be determined based on the correspondence between the narrative rhythm and the story development speed. The story rhythm of the initial video script is compared with the story rhythm corresponding to the narrative rhythm to determine the target narrative rhythm that is the same as the story rhythm. The target candidate auditory element corresponding to the target narrative rhythm is determined as the auditory element associated with the initial video script.
[0069] The technical solution provided in this application determines multiple candidate auditory elements and the narrative rhythm corresponding to each candidate auditory element by the story type of the initial video script, determines the story rhythm based on the number of action keywords and the video design duration, determines the target narrative rhythm that is the same as the story rhythm and the target candidate auditory elements corresponding to the target narrative rhythm, which can achieve the purpose of determining auditory elements based on the story rhythm and the narrative rhythm, and improve the fit between the auditory element determination result and the video script.
[0070] Figure 5 This is a structural block diagram of a video script generation device provided in an embodiment of this application. Figure 5 As shown, it specifically includes the following: The keyword acquisition module 501 is used to acquire multiple scene description texts of the initial video script, and to perform semantic parsing on each scene description text to obtain character keywords, action keywords and background keywords; The text integration module 502 is used to determine the character design information associated with the character keywords and the posture design information associated with the action keywords. It performs text logic integration of the character design information, posture design information and background keywords to generate the storyboard description text corresponding to the storyboard content description text. The script generation module 503 is used to determine the visual and auditory elements associated with the initial video script, and to align the storyboard content description text, storyboard scene description text, auditory elements, and visual elements to obtain the final video script.
[0071] Furthermore, the script generation module 503 is specifically used for: The action keywords in the initial video script are grouped into action groups, and the action range corresponding to each group of actions in the action grouping results is determined. Determine the first shot range corresponding to the action range, match the first shot range with the second shot range of the preset shot type, and determine the target shot type for each group of actions based on the matching result; Obtain reference images for the initial video script, perform image analysis on the reference images to determine the visual style, and combine the target shot type and visual style to obtain the visual elements associated with the initial video script.
[0072] Furthermore, the script generation module 503 is specifically used for: Based on the range of motion, determine the limbs associated with each group of motions in the motion grouping results. Based on the mapping relationship between the preset shot range and the range of motion and the maximum extension of the limbs, determine the first shot range corresponding to the range of motion.
[0073] Furthermore, the script generation module 503 is specifically used for: Obtain the story type and video design duration from the initial video script, determine multiple candidate auditory elements corresponding to the story type in the preset auditory database, and determine the narrative rhythm corresponding to each candidate auditory element based on the music rhythm of each candidate auditory element. Identify the number of action keywords and determine the story rhythm of the initial video script based on the number of action keywords and the video design duration; Identify the target narrative rhythm that is the same as the story rhythm in the narrative rhythm, and determine the auditory elements associated with the initial video script based on the target candidate auditory elements corresponding to the target narrative rhythm.
[0074] Furthermore, the text integration module 502 is specifically used for: Identify character skills from character keywords and determine the skill types corresponding to those skills; Obtain character reference images corresponding to skill types from the preset character reference image library, extract features from the character reference images, and obtain character prop design information and character clothing design information; Multiple character personalities corresponding to the clothing design information are identified, and the facial parameters corresponding to each character personality are combined to obtain the character's facial feature design information.
[0075] Furthermore, the text integration module 502 is specifically used for: Identify action adjectives and verbs in action keywords, determine the action speed corresponding to the action adjectives, and the action posture corresponding to the verbs; The instantaneous motion amplitude is determined based on the motion speed, and the motion posture is adjusted according to the instantaneous motion amplitude to obtain the posture design information.
[0076] Furthermore, the device also includes: The text splitting module is used to obtain the video requirement text and the voice-over text in the initial video script, perform first-scene text splitting on the video requirement text according to the preset scene identifier, and perform second-scene text splitting on the voice-over text. The storyboard text generation module is used to arrange and combine the first storyboard text splitting result and the second storyboard text splitting result according to the storyboard sequence corresponding to the preset storyboard identifier, so as to obtain multiple storyboard content description texts of the initial video script.
[0077] The technical solution provided in this application includes a keyword acquisition module, which acquires multiple scene description texts of the initial video script and performs semantic parsing on each scene description text to obtain character keywords, action keywords, and background keywords; a text integration module, which determines the character design information associated with the character keywords and the posture design information associated with the action keywords, and performs text logic integration on the character design information, posture design information, and background keywords to generate scene description texts corresponding to each scene description text; and a script generation module, which determines the visual and auditory elements associated with the initial video script, aligns the scene description texts, scene description texts, auditory elements, and visual elements to obtain the final video script. The aforementioned video script generation device solves the problem of insufficient flexibility and diversity in existing video script generation methods. By integrating character design information, posture design information, and background keywords into text logic, it generates storyboard description text corresponding to each storyboard content description text. It also determines the visual and auditory elements associated with the initial video script. By aligning the storyboard content description text, storyboard description text, auditory elements, and visual elements, it obtains the final video script. This achieves the goal of designing elements based on video script reference information, thus improving the flexibility and diversity of script generation.
[0078] The video script generation device in this application embodiment can be configured in a device, or in a component, integrated circuit, or chip in a terminal. The device can be a mobile electronic device or a non-mobile electronic device. For example, mobile electronic devices can be mobile phones, tablets, laptops, PDAs, in-vehicle electronic devices, wearable devices, ultra-mobile personal computers (UMPCs), netbooks, or personal digital assistants (PDAs), etc., while non-mobile electronic devices can be servers, network-attached storage (NAS), personal computers (PCs), televisions (TVs), ATMs, or self-service machines, etc. This application embodiment does not impose specific limitations.
[0079] The video script generation device in this application embodiment can be an operating system. The operating system can be Android, iOS, or other possible operating systems; this application embodiment does not specifically limit the specific operating system.
[0080] The video script generation apparatus provided in this application embodiment can implement the various processes implemented in the above method embodiments, and will not be described again here to avoid repetition.
[0081] like Figure 6 As shown, this application embodiment also provides an electronic device 600, including a processor 601, a memory 602, and a program or instructions stored in the memory 602 and executable on the processor 601. When the program or instructions are executed by the processor 601, they implement the various processes of the above-described video script generation method embodiment and achieve the same technical effect. To avoid repetition, they will not be described again here.
[0082] It should be noted that the electronic devices in the embodiments of this application include the mobile electronic devices and non-mobile electronic devices described above.
[0083] This application also provides a readable storage medium storing a program or instructions. When the program or instructions are executed by a processor, they implement the various processes of the above-described video script generation method embodiment and achieve the same technical effect. To avoid repetition, they will not be described again here.
[0084] The processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.
[0085] This application also provides a program product including program code. When the program product is run on a computer device, the program code causes the computer device to perform the steps of the methods described above according to various exemplary embodiments of this application. For example, the computer device can execute a video script generation method described in an embodiment of this application. The program product can be implemented using any combination of one or more readable media.
[0086] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of this application is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.
[0087] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a computer software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of this application.
[0088] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.
[0089] The above description is merely a preferred embodiment and the technical principles employed in this application. This application is not limited to the specific embodiments described herein, and various obvious changes, readjustments, and substitutions that can be made by those skilled in the art will not depart from the scope of protection of this application. Therefore, although this application has been described in detail through the above embodiments, this application is not limited to the above embodiments, and may include more other equivalent embodiments without departing from the concept of this application, the scope of which is determined by the scope of the claims.
Claims
1. A video script generation method, characterized in that, The method includes: Obtain multiple scene description texts from the initial video script, and perform semantic parsing on each scene description text to obtain character keywords, action keywords, and background keywords; Determine the character design information associated with the character keywords and the posture design information associated with the action keywords. Perform text logic integration on the character design information, the posture design information and the background keywords to generate storyboard description text corresponding to each storyboard content description text. The visual and auditory elements associated with the initial video script are determined, and the storyboard content description text, the storyboard scene description text, the auditory elements, and the visual elements are aligned to obtain the final video script.
2. The video script generation method according to claim 1, characterized in that, The determination of the visual elements associated with the initial video script includes: The action keywords in the initial video script are grouped into action groups, and the action range corresponding to each group of actions in the action grouping results is determined. Determine the first shot range corresponding to the action range, match the first shot range with the second shot range of the preset shot type, and determine the target shot type for each group of actions based on the matching result; Obtain reference images for the initial video script, perform image analysis on the reference images to determine the visual style, and combine the target shot type and the visual style to obtain visual elements associated with the initial video script.
3. The video script generation method according to claim 2, characterized in that, Determining the first camera range corresponding to the range of motion includes: Based on the range of motion, determine the limbs associated with each group of motions in the motion grouping results, and determine the first shot range corresponding to the range of motion based on the mapping relationship between the preset shot range and the range of motion and the maximum extension of the limbs.
4. The video script generation method according to claim 1, characterized in that, Determining the auditory elements associated with the initial video script includes: Obtain the story type and video design duration from the initial video script, determine multiple candidate auditory elements corresponding to the story type in the preset auditory database, and determine the narrative rhythm corresponding to each candidate auditory element based on the musical rhythm of each candidate auditory element. Identify the number of action keywords, and determine the story rhythm of the initial video script based on the number of action keywords and the video design duration; Identify the target narrative rhythm that is the same as the story rhythm in the narrative rhythm, and determine the auditory elements associated with the initial video script based on the target candidate auditory elements corresponding to the target narrative rhythm.
5. The video script generation method according to claim 1, characterized in that, The determination of the character design information associated with the character keywords includes: Identify the character skills in the character keywords and determine the skill type corresponding to the character skills; Obtain character reference images corresponding to the skill type from a preset character reference image library, extract features from the character reference images, and obtain character prop design information and character costume design information; Multiple character personalities corresponding to the clothing design information are determined, and the facial parameters corresponding to each character personality are combined to obtain the character's facial feature design information.
6. The video script generation method according to claim 1, characterized in that, Determine the posture design information associated with the action keyword, including: Identify the action adjectives and verbs in the action keywords, determine the action speed corresponding to the action adjectives, and the action posture corresponding to the verbs; The instantaneous motion amplitude is determined based on the motion speed, and the motion posture is adjusted according to the instantaneous motion amplitude to obtain posture design information.
7. The video script generation method according to claim 1, characterized in that, Before obtaining the multiple scene description texts of the initial video script, the method further includes: Obtain the video requirement text and the voiceover text from the initial video script; perform a first segmentation text split on the video requirement text according to the preset segmentation identifier; and perform a second segmentation text split on the voiceover text. The first and second storyboard text splitting results are arranged and combined according to the storyboard sequence corresponding to the preset storyboard identifier to obtain multiple storyboard content description texts of the initial video script.
8. A video script generation device, characterized in that, The device includes: The keyword acquisition module is used to acquire multiple scene content description texts of the initial video script, and to perform semantic parsing on each scene content description text to obtain character keywords, action keywords and background keywords; The text integration module is used to determine the character design information associated with the character keywords and the posture design information associated with the action keywords, and to perform text logic integration on the character design information, the posture design information and the background keywords to generate a storyboard description text corresponding to each storyboard content description text; The script generation module is used to determine the visual and auditory elements associated with the initial video script, and to align the storyboard content description text, the storyboard scene description text, the auditory elements, and the visual elements to obtain the final video script.
9. An electronic device, characterized in that, It includes a processor, a memory, and a program or instructions stored in the memory and executable on the processor, wherein the program or instructions, when executed by the processor, implement the steps of a video script generation method as described in any one of claims 1-7.
10. A readable storage medium, characterized in that, The readable storage medium stores a program or instructions that, when executed by a processor, implement the steps of a video script generation method as described in any one of claims 1-7.