A video script generation method and related apparatus

By acquiring multi-frame images and character feature information from video clips and using a generative model to analyze the video script, the problem of time-consuming and inaccurate character recognition in existing technologies is solved, achieving efficient and accurate script generation.

CN122372812APending Publication Date: 2026-07-10TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
TENCENT TECHNOLOGY (SHENZHEN) CO LTD
Filing Date
2026-05-06
Publication Date
2026-07-10

AI Technical Summary

Technical Problem

Existing video script generation methods require the identification of characters in each frame of a video clip in advance, which results in a significant time consumption and a poor match between characters and dialogue, reducing the efficiency and accuracy of script generation.

Method used

By acquiring multiple frames of images of the target video clip and corresponding character feature information, the target generation model analyzes the scene, characters, and dialogue of each frame under the guidance of the first prompt text, reducing character recognition time and improving analysis accuracy through thought chain, and outputting script text that conforms to the script format.

Benefits of technology

It enables accurate representation of target video segments in a standard script format, improving script generation efficiency and accuracy, reducing character recognition time, and enhancing the accuracy of script generation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122372812A_ABST
    Figure CN122372812A_ABST
Patent Text Reader

Abstract

The application discloses a video script generation method and related device. A target video segment formed by multiple target video images is obtained, and target character feature information corresponding to the target video segment is obtained; the target character feature information and the target video segment are input into a target generation model, scenes, characters and dialogues in each target video image are analyzed based on the target character feature information under the guidance of a first prompt text, character recognition is realized without traversing each target video image in advance, character recognition is realized when the model generates a script, character recognition time is reduced, errors are analyzed and the script is integrated, the analysis accuracy of characters and dialogues is improved through a thinking chain, the video segment is accurately understood, a script text conforming to a script format is output, a video generation script of the target video segment is obtained, the video generation script accurately represents the target video segment under the standard script format, and therefore the script generation efficiency and script generation accuracy are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image-to-text technology, and in particular to a video script generation method and related apparatus. Background Technology

[0002] Currently, many videos fail to save the corresponding scripts. In order to restore the scripts corresponding to the videos, we can use a visual language model to understand the videos and output the scripts according to the script format based on the recognition of scenes, characters and dialogues in the videos.

[0003] In related technologies, methods for generating scripts from videos typically involve identifying characters in video clips in advance, inputting the video clips into a visual language model to generate a script, and outputting the script of the video clips according to the script format.

[0004] However, the above method requires identifying characters in video clips in advance, which involves traversing every frame of the video clip to identify characters. This not only consumes a lot of time, but may also lead to inaccurate understanding of the video clips due to insufficient matching between characters and dialogues, thereby reducing the efficiency and accuracy of script generation. Summary of the Invention

[0005] To address the aforementioned technical problems, this application provides a video script generation method and related apparatus, which improves script generation efficiency and accuracy.

[0006] The embodiments of this application disclose the following technical solutions:

[0007] On one hand, embodiments of this application provide a video script generation method, the method comprising:

[0008] Acquire the target video segment and the target character feature information corresponding to the target video segment; the target video segment includes multiple frames of target video images;

[0009] Using a target generation model, a script is generated for the target video segment based on the target character feature information and the first prompt text, resulting in a video-generated script for the target video segment. The first prompt text guides the target generation model to analyze the scene, characters, and dialogue in each frame of the target video image based on the target character feature information, and to analyze errors and integrate the script to output script text that conforms to the script format.

[0010] On the other hand, embodiments of this application provide a video script generation apparatus, the apparatus comprising: an acquisition unit and a generation unit;

[0011] The acquisition unit is used to acquire a target video segment and target character feature information corresponding to the target video segment; the target video segment includes multiple frames of target video images;

[0012] The generation unit is used to generate a script for the target video segment by using a target generation model based on the target character feature information and the first prompt text, thereby obtaining a video-generated script for the target video segment; the first prompt text is used to guide the target generation model to analyze the scene, characters and dialogue in each frame of the target video image based on the target character feature information, and to analyze errors and integrate the script to output script text that conforms to the script format.

[0013] On the other hand, embodiments of this application provide a computer device, the computer device including a processor and a memory:

[0014] The memory is used to store computer programs and to transfer the computer programs to the processor;

[0015] The processor is configured to execute the method described in any of the foregoing aspects according to instructions in the computer program.

[0016] On the other hand, embodiments of this application provide a computer-readable storage medium for storing a computer program that, when run on a computer device, causes the computer device to perform the methods described in any of the foregoing aspects.

[0017] On the other hand, embodiments of this application provide a computer program product, including a computer program that, when run on a computer device, causes the computer device to perform the method described in any of the foregoing aspects.

[0018] As can be seen from the above technical solution, a target video segment is formed by acquiring multiple frames of target video images, and the target character feature information corresponding to the target video segment is obtained. This target character feature information can be provided to the model to identify the characters in each frame of the target video image when generating the video script. The target character feature information and the target video segment are input into the target generation model. Guided by the first prompt text, the model analyzes the scene, characters, and dialogue in each frame of the target video image based on the target character feature information. It is not necessary to traverse each frame of the target video image in advance to achieve character recognition. Instead, character recognition is achieved when the model generates the script, reducing the character recognition time. Furthermore, it analyzes errors and integrates the script, improving the accuracy of character and dialogue analysis through thought chain, achieving accurate understanding of the video segment, and outputting script text that conforms to the script format. This results in a video-generated script for the target video segment, ensuring that the video-generated script accurately represents the target video segment in a standard script format, thereby improving script generation efficiency and accuracy. Attached Figure Description

[0019] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0020] Figure 1 A schematic diagram of a computer system for a video script generation method provided in an embodiment of this application;

[0021] Figure 2 A schematic diagram illustrating an application scenario of a video script generation method provided in this application embodiment;

[0022] Figure 3 A flowchart illustrating a video script generation method provided in this application embodiment;

[0023] Figure 4 A schematic diagram illustrating a target video segment and the target character feature information corresponding to the target video segment, provided in an embodiment of this application;

[0024] Figure 5 A schematic diagram of a first prompt text provided for an embodiment of this application;

[0025] Figure 6 A schematic diagram illustrating a video generation script for a target video segment provided in an embodiment of this application;

[0026] Figure 7 A schematic diagram illustrating the generation of a script from a video using a target generation model, as provided in an embodiment of this application;

[0027] Figure 8 A schematic diagram illustrating an image analysis output text, an error analysis output text, and a script integration output text provided in an embodiment of this application;

[0028] Figure 9 A schematic diagram illustrating how to train an initial generative model to obtain a target generative model, as provided in an embodiment of this application;

[0029] Figure 10 A flowchart illustrating steps for obtaining a first video segment based on a first sample video, provided in an embodiment of this application;

[0030] Figure 11 A schematic diagram illustrating the optimized target generation model obtained from an optimized target generation model provided in an embodiment of this application;

[0031] Figure 12 This is a schematic diagram illustrating the generation of a script from a video using an optimized target generation model, as provided in an embodiment of this application.

[0032] Figure 13 A schematic diagram of an application product interface in a video script generation scenario provided in an embodiment of this application;

[0033] Figure 14 A structural diagram of a video script generation apparatus provided in an embodiment of this application;

[0034] Figure 15 A structural diagram of a server provided in an embodiment of this application;

[0035] Figure 16 This is a structural diagram of a terminal provided in an embodiment of this application. Detailed Implementation

[0036] The embodiments of this application will now be described with reference to the accompanying drawings.

[0037] Currently, it is necessary to reconstruct the corresponding script from videos that have not saved the script. Specifically, after acquiring video clips, the characters in the video clips are identified in advance, the video clips are input into a visual language model to generate a script, and the script of the video clips is output according to the script format. However, research has found that identifying the characters in the video clips in advance requires traversing every frame of the video clip to perform character recognition, which is not only time-consuming, but may also lead to inaccurate understanding of the video clips due to insufficient matching between characters and dialogue, thereby reducing the efficiency and accuracy of script generation.

[0038] This application provides a video script generation method. It acquires a target video segment formed by multiple frames of target video images and obtains target character feature information corresponding to the target video segment. This target character feature information can be provided to a model to identify characters in each frame of the target video image during video script generation. The target character feature information and the target video segment are input into the target generation model. Guided by a first prompt text, the model analyzes the scene, characters, and dialogue in each frame of the target video image based on the target character feature information. This eliminates the need to pre-traverse each frame of the target video image for character recognition; instead, character recognition is achieved during script generation, reducing character recognition time. Furthermore, the method analyzes errors and integrates the script, improving the accuracy of character and dialogue analysis through thought processes. This leads to an accurate understanding of the video segment, outputting script text that conforms to the script format, resulting in a video-generated script for the target video segment. This ensures that the video-generated script accurately represents the target video segment in a standard script format, thereby improving script generation efficiency and accuracy.

[0039] To facilitate understanding of the video script generation method provided in this application embodiment, the computer system of the video script generation method will be described first below.

[0040] See Figure 1This figure is a schematic diagram of a computer system for a video script generation method provided in an embodiment of this application. The computer system 100 includes multiple devices, such as multiple terminals 1600 and multiple servers 1500. The terminals 1600 and the servers 1500 can communicate with each other through a communication network, and the servers 1500 obtain the data they need from the database 1700.

[0041] The communication network uses standard communication technologies and / or protocols, typically the Internet, but can also be any network, including but not limited to Bluetooth, local area network (LAN), metropolitan area network (MAN), wide area network (WAN), mobile, private network, or any combination of virtual private network. In some embodiments, custom or dedicated data communication technologies may be used to replace or supplement the aforementioned data communication technologies.

[0042] The terminal can be an electronic device such as a smartphone, wearable device, personal computer (PC), intelligent voice interaction device, smart home appliance, vehicle terminal, aircraft, unmanned vending terminal, extended reality (XR) device, etc. XR devices can include virtual reality (VR) devices, augmented reality (AR) devices, and mixed reality (MR) devices. A client application for the target application can be installed and run on the terminal. This target application can be an application that supports video script generation, or other applications that support image-to-text processing, etc., and this application does not limit its scope. Furthermore, this application does not limit the form of the target application, including but not limited to applications (Apps), mini-programs, etc., installed on the terminal, and can also be in the form of a webpage.

[0043] A server can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services such as cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and basic cloud computing services such as big data. A server can also be a backend server for the target application, providing backend services to the client of the target application, such as an image-to-text service.

[0044] To facilitate understanding of the video script generation method provided in this application embodiment, the following uses a server as the execution subject of the video script generation method as an example to illustrate the application scenarios of the video script generation method.

[0045] See Figure 2 This figure is a schematic diagram illustrating an application scenario of a video script generation method provided in an embodiment of this application. Figure 2 In this application scenario, the terminal 1600, server 1500, and database 1700 are included.

[0046] The aforementioned target application, used for generating scripts for videos, is installed on terminal 1600. Server 1500 provides support for the operation of the target application. Database 1700 stores the data required by server 1500. Details are provided below using A1-A7.

[0047] In step A1, in response to the input operation on the video script generation interface for the target video segment and the target character feature information corresponding to the target video segment, the terminal 1600 obtains the target video segment and the target character feature information corresponding to the target video segment, wherein the target video segment includes multiple frames of target video images.

[0048] For example, the target video segment includes video image T1, video image T2, ..., video image T8. The target character feature information corresponding to video image T1, video image T2, ..., video image T8 includes the character Ta makeup image and the character Tb makeup image. Based on the input operation on the video script generation interface for video image T1, video image T2, ..., video image T8, and the corresponding character Ta makeup image and character Tb makeup image, the terminal 1600 obtains video image T1, video image T2, ..., video image T8, and the corresponding character Ta makeup image and character Tb makeup image.

[0049] In step A2, terminal 1600 sends the target video segment and the corresponding target character feature information to server 1500. For example, terminal 1600 sends video image T1, video image T2, ..., video image T8 to server 1500, and sends the corresponding character Ta makeup image and character Tb makeup image.

[0050] In step A3, server 1500 acquires the target video segment and the corresponding target character feature information. For example, server 1500 acquires video image T1, video image T2, ..., video image T8, and acquires the corresponding character Ta makeup image and character Tb makeup image.

[0051] In step A4, server 1500 uses the target generation model to generate a script for the target video segment based on the target character feature information and the first prompt text, thereby obtaining the video generated script for the target video segment. The first prompt text is used to guide the target generation model to analyze the scene, characters and dialogue in each frame of the target video image based on the target character feature information, and to analyze errors and integrate the script to output script text that conforms to the script format.

[0052] For example, the first prompt text is prompt text 1; server 1500 inputs the character Ta makeup image, character Tb makeup image, video image T1, video image T2, ..., video image T8 into the target generation model. Guided by prompt text 1, based on the character Ta makeup image and character Tb makeup image, it analyzes the scene, characters and dialogue in video image Ti, i=1, 2, ..., 8, and analyzes errors and integrates the script to output script text that conforms to the script format. The video generated script of video image T1, video image T2, ..., video image T8 is script T.

[0053] In step A5, server 1500 sends a video-generated script to terminal 1600. For example, server 1500 sends script T to terminal 1600.

[0054] In step A6, terminal 1600 acquires the video-generated script. For example, terminal 1600 acquires script T.

[0055] In step A7, terminal 1600 outputs a video generation script for the target video segment. For example, terminal 1600 outputs script T for video image T1, video image T2, ..., video image T8.

[0056] Therefore, a target video segment is formed by acquiring multiple frames of target video images, and the target character feature information corresponding to the target video segment is obtained. This target character feature information can be provided to the model to identify the characters in each frame of the target video image when generating the video script. The target character feature information and the target video segment are input into the target generation model. Guided by the first prompt text, the model analyzes the scene, characters, and dialogue in each frame of the target video image based on the target character feature information. It is not necessary to traverse each frame of the target video image in advance to achieve character recognition. Instead, character recognition is achieved when the model generates the script, reducing character recognition time. Errors are analyzed and the script is integrated. The accuracy of character and dialogue analysis is improved through the thought chain, so as to accurately understand the video segment and output a script text that conforms to the script format. This results in a video-generated script for the target video segment, which accurately represents the target video segment in a standard script format, thereby improving the script generation efficiency and accuracy.

[0057] The video script generation method provided in this application embodiment can be executed by a server. However, in other embodiments of this application, the terminal may also have similar functions to the server to execute the video script generation method provided in this application embodiment, or the terminal and the server may jointly execute the video script generation method provided in this application embodiment. This embodiment does not limit this.

[0058] The following describes in detail a video script generation method provided in this application through method embodiments.

[0059] See Figure 3 ,Should Figure 3 The flowchart below illustrates a video script generation method provided in this application embodiment. For ease of description, the following embodiment uses a computer device as the execution subject of the video script generation method. The computer device may be the aforementioned terminal or the aforementioned server. The video script generation method includes the following S301-S302.

[0060] S301: Obtain the target video segment and the target character feature information corresponding to the target video segment; the target video segment includes multiple frames of target video images.

[0061] In related technologies, generating video scripts from video clips requires pre-identifying the characters in the video clips, inputting the video clips into a visual language model for script generation, and outputting the video script according to the script format. However, pre-identifying the characters in the video clips requires traversing every frame of the video clip to perform character recognition, which is not only time-consuming but may also lead to inaccurate understanding of the video clips due to mismatches between characters and dialogue, thereby reducing the efficiency and accuracy of script generation.

[0062] In this embodiment of the application, to solve the above-mentioned problems, corresponding character feature information can be obtained from video clips for character recognition when generating a script through a model, thereby avoiding the need to traverse every frame of the video clip in advance to achieve character recognition. Based on this, it is first necessary to obtain a target video clip formed by multiple frames of target video images, and then obtain the target character feature information corresponding to the target video clip.

[0063] Among them, the target video segment is a video segment formed by multiple frames of target video images to be generated script; the target character feature information is the feature information provided for the target video segment that can identify the characters in the target video segment.

[0064] In addition, the steps for obtaining target video segments may include: extracting frames and identifying keyframes from the target video lines of the script to be generated to obtain multiple target key images; segmenting the multiple target key images into plot segments based on the background similarity between each pair of adjacent target key images to obtain multiple target plot segments; and obtaining target video segments from the multiple target plot segments.

[0065] This method, based on acquiring target video segments formed by multiple frames of target video images, further acquires target character feature information corresponding to the target video segments. This target character feature information can be provided to the model to identify the characters in each frame of the target video image when generating the video script, providing a basis for character identification when generating the video script through the model in the future. It does not require traversing each frame of the target video image in advance to achieve character identification, but rather to achieve character identification when generating the script through the model in the future, which can reduce the character identification time.

[0066] As an example, the computer device acquires target video segments including video images T1, T2, ..., T8, and acquires target character feature information corresponding to video images T1, T2, ..., T8, including character Ta's makeup image and character Tb's makeup image.

[0067] See Figure 4 ,Should Figure 4 This is a schematic diagram illustrating a target video segment and corresponding target character feature information provided in an embodiment of this application. Based on the above example, the target video segment includes video image T1, video image T2, ..., video image T8, and the corresponding target character feature information includes a character Ta's makeup image and a character Tb's makeup image.

[0068] S302: Using the target generation model, based on the target character feature information and the first prompt text, a script is generated for the target video segment to obtain the video generated script for the target video segment; the first prompt text is used to guide the target generation model to analyze the scene, characters and dialogue in each frame of the target video image based on the target character feature information, and to analyze errors and integrate the script to output script text that conforms to the script format.

[0069] In this embodiment of the application, after executing the above-mentioned S301 to obtain the target video segment formed by multiple frames of target video images and the corresponding target character feature information, the script can be generated under the guidance of the first prompt text through the target generation model. In fact, it is based on the analysis of the scene, character and dialogue in each frame of target video image according to the target character feature information. Considering that after error analysis and script integration, it can avoid inaccurate understanding of video segments caused by insufficient matching of characters and dialogues. Finally, the video generated script of the target video segment is output according to the script format.

[0070] Based on this, the target character feature information and target video clips are input into the target generation model. Guided by the first prompt text, the model analyzes the scene, characters and dialogue in each frame of the target video image based on the target character feature information, and analyzes errors and integrates the script to output script text that conforms to the script format, thus obtaining the video-generated script for the target video clip.

[0071] The target generation model is a trained video script generation model. Specifically, it identifies the scene, characters, and dialogue in each video frame based on the corresponding character feature information, analyzes errors, and integrates the script to output script text that conforms to the script format. The first prompt text is a prompt text that guides the model to analyze the scene, characters, and dialogue in each video frame based on character feature information, analyzes errors, and integrates the script to output script text that conforms to the script format. The video generated script is the script-formatted text of the scene, time, characters, atmosphere, and dialogue in the target video segment.

[0072] This method uses the target character feature information and the target video clip as input to the target generation model. Guided by the first prompt text, the target generation model analyzes the scene, characters, and dialogue in each frame of the target video image based on the target character feature information. It does not require traversing each frame of the target video image in advance to achieve character recognition, and can complete character recognition when the model generates the script, reducing character recognition time. The target generation model also analyzes errors and integrates the script, and improves the accuracy of character and dialogue analysis through thought chain, so as to accurately understand the video clip and output script text that conforms to the script format, so as to obtain the video-generated script of the target video clip. This makes the video-generated script accurately represent the target video clip in a standard script format, thereby improving the script generation efficiency and accuracy.

[0073] As an example, based on the above example, the first prompt text is prompt text 1; the computer device inputs the character Ta makeup image, character Tb makeup image, video image T1, video image T2, ..., video image T8 into the target generation model. Guided by prompt text 1, it analyzes the scene, characters and dialogue in video image Ti based on the character Ta makeup image and character Tb makeup image, i=1, 2, ..., 8, and analyzes errors and integrates the script to output script text that conforms to the script format. The video-generated script of video image T1, video image T2, ..., video image T8 is script T.

[0074] See Figure 5 ,Should Figure 5This is a schematic diagram of a first prompt text provided in an embodiment of this application. Based on the above example, the first prompt text, namely prompt text 1, guides the target generation model to analyze the scene, characters, and dialogue of video image Ti based on the makeup images of character Ta and character Tb, i=1, 2, ..., 8, and analyzes errors and integrates the script to output script text that conforms to the script format. The specific prompt text 1 is as follows:

[0075] "Given reference characters—Character Ta: Character Ta's makeup image, Character Tb: Character Tb's makeup image. Given multiple target video images (video image T1, video image T2, ..., video image T8) in chronological order of the target video clip, convert them into a script format and output them. Note that familiar characters should be addressed by their names; unfamiliar characters can be referred to as 'Guest A,' 'Passerby B,' etc. The script should be generated according to the given image order; do not change the event order of the original keyframe plot."

[0076] Output in the following format: <think>Image Analysis: xxxx\nError Analysis: xxxx\nScript Integration: xxxx\nScript Output: <\think>xxx, In image analysis, scene, character, and dialogue analysis is performed on each frame of the target video image, outputting: "Image Identifier: x, Scene: x, Character: x, x, Dialogue: xxx.". In error analysis, the error descriptions from the image analysis are output, requiring concise, accurate, and information-rich wording. In script integration, the points to note when integrating the above information into the script are output. The final script is output.

[0077] See Figure 6 ,Should Figure 6 This is a schematic diagram of a video generation script for a target video segment provided in an embodiment of this application. Based on the above example, the video generation script, i.e., script T, is used to represent the scene, time, characters, atmosphere, and dialogue of video images T1, T2, ..., T8. The specific script T is as follows:

[0078] Script output: <\think>

[0079] Scene: XXX

[0080] Time: YYY

[0081] Characters: Character Ta, Character Tb

[0082] Atmosphere: ZZZ

[0083] Dialogue: WWW.

[0084] In summary, see [link / reference] Figure 7 ,Should Figure 7 This is a schematic diagram illustrating the generation of a script from a video using a target generation model, as provided in this application embodiment. After acquiring a target video segment 701 formed from multiple frames of target video images and corresponding target character feature information 702, the target video segment 701 and the corresponding target character feature information 702 are input into a target generation model 703. Guided by a first prompt text 704, the target generation model 703 analyzes the scene, characters, and dialogue in each frame of the target video segment 701 based on the target character feature information 702, and analyzes errors and integrates the script to output a script text conforming to the script format, thus obtaining a video-generated script 705 for the target video segment.

[0085] As can be seen from the above technical solution, a target video segment is formed by acquiring multiple frames of target video images, and the target character feature information corresponding to the target video segment is obtained. This target character feature information can be provided to the model to identify the characters in each frame of the target video image when generating the video script. The target character feature information and the target video segment are input into the target generation model. Guided by the first prompt text, the model analyzes the scene, characters, and dialogue in each frame of the target video image based on the target character feature information. It is not necessary to traverse each frame of the target video image in advance to achieve character recognition. Instead, character recognition is achieved when the model generates the script, reducing the character recognition time. Furthermore, it analyzes errors and integrates the script, improving the accuracy of character and dialogue analysis through thought chain, achieving accurate understanding of the video segment, and outputting script text that conforms to the script format. This results in a video-generated script for the target video segment, ensuring that the video-generated script accurately represents the target video segment in a standard script format, thereby improving script generation efficiency and accuracy.

[0086] In this embodiment, when executing S302 above, which uses a target generation model to generate a script for a target video segment based on target character feature information and a first prompt text, and obtains the specific implementation of the video generation script for the target video segment, the process actually begins by analyzing the scene, characters, and dialogue in each frame of the target video image using a thought chain approach, and then analyzing errors and integrating the script to obtain image analysis output text representing the scene, characters, and dialogue in each frame of the target video image, as well as error analysis output text and script integration output text. Then, based on the image analysis output text, combined with the error analysis output text and script integration output text, a script text conforming to the script format is output as the video generation script for the target video segment. Based on this, this application provides a possible implementation method, and S302 above may include, for example, the following S302a-S302b (not shown in the figure).

[0087] S302a: Using the target generation model, based on the target character feature information and the first prompt text, script generation thinking is performed on the target video segment to obtain image analysis output text, error analysis output text, and script integration output text; the image analysis output text is used to represent the scene, characters, and dialogue in each frame of the target video image.

[0088] S302b: Based on the image analysis output text, error analysis output text, and script integration output text, generate a script for the target video segment to obtain a video-generated script; the video-generated script is used to represent the scene, time, characters, atmosphere, and dialogue in the target video segment.

[0089] The image analysis output text includes representations of scenes, characters, and dialogues in each frame of the target video image; the error analysis output text indicates whether the characters and dialogues in each frame of the target video image are incorrectly identified, as well as the plot content of the target video segment; and the script integration output text indicates the script integration direction based on the image analysis output text and the error analysis output text.

[0090] This method inputs the target character feature information and the target video clip into the target generation model. Guided by the first prompt text, the target generation model first analyzes the scene, characters, and dialogue in each frame of the target video image through a thought chain approach to achieve video understanding. It then analyzes errors and integrates the script to achieve error analysis and script integration analysis. This not only yields image analysis output text representing the scene, characters, and dialogue in each frame of the target video image, but also error analysis output text to determine whether the characters and dialogue in each frame of the target video image have been incorrectly identified and to represent the plot content of the target video clip. In addition, it yields script integration output text to determine the script integration direction based on the image analysis output text and error analysis output text. In this way, a script text conforming to the script format can be output as the video-generated script for the target video clip. This ensures that the video-generated script accurately represents the scene, time, characters, atmosphere, and dialogue in the target video clip under the standard script format, thereby guaranteeing the accuracy of script generation.

[0091] As an example, based on the above example, the computer device inputs the makeup images of character Ta and Tb, video images T1, T2, ..., T8 into the target generation model. Guided by prompt text 1, it analyzes the scene, characters, and dialogue in video image Ti based on the makeup images of character Ta and character Tb, i=1, 2, ..., 8, and analyzes errors and integrates the script, obtaining image analysis output text, error analysis output text, and script integration output text, respectively: image analysis text T, error analysis text T, and script integration text T. Image analysis text T represents the scene, characters, and dialogue in video image Ti. Error analysis text T indicates whether the characters and dialogue in video image Ti are incorrectly identified, as well as the plot content of video images T1, T2, ..., T8. Script integration text T represents the script integration direction based on image analysis text T and error analysis text T. The computer device, based on image analysis text T, combined with error analysis text T and script integration text T, outputs script text conforming to the script format, as the script T for video images T1, T2, ..., T8.

[0092] See Figure 8 ,Should Figure 8 This is a schematic diagram illustrating an image analysis output text, an error analysis output text, and a script integration output text provided in an embodiment of this application. In the above example, the above... Figure 4 and the above Figure 5 Based on this, the image analysis output text, i.e., image analysis text T, represents the scene, characters, and dialogue in video image Ti, i=1, 2, ..., 8; the error analysis output text, i.e., error analysis text T, represents whether the characters and dialogue in video image Ti were identified incorrectly, as well as the plot content of video images T1, T2, ..., T8; the script integration output text, i.e., script integration text T, represents the script integration direction based on image analysis text T and error analysis text T. The specific details of image analysis text T, error analysis text T, and script integration text T are as follows:

[0093] " <think>Image analysis:

[0094] Image identifier: T1

[0095] Scene: x1

[0096] Characters: Character Ta, Character Tb

[0097] Dialogue: 111111

[0098] Image identifier: T2

[0099] Scene: x2

[0100] Character: Character Tb

[0101] Dialogue: 222222

[0102]

[0103] Image identifier: T8

[0104] Scene: x8

[0105] Characters: Character Ta, Character Tb

[0106] Dialogue: 888888

[0107] Error Analysis:

[0108] Character error The character in the video image matches the reference character, and no other unnamed characters appear, so this is assumed to be correct.

[0109] Dialogue error No typos were found in the dialogue.

[0110] Plot Summary :vvv.

[0111] Script Integration:

[0112] Inject actions into dialogue, retain key expressions, and remove redundant descriptions. Use △ to inject environmental character action information.

[0113] In this embodiment, based on the fact that the target character feature information is the feature information of the character in the target video segment that can be identified, and considering that both character makeup images and character description text can represent character features to identify the character, the target character feature information can be the target character makeup image corresponding to the target video segment, or the target character description text corresponding to the target video segment, or the target character makeup image plus the target character description text corresponding to the target video segment. Based on this, this application provides a possible implementation method. The step of determining the target character feature information in S301 above may include, for example, S1 (not shown in the figure): determining at least one of the target character makeup image or the target character description text corresponding to the target video segment as the target character feature information.

[0114] Among them, the target character makeup image is a makeup image provided for the target video clip that can identify the character in the target video clip; the target character description text is a description text provided for the target video clip that can identify the character in the target video clip.

[0115] This method uses the makeup image of the target character corresponding to the target video segment as the feature information of the target character. The makeup image of the target character accurately represents the target character corresponding to the target video segment in terms of visual details, providing accurate and subtle basis for character identification when the video script is generated through the model. It is convenient to accurately identify whether the character in each frame of the target video image is the target character.

[0116] The target character description text corresponding to the target video segment is used as the target character feature information. This target character description text accurately represents the target character corresponding to the target video segment in terms of semantic attributes, providing an accurate and generalized basis for character identification when the video script is generated through the model. This facilitates the rapid identification of whether the character in each frame of the target video image is the target character.

[0117] The target character's makeup image and description text corresponding to the target video segment are used as the target character's feature information. The target character's makeup image accurately represents the target character corresponding to the target video segment in terms of visual details, and the target character's description text accurately represents the target character corresponding to the target video segment in terms of semantic attributes. This provides an accurate and comprehensive basis for character identification when the model is used to generate video scripts, making it easier to more accurately identify whether the character in each frame of the target video image is the target character.

[0118] As an example, based on the above example, the computer device uses the character Ta makeup image and character Tb makeup image corresponding to video image T1, video image T2, ..., video image T8 as target character feature information.

[0119] As another example, the computer device uses the character description text Ta and character description text Tb corresponding to video images T1, T2, ..., T8 as target character feature information.

[0120] As another example, the computer device uses the character Ta makeup image + character Ta description text and the character Tb makeup image + character Tb description text corresponding to video image T1, video image T2, ..., video image T8 as target character feature information.

[0121] Furthermore, in this embodiment, the target dialogue text corresponding to the target video segment can be obtained in advance. When executing the above-mentioned S302, the target generation model generates a script for the target video segment based on the target character feature information and the first prompt text to obtain the specific implementation of the video generation script for the target video segment. The target character feature information, the target video segment, and the target dialogue text are input into the target generation model. Under the guidance of the first prompt text, when analyzing the scene, character, and dialogue in each frame of the target video image based on the target character feature information, it is not necessary to identify the dialogue. After analyzing errors and integrating the script, a script text conforming to the script format is output to obtain the video generation script for the target video segment. Based on this, this application provides a possible implementation method. The method may also include S2 (not shown in the figure): obtaining the target dialogue text corresponding to the target video segment; correspondingly, the above-mentioned S302 may include S302c (not shown in the figure): generating a script for the target video segment based on the target character feature information, the target dialogue text, and the first prompt text through the target generation model to obtain the video generation script.

[0122] The target dialogue text includes the dialogue text in each frame of the target video image; the steps to obtain the target dialogue text may be: performing text recognition on each frame of the target video image in multiple frames of target video images to obtain the target dialogue text corresponding to the target video segment.

[0123] This method, before generating a video script for a target video segment using a model, first obtains the target dialogue text corresponding to the target video segment. This target dialogue text represents the dialogue in each frame of the target video image, which can reduce the dialogue recognition steps during video script generation. Thus, the target character feature information, the target video segment, and the target dialogue text serve as the model input for the target generation model. Guided by the first prompt text, the target generation model analyzes the scene, characters, and dialogue in each frame of the target video image based on the target character feature information without needing to recognize the dialogue, reducing dialogue recognition time and ensuring the accuracy of the dialogue to a certain extent. Further, through error analysis and script integration, the model analyzes whether there are errors in character and dialogue recognition through thought chain analysis, avoiding errors in understanding video segments, and outputs script text that conforms to the script format, resulting in a video-generated script for the target video segment. This allows the video-generated script to more accurately represent the target video segment under a standard script format, thereby further improving the script generation efficiency and accuracy.

[0124] As an example, based on the above example, the computer device acquires the target dialogue text T corresponding to video images T1, T2, ..., T8. The computer device inputs the character Ta's makeup image, the character Tb's makeup image, video images T1, T2, ..., T8 into the target generation model. Guided by the prompt text 1, the computer device analyzes the scene, characters, and dialogue in video image Ti based on the character Ta's makeup image and the character Tb's makeup image. It does not need to identify the dialogue, i=1, 2, ..., 8. After analyzing errors and integrating the script, the computer device outputs script text that conforms to the script format, thus obtaining the script T for video images T1, T2, ..., T8.

[0125] In this embodiment, considering the continuity between the video-generated script of the target video segment and the historical video script when generating a video script for the target video segment using a model, the historical video script corresponding to the target video segment can be obtained in advance. When executing the above-mentioned S302, the target generation model generates a script for the target video segment based on the target character feature information and the first prompt text to obtain the specific implementation of the video-generated script for the target video segment. The target character feature information, the target video segment, and the historical video script are input into the target generation model. Under the guidance of the first prompt text, the historical video script is referenced when analyzing the scene, character, and dialogue in each frame of the target video image based on the target character feature information. After analyzing errors and integrating the script, a script text conforming to the script format is output to obtain the video-generated script for the target video segment. Based on this, this application provides a possible implementation method, which may also include S3 (not shown in the figure): obtaining the historical video script corresponding to the target video segment; correspondingly, the above-mentioned S302 may include S302d (not shown in the figure): generating a script for the target video segment using the target generation model based on the target character feature information, the historical video script, and the first prompt text to obtain the video-generated script.

[0126] Among them, the historical video script is the script of a historical video that is related to the plot content of the target video segment; the first prompt text is used to guide the target generation model to refer to the historical video script when analyzing the scene, characters and dialogue in each frame of the target video image based on the target character feature information, and to analyze errors and integrate the script to output script text that conforms to the script format.

[0127] This method, before generating a video script for a target video segment using a model, first obtains the historical video script corresponding to the target video segment. This historical video script represents the historical plot content related to the plot content of the target video segment and can be provided to the model for reference during video script generation. Thus, the target character feature information, the target video segment, and the historical video script serve as the model input for the target generation model. Guided by the first prompt text, the target generation model analyzes the scenes, characters, and dialogues in each frame of the target video image based on the target character feature information, while referring to the historical video script. This ensures that the analyzed scenes, characters, and dialogues follow the historical plot content and avoid fantasies. After analyzing errors and integrating the script, the model outputs script text that conforms to the script format, resulting in the video-generated script for the target video segment. This ensures that the video-generated script coherently and accurately represents the target video segment based on the historical plot content in a standard script format, guaranteeing the continuity between the video-generated script and the historical video script, thereby further improving the accuracy of script generation.

[0128] As an example, based on the above example, the computer device acquires the historical video scripts corresponding to video images T1, T2, ..., T8 as script H; the computer device inputs the makeup images of character Ta and Tb, video images T1, T2, ..., T8 into the target generation model, and under the guidance of prompt text 1, it analyzes the scene, characters and dialogue in video image Ti based on the makeup images of character Ta and character Tb, referring to script H, i=1, 2, ..., 8. After analyzing errors and integrating the script, it outputs script text that conforms to the script format, thus obtaining the script T of video images T1, T2, ..., T8.

[0129] Correspondingly, the prompt text 1 is updated to include "In order to ensure the continuity of the script, the historical video script preceding this plot is provided as 'xxxx'. After the historical video script, please continue the script according to the information provided above. The continued script must strictly follow the information provided and is not allowed to include non-existent plots."

[0130] In this embodiment, the target generation model is based on a model that identifies the scene, characters, and dialogue in each frame of a multi-frame video image based on corresponding character feature information, analyzes errors, and integrates the script to output script text conforming to the script format. The training method of the target generation model can be as follows: taking sample video segments formed by multiple sample video images and sample character feature information corresponding to the sample video segments as model input; achieving script generation through the initial generation model under the guidance of prompt text, that is, analyzing the scene, characters, and dialogue in each frame of the sample video image based on sample character feature information, and outputting script text conforming to the script format through error analysis and script integration to obtain the model output; taking the script content formed by image analysis content text, error analysis content text, script integration content text, and video script content text based on the first video segment as the training target; and training the initial generation model based on the model output and the training target to obtain the target generation model.

[0131] Therefore, the training process of the target generation model is as follows:

[0132] First, it is necessary to obtain a first video segment formed by multiple frames of first video images, and then obtain the first character feature information corresponding to the first video segment.

[0133] Then, the first character feature information and the first video clip are input into the initial generation model. Guided by the first prompt text, the model analyzes the scene, characters and dialogue in each frame of the first video image based on the first character feature information, and analyzes errors and integrates the script to output script text that conforms to the script format, thus obtaining the first output content of the first video clip.

[0134] Finally, based on the first script content formed from the image analysis content text, error analysis content text, script integration content text, and video script content text of the first video segment, the model parameters of the initial generation model are trained by the difference between the first output content and the first script content, and the trained initial model is used as the target generation model. Based on this, this application provides a possible implementation method, and the training steps of the target generation model in S302 above may include, for example, the following S4-S6 (not shown in the figure).

[0135] S4: Obtain the first video segment and the first character feature information corresponding to the first video segment; the first video segment includes multiple frames of first video images.

[0136] S5: Using the initial generation model, based on the first character feature information and the second prompt text, a script is generated for the first video segment to obtain the first output content of the first video segment; the second prompt text is used to guide the initial generation model to analyze the scene, characters and dialogue in each frame of the first video image based on the first character feature information, and to analyze errors and integrate the script to output script text that conforms to the script format.

[0137] S6: Based on the difference between the first output content and the first script content of the first video segment, train the initial generation model to obtain the target generation model; the first script content includes image analysis content text, error analysis content text, script integration content text and video script content text based on the first video segment.

[0138] The first video segment is a video segment formed by multiple frames of first video images in the video to be trained; the first character feature information is the feature information provided for the first video segment that can identify the character in the first video segment; the initial generation model is a pre-trained generation model; the second prompt text is the prompt text that guides the model to analyze the scene, character and dialogue in each frame of video image based on the character feature information, and to analyze errors and integrate the script to output script text that conforms to the script format; the first output content is the output content obtained by performing image analysis, error analysis, script integration and script output based on the first character feature information of the first video segment.

[0139] The step of determining the first character feature information may be: determining at least one of the first character makeup image or the first character description text corresponding to the first video segment as the first character feature information. The first character makeup image is a makeup image provided for the first video segment that can identify the character in the first video segment; the first character description text is a description text provided for the first video segment that can identify the character in the first video segment.

[0140] Furthermore, the representation format of the first video clip, the first character feature information corresponding to the first video clip (which includes the first character's makeup image), the second prompt text, and the first script content of the first video clip is as follows:

[0141] {

[0142] "sample_id (sample identifier)": aaa;

[0143] "prompt (prompt text)": Secondary prompt text;

[0144] "image (input image)": First character makeup image, first video clip;

[0145] "output (training objective)": The content of the first scenario;

[0146] }

[0147] This method acquires a first video segment formed by multiple frames of first video images and obtains the first character feature information corresponding to the first video segment. This first character feature information can be provided to the model to identify the characters in each frame of the first video image when generating the video script. The first character feature information and the first video segment are input into the initial generation model. Guided by the second prompt text, the model analyzes the scene, characters, and dialogue in each frame of the first video image based on the first character feature information. It is not necessary to traverse each frame of the first video image in advance to achieve character recognition. Instead, character recognition is achieved when the model generates the script, reducing character recognition time. It also analyzes errors and integrates the script, improving the accuracy of character and dialogue analysis through thought chain, achieving accurate understanding of the video segment, and outputting script text that conforms to the script format, thus obtaining the first output content representing the first video segment.

[0148] For the case where the first script content is formed from image analysis content text, error analysis content text, script integration content text, and video script content text of the first video segment, the model parameters of the initial generation model are trained by the difference between the first output content and the first script content, so that the first output content is close to the first script content. This allows the target generation model trained based on the initial generation model to learn the image analysis content text, error analysis content text, script integration content text, and video script content text based on the first video segment. As a result, the target generation model can identify the scene, characters, and dialogue in each frame of the video segment based on the corresponding character feature information, and analyze errors and integrate the script to output script text that conforms to the script format and accurately represents the video segment.

[0149] As an example, the computer device acquires a first video segment including video image S1_1, video image S1_2, ..., and acquires first character feature information corresponding to video image S1_1, video image S1_2, ... including character S1_a makeup image, character S1_b makeup image, ...

[0150] The second prompt text is prompt text 2. The computer device inputs the makeup images of character S1_a, character S1_b, ..., video images S1_1, video images S1_2, ... into the initial generation model. Guided by prompt text 2, it analyzes the scenes, characters, and dialogues in each frame of video images S1_1, S1_2, ... based on the makeup images of character S1_a, character S1_b, ..., and analyzes errors and integrates the script to output script text that conforms to the script format. The first output content of video images S1_1, S1_2, ... is output content S1'.

[0151] The computer device, based on the image analysis content text, error analysis content text, script integration content text, and video script content text formed by video image S1_1, video image S1_2, ..., and the script content text, trains the model parameters of the initial generation model by outputting the difference between content S1' and script content S1, and uses the trained initial model as the target generation model.

[0152] See Figure 9 ,Should Figure 9 This is a schematic diagram illustrating the process of training an initial generation model to obtain a target generation model, as provided in an embodiment of this application. Based on the above example, after acquiring a first video segment 901 formed from multiple frames of first video images and corresponding first character feature information 902, the first video segment 901 and the corresponding first character feature information 902 are input into an initial generation model 903. Guided by a second prompt text 904, the initial generation model 903 analyzes the scene, characters, and dialogue in each frame of the first video image included in the first video segment 901 based on the first character feature information 902, and analyzes errors and integrates the script to output script text conforming to the script format, obtaining the first output content 905 of the first video segment. Through the difference 907 between the first output content 905 and the first script content 906 of the first video segment, the model parameters of the initial generation model 903 are trained, and the trained initial generation model 903 is used as the target generation model 703. The first script content 906 includes image analysis content text based on the first video segment 901, error analysis content text, script integration content text, and video script content text.

[0153] In this embodiment, when executing the above-mentioned S5, which generates a script for the first video segment based on the first character feature information and the second prompt text using an initial generation model, and obtains the first output content of the first video segment, the specific implementation actually involves first analyzing the scene, characters, and dialogue in each frame of the first video image through a thought chain approach, and analyzing errors and integrating the script to obtain the first image analysis text representing the scene, characters, and dialogue in each frame of the first video image, and obtaining the first error analysis text and the first script integration text; then, based on the first image analysis text, combined with the first error analysis text and the first script integration text, a script text conforming to the script format is output as the first output script; thus, the output content formed by the first image analysis text, the first error analysis text, the first script integration text, and the first output script is used as the first output content of the first video segment. Based on this, this application provides a possible implementation method, and the above-mentioned S5 may include, for example, the following S5a-S5c (not shown in the figure).

[0154] S5a: Using the initial generation model, based on the first character feature information and the second prompt text, the script generation thinking is performed on the first video segment to obtain the first image analysis text, the first error analysis text, and the first script integration text; the first image analysis text is used to represent the scene, characters, and dialogue in each frame of the first video image.

[0155] S5b: Based on the first image analysis text, the first error analysis text, and the first script integration text, generate a script output for the first video segment to obtain the first output script.

[0156] S5c: The first image analysis text, the first error analysis text, the first script integration text, and the first output script are determined as the first output content.

[0157] The first image analysis text includes the representation text of the scene, characters and dialogue in each frame of the first video image; the first error analysis text is used to indicate whether the characters and dialogue in each frame of the first video image are identified incorrectly, as well as the plot content of the first video segment; the first script integration text is used to indicate the script integration direction based on the first image analysis text and the first error analysis text.

[0158] This method involves inputting the first character feature information and the first video segment into the initial generation model. Guided by the second prompt text, the initial generation model first analyzes the scene, characters, and dialogue in each frame of the first video image through a thought chain approach to achieve video understanding. It also analyzes errors and integrates the script to achieve error analysis and script integration analysis. This not only yields the first image analysis text representing the scene, characters, and dialogue in each frame of the first video image, but also the first error analysis text to determine whether the characters and dialogue in each frame of the first video image are identified incorrectly, and to represent the plot content of the first video segment. In addition, it yields the first script integration text to determine the script integration direction based on the first image analysis text and the first error analysis text. In this way, a script text conforming to the script format can be output as the first output script. The first image analysis text, the first error analysis text, the first script integration text, and the first output script together form the first output content of the first video segment, so that the first output content is aligned with the first script content of the first video segment.

[0159] In this embodiment, when performing S6 to train the initial generation model based on the difference between the first output content and the first script content of the first video segment to obtain the specific implementation of the target generation model, the probability that the first output content is the first script content is actually determined first, that is, the cumulative value of the probability that each predicted word in the first output content is the corresponding first word in the first script content; then, the cumulative value is maximized to make the first output content close to the first script content, so as to train the model parameters of the initial model, and the trained initial model is used as the target generation model. Based on this, this application provides a possible implementation method, and the above S6 may include, for example, the following S6a-S6b (not shown in the figure).

[0160] S6a: Determine the sum of probabilities that each predicted word in the first output content is the corresponding first word in the first script content.

[0161] S6b: Obtain the target generative model by maximizing the probability and training the initial generative model.

[0162] In the first output content, each predicted word is the sum of probabilities of the corresponding first word in the first script content, specifically the cumulative value of the probability that each predicted word in the first output content is the corresponding first word in the first script content.

[0163] This method uses the probability sum of each predicted word in the first output content to determine if it corresponds to a first word in the first script content. A smaller probability sum indicates a greater difference between the first output content and the first script content, while a larger probability sum indicates a smaller difference. By maximizing the probability sum, the model parameters of the initial generation model are trained, making the first output content close to the first script content. This allows the target generation model, trained based on the initial generation model, to learn multiple first words in the first script content. Consequently, the target generation model can identify the scene, characters, and dialogue in each frame of a video segment formed from multiple video images based on corresponding character feature information, analyze errors, and integrate the script to output script text that accurately represents the video segment in a script format.

[0164] As an example, the probability that each predicted word in the first output content is the corresponding first word in the first script content and the corresponding loss formula are as follows:

[0165] ;

[0166] Where p(m, n) represents the probability that the nth predicted word in the first output content of the mth first video segment in a batch is the corresponding first word in the first script content; bs represents the batch size, i.e. the number of first video segments in a batch; Nm represents the number of predicted words in the first output content of the first video segment.

[0167] In this embodiment, the process of obtaining the first video segment in S4 is as follows: for a first sample video randomly selected from a first candidate video set, multiple first key images are obtained by first extracting frames from the video and then identifying key frames; considering whether adjacent first key images are similar in plot, the multiple first key images are divided into multiple plot segments as multiple first plot segments; one of the multiple first plot segments is taken as the first video segment. Based on this, this application provides a possible implementation method, and the step of obtaining the first video segment in S4 may include, for example, the following S7-S9 (not shown in the figure).

[0168] S7: Perform video frame extraction and keyframe recognition on the first sample video to obtain multiple first key images; the first sample video is randomly extracted from the first candidate video set.

[0169] S8: Based on the background similarity between any two adjacent first key images, perform plot segmentation on multiple first key images to obtain multiple first plot segments.

[0170] S9: Obtain the first video segment from multiple first plot segments.

[0171] The first candidate video set is a video set formed by multiple first videos to be trained; the first sample video is a video randomly selected from the first candidate video set; the first key image is a key image obtained by extracting frames from the first sample video; the background similarity between any two adjacent first key images is the similarity of the plot between any two adjacent first key images; each first plot segment includes a video segment formed by multiple first key images of the same plot in multiple first key images; different first plot segments are video segments of different plots in multiple first key images.

[0172] In scenarios where video script generation is achieved through a model, this method takes into account the model's processing capability for video image frames. After randomly extracting the first sample video from the first candidate video set, multiple first key images are obtained through video frame extraction and key frame recognition, achieving sparse frame extraction and key recognition for the video. Furthermore, considering whether multiple first key images represent the same plot, the similarity between any two adjacent first key images indicates whether the plots of any two adjacent first key images are similar. The multiple first key images are then divided into multiple first plot segments, thus dividing the video images after sparse frame extraction and key recognition into different plots. In this way, the first video segment to be generated by the model can be obtained from the multiple first plot segments.

[0173] See Figure 10 ,Should Figure 10 This application provides a step diagram for obtaining a first video segment based on a first sample video, as provided in an embodiment. The step includes the following steps: S1001-S1004.

[0174] S1001: Obtain the first sample video.

[0175] S1002: The first sample video is processed by video frame extraction and keyframe recognition to obtain multiple first key images.

[0176] S1003: Based on the background similarity between every two adjacent first key images, the multiple first key images are divided into multiple first plot segments.

[0177] S1004: Obtain the first video segment from multiple first plot segments.

[0178] S1003 includes the following S1003a-S1003e.

[0179] S1003a: Remove the foreground from the first key image of each frame to obtain the background image of the first key image of each frame.

[0180] S1003b: Extract features from the background image of the first key image in each frame to obtain the background features of the first key image in each frame.

[0181] S1003c: Determine the plot segmentation points in multiple first key images based on the background feature similarity between any two adjacent first key images.

[0182] If the background feature similarity of two adjacent first key images is greater than or equal to a preset similarity, it means that the two adjacent first key images belong to the same plot. If the background feature similarity of two adjacent first key images is less than the preset similarity, it means that the two adjacent first key images belong to different plots. Based on the two adjacent first key images whose background feature similarity is less than the preset similarity, plot segmentation points in multiple first key images are determined.

[0183] S1003d: If the plot segmentation point in the first key image of multiple frames is within a dialogue time corresponding to the first key image of multiple frames, move the plot segmentation point to the end point of the dialogue to obtain the updated plot segmentation point in the first key image of multiple frames.

[0184] S1003e: Based on the plot segmentation points in the updated multi-frame first key image, the multi-frame first key image is segmented into multiple first plot segments.

[0185] In this embodiment of the application, the first script content in S6 above can be obtained by first inputting the first video clip into the script generation model to generate a script, and then revising the first generated script. Based on this, the application provides a possible implementation method. The step of obtaining the first script content in S6 above may include, for example, the following S10-S11 (not shown in the figure).

[0186] S10: Generate a script for the first video clip using a script generation model to obtain the first generated script for the first video clip.

[0187] S11: Revise the first generated script to obtain the content of the first script.

[0188] The first generated script is the script output by the script generation model after the first video clip is input into it.

[0189] This method inputs the first video clip into the script generation model. Based on the powerful generation capabilities of the script generation model, it can quickly and automatically generate a first generated script for the first video clip to represent the first video clip in a script format. The first generated script is revised through script revision, realizing image analysis, error analysis, script integration, and script output. This ensures that the content of the first script accurately represents the image analysis content, error analysis content, script integration content, and script output content of the first video clip in a standard script format.

[0190] Furthermore, the specific training process for obtaining the target generative model based on the initial generative model is as follows:

[0191] (1) Parameter initialization: The model parameters of the initial generated model are initialized by using the weights of the open-source pre-trained small model network.

[0192] (2) Configure learning parameters: Configure learning rate, batch size, number of training rounds, optimizer and other learning parameters.

[0193] (3) Learning process: During each batch of learning, for each first video segment (sample), determine the probability that each predicted word in the first output content is the corresponding first word in the first script content, calculate the loss under certain rules, calculate the total loss of the batch, and feed the total loss back to the model to calculate the gradient of the model parameters in order to update the model parameters.

[0194] Furthermore, in this embodiment, the optimization method of the target generation model can be as follows: taking other video segments formed by multiple frames of other video images, and other character feature information corresponding to other video segments as model input; achieving script generation through the target generation model under the guidance of prompt text, that is, analyzing the scene, characters and dialogue in each frame of other video images based on other character feature information, and after error analysis and script integration, outputting script text that conforms to the script format to obtain the model output; taking the script content formed by the image analysis content text, error analysis content text, script integration content text and video script content text based on the second video segment as the training target; determining the reward score based on the model output and the training target; and optimizing the target generation model to obtain the optimized target generation model.

[0195] Therefore, the optimization process of the target generation model is as follows:

[0196] First, it is necessary to obtain a second video segment formed by multiple frames of second video images, and then obtain the second character feature information corresponding to the second video segment.

[0197] Then, the second character feature information and the second video clip are input into the initial generation model. Guided by the third prompt text, the scene, characters and dialogue in each frame of the second video image are analyzed based on the second character feature information. Errors are analyzed and the script is integrated to output script text that conforms to the script format, thus obtaining the second output content of the second video clip.

[0198] Finally, based on the second script content formed by the image analysis content text, error analysis content text, script integration content text, and video script content text of the second video clip, a reward score corresponding to the second output content is determined based on the second script content; the model parameters of the target generation model are optimized through the reward score to obtain the optimized target generation model. Based on this, this application provides a possible implementation method, which may further include S12-S15 (not shown in the figure).

[0199] S12: Obtain the second video segment and the second character feature information corresponding to the second video segment; the second video segment includes multiple frames of second video images.

[0200] S13: Using the target generation model, based on the second character feature information and the third prompt text, a script is generated for the second video segment to obtain the second output content of the second video segment; the third prompt text is used to guide the target generation model to analyze the scene, characters and dialogue in each frame of the second video image based on the second character feature information, and to analyze errors and integrate the script to output script text that conforms to the script format.

[0201] S14: Determine the reward score based on the second output content and the second script content of the second video clip; the second script content includes image analysis content text, error analysis content text, script integration content text, and video script content text based on the second video clip.

[0202] S15: Optimize the target generation model based on the reward score to obtain the optimized target generation model.

[0203] The second video segment is a video clip formed by multiple frames of second video images from the training video; the second character feature information is the feature information provided for the second video segment that can identify the characters in the second video segment; the third prompt text is the prompt text that guides the model to analyze the scene, characters and dialogue in each frame of video image based on the character feature information, and to analyze errors and integrate the script to output script text that conforms to the script format; the second output content is the output content obtained by performing image analysis, error analysis, script integration and script output based on the second character feature information of the second video segment; the reward score is a quality score of the second output content in terms of format and content based on the second script content.

[0204] The step of determining the second character feature information may include: identifying at least one of the second character makeup image or the second character description text corresponding to the second video segment as the second character feature information. The second character makeup image is a makeup image provided for the second video segment that can identify the character in the second video segment; the second character description text is a descriptive text provided for the second video segment that can identify the character in the second video segment.

[0205] Furthermore, the representation format of the second video clip, the corresponding second character feature information (including the second character's makeup image, the second prompt text), and the second script content of the second video clip is as follows:

[0206] {

[0207] "sample_id (sample identifier)": bbb;

[0208] "prompt (prompt text)": Third prompt text;

[0209] "image (input image)": Second character makeup image, second video clip;

[0210] "output (training objective)": Content of the second scenario;

[0211] "output1": Image analysis content text;

[0212] "output2": Error analysis content text;

[0213] "output3": Integrated script content text;

[0214] "output4": Video script content text;

[0215] }

[0216] This method acquires a second video segment formed by multiple frames of second video images and obtains the second character feature information corresponding to the second video segment. This second character feature information can be provided to the model to identify the characters in each frame of the second video image when generating the video script. The second character feature information and the second video segment are input into the initial generation model. Guided by the third prompt text, the model analyzes the scene, characters, and dialogue in each frame of the second video image based on the second character feature information. It is not necessary to traverse each frame of the second video image in advance to achieve character recognition. Instead, character recognition is achieved when the model generates the script, reducing character recognition time. It also analyzes errors and integrates the script, improving the accuracy of character and dialogue analysis through thought chain, and achieving accurate understanding of the video segment to output script text that conforms to the script format, thus obtaining the second output content representing the second video segment.

[0217] For the second script content formed from image analysis text, error analysis text, script integration text, and video script content text of the second video clip, a reward score is determined based on the second script content to determine the quality score of the second output content in terms of format and content. The model parameters of the target generation model are optimized through the reward score so that the second output content is close to the second script content in terms of format and content. This allows the optimized target generation model to focus on learning the format and content of the image analysis text, error analysis text, script integration text, and video script content text based on the second video clip. As a result, the optimized target generation model can identify the scene, characters, and dialogue in each frame of the video clip formed from multiple video images based on the corresponding character feature information, and analyze errors and integrate the script to output script text that is more in line with the script format and more accurately represents the video clip.

[0218] As an example, the computer device acquires the second video segment, including video image S2_1, video image S2_2, ..., and acquires the second character feature information corresponding to video image S2_1, video image S2_2, ..., including character S2_a makeup image, character S2_b makeup image, ...

[0219] The third prompt text is prompt text 3. The computer device inputs the makeup images of character S2_a, character S2_b, ..., video images S2_1, video images S2_2, ... into the target generation model. Guided by prompt text 3, the computer analyzes the scenes, characters, and dialogues in each frame of video images S2_1, S2_2, ... based on the makeup images of character S2_a, character S2_b, ..., and analyzes errors and integrates the script to output script text that conforms to the script format. The second output content of video images S2_1, S2_2, ... is output content S2'.

[0220] The computer device, based on the second script content S2 formed by the image analysis content text, error analysis content text, script integration content text, and video script content text of video image S2_1, video image S2_2, ..., determines the reward score corresponding to the output content S2' based on the script content S2; and optimizes the model parameters of the target generation model through the reward score to obtain the optimized target generation model.

[0221] See Figure 11 ,Should Figure 11 This is a schematic diagram illustrating an optimized target generation model obtained from an embodiment of this application. Based on the above example, after acquiring a second video segment 1101 formed from multiple frames of second video images and corresponding second character feature information 1102, the second video segment 1101 and corresponding second character feature information 1102 are input into a target generation model 703. Guided by a third prompt text 1103, the target generation model 703 analyzes the scene, characters, and dialogue in each frame of the second video image included in the second video segment 1101 based on the second character feature information 1102, and analyzes errors and integrates the script to output script text conforming to the script format, obtaining the second output content 1104 of the second video segment. A reward score 1106 is determined by the second output content 1104 and the second script content 1105 of the second video segment. The second output content 1104 includes image analysis content text, error analysis content text, script integration content text, and video script content text based on the second video segment. The model parameters of the target generation model 703 are optimized using the reward score 1106 to obtain an optimized target generation model 1107.

[0222] In this embodiment, when executing the above-described S13, which uses a target generation model to generate a script for the second video segment based on the second character feature information and the third prompt text, and obtains the second output content of the second video segment, the specific implementation actually involves first analyzing the scene, characters, and dialogue in each frame of the second video image using a thought chain approach, and analyzing errors and integrating the script to obtain a second image analysis text representing the scene, characters, and dialogue in each frame of the second video image, and obtaining a second error analysis text and a second script integrated text; then, based on the second image analysis text, combined with the second error analysis text and the second script integrated text, a script text conforming to the script format is output as the second output script; thus, the output content formed by the second image analysis text, the second error analysis text, the second script integrated text, and the second output script is used as the second output content of the second video segment. Based on this, this application provides a possible implementation method, and the above-described S13 may include, for example, the following S13a-S13c (not shown in the figure).

[0223] S13a: Using the target generation model, based on the second character feature information and the third prompt text, script generation thinking is performed on the second video segment to obtain the second image analysis text, the second error analysis text, and the second script integration text; the second image analysis text is used to represent the scene, characters, and dialogue in each frame of the second video image.

[0224] S13b: Based on the second image analysis text, the second error analysis text, and the second script integration text, generate a script output for the second video segment to obtain the second output script.

[0225] S13c: The second image analysis text, the second error analysis text, the second script integration text, and the second output script are determined as the second output content.

[0226] The second image analysis text includes representation texts of scenes, characters, and dialogues in each frame of the second video images; the second error analysis text is used to indicate whether the characters and dialogues in each frame of the second video images are incorrectly identified, as well as the plot content of the second video segment; and the second script integration text is used to indicate the script integration direction based on the second image analysis text and the second error analysis text.

[0227] This method, after inputting the second character feature information and the second video segment into the initial generation model, guides the initial generation model to first analyze the scene, characters, and dialogue in each frame of the second video image through a thought chain approach to achieve video understanding. It also analyzes errors and integrates the script to achieve error analysis and script integration analysis. This not only yields second image analysis text representing the scene, characters, and dialogue in each frame of the second video image, but also second error analysis text to determine whether the characters and dialogue in each frame of the second video image are identified incorrectly, and to represent the plot content of the second video segment. In addition, it yields second script integration text to determine the script integration direction based on the second image analysis text and the second error analysis text. In this way, a script text conforming to the script format can be output as the second output script. The second image analysis text, the second error analysis text, the second script integration text, and the second output script together form the second output content of the second video segment, so that the second output content is aligned with the second script content of the second video segment.

[0228] In this embodiment of the application, when performing the above-mentioned S14 to determine the reward score based on the second script content of the second output content and the second video segment, it is considered that the reward score is a quality score of the second output content in terms of format and content based on the second script content. On the basis that the second output content includes the second image analysis text, the second error analysis text, the second script integration text and the second output script, the second output content has a fixed overall format. The second image analysis text, the second error analysis text and the second script integration text in the second output content are thinking content texts and also have a fixed thinking format. The second output script in the second output content has a fixed output format.

[0229] Therefore, not only is an overall format score determined based on the overall format of the second output content; but also a thinking format score and a thinking content score are determined through the format and content of the second image analysis text, the second error analysis text, and the second script integration text, as well as the image analysis content text, error analysis content text, and script integration content text in the second script content; furthermore, a script format score and a script content score are determined through the format and content of the second output script, as well as the video script content text in the second script content; and a reward score is determined by combining the overall format score, thinking format score, thinking content score, script format score, and script content score. Based on this, this application provides a possible implementation, and the above S14 may include, for example, the following S14a-S14d (not shown in the figure).

[0230] S14a: Determine the overall format score based on the overall format of the second output content and the overall format of the second script content.

[0231] S14b: Determine the thinking format score and thinking content score based on the format and content of the second image analysis text, the second error analysis text, and the second script integration text, as well as the format and content of the image analysis content text, error analysis content text, and script integration content text in the second script content.

[0232] S14c: Determine the script format score and script content score based on the format and content of the second output script, as well as the format and content of the video script content text in the second script content.

[0233] S14d: Determine the bonus score based on the overall format score, thinking format score, thinking content score, script format score, and script content score.

[0234] The overall format score measures whether the second output content includes both script generation thinking and script generation output. The thinking format score measures whether the second output content includes the three keywords "image analysis," "error analysis," and "script integration," and whether the second image analysis text includes the four keywords "image identifier," "scene," "character," and "dialogue." The thinking content score measures the quality of the second image analysis text, second error analysis text, and second script integration text based on the image analysis text, error analysis text, and script integration text in the second script content. The script format score measures whether the second output script includes the five keywords "scene," "time," "character," "atmosphere," and "dialogue." The script content score measures the quality of the second output script based on the video script content text in the second script content.

[0235] The reward score includes: an overall format score for the second output content determined by the overall format of the second script content, which facilitates the learning of the overall format of the second script content by the target generation model optimized by the reward score; a thinking format score and a thinking content score for the second image analysis text, the second error analysis text, and the second script integration text determined by the format and content of the image analysis content text, the error analysis content text, and the script integration content text in the second script content, which facilitates the learning of the thinking format and thinking content of the second script content by the target generation model optimized by the reward score; and a script format score and a script content score for the second output script determined by the format and content of the video script content text in the second script content, which facilitates the learning of the script format and script content of the second script content by the target generation model optimized by the reward score. In this way, when optimizing the target generation model, it not only learns the format and content of the required output content, but also focuses on learning the format first, and then focuses on learning the content.

[0236] As an example, the reward points, with a maximum of 100 points, can be divided as follows:

[0237] 1. Overall score format, maximum score is 30 points.

[0238] If the second output includes <think>…<\think>…, the overall format score is 30 points, excluding… <think>…<think>…, the overall format score is 0 points.

[0239] 2. Format score: 20 points; Content score: 20 points.

[0240] (1) The total score for the overall format score is 15 points.

[0241] If the second output includes the keywords "Image Analysis: ... Error Analysis: ... Script Integration:", the overall format score is 15 points. If the second output does not include the keywords "Image Analysis: ... Error Analysis: ... Script Integration:", the overall format score is 0 points. Whether the second output includes the keywords "Image Analysis: ... Error Analysis: ... Script Integration:" is determined using regular expression matching.

[0242] (2) Image analysis format score, full marks are 5 points; image analysis content score, full marks are 10 points.

[0243] Image analysis format score: If the number of second video images in multiple frames is 'a', and the second output content includes 'b' sets of keywords "image identifier: ... scene: ... character: ... dialogue: ...", the image analysis format score is (b ÷ a) × 5 points; if the second output content does not include the keywords "image identifier: ... scene: ... character: ... dialogue: ...", the image analysis format score is 0 points, i.e., b = 0. Whether the second output content includes "image identifier: ... scene: ... character: ... dialogue: ..." is determined using regular expression matching.

[0244] Image analysis content score: Determine the weights w1 for image identifiers, w2 for scenes, w3 for characters, and w4 for dialogue in the second image analysis text; based on the image analysis content text in the second script content, determine the accuracy rates r1 for image identifier content, r2 for scene content, r3 for character content, and r4 for dialogue content in the second image analysis text; the image analysis content score is r1×(w1×10) + r2×(w2×10) + r3×(w3×10) + r4×(w4×10). The accuracy rate of image identifier content (character content, dialogue content) is determined based on whether the image identifier content (character content, dialogue content) in the second image analysis text completely matches the image identifier content (character content, dialogue content) included in the image analysis content text of the second script content; the accuracy rate of scene content is determined based on the similarity between the scene content in the second image analysis text and the scene content included in the image analysis content text of the second script content.

[0245] (3) Error analysis score, with a maximum score of 8 points.

[0246] The correct clauses are determined by analyzing the error analysis content in the second script using a large language model. The accuracy rate r5 is determined based on the number of correct clauses and the total number of clauses. The error analysis content score is r5 × 5.

[0247] (4) The score for the integrated content of the script is 2 points.

[0248] The correct clauses are determined by using a large language model to identify each clause in the integrated text of the second script. The correctness is determined by the number of correct clauses and the total number of clauses. The error analysis score is r6 × 2.

[0249] 3. Script format score, maximum 15 points; Script content score, maximum 15 points.

[0250] Script Format Score: If the second output script includes "Scene:…\nTime:…\nCharacters:…\nAtmosphere:…\nDialogue:…", the script format score is 15 points; if the second output script does not include "Scene:…\nTime:…\nCharacters:…\nAtmosphere:…\nDialogue:…", the script format score is 0 points. Whether the second output script includes "Scene:…\nTime:…\nCharacters:…\nAtmosphere:…\nDialogue:…" is determined using regular expression matching.

[0251] Script content score: Based on the similarity between the second output script and the video script content text in the second script content, the normalized correct score of the second script content is determined. The script content score is score × 15 points.

[0252] In this embodiment, the process of obtaining the second video segment in S12 is as follows: for the second sample video randomly selected based on the second candidate video set, multiple frames of second key images are obtained by first extracting frames from the video and then identifying key frames; considering whether each pair of adjacent second key images is similar in plot, the multiple frames of second key images are divided into multiple plot segments as multiple second plot segments; one of the multiple second plot segments is taken as the second video segment. Based on this, this application provides a possible implementation method, and the step of obtaining the second video segment in S12 may include, for example, the following S16-S18 (not shown in the figure).

[0253] S16: Perform video frame extraction and keyframe recognition on the second sample video to obtain multiple frames of second key images; the second sample video is randomly selected from the second candidate video set.

[0254] S17: Based on the background similarity between any two adjacent second key images, perform plot segmentation on multiple second key images to obtain multiple second plot segments.

[0255] S18: Obtain a second video clip from multiple second plot segments.

[0256] The second candidate video set is a video set formed by multiple second videos to be trained; the second sample video is a video randomly selected from the second candidate video set; the second key image is a key image obtained by extracting frames from the second sample video; the background similarity between any two adjacent second key images is the similarity of any two adjacent second key images in terms of plot; each second plot segment includes multiple second key images of the same plot forming a video segment; different second plot segments are video segments of different plots in multiple second key images.

[0257] S17 may include, for example, S17a-S17e (not shown in the figure).

[0258] S17a: Remove the foreground from the second key image of each frame to obtain the background image of the second key image of each frame.

[0259] S17b: Extract features from the background image of the second key image in each frame to obtain the background features of the second key image in each frame.

[0260] S17c: Determine the plot segmentation points in multiple frames of second key images based on the background feature similarity between any two adjacent frames of second key images.

[0261] If the background feature similarity of two adjacent frames of second key images is greater than or equal to a preset similarity, it means that the two adjacent frames of second key images belong to the same plot. If the background feature similarity of two adjacent frames of second key images is less than the preset similarity, it means that the two adjacent frames of second key images belong to different plots. Based on the two adjacent frames of second key images whose background feature similarity is less than the preset similarity, plot segmentation points in multiple frames of second key images are determined.

[0262] S17d: If the plot segmentation point in the second key image of multiple frames is within a dialogue time corresponding to the second key image of multiple frames, move the plot segmentation point to the end point of the dialogue to obtain the updated plot segmentation point in the second key image of multiple frames.

[0263] S17e: Based on the plot segmentation points in the updated multi-frame second key image, divide the multi-frame second key image into multiple second plot segments.

[0264] In scenarios where video script generation is achieved through a model, this method takes into account the model's processing capability for video image frames. After randomly selecting a second sample video from the second candidate video set, multiple second key images are obtained through video frame extraction and key frame recognition, achieving sparse frame extraction and key recognition for the video. Furthermore, considering whether multiple second key images represent the same plot, the similarity between any two adjacent second key images indicates whether they are similar in plot. The multiple second key images are then segmented into multiple second plot segments, thus dividing the video images after sparse frame extraction and key recognition into different plots. In this way, the second video segment to be generated by the model can be obtained from the multiple second plot segments.

[0265] In this embodiment of the application, the second script content in S14 can be obtained by first inputting the second video clip into the script generation model to generate a second script, and then revising the second script. Based on this, the application provides a possible implementation method, and the step of obtaining the second script content in S14 may include, for example, the following S19-S20 (not shown in the figure).

[0266] S19: Generate a script for the second video clip using a script generation model to obtain a second generated script for the second video clip.

[0267] S20: Revise the second generated script to obtain the content of the second script.

[0268] The second generated script is the script output by the script generation model after the second video clip is input into it.

[0269] This method inputs the second video clip into the script generation model. Based on the powerful generation capabilities of the script generation model, it can quickly and automatically generate a second generated script for the second video clip to represent the second video clip in a script format. The second generated script is revised through script revision, realizing image analysis, error analysis, script integration, and script output. This ensures that the content of the second script accurately represents the image analysis content, error analysis content, script integration content, and script output content of the second video clip in a standard script format.

[0270] In addition, the base models for the initial generation model, the target generation model, or the optimized target generation model include a visual encoder and a text processing model. Among them, the visual encoder adopts a redesigned Vision Transformer (ViT) architecture, incorporating a two-dimensional rotational position encoding mechanism and a window attention mechanism. During training and inference, the height and width of the image are adjusted to multiples of 28 before being input into ViT, which segments the image into blocks with a stride of 14 to generate image features.

[0271] Furthermore, instead of directly using the original block features generated by ViT, the four spatially adjacent block features are grouped, concatenated, and then projected onto a dimension aligned with the text embeddings in the large language model upon which the text processing model is based using two layers of multilayer perceptron (MLP). This reduces computational costs and dynamically compresses image feature sequences of different lengths, facilitating their fusion with text features.

[0272] The text processing model is based on a large language model, which is initialized using its pre-trained weights. The one-dimensional rotational positional encoding is modified to a multimodal rotational positional encoding aligned to absolute time, which enables visual and textual information to be fused under a unified positional encoding system, thus improving visual-language compatibility.

[0273] In this embodiment, based on the optimized target generation model's ability to identify the scene, characters, and dialogue in each frame of a video segment formed from multiple video images, based on corresponding character feature information, and to analyze errors and integrate the script to output a script text that more accurately represents the video segment and conforms to the script format, in the specific implementation of S302 above, the target generation model generates a script for the target video segment based on the target character feature information and the first prompt text to obtain the video-generated script for the target video segment. Specifically, the target character feature information and the target video segment are input into the optimized target generation model. Guided by the first prompt text, the model analyzes the scene, characters, and dialogue in each frame of the target video image based on the target character feature information, analyzes errors, and integrates the script to output a script text that conforms to the script format, thus obtaining the video-generated script for the target video segment. Based on this, this application provides a possible implementation method. For example, S302 above may include S302e (not shown in the figure): generating a script for the target video segment based on the target character feature information and the first prompt text using the optimized target generation model to obtain the video-generated script.

[0274] This method uses the target character feature information and the target video clip as model inputs to the optimized target generation model. Guided by the first prompt text, the optimized target generation model outputs script text that is more in line with the script format based on the target character feature information for the target video clip, thus obtaining the video-generated script for the target video clip. This further ensures the standardization of the script format and the accuracy of the content of the video-generated script, thereby further improving the script generation accuracy.

[0275] See Figure 12 ,Should Figure 12 This is a schematic diagram illustrating the generation of a script from a video using an optimized target generation model, as provided in this application embodiment. After acquiring a target video segment 701 formed from multiple frames of target video images and corresponding target character feature information 702, the target video segment 701 and the corresponding target character feature information 702 are input into the optimized target generation model 1107. Guided by a first prompt text 704, the optimized target generation model 1107 analyzes the scene, characters, and dialogue in each frame of the target video segment 701 based on the target character feature information 702, and analyzes errors and integrates the script to output script text conforming to the script format, thus obtaining a video-generated script 705 for the target video segment.

[0276] In summary, the video script generation method provided in this application integrates character recognition into the script generation process, addressing the limited character information corresponding to video clips. This eliminates the need for pre-identification, avoiding high script generation latency and enabling efficient script generation. The model uses a thought-connected approach to perform image analysis, error analysis, and script integration, outputting a script that conforms to the script format. This ensures accurate script generation and guarantees that the script's representation of the video clip's scene, time, characters, atmosphere, and dialogue is directly usable. Leveraging the model's powerful generation capabilities, combined with supervised fine-tuning training of video clip-script content, the model can be quickly trained to efficiently and accurately generate scripts from videos. Furthermore, by incorporating reinforcement learning training based on format and content of video clip-script content, the model outputs scripts with more standardized format and more accurate content.

[0277] It should be noted that, based on the implementation methods provided in the above aspects, this application can be further combined to provide more implementation methods.

[0278] See Figure 13 ,Should Figure 13 This is a schematic diagram of an application product interface in a video script generation scenario provided by an embodiment of this application. The application product interface can be a video analysis interface, which displays the main plot content, character settings, script reading content, and script value analysis content of the video. The script reading content is abridged and can specifically display the script text content and scene analysis content under each scene identifier. The script text content is obtained through the video script generation method provided by an embodiment of this application.

[0279] In this process, after generating a video script from a previous video segment, if the current video segment has the same scene identifier as a historical video segment, the video script generated for the current video segment needs to be merged into the video script generated for the historical video segment.

[0280] The following describes in detail a video script generation apparatus provided in this application through an embodiment of the apparatus.

[0281] based on Figure 2 In accordance with the video script generation method provided in the corresponding embodiments, this application also provides a video script generation apparatus, see [link to relevant documentation]. Figure 14 ,Should Figure 14 This is a structural diagram of a video script generation device provided in an embodiment of the present application. The video script generation device 1400 includes: an acquisition unit 1401 and a generation unit 1402.

[0282] The acquisition unit 1401 is used to acquire a target video segment and the target character feature information corresponding to the target video segment; the target video segment includes multiple frames of target video images;

[0283] The generation unit 1402 is used to generate a script for a target video segment by using a target generation model based on the target character feature information and the first prompt text, thereby obtaining a video-generated script for the target video segment. The first prompt text is used to guide the target generation model to analyze the scene, characters and dialogue in each frame of the target video image based on the target character feature information, and to analyze errors and integrate the script to output script text that conforms to the script format.

[0284] In one possible implementation, the generating unit 1402 is specifically used for:

[0285] Using the target generation model, based on the target character feature information and the first prompt text, the target video clip is analyzed to generate a script, resulting in image analysis output text, error analysis output text, and script integration output text. The image analysis output text is used to represent the scene, characters, and dialogue in each frame of the target video image.

[0286] Based on the image analysis output text, error analysis output text, and script integration output text, a script is generated for the target video clip to obtain a video-generated script. The video-generated script is used to represent the scene, time, characters, atmosphere, and dialogue in the target video clip.

[0287] In one possible implementation, the device 1400 further includes: a determining unit;

[0288] Determine the unit, used for:

[0289] The target character feature information is determined by either the target character's makeup image or the target character's descriptive text corresponding to the target video clip.

[0290] In one possible implementation, the acquisition unit 1401 is also used for:

[0291] Obtain the target dialogue text corresponding to the target video segment;

[0292] Generation unit 1402 is specifically used for:

[0293] Using a target generation model, a script is generated for the target video segment based on the target character feature information, target dialogue text, and first prompt text, thus obtaining a video-generated script.

[0294] In one possible implementation, the acquisition unit 1401 is also used for:

[0295] Obtain the historical video script corresponding to the target video segment;

[0296] Generation unit 1402 is specifically used for:

[0297] Using a target generation model, scripts are generated for target video segments based on target character feature information, historical video scripts, and initial prompt text, thus obtaining a video-generated script.

[0298] In one possible implementation, the device 1400 further includes: a training unit;

[0299] Training units are used for:

[0300] Obtain the first video segment and the first character feature information corresponding to the first video segment; the first video segment includes multiple frames of first video images;

[0301] The initial generation model generates a script for the first video segment based on the first character feature information and the second prompt text, and obtains the first output content of the first video segment. The second prompt text is used to guide the initial generation model to analyze the scene, characters and dialogue in each frame of the first video image based on the first character feature information, and to analyze errors and integrate the script to output script text that conforms to the script format.

[0302] Based on the difference between the first output content and the first script content of the first video clip, the initial generation model is trained to obtain the target generation model; the first script content includes image analysis content text, error analysis content text, script integration content text, and video script content text based on the first video clip.

[0303] In one possible implementation, the training unit is specifically used for:

[0304] Using the initial generation model, based on the first character feature information and the second prompt text, the script generation thinking is performed on the first video segment to obtain the first image analysis text, the first error analysis text, and the first script integration text; the first image analysis text is used to represent the scene, characters, and dialogue in each frame of the first video image.

[0305] Based on the first image analysis text, the first error analysis text, and the first script integration text, a script is generated and output for the first video segment to obtain the first output script.

[0306] The first image analysis text, the first error analysis text, the first script integration text, and the first output script are identified as the first output content.

[0307] In one possible implementation, the training unit is specifically used for:

[0308] Determine the sum of probabilities that multiple predicted terms in the first output content correspond to multiple first terms in the first script content;

[0309] The target generative model is obtained by maximizing the probability and training the initial generative model.

[0310] In one possible implementation, the device 1400 further includes: a first obtaining unit;

[0311] The first acquisition unit is used for:

[0312] The first sample video is subjected to video frame extraction and keyframe recognition to obtain multiple first key images; the first sample video is randomly selected from the first candidate video set.

[0313] Based on the background similarity between any two adjacent first key images, multiple first key images are segmented into plot segments to obtain multiple first plot fragments.

[0314] Obtain the first video clip from multiple first-episode segments.

[0315] In one possible implementation, the device 1400 further includes: a second obtaining unit;

[0316] The second obtaining unit is used for:

[0317] The script generation model is used to generate a script for the first video clip, resulting in the first generated script for the first video clip.

[0318] Revise the first generated script to obtain the content of the first script.

[0319] In one possible implementation, the device 1400 further includes an optimization unit;

[0320] Optimization unit, used for:

[0321] Obtain the second video segment and the second character feature information corresponding to the second video segment; the second video segment includes multiple frames of second video images;

[0322] The target generation model generates a script for the second video clip based on the second character feature information and the third prompt text, and obtains the second output content of the second video clip. The third prompt text is used to guide the target generation model to analyze the scene, characters and dialogue in each frame of the second video image based on the second character feature information, and to analyze errors and integrate the script to output script text that conforms to the script format.

[0323] The reward score is determined based on the second output content and the second script content of the second video clip; the second script content includes image analysis content text based on the second video clip, error analysis content text, script integration content text, and video script content text;

[0324] The target generation model is optimized based on the reward score to obtain the optimized target generation model.

[0325] In one possible implementation, the optimization unit is specifically used for:

[0326] Using a target generation model, based on the second character feature information and the third prompt text, a script generation process is performed on the second video segment to obtain the second image analysis text, the second error analysis text, and the integrated second script text. The second image analysis text is used to represent the scene, characters, and dialogue in each frame of the second video image.

[0327] Based on the second image analysis text, the second error analysis text, and the second script integration text, a script is generated and output for the second video segment to obtain the second output script.

[0328] The second image analysis text, the second error analysis text, the second script integration text, and the second output script are identified as the second output content.

[0329] In one possible implementation, the optimization unit is specifically used for:

[0330] The overall format score is determined based on the overall format of the second output content and the overall format of the second script content.

[0331] Based on the format and content of the second image analysis text, the second error analysis text, and the second script integration text, as well as the format and content of the image analysis content text, error analysis content text, and script integration content text in the second script content, determine the thinking format score and the thinking content score;

[0332] Based on the format and content of the second output script, as well as the format and content of the video script text within the second script content, determine the script format score and script content score;

[0333] The bonus score is determined based on the overall format score, the thinking format score, the thinking content score, the script format score, and the script content score.

[0334] In one possible implementation, the device 1400 further includes: a third obtaining unit;

[0335] The third acquisition unit is used for:

[0336] The second sample video is subjected to video frame extraction and keyframe recognition to obtain multiple frames of second key images; the second sample video is randomly selected from the second candidate video set.

[0337] Based on the background similarity between any two adjacent second key images, multiple frames of second key images are segmented into multiple second plot segments.

[0338] Obtain the second video clip from multiple second plot segments.

[0339] In one possible implementation, the device 1400 further includes: a fourth obtaining unit;

[0340] The fourth acquisition unit is used for:

[0341] The script generation model is used to generate a script for the second video clip, resulting in a second generated script for the second video clip.

[0342] Revise the second generated script to obtain the content of the second script.

[0343] In one possible implementation, the generating unit 1402 is specifically used for:

[0344] The optimized target generation model generates a script for the target video segment based on the target character feature information and the first prompt text, thus obtaining the video-generated script.

[0345] As can be seen from the above technical solution, a target video segment is formed by acquiring multiple frames of target video images, and the target character feature information corresponding to the target video segment is obtained. This target character feature information can be provided to the model to identify the characters in each frame of the target video image when generating the video script. The target character feature information and the target video segment are input into the target generation model. Guided by the first prompt text, the model analyzes the scene, characters, and dialogue in each frame of the target video image based on the target character feature information. It is not necessary to traverse each frame of the target video image in advance to achieve character recognition. Instead, character recognition is achieved when the model generates the script, reducing the character recognition time. Furthermore, it analyzes errors and integrates the script, improving the accuracy of character and dialogue analysis through thought chain, achieving accurate understanding of the video segment, and outputting script text that conforms to the script format. This results in a video-generated script for the target video segment, ensuring that the video-generated script accurately represents the target video segment in a standard script format, thereby improving script generation efficiency and accuracy.

[0346] This application also provides a computer device, which may be a server, see [link to previous document]. Figure 15 ,Should Figure 15 This application provides a structural diagram of a server. The server 1500 can vary significantly due to different configurations or performance. It may include one or more processors, such as a central processing unit (CPU) 1522, and a memory 1532, as well as one or more storage media 1530 (e.g., one or more mass storage devices) for storing application programs 1542 or data 1544. The memory 1532 and storage media 1530 can be temporary or persistent storage. The program stored in the storage media 1530 may include one or more modules (not shown in the diagram), each module including a series of instruction operations on the server. Furthermore, the CPU 1522 may be configured to communicate with the storage media 1530 and execute the series of instruction operations stored in the storage media 1530 on the server 1500.

[0347] Server 1500 may also include one or more power supplies 1526, one or more wired or wireless network interfaces 1550, one or more input / output interfaces 1558, and / or one or more operating systems 1541, such as Windows Server. TM Mac OS X TM Unix TM Linux TM FreeBSD TM etc.

[0348] In this embodiment, the central processing unit 1522 in the server 1500 can execute the methods provided in the various optional implementations of the above embodiments.

[0349] The computer device provided in this application embodiment can also be a terminal, see [link to relevant documentation]. Figure 16 ,Should Figure 16 This is a structural diagram of a terminal provided in an embodiment of this application. Taking a smartphone as an example, the smartphone includes: a radio frequency (RF) circuit 1610, a memory 1620, an input unit 1630, a display unit 1640, a sensor 1650, an audio circuit 1660, a wireless Fidelity (WiFi) module 1670, a processor 1680, and a power supply 1690, etc. The input unit 1630 may include a touch panel 1631 and other input devices 1632, the display unit 1640 may include a display panel 1641, and the audio circuit 1660 may include a speaker 1661 and a microphone 1662. Those skilled in the art will understand that... Figure 16 The smartphone structure shown does not constitute a limitation on smartphones and may include more or fewer components than shown, or combine certain components, or have different component arrangements.

[0350] The memory 1620 can be used to store software programs and modules. The processor 1680 executes various functions and data processing of the smartphone by running the software programs and modules stored in the memory 1620. The memory 1620 may mainly include a program storage area and a data storage area. The program storage area may store the operating system, applications required for at least one function (such as sound playback function, image playback function, etc.), etc.; the data storage area may store data created according to the use of the smartphone (such as audio data, phonebook, etc.). In addition, the memory 1620 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device.

[0351] The processor 1680 is the control center of the smartphone, connecting various parts of the smartphone via various interfaces and lines. It performs various functions and processes data by running or executing software programs and / or modules stored in the memory 1620 and by accessing data stored in the memory 1620. Optionally, the processor 1680 may include one or more processing units; preferably, the processor 1680 may integrate an application processor and a modem processor, wherein the application processor mainly handles the operating system, user interface, and applications, and the modem processor mainly handles wireless communication. It is understood that the modem processor may not be integrated into the processor 1680.

[0352] In this embodiment, the processor 1680 in the smartphone can execute the methods provided in the various optional implementations of the above embodiments.

[0353] According to one aspect of this application, a computer-readable storage medium is provided for storing a computer program that, when run on a computer device, causes the computer device to perform the methods provided in various optional implementations of the above embodiments.

[0354] According to one aspect of this application, a computer program product is provided, comprising a computer program stored in a computer-readable storage medium. A processor of a computer device reads the computer program from the computer-readable storage medium and executes the computer program, causing the computer device to perform the methods provided in various optional implementations of the above embodiments.

[0355] The descriptions of the processes or structures corresponding to the above figures each have their own emphasis. For parts of a process or structure that are not described in detail, please refer to the relevant descriptions of other processes or structures.

[0356] The terms "first," "second," etc., used in this application's specification and the foregoing drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented, for example, in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0357] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be an indirect coupling or communication connection between apparatuses or units through some interfaces, and may be electrical, mechanical, or other forms.

[0358] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0359] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0360] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing computer programs, such as USB flash drives, portable hard drives, read-only memory (ROM), RAM, magnetic disks, or optical disks.

[0361] In the embodiments of this application, the terms "module" or "unit" refer to a computer program or part of a computer program that has a predetermined function and works with other related parts to achieve a predetermined goal, and can be implemented wholly or partially using software, hardware (such as processing circuitry or memory), or a combination thereof. Similarly, a processor (or multiple processors or memory) can be used to implement one or more modules or units. Furthermore, each module or unit can be part of an overall module or unit that includes the functionality of that module or unit.

[0362] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.< / think> < / think> < / think> < / think>

Claims

1. A method for generating video scripts, characterized in that, The method includes: Acquire the target video segment and the target character feature information corresponding to the target video segment; the target video segment includes multiple frames of target video images; Using a target generation model, a script is generated for the target video segment based on the target character feature information and the first prompt text, resulting in a video-generated script for the target video segment. The first prompt text guides the target generation model to analyze the scene, characters, and dialogue in each frame of the target video image based on the target character feature information, and to analyze errors and integrate the script to output script text that conforms to the script format.

2. The method according to claim 1, characterized in that, The step of generating a script for the target video segment using a target generation model, based on the target character feature information and the first prompt text, to obtain the video-generated script for the target video segment includes: Using the target generation model, based on the target character feature information and the first prompt text, script generation thinking is performed on the target video segment to obtain image analysis output text, error analysis output text, and script integration output text; the image analysis output text is used to represent the scene, character, and dialogue in each frame of the target video image. Based on the image analysis output text, the error analysis output text, and the script integration output text, a script is generated for the target video segment to obtain the video-generated script; the video-generated script is used to represent the scene, time, characters, atmosphere, and dialogue in the target video segment.

3. The method according to claim 1, characterized in that, The steps for determining the target character feature information include: The target character feature information is determined by at least one of the target character makeup image or target character description text corresponding to the target video segment.

4. The method according to claim 1, characterized in that, The method further includes: Obtain the target dialogue text corresponding to the target video segment; The step of generating a script for the target video segment using a target generation model, based on the target character feature information and the first prompt text, to obtain the video-generated script for the target video segment includes: Using the target generation model, the target video segment is scripted based on the target character feature information, the target dialogue text, and the first prompt text, thereby obtaining the video-generated script.

5. The method according to claim 1, characterized in that, The method further includes: Obtain the historical video script corresponding to the target video segment; The step of generating a script for the target video segment using a target generation model, based on the target character feature information and the first prompt text, to obtain the video-generated script for the target video segment includes: Using the target generation model, the target video segment is scripted based on the target character feature information, the historical video script, and the first prompt text, thus obtaining the video-generated script.

6. The method according to claim 1, characterized in that, The training steps of the target generation model include: Obtain a first video segment and the first character feature information corresponding to the first video segment; the first video segment includes multiple frames of first video images; The initial generation model generates a script for the first video segment based on the first character feature information and the second prompt text, thereby obtaining the first output content of the first video segment. The second prompt text is used to guide the initial generation model to analyze the scene, characters and dialogue in each frame of the first video image based on the first character feature information, and to analyze errors and integrate the script to output script text that conforms to the script format. Based on the difference between the first output content and the first script content of the first video segment, the initial generation model is trained to obtain the target generation model; the first script content includes image analysis content text, error analysis content text, script integration content text, and video script content text based on the first video segment.

7. The method according to claim 6, characterized in that, The step of generating a script for the first video segment using an initial generation model, based on the first character feature information and the second prompt text, to obtain the first output content of the first video segment includes: Using the initial generation model, based on the first character feature information and the second prompt text, script generation thinking is performed on the first video segment to obtain the first image analysis text, the first error analysis text, and the first script integration text; the first image analysis text is used to represent the scene, character, and dialogue in each frame of the first video image. Based on the first image analysis text, the first error analysis text, and the first script integration text, a script is generated and output for the first video segment to obtain the first output script. The first image analysis text, the first error analysis text, the first script integration text, and the first output script are determined as the first output content.

8. The method according to claim 6, characterized in that, The step of training the initial generation model based on the difference between the first output content and the first script content of the first video segment to obtain the target generation model includes: The sum of probabilities that multiple predicted words in the first output content are the corresponding multiple first words in the first script content; The target generative model is obtained by maximizing the probability and training the initial generative model.

9. The method according to claim 6, characterized in that, The steps for obtaining the first video segment include: The first sample video is subjected to video frame extraction and keyframe recognition to obtain multiple first key images; the first sample video is randomly extracted from the first candidate video set. Based on the background similarity between any two adjacent first key images, the multiple first key images are segmented into multiple first plot segments. The first video segment is obtained from the plurality of first plot segments.

10. The method according to claim 6, characterized in that, The steps for obtaining the content of the first script include: The script generation model is used to generate a script for the first video segment, thereby obtaining a first generated script for the first video segment. Revise the first generated script to obtain the content of the first script.

11. The method according to any one of claims 6-10, characterized in that, The method further includes: Obtain the second video segment and the second character feature information corresponding to the second video segment; the second video segment includes multiple frames of second video images; The target generation model generates a script for the second video segment based on the second character feature information and the third prompt text, thereby obtaining the second output content of the second video segment. The third prompt text is used to guide the target generation model to analyze the scene, characters and dialogue in each frame of the second video image based on the second character feature information, and to analyze errors and integrate the script to output script text that conforms to the script format. The reward score is determined based on the second output content and the second script content of the second video clip; the second script content includes image analysis content text, error analysis content text, script integration content text, and video script content text based on the second video clip; The target generation model is optimized based on the reward score to obtain an optimized target generation model.

12. The method according to claim 11, characterized in that, The step of generating a script for the second video segment using the target generation model, based on the second character feature information and the third prompt text, to obtain the second output content of the second video segment, includes: Using the target generation model, based on the second character feature information and the third prompt text, script generation thinking is performed on the second video segment to obtain second image analysis text, second error analysis text, and second script integration text; the second image analysis text is used to represent the scene, characters, and dialogue in each frame of the second video image. Based on the second image analysis text, the second error analysis text, and the second script integration text, a script is generated and output for the second video segment to obtain the second output script; The second image analysis text, the second error analysis text, the second script integration text, and the second output script are determined as the second output content.

13. The method according to claim 12, characterized in that, The step of determining the reward score based on the second script content of the second output content and the second video clip includes: The overall format score is determined based on the overall format of the second output content and the overall format of the second script content; Based on the format and content of the second image analysis text, the second error analysis text, and the second script integration text, as well as the format and content of the image analysis content text, error analysis content text, and script integration content text in the second script content, determine the thinking format score and the thinking content score; Based on the format and content of the second output script, and the format and content of the video script content text in the second script content, determine the script format score and script content score; The reward score is determined based on the overall format score, the thinking format score, the thinking content score, the script format score, and the script content score.

14. The method according to claim 11, characterized in that, The steps for obtaining the second video segment include: The second sample video is subjected to video frame extraction and keyframe recognition to obtain multiple frames of second key images; the second sample video is randomly extracted from the second candidate video set. Based on the background similarity between any two adjacent frames of the second key image, the multiple frames of the second key image are segmented into multiple second plot segments. The second video segment is obtained from the plurality of second plot segments.

15. The method according to claim 11, characterized in that, The steps for obtaining the second script content include: The script generation model is used to generate a script for the second video segment, resulting in a second generated script for the second video segment. Revise the second generated script to obtain the content of the second script.

16. The method according to claim 11, characterized in that, The step of generating a script for the target video segment using a target generation model, based on the target character feature information and the first prompt text, to obtain the video-generated script for the target video segment includes: Using the optimized target generation model, the target video segment is scripted based on the target character feature information and the first prompt text to obtain the video-generated script.

17. A video script generation device, characterized in that, The device includes: an acquisition unit and a generation unit; The acquisition unit is used to acquire a target video segment and target character feature information corresponding to the target video segment; the target video segment includes multiple frames of target video images; The generation unit is used to generate a script for the target video segment by using a target generation model based on the target character feature information and the first prompt text, thereby obtaining a video-generated script for the target video segment; the first prompt text is used to guide the target generation model to analyze the scene, characters and dialogue in each frame of the target video image based on the target character feature information, and to analyze errors and integrate the script to output script text that conforms to the script format.

18. A computer device, characterized in that, The computer device includes a processor and memory: The memory is used to store computer programs and to transfer the computer programs to the processor; The processor is configured to execute the method according to any one of claims 1-16 according to instructions in the computer program.

19. A computer-readable storage medium, characterized in that, The computer-readable storage medium is used to store a computer program that, when run on a computer device, causes the computer device to perform the method according to any one of claims 1-16.

20. A computer program product, comprising a computer program, characterized in that, When the computer program is run on a computer device, it causes the computer device to perform the method according to any one of claims 1-16.