Scenario generation method and related device
By strategizing the original video and tracking and marking the character, a target script containing characters is generated, which solves the problem of direct generation from video to script in the existing technology, and achieves high-quality and consistent script generation.
Patent Information
- Application Number
- CN202510496938.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-21
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2045-04-21
AI Technical Summary
The existing script generation methods are mainly the conversion from text to script. They lack direct generation methods from video to script, and it is difficult to effectively utilize the scenes and character information in the video.
By strategizing the original video, multiple strategizing videos are obtained, and then the characters in these storyboard videos are tracked and marked, multiple character marking videos are generated, and the videos are finally described in a scenario to generate a target script containing the characters.
It improves the accuracy and reliability of character tracking and marking, enhances the accuracy of scenario description, thereby improving the quality of the target script and consistency with the original video.
Smart Images

Figure CN120032374A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of video processing technology, and in particular to a script generation method and related devices. Background Art
[0002] In some scenarios, it is necessary to generate a complete script based on a given video. For example, a corresponding script is generated based on a short video published on a short video creation platform for use in short video secondary creation, screenwriting and other fields.
[0003] However, current script generation methods are basically text-to-script generation methods, and few involve video-to-script generation methods. Summary of the invention
[0004] In view of the above problems, the present application provides a script generation method and related devices to achieve the purpose of generating a target script based on the original video. The specific scheme is as follows:
[0005] The first aspect of the present application provides a script generation method, comprising:
[0006] Get the original video to be processed;
[0007] Performing storyboard processing on the original video to obtain multiple storyboard videos;
[0008] Tracking and marking characters in the multiple storyboard videos to obtain multiple character-marked videos;
[0009] The multiple character-labeled videos are processed for scenario description to obtain a target script containing the characters.
[0010] In a possible implementation, the performing storyboard processing on the original video to obtain a plurality of storyboard videos includes:
[0011] Detecting scene change information in the original video;
[0012] According to the scene change information, the original video is divided into the multiple storyboard videos, wherein the multiple storyboard videos each belong to a different scene.
[0013] In a possible implementation, tracking and marking characters in the multiple storyboard videos to obtain multiple character marked videos includes:
[0014] Tracking characters in the plurality of storyboard videos respectively to obtain character trajectories corresponding to the plurality of storyboard videos respectively;
[0015] According to the character trajectories respectively corresponding to the multiple storyboard videos, the characters in the multiple storyboard videos are marked with characters to obtain the multiple character marked videos.
[0016] In a possible implementation, the character-labeling of characters in the multiple storyboard videos according to the character trajectories respectively corresponding to the multiple storyboard videos to obtain the multiple character-labeled videos includes:
[0017] Obtaining a preconfigured character face mapping table, wherein the character face mapping table includes a correspondence between the character identifications and facial images of all characters in the original video;
[0018] For each storyboard video in the plurality of storyboard videos:
[0019] According to the character trajectory corresponding to the storyboard video, a facial image of each character is obtained from the storyboard video;
[0020] Performing similarity matching between the facial image of each character and the facial images in the character face mapping table to obtain a matching result, and determining the role identification of each character according to the matching result;
[0021] Mark each character in the storyboard video according to the character identification of each character, and obtain a character role-marked video corresponding to the storyboard video;
[0022] The character marking videos corresponding to the multiple storyboard videos are obtained as the multiple character marking videos.
[0023] In a possible implementation, performing scenario description processing on the multiple character-labeled videos to obtain a target script containing the characters includes:
[0024] Performing situation understanding on the plurality of character-labeled videos respectively to obtain video content description information corresponding to the plurality of character-labeled videos respectively, wherein the video content description information includes role identifiers of characters in the corresponding character-labeled videos;
[0025] The video content description information corresponding to the multiple character marking videos is integrated to obtain the target script.
[0026] In a possible implementation, performing context understanding on the multiple character-labeled videos to obtain video content description information corresponding to the multiple character-labeled videos respectively includes:
[0027] Acquire a first text prompt word, wherein the first text prompt word includes basic script elements required to generate the target script;
[0028] Inputting the plurality of character-labeled videos and the first text prompt word into a pre-trained multimodal large model for video understanding, and obtaining video content description information corresponding to the plurality of character-labeled videos;
[0029] Among them, the training data used by the video understanding multimodal large model in the training stage includes: multiple character labeling training videos and corresponding video content description information, and the multiple character labeling training videos are obtained by tracking and labeling the characters in multiple training storyboard videos.
[0030] In a possible implementation, the step of integrating the video content description information corresponding to the plurality of character-labeled videos to obtain the target script includes:
[0031] Obtaining a second text prompt word, wherein the second text prompt word is used to prompt the large language model to integrate the video content description information corresponding to the multiple character labeling videos respectively;
[0032] The video content description information corresponding to the multiple character-labeled videos and the second text prompt words are input into the large language model to obtain the target script output by the large language model.
[0033] A second aspect of the present application provides a script generation device, comprising:
[0034] A video acquisition module is used to acquire the original video to be processed;
[0035] A video storyboard module, used for storyboarding the original video to obtain multiple storyboard videos;
[0036] A tracking and marking module, used for tracking and marking characters in the plurality of storyboard videos to obtain a plurality of character marking videos;
[0037] The script generation module is used to perform scenario description processing on the multiple character marking videos to obtain a target script containing the characters.
[0038] A third aspect of the present application provides an electronic device, comprising at least one processor and a memory connected to the processor, wherein:
[0039] The memory is used to store computer programs;
[0040] The processor is used to execute the computer program so that the electronic device can implement the script generation method of the above-mentioned first aspect or any implementation method of the first aspect.
[0041] The fourth aspect of the present application provides a computer storage medium, which carries one or more computer programs. When the one or more computer programs are executed by an electronic device, the electronic device can implement the script generation method of the above-mentioned first aspect or any implementation method of the first aspect.
[0042] By means of the above technical solution, the script generation method provided by the present application takes into account that short video clips are more helpful for understanding the video content and character tracking than complete videos. The present application performs storyboard processing on the original video to obtain multiple storyboard videos, and then tracks and marks the characters in the multiple storyboard videos to obtain multiple character role marked videos. Further, the multiple character role marked videos are processed for scene description to obtain a target script containing the character roles. The present application can avoid the loss of tracking targets during character tracking through storyboarding, improve the accuracy and reliability of tracking and marking, and the storyboard makes the content of the character role marked videos more understandable, improves the accuracy of scene description processing, and thus improves the quality of the target script. At the same time, through character tracking and marking, the character roles can be integrated into the target script, further improving the quality of the target script and ensuring the consistency of the target script with the original video. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] The above and other features, advantages and aspects of the embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the accompanying drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and the originals and elements are not necessarily drawn to scale.
[0044] Figure 1 A schematic diagram of a system architecture provided for this application;
[0045] Figure 2 A schematic diagram of an optional hardware structure of the terminal 100 provided in this application;
[0046] Figure 3 A schematic diagram of the structure of a server 200 provided in this application;
[0047] Figure 4 A flowchart of a script generation method provided for this application;
[0048] Figure 5 A schematic diagram of marking a person's head contained in a video frame provided by the present application;
[0049] Figure 6 A schematic diagram of the structure of a script generation device provided in this application;
[0050] Figure 7 A schematic diagram of the structure of an electronic device provided in this application. DETAILED DESCRIPTION
[0051] The following describes the embodiments of the present application in conjunction with the drawings in the embodiments of the present application. The terms used in the implementation method section of the present application are only used to explain the specific embodiments of the present application, and are not intended to limit the present application.
[0052] The embodiments of the present application are described below in conjunction with the accompanying drawings. Those skilled in the art will appreciate that, with the development of technology and the emergence of new scenarios, the technical solutions provided in the embodiments of the present application are also applicable to similar technical problems.
[0053] The terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and need not be used to describe a specific order or sequential order. It should be understood that the terms used in this way can be interchangeable under appropriate circumstances, which is only to describe the distinction mode adopted by the objects of the same attributes when describing in the embodiments of the present application. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, so that the process, method, system, product or equipment comprising a series of units need not be limited to those units, but may include other units that are not clearly listed or inherent to these processes, methods, products or equipment.
[0054] See also Figure 1 , Figure 1 A schematic diagram of a system architecture is shown. The system may include a terminal 100 and a server 200. The server 200 may include one or more servers ( Figure 1 In the example, a server is included, and the server 200 can provide the method provided in the embodiment of the present application for one or more terminals.
[0055] Among them, an application can be installed on the terminal 100, and the above application and web page can provide an interface. The terminal 100 can receive relevant parameters entered by the user on the interface and send the above parameters to the server 200. The server 200 can obtain processing results based on the received parameters and return the processing results to the terminal 100.
[0056] It should be understood that in some optional implementations, the terminal 100 can also complete the action of obtaining the processing result based on the received parameters by itself without the cooperation of the server, and the embodiments of the present application are not limited to this.
[0057] Next describe Figure 1 The product form of the mid-terminal 100;
[0058] The terminal 100 in the embodiment of the present application can be a mobile phone, a tablet computer, a wearable device, a vehicle-mounted device, an augmented reality (AR) / virtual reality (VR) device, a laptop computer, an ultra-mobile personal computer (UMPC), a netbook, a personal digital assistant (PDA), etc., and the embodiment of the present application does not impose any restrictions on this.
[0059] Figure 2 An optional hardware structure diagram of the terminal 100 is shown.
[0060] refer to Figure 2 As shown, the terminal 100 may include a radio frequency unit 110, a first memory 120, an input unit 130, a display unit 140, a camera 150 (optional), an audio circuit 160 (optional), a speaker 161 (optional), a microphone 162 (optional), an earphone jack 163 (optional), a first processor 170, an external interface 180, a power supply 190 and other components. Those skilled in the art will appreciate that Figure 2 These are merely examples of terminals or multi-function devices and do not constitute limitations on the terminals or multi-function devices, which may include more or fewer components than those shown in the figures, or combinations of certain components, or different components.
[0061] The input unit 130 can be used to receive input digital or character information, and generate key signal input related to the user settings and function control of the portable multifunctional device. Specifically, the input unit 130 may include a touch screen 131 (optional) and / or other input devices 132. The touch screen 131 can collect the user's touch operations on or near it (such as the user's operation on or near the touch screen using any suitable object such as fingers, joints, stylus, etc.), and drive the corresponding connection device according to a pre-set program. The touch screen can detect the user's touch action on the touch screen, convert the touch action into a touch signal and send it to the first processor 170, and can receive and execute the command sent by the first processor 170; the touch signal at least includes the touch point coordinate information. The touch screen 131 can provide an input interface and an output interface between the terminal 100 and the user. In addition, the touch screen can be implemented using multiple types such as resistive, capacitive, infrared and surface acoustic wave. In addition to the touch screen 131, the input unit 130 can also include other input devices. Specifically, other input devices 132 may include, but are not limited to, one or more of a physical keyboard, function keys (such as a volume control key, a switch key, etc.), a trackball, a mouse, a joystick, and the like.
[0062] Among them, the input device 132 can receive input data and the like.
[0063] The display unit 140 may be used to display information input by a user or provided to a user, various menus of the terminal 100, an interactive interface, file display, and / or playback of any multimedia file.
[0064] The first memory 120 can be used to store instructions and data. The first memory 120 can mainly include an instruction storage area and a data storage area. The data storage area can store various data, such as multimedia files, texts, etc.; the instruction storage area can store software units such as operating systems, applications, instructions required for at least one function, or their subsets and extensions. It can also include a non-volatile random access memory; provide the first processor 170 with hardware, software and data resources including management computing and processing equipment, and support control software and applications. It is also used for the storage of multimedia files, and the storage of running programs and applications.
[0065] The first processor 170 is the control center of the terminal 100. It uses various interfaces and lines to connect various parts of the entire terminal 100. By running or executing instructions stored in the first memory 120 and calling data stored in the first memory 120, it executes various functions of the terminal 100 and processes data, thereby controlling the terminal device as a whole. Optionally, the first processor 170 may include one or more processing units; preferably, the first processor 170 may integrate an application processor and a modem processor, wherein the application processor mainly processes an operating system, a user interface, and an application program, and the modem processor mainly processes wireless communication. It is understandable that the above-mentioned modem processor may not be integrated into the first processor 170. In some embodiments, the first processor 170 and the first memory 120 may be implemented on a single chip, and in some embodiments, they may also be implemented separately on separate chips. The first processor 170 may also be used to generate corresponding operation control signals, send them to corresponding components of the computing and processing device, read and process data in the software, especially read and process data and programs in the first memory 120, so that each functional module therein performs corresponding functions, thereby controlling the corresponding components to act according to the requirements of the instructions.
[0066] Among them, the first memory 120 can be used to store software codes related to the script generation method, the first processor 170 can execute the steps of the script generation method, and can also schedule other units (such as the above-mentioned input unit 130 and display unit 140) to implement corresponding functions.
[0067] The radio frequency unit 110 (optional) can be used for receiving and sending information or receiving and sending signals during a call, for example, after receiving the downlink information of the base station, it is sent to the first processor 170 for processing; in addition, the designed uplink data is sent to the base station. Generally, the RF circuit includes but is not limited to an antenna, at least one amplifier, a transceiver, a coupler, a low noise amplifier (LowNoiseAmplifier, LNA), a duplexer, etc. In addition, the radio frequency unit 110 can also communicate with network devices and other devices through wireless communication. The wireless communication can use any communication standard or protocol, including but not limited to Global System of Mobile Communication (Global System of Mobile Communication, GSM), General Packet Radio Service (General PacketRadio Service, GPRS), Code Division Multiple Access (Code Division Multiple Access, CDMA), Wideband Code Division Multiple Access (Wideband Code Division Multiple Access, WCDMA), Long Term Evolution (Long Term Evolution, LTE), email, Short Messaging Service (SMS), etc.
[0068] In this embodiment of the present application, the RF unit 110 can send data to the server 200 and receive processing results sent by the server 200.
[0069] It should be understood that the radio frequency unit 110 is optional and can be replaced by other communication interfaces, such as a network port.
[0070] The terminal 100 also includes a power supply 190 (such as a battery) for supplying power to various components. Preferably, the power supply can be logically connected to the first processor 170 through a power management system, so that the power management system can manage functions such as charging, discharging, and power consumption.
[0071] The terminal 100 further includes an external interface 180 , which may be a standard Micro USB interface or a multi-pin connector, and may be used to connect the terminal 100 to communicate with other devices, or to connect a charger to charge the terminal 100 .
[0072] Although not shown, the terminal 100 may also include a flashlight, a wireless fidelity (WiFi) module, a Bluetooth module, sensors with different functions, etc., which are not described in detail here. Some or all of the methods described below may be applied in the following embodiments. Figure 2 In the terminal 100 shown.
[0073] Next describe Figure 1 The product form of the server 200;
[0074] Figure 3 A structural diagram of a server 200 is provided, such as Figure 3 As shown, the server 200 includes a first bus 201, a second processor 202, a communication interface 203, and a second memory 204. The second processor 202, the second memory 204, and the communication interface 203 communicate with each other via the first bus 201.
[0075] The first bus 201 may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus. The first bus 201 may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 3 Only one thick line is used in the diagram, but this does not mean that there is only one bus or only one type of bus.
[0076] The second processor 202 may be any one or more of a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).
[0077] The second memory 204 may include a volatile memory, such as a random access memory (RAM). The second memory 204 may also include a non-volatile memory, such as a read-only memory (ROM), a flash memory, a hard disk drive (HDD), or a solid state drive (SSD).
[0078] Among them, the second memory 204 can be used to store software codes related to the script generation method, the second processor 202 can execute the steps of the script generation method of the chip, and can also schedule other units to implement corresponding functions.
[0079] It should be understood that the above-mentioned terminal 100 and server 200 can be centralized or distributed devices, and the first processor 170 in the above-mentioned terminal 100 and the second processor 202 in the server 200 can be hardware circuits (such as application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), general-purpose processors, DSPs, microprocessors or microcontrollers, etc.), or a combination of these hardware circuits. For example, the first processor 170 and the second processor 202 can be hardware systems with the function of executing instructions, such as CPU, DSP, etc., or hardware systems without the function of executing instructions, such as ASIC, FPGA, etc., or a combination of the above-mentioned hardware systems without the function of executing instructions and hardware systems with the function of executing instructions.
[0080] The present application provides a script generation method. The script generation method of the embodiment of the present application is described in detail below in conjunction with the accompanying drawings.
[0081] Reference Figure 4 , Figure 4 A flowchart of a script generation method provided in an embodiment of the present application, the method may include:
[0082] Step S401: Obtain the original video to be processed.
[0083] Here, the original video refers to the video for which the script is to be created.
[0084] Optionally, the original video can be a video on a video creation platform, a video on the Internet, or a video shot on an electronic device such as a mobile phone.
[0085] Of course, the source of the original video may be other sources, which is not specifically limited in this application.
[0086] Step S402: perform storyboard processing on the original video to obtain multiple storyboard videos.
[0087] Considering that the amount of video data is generally large, the video content is difficult to understand and process, and at the same time, long videos are prone to target tracking loss. In order to facilitate subsequent processing, this embodiment can perform storyboard processing on the original video to obtain multiple storyboard videos contained in the original video.
[0088] Here, there are many bases for storyboarding, and this application provides but is not limited to the following.
[0089] The first method is to perform storyboard processing according to a fixed number of video frames. For example, if the number of video frames is x, the original video is divided into a video segment every x frames, thereby obtaining multiple storyboard videos.
[0090] The second method: storyboard processing is performed according to scenes. For example, if the original video contains y scenes, then y scenes will divide the original video into y video clips, that is, y storyboard videos are obtained.
[0091] Based on this, optionally, this embodiment can detect scene change information in the original video, and divide the original video into multiple storyboard videos according to the detected scene change information, wherein the multiple storyboard videos each belong to a different scene.
[0092] Optionally, the scene change information may be a start time code and / or an end time code of a scene.
[0093] Optionally, you can use the scenedetect library to detect scene change information in the original video, so as to split the original video into multiple storyboard videos based on the scene change information. Here, the scenedetect library is a Python library for video scene detection, which uses a variety of algorithms to detect scene changes by analyzing the visual differences between video frames. It can identify different scenes in the video and return the start and end time codes of each scene, which allows users to split the original video into multiple sub-videos based on these time codes.
[0094] Of course, the basis for storyboarding can also be other, such as storyboarding according to themes, etc., which is not specifically limited in this application.
[0095] Step S403: Track and mark the characters in the multiple storyboard videos to obtain multiple character marked videos.
[0096] Considering that the script is generated based on the content understanding of the storyboard video, there may be a situation where the script only objectively describes the video content, but the video content is not associated with the character. For example, the generated script may be "a man in a white shirt gave an apple to a woman in a blue suit". Since it is unclear which character "the man in the white shirt" and "the woman in the blue suit" correspond to, the generated script is unusable.
[0097] In order to avoid the above problems, the embodiment of the present application can track and mark the characters in each storyboard video to obtain a character role marked video corresponding to the storyboard video, thereby obtaining multiple character role marked videos.
[0098] It should be noted that, when labeling, the same character in multiple storyboard videos is labeled with the same label, that is, the same character in multiple character labeling videos is labeled with the same label, and different characters are labeled with different labels.
[0099] For example, storyboard video 1 contains characters 1 and 2, storyboard video 2 contains characters 3 and 4, and characters 1 and 3 are the same character. In the character labeling video corresponding to storyboard video 1, character 1 is labeled 1 and character 2 is labeled 2. In the character labeling video corresponding to storyboard video 2, character 3 is labeled 1 and character 4 is labeled 3.
[0100] It should also be noted that, in this embodiment, the person being tracked and marked can be a person or an object, which can be determined according to the actual application scenario, and this application does not make any specific limitation.
[0101] Step S404: Perform scenario description processing on multiple character-labeled videos to obtain a target script containing the characters.
[0102] Specifically, this embodiment can understand the video content of multiple character-tagged videos and perform scenario description based on the tagged tags to obtain a target script containing the character.
[0103] The script generation method provided by the present application takes into account that short video clips are more helpful for understanding video content and character tracking than complete videos. The present application performs storyboard processing on the original video to obtain multiple storyboard videos, and then tracks and marks the characters in the multiple storyboard videos to obtain multiple character role marking videos. Further, the multiple character role marking videos are processed for scene description to obtain a target script containing the character roles. The present application can avoid the loss of tracking targets during character tracking through storyboarding, improve the accuracy and reliability of tracking and marking, and the storyboard makes the content of the character role marking video more understandable, improves the accuracy of scene description processing, and thus improves the quality of the target script. At the same time, through character tracking and marking, the character roles can be integrated into the target script, further improving the quality of the target script and ensuring the consistency of the target script with the original video.
[0104] In some embodiments of the present application, the process of the above step S403 "tracking and marking characters in multiple storyboard videos to obtain multiple character marked videos" is introduced.
[0105] In a possible implementation, this embodiment can track and mark characters in multiple storyboard videos through a storyboard-based character marking system to obtain multiple character-marked videos.
[0106] Specifically, firstly, character tracking is performed on multiple storyboard videos respectively to obtain character trajectories corresponding to the multiple storyboard videos respectively.
[0107] Optionally, a pre-trained head detection model may be used to track the person in each storyboard video to obtain the person trajectory corresponding to each storyboard video. Here, the pre-trained head detection model is trained using training videos with person trajectory labels as training data.
[0108] Optionally, the pre-trained head detection model can be a YOLOv8 model. Here, the YOLOv8 model can use the multi-target tracking DeepSORT algorithm to simultaneously track multiple characters in the storyboard video, so that when a storyboard video contains multiple characters, multiple character trajectories can be obtained. Here, the DeepSORT algorithm is mainly based on the SORT (Simple Online and Realtime Tracking, multi-target tracking) algorithm, and combines deep learning technology to improve tracking accuracy.
[0109] Furthermore, the characters in the multiple storyboard videos may be labeled according to the character trajectories respectively corresponding to the multiple storyboard videos, thereby obtaining multiple character labeled videos.
[0110] For example, characters belonging to the same character trajectory are marked with the same identification (ID). Optionally, the identification may be a role identifier, such as an actor's name, or a predefined character serial number, etc., which is not limited in this application.
[0111] In an optional embodiment, the process of "labeling characters in a plurality of storyboard videos according to character trajectories corresponding to the plurality of storyboard videos, and obtaining a plurality of character-labeled videos" may include: obtaining a preconfigured character face mapping table, wherein the character face mapping table includes the correspondence between the character identification (such as actor name) and facial image of each character in the original video, for each storyboard video in the plurality of storyboard videos, according to the character trajectory corresponding to the storyboard video, obtaining the facial image of each character from the storyboard video, performing similarity matching between the facial image of each character and the facial image in the character face mapping table to obtain a matching result, determining the character identification of each character according to the matching result, labeling each character in the storyboard video according to the character identification of each character, and obtaining a character-labeled video corresponding to the storyboard video, so as to obtain character-labeled videos corresponding to the plurality of storyboard videos as the plurality of character-labeled videos.
[0112] Optionally, the process of "performing similarity matching between the facial image of each person and the facial images in the character facial mapping table" may include: using a face recognition model to perform cosine similarity matching between the facial image of each person and the facial images in the character facial mapping table, where the face recognition model is trained with the training facial images labeled with cosine similarity tags and the corresponding character facial mapping table as training data.
[0113] Taking the character facial mapping table including 4 groups of mapping relationships of character identifier 1 - facial image 1, character identifier 2 - facial image 2, character identifier 3 - facial image 3, and character identifier 4 - facial image 4 as an example, the facial image of each person (taking the facial image of person a as an example) can be used to calculate the similarity with facial images 1 to 4 respectively. Assuming the similarities are 0.5, 0.6, 0.7, and 0.8 respectively, then it can be determined that the character identifier of this person a is character identifier 4.
[0114] Optionally, the process of "performing character tagging on each person in the storyboard video according to the character identifier of each person" may include: printing the character identifier of each person as a tag to the head area of the corresponding person in the storyboard video.
[0115] For example, referring to Figure 5 As shown, it is a schematic diagram of tagging the head of a person included in a video frame. Taking the character identifier as the actor's name as an example, in this embodiment, the actor's name can be tagged to the corresponding person's head area. For example, "Zhang San" is in Figure 5 the head area of the left person, and "Li Si" is in Figure 5 the head area of the right person.
[0116] Of course, the position of character tagging can be other areas besides the head area, and this application does not make a limitation.
[0117] In summary, in the embodiments of this application, character tracking and tagging are performed on the people in each storyboard video respectively, so that the character tagging videos corresponding to each storyboard video are marked with the character identifiers of each person, such as the actor's name, which improves the character recognition rate, helps to distinguish character roles in subsequent scenario description processing, and thus improves the generation effect of the target script.
[0118] In other embodiments of this application, the process of step S404 "performing scenario description processing on multiple character tagging videos to obtain a target script including character roles" is introduced.
[0119] In one possible implementation, in this embodiment, a scenario description system based on a multi-modal large model integrating character information can be used to perform scenario description processing on multiple character tagging videos to obtain a target script including character roles.
[0120] Specifically, firstly, scenario understanding is performed on multiple character labeling videos respectively to obtain video content description information corresponding to the multiple character labeling videos respectively, wherein the video content description information includes role identifications of the characters in the corresponding character labeling videos.
[0121] Optionally, the process of "performing situational understanding on multiple character-labeled videos respectively to obtain video content description information corresponding to the multiple character-labeled videos respectively" may include: it may be implemented by a pre-trained video understanding multimodal large model. Here, the training data used by the video understanding multimodal large model in the training phase includes: multiple character-labeled training videos and corresponding video content description information, and the multiple character-labeled training videos are obtained by tracking and labeling the characters in multiple training storyboard videos.
[0122] Specifically, this embodiment can obtain a certain number of videos from various channels such as the Internet and databases, and then obtain multiple training storyboard videos according to the storyboard processing method provided above. For example, the obtained videos are storyboarded using the scenedetect library to obtain 1,000 training storyboard videos.
[0123] The multiple training storyboard videos are sent to the storyboard-based character labeling system to track and label the characters in the multiple training storyboard videos, thereby obtaining multiple character character labeling training videos.
[0124] Next, the present application can manually describe the video content of multiple character labeling training videos to obtain video content description information corresponding to the multiple character labeling training videos. It is worth noting that the manually generated video content description information needs to include the character identifiers in the character labeling training videos, for example, Figure 5 The corresponding video content description information is "Zhang San and Li Si are talking."
[0125] After the above preprocessing, this embodiment obtains multiple character role labeling training videos and corresponding video content description information, which can be used as training sets to fine-tune and pre-train the video understanding multimodal base large model, and obtain the above-mentioned pre-trained video understanding multimodal large model (the video understanding multimodal large model trained by the training set can already meet the requirements of this embodiment to output video content description information containing character identification. Therefore, there is no need to fine-tune and pre-train the video understanding multimodal large model again based on a new training set in the future).
[0126] Optionally, the video understanding multimodal base large model can be Qwen2-VL video understanding multimodal large model, where Qwen2-VL is a large visual language model with three models of different parameter sizes: 2 billion, 7 billion, and 72 billion. Qwen2-VL is optimized for dialogue scenarios and outperforms other open source dialogue models in most human evaluations of usefulness and safety. Optionally, this embodiment selects the 7 billion (7B) model Qwen2-VL-7B-Instruct as the video understanding multimodal base large model.
[0127] Optionally, the LORA (Low-Rank Adaptation) method in PEFT (Parameter Efficient Fine-tuning) can be used for fine-tuning the pre-training of the large multimodal base model for video understanding.
[0128] Optional, Lora configuration parameters include: matrix rank r=64, scaling factor lora_alpha=16, lora_dropout=0.05 (lora_dropout is used to randomly "discard" (i.e. temporarily remove) some neurons in the network during training to prevent overfitting), training epochs=20, learning rate learning_rate=0.001.
[0129] After training the video understanding multimodal large model in the previous article, this embodiment can obtain the first text prompt words, wherein the first text prompt words include the basic elements of the script required to generate the target script; multiple character role labeled videos and the first text prompt words are input into the video understanding multimodal large model together to obtain the video content description information corresponding to the multiple character role labeled videos.
[0130] Optionally, the basic elements of the script include one or more of the following information: character dialogue information, surrounding environment information, and character action description information.
[0131] Of course, the basic elements of the script may also be other, and this application does not limit them.
[0132] Furthermore, this embodiment can integrate the video content description information corresponding to multiple character marking videos to obtain the target script.
[0133] Optionally, the process of “integrating video content description information corresponding to multiple character labeled videos to obtain a target script” can be achieved through a large language model.
[0134] Specifically, this embodiment can obtain a second text prompt word, wherein the second text prompt word is used to prompt the large language model to integrate the video content description information corresponding to multiple character labeling videos; the video content description information corresponding to the multiple character labeling videos and the second text prompt word are input into the large language model together to obtain the target script output by the large language model.
[0135] Optionally, the second text prompt word includes prompt words integrating video content description information corresponding to all storyboard videos. The second text prompt word may also include other optimized prompt words, which are not limited in this application.
[0136] Optionally, the large language model may be GPT-4 (Generative Pre-trained Transformer 4). Of course, the large language model may also be other, which is not limited in this application.
[0137] In summary, this embodiment uses a large multimodal model for video understanding to generate video content description information. Compared with other large models, it can more accurately understand the video content of the character marking video, and thus generate more accurate video content description information. Furthermore, a large language model is used to integrate the video content description information corresponding to multiple storyboard videos, so that the generated target script contains the role identification of each character, so that the target script has higher usability.
[0138] It has been verified experimentally that the script generation method provided in this embodiment can accurately integrate the characters in the original video into the script content description (label the characters), and automatically generate a high-quality, strongly correlated, content-rich and professional target script. The target script does not require secondary manual calibration, and the effect is equivalent to that of a manually produced script, which greatly saves manpower and improves the generation effect from video to script.
[0139] The above introduces a script generation method provided by an embodiment of the present application, and the following will introduce a device for executing the above-mentioned script generation method.
[0140] See also Figure 6 , Figure 6 This is a schematic diagram of the structure of a script generation device provided in an embodiment of the present application. Figure 6 As shown, the device may include:
[0141] The video acquisition module 501 is used to acquire the original video to be processed;
[0142] The video splitting module 502 is used to split the original video to obtain multiple splitting videos;
[0143] A tracking and marking module 503 is used to track and mark the characters in the multiple storyboard videos to obtain multiple character marking videos;
[0144] The script generation module 504 is used to perform scenario description processing on multiple character marking videos to obtain a target script containing the characters.
[0145] In a possible implementation, the video storyboard module may include: a scene detection module and a scene storyboard module;
[0146] A scene detection module is used to detect scene change information in the original video;
[0147] The scene storyboard module is used to divide the original video into multiple storyboard videos according to scene change information, wherein the multiple storyboard videos each belong to a different scene.
[0148] In a possible implementation, the tracking and marking module may include: a person tracking module and a character marking module;
[0149] A character tracking module is used to track characters in multiple storyboard videos respectively, and obtain character trajectories corresponding to the multiple storyboard videos respectively;
[0150] The character labeling module is used to label the characters in the multiple storyboard videos according to the character trajectories corresponding to the multiple storyboard videos, so as to obtain multiple character labeling videos.
[0151] In a possible implementation, the above-mentioned character marking module may include: a mapping table acquisition module and a face matching module;
[0152] A mapping table acquisition module, used to acquire a pre-configured character face mapping table, wherein the character face mapping table includes the corresponding relationship between the character identification and the facial image of each character in the original video;
[0153] The face matching module is used for obtaining the face image of each character from each storyboard video in a plurality of storyboard videos according to the character trajectory corresponding to the storyboard video; performing similarity matching between the face image of each character and the face image in the character face mapping table to obtain a matching result, and determining the role identification of each character according to the matching result; performing role labeling on each character in the storyboard video according to the role identification of each character to obtain a role labeling video corresponding to the storyboard video; and obtaining the role labeling videos corresponding to the plurality of storyboard videos as the plurality of role labeling videos.
[0154] In a possible implementation, the script generation module may include: a scenario understanding module and an information integration module;
[0155] A scenario understanding module is used to perform scenario understanding on a plurality of character-marked videos, respectively, to obtain video content description information corresponding to the plurality of character-marked videos, wherein the video content description information includes role identifiers of the characters in the corresponding character-marked videos;
[0156] The information integration module is used to integrate the video content description information corresponding to multiple character labeling videos to obtain the target script.
[0157] In a possible implementation, the scenario understanding module may include: a first prompt acquisition module and a video description module;
[0158] A first prompt acquisition module, used to acquire a first text prompt word, wherein the first text prompt word includes a basic script element required to generate a target script;
[0159] A video description module, used to input the multiple character-labeled videos and the first text prompt word into a pre-trained video understanding multimodal large model to obtain video content description information corresponding to the multiple character-labeled videos;
[0160] Among them, the training data used by the video understanding multimodal large model in the training stage includes: multiple character labeling training videos and corresponding video content description information. The multiple character labeling training videos are obtained by tracking and labeling the characters in multiple training storyboard videos.
[0161] In a possible implementation, the information integration module may include: a second prompt acquisition module and a video description integration module;
[0162] A second prompt acquisition module is used to acquire a second text prompt word, wherein the second text prompt word is used to prompt the large language model to integrate video content description information corresponding to the multiple character-labeled videos;
[0163] The video description integration module is used to input the video content description information and the second text prompt words corresponding to multiple character-labeled videos into the large language model to obtain the target script output by the large language model.
[0164] The script generation device provided in the embodiment of the present application corresponds to the script generation method provided above. For details, please refer to the above introduction and will not be repeated here.
[0165] The present application also provides an electronic device in an embodiment. Figure 7As shown, it shows a schematic diagram of the structure of an electronic device suitable for implementing the embodiment of the present application. The electronic device in the embodiment of the present application may include but is not limited to fixed terminals such as mobile phones, laptops, PDAs (personal digital assistants), PADs (tablet computers), desktop computers, etc. Figure 7 The electronic device shown is merely an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.
[0166] like Figure 7 As shown, the electronic device may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 601, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 602 or a program loaded from a storage device 608 to a random access memory (RAM) 603. When the electronic device is powered on, various programs and data required for the operation of the electronic device are also stored in the RAM 603. The processing device 601, the ROM 602, and the RAM 603 are connected to each other via a second bus 604. An input / output (I / O) interface 605 is also connected to the second bus 604.
[0167] Typically, the following devices may be connected to the I / O interface 605: an input device 606 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 607 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 608 including, for example, a memory card, a hard disk, etc.; and a communication device 609. The communication device 609 may allow the electronic device to communicate with other devices wirelessly or by wire to exchange data. Although Figure 7 An electronic device having various devices is shown, but it should be understood that it is not required to implement or possess all the devices shown. More or fewer devices may be implemented or possessed instead.
[0168] Also provided in an embodiment of the present application is a computer program product, including computer-readable instructions. When the computer-readable instructions are executed on an electronic device, the electronic device implements any one of the script generation methods provided in the embodiments of the present application.
[0169] A computer-readable storage medium is also provided in an embodiment of the present application. The storage medium carries one or more computer programs. When one or more computer programs are executed by an electronic device, the electronic device can implement any script generation method provided in an embodiment of the present application.
[0170] It should also be noted that the device embodiments described above are merely schematic, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed over multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. In addition, in the drawings of the device embodiments provided by the present application, the connection relationship between the modules indicates that there is a communication connection between them, which may be specifically implemented as one or more communication buses or signal lines.
[0171] Through the description of the above implementation mode, the technicians in the field can clearly understand that the present application can be implemented by means of software plus necessary general hardware, and of course, it can also be implemented by special hardware including special integrated circuits, special CPUs, special memories, special components, etc. In general, all functions completed by computer programs can be easily implemented by corresponding hardware, and the specific hardware structure used to implement the same function can also be various, such as analog circuits, digital circuits or special circuits. However, for the present application, software program implementation is a better implementation mode in more cases. Based on such an understanding, the technical solution of the present application is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which is stored in a readable storage medium, such as a computer floppy disk, a U disk, a mobile hard disk, a ROM, a RAM, a disk or an optical disk, etc., including a number of instructions to enable a computer device (which can be a personal computer, a training device, or a network device, etc.) to execute the methods described in each embodiment of the present application.
[0172] In the above embodiments, all or part of the embodiments may be implemented by software, hardware, firmware or any combination thereof. When implemented by software, all or part of the embodiments may be implemented in the form of a computer program product.
[0173] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions may be transmitted from a website site, a computer, a training device, or a data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) mode to another website site, computer, training device, or data center. The computer-readable storage medium may be any available medium that a computer can store or a data storage device such as a training device, a data center, etc. that includes one or more available media integrations. The available medium may be a magnetic medium, (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive (SSD)), etc.
Claims
1. A script generation method, characterized in that: include: Get the original video to be processed; Performing storyboard processing on the original video to obtain multiple storyboard videos; Tracking and marking characters in the multiple storyboard videos to obtain multiple character-marked videos; The multiple character-labeled videos are processed for scenario description to obtain a target script containing the characters.
2. The script generation method according to claim 1, characterized in that: The process of performing storyboard processing on the original video to obtain a plurality of storyboard videos includes: Detecting scene change information in the original video; According to the scene change information, the original video is divided into the multiple storyboard videos, wherein the multiple storyboard videos each belong to a different scene.
3. The script generation method according to claim 1, characterized in that: The step of tracking and marking the characters in the plurality of storyboard videos to obtain a plurality of character marked videos includes: Tracking characters in the plurality of storyboard videos respectively to obtain character trajectories corresponding to the plurality of storyboard videos respectively; According to the character trajectories respectively corresponding to the multiple storyboard videos, the characters in the multiple storyboard videos are marked with characters to obtain the multiple character marked videos.
4. The script generation method according to claim 3, characterized in that: The step of labeling characters in the multiple storyboard videos according to the character trajectories respectively corresponding to the multiple storyboard videos to obtain the multiple character labeling videos includes: Obtaining a preconfigured character face mapping table, wherein the character face mapping table includes a correspondence between the character identifications and facial images of all characters in the original video; For each storyboard video in the plurality of storyboard videos: According to the character trajectory corresponding to the storyboard video, a facial image of each character is obtained from the storyboard video; Performing similarity matching between the facial image of each character and the facial images in the character face mapping table to obtain a matching result, and determining the role identification of each character according to the matching result; Mark each character in the storyboard video according to the character identification of each character, and obtain a character role-marked video corresponding to the storyboard video; The character marking videos corresponding to the multiple storyboard videos are obtained as the multiple character marking videos.
5. The script generation method according to any one of claims 1 to 4, characterized in that: The step of performing scenario description processing on the plurality of character-labeled videos to obtain a target script containing the characters includes: Performing situation understanding on the plurality of character-labeled videos respectively to obtain video content description information corresponding to the plurality of character-labeled videos respectively, wherein the video content description information includes role identifiers of characters in the corresponding character-labeled videos; The video content description information corresponding to the multiple character marking videos is integrated to obtain the target script.
6. The script generation method according to claim 5, characterized in that: The performing scene understanding on the plurality of character-labeled videos to obtain video content description information corresponding to the plurality of character-labeled videos respectively includes: Acquire a first text prompt word, wherein the first text prompt word includes basic script elements required to generate the target script; Inputting the plurality of character-labeled videos and the first text prompt word into a pre-trained multimodal large model for video understanding, and obtaining video content description information corresponding to the plurality of character-labeled videos; Among them, the training data used by the video understanding multimodal large model in the training stage includes: multiple character labeling training videos and corresponding video content description information, and the multiple character labeling training videos are obtained by tracking and labeling the characters in multiple training storyboard videos.
7. The script generation method according to claim 5, characterized in that: The step of integrating the video content description information corresponding to the plurality of character-labeled videos to obtain the target script includes: Obtaining a second text prompt word, wherein the second text prompt word is used to prompt the large language model to integrate the video content description information corresponding to the multiple character labeling videos respectively; The video content description information corresponding to the multiple character-labeled videos and the second text prompt words are input into the large language model to obtain the target script output by the large language model.
8. A script generation device, characterized in that: include: A video acquisition module is used to acquire the original video to be processed; A video storyboard module, used for storyboarding the original video to obtain multiple storyboard videos; A tracking and marking module, used for tracking and marking characters in the plurality of storyboard videos to obtain a plurality of character marking videos; The script generation module is used to perform scenario description processing on the multiple character marking videos to obtain a target script containing the characters.
9. An electronic device, characterized in that: The method comprises at least one processor and a memory connected to the processor, wherein: The memory is used to store computer programs; The processor is used to execute the computer program so that the electronic device can implement the script generation method as described in any one of claims 1 to 7.
10. A computer storage medium, characterized in that: The storage medium carries one or more computer programs, which, when executed by an electronic device, enable the electronic device to implement the script generation method as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Generation method and generation device of character playing locus in video, and client
CN103336955A
Method and device for tracking character in scene
CN106340035A
Method and device for implementing video character labeling
CN108401176A
A tagging method for face tracking in image recognition
CN109215058A
Video synthesis method, client and system
CN112188117A