Method, related device and medium for generating character perspective video

By extracting the source script information and background descriptions of film and television works, combining the character script conversion model and diffusion model, we generate auxiliary videos of different perspectives of characters in film and television works, solving the problem of single perspective display in the existing technology and improving the efficiency of information transmission.

CN119545119BActive Publication Date: 2025-06-10TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510032966.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-09
Publication Date
2025-06-10
Estimated Expiration
2045-01-09

AI Technical Summary

Technical Problem

It is difficult for the existing technology to efficiently generate videos of different perspectives of characters in film and television works, especially in film and television works. The single perspective display limits the audience's ability to experience film and television works from different perspectives, resulting in inefficient information transmission.

Method used

By obtaining the source video and source background description information, extracting the source script information, and using the character script transformation model to generate the character script information from the target character perspective, combining the target character picture information and source background description information to generate diffusion guidance information, and finally generating the character view video from the target character perspective.

Benefits of technology

It improves the accuracy of film and television works to generate auxiliary videos with different perspectives, enhances the efficiency of information transmission, and allows viewers to experience the content of film and television works from multiple angles.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119545119B_ABST
    Figure CN119545119B_ABST
Patent Text Reader

Abstract

The present disclosure provides a method for generating a role perspective video, related devices, and a medium. The method includes: obtaining a source video with a target role and source background description information of the source video; extracting source script information from the source video; obtaining guiding language template information corresponding to a role script conversion model; generating role script information from the perspective of the target role by using the role script conversion model based on the source script information, the source background description information, and the guiding language template information; generating diffusion guiding information based on the role script information, picture information corresponding to the target role, and the source background description information, and generating a role perspective video from the perspective of the target role based on the diffusion guiding information. The present disclosure can improve the accuracy of generating auxiliary videos from different character perspectives for film and television works. The present disclosure can be applied to various scenarios such as big data, artificial intelligence, cloud technology, etc.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of big data technology, and in particular to a method for generating a character perspective video, a related device and a medium. Background Art

[0002] At present, in some business scenarios of showing narrative movies or TV series, the film and television works often need to be shown to each object from a single perspective. This display method often limits the object from experiencing the film and television works from different perspectives, such as the psychological activities of the characters and the picture information from the perspective of the characters. This information is sometimes critical to fully understand the content of the film and television works. This key information cannot be transmitted to the object, resulting in low efficiency of information transmission.

[0003] Therefore, there is an urgent need to generate auxiliary videos from different perspectives for film and television works to improve the efficiency of information transmission. The related technology is a technology that generates short videos from different perspectives for characters in a short video based on the description information of a short video, but this technology is limited to short videos, and it is generated based on the description information of the short video, which cannot be accurately obtained, resulting in the low accuracy of the auxiliary short videos generated from different perspectives of the characters in the technology. How to generate auxiliary videos from different perspectives of characters for film and television works with high accuracy is still a problem that needs to be solved urgently. Summary of the invention

[0004] The embodiments of the present disclosure provide a method, a related device and a medium for generating a character perspective video, which can improve the accuracy of generating auxiliary videos from different character perspectives for film and television works.

[0005] According to one aspect of the present disclosure, a method for generating a character perspective video is provided, the method comprising:

[0006] Acquire a source video having a target character and source background description information of the source video;

[0007] Extracting source script information from the source video;

[0008] Obtain the introductory text template information corresponding to the role script conversion model;

[0009] Based on the source script information, the source background description information and the guide template information, the role script information from the perspective of the target character is generated by using the role script conversion model;

[0010] Based on the role script information, the picture information corresponding to the target role, and the source background description information, diffusion guidance information is generated, and based on the diffusion guidance information, a role perspective video from the perspective of the target role is generated.

[0011] According to one aspect of the present disclosure, a device for generating a character perspective video is provided, the device comprising:

[0012] A first acquisition unit, configured to acquire a source video having a target character and source background description information of the source video;

[0013] An extraction unit, used to extract source script information from the source video;

[0014] The second acquisition unit is used to acquire the introductory language template information corresponding to the role script conversion model;

[0015] A first generating unit, configured to generate role script information from the perspective of the target role using the role script conversion model based on the source script information, the source background description information and the guide template information;

[0016] The second generating unit is used to generate diffusion guidance information based on the role script information, the picture information corresponding to the target role, and the source background description information, and generate a role perspective video from the perspective of the target role based on the diffusion guidance information.

[0017] Optionally, the first generating unit includes:

[0018] An acquisition module, used to acquire image information corresponding to the target character;

[0019] The first generation module is used to generate introductory words based on the source script information, the picture information, and the source background description information using the introductory words template information, and input the introductory words into the role script conversion model to obtain the role script information.

[0020] Optionally, the source script information includes source line information and source scene information, the role script information includes role script line information and role script scene information, the introductory language template information includes first introductory language template information and second introductory language template information, and the introductory language includes first introductory language and second introductory language;

[0021] The first generating module is used for:

[0022] Based on the source dialogue information, the source scene information, the picture information and the source background description information, the first guide language template information is used to generate a first guide language, and the first guide language is input into the role script conversion model to obtain the role script dialogue information;

[0023] Based on the source scene information and the source background description information, the second guide language is generated using the second guide language template information, and the second guide language is input into the role-script conversion model to obtain the role-script scene information.

[0024] Optionally, the role-script conversion model includes a first sub-model and a second sub-model;

[0025] The first generating unit is used for:

[0026] Based on the guide template information, the source script information, and the source background description information, the first sub-model is used to generate an instance to obtain a role script prompt instance;

[0027] Based on the role script prompt instance, the script is generated through the second sub-model to obtain the role script information.

[0028] Optionally, the second acquiring unit includes:

[0029] A segmentation module, used for segmenting the source video to obtain a plurality of source video frames arranged in time sequence;

[0030] A second generation module is used to input the plurality of source video frames into a pre-trained video script description generation model for description generation, so as to obtain video frame script description information of each of the plurality of source video frames;

[0031] An integration module is used to integrate the video frame script description information of each of the multiple source video frames to obtain the source script information.

[0032] Optionally, the video script description generation model includes a feature extraction network and a script description generation network;

[0033] The second generation module comprises:

[0034] An extraction submodule, configured to extract features of each of the source video frames based on the feature extraction network to obtain target video frame features of the source video frame;

[0035] A generation submodule is used to generate a script based on the target video frame features using the script description generation network to obtain the video frame script description information of the source video frame.

[0036] Optionally, the extraction submodule is used to:

[0037] For each of the source video frames, performing multi-scale downsampling on the source video frame based on the feature extraction network to obtain a first downsampling feature, a second downsampling feature, and a third downsampling feature, wherein the feature dimensions of the first downsampling feature, the second downsampling feature, and the third downsampling feature are increased in sequence;

[0038] Performing dimensionality reduction processing on the first down-sampling feature, the second down-sampling feature, and the third down-sampling feature, respectively, to obtain a first dimensionality reduction feature corresponding to the first down-sampling feature, a second dimensionality reduction feature corresponding to the second down-sampling feature, and a third dimensionality reduction feature corresponding to the third down-sampling feature;

[0039] Performing splicing processing on the first dimensionality reduction feature, the second dimensionality reduction feature, and the third dimensionality reduction feature to obtain an initial video frame feature of the source video frame;

[0040] The initial video frame features of the source video frame and the initial video frame features of a previous source video frame of the source video frame are spliced ​​to obtain target video frame features of the source video frame.

[0041] Optionally, the integration module is used to:

[0042] splicing the video frame script description information of each of the multiple source video frames to obtain preliminary script information;

[0043] The preliminary script information is polished to obtain the source script information.

[0044] Optionally, the second generating unit is used for:

[0045] Encoding the source background description information to obtain background description encoding features;

[0046] Encoding the role script information to obtain role script encoding features;

[0047] Extracting features from the image information to obtain features of the target character image;

[0048] Based on the background description coding features, the role script coding features, and the target role image features, conditional information is generated through a preset conditional fusion network to obtain the diffusion guidance information.

[0049] Optionally, the second generating unit is used to: guide a diffusion model based on the diffusion guiding information to generate a character perspective video of the target character perspective;

[0050] The diffusion model is trained in the following way:

[0051] A sample acquisition unit, used to acquire a sample video having a sample character, sample background description information of the sample video, and an expected character perspective video of the sample character;

[0052] A sample extraction unit, used to extract sample script information from the sample video;

[0053] A sample script generating unit, configured to generate sample role script information from the perspective of the sample role based on the sample script information, the sample background description information and the preset introductory language template information and using a role script conversion model;

[0054] A sample guide information generating unit, configured to generate sample guide information based on the sample role script information, the sample picture information corresponding to the sample role, and the sample background description information;

[0055] A sample video generating unit, configured to generate a video through a target model based on the sample guiding information to obtain a sample character perspective video of the sample character perspective;

[0056] The first training unit is used to train the target model based on the sample role perspective video and the expected role perspective video to obtain the diffusion model.

[0057] Optionally, the sample video includes a plurality of sample video frames; the target model includes a diffusion network, a denoising network, and a decoding network;

[0058] The sample video generating unit is used for:

[0059] A diffusion module, configured to perform diffusion processing on the sample video frame encoding features of each of the sample video frames based on the diffusion network to obtain a sample latent space feature vector;

[0060] A denoising module, configured to perform denoising processing on the sample latent space feature vector through the denoising network based on the sample guide information to obtain a sample video frame denoising feature;

[0061] A decoding module, used for decoding the denoising features of the sample video frames based on the decoding network to obtain predicted video frames of each of the sample video frames;

[0062] The video generation module is used to obtain the sample character perspective video based on the predicted video frames of each of the multiple sample video frames.

[0063] Optionally, the denoising network includes an upsampling module and a downsampling module;

[0064] The denoising module is used to:

[0065] Performing splicing processing on the sample guide information and the sample latent space feature vector to obtain a sample denoising splicing feature;

[0066] Downsampling the sample denoising splicing features based on the downsampling module to obtain video frame downsampling features;

[0067] The down-sampled features of the video frame are up-sampled based on the up-sampling module to obtain the denoising features of the sample video frame.

[0068] Optionally, the denoising module is used to:

[0069] Performing splicing processing on the sample guide information and the sample latent space feature vector to obtain a sample denoising splicing feature;

[0070] Based on the sample denoising splicing features, performing a first denoising process through the denoising network to obtain a first denoising feature;

[0071] Based on the sample latent space feature vector, performing a second denoising process through the denoising network to obtain a second denoising feature;

[0072] The first denoising feature and the second denoising feature are weighted to obtain the sample video frame denoising feature.

[0073] Optionally, the sample video includes a plurality of sample video frames;

[0074] The first training unit is used for:

[0075] Determining a first loss sub-function based on the sample character perspective video and the expected character perspective video;

[0076] Determining a second loss sub-function based on the sample video frame denoising features corresponding to each of the plurality of sample video frames and a preset benchmark denoising feature;

[0077] Based on the first loss sub-function and the second loss sub-function, obtaining a first loss function;

[0078] Based on the first loss function, the target model is trained to obtain the diffusion model.

[0079] Optionally, determining the second loss sub-function based on the sample video frame denoising features corresponding to each of the plurality of sample video frames and a preset reference denoising feature includes:

[0080] For each of the sample video frames, a regularization term is calculated based on the sample video frame denoising feature and the reference denoising feature to obtain a regularization term calculation result;

[0081] The second loss sub-function is determined based on the calculation results of the regularization term of each of the multiple sample video frames.

[0082] Optionally, the sample character perspective video includes a plurality of character perspective video frames, and the expected character perspective video includes reference video frames corresponding to the plurality of character perspective video frames one by one;

[0083] The determining of a first loss sub-function based on the sample character perspective video and the expected character perspective video includes:

[0084] Determine a first pixel value of a pixel point of each of the character perspective video frames, and determine a second pixel value of a pixel point of each of the reference video frames;

[0085] For each of the character perspective video frames, determining a pixel difference value based on the first pixel value and a second pixel value of the reference video frame corresponding to the character perspective video frame;

[0086] The first loss sub-function is determined based on the pixel difference values ​​of each of the multiple character perspective video frames.

[0087] Optionally, the target model is obtained by adjusting the first model through the following steps:

[0088] A reference video frame acquisition unit, used to acquire a plurality of reference video frames of a reference video;

[0089] A convolution unit, configured to perform three-dimensional convolution processing on the multiple reference video frames to obtain predicted means and predicted standard deviations of the multiple reference video frames;

[0090] A reparameterization unit, configured to reparameterize the plurality of reference video frames respectively based on the predicted mean and the predicted standard deviation to obtain video frame hidden state features of the plurality of reference video frames;

[0091] A feature decoding unit, configured to decode the video frame hidden state features of each of the plurality of reference video frames to obtain a predicted video frame corresponding to each of the plurality of reference video frames;

[0092] An adjustment unit is used to adjust the parameters of the first model based on the predicted video frame, the reference video frame, the predicted mean and the predicted standard deviation to obtain the target model.

[0093] Optionally, the adjustment unit is used to:

[0094] Performing mean square error calculation based on the multiple reference video frames and the predicted video frames corresponding to the multiple reference video frames to obtain a video frame error function;

[0095] Performing logarithmic calculation based on the predicted mean and the predicted standard deviation to obtain a log-likelihood function;

[0096] Based on the video frame error function and the log-likelihood function, parameters of the first model are adjusted to obtain the target model.

[0097] Optionally, the video script description generation model is trained in the following manner:

[0098] A reference video acquisition unit, configured to acquire a plurality of reference video frames arranged in time sequence of a reference video, and reference script description information corresponding to each of the plurality of reference video frames;

[0099] a feature extraction unit, configured to extract reference video frame features of each of the plurality of reference video frames;

[0100] A script generation unit, configured to generate a script by using a preset second model based on the reference video frame features of each of the multiple reference video frames and the reference script description information, to obtain predicted script information of each of the multiple reference video frames;

[0101] The second training unit is used to train the second model based on the reference script description information and the predicted script information of each of the multiple reference video frames to obtain the video script description generation model.

[0102] Optionally, the script generation unit is used to:

[0103] For each of the reference video frames, embedding the reference script description information to obtain a reference script description embedding feature;

[0104] Based on the reference script description embedding feature, performing self-attention calculation through the second model to obtain a first attention calculation result;

[0105] Based on the first attention calculation result and the reference video frame feature, performing cross attention calculation through the second model to obtain a second attention calculation result;

[0106] Based on the second attention calculation result and the preset vocabulary in the second model, a script description is generated to obtain the predicted script information of the reference video frame.

[0107] Optionally, the second training unit is used for:

[0108] For each of the reference video frames, based on the reference video frame script description information and the predicted script description information, determine a third loss sub-function;

[0109] Determine a second loss function based on the plurality of third loss sub-functions;

[0110] Based on the second loss function, the second model is trained to obtain the video script description generation model.

[0111] Optionally, the predicted script description information includes a first probability and a first sequence number of each of the plurality of predicted words, and the reference video frame script description information includes a second sequence number of each of the plurality of reference words;

[0112] The determining, for each of the reference video frames, a third loss sub-function based on the reference video frame script description information and the predicted script description information comprises:

[0113] For a single predicted word, if the predicted word is consistent with the reference word whose second serial number is the same as the first serial number of the single predicted word, determining the predicted word as a target word;

[0114] The third loss sub-function is determined based on the negative logarithm of the first probability of the target word.

[0115] According to one aspect of the present disclosure, an electronic device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor implements the above-mentioned method for generating a character perspective video when executing the computer program.

[0116] According to one aspect of the present disclosure, a computer-readable storage medium is provided, wherein the storage medium stores a computer program, and when the computer program is executed by a processor, the method for generating a character perspective video as described above is implemented.

[0117] According to one aspect of the present disclosure, a computer program product is provided. The computer program product includes a computer program. The computer program is read and executed by a processor of an electronic device, so that the electronic device executes the character perspective video generation method as described above.

[0118] In the disclosed embodiment, source script information with a target role is extracted from a source video. The source script information is described from the perspective of a third party other than any role in the source script information. The source background description information reflects the background in which the source script exists, and the psychological activities and actions generated by any role are closely related to the background. With the source script information and the source background description information, the source script can be adjusted using the role script conversion model, and the source script information from the third party's perspective can be converted into the role script information from the perspective of the target role using the role script conversion model. The role script information reflects the adjustment of the psychological activities, actions generated, and pictures seen by the role under the source background description information. Then, based on the role script information, the picture information corresponding to the target role, and the source background description information, diffusion guidance information is generated, and based on the diffusion guidance information, a role perspective video from the perspective of the target role is generated. The picture information corresponding to the target role concretizes the image of the role in the role perspective video, and the source background description information determines the background and color tone of the picture in the role perspective video. Therefore, the specific video is generated by integrating the role script information, the role picture information and the source background description information, taking into account the influence of the role's appearance and the background of the picture on the specific content in the video, and improving the accuracy of the generated role perspective video, thereby improving the information transmission efficiency of the role perspective video. The disclosed embodiment can obtain the role perspective video of the target role perspective from the source video of the third party perspective with high accuracy.

[0119] Other features and advantages of the present disclosure will be described in the following description, and partly become apparent from the description, or understood by practicing the present disclosure. The purpose and other advantages of the present disclosure can be realized and obtained by the structures particularly pointed out in the description, claims and drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0120] The accompanying drawings are used to provide further understanding of the technical solution of the present disclosure and constitute a part of the specification. Together with the embodiments of the present disclosure, they are used to explain the technical solution of the present disclosure and do not constitute a limitation on the technical solution of the present disclosure.

[0121] Figure 1 is a system architecture diagram of a system to which the method for generating a character perspective video according to an embodiment of the present disclosure is applied;

[0122] Figure 2A-2E A schematic diagram showing the application of the method for generating a character perspective video according to an embodiment of the present disclosure in a video production scenario;

[0123] Figure 3A-3B A schematic diagram showing the application of the method for generating a character perspective video in a video playback scenario according to an embodiment of the present disclosure;

[0124] Figure 4is a flowchart of a method for generating a character perspective video according to an embodiment of the present disclosure;

[0125] Figure 5 is a flowchart of extracting source script information according to an embodiment of the present disclosure;

[0126] Figure 6 is a flow chart of generating video frame script description information according to an embodiment of the present disclosure;

[0127] Figure 7 is a flow chart of generating target video frame features of a source video frame according to an embodiment of the present disclosure;

[0128] Figure 8 is a flow chart of generating source script information according to video frame script description information according to an embodiment of the present disclosure;

[0129] Fig. 9 It is a schematic diagram of the overall implementation process of extracting source script information according to an embodiment of the present disclosure;

[0130] Fig.10 It is a schematic diagram of the implementation process of generating role script information according to an embodiment of the present disclosure;

[0131] Fig.11 is a flow chart of generating role script information according to an embodiment of the present disclosure;

[0132] Fig.12 is a flow chart of generating character script line information and character script scene information according to an embodiment of the present disclosure;

[0133] Fig.13 It is a schematic diagram of the implementation process of generating character script line information and character script scene information according to an embodiment of the present disclosure;

[0134] Fig.14 is a flow chart of generating diffusion guidance information according to an embodiment of the present disclosure;

[0135] Fig.15 is a schematic diagram of an implementation process of generating a character perspective video using a diffusion model according to an embodiment of the present disclosure;

[0136] Fig.16 is a schematic diagram of the overall implementation process of generating a character perspective video according to an embodiment of the present disclosure;

[0137] Fig.17 is a flow chart of training a diffusion model according to one embodiment of the present disclosure;

[0138] Fig.18 is a flow chart for generating a sample character perspective video according to one embodiment of the present disclosure;

[0139] Fig.19 is a flow chart for generating sample video frame denoising features according to one embodiment of the present disclosure;

[0140] Fig. 20 is a schematic diagram of an implementation process of generating sample video frame denoising features according to an embodiment of the present disclosure;

[0141] Fig.21 is a flow chart of generating sample video frame denoising features according to another embodiment of the present disclosure;

[0142] Fig. 22 is a flowchart of training a target model based on a first loss function according to an embodiment of the present disclosure;

[0143] Fig.23 is a flow chart of adjusting parameters of a first model according to an embodiment of the present disclosure;

[0144] Figure 24A-Figure 24C is a schematic diagram of the implementation process of the training diffusion model of an embodiment of the present disclosure;

[0145] Fig.25 is a flow chart of a training video script description generation model according to one embodiment of the present disclosure;

[0146] Fig.26 is a flowchart of generating predicted script information of a reference video frame according to an embodiment of the present disclosure;

[0147] Fig. 27 is a module diagram of a device for generating a character perspective video according to an embodiment of the present disclosure;

[0148] Fig.28 is a terminal structure diagram of a method for generating a character perspective video according to an embodiment of the present disclosure;

[0149] Fig.29 It is a server structure diagram of a method for generating a character perspective video according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0150] In order to make the purpose, technical solution and advantages of the present disclosure more clear, the present disclosure is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present disclosure and are not used to limit the present disclosure.

[0151] The following describes the system architecture and scenarios for application of the embodiments of the present disclosure.

[0152] Figure 11 is a system architecture diagram used by the method for generating a character perspective video according to an embodiment of the present disclosure, and includes an object terminal 140, the Internet 130, a gateway 120, a video processing server 110, and a database 150, etc.

[0153] The object terminal 140 includes various forms such as desktop computers, laptop computers, PDAs (personal digital assistants), tablet computers, mobile phones, vehicle-mounted terminals, home theater terminals, smart TVs, and dedicated terminals. In addition, it can be a single device or a collection of multiple devices. The object terminal 140 can communicate with the Internet 130 in a wired or wireless manner to exchange data. The object terminal 140 includes a multimedia platform, which is used to play the film and television works selected by the object. The multimedia platform is also used to receive the film and television works selected by the object, the video clips of the film and television works, and the background description information, and submit the video clips and background description information of the film and television works to the video processing server, so that the video processing server generates a character perspective video with a certain character in the film and television works as the main perspective.

[0154] The video processing server 110 refers to a computer system that can provide certain services to the object terminal 140. Compared with the ordinary object terminal 140, the video processing server 110 has higher requirements in terms of stability, security, performance, etc. The video processing server 110 can be a high-performance computer in a network platform, a cluster of multiple high-performance computers, a part of a high-performance computer (such as a virtual machine), a combination of parts of multiple high-performance computers (such as virtual machines), or a cloud server. The video processing server 110 includes various types of services, wherein the implementation of each service of the video processing server 110 is often associated with some intermediate databases or storage media. The video processing server 110 is used to generate a video script corresponding to the video clip according to the video clip of the film and television work selected by the object, and generate a supplementary script for a certain role using the video script and background description, and generate a role perspective video with the perspective of the role as the main perspective according to the background description submitted by the object, the supplementary script, and the role image of the role extracted from the database 150. The database 150 is used to store various images such as scene images of film and television works, character images of various characters, etc., wherein the database 150 can be set up separately or integrated in the video processing server 110 or other electronic devices.

[0155] The gateway 120 is also called an internetwork connector or a protocol converter. The gateway realizes network interconnection at the transport layer and is a computer system or device that acts as a converter. The gateway is a translator between two systems that use different communication protocols, data formats or languages, or even completely different architectures. At the same time, the gateway can also provide filtering and security functions. The message sent by the object terminal 140 to the video processing server 110 is sent to the corresponding server through the gateway 120. The message sent by the video processing server 110 to the object terminal 140 is also sent to the corresponding object terminal 140 through the gateway 120.

[0156] The embodiments of the present disclosure can be applied in various scenarios, such as Figure 2A-2E The video production scene shown, Figure 3A-3B The video playback scene shown, etc.

[0157] 1. Video production scene.

[0158] like Figure 2A As shown, when the subject wants to generate a role perspective video for a certain role in a certain film or TV work, the subject will log in to the multimedia platform on the subject terminal and enter the video production process. At this time, a prompt field "Please provide the following information" will be displayed on the page, and an editing area for entering the name of the film or TV series, an editing area for selecting a film or TV series clip, an editing area for entering a brief description of the film or TV series, an editing area for selecting the role corresponding to the target perspective, and an editing area for selecting the video storage location will be provided. Based on this, the object enters "Xxxx1" in the editing area for entering the name of the film and television drama; the object selects "D:\Film and Television Drama Xxxx1\Clip 1" in the editing area for selecting the film and television drama clip; the object enters "In the xxx era, character Axxx, character Bxx, xxxx" in the editing area for entering a brief description of the film and television drama, selects "Role B" in the editing area for selecting the character corresponding to the target perspective, selects "D:\Film and Television Drama Xxxx1\Character Video" in the editing area for selecting the video storage location, and clicks the "OK" button to determine the generation of a character perspective video with character B as the main perspective for clip 1 of the film and television drama Xxxx1, and store the generated character perspective video under the path "D:\Film and Television Drama Xxxx1\Character Video".

[0159] like Figure 2B As shown, after the subject clicks the "OK" button, a prompt window will be displayed on the page, wherein the prompt window has a prompt field "The character image of character B is being extracted from the database, please wait patiently..." to indicate to the subject the extraction process of the character image for character B before the video is generated.

[0160] like Figure 2CAs shown, after the character images are extracted, a prompt window will be displayed on the page, wherein the prompt window has a prompt field "10 character images of character B have been extracted, and the supplementary character script of character B is being generated, please wait patiently..." to indicate the extraction results of the character images to the subject.

[0161] like Figure 2D As shown, it is a schematic diagram of the multimedia platform generating a supplementary role script for character B in video generation. In the page, a prompt field is displayed, "The supplementary role script corresponding to character B in clip 1 has been generated, see the attached figure for details. In addition, the generated supplementary role script has been saved to the path: M:\Film and Television Scripts\Supplementary Role Scripts.", and the supplementary role script is displayed as "(It is late at night, and character B is the only one in the company's conference room. He sits at the table with project resources spread out in front of him, looking tired and solemn.); Character B: (talking to himself) What went wrong? Why is this happening? (He flips through the information, trying to find the source of the error)", to indicate to the subject the generated supplementary role script for character B.

[0162] like Figure 2E As shown, after the supplementary role script is generated, a prompt window will be displayed on the page, wherein the prompt window has a prompt field "Generating the role perspective video of role B, please wait patiently..." to indicate the generation process of the role perspective video to the subject.

[0163] (2) Video playback scenarios.

[0164] like Figure 3A As shown, the subject is using the multimedia application on the subject's terminal to play a video clip of a film and television work, which is Episode 01 of the film and television work named "xxx Season 2". The total length of the video clip is 47:21, and at this time, the video clip is currently played to the 37th second. The video frame at the 37th second is a video screen displayed from a third-party perspective, with characters 1, 2, and 3 located on the left side of the video screen, characters 4, 5, and 6 located on the right side of the video screen, and character 7 located directly in front of the video screen. Based on this, when the subject wants to switch the video screen from the third-party perspective to the perspective of character 1, the subject will click the "Multi-Perspective" control below the video screen, and select "Role 1 Perspective" from the multiple character perspectives displayed by triggering the "Multi-Perspective" control to switch the video screen from the third-party perspective to the perspective of character 1.

[0165] like Figure 3B As shown, when the subject selects "Perspective of Character 1", the video screen will switch from the third-party perspective to the perspective of Character 1. At this time, the video screen is that Character 7 is displayed on the left side of the video screen, and Characters 4, 5 and 6 are displayed in front of the video screen.

[0166] It should be noted that the character perspective video generation method of the embodiment of the present disclosure can be applied not only to the above-mentioned video production scenarios and video playback scenarios, but also to various application scenarios such as video editing scenarios in self-media and personalized video production scenarios in social media.

[0167] The following is a general description of the embodiments of the present disclosure.

[0168] According to an embodiment of the present disclosure, a method for generating a character perspective video is provided.

[0169] The character perspective video generation method is generally used in business scenarios where auxiliary videos viewed from different character perspectives need to be generated for film and television works, such as Figure 2A-2E The video production scene shown, Figure 3A-3B The video playback scene shown, etc. The disclosed embodiment provides a solution for obtaining a character perspective video from the perspective of a target character from a source video from a third-party perspective by integrating character script information, character picture information, and source background description information, which can improve the accuracy of generating auxiliary videos from different character perspectives for film and television works.

[0170] like Figure 4 As shown, the character perspective video generation method according to an embodiment of the present disclosure can be executed by an electronic device, which can be Figure 1 The video processing server or object terminal shown. A method for generating a role perspective video according to an embodiment of the present disclosure may include:

[0171] Step 410: Obtain a source video having a target character and source background description information of the source video;

[0172] Step 420: extracting source script information from the source video;

[0173] Step 430: Obtaining the introductory text template information corresponding to the role script conversion model;

[0174] Step 440: Generate role script information from the perspective of the target role using a role script conversion model based on the source script information, the source background description information and the guide template information;

[0175] Step 450: Generate diffusion guidance information based on the role script information, the image information corresponding to the target role, and the source background description information, and generate a role perspective video from the perspective of the target role based on the diffusion guidance information.

[0176] Steps 410 - 450 are described in detail below.

[0177] In step 410, a source video having a target character and source background description information of the source video are obtained.

[0178] The target character refers to a character in the source video, where the target character can be a person or an animal in the source video.

[0179] The source video refers to the original video data of the film or television work for which the character's perspective video needs to be generated.

[0180] The source background description information is used to briefly introduce the background of the entire story of the source video in text form. The source background description information is generally a simplified version of the entire video script of the source video, often only including key plots, historical background, location, cultural environment or summary information.

[0181] When this embodiment is implemented, since the source video is often determined and published by the producer of the source video, based on this, with authorization and permission, the source video can be downloaded from official channels or public platforms to obtain the source video with the target character.

[0182] Similarly, when producing and publishing the source video, in order to let the public know the general outline of the source video, the producer of the source video often determines the source background description information of the source video and publishes the source background description information of the source video. Based on this, with authorization, the source background description information of the source video can be obtained from official channels or credible public platforms.

[0183] In step 420, source script information is extracted from the source video.

[0184] The source script information refers to the script content that the source video relies on when shooting. The source script information often includes scene descriptions, lines and actions of each character, and so on.

[0185] To save space, the specific process of extracting source script information from the source video in the embodiment of the present disclosure will be described in detail below and will not be repeated here.

[0186] In step 430, the introductory text template information corresponding to the role-playing script conversion model is obtained.

[0187] The role-script conversion model refers to a large language model that can automatically generate a role script for a target character based on information such as video clips, background descriptions, and preset prompts of a film or television work.

[0188] The role-playing script conversion model in the disclosed embodiment may be composed of one language model or a collection of multiple language models, without limitation.

[0189] The introductory language template information is used to guide the role script conversion model to generate the framework of the role script information of the target role. The introductory language template information often includes prompt words related to the role name, emotional state, background information, scene setting, and plot development.

[0190] When this embodiment is implemented, due to different task requirements, the prompt information required by the role script conversion model to generate a specific text is different. In order to improve the efficiency of task processing, corresponding introductory language templates are often pre-set for various tasks, and multiple introductory language templates are stored in a template library. Based on this, with authorization, according to the task requirements of generating the role script information of the target role, the introductory language template information corresponding to the role script conversion model can be directly called from the template library.

[0191] In step 440, based on the source script information, the source background description information and the guide template information, the role script conversion model is used to generate the role script information from the perspective of the target character.

[0192] The role script information is used to indicate the script content generated to supplement the video content of the source video, with the target character's perspective as the main perspective. The role script information includes but is not limited to some possible plot developments, interactions between characters, etc. that are not in the source video.

[0193] To save space, the specific process of generating the role script information from the perspective of the target role based on the source script information, source background description information and guide template information in the embodiment of the present disclosure will be described in detail below and will not be repeated here.

[0194] In step 450, diffusion guidance information is generated based on the role script information, the picture information corresponding to the target role, and the source background description information, and a role perspective video from the perspective of the target role is generated based on the diffusion guidance information.

[0195] Among them, the step of generating a character perspective video from the perspective of the target character based on the diffusion guidance information may include but is not limited to: guiding the diffusion model based on the diffusion guidance information to generate a character perspective video from the perspective of the target character.

[0196] The diffusion guidance information is used to guide the video generation process of the diffusion model. The diffusion guidance information is used as a text condition constraint to assist the diffusion model in performing video frame denoising during video generation, so that the video frame effect meets the actual requirements.

[0197] The diffusion model (SD) of the embodiment of the present disclosure is a stable diffusion model from text to video. The diffusion model of the embodiment of the present disclosure often takes the background description of the film and television work, the text description of the character image, and the supplementary generated character script information as input, and outputs a video with the target character's perspective as the main perspective. The diffusion model can achieve the goal of obtaining the target character's perspective video from the source video of the third party's perspective.

[0198] A character perspective video refers to a video that uses the target character's perspective as the main perspective. The plot described in the character perspective video is similar to the plot described in the source video. The difference is that the character perspective video is presented from the perspective of the target character, while the source video is presented from a third-party perspective.

[0199] To save space, the specific process of generating diffusion guidance information based on role script information, image information corresponding to the target role, and source background description information, as well as the specific process of generating a role perspective video based on the diffusion guidance information to guide the diffusion model will be described in detail below and will not be elaborated here.

[0200] Through the above steps 410-450, the embodiment of the present disclosure extracts source script information with a target role from a source video. The source script information is described from the perspective of a third party other than any character in the source script information. The source background description information reflects the background in which the source script exists, and the psychological activities and actions generated by any character are closely related to the background. With the source script information and the source background description information, the source script can be adjusted using the role script conversion model, and the source script information from the perspective of the third party can be converted into the role script information from the perspective of the target character using the role script conversion model. The role script information reflects the adjustment of the psychological activities, actions generated, and pictures seen by the character under the source background description information. Then, based on the role script information, the picture information corresponding to the target character, and the source background description information, diffusion guidance information is generated, and based on the diffusion guidance information, a role perspective video from the perspective of the target character is generated. The picture information corresponding to the target character makes the image of the character in the role perspective video concrete, and the source background description information determines the background and color tone of the picture in the role perspective video. Therefore, the specific video is generated by integrating the role script information, role picture information and source background description information using the diffusion model, taking into account the influence of the role's appearance and the background of the picture on the specific content in the video, improving the accuracy of the generated role perspective video, thereby improving the information transmission efficiency of the role perspective video. The disclosed embodiment can obtain the role perspective video of the target role perspective from the source video of the third party perspective with high accuracy.

[0201] In addition, according to the embodiments of the present disclosure, the subject can understand the story of the film or television work to which the source video and the character perspective video belong from the perspective of different characters that the subject prefers, which can deepen the subject's understanding and involvement in the plot and increase the subject's immersion in watching these narrative film or television works, thereby increasing the user stickiness of the film or television works.

[0202] The above is a general description of steps 410 to 450. Since steps 410 and 430 have been described in detail in the above general description, the specific implementation of steps 420, 440 and 450 will be described in detail below.

[0203] Step 420 is described in detail below.

[0204] In step 420, source script information is extracted from the source video.

[0205] Please refer to Figure 5 In one embodiment, step 420 specifically includes but is not limited to the following steps 510-530:

[0206] Step 510: segment the source video to obtain a plurality of source video frames arranged in time sequence;

[0207] Step 520: Input the plurality of source video frames into a pre-trained video script description generation model for description generation, and obtain video frame script description information of the plurality of source video frames;

[0208] Step 530: Integrate the video frame script description information of each of the multiple source video frames to obtain source script information.

[0209] Steps 510 - 530 are described in detail below.

[0210] In step 510, the source video frame refers to a video segment in the source video.

[0211] In the specific implementation of this embodiment, first, the video frame length of the source video frame to be segmented is set. Then, according to the set video frame length, a command line tool is used to segment the source video into multiple source video frames arranged in time sequence, and the multiple source video frames are respectively saved in image form, wherein the command line tool of the embodiment of the present disclosure can be FFmpeg.

[0212] In step 520, the video script description generation model refers to a neural network model that can generate a video script corresponding to the video of the film and television work. The video script description generation model of the embodiment of the present disclosure often takes the source video of the film and television work as input and outputs the script content corresponding to the source video of the film and television work.

[0213] Furthermore, the video script description generation model of the embodiment of the present disclosure may be a deep learning model (Vision Transformer, ViT) based on a transformer architecture.

[0214] The video frame script description information is used to indicate the script content corresponding to a single source video frame.

[0215] To save space, the specific process of inputting multiple source video frames into a pre-trained video script description generation model for description generation in the embodiment of the present disclosure will be described in detail below and will not be repeated here.

[0216] In step 530, in terms of time sequence, the video frame script description information of each of the multiple source video frames is spliced ​​in sequence to achieve integration of the multiple video frame script description information to obtain the source script information.

[0217] The benefit of this embodiment is that when generating source script information, the source video is first divided into multiple source video frames, and the script is generated with the source video frames as the basic unit, which can achieve fine-grained processing of the source video and make the generated source script information more accurate. Furthermore, when specifically generating the video frame script description information corresponding to each source video frame, it is achieved by relying on a pre-trained video script description generation model, which can improve the accuracy and efficiency of generating video frame script description information to a certain extent, and thus improve the degree of automation of character video generation.

[0218] In the disclosed embodiment, the video script description generation model includes a feature extraction network and a script description generation network. The feature extraction network is used to extract video feature information of each source video frame of the source video. The script description generation network is used to generate script description content corresponding to the source video frame according to the video feature information of each source video frame.

[0219] Please refer to Figure 6 In one embodiment, step 520 specifically includes but is not limited to the following steps 610-620:

[0220] Step 610: For each source video frame, extract features of the source video frame based on a feature extraction network to obtain target video frame features of the source video frame;

[0221] Step 620: Based on the target video frame features, a script description generation network is used to generate a script to obtain video frame script description information of the source video frame.

[0222] Steps 610-620 are described in detail below.

[0223] In step 610, the target video frame feature is used to indicate video feature information contained in a single source video frame.

[0224] When this embodiment is implemented, the feature extraction network may be a convolutional neural network structure. Specifically, for each source video frame, the source video frame is input into the feature extraction network, and multiple convolutional layers of the feature extraction network sequentially extract features from the source video frame, each convolutional layer extracts video feature information of the source video frame at different levels, and the output result of the last convolutional layer is used as the target video frame feature of the source video frame.

[0225] It should be noted that when multiple convolutional layers of the feature extraction network extract features from the source video frame in sequence, the input of the first convolutional layer is the source video frame, and the output is the video frame features corresponding to the source video frame; the input of the second convolutional layer is the video frame features output by the first convolutional layer, and the output is the video frame features, and so on. The video frame features output by the second to last convolutional layer are used as the input of the last convolutional layer, and the last convolutional layer outputs the target video frame features.

[0226] In step 620, the script description generation network can be a Transformer-based decoder structure, wherein the script description generation network is composed of an attention layer and a feed-forward layer.

[0227] In the specific implementation of this embodiment, for each source video frame, the target video frame features are input into the script description generation network. The attention layer of the script description generation network first performs self-attention calculation on the target video frame features to obtain the video frame attention features, and then the feedforward layer performs feedforward processing on the video frame attention features to achieve nonlinear transformation of the video frame attention features to obtain the transformed video frame features. Finally, the preset function of the script description generation network calculates the probability of each feature element in the transformed video frame feature pointing to each reference word in the preset vocabulary, and uses the reference word with the highest probability as the predicted word corresponding to the feature element, and splices the predicted words corresponding to each feature element together, thereby obtaining the video frame script description information of the source video frame. Among them, the preset function is generally an activation function such as a softmax function that can realize feature classification. The prediction vocabulary is a list consisting of multiple reference words that is pre-set according to business needs.

[0228] Among them, when using the attention layer to perform self-attention calculation on the target video frame features, the target video frame features are first linearly projected, and the query channel of the attention layer is used to project the video frame features into query features, the key channel of the attention layer is used to project the video frame features into key features, and the value channel of the attention layer is used to project the video frame features into value features. Then, the query feature and the transposed result of the key feature are inner-product processed to obtain the inner-product result, and the inner-product result is divided by the square root of the feature dimension of the key feature to obtain the attention score. Furthermore, the softmax function is used to convert the attention score into an attention weight, and the attention weight and the value feature are multiplied to obtain the video frame attention feature.

[0229] The benefit of this embodiment is that the script is generated by using the video script description generation model, and the automatic feature extraction and script description generation can significantly improve the generation efficiency of video frame script description information, and reduce the time cost of manual editing and creation. At the same time, the target video frame features of each source video frame are extracted by the feature extraction network, which can capture the spatial features (such as edges, textures, colors, etc.) in the source video frame, improve the consistency of the feature description of each source video frame in quality, and reduce the errors of human operation; at the same time, the attention operation and feedforward operation are performed on the extracted target video frame features, which can capture the dependency relationship between the various feature elements in the source video frame for analysis, and at the same time, the prediction words corresponding to each feature element are generated in combination with the probability calculation method, which can improve the accuracy of the generated video frame script description information.

[0230] Since the target video frame features in the above embodiment are the output of the last convolutional layer of the feature extraction network, this will make the feature information of the target video frame features not comprehensive enough, which will affect the accuracy of the generated video frame script description information. Based on this, the embodiment of the present disclosure provides a solution for extracting features from source video frames based on feature fusion, which can improve the feature comprehensiveness and feature richness of the target video frame features.

[0231] Please refer to Figure 7 In one embodiment, step 610 specifically includes but is not limited to the following steps 710-740:

[0232] Step 710: for each source video frame, perform multi-scale downsampling on the source video frame based on a feature extraction network to obtain a first downsampling feature, a second downsampling feature, and a third downsampling feature;

[0233] Step 720: Perform dimensionality reduction processing on the first down-sampled feature, the second down-sampled feature, and the third down-sampled feature, respectively, to obtain a first dimensionality reduction feature corresponding to the first down-sampled feature, a second dimensionality reduction feature corresponding to the second down-sampled feature, and a third dimensionality reduction feature corresponding to the third down-sampled feature;

[0234] Step 730: performing splicing processing on the first dimensionality reduction feature, the second dimensionality reduction feature, and the third dimensionality reduction feature to obtain an initial video frame feature of the source video frame;

[0235] Step 740: perform splicing processing on the initial video frame features of the source video frame and the initial video frame features of the previous source video frame of the source video frame to obtain the target video frame features of the source video frame.

[0236] Steps 710 - 740 are described in detail below.

[0237] In step 710, the first down-sampling feature, the second down-sampling feature and the third down-sampling feature respectively indicate video feature information of the source video frame at different scales, wherein there are certain differences in feature information contained in the first down-sampling feature, the second down-sampling feature and the third down-sampling feature.

[0238] Among them, the feature dimensions of the first down-sampling feature, the second down-sampling feature and the third down-sampling feature increase successively.

[0239] In the specific implementation of this embodiment, for each source video frame, multiple convolutional layers based on the feature extraction network sequentially perform convolution processing on the source video frame features, and each convolutional layer takes the output features of the previous convolutional layer as input and outputs a new video frame feature. Based on this, three convolutional layers can be selected from the multiple convolutional layers as target layers, and the output features of the three target layers are respectively used as the first down-sampling feature, the second down-sampling feature, and the third down-sampling feature.

[0240] For example, in the extraction network including six convolutional layers, the feature size of the output feature of the first convolutional layer is 1 / 2 of the video frame size of the source video frame; the feature size of the output feature of the second convolutional layer is 1 / 4 of the video frame size of the source video frame; the feature size of the output feature of the third convolutional layer is 1 / 8 of the video frame size of the source video frame; the feature size of the output feature of the fourth convolutional layer is 1 / 16 of the video frame size of the source video frame; the feature size of the output feature of the fifth convolutional layer is 1 / 32 of the video frame size of the source video frame; the feature size of the output feature of the sixth convolutional layer is 1 / 64 of the video frame size of the source video frame. Based on this, the output feature of the third convolutional layer is determined as the first down-sampling feature, the output feature of the fourth convolutional layer is determined as the second down-sampling feature, and the output feature of the fifth convolutional layer is determined as the third down-sampling feature.

[0241] It should be noted that the output features selected in the above examples are from three consecutive convolutional layers. In other specific examples, the extracted multiple down-sampled features may also come from multiple discontinuous convolutional layers.

[0242] In step 720, since the feature dimensions of the first downsampled feature, the second downsampled feature, and the third downsampled feature are different. In order to fuse the feature information contained in the first downsampled feature, the second downsampled feature, and the third downsampled feature, it is necessary to map the first downsampled feature, the second downsampled feature, and the third downsampled feature to a preset one-dimensional vector space respectively, convert the first downsampled feature, the second downsampled feature, and the third downsampled feature from multi-dimensional arrays into one-dimensional arrays respectively, and obtain the first dimensionality-reduced feature corresponding to the first downsampled feature, the second dimensionality-reduced feature corresponding to the second downsampled feature, and the third dimensionality-reduced feature corresponding to the third downsampled feature. Among them, the first dimensionality-reduced feature, the second dimensionality-reduced feature, and the third dimensionality-reduced feature are all one-dimensional feature vectors.

[0243] In step 730, the initial video frame feature is used to indicate the comprehensive feature information of the source video frame at multiple scales.

[0244] In the specific implementation of this embodiment, since the first dimensionality-reduced feature, the second dimensionality-reduced feature, and the third dimensionality-reduced feature are features of the same feature dimension. Based on this, the first dimensionality-reduced feature, the second dimensionality-reduced feature, and the third dimensionality-reduced feature can be directly concatenated to obtain the initial video frame feature of the source video frame.

[0245] In step 740, the method for obtaining the initial video frame feature of the previous source video frame of the source video frame is basically the same as that in steps 710-730, and will not be elaborated.

[0246] In the specific implementation of this embodiment, since the initial video frame feature of the source video frame and the initial video frame feature of the previous source video frame of the source video frame are features of the same feature dimension. Based on this, the initial video frame feature of the source video frame and the initial video frame feature of the previous source video frame of the source video frame can be directly concatenated to obtain a concatenated feature with a longer feature length, and the obtained concatenated feature is used as the target video frame feature of the source video frame.

[0247] In other embodiments, the initial video frame feature of the source video frame and the initial video frame feature of the previous source video frame of the source video frame can also be directly added vectorially to obtain the target video frame feature of the source video frame.

[0248] It should be noted that if the current source video frame is the first source video frame of the source video, then when concatenating the initial video frame feature of the source video frame and the initial video frame feature of the previous source video frame of the source video frame, the value of the initial video frame feature of the previous source video frame of the source video frame can be set to zero, and the initial video frame feature of the current source video frame can be directly used as the target video frame feature of the source video frame.

[0249] The advantage of this embodiment is that when extracting features from the source video frames, feature fusion of the output features of multiple convolutional layers of the feature extraction network is considered. At the same time, considering that the feature dimensions of the output features of different convolutional layers are different and cannot be directly fused, downsampled feature maps of multiple scales are mapped to the same one-dimensional vector space, and the first dimension-reduced feature, the second dimension-reduced feature, and the third dimension-reduced feature are concatenated into the initial video frame feature of the source video frame, which can improve the feature comprehensiveness and richness of the initial video frame feature. At the same time, due to the temporal dependence relationship between source video frames, the present disclosure embodiment also considers using the concatenation result of the initial video frame feature of the current source video frame and the initial video frame feature of the previous source video frame as the target video frame feature. When extracting, the feature dependence relationship and feature correlation of the video frame features of each source video frame in time series are considered, which can better improve the comprehensiveness and accuracy of the feature information of the extracted target video frame feature.

[0250] Since the video script description generation model is often trained based on information such as videos and background descriptions corresponding to various categories of film and television works, when the video script description generation model generates video frame script description information, it tends to meet the most basic requirements, rather than focusing on the video frame script description requirements in a specific field. However, the video scripts of different fields or different categories of film and television works often have different characteristics. Determining the source script information directly based on the video frame script description information generated by the video script description generation model will result in poor accuracy and coherence of the source script information. Based on this, the present disclosure embodiment provides a solution for polishing the video frame script description information, which can improve the accuracy and coherence of the source script information.

[0251] Please refer to Figure 8 , in one embodiment, step 530 specifically includes but is not limited to the following steps 810-820:

[0252] Step 810, concatenate the video frame script description information of each of the multiple source video frames to obtain preliminary script information;

[0253] Step 820, polish the preliminary script information to obtain source script information.

[0254] The following is a detailed description of steps 810-820.

[0255] In step 810, the preliminary script information is used to indicate the overall script content corresponding to the source video generated by the video script description generation model according to the video content of the source video.

[0256] In the specific implementation of this embodiment, according to the chronological order of multiple source video frames, the multiple video frame script description information is spliced in this chronological order to form an overall description information, and the obtained overall description information is used as the preliminary script information.

[0257] In step 820, in order to improve the polishing efficiency and quality, when polishing the preliminary script information, the preliminary script information and a preset prompt can be input into a preset large language model, and the large language model automatically polishes the preliminary script information. Among them, the preset prompt can be "Please polish the following plot description to make it more coherent and engaging: [plot description text]" and so on. Specifically, first, preprocess the preliminary script information. The preprocessing steps include but are not limited to the following steps: removing redundant spaces in the preliminary script information, correcting punctuation marks, unifying terms and styles in the entire preliminary script information, etc., to improve the consistency of the entire preliminary script information. Then, input the preprocessed preliminary script information and the preset prompt word into the large language model, and the large language model performs multiple iterative polishings and outputs a polished preliminary script information. Further, in order to make the script description conform to a specific style or requirement, the review end makes adjustments such as addition, deletion, and modification to the sentence structure and content details of the polished preliminary script information, so as to obtain the source script information with no logical errors and coherent plot.

[0258] The advantage of this embodiment is that it takes into account the polishing of the video frame script description information corresponding to the source video frames, and also takes into account the overall coherence of the source script information. First, the video frame script description information corresponding to each of the multiple source video frames is spliced into an overall one to obtain the preliminary script information, and then the preliminary script information is polished. In this way, compared with polishing each video frame script description information separately, the large language model can pay attention to the dependency and correlation relationships of the video frame script description information during polishing, thereby improving the accuracy of the polishing process, and further making the generated source script information have better accuracy and coherence.

[0259] The following combines Fig. 9 An example description is given of the specific process of extracting source script information from the source video in the embodiments of the present disclosure.

[0260] Such as Fig. 9As shown, for the source video frame in time window t, first use multiple convolutional layers of the feature extraction network to perform multi-scale feature extraction on the source video frame to obtain the first downsampled feature, the second downsampled feature, and the third downsampled feature. Then, concatenate the first downsampled feature, the second downsampled feature, and the third downsampled feature into the initial video frame feature, where the concatenation process includes two steps: dimensionality reduction processing and concatenation processing. Further, concatenate the initial video feature of the source video frame in time window t and the initial video frame feature of the source video frame in time window t-1 to obtain the target video frame feature of the source video frame in time window t. Then, input the target video frame feature into the script description generation network, and perform feature processing on the target video frame feature through the attention layer and the feedforward layer in sequence to obtain the video frame script description information corresponding to the source video frame in time window t. Further, input the video frame script description information of each of the multiple source video frames within the entire time interval (from 0 to t) into the large language model for polishing to obtain the character script information of the source video.

[0261] The following details step 440.

[0262] In step 440, based on the source script information, the source background description information, and the guiding language template information, use the character script conversion model to generate the character script information from the perspective of the target character.

[0263] In the embodiments of the present disclosure, the character script conversion model can be a general large language model, where the character script conversion model includes a first sub-model and a second sub-model.

[0264] The first sub-model is used to fill the input guiding language template information according to the input source script information and source background description information to form a complete script prompt instance.

[0265] The second sub-model is used to automatically generate a complete character script according to the complete script prompt instance.

[0266] In one embodiment, step 440 specifically includes but is not limited to the following steps:

[0267] Based on the guiding language template information, the source script information, and the source background description information, generate a character script prompt instance through the first sub-model;

[0268] Based on the character script prompt instance, generate character script information through the second sub-model.

[0269] Among them, the character script prompt instance is used to indicate the overall outline content for generating the character script information. The character script prompt instance often consists of multiple parts, and the specific content of each part is refined according to the source script information and the source background description information.

[0270] In the specific implementation of this embodiment, first, the guiding language template information, the source script information, and the source background description information are input into the character script conversion model. The first sub-model first determines the main composition of the guiding language template information, divides the guiding language template information into multiple template modules, and determines the guiding language corresponding to each template module. Then, for each template module, according to the guiding language, the script elements corresponding to the guiding language are extracted from the source script information and the source background description information, and the extracted script elements are filled into the template module. Thus, a character script prompt instance is obtained based on multiple template modules filled with script elements. Further, the second sub-model automatically creates complete script elements such as character dialogues and plot developments that meet the specific requirements of the character script prompt instance, and expands the character script prompt instance into a coherent and logical script text, thereby obtaining the character script information.

[0271] As Fig.10 shown, it is a specific schematic of generating character script information based on the character script conversion model. Specifically, the character script conversion model consists of 5 language models. Language model 1 can be a large language model ChatGPT developed by a certain entity, and language model 1 is mainly used to process the input source script information; language model 2 can be a language model Gemini developed by another entity, and language model 2 is mainly used to process the input source background description information; language model 3 can be the language model Claude, language model 4 can be the language model Coze, and language model 5 can be the language model Llama. Language models 3, 4, and 5 are mainly used to process the input guiding language template information. Therefore, after the guiding language template information, the source script information, and the source background description information are input into the character script conversion model, first, language models 3, 4, and 5 generate the guiding language according to the guiding language template information, the source script information, and the source background description information. Then, language models 1 and 2 extract the script elements corresponding to the guiding language from the source script information and the source background description information to generate the character script information.

[0272] The advantage of this embodiment is that when generating the character script information, the first sub-model in the character script conversion model can automatically group the source script information and the source background description information according to the modules set in the guiding language template information and the guiding language of each module, and fill the source script information and the source background description information into each module in the guiding language template information to form a script prompt instance corresponding to the source video. This script prompt instance is the overall outline of a generated script. Further, the second sub-model in the character script conversion model will automatically generate character script content with character dialogues, descriptions of character mental activities, and scene descriptions. This method generates character script information in two stages, allowing the script prompt instance generated in the first stage to be modified and adjusted when generating the character script information, improving the flexibility and accuracy of generating the character script information.

[0273] In a specific example, the preset guiding language template information can be shown in Table 1. The script prompt instance generated according to the guiding language template information, the source script information, and the source background description information can be shown in Table 2. The character script information generated according to the script prompt instance can be shown in Table 3.

[0274] Table 1 - Guiding Language Template Information

[0275]

[0276] Table 2 - Script Prompt Instance

[0277]

[0278] Table 3 - Character Script Information

[0279]

[0280] Since the lines and mental activities of different character images in the same plot scene often vary greatly. If the character script conversion model generates the character script without knowing the character characteristics such as the character personality and appearance of the target character, it often leads to the generated character script information not matching the character characteristics of the target character. Based on this, in order to improve the accuracy of the generated character script information, the embodiments of the present disclosure provide a solution for generating character script information based on character images, which can make the character lines and character actions in the character script information more conform to the character characteristics of the target character, thereby improving the accuracy of the generated character script information.

[0281] Please refer to Fig.11 In another embodiment, step 440 specifically includes but is not limited to the following steps 1110 - 1120:

[0282] Step 1110: Obtain the picture information corresponding to the target character;

[0283] Step 1120: Based on the source script information, picture information, and source background description information, use the guiding statement template information to generate a guiding statement, and input the guiding statement into the character script conversion model to obtain the character script information.

[0284] The following is a detailed description of steps 1110-1120.

[0285] In step 1110, the picture information is used to indicate the visual information of the target character in various scenes and from various perspectives. Among them, the picture information includes feature information such as the makeup, facial expressions, and body movements of the target character in various scenes and from various perspectives.

[0286] In the specific implementation of this embodiment, since the characteristics reflecting the character traits of the target character, such as the clothing, hairstyle, expression, and accessories of the target character, are intuitively reflected in the character image of the target character, and the character image of the target character is often stored in the database. Based on this, with the permission of authorization, extract the character images of the target character in different perspectives and different scenes from the database, and integrate the obtained multiple character images to obtain the picture information corresponding to the target character.

[0287] In step 1120, first, integrate the source script information, picture information, and source background description information to obtain the integrated information. Then, use the guiding statement template information and combine it with the integrated information to generate a more specific guiding statement. Further, input the generated guiding statement into the character script conversion model, so that the character script conversion model generates the character script information according to the guiding statement and combines the various knowledge it has learned.

[0288] The advantage of this embodiment is that when generating the character script information, considering the influence of the character image of the target character on the character lines and character actions of the target character, input the picture information, source script information, and source background description information describing the character traits of the target character into the character script conversion model together, making the generated guiding statement more accurate. In this way, when the character script conversion model generates the character script, it will customize the script according to the appearance characteristics of the target character, and construct a more vivid and three-dimensional character image in the character script information by means of narration description, etc. according to the character introduction, etc., so that the character lines and character actions in the generated character script information are more in line with the characteristics of the target character, and can improve the accuracy of the generated character script information to a certain extent.

[0289] Since the role script information contains various types of script elements, for example, character lines, scene descriptions, story backgrounds, character introductions, etc. all belong to the category of role script information. In order to generate various types of script elements in a targeted manner, the embodiments of the present disclosure provide a solution for parallelly generating multiple types of script information, which can improve the generation accuracy of role script information.

[0290] In the embodiments of the present disclosure, the source script information includes source line information and source scene information, the role script information includes role script line information and role script scene information, the guiding language template information includes first guiding language template information and second guiding language template information, and the guiding language includes a first guiding language and a second guiding language.

[0291] The source line information is used to indicate the dialogue lines, monologue lines, etc. of each character in the source video.

[0292] The source scene information is used to indicate what the scenes are like when each plot in the source video occurs.

[0293] The role script line information is used to indicate the dialogue lines and monologue lines of each character in the role perspective video with the target character as the main perspective.

[0294] The role script scene information is used to indicate what the scenes are like when each plot in the role perspective video with the target character as the main perspective occurs.

[0295] The first guiding language template information is used to indicate the specific structure and content of the guiding language for generating character lines.

[0296] The second guiding language template information is used to indicate the specific structure and content of the guiding language for generating scene information.

[0297] The first guiding language is used to guide the role script conversion model to generate the role line information of the role perspective video with the target character as the main perspective.

[0298] The second guiding language is used to guide the role script conversion model to generate the scene information of the role perspective video with the target character as the main perspective.

[0299] Please refer to Fig.12 , in another embodiment, step 1120 specifically includes but is not limited to the following steps 1210-1220:

[0300] Step 1210, based on the source line information, source scene information, picture information, and source background description information, use the first guiding language template information to generate a first guiding language, and input the first guiding language into the role script conversion model to obtain the role script line information;

[0301] Step 1220: Based on the source scenario information and the source background description information, use the second guiding language template information to generate the second guiding language, and input the second guiding language into the character script conversion model to obtain the character script scenario information.

[0302] The following is a detailed description of steps 1210 - 1220.

[0303] In step 1210, based on the source line information, source scenario information, picture information, and source background description information, use the first guiding language template information to generate the guiding language. Extract the information related to the character lines from the source line information, source scenario information, picture information, and source background description information and fill in each part of the first guiding language template to generate the first guiding language. Then, input the first guiding language into the character script conversion model, so that the character script conversion model generates the character script line information based on the first guiding language using its own corpus foundation.

[0304] The specific process of step 1220 is similar to the above step 1210. The difference is that the generation of the character script scenario information does not depend on the picture information and the source line information, but the generation of the character script line information depends on the source scenario information. To save space, it will not be elaborated here.

[0305] As Fig.13 shown, when generating the character script information, the generation of the character script line information and the character script scenario information is executed simultaneously. Specifically, the first branch is to generate the guiding language based on the source line information, source scenario information, picture information, and source background description information using the first guiding language template information. Extract the information related to the character lines from the source line information, source scenario information, picture information, and source background description information and fill in each part of the first guiding language template to generate the first guiding language, and input the first guiding language into the character script conversion model, so that the character script conversion model generates the character script line information based on the first guiding language. The second branch is to generate the guiding language based on the source scenario information and the source background description information using the second guiding language template information. Extract the information related to the scenario from the source scenario information and the source background description information and fill in each part of the second guiding language template to generate the second guiding language, and input the second guiding language into the character script conversion model, so that the character script conversion model generates the character script scenario information based on the second guiding language.

[0306] The advantage of this embodiment is that since the character lines and scene information in the character script information are two relatively important parts, for these two different types of script information, two different guiding language template information are considered to be set separately: the first guiding language template information applicable to generating character script lines and the second guiding language template information applicable to generating character script scenes; further, based on the first guiding language template information and the second guiding language template information, the character script conversion model is used to generate the character script line information and the character script scene information in parallel. At the same time, when generating the character script line information and the character script scene information, it is considered that there will be differences in the description information associated with the lines and scenes. The source line information, source scene information, picture information, and source background description information are used to generate the lines, while only the source scene information and the source background description information are used to generate the description information, so as to eliminate the interference of the source line information and the picture information on the generation of the character scene information while optimizing the generation of the character lines using the source line information and the picture information. This method can improve the generation accuracy of the character script line information and the character script scene information in the character script information to a certain extent.

[0307] The following details step 450.

[0308] In step 450, based on the character script information, the picture information corresponding to the target character, and the source background description information, diffusion guiding information is generated, and based on the diffusion guiding information, a character perspective video from the target character's perspective is generated.

[0309] Please refer to Fig.14 , in one embodiment, the diffusion guiding information is generated in the following manner:

[0310] Step 1410: Perform encoding processing on the source background description information to obtain background description encoding features;

[0311] Step 1420: Perform encoding processing on the character script information to obtain character script encoding features;

[0312] Step 1430: Extract features from the picture information to obtain target character image features;

[0313] Step 1440: Based on the background description encoding features, the character script encoding features, and the target character image features, perform conditional information generation through a preset conditional fusion network to obtain diffusion guiding information.

[0314] The following details steps 1410 - 1440.

[0315] In step 1410, the background description encoding features are used to represent the plot background of the source video in the form of vector features.

[0316] In the specific implementation of this embodiment, the source background description information can be encoded based on a preset text encoder to obtain the role script encoding features.

[0317] In the embodiment of the present disclosure, the text encoder includes a tokenizer, an embedding layer, and a text attention calculation module (text transformer).

[0318] Specifically, first, the source background description information is input into the tokenizer, and the source background description information is tokenized based on the tokenizer according to spaces, punctuation marks, or delimiters to obtain multiple words. Further, stop words in the multiple words are removed to obtain multiple descriptive words, and the index corresponding to each descriptive word is found in a preset index table, and the descriptive words are represented by the indexes. The preset index table is used to indicate the correspondence between words and indexes. Then, the indexes of each descriptive word are input into the embedding layer for word embedding processing to obtain the descriptive word embedding features corresponding to each descriptive word. Finally, the descriptive word embedding features of each descriptive word are input into the text attention calculation module for attention calculation, and the role script encoding features corresponding to the source background description information are output by the text attention calculation module.

[0319] In step 1420, the role script encoding features are used to represent the supplementary script content that the target role needs to perform in the form of vector features.

[0320] In the specific implementation of this embodiment, the specific process of step 1420 is similar to the above step 1410. For the sake of brevity, it will not be elaborated here.

[0321] In step 1430, the target role image features are used to represent the role characteristics of the target role in the form of vector features, where the role characteristics include but are not limited to the expression characteristics, makeup and styling characteristics, and action characteristics of the target role, etc.

[0322] In the specific implementation of this embodiment, the image information can be feature-extracted based on a preset image encoder to obtain the target role image features. The preset image encoder can be an image encoder based on the transformer structure (Vision Transformer, ViT).

[0323] Specifically, when extracting features from picture information based on an image encoder, first, the picture information is segmented into multiple image patches of the same area. Then, the image encoder is used to perform attention calculation and feed-forward processing on each image patch to obtain the image patch encoding features corresponding to each image patch. Finally, the encoding features of multiple image patches are concatenated into the target character image features. Among them, the specific processes of attention calculation and feed-forward processing are similar to the self-attention calculation and feed-forward processing in step 620 above. The difference is that the input of the feed-forward processing in step 1430 is the concatenation result of the attention calculation result and the image patch encoding features; while the input of the feed-forward processing in step 620 is the video frame attention feature (self-attention calculation result). For the sake of brevity, it will not be elaborated here.

[0324] In step 1440, the conditional fusion network refers to a neural network structure used to integrate multiple conditional constraint vectors into a whole.

[0325] In the specific implementation of this embodiment, since the feature dimensions of the background description encoding features, the character script encoding features, and the target character image features may vary. To integrate multiple conditional constraints, first, the background description encoding features, the character script encoding features, and the target character image features are input into the conditional fusion network. The conditional fusion network first maps the background description encoding features, the character script encoding features, and the target character image features into the same feature vector space to obtain the background description encoding features, the character script encoding features, and the target character image features with the same feature dimension. Then, the conditional fusion network will perform feature fusion on the background description encoding features, the character script encoding features, and the target character image features with the same feature dimension by means of feature concatenation or feature addition, etc., to integrate multiple conditional constraint vectors into a whole, thereby obtaining the diffusion guidance information. Among them, the diffusion guidance information is often represented in the form of a vector.

[0326] The advantage of this embodiment is that the three types of feature information of the character script information, the picture information corresponding to the target character, and the source background description information are aggregated into a whole and jointly used as the conditional constraints for denoising processing, which can improve the comprehensiveness and effectiveness of the content of the diffusion guidance information. This method takes into account the influence of the appearance of the character and the background of the picture on the specific content in the video, so that under the guidance of the conditional guidance information, the diffusion model can generate a more accurate character perspective video.

[0327] In the embodiments of the present disclosure, the diffusion model includes an encoding network, a decoding network, a diffusion network, and a denoising network. Among them, the encoding network is used to encode a random noise map into encoded features corresponding to the random noise map. The diffusion network is used to gradually add noise to the encoded features until the input encoded features approach pure noise. The denoising network is used to gradually denoise the latent space feature vectors output by the diffusion network under given conditional constraints until a denoising result that meets the conditional constraints is obtained. The decoding network is used to convert the character video frame features in the denoising result from the latent vector space to the pixel space, thereby generating a character perspective video.

[0328] In one embodiment, the process of guiding the diffusion model based on diffusion guidance information to generate a character perspective video of a target character perspective includes but is not limited to the following steps:

[0329] Encode the random noise map based on the encoding network to obtain target video frame encoded features;

[0330] Perform diffusion processing on the target video frame encoded features based on the diffusion network to obtain target latent space feature vectors;

[0331] Based on the diffusion guidance information, denoise the target latent space feature vectors through the denoising network to obtain video frame denoised features;

[0332] Decode the video frame denoised features based on the decoding network to obtain multiple character video frames, and generate a character perspective video of the target character perspective based on the multiple character video frames.

[0333] Specifically, the specific process of guiding the diffusion model based on the diffusion guidance information to generate a character perspective video of the target character perspective in this embodiment is similar to steps 1810-1840 of the embodiments of the present disclosure. The specific processes of steps 1810-1840 will be described in detail below. For the sake of brevity, it will not be elaborated here.

[0334] It should be noted that the random noise map can be generated by a library function of a random number generator (for example, the numpy.random.normal() function). According to the Gaussian distribution with a mean of 0 and a standard deviation of 1, a random number is generated, and this random number follows the Gaussian distribution. Further, converting this random number that follows the Gaussian distribution into the form of a random noise map, the random noise map required for video generation by the diffusion model is obtained.

[0335] As Fig.15 shown, it is the overall flowchart of video generation based on the diffusion model in the embodiments of the present disclosure. Specifically, the character script information, the picture information corresponding to the target character, and the source background description information are used as the inputs of the diffusion model.

[0336] First, a random noise map is generated using a random number i.

[0337] Next, the encoding network is used to encode the random noise map to obtain the encoded features Z of the target video frame, and the diffusion network is used to perform diffusion processing on the encoded features Z of the target video frame. The diffusion network adds noise to the encoded features Z of the target video frame T times to obtain the target latent space feature vector Z T , and its specific process is similar to the following step 1810. Further, the source background description information, the character script information, and the picture information are used as conditional constraints and converted into a form that meets the input requirements of the conditional fusion network τ, and the source background description information, the character script information, and the picture information are input into the conditional fusion network, and the conditional fusion network outputs the diffusion guidance information, and its specific process is similar to the above steps 1410-1440.

[0338] Next, when the denoising network denoises the target latent space feature vector Z T , first, based on the diffusion guidance information, the target latent space feature vector Z T is downsampled through the downsampling module of the denoising network to obtain the target downsampling result. Then, based on the diffusion guidance information, the target downsampling result is upsampled through the upsampling module of the denoising network to obtain the denoising result Z after one denoising operation T-1' . Further, the denoising network is used to denoise the denoising result Z at the prediction time step T T-1' according to the above process. After T-1 denoising operations, the video frame denoising features Z ’ are obtained. Finally, based on the decoding network, the video frame denoising features Z ’ are decoded to obtain multiple character video frames I, and a character perspective video from the target character's perspective is generated based on the multiple character video frames I, and its specific process is similar to the following steps 1820-1840. For the sake of brevity, it will not be elaborated here.

[0339] The advantage of this embodiment is that the diffusion network is first used to add noise to the encoded features of the target video frame corresponding to the random noise map multiple times to obtain the target latent space feature vector. Further, the denoising network is used to denoise the target latent space feature vector multiple times to realize the feature prediction of the character perspective video to be generated; in this process, the diffusion guidance information is used to fine-tune the predicted video frame denoising features to realize the adjustment of the feature information in terms of image and text, so that the information such as character lines and character scenes in the finally predicted denoising result is more accurate. Finally, the decoding network is used to decode the video frame denoising features, thereby obtaining the character perspective video from the target character's perspective. This method can obtain the character perspective video from the target character's perspective from the source video from a third-party perspective with relatively high accuracy.

[0340] Next, in combination with Fig.16 An example description is given of the overall process of generating a character perspective video from a source video for a target character's perspective.

[0341] As Fig.16 shown, when generating a character perspective video from a source video for a target character's perspective, first, multiple source video frames of the source video are respectively input into a video script description generation model. The feature extraction network and the script description generation network of the video script description generation model are used to sequentially perform feature extraction and script description generation on the source video frames, obtaining video frame script description information for each of the multiple source video frames, and integrating the multiple video frame script description information into source script information corresponding to the source video. Further, the source script information, source background description information, and guiding language template information are input into a character script conversion model assembled by language model 1, language model 2, language model 3, language model 4, and language model 5, and the character script conversion model outputs the character script information of the target character. Finally, the character script information, the picture information corresponding to the target character, and the source background description information are jointly used as conditional constraints (diffusion guiding information) and input into a diffusion model, and video generation is performed based on the encoding network, diffusion network, denoising network, and decoding network of the diffusion network, finally obtaining the character perspective video corresponding to the source video, and this character perspective video is a video with the perspective of the target character as the main perspective.

[0342] The training process of the diffusion model in the embodiments of the present disclosure is described in detail below.

[0343] Please refer to Fig.17 , in one embodiment, the diffusion model is trained through the following steps:

[0344] Step 1710: Obtain a sample video with a sample character, sample background description information of the sample video, and an expected character perspective video of the sample character;

[0345] Step 1720: Extract sample script information from the sample video;

[0346] Step 1730: Based on the sample script information, sample background description information, and a preset guiding language template information, use the character script conversion model to generate sample character script information from the perspective of the sample character;

[0347] Step 1740: Based on the sample character script information, sample picture information corresponding to the sample character, and sample background description information, generate sample guiding information;

[0348] Step 1750: Based on the sample guiding information, perform video generation through a target model to obtain a sample character perspective video of the sample character;

[0349] Step 1760: Based on the sample character perspective video and the expected character perspective video, train the target model to obtain a diffusion model.

[0350] The following provides a detailed description of steps 1710 - 1760.

[0351] In step 1710, the sample character refers to a certain character in the sample video. Among them, the target character can be a person, an animal, etc. in the source video.

[0352] The sample video refers to the video data corresponding to some narrative film and television works used to train the diffusion model.

[0353] The sample background description information is used to briefly introduce the background in which the whole story of the sample video takes place in written form.

[0354] The expected character perspective video refers to the video that is expected to be generated with the perspective of the sample character as the main perspective for the sample video.

[0355] In the specific implementation of this embodiment, the specific processes of the sample video with the sample character and the sample background description information of the sample video in step 1710 are similar to the above step 410. For the sake of brevity, it will not be elaborated here.

[0356] Furthermore, the expected character perspective video can be specially shot by the producer of the sample video, or can be pre-generated according to video editing software, etc., without limitation.

[0357] In step 1720, the sample script information refers to the script content on which the sample video depends during shooting.

[0358] In the specific implementation of this embodiment, the specific process of step 1720 is similar to the specific processes of the above steps 510 - 530. For the sake of brevity, it will not be elaborated here.

[0359] In step 1730, the sample character script information is used to indicate the script content that is supplemented and generated from the perspective of the sample character for the video content of the sample video.

[0360] Among them, the preset guiding language template information and the character script conversion model are basically the same as those in the above step 430.

[0361] In the specific implementation of this embodiment, the specific process of step 1730 is similar to the specific processes of the above steps 1110 - 1120 or steps 1210 - 1220. For the sake of brevity, it will not be elaborated here.

[0362] In step 1740, the sample guidance information is used as a text condition constraint to assist the target model in video denoising during video generation, so that the feature information of the sample character video generated by the target model can be as close as possible to the feature information of the expected character perspective video.

[0363] When this embodiment is specifically implemented, the specific process of step 1740 is similar to the specific processes of the above steps 1410 - 1440. To save space, it will not be elaborated here.

[0364] In step 1750, the target model is an untrained neural network model. This target model takes the background description of a film and television work, the text description of a character image, and the supplemented generated character script information, etc. as inputs, and outputs a video with the perspective of the sample character as the main perspective. Among them, the model structure of the target model is the same as that of the diffusion model, but there are certain differences in the model parameters between the target model and the diffusion model.

[0365] The sample character perspective video refers to a video with the perspective of the sample character as the main perspective. Among them, the plot described in the sample character perspective video is relatively similar to the plot described in the sample video. The difference is that the sample character perspective video is shown from the perspective of the sample character, while the sample video is shown from a third - party perspective.

[0366] To save space, the specific process of generating a sample character perspective video based on the sample guidance information through the target model in the embodiments of the present disclosure will be described in detail below. It will not be elaborated here.

[0367] In step 1760, with the goal of minimizing the difference between the sample character perspective video and the expected character perspective video, through iterative training, in each iterative training round, the model parameters of the target model are continuously adjusted, and the above steps 1710 - 1750 are continuously repeated until the difference between the sample character perspective video and the expected character perspective video reaches the maximum in a certain iterative training round. The model parameters of the target model in this iterative training round are determined as the final model parameters, and the target model with the final model parameters is determined as the trained diffusion model.

[0368] The advantage of this embodiment is that by comparing and evaluating the sample character perspective video generated by the target model under the conditional constraints of the sample character script information, sample picture information, and sample background description information with the expected character perspective video preset for the sample character of the sample video, it is possible to improve the learning and mining of the video content from the perspective of a specific character by the target model through iterative training, improve the accuracy of the diffusion model trained to generate the character perspective video from the perspective of a specific character, meet the video playback needs of different people, and thus improve the display diversity of video information in various film and television works and the information transmission efficiency of the character perspective video.

[0369] In the embodiment of the present disclosure, the sample video includes a plurality of sample video frames. Among them, the acquisition method of the plurality of sample video frames is basically the same as that in step 510 above. For the sake of brevity, it will not be elaborated here.

[0370] At the same time, the target model in the embodiment of the present disclosure includes a diffusion network, a denoising network, and a decoding network. Among them, the specific introduction of the diffusion network, denoising network, and decoding network in the target model is basically the same as the specific structure and specific introduction of the diffusion model above. For the sake of brevity, it will not be elaborated here.

[0371] Please refer to Fig.17 , in one embodiment, step 1750 specifically includes but is not limited to the following steps 1810-1840:

[0372] Step 1810, for each sample video frame, perform diffusion processing on the sample video frame encoding feature of the sample video frame based on the diffusion network to obtain a sample latent space feature vector;

[0373] Step 1820, based on the sample guidance information, perform denoising processing on the sample latent space feature vector through the denoising network to obtain the sample video frame denoising feature;

[0374] Step 1830, perform decoding processing on the sample video frame denoising feature based on the decoding network to obtain the predicted video frame of each sample video frame;

[0375] Step 1840, based on the predicted video frames of each of the plurality of sample video frames, obtain the sample character perspective video.

[0376] The following will describe steps 1810-1840 in detail.

[0377] In step 1810, the sample video frame encoding feature is used to indicate the feature representation of the sample video frame in the vector space (latent latent space).

[0378] The sample latent space feature vector is used to indicate the result of adding noise to the sample video frame encoding feature by the diffusion network at a fixed time step.

[0379] In the specific implementation of this embodiment, first, with authorization, a preset image encoder is called. Then, the sample video frame is input into the image encoder, and the sample video frame is encoded based on the image encoder to realize the mapping of the sample video frame from the pixel space to the vector space (latent space), thereby obtaining the encoded features of the sample video frame.

[0380] Further, according to the preset time steps, noise is successively added to the encoded features of the sample video frame through the diffusion network, so that the encoded features of the sample video frame gradually lose their original features. When the number of noise addition times reaches the preset time steps, the encoded features of the sample video frame will become a latent space representation without any features. At this time, the obtained latent space representation is determined as the sample latent space feature vector.

[0381] For example, the preset time steps are set to T, where T is a positive integer. The diffusion network is used to add noise to the encoded features of the sample video frame T times, and the result after T times of noise addition is used as the sample latent space feature vector.

[0382] In step 1820, the denoised features of the sample video frame are used to indicate the video frame features that meet the constraint requirements of the sample guidance information generated by denoising the sample latent space feature vector at a fixed time step through the denoising network.

[0383] In the specific implementation of this embodiment, according to the preset time steps, the sample latent space feature vector is successively denoised through the denoising network, and each time during denoising, the sample guidance information is used as a conditional constraint to make the denoising of the sample latent space feature vector meet the requirements of the sample guidance information. Further, during the T times of denoising, except that the first denoising operation uses the sample latent space feature vector as the denoising object, the subsequent T - 1 denoising operations use the denoising result of the previous denoising operation as the denoising object, and the denoising result of the last denoising operation is determined as the denoised features of the sample video frame.

[0384] In step 1830, the predicted video frame refers to a single video frame constructed by the decoding network based on the denoised features of the sample video frame from the perspective of the sample character as the main perspective.

[0385] In the specific implementation of this embodiment, for each sample video frame, the denoised features of the sample video frame are input into the decoding network, and through the decoding network, the denoised features of the sample video frame are mapped back from the latent vector space to the original pixel space to generate a noise-free predicted video frame that conforms to the sample guidance information.

[0386] In step 1840, according to the temporal order of multiple sample video frames, the predicted video frames are sequentially spliced in this order, so that multiple predicted video frames are spliced into a whole, thereby obtaining the sample character perspective video.

[0387] The advantage of this embodiment is that the sample video frames are mapped from the pixel space to the latent space, so that the sample video frames become the sample video frame encoding features that meet the input requirements of the target model, which can improve the usability of the sample video frames. Further, the diffusion network of the target model is used to add noise to the sample video frame encoding features multiple times, realizing the pure noise processing of the sample video frame encoding features, which can effectively strip the effective video frame features in the sample video frame encoding features, enabling the model to perform feature recovery according to the sample guidance information in the subsequent denoising process, and decoding the recovered sample video frame denoising features into predicted video frames, and generating the sample character perspective video according to the predicted video frames. This method is beneficial to improving the processing accuracy of the model for the task of generating the character perspective video.

[0388] In the embodiment of the present disclosure, the denoising network can be a U-shaped network structure (U-net network), and the denoising network includes an upsampling module and a downsampling module.

[0389] The upsampling module is used to perform downsampling processing on the sample latent space feature vector to obtain more low-dimensional features.

[0390] The downsampling module is used to perform upsampling processing on the output result of the upsampling module, and restore the output result to the sample video frame denoising features with the same feature dimension as the sample latent space feature vector.

[0391] Please refer to Fig.19 , in one embodiment, step 1820 specifically includes but is not limited to the following steps 1910-1930:

[0392] Step 1910: Perform splicing processing on the sample guidance information and the sample latent space feature vector to obtain the sample denoising splicing features;

[0393] Step 1920: Based on the downsampling module, perform downsampling processing on the sample denoising splicing features to obtain the video frame downsampling features;

[0394] Step 1930: Based on the upsampling module, perform upsampling processing on the video frame downsampling features to obtain the sample video frame denoising features.

[0395] The following will describe steps 1910-1930 in detail.

[0396] In step 1910, the sample denoising splicing features are used to represent the feature representation of fusing the sample guidance information into the sample latent space feature vector.

[0397] In the specific implementation of this embodiment, the sample guidance information and the sample latent space feature vector are concatenated to obtain a feature vector with a longer feature length, and the concatenated feature vector is determined as the sample denoising concatenated feature.

[0398] In step 1920, the video frame downsampling feature is used to indicate the video frame feature generated by denoising the sample latent space feature vector by means of downsampling.

[0399] In the specific implementation of this embodiment, in order to improve the learning effect of the model on the script information, background information, and character image information in the sample guidance information during the downsampling process, a self-attention mechanism can be introduced during the downsampling process. Among them, the self-attention calculation method in this downsampling process is similar to the self-attention calculation method in step 620 above. For the sake of brevity, it will not be elaborated here.

[0400] In another embodiment, step 1820 can also introduce a cross-attention mechanism during the downsampling process. Specifically, first, the sample latent space feature vector is linearly projected to obtain a query vector, and the sample guidance information is linearly projected to obtain a key vector and a value vector. Then, cross-attention calculation is performed based on the query vector, key vector, and value vector to obtain the video frame downsampling feature. This method enables the target model to guide the generation process of the sample character perspective video according to the sample guidance information.

[0401] In step 1930, the specific process of step 1920 is similar to the specific process of the above step 1910. The difference is that step 1910 is for the downsampling of the sample latent space feature vector; while step 1920 is for the upsampling of the sample downsampling result, and their specific processes are symmetric. In addition, the feature dimension of the sample video frame denoising feature obtained in step 1920 is the same as the feature dimension of the sample latent space feature vector input in step 1920.

[0402] Such as Fig. 20As shown, it is a specific schematic diagram of the denoising network performing denoising operations according to sample guidance information (when the predetermined time step is T = 1). Specifically, the denoising network includes a downsampling module and an upsampling module. Among them, the downsampling module includes an attention downsampling module 1 with an attention downsampling layer and a residual block structure, and an attention downsampling module 2 with an attention downsampling layer and a residual block structure. The upsampling module includes an attention upsampling module 1 with an attention upsampling layer and a residual block structure, and an attention upsampling module 2 with an attention upsampling layer and a residual block structure. First, the attention downsampling layer in the attention downsampling module 1 performs attention calculation on the sample latent space feature vector based on the sample guidance information to obtain a first attention calculation result, and the residual block structure in the attention downsampling module 1 performs residual processing on the first attention calculation result to obtain a first downsampling result. Further, the attention downsampling layer in the attention downsampling module 2 performs attention calculation on the first downsampling result based on the sample guidance information to obtain a second attention calculation result, and the residual block structure in the attention downsampling module 2 performs residual processing on the second attention calculation result to obtain the video frame downsampled feature. Then, the spatial transformation structure performs spatial transformation on the video frame downsampled feature to obtain the spatially transformed video frame downsampled feature. Further, the attention upsampling layer in the attention upsampling module 2 performs attention calculation on the spatially transformed video frame downsampled feature based on the sample guidance information to obtain a third attention calculation result, and the residual block structure in the attention downsampling module 2 performs residual processing on the third attention calculation result to obtain a first upsampling result. Finally, the attention upsampling layer in the attention upsampling module 1 performs attention calculation on the first upsampling result based on the sample guidance information to obtain a fourth attention calculation result, and the residual block structure in the attention upsampling module 2 performs residual processing on the fourth attention calculation result to obtain the sample denoising splicing feature.

[0403] It should be noted that the above denoising operation is illustrated with T = 1 as an example. When the predetermined time step is greater than 1, the first denoising operation is exactly the same as the above process. Starting from the second denoising operation, the denoising result of the previous denoising operation is used to replace the sample latent space feature vector in the first denoising operation, and the same process is executed. At this time, the sample denoising splicing feature refers to the final output of the T-th denoising operation (the last denoising operation).

[0404] The advantage of this embodiment is that the denoising process is divided into a downsampling stage and an upsampling stage, and the same sample guidance information is relied on in the upsampling stage and the downsampling stage, which can improve the consistency of feature fine-tuning in the downsampling stage and the upsampling stage. Specifically, in the downsampling stage, the downsampling module performs feature downsampling according to the sample guidance information, and in the upsampling stage, the upsampling module performs feature upsampling according to the sample guidance information. In this way, in both the downsampling stage and the upsampling stage, the three aspects of control information, namely the role script information, the role picture information, and the source background description information, are jointly used for guidance and constraint, enabling precise fine-tuning and flexible control of video frame denoising. This method enables the model to perform denoising under various conditional constraints, improves the denoising ability of the model, and is conducive to the model constructing specific content in the video according to the appearance of the role and the background of the picture, thereby improving the accuracy of the model in generating videos from the perspective of the role.

[0405] Since the denoising network is conditioned on the sample guidance information every time it performs denoising, the target model will rely to a large extent on the given conditional constraints throughout the denoising process. This method is not conducive to the autonomous learning of the target model and will result in low denoising ability of the target model under unconstrained conditions. Based on this, the embodiments of the present disclosure provide a denoising processing solution based on a conditional masking strategy, which can improve the denoising ability of the trained diffusion model under unconstrained conditions to a certain extent, and further improve the accuracy of the trained diffusion model in generating videos from the perspective of the role.

[0406] Please refer to Fig.21 , in another embodiment, step 1820 specifically includes but is not limited to the following steps 2110-2140:

[0407] Step 2110: Concatenate the sample guidance information and the sample latent space feature vector to obtain a sample denoising concatenated feature;

[0408] Step 2120: Based on the sample denoising concatenated feature, perform a first denoising process through the denoising network to obtain a first denoising feature;

[0409] Step 2130: Based on the sample latent space feature vector, perform a second denoising process through the denoising network to obtain a second denoising feature;

[0410] Step 2140: Perform feature weighting on the first denoising feature and the second denoising feature to obtain a sample video frame denoising feature.

[0411] The following will describe steps 1910-1930 in detail.

[0412] In steps 2110-2120, the first denoising feature is used to indicate the result of the denoising network performing a single denoising on the sample latent space feature under the constraint of the sample guidance information.

[0413] In the specific implementation of this embodiment, the specific processes of steps 2110 - 2120 are similar to those of 1910 - 1930. For the sake of brevity, they will not be elaborated here.

[0414] In step 2130, the second denoised feature is used to indicate the result of the denoising network performing a single denoising on the sample latent space feature without the constraint of sample guidance information.

[0415] In the specific implementation of this embodiment, since the denoising network performs denoising under no conditional constraints, that is, it gradually eliminates the noise added by the diffusion network to the sample latent space feature vector, the feature information in the sample latent space feature vector will not be changed due to the influence of any conditional constraints. This denoising operation only reduces the noise information carried by the sample latent space feature vector without changing the original feature information. Therefore, after the simple denoising process of the denoising network on the sample latent space feature vector, the sample latent space feature vector will be restored to the state before noise addition, thereby obtaining the second denoised feature. Among them, the second denoised feature is basically the same as the sample video frame encoding feature.

[0416] In step 2140, first, determine the first weight and the second weight, where the sum of the first weight and the second weight is 1. The first weight is used to indicate the importance of the conditional-constrained denoising process in model training, and the second weight is used to indicate the importance of the unconditional-constrained denoising process in model training. Then, multiply the first weight by the first denoised feature to obtain the first weighted denoised feature, and multiply the second weight by the second denoised feature to obtain the second weighted denoised feature. Finally, add the first weighted denoised feature and the second weighted denoised feature to obtain the sample video frame denoised feature.

[0417] It should be noted that this embodiment is illustrated with a predetermined number of time steps T = 1 (that is, the total number of denoising times is 1). When the total number of denoising times is greater than 1, each denoising process is executed according to steps 2110 - 2140, and the result of the denoising process is the result of feature weighting of the first denoised feature and the second denoised feature. The object of denoising for each denoising process is the result of the previous denoising process, and the sample video frame denoised feature is the weighted result of the first denoised feature and the second denoised feature generated by the last denoising.

[0418] Furthermore, the denoising process of the denoising network in the embodiments of the present disclosure can be expressed as the following formula:

[0419] ;

[0420] Among them, Describes the denoising process of the denoising network, which is equivalent to a conditional Gaussian distribution and is used to infer the denoising result of the previous time step t - 1 from the denoising result at the current time step t and the sample guidance information . This conditional Gaussian distribution approximates the posterior distribution represents Gaussian distribution sampling is a predefined constant represents the mean of the conditional Gaussian distribution represents the variance of the conditional Gaussian distribution is a preset identity matrix represents the denoising result of the previous time step t - 1 conforms to a Gaussian distribution with a mean of and a variance of

[0421] Among them, the mean of this conditional Gaussian distribution can be expressed as the following formula

[0422] ;

[0423] Among them is a preset constant, and the constant is used to control the proportion of the first denoising feature and the second denoising feature in feature weighting to balance the influence of conditional constraint information and unconditional constraint information on generating the denoising result of the previous time step t - 1 represents the denoising result of the previous time step calculated by the target model with model parameters θ given the sample guidance information c and time step t represents the denoising result of the previous time step calculated by the target model with model parameters θ by masking the given sample guidance information c and according to the given time step t is the attention weight at time step t refers to the cumulative average attention weight at time step t and are explained and described in detail below and will not be elaborated here

[0424] ​The advantage of this embodiment is that it takes into account the introduction of conditional masking with a certain probability in the denoising process of the sample latent space feature vector. Specifically, in each denoising operation, the sample latent space feature vector is denoised under the conditional constraint of the sample guidance information to obtain a denoising result (the first denoising feature) affected by the conditional constraint; at the same time, the sample latent space feature vector is denoised without conditional constraint to obtain a denoising result (the second denoising feature) affected by the unconditional constraint. Further, in each denoising operation, the first denoising feature and the second denoising feature are fused by means of feature weighting, so that after multiple denoising operations combining conditional constraint and unconditional constraint, the sample video frame denoising feature is generated. This method can, to a certain extent, improve the denoising ability of the trained diffusion model under unconditional constraint, and further improve the accuracy of the trained diffusion model in generating the character perspective video.

[0425] Please refer to Fig. 22 , in one embodiment, step 1760 specifically includes but is not limited to the following steps 2210-2240:

[0426] Step 2210, based on the sample character perspective video and the expected character perspective video, determine the first loss sub-function;

[0427] Step 2220, based on the sample video frame denoising features corresponding to each of the multiple sample video frames and a preset reference denoising feature, determine the second loss sub-function;

[0428] Step 2230, based on the first loss sub-function and the second loss sub-function, obtain the first loss function;

[0429] Step 2240, based on the first loss function, train the target model to obtain the diffusion model.

[0430] The following is a detailed description of steps 2210-2240.

[0431] In step 2210, the first loss sub-function is used to indicate the overall difference degree of the video information between the sample character perspective video and the expected character perspective video.

[0432] In the specific implementation of this embodiment, the sample character perspective video includes multiple character perspective video frames, and the expected character perspective video includes reference video frames corresponding one-to-one to the multiple character perspective video frames.

[0433] This step 2210 may include but is not limited to the following steps:

[0434] Determine the first pixel value of the pixel points of each character perspective video frame, and determine the second pixel value of the pixel points of each reference video frame;

[0435] For each video frame from a character's perspective, determine a pixel difference based on the first pixel value and the second pixel value of the reference video frame corresponding to the video frame from the character's perspective.

[0436] Based on the pixel differences of each of the multiple video frames from the character's perspective, determine a first loss sub-function.

[0437] Among them, the first pixel value is used to indicate the pixel value of a pixel point in the video frame from the character's perspective. The second pixel value is used to indicate the pixel value of a pixel point in the reference video frame. The pixel difference is used to indicate the magnitude of the difference in the pixel values of the pixel points at the same position in the video frame from the character's perspective and the reference video frame.

[0438] Specifically, for each video frame from the character's perspective, calculate the pixel difference for the pixel points at the same position in the video frame from the character's perspective and the corresponding reference video frame, subtract the second pixel value from the first pixel value of the pixel point to obtain the pixel difference. Then, calculate the mean squared error of all the pixel differences to obtain the mean squared error of the pixels of a single video frame from the character's perspective. Finally, average the mean squared errors of all the video frames from the character's perspective to obtain the first loss sub-function.

[0439] In step 2220, the reference denoising feature is used to indicate the degree of noise added to the reference video frame corresponding to each sample video frame. Among them, in the embodiments of the present disclosure, the reference denoising feature may be a feature that conforms to the standard normal distribution.

[0440] The second loss sub-function is used to indicate the overall difference degree between the predicted noise and the reference noise of multiple sample video frames.

[0441] When specifically implementing this embodiment, this step 2220 may include but is not limited to the following steps:

[0442] For each sample video frame, perform a regularization term calculation based on the sample video frame denoising feature and the reference denoising feature to obtain a regularization term calculation result;

[0443] Based on the regularization term calculation results of each of the multiple sample video frames, determine the second loss sub-function.

[0444] Among them, the regularization term calculation result is used to indicate the difference degree between the predicted noise and the reference noise of a single sample video frame.

[0445] Specifically, first, for each sample video frame, calculate the noise difference between the sample video frame denoising feature and the reference denoising feature to obtain the noise difference; then, perform a regularization term calculation on the noise difference to obtain the regularization term calculation result. Finally, average the regularization term calculation results of all the sample video frames to obtain the second loss sub-function.

[0446] Among them, the calculation results of the regularization terms for each sample video frame can be expressed as shown in the following formula:

[0447] ;

[0448] Among them, loss1 refers to the calculation result of the regularization term corresponding to a single sample video frame (character perspective video frame). refers to the reference denoising feature at time step t, refers to the sample video frame denoising feature at time step t, , (.) is the target model with model parameters θ, refers to the result obtained by the target model with model parameters θ performing t denoising operations on the sample latent space feature vector according to the sample guidance information c, is the sample latent space feature vector at time step t, is the sample guidance information, represents the time step, used to indicate the currently accumulated number of noise times.

[0449] Furthermore, the sample latent space feature vector can be expressed as shown in the following formula:

[0450] ;

[0451] Among them, ; ;

[0452] Among them, is a real number close to 0 set in advance, will gradually decrease as the time step t increases. The reference denoising feature is a randomly generated sampling value of the standard normal distribution. The tensor shape of the reference denoising feature is the same as the tensor shape of the sample latent space feature vector . T is the total number of time steps, and t is any integer between 1 and T. is the attention weight at time step t. refers to the sample video frame encoding feature. refers to the cumulative average attention weight at time step t.

[0453] Based on this, the sample latent space feature vector at time step t can be understood as multiplying the square root of the sample video frame encoding feature by the cumulative average attention weight , then adding the reference denoising feature multiplied by 1 minus the cumulative average attention weight obtained by multiplying the square roots of

[0454] In step 2230, the first loss function is used to indicate the overall difference degree between the sample character perspective video generated by the target model and the expected character perspective video in terms of pixels and noise.

[0455] In the specific implementation of this embodiment, the first loss sub-function and the second loss sub-function are weighted and calculated to obtain the first loss function. Among them, the specific process of weighting and calculating the first loss sub-function and the second loss sub-function is similar to the specific process of feature weighting of the first denoising feature and the second denoising feature in step 2140 above. The difference is that the weights of the first loss sub-function and the second loss sub-function in step 2230 can be freely set, and the sum of the weights of the two can be not equal to 1. For the sake of brevity, it will not be elaborated here.

[0456] In step 2240, taking the minimization of the first loss function as the training objective, the model parameters of the target model are adjusted, and the above steps 1710-1760 are repeated to implement iterative training of the target model. The model parameters that minimize the first loss function are used as the final model parameters, and the target model with the final model parameters is used as the trained diffusion model.

[0457] The advantage of this embodiment is that it takes into account the pixel value difference and noise difference between the sample character perspective video generated by the target model in each iterative training round and the expected character perspective video; based on the supervised learning method, according to the pixel difference between the first pixel value of each pixel point in each character perspective video frame and the second pixel value of this pixel point in the reference video frame, as well as the noise feature difference between the sample video frame denoising feature and the reference denoising feature of each character video frame, the first loss function is jointly determined. This method can train the target model by minimizing the pixel value difference and noise difference, which is beneficial to improving the model training effect and further improving the accuracy of the model in generating character perspective videos.

[0458] Please refer to Fig.23 , in one embodiment, the target model is obtained by adjusting the first model through the following steps:

[0459] Step 2310, obtain multiple reference video frames of the reference video;

[0460] Step 2320, perform three-dimensional convolution processing on the multiple reference video frames to obtain the predicted mean and predicted standard deviation of the multiple reference video frames;

[0461] Step 2330, based on the predicted mean and predicted standard deviation, reparameterize each of the multiple reference video frames to obtain the video frame hidden state features of each of the multiple reference video frames;

[0462] Step 2340: Decode the video frame hidden state features of each of the multiple reference video frames to obtain the predicted video frames corresponding to each of the multiple reference video frames;

[0463] Step 2350: Based on the predicted video frames, reference video frames, predicted mean, and predicted standard deviation, adjust the parameters of the first model to obtain the target model.

[0464] The following provides a detailed description of steps 2310 - 2350.

[0465] In step 2310, the reference video refers to the video data corresponding to some narrative film and television works used to adjust the first model. The multiple reference video frames are the video frames obtained by segmenting the reference video.

[0466] For example, for a reference video corresponding to a certain reference film and television work, it is segmented at fixed time steps (e.g., 30 seconds) to obtain the set of reference video frames corresponding to the reference video. This set of reference video frames is represented as , , where represents the reference video frame at time window t, n is the total number of reference video frames, n is an integer, and t is any integer from 1 to n. refers to the model parameters of the first model. indicates that the set of reference video frames is the set of video frames generated under the influence of the model parameters of the first model.

[0467] In the specific implementation of this embodiment, step 2310 is similar to the above-mentioned step 1710. For the sake of brevity, it will not be elaborated here.

[0468] In step 2320, the predicted mean refers to the expected value of the Gaussian distribution corresponding to each of the multiple reference video frames in the latent vector space. The predicted standard deviation refers to the standard deviation of the Gaussian distribution corresponding to each of the multiple reference video frames in the latent vector space.

[0469] It should be noted that each reference video frame has a corresponding predicted mean and predicted standard deviation. There will be certain differences in the predicted means and predicted standards of different reference video frames.

[0470] In the specific implementation of this embodiment, first, the encoding network of the first model performs three-dimensional convolution processing on each reference video frame, causing the three-dimensional convolution kernel of the encoding network to slide in the temporal dimension, height, and width of the reference video frame to capture the spatio-temporal features of the reference video frame, obtaining the video frame convolution features of each reference video frame, where there are multiple video frame convolution features for each reference video frame feature. Then, for each reference video frame, pixel averaging is performed on the multiple video frame convolution features of this reference video frame to obtain the predicted mean corresponding to this reference video frame. Further, the square of the difference between each video frame convolution feature of this reference video frame and the predicted mean is calculated to obtain a squared value, and the square root is taken after averaging multiple squared values to obtain the predicted standard deviation of this reference video frame.

[0471] Among them, the predicted mean can be expressed as , and the predicted standard deviation can be expressed as . The tensor shape of the predicted mean and the predicted standard deviation can be expressed as 64×64×n, where n refers to the total number of reference video frames.

[0472] In step 2330, the video frame hidden state feature is used to indicate the Gaussian distribution corresponding to the reference video frame in the latent vector space.

[0473] In the specific implementation of this embodiment, for each reference video frame feature, a random variable is sampled from a standard normal distribution with a mean of 0 and a standard deviation of 1; then, the random variable is multiplied by the predicted standard deviation to obtain a product result, and the product result is added to the predicted mean to obtain a Gaussian distribution that meets the requirements, and this generated Gaussian distribution is determined as the video frame hidden state feature corresponding to the reference video frame.

[0474] Among them, the reference video frame The corresponding video frame hidden state feature can be expressed as the following formula:

[0475] ;

[0476] Among them, refers to the video frame hidden state feature, refers to the predicted mean, refers to the predicted standard deviation. refers to the standard normal distribution with a mean of 0 and a standard deviation of 1.

[0477] In step 2340, the predicted video frame refers to the prediction result of the reference video frame generated after the decoding network of the first model decodes the video frame hidden state feature.

[0478] In the specific implementation of this embodiment, for each reference video frame, the video frame hidden state feature of the reference video frame is input into the decoding network of the first model, and the decoding network performs inverse three-dimensional convolution processing on the video frame hidden state feature to obtain the predicted video frame corresponding to the reference video frame. Among them, the reference video frame The corresponding predicted video frame can be expressed as .

[0479] In step 2350, the specific implementation of this embodiment may include but is not limited to the following steps:

[0480] Based on multiple reference video frames and the predicted video frames corresponding to each of the multiple reference video frames, mean square error calculation is performed to obtain a video frame error function;

[0481] Based on the predicted mean and predicted standard deviation, logarithmic calculation is performed to obtain a log-likelihood function;

[0482] Based on the video frame error function and the log-likelihood function, the parameters of the first model are adjusted to obtain a target model.

[0483] Among them, the video frame error function is used to indicate the overall difference degree between the predicted video frame and the reference video frame, and the log-likelihood function is used to indicate the overall difference degree between the video frame hidden state feature of the reference video frame and the standard normal distribution.

[0484] Specifically, when determining the video frame error function, for each reference video frame, for each pixel point of the reference video frame, the square of the difference between the pixel value of the pixel point in the reference video frame and the pixel value of the pixel point in the predicted video frame is calculated to obtain a pixel square value, and the average of the pixel square values of all pixel points of the reference video frame is obtained to obtain a sub-error function, and the average of the sub-error functions of all reference video frames is obtained to obtain the video frame error function. Among them, the reference video frame The sub-error function of can be expressed as . MSE(.) is the mean squared error function.

[0485] When determining the log-likelihood function, for each reference video frame, logarithmic calculation is performed on the predicted mean and predicted standard deviation according to the first formula to obtain the logarithmic result of the reference video frame. Then, the logarithmic results of multiple reference video frames are averaged to obtain the log-likelihood function. Among them, the first formula can be expressed as the following formula:

[0486] ;

[0487] Among them, represents the logarithmic result of the reference video frame , represents the predicted mean of the reference video frame , Indicates the reference video frame of the predicted standard deviation.

[0488] Further, subtract the log-likelihood function from the video frame error function to obtain the total loss function, and adjust the parameters of the first model according to the total loss function to obtain the target model. Among them, the specific method of adjusting the parameters of the first model based on the video frame error function and the log-likelihood function is similar to the above steps 2230-2240. For the sake of brevity, it will not be elaborated here.

[0489] The advantage of this embodiment is that it allows the first model to learn the predicted mean and the predicted standard deviation values during the training process through the backpropagation algorithm, and maintain a certain degree of randomness. This randomness is due to regarding the randomly sampled part (i.e., normal(0,1)) as a constant during the backpropagation process, and taking the predicted mean and the predicted standard deviation as model parameters to be continuously updated, so that the first model can optimize the predicted mean and the predicted standard deviation while maintaining randomness, thereby learning the latent distribution of the video data of the reference video frame, and improving the learning ability of the first model. Further, when continuously adjusting the model parameters of the first model, the mean squared error and the log-likelihood result of the hidden state features and the standard normal distribution are combined, which can improve the adjustment accuracy of the first model, and thus improve the performance of the basic model of the target model used to train the diffusion model.

[0490] Next, in conjunction with Figure 24A-Figure 24C an example description will be given of the specific process of training the diffusion model in the embodiments of the present disclosure.

[0491] As Fig.24A shown, it is a specific illustration of using the target model to generate the role perspective video frame corresponding to the sample video frame in one iterative training round. Specifically, first, input the sample video frame into the encoding network, and the encoding network generates the sample video frame encoding feature , and this sample video frame encoding feature is a hidden state feature. Then, the diffusion network performs forward diffusion on the sample video frame encoding feature , adds noise to the sample video frame encoding feature t times to obtain the sample latent space feature vector . Further, input the sample role script information, sample picture information, and sample background description information as conditional constraints into the conditional fusion network , and by the conditional fusion network Output sample guidance information. Then, the denoising network performs reverse diffusion on the sample latent space feature vector. At this time, the denoising network performs t denoising processes on the sample latent space feature vector according to the sample guidance information to obtain the denoised features of the sample video frames. Finally, the decoding network decodes the character perspective video frames based on the denoised features of the sample video frames.

[0492] As Fig. 24B shown, it is a specific illustration of using the first model to generate predicted video frames corresponding to the reference video frames in an iteration when adjusting the first model. Specifically, first, the sample video frames are input into the encoding network, and the encoding network generates video frame hidden state features. Then, the decoding network decodes the predicted video frames based on the video frame hidden state features.

[0493] As Fig.24C shown, it is a specific illustration of the diffusion process of the diffusion network and the denoising process of the denoising network in the target model in Fig.24A . First, the diffusion network performs forward diffusion on the encoded features of the sample video frames , adds noise to the encoded features of the sample video frames t times to obtain the sample latent space feature vector . Further, the sample character script information, sample picture information, and sample background description information are input into the conditional fusion network as conditional constraints, and the conditional fusion network outputs sample guidance information. Then, the denoising network performs reverse diffusion on the sample latent space feature vector. At this time, the denoising network performs t denoising processes on the sample latent space feature vector according to the sample guidance information to obtain the denoised features of the sample video frames . Among them, when the denoising network performs denoising processing, according to the conditional masking strategy, it will set and switch the switch. When the switch is set on the left, it performs unconditional constraint processing on the sample latent space feature vector, that is, it directly performs denoising operations on the sample latent space feature vector without being affected by the sample guidance information. At the same time, when the switch is set on the right, it performs conditional constraint processing on the sample latent space feature vector, that is, it performs denoising operations on the sample latent space feature vector according to the sample guidance information.

[0494] Next, the training process of the video script description generation model of the present disclosure embodiment will be described in detail.

[0495] Please refer to Fig.25 , in an embodiment, the video script description generation model is trained through the following steps:

[0496] Step 2510: Obtain a plurality of reference video frames arranged in time sequence of the reference video, and reference script description information corresponding to each of the plurality of reference video frames;

[0497] Step 2520: Extract the reference video frame features of each of the multiple reference video frames;

[0498] Step 2530: Based on the reference video frame features of each of the multiple reference video frames and the reference script description information, perform script generation through a preset second model to obtain the predicted script information of each of the multiple reference video frames;

[0499] Step 2540: Based on the reference script description information and the predicted script information of each of the multiple reference video frames, train the second model to obtain a video script description generation model.

[0500] The following provides a detailed description of steps 2510 - 2540.

[0501] In step 2510, the reference script description information is used to indicate the script content corresponding to each reference video frame obtained through manual annotation.

[0502] In the specific implementation of this embodiment, the obtaining of multiple reference video frames arranged in time sequence in step 2510 is similar to the above step 2310. For the sake of brevity, it will not be elaborated here.

[0503] Furthermore, based on the manual annotation method, the plot corresponding to each reference video frame in the reference video frame set is described and generated to obtain the reference script description information corresponding to each reference video frame. Among them, the reference video frame set The corresponding reference script description set can be expressed as , where is the reference script description information corresponding to the reference video frame .

[0504] In step 2520, the reference video frame feature is used to indicate the video feature information contained in a single reference video frame.

[0505] In the specific implementation of this embodiment, the reference video frame feature is extracted through the feature extraction network of the second model, and this feature extraction process is similar to the above steps 710 - 740. For the sake of brevity, it will not be elaborated here.

[0506] In step 2530, the second model refers to a neural network model that can generate corresponding script description information for the input video frame. Among them, the second model mainly includes a feature extraction network based on a convolutional neural network structure and a script description generation network based on a transformer decoder.

[0507] The predicted script information is used to indicate the script content generated by the second model for a single reference video frame.

[0508] For the sake of brevity, the specific process of generating a script through a preset second model based on the reference video frame features of each of multiple reference video frames and the reference script description information will be described in detail below. It will not be elaborated here.

[0509] In step 2540, the specific process of training the second model based on the reference script description information and the predicted script information of each of multiple reference video frames may include but is not limited to the following steps:

[0510] For each reference video frame, determine a third loss sub-function based on the reference video frame script description information and the predicted script description information;

[0511] Based on multiple third loss sub-functions, determine a second loss function;

[0512] Based on the second loss function, train the second model to obtain a video script description generation model.

[0513] Among them, the third loss sub-function is used to indicate the degree of difference between the predicted script description information of a single reference video frame and the reference video frame script description information. The second loss function is used to indicate the overall degree of difference between the predicted script description information predicted by the second model for multiple reference video frames and the respective reference video frame script description information.

[0514] Specifically, the predicted script description information of each reference video frame includes the first probability and the first serial number of each of multiple predicted words, and the reference video frame script description information includes the second serial number of each of multiple reference words.

[0515] A reference word refers to each word in the reference video frame script description information.

[0516] A predicted word is the prediction result of the second model for each element of the reference video frame feature. Among them, the predicted word is often a word taken from the preset word list of the second model.

[0517] The first probability is used to indicate the probability of predicting the element of the reference video frame feature as this predicted word.

[0518] The first serial number is used to indicate the position order of the predicted word in the predicted script description information.

[0519] The second serial number is used to indicate the position order of the reference word in the reference video frame script description information.

[0520] In this embodiment, the specific process of determining the third loss sub-function for each reference video frame based on the reference video frame script description information and the predicted script description information may include but is not limited to the following steps:

[0521] For a single predicted word, if the predicted word is identical to the reference word with the same first serial number as the second serial number of the single predicted word, determine the predicted word as the target word;

[0522] Determine the third loss sub-function based on the negative logarithm of the first probability of the target word.

[0523] Specifically, for each reference video frame, according to the position order of each predicted word in the predicted script description information, start the inspection from the predicted word with the first serial number of 1, and determine whether the predicted word with the first serial number of 1 is the same as the reference word with the second serial number of 1. Only when the predicted word with the first serial number of 1 is the same as the reference word with the second serial number of 1, use the predicted word with the first serial number of 1 as the target word, and so on, inspect the predicted word with the first serial number of 2,..., until all predicted words with the first serial number equal to the total number of predicted words (the total number of reference words) are inspected to obtain multiple target words. Further, for each target word, take the negative logarithm of the first probability of the target word to obtain the negative logarithm of the first probability, and add up the negative logarithms of the first probabilities of all target words to obtain the third loss sub-function of this reference video frame.

[0524] Among them, the third loss sub-function can be expressed as the following formula:

[0525] ;

[0526] Among them, refers to the reference video frame corresponding third loss sub-function, refers to the one-hot vector at the i-th position of the reference script description information. When the predicted word at the i-th position is the same as the reference word, the value of is 1; when the predicted word at the i-th position is different from the reference word, the value of is 0. refers to the probability value corresponding to each predicted word. Only when the predicted word at the i-th position is the same as the reference word, the probability value of the predicted word at the i-th position is the first probability. represents the total number of reference words in the reference script description information. (.) represents the cross-entropy loss function.

[0527] Further, average the multiple third loss sub-functions to obtain the second loss function. Finally, based on the second loss function, train the second model to obtain the video script description generation model, where the specific method of training the second model based on the second loss function is similar to the above step 2240. For the sake of brevity, it will not be elaborated here.

[0528] The advantage of this embodiment is that in the way of supervised training, the second model generates a corresponding predicted script information for each reference video frame, and takes minimizing the difference between the predicted script information of each reference video frame and the reference script description information as the training objective, and iteratively adjusts the model parameters of the second model. This way can improve the accuracy of the trained video script description generation model to generate script information according to video frame features.

[0529] Please refer to Fig.26 , in one embodiment, step 2530 specifically includes but is not limited to the following steps 2610-2640:

[0530] Step 2610: For each reference video frame, perform embedding processing on the reference script description information to obtain reference script description embedding features;

[0531] Step 2620: Based on the reference script description embedding features, perform self-attention calculation through the second model to obtain a first attention calculation result;

[0532] Step 2630: Based on the first attention calculation result and the reference video frame features, perform cross-attention calculation through the second model to obtain a second attention calculation result;

[0533] Step 2640: Based on the second attention calculation result and the preset vocabulary in the second model, generate a script description to obtain the predicted script information of the reference video frame.

[0534] The following will describe steps 2610-2640 in detail.

[0535] In step 2610, the reference script description embedding features are used to indicate the latent representation of the reference script description information in the latent vector space.

[0536] When this embodiment is specifically implemented, the specific process of step 2610 is similar to the above step 1410. For the sake of brevity, it will not be elaborated here.

[0537] In step 2620, the specific process of performing self-attention calculation on the reference script description embedding features through the second model is similar to the specific process of performing self-attention calculation on the target video frame features by using the attention layer in the above step 620. For the sake of brevity, it will not be elaborated here.

[0538] In step 2630, first, linearly project the first attention calculation result. Use the query channel of the attention layer to project the first attention calculation result into query features, use the key channel of the attention layer to project the reference video frame features into key features, and use the value channel of the attention layer to project the reference video frame features into value features. Then, perform an inner product operation on the query features and the transposed result of the key features to obtain an inner product result, and divide the inner product result by the square root of the feature dimension of the key features to obtain attention scores. Further, use the softmax function to convert the attention scores into attention weights, and perform a multiplication operation on the attention weights and the value features to obtain the second attention calculation result.

[0539] In step 2640, the preset vocabulary is a list composed of multiple reference words preset according to business requirements.

[0540] In the specific implementation of this embodiment, the specific process of generating a script description based on the second attention calculation result and the preset vocabulary in the second model is basically the same as the process of the feed-forward processing in step 620 above and calculating the probability magnitudes of the reference words of each element in the second attention calculation result. To save space, it will not be elaborated here.

[0541] The advantage of this embodiment is that the self-attention mechanism and the cross-attention mechanism are introduced, enabling the second model to identify and emphasize the key features of the reference script description information of the reference video frames. At the same time, using the reference video frame features corresponding to the reference video frames as key features and value features can enable the second model to pay more attention to the video content related to the current plot when generating a script description, improving the relevance and accuracy of the generated predicted script information. In addition, the second model relies on the transform Decoder structure when generating a script description, which makes the process of generating a script description have good flexibility and can adapt to different video contents and plot description requirements. Through iterative training, the second model can fully learn how to generate script description information that meets the requirements based on video frame features.

[0542] The devices and equipment of the embodiments of the present disclosure will be described below.

[0543] It can be understood that although the steps in the above-mentioned various flowcharts are sequentially shown according to the representation of the arrows, these steps are not necessarily executed sequentially in the order represented by the arrows. Unless there is a clear description in this embodiment, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the above-mentioned flowchart may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or alternately with at least a part of the steps or stages in other steps or other steps.

[0544] It should be noted that in each specific embodiment of the present application, when it comes to performing relevant processing based on data related to the characteristics of the target object, such as target object attribute information or a set of attribute information, the permission or consent of the target object will be obtained first. Moreover, the collection, use, and processing of these data will comply with relevant laws, regulations, and standards. In addition, when the embodiment of the present application needs to obtain the target object attribute information, it will obtain the separate permission or separate consent of the target object through methods such as pop-up windows or jumping to a confirmation page. After clearly obtaining the separate permission or separate consent of the target object, the necessary target object-related data for the normal operation of the embodiment of the present application will be obtained.

[0545] Fig. 27 FIG. 2700 is a schematic structural diagram of a role perspective video generation device provided by an embodiment of the present disclosure. The role perspective video generation device 2700 includes:

[0546] A first acquisition unit 2710, configured to acquire a source video with a target role and source background description information of the source video;

[0547] An extraction unit 2720, configured to extract source script information from the source video;

[0548] A second acquisition unit 2730, configured to acquire guiding language template information corresponding to a role script conversion model;

[0549] A first generation unit 2740, configured to generate role script information from the perspective of the target role by using the role script conversion model based on the source script information, the source background description information, and the guiding language template information;

[0550] A second generation unit 2750, configured to generate diffusion guiding information based on the role script information, the picture information corresponding to the target role, and the source background description information, and generate a role perspective video from the perspective of the target role based on the diffusion guiding information.

[0551] Optionally, the first generation unit 2740 includes:

[0552] An acquisition module (not shown) for acquiring picture information corresponding to a target character;

[0553] A first generation module (not shown) for generating a guiding statement based on source script information, picture information, and source background description information by using guiding statement template information, and inputting the guiding statement into a character script conversion model to obtain character script information.

[0554] Optionally, the source script information includes source line information and source scene information, the character script information includes character script line information and character script scene information, the guiding statement template information includes first guiding statement template information and second guiding statement template information, and the guiding statement includes a first guiding statement and a second guiding statement;

[0555] The first generation module (not shown) is used for:

[0556] Generating a first guiding statement based on the source line information, source scene information, picture information, and source background description information by using the first guiding statement template information, and inputting the first guiding statement into the character script conversion model to obtain character script line information;

[0557] Generating a second guiding statement based on the source scene information and source background description information by using the second guiding statement template information, and inputting the second guiding statement into the character script conversion model to obtain character script scene information.

[0558] Optionally, the character script conversion model includes a first sub-model and a second sub-model;

[0559] The first generation unit 2740 is used for:

[0560] Generating a character script prompt instance through the first sub-model based on the guiding statement template information, source script information, and source background description information;

[0561] Generating character script information through the second sub-model based on the character script prompt instance.

[0562] Optionally, the second acquisition unit 2730 includes:

[0563] A segmentation module (not shown) for performing segmentation processing on the source video to obtain a plurality of source video frames arranged in time sequence;

[0564] A second generation module (not shown) for inputting the plurality of source video frames into a pre-trained video script description generation model for description generation to obtain video frame script description information for each of the plurality of source video frames;

[0565] An integration module (not shown) for integrating the video frame script description information of each of multiple source video frames to obtain source script information.

[0566] Optionally, the video script description generation model includes a feature extraction network and a script description generation network;

[0567] The second generation module (not shown) includes:

[0568] An extraction sub-module (not shown) for extracting features of each source video frame based on the feature extraction network to obtain target video frame features of the source video frame;

[0569] A generation sub-module (not shown) for generating a script based on the target video frame features using the script description generation network to obtain video frame script description information of the source video frame.

[0570] Optionally, the extraction sub-module (not shown) is used for:

[0571] For each source video frame, perform multi-scale downsampling on the source video frame based on the feature extraction network to obtain a first downsampled feature, a second downsampled feature, and a third downsampled feature, where the feature dimensions of the first downsampled feature, the second downsampled feature, and the third downsampled feature increase in sequence;

[0572] Perform dimensionality reduction processing on the first downsampled feature, the second downsampled feature, and the third downsampled feature respectively to obtain a first dimensionality reduction feature corresponding to the first downsampled feature, a second dimensionality reduction feature corresponding to the second downsampled feature, and a third dimensionality reduction feature corresponding to the third downsampled feature;

[0573] Perform splicing processing on the first dimensionality reduction feature, the second dimensionality reduction feature, and the third dimensionality reduction feature to obtain initial video frame features of the source video frame;

[0574] Perform splicing processing on the initial video frame features of the source video frame and the initial video frame features of the previous source video frame of the source video frame to obtain target video frame features of the source video frame.

[0575] Optionally, the integration module (not shown) is used for:

[0576] Splice the video frame script description information of each of multiple source video frames to obtain preliminary script information;

[0577] Polish the preliminary script information to obtain source script information.

[0578] Optionally, the second generation unit 2750 is used for:

[0579] Encode the source background description information to obtain background description encoded features;

[0580] Encode the character script information to obtain the character script encoding features;

[0581] Extract features from the picture information to obtain the target character image features;

[0582] Based on the background description encoding features, the character script encoding features, and the target character image features, generate conditional information through a preset conditional fusion network to obtain diffusion guidance information.

[0583] Optionally, the second generation unit 2750 is used to: guide the diffusion model based on the diffusion guidance information to generate a character perspective video from the target character's perspective;

[0584] The diffusion model is trained as follows:

[0585] A sample acquisition unit (not shown) is used to acquire a sample video with a sample character, sample background description information of the sample video, and an expected character perspective video of the sample character;

[0586] A sample extraction unit (not shown) is used to extract sample script information from the sample video;

[0587] A sample script generation unit (not shown) is used to generate sample character script information from the sample character's perspective based on the sample script information, the sample background description information, and preset guiding language template information, using a character script conversion model;

[0588] A sample guidance information generation unit (not shown) is used to generate sample guidance information based on the sample character script information, the sample picture information corresponding to the sample character, and the sample background description information;

[0589] A sample video generation unit (not shown) is used to generate a sample character perspective video from the sample character's perspective through a target model based on the sample guidance information;

[0590] A first training unit (not shown) is used to train the target model based on the sample character perspective video and the expected character perspective video to obtain the diffusion model.

[0591] Optionally, the sample video includes multiple sample video frames; the target model includes a diffusion network, a denoising network, and a decoding network;

[0592] The sample video generation unit (not shown) is used to:

[0593] A diffusion module (not shown) is used to perform diffusion processing on the sample video frame encoding features of each sample video frame based on the diffusion network to obtain sample latent space feature vectors;

[0594] A denoising module (not shown) is configured to denoise a sample latent space feature vector through a denoising network based on sample guidance information to obtain a denoised feature of a sample video frame;

[0595] A decoding module (not shown) is configured to decode the denoised feature of the sample video frame through a decoding network to obtain a predicted video frame for each sample video frame;

[0596] A video generation module (not shown) is configured to obtain a sample character perspective video based on the predicted video frames of the respective sample video frames.

[0597] Optionally, the denoising network includes an upsampling module and a downsampling module;

[0598] The denoising module (not shown) is configured to:

[0599] Concatenate the sample guidance information and the sample latent space feature vector to obtain a denoised concatenated feature of the sample;

[0600] Perform downsampling on the denoised concatenated feature of the sample through the downsampling module to obtain a downsampled feature of the video frame;

[0601] Perform upsampling on the downsampled feature of the video frame through the upsampling module to obtain a denoised feature of the sample video frame.

[0602] Optionally, the denoising module (not shown) is configured to:

[0603] Concatenate the sample guidance information and the sample latent space feature vector to obtain a denoised concatenated feature of the sample;

[0604] Perform a first denoising process on the denoised concatenated feature of the sample through the denoising network to obtain a first denoised feature;

[0605] Perform a second denoising process on the sample latent space feature vector through the denoising network to obtain a second denoised feature;

[0606] Perform feature weighting on the first denoised feature and the second denoised feature to obtain a denoised feature of the sample video frame.

[0607] Optionally, the sample video includes a plurality of sample video frames;

[0608] A first training unit (not shown) is configured to:

[0609] Determine a first loss sub-function based on the sample character perspective video and the expected character perspective video;

[0610] Determine a second loss sub-function based on the denoised features of the respective sample video frames corresponding to the plurality of sample video frames and a preset reference denoised feature;

[0611] Based on the first loss sub-function and the second loss sub-function, obtain the first loss function;

[0612] Based on the first loss function, train the target model to obtain the diffusion model.

[0613] Optionally, based on the sample video frame denoising features corresponding to each of the multiple sample video frames and the preset reference denoising features, determine the second loss sub-function, including:

[0614] For each sample video frame, perform a regularization term calculation based on the sample video frame denoising feature and the reference denoising feature to obtain a regularization term calculation result;

[0615] Based on the regularization term calculation results of each of the multiple sample video frames, determine the second loss sub-function.

[0616] Optionally, the sample role perspective video includes multiple role perspective video frames, and the expected role perspective video includes reference video frames corresponding one-to-one to the multiple role perspective video frames;

[0617] Based on the sample role perspective video and the expected role perspective video, determine the first loss sub-function, including:

[0618] Determine the first pixel value of the pixel points of each role perspective video frame, and determine the second pixel value of the pixel points of each reference video frame;

[0619] For each role perspective video frame, determine the pixel difference based on the first pixel value and the second pixel value of the reference video frame corresponding to the role perspective video frame;

[0620] Based on the pixel differences of each of the multiple role perspective video frames, determine the first loss sub-function.

[0621] Optionally, the target model is obtained by adjusting the first model through the following steps:

[0622] A reference video frame acquisition unit (not shown), configured to acquire multiple reference video frames of the reference video;

[0623] A convolution unit (not shown), configured to perform three-dimensional convolution processing on the multiple reference video frames to obtain the predicted mean and predicted standard deviation of the multiple reference video frames;

[0624] A reparameterization unit (not shown), configured to perform reparameterization on each of the multiple reference video frames based on the predicted mean and predicted standard deviation to obtain the video frame hidden state features of each of the multiple reference video frames;

[0625] A feature decoding unit (not shown), configured to perform decoding processing on the video frame hidden state features of each of the multiple reference video frames to obtain the predicted video frames corresponding to each of the multiple reference video frames;

[0626] An adjustment unit (not shown) for adjusting the parameters of the first model based on a predicted video frame, a reference video frame, a predicted mean, and a predicted standard deviation to obtain a target model.

[0627] Optionally, the adjustment unit (not shown) is configured to:

[0628] Calculate the mean square error based on multiple reference video frames and the corresponding predicted video frames for each of the multiple reference video frames to obtain a video frame error function;

[0629] Perform a logarithmic calculation based on the predicted mean and the predicted standard deviation to obtain a log-likelihood function;

[0630] Adjust the parameters of the first model based on the video frame error function and the log-likelihood function to obtain a target model.

[0631] Optionally, the video script description generation model is trained as follows:

[0632] A reference video acquisition unit (not shown) for acquiring multiple reference video frames arranged in time sequence of a reference video and reference script description information corresponding to each of the multiple reference video frames;

[0633] A feature extraction unit (not shown) for extracting reference video frame features for each of the multiple reference video frames;

[0634] A script generation unit (not shown) for generating a predicted script information for each of the multiple reference video frames by performing script generation through a preset second model based on the reference video frame features for each of the multiple reference video frames and the reference script description information;

[0635] A second training unit (not shown) for training the second model based on the reference script description information and the predicted script information for each of the multiple reference video frames to obtain a video script description generation model.

[0636] Optionally, the script generation unit (not shown) is configured to:

[0637] Perform an embedding process on the reference script description information for each reference video frame to obtain reference script description embedding features;

[0638] Perform self-attention calculation through the second model based on the reference script description embedding features to obtain a first attention calculation result;

[0639] Perform cross-attention calculation through the second model based on the first attention calculation result and the reference video frame features to obtain a second attention calculation result;

[0640] Generate a script description based on the second attention calculation result and a preset vocabulary in the second model to obtain the predicted script information for the reference video frame.

[0641] Optionally, a second training unit (not shown) is used for:

[0642] For each reference video frame, determine a third loss sub-function based on the reference video frame script description information and the predicted script description information;

[0643] Determine a second loss function based on multiple third loss sub-functions;

[0644] Train the second model based on the second loss function to obtain a video script description generation model.

[0645] Optionally, the predicted script description information includes the first probability and the first serial number of each of multiple predicted words, and the reference video frame script description information includes the second serial number of each of multiple reference words;

[0646] For each reference video frame, determining a third loss sub-function based on the reference video frame script description information and the predicted script description information includes:

[0647] For a single predicted word, if the predicted word is consistent with the reference word whose second serial number is the same as the first serial number of the single predicted word, determine the predicted word as the target word;

[0648] Determine the third loss sub-function based on the negative logarithm of the first probability of the target word.

[0649] Refer to Fig.28 , Fig.28 As shown in the structural block diagram of a part of the terminal for implementing the role perspective video generation method of the embodiments of the present disclosure, the terminal may be Figure 1 the object terminal shown. The terminal includes: a Radio Frequency (RF) circuit 2810, a memory 2815, an input unit 2830, a display unit 2840, a sensor 2850, an audio circuit 2860, a wireless fidelity (WiFi) module 2870, a processor 2880, and a power supply 2890, etc. Those skilled in the art can understand that Fig.28 the shown terminal structure does not constitute a limitation on a mobile phone or a computer, and may include more or fewer components than shown, or combine some components, or have different component arrangements.

[0650] The RF circuit 2810 can be used for receiving and sending signals during information reception or call processes. Specifically, after receiving the downlink information from the base station, it is given to the processor 2880 for processing; in addition, the designed uplink data is sent to the base station.

[0651] The memory 2815 can be used to store software programs and modules. The processor 2880 executes various functional applications and data processing of the object terminal by running the software programs and modules stored in the memory 2815.

[0652] The input unit 2830 can be used to receive input digital or character information, and generate key signal inputs related to the settings and function controls of the object terminal. Specifically, the input unit 2830 can include a touch panel 2831 and other input devices 2832.

[0653] The display unit 2840 can be used to display the input information or provided information and various menus of the object terminal. The display unit 2840 can include a display panel 2841.

[0654] The audio circuit 2860, the speaker 2861, and the microphone 2862 can provide an audio interface.

[0655] In this embodiment, the processor 2880 included in the terminal can execute the role perspective video generation method of the previous embodiment.

[0656] Fig.29 It is a structural block diagram of a part of the server for implementing the role perspective video generation method of the embodiments of the present disclosure. The server can be Figure 1 the video processing server shown. The server can vary greatly due to configuration or performance differences, and can include one or more central processing units (CPUs) 2922 (for example, one or more processors) and a memory 2932, and one or more storage media 2930 (for example, one or more mass storage devices) for storing application programs 2942 or data 2944. Among them, the memory 2932 and the storage media 2930 can be transient storage or persistent storage. The programs stored in the storage media 2930 can include one or more modules (not shown in the figure), and each module can include a series of instruction operations on the server. Further, the central processor 2922 can be set to communicate with the storage media 2930 and execute a series of instruction operations in the storage media 2930 on the server.

[0657] The server can also include one or more power supplies 2926, one or more wired or wireless network interfaces 2950, one or more input / output interfaces 2958, and / or one or more operating systems 2941, such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM, and so on.

[0658] The central processing unit 2922 in the server can be used to execute the role perspective video generation method of the embodiments of the present disclosure.

[0659] The embodiments of the present disclosure also provide a computer-readable storage medium for storing a computer program for executing the role perspective video generation method of the foregoing various embodiments.

[0660] The embodiments of the present disclosure also provide a computer program product including a computer program. The processor of the electronic device reads and executes the computer program, so that the electronic device executes the role perspective video generation method described above.

[0661] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of the present disclosure and the above drawings are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that such used data may be interchanged under appropriate circumstances, so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "including" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily limit to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0662] It should be understood that in the present disclosure, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects and indicates that there can be three relationships. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally means that the associated objects before and after are in an "or" relationship. "At least one (piece) of the following" or similar expressions refer to any combination of these items, including any combination of single item (piece) or plural items (pieces). For example, at least one (piece) of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.

[0663] It should be understood that in the description of the embodiments of the present disclosure, the meaning of multiple (or multiple items) is more than two. Understandings such as greater than, less than, exceeding, etc. do not include the present number, and understandings such as above, below, within, etc. include the present number.

[0664] In several embodiments provided by the present disclosure, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections between each other can be through some interfaces, and the indirect couplings or communication connections of devices or units can be in electrical, mechanical, or other forms.

[0665] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0666] In addition, each functional unit in various embodiments of the present disclosure can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.

[0667] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present disclosure, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in various embodiments of the present disclosure. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs that can store program codes.

[0668] It should also be understood that the various implementation manners provided in the embodiments of the present disclosure can be combined arbitrarily to achieve different technical effects.

[0669] The above is a specific description of the embodiments of the present disclosure. However, the present disclosure is not limited to the above embodiments. Those skilled in the art can also make various equivalent deformations or substitutions without departing from the spirit of the present disclosure, and these equivalent deformations or substitutions are all included within the scope defined by the claims of the present disclosure.

Claims

1. A method for generating a character perspective video, characterized in that: The method comprises: Acquire a source video having a target character and source background description information of the source video, wherein the source video is a video from a third-party perspective; Extracting source script information from the source video; Obtain the introductory template information corresponding to the role script conversion model; Based on the source script information, the source background description information and the guide template information, the role script information from the perspective of the target character is generated by using the role script conversion model; Generate diffusion guidance information based on the role script information, the picture information corresponding to the target role, and the source background description information, and generate a role perspective video from the perspective of the target role based on the diffusion guidance information, wherein the perspective of the target role is different from the perspective of the third party; The source script information includes source lines information and source scene information, the role script information includes role script lines information and role script scene information; the introductory language template information includes first introductory language template information and second introductory language template information; the role script information from the perspective of the target role is generated based on the source script information, the source background description information and the introductory language template information and using the role script conversion model, including: Obtaining picture information corresponding to the target character; Based on the source dialogue information, the source scene information, the picture information and the source background description information, the first guide language template information is used to generate a first guide language, and the first guide language is input into the role script conversion model to obtain the role script dialogue information; Based on the source scene information and the source background description information, the second guide language is generated using the second guide language template information, and the second guide language is input into the role-script conversion model to obtain the role-script scene information.

2. The method for generating a character perspective video according to claim 1, characterized in that: The extracting source script information from the source video includes: Segmenting the source video to obtain a plurality of source video frames arranged in time sequence; Inputting the plurality of source video frames into a pre-trained video script description generation model for description generation, and obtaining video frame script description information of each of the plurality of source video frames; The video frame script description information of each of the multiple source video frames is integrated to obtain the source script information.

3. The method for generating a character perspective video according to claim 2, characterized in that: The video script description generation model includes a feature extraction network and a script description generation network; The step of inputting the plurality of source video frames into a pre-trained video script description generation model for description generation to obtain video frame script description information of each of the plurality of source video frames includes: For each of the source video frames, extract features of the source video frame based on the feature extraction network to obtain target video frame features of the source video frame; Based on the target video frame features, the script description generation network is used to generate a script to obtain the video frame script description information of the source video frame.

4. The method for generating a character perspective video according to claim 3, characterized in that: The step of extracting features of each of the source video frames based on the feature extraction network to obtain target video frame features of the source video frames includes: For each of the source video frames, performing multi-scale downsampling on the source video frame based on the feature extraction network to obtain a first downsampling feature, a second downsampling feature, and a third downsampling feature, wherein the feature dimensions of the first downsampling feature, the second downsampling feature, and the third downsampling feature are increased in sequence; Performing dimensionality reduction processing on the first down-sampling feature, the second down-sampling feature, and the third down-sampling feature, respectively, to obtain a first dimensionality reduction feature corresponding to the first down-sampling feature, a second dimensionality reduction feature corresponding to the second down-sampling feature, and a third dimensionality reduction feature corresponding to the third down-sampling feature; Performing splicing processing on the first dimensionality reduction feature, the second dimensionality reduction feature, and the third dimensionality reduction feature to obtain an initial video frame feature of the source video frame; The initial video frame features of the source video frame and the initial video frame features of a previous source video frame of the source video frame are spliced ​​to obtain target video frame features of the source video frame.

5. The method for generating a character perspective video according to claim 1, characterized in that: The step of generating diffusion guidance information based on the role script information, the picture information corresponding to the target role, and the source background description information includes: Encoding the source background description information to obtain background description encoding features; Encoding the role script information to obtain role script encoding features; Extracting features from the image information to obtain features of the target character image; Based on the background description coding features, the role script coding features, and the target role image features, conditional information is generated through a preset conditional fusion network to obtain the diffusion guidance information.

6. The method for generating a character perspective video according to claim 2, characterized in that: The step of integrating the video frame script description information of each of the plurality of source video frames to obtain the source script information includes: splicing the video frame script description information of each of the multiple source video frames to obtain preliminary script information; The preliminary script information is polished to obtain the source script information.

7. The method for generating a character perspective video according to claim 1, characterized in that: The generating the character perspective video of the target character perspective based on the diffusion guidance information includes: guiding a diffusion model based on the diffusion guidance information to generate the character perspective video of the target character perspective; The diffusion model is trained in the following way: Acquire a sample video having a sample character, sample background description information of the sample video, and an expected character perspective video of the sample character; Extracting sample script information from the sample video; Based on the sample script information, the sample background description information and the preset introductory language template information, using the role script conversion model, generate the sample role script information from the perspective of the sample role; Generate sample guidance information based on the sample role script information, the sample picture information corresponding to the sample role, and the sample background description information; Based on the sample guidance information, a video is generated through a target model to obtain a sample character perspective video of the sample character perspective; The target model is trained based on the sample character perspective video and the expected character perspective video to obtain the diffusion model.

8. The method for generating a character perspective video according to claim 7, characterized in that: The sample video includes a plurality of sample video frames; the target model includes a diffusion network, a denoising network, and a decoding network; The step of generating a video based on the sample guide information by using a target model to obtain a sample role perspective video of the sample role perspective includes: For each of the sample video frames, diffusion processing is performed on the sample video frame encoding features of the sample video frame based on the diffusion network to obtain a sample latent space feature vector; Based on the sample guidance information, the sample latent space feature vector is denoised by the denoising network to obtain a sample video frame denoising feature; Decoding the denoising features of the sample video frames based on the decoding network to obtain predicted video frames of each of the sample video frames; The sample character perspective video is obtained based on the predicted video frames of each of the multiple sample video frames.

9. The method for generating a character perspective video according to claim 8, characterized in that: The denoising network includes an upsampling module and a downsampling module; The denoising process is performed on the sample latent space feature vector based on the sample guide information by the denoising network to obtain the sample video frame denoising feature, including: Performing splicing processing on the sample guide information and the sample latent space feature vector to obtain a sample denoising splicing feature; Downsampling the sample denoising splicing features based on the downsampling module to obtain video frame downsampling features; The down-sampled features of the video frame are up-sampled based on the up-sampling module to obtain the denoising features of the sample video frame.

10. The method for generating a character perspective video according to claim 8, characterized in that: The denoising process is performed on the sample latent space feature vector based on the sample guide information by the denoising network to obtain the sample video frame denoising feature, including: Performing splicing processing on the sample guide information and the sample latent space feature vector to obtain a sample denoising splicing feature; Based on the sample denoising splicing features, performing a first denoising process through the denoising network to obtain a first denoising feature; Based on the sample latent space feature vector, performing a second denoising process through the denoising network to obtain a second denoising feature; The first denoising feature and the second denoising feature are weighted to obtain the sample video frame denoising feature.

11. The method for generating a character perspective video according to claim 8, characterized in that: The sample video includes a plurality of sample video frames; The step of training the target model based on the sample role perspective video and the expected role perspective video to obtain the diffusion model includes: Determining a first loss sub-function based on the sample character perspective video and the expected character perspective video; Determining a second loss sub-function based on the sample video frame denoising features corresponding to each of the plurality of sample video frames and a preset benchmark denoising feature; Based on the first loss sub-function and the second loss sub-function, a first loss function is obtained; Based on the first loss function, the target model is trained to obtain the diffusion model.

12. The method for generating a character perspective video according to claim 7, characterized in that: The target model is obtained by adjusting the first model through the following steps: Acquire multiple reference video frames of a reference video; Performing three-dimensional convolution processing on the multiple reference video frames to obtain predicted means and predicted standard deviations of the multiple reference video frames; Based on the predicted mean and the predicted standard deviation, the multiple reference video frames are respectively re-parameterized to obtain video frame hidden state features of the multiple reference video frames; Decoding the video frame hidden state features of each of the multiple reference video frames to obtain predicted video frames corresponding to each of the multiple reference video frames; Based on the predicted video frame, the reference video frame, the predicted mean and the predicted standard deviation, the parameters of the first model are adjusted to obtain the target model.

13. The method for generating a character perspective video according to claim 2, characterized in that: The video script description generation model is trained in the following way: Acquire a plurality of reference video frames arranged in time sequence of a reference video, and reference script description information corresponding to each of the plurality of reference video frames; extracting reference video frame features of each of the multiple reference video frames; Based on the reference video frame features of each of the multiple reference video frames and the reference script description information, script generation is performed through a preset second model to obtain predicted script information of each of the multiple reference video frames; Based on the reference script description information and the predicted script information of each of the multiple reference video frames, the second model is trained to obtain the video script description generation model.

14. The method for generating a character perspective video according to claim 13, characterized in that: The step of generating a script based on the reference video frame features of each of the plurality of reference video frames and the reference script description information by using a preset second model to obtain predicted script information of each of the plurality of reference video frames includes: For each of the reference video frames, embedding the reference script description information to obtain a reference script description embedding feature; Based on the reference script description embedding feature, performing self-attention calculation through the second model to obtain a first attention calculation result; Based on the first attention calculation result and the reference video frame feature, performing cross attention calculation through the second model to obtain a second attention calculation result; Based on the second attention calculation result and the preset vocabulary in the second model, a script description is generated to obtain the predicted script information of the reference video frame.

15. A device for generating a character perspective video, characterized in that: The device comprises: A first acquisition unit is used to acquire a source video having a target character and source background description information of the source video, wherein the source video is a video from a third-party perspective; An extraction unit, used for extracting source script information from the source video; The second acquisition unit is used to acquire the introductory language template information corresponding to the role script conversion model; A first generating unit, configured to generate role script information from the perspective of the target role using the role script conversion model based on the source script information, the source background description information and the guide template information; A second generating unit is used to generate diffusion guidance information based on the role script information, the picture information corresponding to the target role, and the source background description information, and generate a role perspective video from the perspective of the target role based on the diffusion guidance information, wherein the perspective of the target role is different from the perspective of the third party; The source script information includes source lines information and source scene information, the role script information includes role script lines information and role script scene information; the introductory language template information includes first introductory language template information and second introductory language template information; the role script information from the perspective of the target role is generated based on the source script information, the source background description information and the introductory language template information and using the role script conversion model, including: Obtaining picture information corresponding to the target character; Based on the source dialogue information, the source scene information, the picture information and the source background description information, the first guide language template information is used to generate a first guide language, and the first guide language is input into the role script conversion model to obtain the role script dialogue information; Based on the source scene information and the source background description information, the second guide language is generated using the second guide language template information, and the second guide language is input into the role-script conversion model to obtain the role-script scene information.

16. An electronic device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, it implements the character perspective video generation method according to any one of claims 1 to 14.

17. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method for generating a character perspective video according to any one of claims 1 to 14 is implemented.

18. A computer program product, comprising a computer program, wherein the computer program is read and executed by a processor of an electronic device, so that the electronic device executes the method for generating a character perspective video according to any one of claims 1 to 14.

Citation Information

Patent Citations

  • Data processing method and device, electronic equipment and storage medium

    CN113821690A