Video feature extraction method and apparatus, video generation method and apparatus, and medium and device

By receiving the target video and extracting features in multiple target dimensions during the video understanding process, stratifying and feature extraction are performed, the problem of low feature extraction efficiency and accuracy in video generation is solved, and more efficient and accurate video feature extraction and video generation are achieved.

WO2025092911A1PCT designated stage expired Publication Date: 2025-05-08BEIJING YOUZHUJU NETWORK TECH CO LTD

Patent Information

Application Number
PCT/CN2024/128921
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-10-31
Filing Date
2024-10-31
Publication Date
2025-05-08

AI Technical Summary

Technical Problem

In the process of video understanding, it is difficult to extract video features efficiently and accurately, resulting in low video generation efficiency and accuracy.

Method used

By receiving the target video to be processed, it determines its characteristic information under multiple target dimensions (subject content dimension and object content dimension), performs storyboard detection, extracts the video frame characteristics in the storyboard segment, and generates storyboard script information.

Benefits of technology

It improves the accuracy and comprehensiveness of video feature extraction, achieves the matching degree between video features and target video content, and provides accurate training data for the video generation model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024128921_08052025_PF_FP_ABST
    Figure CN2024128921_08052025_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure relates to a video feature extraction method and apparatus, a video generation method and apparatus, and a medium and a device. The video feature extraction method comprises: receiving a target video to be processed; determining feature information of the target video in a plurality of target dimensions, wherein the target dimensions at least include a subject content dimension and an object content dimension; performing shot detection on the target video on the basis of the feature information in the plurality of target dimensions, so as to determine shot segments in the target video; and for each shot segment, performing feature extraction on video frames in the shot segment, and generating on the basis of extracted information shot script information corresponding to the shot segment. Thus, the embodiments of the present disclosure can realize shot detection on the basis of changes in subject content and object content in the target video, such that the accuracy of video shots can be improved; and feature extraction can be performed from the levels of the shot segments, thus improving the accuracy and comprehensiveness of video features, and realizing automatic conversion from the target video to a shot script.
Need to check novelty before this filing date? Find Prior Art

Description

Video feature extraction method, video generation method, device, medium and equipment

[0001] This disclosure claims priority to Chinese Patent Application No. 202311436271.7 filed on October 31, 2023, and the contents of the above-mentioned Chinese patent application disclosure are hereby incorporated by reference in their entirety as part of this disclosure. Technical Field

[0002] The present disclosure relates to a video feature extraction method, a video generation method, an apparatus, a medium and a device. Background Art

[0003] With the development of computer technology, video applications have gradually facilitated people's daily lives. In related technologies, video generation can be achieved by manually uploading images or videos and then editing them. Another method is automatic video generation based on video understanding. However, video understanding involves temporal and dynamically changing information, making it more difficult to understand than natural language or image content. This results in lower efficiency and accuracy in video feature extraction.

[0004] Summary of the Invention

[0005] This summary is provided to briefly introduce concepts that will be described in detail in the detailed description below. This summary is not intended to identify key features or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.

[0006] In a first aspect, the present disclosure provides a video feature extraction method, the method comprising:

[0007] Receive the target video to be processed;

[0008] Determining feature information of the target video in multiple target dimensions, wherein the target dimensions include at least a subject content dimension and an object content dimension;

[0009] Performing frame detection on the target video according to the feature information under the multiple target dimensions to determine frame segments in the target video;

[0010] For each of the storyboard segments, feature extraction is performed on the video frames in the storyboard segment, and storyboard script information corresponding to the storyboard segment is generated based on the extracted information.

[0011] In a second aspect, the present disclosure provides a video generation method, the method comprising:

[0012] receiving a target text, wherein the target text includes at least one storyboard;

[0013] According to the target text and video generation model, a target video corresponding to the target text is generated, wherein the video generation model is trained based on training video samples and annotation scripts corresponding to the training video samples, and the annotation scripts corresponding to the training video samples are generated based on the video feature extraction method described in the first aspect.

[0014] In a third aspect, the present disclosure provides a video feature extraction device, the device comprising:

[0015] A first receiving module, configured to receive a target video to be processed;

[0016] A first determining module is configured to determine feature information of the target video in multiple target dimensions, wherein the target dimensions include at least a subject content dimension and an object content dimension;

[0017] A detection module, configured to perform frame detection on the target video based on feature information under the plurality of target dimensions to determine frame segments in the target video;

[0018] The processing module is used to extract features of the video frames in each storyboard segment and generate storyboard script information corresponding to the storyboard segment based on the extracted information.

[0019] In a fourth aspect, the present disclosure provides a video generation device, comprising:

[0020] A second receiving module is configured to receive a target text, wherein the target text includes at least one storyboard script;

[0021] A generation module is used to generate a target video corresponding to the target text based on the target text and the video generation model, wherein the video generation model is trained based on training video samples and annotation scripts corresponding to the training video samples, and the annotation scripts corresponding to the training video samples are generated based on the video feature extraction method described in the first aspect.

[0022] In a fifth aspect, the present disclosure provides a computer-readable medium having a computer program stored thereon, which, when executed by a processing device, implements the steps of the method described in the first aspect or the second aspect.

[0023] In a sixth aspect, the present disclosure provides an electronic device, including:

[0024] a storage device having a computer program stored thereon;

[0025] A processing device is used to execute the computer program in the storage device to implement the steps of the method of the first aspect or the second aspect.

[0026] Other features and advantages of the present disclosure will be described in detail in the following detailed description. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] The above and other features, advantages and aspects of the various embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and that the originals and elements are not necessarily drawn to scale. In the drawings:

[0028] FIG1 is a flowchart of a video feature extraction method provided according to an embodiment of the present disclosure.

[0029] FIG2 is a block diagram of a video feature extraction apparatus according to an embodiment of the present disclosure.

[0030] FIG3 shows a schematic structural diagram of an electronic device suitable for implementing the embodiments of the present disclosure. DETAILED DESCRIPTION

[0031] The following describes embodiments of the present disclosure in more detail with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments described herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.

[0032] It should be understood that the various steps described in the method embodiments of the present disclosure may be performed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this respect.

[0033] As used herein, the term "including" and its variations are open-ended, i.e., "including but not limited to." The term "based on" means "based, at least in part, on." The term "one embodiment" means "at least one embodiment," the term "another embodiment" means "at least one additional embodiment," and the term "some embodiments" means "at least some embodiments." Other terms are defined in the following description.

[0034] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.

[0035] It should be noted that the modifications of "one" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, they should be understood as "one or more".

[0036] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.

[0037] It is understandable that before using the technical solutions disclosed in the various embodiments of this disclosure, the type, scope of use, usage scenarios, etc. of the personal information involved in this disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.

[0038] For example, in response to a user's active request, a prompt message is sent to the user to clearly inform the user that the operation requested will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the electronic device, application, server, storage medium, or other software or hardware that performs the operations of the disclosed technical solution based on the prompt message.

[0039] As an optional but non-limiting implementation, in response to receiving a user's active request, the prompt information may be sent to the user in the form of a pop-up window, in which the prompt information may be presented in text form. Furthermore, the pop-up window may also contain a selection control for the user to select "agree" or "disagree" to provide personal information to the electronic device.

[0040] It is understandable that the above notification and user authorization process are merely illustrative and do not limit the implementation of the present disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of the present disclosure.

[0041] At the same time, it is understandable that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) must comply with the requirements of relevant laws, regulations and relevant provisions.

[0042] FIG1 is a flowchart of a video feature extraction method according to an embodiment of the present disclosure. As shown in FIG1 , the method may include:

[0043] In step 11, a target video to be processed is received.

[0044] The target video may be a video uploaded by a user or a video received from other service interfaces.

[0045] In step 12, feature information of the target video in multiple target dimensions is determined, wherein the target dimensions at least include a subject content dimension and an object content dimension.

[0046] The subject content dimension is used to represent the characteristics of the subject object in the target video, and the object content dimension is used to represent the characteristics of the object object in the target video. The applicant has discovered through research that shot switching typically occurs when the subject object and object object change. Therefore, in this embodiment, the characteristics of the target video in the subject content dimension and the object content dimension can be extracted separately to provide data support for subsequent accurate shot detection.

[0047] In step 13, based on the feature information under multiple target dimensions, the target video is subjected to frame detection to determine the frame segments in the target video.

[0048] A storyboard, also known as a storyboard, is a diagram used to illustrate the composition of various visual media, such as movies, animations, TV series, commercials, and music videos, before actual filming or drawing. It breaks down continuous footage into individual shot movements and labels the shot movement method, duration, dialogue, special effects, etc. Therefore, a single shot can be used to represent continuous and related content within a short period of time in a video. In this embodiment, a target video can be split into multiple shot segments, which are then concatenated to form the complete target video.

[0049] Therefore, in this embodiment, the target video can be split into storyboards to obtain multiple storyboard segments contained within the target video. As described above, each storyboard segment can contain continuous and interrelated content, facilitating analysis and feature extraction of the target video based on the content of each storyboard segment. Extracting features from the video at the storyboard segment level can, to a certain extent, ensure the correlation between the extracted features, thereby improving the matching between the extracted features and the content of the target video.

[0050] In step 14, for each storyboard segment, feature extraction is performed on the video frames in the storyboard segment, and storyboard script information corresponding to the storyboard segment is generated based on the extracted information.

[0051] Feature extraction can be performed on the video frames in the storyboard segment using a multimodal feature extraction approach, such as visual feature extraction based on the image content in the video frames, sound feature extraction based on the sound content in the video frames, and scene change feature extraction based on the consecutive video frames in the storyboard segment. This can be implemented based on image feature extraction and sound feature extraction algorithms or models known in the art, and will not be further elaborated here.

[0052] After the corresponding information is extracted, the information can be added to the storyboard script information to perform standardized storage and display of the information of the storyboard segment.

[0053] Thus, in the above technical solution, by determining the feature information of the target video under multiple target dimensions, the target video can be subjected to frame detection based on the feature information under the multiple target dimensions to determine the frame segments in the target video. Then, for each frame segment, feature extraction can be performed on the video frames in the frame segment, and storyboard information corresponding to the frame segment can be generated based on the extracted information. Through the above technical solution, frame detection can be achieved based on the changes in the subject content and object content in the target video, thereby improving the accuracy of video frame detection. Then, feature extraction can be performed on each frame segment at the frame segment level. On the one hand, the correlation between the extracted video features and the matching degree between the video features and the overall content of the target video can be ensured, thereby improving the accuracy and comprehensiveness of the video features. On the other hand, the frame segment information corresponding to the frame segment can be generated by extracting the video features, realizing the automatic conversion of the target video to the frame segment without manual recording and annotation, saving manual workload and providing a large amount of accurate training data for the subsequent video generation model based on the script generation, providing effective data support for the subsequent training of the video generation model. At the same time, the timeline association between multiple storyboard segments can be obtained, and the arrangement characteristics of the target video and the multiple storyboard segments in the time dimension can be obtained, providing a reference for ensuring the logical accuracy of the video generation process in the subsequent video generation model.

[0054] In a possible embodiment, an exemplary implementation of performing frame detection on the target video based on the feature information under the multiple target dimensions to determine frame segments in the target video may include:

[0055] A multimodal content matrix is ​​generated based on feature information of the target video in multiple target dimensions.

[0056] The feature information under multiple target dimensions may be spliced ​​together to obtain the multimodal content matrix, so as to comprehensively characterize the features of the target video based on the multimodal content matrix.

[0057] Afterwards, storyboard detection is performed based on a storyboard detection model and the multimodal content matrix to obtain a plurality of storyboard segments, wherein the storyboard detection model is determined based on a training video and training annotated storyboards corresponding to the training video, and the training annotated storyboards are annotated based on the multiple target dimensions.

[0058] The scene detection model may be a pre-trained model. As an example, the scene detection model may be determined by:

[0059] A training sample set is obtained, where the training sample set may include a training video and a training annotated storyboard corresponding to the training video.

[0060] As an example, when annotating the storyboards corresponding to the training video, standard divisions can be performed according to the multiple target dimensions, so that the annotators can perform corresponding annotations based on the characteristics of the training video under the multiple target dimensions and determine the multiple storyboards corresponding to the training video.

[0061] As another example, process files during the video generation process may be obtained, such as a training video and an unedited video corresponding to the training video, so as to determine a plurality of storyboards in the training video based on the content of the unedited video and the training video.

[0062] In a possible embodiment, the target dimension further includes at least one of a camera dimension, a scene dimension, a music dimension, and an aesthetic dimension.

[0063] For example, scene annotation can be performed based on changes in features in the target dimension. For example, in the subject content dimension, whether to perform scene splitting can be determined based on changes in the subject object, changes in the number of subject objects, and changes in the relationship between subject objects. For example, in the object content dimension, whether to perform scene splitting can be determined based on changes in the object object, changes in the number of object objects, and changes in the relationship between object objects.

[0064] The camera movement dimension can be used to represent changes in shooting methods during video shooting. Usually, continuous images can be decomposed into a storyboard based on one camera movement. The changes in camera movement methods can be used to provide a reference for whether to split the storyboards, such as camera movement changes, scene changes, etc.

[0065] The scene dimension can be used to represent scenes during video shooting. Different scenes usually correspond to different storyboards. The changes in the scene can be determined to provide a reference for whether to split the storyboards, such as changes in location and composition.

[0066] The music dimension can be used to represent the background music in the video. Different background music usually corresponds to different storyboards. By determining the changes in the music, you can provide a reference for whether to split the storyboards, switch, pause or add background music, etc.

[0067] The aesthetic dimension can be used to represent the style and atmosphere of the video. Different styles or atmospheres usually correspond to different storyboards. The changes in the aesthetic dimension can be used to provide a reference for whether to split the storyboards, such as changes in picture style or picture atmosphere.

[0068] Among them, the standards and rules for determining whether storyboard splitting is required based on the features under multiple target dimensions can be pre-set based on the actual application scenario. This disclosure does not limit this, and it is sufficient to ensure that the standards and rules for splitting are the same in the same model.

[0069] In this embodiment, the training and annotated storyboards can be referred to and annotated from the above-mentioned feature dimensions when annotating, so that the segmentation standard of the training and annotated storyboards conforms to the actual storyboard shooting process, thereby improving the accuracy of multiple storyboards obtained by the storyboard detection model to a certain extent, and making the predicted division of multiple storyboard segments match the actual video shooting process, thereby improving the accuracy and rationality of the storyboard segment division.

[0070] Afterwards, the training video can be used as the input of the model and the training labeled storyboards can be used as the target output of the model for training, thereby obtaining a storyboard detection model.

[0071] The frame detection model can be implemented based on a neural network model with an attention mechanism. Based on the input training video, a corresponding multimodal content matrix can be obtained and processed with the attention mechanism to obtain multiple predicted frames as output by the model. Subsequently, the rationality of the predicted frames can be determined based on the division criteria and rules of multiple target dimensions. The model can be adjusted based on the similarity between the predicted frames and the training annotated frames until a frame detection model that meets the training requirements is obtained. Meeting the training requirements can mean that the accuracy of the predicted frames reaches a preset threshold, or that the number of model iterations reaches a threshold, which is not limited in this disclosure.

[0072] Thus, through the above technical solution, the target video can be split into storyboards using the storyboard detection model to obtain multiple storyboard segments, thereby enabling video feature extraction and understanding in units of storyboard segments. The training annotated storyboards of the training video corresponding to the storyboard detection model are annotated based on the multiple target dimensions, thereby ensuring the interpretability of the storyboard segments obtained based on the storyboard detection model, ensuring the accuracy and effectiveness of the storyboard splitting, and improving the effective guidance of the tuning direction of the video generation model when the obtained storyboard script is used in the subsequent video generation model.

[0073] In a possible embodiment, determining feature information of the target video in multiple target dimensions may include:

[0074] Obtain a target video frame in the target video.

[0075] Among them, all video frames in the target video can be used as the target video frame, or frames can be extracted from the target video at intervals of a preset time length to obtain the target video frame. The preset time length can be set based on the actual application scenario, and this disclosure does not limit this.

[0076] Object recognition is performed on the target video frame to determine candidate objects in the target video.

[0077] Among them, for each target video frame, object recognition can be performed on it. As an example, the object can include people, animals, static objects, etc., and the objects in the target video frame can be classified and recognized by pre-training an object classification and recognition model, such as can be implemented based on a classification model commonly used in this field.

[0078] As an example, each object identified from the target video can be used as a candidate object. As another example, the categories corresponding to the candidate objects can be preset, and objects belonging to the preset categories among the identified objects can be used as candidate objects.

[0079] Display information of each candidate object in the target video is determined, and a subject object and an object object in the candidate objects are determined according to the display information.

[0080] The display information is used to represent the characteristics of the candidate object when displayed in the target video. For example, the display information may include at least one of the candidate object's display duration, display frequency, and display match with the subject of the target video. The display duration can be represented by the number of video frames containing the candidate object, the display frequency can be determined by the ratio of the number of video frames containing the candidate object to the target video frames, and the display match can be determined based on a vector similarity calculation between the candidate object and the subject of the target video. These details are not elaborated here.

[0081] As an example, for each candidate object, the display components can be determined based on the above display information, and the display components can be weighted based on the weight of each display information to obtain a comprehensive display score for the candidate object. The subject object and the object object are then determined based on the comprehensive display score. For example, the object with the largest comprehensive display score can be used as the subject object, and the objects ranked 2-N in descending order of comprehensive display scores can be used as object objects, where N can be set according to the actual application scenario. For another example, the number S of subject objects and the number T of object objects can be preset respectively, and then the first S objects can be selected as the subject objects in descending order according to the comprehensive display scores, and then T objects can be selected as object objects.

[0082] For another example, the display information may include the ratio of the candidate objects to the overall image. The subject and object objects can then be determined based on the ratio of the candidate objects to the overall image. Alternatively, the above-described weighted approach can be used to combine multiple pieces of display information for comprehensive weighted determination to identify the subject and object objects. This will not be further elaborated here.

[0083] The above embodiment is merely an example and does not limit the present disclosure. In this embodiment, the subject object and the object object in the target video can be determined based on the display information of the object in the video using a unified standard.

[0084] Afterwards, the feature information corresponding to the subject object is determined as the feature information under the subject content dimension, and the feature information corresponding to the object object is determined as the feature information under the object content dimension.

[0085] Accordingly, the feature information of the subject object determined during the object recognition process can be used as the feature information under the subject content dimension. For example, the feature information of the subject object can be the corresponding coding features of the subject object extracted during the model processing process, such as the coding features of the subject object determined in the classification recognition model. Similarly, the feature information under the object content dimension can be determined in a similar manner.

[0086] Therefore, through the above technical solution, the main content and object content in the target video can be extracted and represented, providing support for subsequent storyboard detection based on changes in the main content and object content, so that the process of storyboard splitting corresponds to the process of storyboard splicing during the video generation process, thereby improving the accuracy of storyboard splitting to a certain extent.

[0087] In a possible embodiment, an exemplary implementation of extracting features from the video frames in the storyboard segments and generating storyboard script information corresponding to the storyboard segments based on the extracted information may include:

[0088] Based on the video frames in the storyboard fragments, feature information of the storyboard fragments under the classification identification dimension corresponding to each script dimension is extracted, wherein the script dimension includes a subject content dimension and an object content dimension, and at least one classification identification dimension is preset under each script dimension.

[0089] The script dimension can be pre-set based on the actual application scenario. As an example, the script dimension can also include at least one of a shot dimension, a scene dimension, a sound dimension, a text dimension, and an aesthetic dimension, thereby enabling feature extraction from video frames from multiple dimensions to ensure the comprehensiveness and accuracy of feature extraction. The text dimension can be used to represent subtitle information in video frames, such as through OCR (Optical Character Recognition) technology.

[0090] Classification and identification dimensions under each script dimension can be preset based on actual application scenarios, wherein the classification and identification dimensions are more fine-grained divisions under the script dimension. For example, the classification and identification dimensions under the subject content dimension and corresponding to the object content dimension include at least one of a trait dimension, an action and behavior dimension, a trait change dimension, and a subject-object interaction dimension.

[0091] The following uses the script dimension as an example to illustrate the main content dimension. For example, the first feature information identified is the trait dimension, which can represent the attribute characteristics of the subject object itself, such as a boy wearing black clothes; the action behavior dimension can be used to represent the behavior and action events that cause the subject object's state to change, such as a boy preparing to shoot; the trait change dimension can be used to represent events that cause the subject object's traits to change, such as a boy looking at the basket; and the subject-object interaction dimension is used to represent the relationship between the subject object and the object object, such as attribute associations, relative spatial positions, and behavioral interactions, such as cheerleaders standing in a row on the court.

[0092] As another example, the classification and identification dimension corresponding to the lens dimension may include at least one of the shot dimension, the camera movement dimension, the angle dimension, and the subject-object relationship dimension. Among them, the shot refers to the difference in the size of the range of the subject presented in the camera recorder due to the different distances between the camera and the subject when the focal length is constant. The value of the shot dimension can be determined based on the shot division method commonly used in this field, such as long shot, panorama, etc. The value of the camera movement dimension can be set based on the common camera movement method, such as push, pull, pan, etc.; the value of the angle dimension can be set based on the common angle setting method in this field, such as pitch, tilt, etc.; the subject-object relationship dimension is used to represent the relationship and temporal changes of the shot, camera movement, angle, etc. relative to the subject object. Therefore, the feature information such as shot, camera movement and angle corresponding to the storyboard segment can be obtained by extracting the video features in the storyboard segment.

[0093] The classification and identification dimensions corresponding to the scene dimension can include at least one of the location dimension, the environmental composition dimension, the environmental style dimension, and the subject-object relationship dimension. The location dimension describes the attributes of a location, such as temporal and spatial properties, and place names; the environmental composition dimension represents the elements present in the scene environment, their characteristics, and their relative spatial positions; the environmental style dimension represents the style of the scene environment, and its value can be set through classification; and the subject-object relationship dimension represents the relationship and temporal changes of the location, environmental composition, and other factors relative to the subject object. This allows for a more comprehensive and accurate description of the scenes in the storyboard clips using these classification and identification dimensions.

[0094] The sound dimension can include the human voice dimension, the background sound dimension, and the sound effect dimension, which can be obtained by extracting features from the audio in the storyboard clips to represent the audio features in the storyboard clips; the text dimension can include the subtitle dimension and the artistic font dimension, so as to represent the text features in the storyboard clips, and the aesthetic dimension can include the color tone dimension and the atmosphere dimension.

[0095] It should be noted that a sub-model for feature extraction under the classification and recognition dimension can be trained for each script dimension. The sub-model can be implemented based on a classification model or encoder model in this field, and this disclosure does not limit this.

[0096] Afterwards, the key video frames in the storyboard segment are determined, and the storyboard description text corresponding to the storyboard segment and the object change information between the key video frames are generated. The storyboard script information includes the key video frames corresponding to the storyboard segment, the storyboard description text and the content change information between the key video frames.

[0097] For example, the script dimension can be used as a first-level directory, the classification identification dimension under the script dimension can be used as a second-level directory, and the feature information under the classification identification dimension can be written into the value of the classification identification dimension to obtain the feature information under the classification identification dimension. The feature information can then be input into the video understanding model to identify the key video frames in the storyboard segment based on the multiple feature information, and generate the storyboard description text corresponding to the storyboard segment and the content change information between the key video frames. Among them, the subject and object objects contained in different video frames may be different or the status of the subject or object objects may be different, so they can be represented based on the content change information between the key video frames to describe the change status of the objects between different video frames.

[0098] The video understanding model can be pre-trained, for example, by obtaining training data for the video understanding model based on static image frames corresponding to each storyboard produced during the filming of the target video, as well as information describing the changes between the storyboard and the multiple static image frames. The target video can then be input into the video understanding model, and training is performed using the static image frames of the multiple storyboard segments contained in the target video, as well as the information describing the changes between the storyboard and the multiple static image frames, as model target values ​​to obtain the video understanding model.

[0099] Therefore, in the process of generating storyboard script information, feature representation of secondary dimensions can be performed. On the one hand, the accuracy of feature extraction can be effectively improved, so that more comprehensive and rich features can be extracted from the storyboard fragments, thereby ensuring the accuracy of the storyboard script information, and at the same time making the storyboard script information more comprehensive and effective in describing the storyboard fragments. On the other hand, the storyboard fragments can be represented as a whole, and features can be extracted at the storyboard fragment level, further improving the matching degree between the feature extraction process and the actual video generation process.

[0100] In a possible embodiment, the method further includes:

[0101] Determine the storyboard display position corresponding to each storyboard segment in the target video.

[0102] The target video may contain multiple storyboard segments, and the display order of the multiple storyboard segments represents the time logic of the target video. For example, the storyboard display position can be used to indicate which storyboard segment the storyboard segment is displayed in the target video.

[0103] Afterwards, the storyboard display position corresponding to the storyboard segment is associated with the storyboard script information corresponding to the storyboard segment.

[0104] In this embodiment, while extracting the storyboard script information corresponding to each storyboard segment in the target video, the corresponding storyboard display position is associated with it, thereby characterizing the combination logic of different storyboard segments in the target video. Accordingly, when training a video generation model based on the storyboard segments and storyboard script information obtained in this process, the combination logic of the storyboard segments in the video can be learned synchronously during the video generation model training process. On the one hand, this ensures that the overall video display logic is considered during the video feature extraction process, thereby ensuring the accuracy of video feature extraction; on the other hand, it can provide a more comprehensive feature reference for subsequent video generation, improving the reasonable arrangement of the generated video in the time dimension.

[0105] The present disclosure also provides a video generation method, which may include:

[0106] A target text is received, wherein the target text includes at least one storyboard script.

[0107] The target text may be a detailed description of the video that the user wants to generate.

[0108] According to the target text and video generation model, a target video corresponding to the target text is generated, wherein the video generation model is trained based on training video samples and annotation scripts corresponding to the training video samples, and the annotation scripts corresponding to the training video samples are generated based on the video feature extraction method described above.

[0109] In this embodiment, feature extraction can be performed on the training video sample based on the video feature extraction method described above, so that the storyboard script corresponding to the training video sample can be obtained as the annotation script. In the process of video feature extraction, the corresponding storyboard segments are first detected, and then the storyboard script information corresponding to each storyboard segment is determined. When the video generation model is trained based on the training video sample and its annotation script, the video generation model can learn the corresponding features in the storyboard detection process and the timeline logic between the multiple storyboard segments contained in the training video sample. While quickly determining a large number of training samples to save the workload of manual annotation, the video generation process of the video generation model is improved. The matching degree between the real video editing processing process is improved, thereby improving the efficiency and accuracy of video generation and enhancing the user experience.

[0110] In a possible embodiment, the target video and the storyboard clips and storyboard script corresponding to the target video can be used as training data to train the script generation model, the picture analysis model, and the video shooting guidance model, etc., which can improve the processing accuracy of the above models to a certain extent, and can also effectively reduce the amount of manual annotation required to generate training data.

[0111] The present disclosure further provides a video feature extraction device, as shown in FIG2 , wherein the device 10 includes:

[0112] A first receiving module 100 is configured to receive a target video to be processed;

[0113] A first determination module 200 is configured to determine feature information of the target video in multiple target dimensions, wherein the target dimensions include at least a subject content dimension and an object content dimension;

[0114] A detection module 300 is configured to perform frame detection on the target video based on the feature information under the multiple target dimensions to determine frame segments in the target video;

[0115] The processing module 400 is used to extract features of the video frames in each storyboard segment, and generate storyboard script information corresponding to the storyboard segment based on the extracted information.

[0116] Optionally, the detection module includes:

[0117] A first generating submodule is configured to generate a multimodal content matrix based on feature information of the target video in multiple target dimensions;

[0118] A detection submodule is used to perform storyboard detection based on a storyboard detection model and the multimodal content matrix to obtain a plurality of storyboard segments, wherein the storyboard detection model is determined based on a training video and training annotated storyboards corresponding to the training video, and the training annotated storyboards are annotated based on the multiple target dimensions.

[0119] Optionally, the first determining module includes:

[0120] An acquisition submodule, configured to acquire a target video frame from the target video;

[0121] an identification submodule, configured to perform object recognition on the target video frame and determine candidate objects in the target video;

[0122] A first determining submodule is configured to determine display information of each candidate object in the target video, and determine a subject object and an object object among the candidate objects according to the display information;

[0123] The second determining submodule is configured to determine the feature information corresponding to the subject object as the feature information under the subject content dimension, and determine the feature information corresponding to the object object as the feature information under the object content dimension.

[0124] Optionally, the target dimension also includes at least one of a camera dimension, a scene dimension, a music dimension and an aesthetic dimension.

[0125] Optionally, the processing module includes:

[0126] an extraction submodule, configured to extract, based on the video frames in the storyboard segments, feature information of the storyboard segments under a classification and identification dimension corresponding to each script dimension, wherein the script dimensions include a subject content dimension and an object content dimension, and each script dimension is preset with at least one classification and identification dimension;

[0127] The second generation submodule is used to determine the key video frames in the storyboard segment based on the feature information under the classification identification dimension, and generate the storyboard description text corresponding to the storyboard segment and the object change information between the key video frames. The storyboard script information includes the key video frames corresponding to the storyboard segment, the storyboard description text and the content change information between the key video frames.

[0128] Optionally, the script dimension also includes at least one of a shot dimension, a scene dimension, a sound dimension, a text dimension and an aesthetic dimension, and the classification and identification dimension under the subject content dimension and corresponding to the object content dimension includes at least one of a trait dimension, an action behavior dimension, a trait change dimension and a subject-object interaction dimension.

[0129] Optionally, the device further comprises:

[0130] A second determining module is used to determine the corresponding storyboard display position of each storyboard segment in the target video;

[0131] An associating module is used to associate the storyboard display position corresponding to the storyboard segment with the storyboard script information corresponding to the storyboard segment.

[0132] The present disclosure also provides a video generation device, the device comprising:

[0133] A second receiving module is configured to receive a target text, wherein the target text includes at least one storyboard script;

[0134] A generation module is used to generate a target video corresponding to the target text based on the target text and the video generation model, wherein the video generation model is trained based on training video samples and the annotation scripts corresponding to the training video samples, and the annotation scripts corresponding to the training video samples are generated based on the video feature extraction method described above.

[0135] Reference is now made to FIG3 , which illustrates a schematic diagram of the structure of an electronic device (e.g., a terminal device or server) 600 suitable for implementing embodiments of the present disclosure. The terminal device in the embodiments of the present disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. The electronic device illustrated in FIG3 is merely an example and should not limit the functionality or scope of use of the embodiments of the present disclosure.

[0136] As shown in Figure 3, the electronic device 600 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 601, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 602 or a program loaded from a storage device 608 into a random access memory (RAM) 603. Various programs and data required for the operation of the electronic device 600 are also stored in the RAM 603. The processing device 601, the ROM 602, and the RAM 603 are connected to each other via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.

[0137] Typically, the following devices may be connected to the I / O interface 605: an input device 606 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 607 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 608 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 609. The communication device 609 may allow the electronic device 600 to communicate with other devices wirelessly or by wire to exchange data. Although FIG3 shows the electronic device 600 with various devices, it should be understood that not all of the devices shown are required to be implemented or present. More or fewer devices may alternatively be implemented or present.

[0138] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication device 609, or installed from the storage device 608, or installed from the ROM 602. When the computer program is executed by the processing device 601, the above-mentioned functions defined in the method of the embodiment of the present disclosure are performed.

[0139] It should be noted that the computer-readable medium mentioned above in the present disclosure may be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, device, or component. In the present disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to wires, optical cables, RF (radio frequency), etc., or any suitable combination thereof.

[0140] In some embodiments, the client and server can communicate using any currently known or later developed network protocol, such as HTTP (HyperText Transfer Protocol), and can be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), an internet (e.g., the Internet), and a peer-to-peer network (e.g., an ad hoc peer-to-peer network), as well as any currently known or later developed network.

[0141] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.

[0142] The above-mentioned computer-readable medium carries one or more programs. When the above-mentioned one or more programs are executed by the electronic device, the electronic device is enabled to: receive a target video to be processed; determine feature information of the target video under multiple target dimensions, wherein the target dimensions at least include a subject content dimension and an object content dimension; perform storyboard detection on the target video based on the feature information under the multiple target dimensions to determine the storyboard segments in the target video; for each storyboard segment, perform feature extraction on the video frames in the storyboard segment, and generate storyboard script information corresponding to the storyboard segment based on the extracted information.

[0143] Alternatively, the computer-readable medium carries one or more programs, and when the one or more programs are executed by the electronic device, the electronic device: receives a target text, wherein the target text contains at least one storyboard script; generates a target video corresponding to the target text based on the target text and a video generation model, wherein the video generation model is trained based on training video samples and annotation scripts corresponding to the training video samples, and the annotation scripts corresponding to the training video samples are generated based on the above-mentioned video feature extraction method.

[0144] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages, or a combination thereof, including, but not limited to, object-oriented programming languages ​​such as Java, Smalltalk, C++, and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on the remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).

[0145] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the module, program segment, or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0146] The modules described in the embodiments of the present disclosure may be implemented in software or hardware. In some cases, the name of a module does not necessarily limit the module itself. For example, the first receiving module may also be described as a "module for receiving a target video to be processed."

[0147] The functions described above herein may be performed, at least in part, by one or more hardware logic components. For example, and without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chip (SOCs), complex programmable logic devices (CPLDs), and the like.

[0148] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by an instruction execution system, device or equipment or used in combination with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0149] According to one or more embodiments of the present disclosure, Example 1 provides a video feature extraction method, wherein the method includes:

[0150] Receive the target video to be processed;

[0151] Determining feature information of the target video in multiple target dimensions, wherein the target dimensions include at least a subject content dimension and an object content dimension;

[0152] Performing frame detection on the target video according to the feature information under the multiple target dimensions to determine frame segments in the target video;

[0153] For each of the storyboard segments, feature extraction is performed on the video frames in the storyboard segment, and storyboard script information corresponding to the storyboard segment is generated based on the extracted information.

[0154] According to one or more embodiments of the present disclosure, Example 2 provides the method of Example 1, wherein the performing frame detection on the target video based on the feature information under the multiple target dimensions to determine the frame segments in the target video includes:

[0155] generating a multimodal content matrix based on feature information of the target video in multiple target dimensions;

[0156] Storyboard detection is performed based on a storyboard detection model and the multimodal content matrix to obtain a plurality of storyboard segments, wherein the storyboard detection model is determined based on a training video and training annotated storyboards corresponding to the training video, and the training annotated storyboards are annotated based on the multiple target dimensions.

[0157] According to one or more embodiments of the present disclosure, Example 3 provides the method of Example 1, wherein determining feature information of the target video in multiple target dimensions includes:

[0158] Obtaining a target video frame in the target video;

[0159] Performing object recognition on the target video frame to determine candidate objects in the target video;

[0160] Determining display information of each candidate object in the target video, and determining a subject object and an object object among the candidate objects according to the display information;

[0161] The characteristic information corresponding to the subject object is determined as the characteristic information under the subject content dimension, and the characteristic information corresponding to the object object is determined as the characteristic information under the object content dimension.

[0162] According to one or more embodiments of the present disclosure, Example 4 provides the method of Example 1, wherein the target dimension further includes at least one of a camera dimension, a scene dimension, a music dimension, and an aesthetic dimension.

[0163] According to one or more embodiments of the present disclosure, Example 5 provides the method of Example 1, wherein extracting features from the video frames in the storyboard segments and generating storyboard script information corresponding to the storyboard segments based on the extracted information includes:

[0164] Extracting feature information of the storyboard segments under a classification and identification dimension corresponding to each script dimension based on the video frames in the storyboard segments, wherein the script dimensions include a subject content dimension and an object content dimension, and each script dimension is preset with at least one classification and identification dimension;

[0165] Based on the feature information under the classification and identification dimension, the key video frames in the storyboard segment are determined, and the storyboard description text corresponding to the storyboard segment and the object change information between the key video frames are generated. The storyboard script information includes the key video frames corresponding to the storyboard segment, the storyboard description text and the content change information between the key video frames.

[0166] According to one or more embodiments of the present disclosure, Example 6 provides the method of Example 5, wherein the script dimension also includes at least one of a shot dimension, a scene dimension, a sound dimension, a text dimension, and an aesthetic dimension, and the classification and identification dimension under the subject content dimension and corresponding to the object content dimension includes at least one of a trait dimension, an action behavior dimension, a trait change dimension, and a subject-object interaction dimension.

[0167] According to one or more embodiments of the present disclosure, Example 7 provides the method of Example 1, wherein the method further includes:

[0168] Determining a corresponding storyboard display position of each storyboard segment in the target video;

[0169] The storyboard display position corresponding to the storyboard segment is associated with the storyboard script information corresponding to the storyboard segment.

[0170] According to one or more embodiments of the present disclosure, Example 8 provides a video generation method, the method including:

[0171] receiving a target text, wherein the target text includes at least one storyboard;

[0172] According to the target text and video generation model, a target video corresponding to the target text is generated, wherein the video generation model is trained based on training video samples and annotation scripts corresponding to the training video samples, and the annotation scripts corresponding to the training video samples are generated based on the video feature extraction method described in any one of Examples 1-7.

[0173] According to one or more embodiments of the present disclosure, Example 9 provides a video feature extraction device, the device comprising:

[0174] A first receiving module, configured to receive a target video to be processed;

[0175] A first determining module is configured to determine feature information of the target video in multiple target dimensions, wherein the target dimensions include at least a subject content dimension and an object content dimension;

[0176] A detection module, configured to perform frame detection on the target video based on feature information under the plurality of target dimensions to determine frame segments in the target video;

[0177] The processing module is used to extract features of the video frames in each storyboard segment and generate storyboard script information corresponding to the storyboard segment based on the extracted information.

[0178] According to one or more embodiments of the present disclosure, Example 10 provides a video generating apparatus, the apparatus comprising:

[0179] A second receiving module is configured to receive a target text, wherein the target text includes at least one storyboard script;

[0180] A generation module is used to generate a target video corresponding to the target text based on the target text and the video generation model, wherein the video generation model is trained based on training video samples and annotation scripts corresponding to the training video samples, and the annotation scripts corresponding to the training video samples are generated based on the video feature extraction method described in any one of Examples 1-7.

[0181] According to one or more embodiments of the present disclosure, Example 11 provides a computer-readable medium having a computer program stored thereon, which implements the steps of any one of the methods described in Examples 1-8 when executed by a processing device.

[0182] According to one or more embodiments of the present disclosure, Example 12 provides an electronic device, including:

[0183] a storage device having a computer program stored thereon;

[0184] A processing device is used to execute the computer program in the storage device to implement the steps of the method described in any one of Examples 1-8.

[0185] The above description is merely a preferred embodiment of the present disclosure and an illustration of the technical principles employed. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above-mentioned technical features, but also includes other technical solutions formed by any combination of the above-mentioned technical features or their equivalents without departing from the above-mentioned disclosed concepts. For example, a technical solution formed by replacing the above-mentioned features with (but not limited to) technical features with similar functions disclosed in this disclosure.

[0186] In addition, although each operation is described in a specific order, this should not be understood as requiring these operations to be performed in the specific order shown or in a sequential order. Under certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although some specific implementation details have been included in the above discussion, these should not be interpreted as limiting the scope of the present disclosure. Some features described in the context of a separate embodiment can also be implemented in a single embodiment in combination. On the contrary, the various features described in the context of a single embodiment can also be implemented in multiple embodiments individually or in any suitable sub-combination mode.

[0187] Although the subject matter has been described using language specific to structural features and / or methodological logical acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms of implementing the claims. Regarding the apparatus in the above-described embodiments, the specific manner in which each module performs operations has been described in detail in the embodiments related to the method and will not be elaborated upon here.

Claims

1. A video feature extraction method, comprising: Receive a target video to be processed; Determining feature information of the target video in multiple target dimensions, wherein the target dimensions at least include a subject content dimension and an object content dimension; Performing frame detection on the target video according to the feature information under the multiple target dimensions to determine frame segments in the target video; For each of the storyboard segments, feature extraction is performed on the video frames in the storyboard segment, and storyboard script information corresponding to the storyboard segment is generated based on the extracted information.

2. The method according to claim 1, wherein: The step of performing frame detection on the target video according to the feature information under the multiple target dimensions to determine frame segments in the target video includes: Generate a multimodal content matrix according to feature information of the target video in multiple target dimensions; Storyboard detection is performed based on a storyboard detection model and the multimodal content matrix to obtain a plurality of storyboard segments, wherein the storyboard detection model is determined based on a training video and training annotated storyboards corresponding to the training video, and the training annotated storyboards are annotated based on the multiple target dimensions.

3. The method according to claim 1 or 2, wherein: The determining feature information of the target video in multiple target dimensions includes: Acquire a target video frame in the target video; Performing object recognition on the target video frame to determine candidate objects in the target video; Determine display information of each candidate object in the target video, and determine the subject object and the object object in the candidate objects according to the display information; The characteristic information corresponding to the subject object is determined as the characteristic information under the subject content dimension, and the characteristic information corresponding to the object object is determined as the characteristic information under the object content dimension.

4. The method according to any one of claims 1 to 3, wherein: The target dimension also includes at least one of a camera dimension, a scene dimension, a music dimension, and an aesthetic dimension.

5. The method according to any one of claims 1 to 4, wherein: The extracting features of the video frames in the storyboard segments and generating storyboard script information corresponding to the storyboard segments according to the extracted information includes: According to the video frames in the storyboard segments, feature information of the storyboard segments under the classification identification dimension corresponding to each script dimension is extracted, wherein the script dimension includes a subject content dimension and an object content dimension, and each of the script dimensions There is at least one classification identification dimension preset below; Based on the feature information under the classification identification dimension, the key video frames in the storyboard segment are determined, and the storyboard description text corresponding to the storyboard segment and the object change information between the key video frames are generated. The storyboard script information includes the key video frames corresponding to the storyboard segment, the storyboard description text and the content change information between the key video frames.

6. The method according to claim 5, wherein: The script dimension also includes at least one of a shot dimension, a scene dimension, a sound dimension, a text dimension and an aesthetic dimension, and the classification and identification dimension under the subject content dimension and corresponding to the object content dimension includes at least one of a trait dimension, an action and behavior dimension, a trait change dimension and a subject-object interaction dimension.

7. The method according to any one of claims 1 to 6, wherein: The method further comprises: Determine a storyboard display position corresponding to each storyboard segment in the target video; The storyboard display position corresponding to the storyboard segment is associated with the storyboard script information corresponding to the storyboard segment.

8. A video generation method, comprising: Receiving a target text, wherein the target text includes at least one storyboard script; According to the target text and the video generation model, a target video corresponding to the target text is generated, wherein the video generation model is trained based on training video samples and annotation scripts corresponding to the training video samples, and the annotation scripts corresponding to the training video samples are generated based on the video feature extraction method described in any one of claims 1-7.

9. A video feature extraction device, comprising: A first receiving module is configured to receive a target video to be processed; A first determination module is configured to determine feature information of the target video in multiple target dimensions, wherein the target dimensions at least include a subject content dimension and an object content dimension; A detection module is configured to perform frame detection on the target video according to the feature information under the multiple target dimensions to determine the frame segments in the target video; The processing module is configured to extract features of the video frames in each storyboard segment, and generate storyboard script information corresponding to the storyboard segment based on the extracted information.

10. A video generating device, comprising: A second receiving module is configured to receive a target text, wherein the target text includes at least one storyboard script; A generation module is configured to generate a target video corresponding to the target text according to the target text and the video generation model, wherein the video generation model is trained based on training video samples and the annotation scripts corresponding to the training video samples. The annotation script corresponding to the training video sample is generated based on the video feature extraction method described in any one of claims 1-7.

11. A computer readable medium having a computer program stored thereon, wherein: When the program is executed by a processing device, the steps of the method according to any one of claims 1 to 8 are implemented.

12. An electronic device comprising: a storage device having a computer program stored thereon; A processing device is configured to execute the computer program in the storage device to implement the steps of the method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Video processing method and device, script generation method and device, computer equipment and storage medium

    CN111629230A

  • Advertisement video picture clipping method and system

    CN111815645A

  • Video scripting method

    CN113923521A

  • Video editing method and electronic equipment

    CN115052201A

  • Video processing method and device, electronic equipment and storage medium

    CN116233534A

Cited By

  • Alarm information processing method, device and system, and alarm information auxiliary processing method, device and system

    CN120640083A

  • Mirror operation description evaluation method and device, electronic equipment and storage medium

    CN120766188A

  • Video generation method and device, equipment and storage medium

    CN121126087A

  • Video propagation effect analysis method and device, electronic equipment and storage medium

    CN121418623A