A supervised data generation method based on video semantic structural analysis
Through video semantic structured analysis, the learnability, diversity and uncertainty scores of unlabeled videos are calculated, and valuable samples are selected for annotation. This solves the problems of scarce supervised data and collective outliers in video manuscript generation, and achieves efficient and low-cost supervised data generation.
Patent Information
- Application Number
- CN202411268855.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-11
- Publication Date
- 2025-10-14
- Estimated Expiration
- 2044-09-11
AI Technical Summary
Existing technologies lack high-quality supervised data in video manuscript generation. Traditional data augmentation and pseudo-labeling methods are ineffective, manual labeling is costly, and collective outliers in the human labeling process affect the effectiveness of active learning.
Through a method based on video semantic structured parsing, a large-scale visual-language model is used to generate image descriptions and texts, calculate the learnability, diversity and uncertainty scores of unlabeled videos, select valuable samples for manual annotation, and update the video text generation model.
It effectively reduces the impact of collective outliers, improves the quality of supervised data and model performance, reduces annotation costs, and achieves efficient supervised data generation.
Smart Images

Figure CN119048964B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application belongs to the technical field of computer vision and machine learning, and particularly relates to a supervised data generation method based on video semantic structured analysis. BACKGROUND
[0002] Video manuscript generation, as a cutting-edge technology, aims to deeply analyze video content and automatically convert it into a text-form manuscript or subtitle. It has great potential in AI-assisted content creation, news reporting automation, and media integration. This technology is leading the innovation of information dissemination by significantly improving work efficiency. However, the core challenge of video manuscript generation lies in its high dependence on large-scale and high-quality supervised data, which usually represents the correspondence between videos and corresponding natural language texts. Unfortunately, due to the complexity of video content and the diversity of natural language labels, directly available supervised data is extremely scarce, especially when building a specific domain such as a news writing system, this problem is more prominent.
[0003] Given the complex generation process of video content and the flexibility of natural language labels, traditional data augmentation or synthesis methods are not suitable for this task. Currently, pseudo-label technology provides a possible solution, that is, by predicting the content of unlabelled videos using a model and using it as a training label. However, its effectiveness is limited by the model's ability to understand complex video content, and excessive reliance on such methods may lead to instability in the training process and a decline in model performance. Therefore, high-quality supervised data mainly relies on manual annotation, which is not only time-consuming and labor-intensive but also costly.
[0004] To alleviate this dilemma, active learning strategies have gradually entered people's field of vision. In the case of limited data annotation resources, active learning aims to intelligently select the most valuable samples from the unlabelled data set for manual annotation, in order to achieve the best model performance improvement with the least annotation cost. This strategy has achieved remarkable results in closed tasks such as image classification and object detection, and its core lies in evaluating the uncertainty or diversity of samples to guide sample selection. However, in the open and complex task of video manuscript generation, the application of active learning is still in the exploratory stage.
[0005] It is worth noting that although uncertainty and diversity are widely studied in active learning, the potential problems in the human annotation process, especially the existence of collective outliers, are often overlooked. As a group of data points that deviate significantly from the data set specification or expected behavior, collective outliers have a significant impact in open tasks, including video manuscript generation. These outliers may be caused by inconsistencies among human annotators, specifically in terms of abstraction level and description granularity. This inconsistency not only increases the complexity of the data set, but also may threaten the effectiveness of active learning.
[0006] To address this issue, the present invention proposes a supervised data generation method based on video semantic structured parsing to solve the above problem. Summary of the Invention
[0007] The purpose of the present invention is to provide a supervised data generation method based on video semantic structured parsing to solve the problems raised in the above background technology.
[0008] To achieve the above objectives, the present invention provides the following technical solution: a method for generating supervised data based on video semantic structured parsing, comprising the following steps:
[0009] S1: Initialization, constructing labeled and unlabeled datasets, training the video-to-text generation model, preparing a large-scale visual-language model, and determining the number of active learning selection stages and the budget for each stage;
[0010] S2: Sample frames from the unlabeled video set, generate image descriptions using a large vision-language model, and generate text using a video text generation model;
[0011] S3: Referring to SPICE, it converts textual image descriptions and documents into scene graphs;
[0012] S4: Compute the learnability score of the unlabeled video by performing internal and external metrics on the samples.
[0013] S5: Calculate the diversity score of the unlabeled videos based on the frequency of samples in the entire dataset.
[0014] S6: Calculate the uncertainty score of the unlabeled video based on the reliability of the sample;
[0015] S7: Obtain the overall score of the unlabeled videos, select and label a portion of the unlabeled videos that improve the performance of the video manuscript generation model, integrate them to obtain a new supervised data set, and update the video manuscript generation model;
[0016] S8: Repeat S2 to S7 until the model achieves the desired effect, the annotation budget is exhausted, or the maximum number of active learning selection stages is reached.
[0017] Preferably, the S1 specifically includes the steps of:
[0018] S11: Select existing supervised data to form a label set, expressed as ,Depend on A textual description of video sequences and their human annotations composition;
[0019] S12: Unmarked a video representation is , ;
[0020] S13: Train an initial video script with supervised data L to generate a model ;
[0021] S14: Prepare a large vision-language model BLIP2;
[0022] S15: According to the annotation budget and actual situation, determine the number of stages selected by active learning and the budget selected by each stage . .
[0023] Preferably, the S2 specifically comprises the following steps:
[0024] S21: Uniformly sample frames from , and call the first frame as ;
[0025] S22: Apply BLIP2 to all sampled frames to obtain a set of , where is the image description on ;
[0026] S23: Obtain the script on through the video script generation model , and call the script with the highest confidence as .
[0027] Preferably, the S3 specifically comprises the following steps:
[0028] S31: Given the image description or the script , use or refer to the Stanford Scene Parser to obtain a scene graph composed of objects, attributes and relationships or ;
[0029] wherein and are composed of objects , attributes and relationships ;
[0030] is the set of all scene graphs generated from , , and .
[0031] Preferably, the S4 specifically includes the following steps:
[0032] S41; Pay attention to abstract inconsistencies, calculate internal metrics, and use the scene graph generated by BLIP2 frame by frame Measuring the internal prediction consistency of each video, the specific formula is as follows:
[0033] ;
[0034] in To calculate an object, attribute, or relationship Number of times it appears in , and Capture The number of unique objects and attributes in , intuitively speaking, Formula 1 prefers In other words, The higher, The more consistent the prediction is with respect to other frames. Deliberately ignoring relationships , because this relationship is not very reliable in current large-scale vision-language models (such as BLIP2);
[0035] S42: Pay attention to the granularity inconsistency of human annotations at different granularities, calculate external metric scores, and use unlabeled samples In the active learning selection stage and The absolute distance between the SPICEs on , the specific formula is as follows:
[0036] ;
[0037] in Yes all The minimum and maximum normalization function on samples.
[0038] Preferably, the S5 specifically includes the following steps:
[0039] S51: Calculate the diversity score by re-weighting TF for longer descriptions. Indicated as containing The collection of scene graphs with the highest confidence predictions , the image description of BLIP2 prediction for the full dataset is represented as , ,in Defined as:
[0040] ,in is a binary indicator function.
[0041] Preferably, the S6 specifically includes the following steps:
[0042] S61: Calculate the uncertainty score. The specific formula is as follows:
[0043] ;
[0044] in, Generate scene graphs for video document predictions and scene graphs for image descriptions predicted by large vision-language models Shared objects, properties, and relationships, where Used as human annotation, To measure the highest confidence.
[0045] Preferably, the step S7 specifically includes the following steps:
[0046] S71: Calculate Sample Overall score on , The calculation formula is as follows:
[0047] ;
[0048] in 、 、 is a hyperparameter;
[0049] S72: In the active learning algorithm Stage, based on overall score Will The unlabeled videos are ranked, with the lower Indicates a greater amount of information, based on the budget selected at this stage choose The lowest batch of unlabeled samples ;
[0050] S73: Feed samples to humans for annotation ;
[0051] S74: Use Update labeled and unlabeled sets;
[0052] S75: Use the newer Retrain the video transcript generation model .
[0053] Preferably, SPICE is a graph-based semantic representation method used to evaluate the quality of machine-generated image descriptions in image annotation tasks.
[0054] Compared with the prior art, the present application has the following advantages:
[0055] (1) The present application
[0056] (2) The present application BRIEF DESCRIPTION OF DRAWINGS
[0057] Figure 1 Flowchart of the pool-based active learning process of the present application;
[0058] Figure 2 Flowchart of the scene graph parser used in the present application;
[0059] Figure 3 Schematic diagram of the active learning process performed in the present application;
[0060] Figure 4 Effect diagram of the present application on the MSVD public dataset;
[0061] Figure 5 Flowchart of the present application. DETAILED DESCRIPTION
[0062] In order to make the personnel in the art better understand the technical solutions in the embodiments of the present application, the technical solutions of the present application will be described clearly and completely below in conjunction with the drawings. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. It should be understood that these descriptions are only exemplary and are not intended to limit the scope of the present application. Based on the embodiments of the present application, all other embodiments obtained by those of ordinary skill in the art without creative labor should fall within the scope of the present application.
[0063] In addition, in the following description, the description of well-known structures and techniques is omitted to avoid unnecessary confusion of the concepts disclosed in the present application.
[0064] The exemplary embodiments will be described in detail below with reference to the accompanying drawings. In the following description, the same numbers refer to the same or similar elements unless otherwise indicated. The implementations described in the following exemplary embodiments are not meant to represent all implementations consistent with the present application. Rather, they are merely examples with respect to methods and systems consistent with some aspects of the present application as detailed in the appended claims.
[0065] Referring to Figures 1-5 As shown in the drawings, the present application provides the following technical solutions:
[0066] Embodiment One
[0067] A supervised data generation method based on video semantic structured analysis, characterized in that it comprises the following steps:
[0068] S1: initialization, constructing a labeled set and an unlabeled set, training a video script generation model, preparing a large visual-linguistic model, determining the number of active learning selection stages and the budget selected in each stage;
[0069] S2: sampling frames from the videos in the unlabeled set, generating image descriptions through the large visual-linguistic model, and generating scripts using the video script generation model;
[0070] S3: converting the image descriptions and scripts in text form into scene graphs, SPICE being a graph-based semantic representation method for evaluating the quality of machine-generated image descriptions in image annotation tasks;
[0071] S4: calculating the learnability score of the unlabeled videos by performing internal and external metrics on the samples.
[0072] S5: calculating the diversity score of the unlabeled videos based on the frequency of the samples in the entire dataset.
[0073] S6: calculating the uncertainty score of the unlabeled videos based on the reliability of the samples;
[0074] S7: obtaining the overall score of the unlabeled videos, selecting a part of the unlabeled videos that can improve the effect of the video script generation model and labeling them, integrating a new supervised data set, and updating the video script generation model;
[0075] S8: repeating S2-S7 until the model reaches the desired effect, the annotation budget is exhausted, or the maximum number of active learning selection stages is reached.
[0076] In addition, in the present application, regarding the above S1, S1 specifically comprises the following steps:
[0077] S11: selecting existing supervised data to form the labeled set, represented as , which is composed of video sequences and their human-annotated script descriptions ;
[0078] S12: representing the unlabeled videos as ; ;
[0079] S13: training an initial video script generation model using the supervised data L;
[0080] S14: Prepare large vision-language model BLIP2;
[0081] S15: According to the annotation budget and the actual situation, determine the number of stages selected by active learning and the budget selected by each stage .
[0082] In addition, in the present application, regarding the above S2, S2 specifically comprises the following steps:
[0083] S21: Uniformly sample frames from , and call the first frame as ;
[0084] S22: Apply BLIP2 to all sampled frames to obtain a set , wherein is the image description on ;
[0085] S23: Obtain the script on through the video script generation model , and express the script with the highest confidence as .
[0086] In addition, in the present application, regarding the above S3, S3 specifically comprises the following steps:
[0087] S31: Given the image description or the script , use or refer to the Stanford scene parser to obtain the scene graph composed of objects, attributes and relationships or ;
[0088] wherein, and are composed of objects , attributes and relationships ;
[0089] is the set of all scene graphs generated from , , and .
[0090] In addition, in the present application, regarding the above S4, S4 specifically comprises the following steps:
[0091] S41: Pay attention to abstract inconsistency, calculate internal metrics, and use scene graphs generated by frame-by-frame BLIP2 The intra-prediction consistency of each video is measured, specifically as follows:
[0092]
[0093] wherein is the number of times the object, attribute or relation appears in the computation , and are the number of unique objects and attributes captured in the respectively, intuitively, formula 1 favors cases where the opinions of are consistent. In other words, the higher is, the more consistent the prediction is with respect to other frames. In addition, the relation is intentionally ignored because this relation is not very reliable in current large-scale vision-language models (e.g., BLIP2);
[0094] S42: Pay attention to the granularity inconsistency when human annotations are at different granularities, and calculate the external metric score. The image description predicted by the large-scale vision-language model based on images and the script generated by the video script generation model tend to maximize the capture of different aspects of , thus approximating the potential granularity inconsistency between and . In particular, granularity inconsistency can be measured by the SPICE score between the two predictions, where a lower score indicates higher inconsistency. Simply relying on SPICE gains may not be high, so attention is paid to the time-varying changes in the SPICE of each to simulate the expected changes in granularity consistency. The absolute distance between the SPICE of the unlabeled sample on the active learning selection stage and is denoted as , specifically as follows:
[0095]
[0096] wherein is the min-max normalization function on all samples. The designed prefers samples with moderate changes, and large changes and small changes may be related to samples with less information and difficult to learn, respectively.
[0097] In addition, in the present application, regarding the above S5, S5 specifically comprises the following steps:
[0098] S51: Calculate diversity score. Diversity is preferred where the content is beyond the current video transcript generation model. Specifically, we use the concept of Term Frequency-Inverse Document Frequency (TF-IDF) to measure the importance of video content. Longer descriptions tend to co-occur more with diverse videos, so we re-weight TF for longer descriptions. Indicated as containing The collection of scene graphs with the highest confidence predictions Similarly, the image description of BLIP2 predictions for the full dataset is represented as Mathematically, ,in Defined as:
[0099]
[0100] in is a binary indicator function if If valid is equal to 1. Mathematically speaking, It only focuses on the content outside the current video manuscript generation model. At the same time, it values unique samples with important content. Overall, higher express The diversity is greater.
[0101] In addition, in the present invention, regarding the above S6, S6 specifically includes the following steps:
[0102] S61: Calculate the uncertainty score. The specific formula is as follows:
[0103] ;
[0104] Specifically, Compute the scene graph of the document predicted by the video document generation model and scene graphs for image descriptions predicted by large vision-language models Shared objects, properties, and relationships, where Used as human annotation to measure the highest confidence How good is the prediction of reflects More certainty.
[0105] In addition, in the present invention, regarding the above S7, S7 specifically includes the following steps:
[0106] S71: Calculate Sample Overall score on , The calculation formula is as follows:
[0107] ;
[0108] wherein , , is a hyperparameter;
[0109] S72: In the first stage of the active learning algorithm, the unannotated videos in the set are sorted according to the overall score , wherein a lower indicates a larger amount of information, and a budget selected in this stage is used to select the batch of unannotated samples with the lowest ; ;
[0110] S73: The samples are fed to a human for annotation ;
[0111] S74: The annotated and unannotated sets are updated ;
[0112] S75: The video script generation model is retrained using the updated .
[0113] And, in order to assist personnel in understanding the present application, the following embodiments are additionally provided:
[0114] Embodiment Two
[0115] As shown in Figure 1 , pool-based active learning is a classic paradigm in active learning, which assumes a large pool of unannotated data and a small set of annotated data. The initial model is trained on this small set of annotated data. This data set can only provide limited information, resulting in a model that may not be accurate enough when facing unannotated data, but it will gradually become complex and accurate as iterations proceed. The algorithm will actively select a number of most valuable samples from the unannotated data pool for annotation according to the active learning strategy each time. The selected samples are sent to human annotators for annotation. The annotation results will be added to the annotated data set. The newly annotated data is combined with the previously annotated data to retrain the model. This step allows the model to gradually learn the features of more samples, thereby improving overall performance. Considering the limited annotation budget and expensive annotation process, the general active learning method does not spend all the budget at once, but rather like Figure 1 The illustrated process of selecting and annotating a portion of the budget each time, continuously cycling, gradually improves the model's effectiveness by selecting and annotating more valuable unlabeled samples step by step, until the entire budget is used up.
[0116] As Figure 2 illustrated, the workflow of the rule-based scene parser used is as follows:
[0117] A1: After giving the manuscript to be parsed, use the probabilistic context-free grammar (PCFG) dependency parser to convert it into a dependency parse tree.
[0118] A2: According to the parsed dependency tree, obtain a semantic graph by simplifying quantitative modifiers and parsing pronouns.
[0119] A3: Parse the semantic graph through language rules to extract lemmatized objects, relationships, and attributes to form a scene graph.
[0120] As Figure 3 illustrated, the specific description of the active learning process performed by the present application is as follows:
[0121] Due to the lack of access to human annotations in active learning, it is impossible to directly identify collective outliers. The present application turns to the main attributes or inconsistencies of collective outliers as a breakthrough, and estimates the abstract and granularity inconsistencies in unlabeled data. Specifically, by using existing large-scale visual-linguistic models based on images (compared to video-based models, the current technology is more mature), the video frames of the unlabeled video are estimated to obtain image descriptions of each frame with consistent granularity to approximate human annotations. By converting the description results into semantic scene graphs, different description results can be compared, which is important because learnability, uncertainty, and diversity are all measured based on scene graphs. The present application not only measures the consistency between the description results of each frame internally according to the image description results of the large-scale visual-linguistic model on the video frames, but also measures the consistency between the manuscript generated by the video manuscript generation model and the description results of each frame externally. Internal measurement captures abstract consistency, and external measurement relies on granularity consistency. Combine the two to measure the learnability of the sample, or the likelihood that this sample belongs to the collective outliers. The present application also measures the diversity and uncertainty of the sample according to its frequency in the entire dataset (find samples that have not been learned and are far from other samples through TF-IDF technology) and reliability (the consistency between the scene graph of the manuscript predicted by the video manuscript generation model and the scene graph of the image description predicted by the large-scale visual-linguistic model). Finally, according to the overall score of the weighted fusion, the samples are sorted, the most valuable samples are selected and annotated according to the sorting, and then combined with the annotated data and update the model. The present application can select reliable, diverse, but learnable samples, and achieve a good trade-off between accuracy and manpower.
[0122] To test the effect of the method of the present application, comparative experiments were conducted on the public data set MSVD and two models SwinBERT and CoCap, and the evaluation indexes included BLEU4, METEOR, ROUGE_L, CIDEr and SPICE. In order to ensure more reliable results, all the compared active learning methods were run three times. The experimental results are shown in FIG. 1, which reports the average performance and the difference, and there are several interesting observations. First of all, it is worth noting that the method of the present application is always superior to all other baselines in all five evaluation indexes of the two models, highlighting the advantage of the method of the present application. It can be foreseen that random sampling is usually the second best choice, consistent with the findings of other open-ended tasks. More importantly, the method of the present application uses less than 25% of the labeled samples to surpass the results of full training using 100% labeled data with SwinBERT as shown by the horizontal dotted line in 4 evaluation indexes. For example, when measured by BLEU4 and ROUGE-L respectively, the method of the present application only uses about 10% and 15% of the labeled data to achieve full effect. For CIDEr, the method of the present application achieves 103.65% of the full effect with 25% of the labeled samples. Even with CoCap, the method of the present application can almost always achieve more than 95% of the full performance with 25% of the labeled samples. This observation supports the concern of the method of the present application about the negative impact of collective outliers on the training process. Obviously, the method of the present application effectively reduces the impact of collective outliers, resulting in significant improvement in the effect of the supervised data generation method. Figure 4
[0123] In addition, in the present application, the supervised data generation method mainly follows the pool-based active learning paradigm, and the video semantics are structurally analyzed, and the unlabeled videos are labeled according to the results, so as to provide rich and high-quality supervised data for the video manuscript generation task, and promote the video manuscript generation method to achieve better effect. Video semantic structural analysis involves image-based Large Vision-Language Models (LVLM) and scene graphs. By reasoning the video through the image-based Large Vision-Language Models and generating the scene graph based on the semantics, the video can be analyzed and distinguished from the learnability, uncertainty and diversity of the scene graph.
[0124] Embodiment Three
[0125] In this specific embodiment, the method of the present application is applied to generate a supervised dataset for sports event videos (e.g., soccer matches). First, a portion of videos with detailed match commentary scripts is selected from an existing sports event video library as the initial annotation set. Then, videos are sampled from the unlabeled sports event video library, and the BLIP2 visual-linguistic model is used to generate image descriptions for each frame, and a preliminary match commentary script is generated by the video script generation model. By calculating internal metrics (e.g., consistency between frames within a video) and external metrics (e.g., SPICE score changes between image descriptions and scripts), as well as diversity and uncertainty scores, the most informative videos are selected for manual annotation. During annotation, focus is placed on the description of key match events (e.g., goals, fouls) to improve the quality and diversity of the dataset. As the model is iteratively trained, the quality of the generated scripts gradually improves, ultimately forming a high-quality supervised dataset for sports event videos.
[0126] Embodiment Four
[0127] For the generation of supervised data for online education videos (e.g., course lectures, experiment demonstrations), the method of the present application is also applicable. First, videos with detailed teaching notes are collected from education platforms as the annotation set. Then, videos are sampled from the unlabeled video library, and the visual-linguistic model is used to describe the video frames, and a preliminary teaching script is generated by the video script generation model. When calculating learnability, diversity, and uncertainty scores, special attention is paid to the coherence and completeness of the teaching content. Videos that can cover a wide range of knowledge points and have a large difference between predicted scripts and image descriptions are selected for manual annotation. During annotation, focus is placed on the accuracy of knowledge points and the clarity of teaching logic to improve the teaching value of the dataset. Through multiple iterations, a supervised dataset suitable for education video understanding is generated.
[0128] Embodiment Five
[0129] In the field of news video reporting, the method of the present application can be used to construct a supervised dataset containing key information such as news events, locations, and characters. First, videos with detailed news report scripts are obtained from news websites or television stations as the annotation set. Then, videos are sampled from the unlabeled news video library, and the visual-linguistic model is used to generate image descriptions for video frames, and a preliminary news report script is generated by the video script generation model. When calculating various indicators, special attention is paid to the accuracy and timeliness of news events. Videos containing important news events and having a high degree of consistency between predicted scripts and image descriptions are selected for manual annotation. During annotation, focus is placed on the completeness and accuracy of news elements to improve the news value of the dataset. Through multiple iterations, a high-quality supervised dataset for news video reporting is constructed, which helps to improve the effectiveness of tasks such as automatic summarization and classification of news videos.
[0130] Embodiment Six
[0131] In the field of nature documentaries, the method of the present application can help generate a supervised dataset containing rich biodiversity and ecological information. First, select videos containing detailed commentary from existing nature documentary libraries as the initial annotation set, which usually describes the habits of animals and plants, ecological environment, etc. Then, sample from the unannotated nature documentary video library, use large visual-linguistic models such as BLIP2 to generate image descriptions for video frames, and generate preliminary commentary scripts through video script generation models. When calculating learnability, diversity and uncertainty scores, pay special attention to the accuracy of biological species identification, the completeness of ecological environment description, and the consistency between predicted scripts and image descriptions. Select videos containing rare species, unique ecological environments, and high uncertainty in predicted results for manual annotation to ensure the uniqueness and richness of the dataset. Through multiple iterations, a high-quality nature documentary video supervised dataset is generated, which helps to promote the progress of biodiversity research and ecological protection work.
[0132] Embodiment seven
[0133] In the advertising creative industry, the method of the present application can be applied to build a supervised dataset containing various advertising creative elements. First, select advertising videos containing detailed creative scripts and visual elements from the advertising database as the annotation set. Then, sample from the unissued advertising video library, use visual-linguistic models to generate image descriptions for advertising video frames, and generate preliminary creative scripts through video script generation models. When calculating each indicator, pay attention to the novelty of advertising creativity, the targeting of target audience, and the relevance between predicted scripts and image descriptions. Select those advertising videos with unique creativity, clear target audience, and high uncertainty in predicted results for manual annotation to enrich the creative diversity and practicality of the dataset. Through multiple iterations, a high-quality advertising video creative supervised dataset is generated, providing strong support for the automated generation and optimization of advertising creativity.
[0134] Embodiment eight
[0135] In the field of movie promotion and marketing, the method of the present application can be used for data enhancement and annotation optimization of movie trailers. First, trailers with detailed plot summaries and character introductions are collected from movie databases as the annotation set. Then, for the movies to be released or in production, samples are taken from the unreleased trailer materials, the visual-linguistic model is used to describe the video frames, and the preliminary trailer script is generated through the video script generation model. When calculating various indicators, special attention is paid to the coherence of the plot, the shaping of the characters, and the matching degree of the trailer script and the visual content. Trailer materials that can attract audience attention, show the core selling points of the movie, and have high diversity in prediction results are selected for manual annotation to enhance the attractiveness and representativeness of the data set. Through multiple iterations, a high-quality movie trailer supervised data set is generated, providing more accurate and effective data support for movie promotion and marketing.
[0136] In addition, to achieve the above-mentioned purpose, an embodiment of the present application also provides a computer readable storage medium, which stores a computer program. The computer program is executed by a processor to realize the method for generating supervised data based on video semantic structured analysis.
[0137] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the embodiments of the present application can be in the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the embodiments of the present application can be in the form of a computer program product implemented on one or more computer usable vehicles (including but not limited to disk memory, CD-ROM, optical memory, etc.) containing computer usable program codes.
[0138] The embodiments of the present application are described with reference to the flowcharts and / or block diagrams according to the method, terminal device (system), and computer program product of the embodiments of the present application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of the flows and / or blocks in the flowcharts and / or block diagrams can be realized by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing terminal device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing terminal device realize the functions specified in the flowcharts and / or block diagrams. Figure 1 The system that realizes the functions specified in one flow or multiple flows and / or blocks Figure 1 The system that realizes the functions specified in one block or multiple blocks
[0139] These computer program instructions can also be stored in a computer- readable memory that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer-readable memory produce an article of manufacture including instructions which implement the Figure 1 function specified in the flow or flows and / or blocks Figure 1 of the block or blocks.
[0140] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the Figure 1 function specified in the flow or flows and / or blocks Figure 1 of the block or blocks.
[0141] Finally, it should be noted that, in the description above, relational terms such as first and second, and the like can be used solely to distinguish one entity or action from another entity or action without necessarily requiring or implying any actual such relationship or order between such entities or actions. The terms "and / or" and / or "and / or" mean any one of the items, optionally including combinations of items. In addition, the term "comprising" is intended to mean including, but not limited to, to the extent permitted by applicable law. Unless otherwise required by context, the use herein or the use of a word such as "comprises" or "comprising" or "includes" or "including" means "including but not limited to".
[0142] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the embodiments of the present application, and not to limit them. Although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or make equivalent replacements to some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application. Any changes or replacements that can be easily thought of by those skilled in the art within the technical scope disclosed by the present application should be covered within the protection scope of the present application.
[0143] While embodiments of the application have been shown and described, it is to be understood that the embodiments described are merely exemplary of the principles and application of the present application. Numerous modifications and adaptions can be effected without departing from the spirit and scope of the present application, which is not limited to the exact construction and arrangement described. It is intended, therefore, to cover all modifications and adaptions that fall within the scope of the claims and their equivalents.
Claims
1. A supervised data generation method based on video semantic structured parsing, characterized by: The steps include: S1: Initialization, constructing labeled and unlabeled datasets, training the video-to-text generation model, preparing a large-scale visual-language model, and determining the number of active learning selection stages and the budget for each stage; S2: Sample frames from the unlabeled video set, generate image descriptions using a large vision-language model, and generate text using a video text generation model; S3: Referring to SPICE, it converts textual image descriptions and documents into scene graphs; S4: Calculate the learnability score of the unlabeled video by performing internal and external metrics on the sample; S5: Calculate the diversity score of the unlabeled video based on the frequency of the sample in the entire dataset; S6: Calculate the uncertainty score of the unlabeled video based on the reliability of the sample; S7: Obtain the overall score of the unlabeled videos, select and label a portion of the unlabeled videos that improve the performance of the video manuscript generation model, integrate them to obtain a new supervised data set, and update the video manuscript generation model; S8: Repeat S2 to S7 until the model achieves the desired effect, the annotation budget is exhausted, or the maximum number of active learning selection stages is reached.
2. The method for generating supervised data based on video semantic structured analysis according to claim 1, characterized in that: Said S1 specifically comprises the steps of: S11: Select existing supervised data to form a label set, expressed as ,Depend on A textual description of video sequences and their human annotations composition; S12: Unmarked Videos are represented as , ; S13: Use supervised data L to train an initial video manuscript and generate a model ; S14: Prepare the large-scale vision-language model BLIP2; S15: Determine the number of stages for active learning based on the annotation budget and actual situation and the budget selected at each stage .
3. The method for generating supervised data based on video semantic structured analysis according to claim 1, characterized in that: The S2 specifically includes the following steps: S21: From uniformly sampled frame, and the Frame is called ; S22: Apply BLIP2 to all sample frames to obtain a set of ,in yes Image description on; S23: Generate Models from Video Documents get The manuscripts on the list are sorted and the manuscript with the highest confidence is denoted as .
4. The method for generating supervised data based on video semantic structured analysis according to claim 1, characterized in that: The S3 specifically includes the following steps: S31: Given an image description or manuscript , use or refer to the Stanford scene graph parser to obtain a scene graph consisting of objects, attributes and relationships or ; in, and By object ,property and relationships composition; It is from The set of all scene graphs generated, , and .
5. The method for generating supervised data based on video semantic structured analysis according to claim 1, characterized in that: The S4 specifically includes the following steps: S41; Pay attention to abstract inconsistencies, calculate internal metrics, and use the scene graph generated by BLIP2 frame by frame Measuring the internal prediction consistency of each video, the specific formula is as follows: ; in To calculate an object, attribute, or relationship Number of times it appears in , and Capture the number of unique objects and properties in ; S42: Pay attention to the granularity inconsistency of human annotations at different granularities, calculate external metric scores, and use unlabeled samples In the active learning selection stage and The absolute distance between the SPICEs on , the specific formula is as follows: ; in Yes all The minimum and maximum normalization function on samples.
6. The method for generating supervised data based on video semantic structured analysis according to claim 1, characterized in that: The S5 specifically includes the following steps: S51: Calculate the diversity score by re-weighting TF for longer descriptions. Indicated as containing The set of scene graphs with the highest confidence predictions , the image description of BLIP2 prediction for the full dataset is represented as , ,in Defined as: ,in is a binary indicator function.
7. The method for generating supervised data based on video semantic structured analysis according to claim 1, characterized in that: The S6 specifically includes the following steps: S61: Calculate the uncertainty score. The specific formula is as follows: ; in, Generate scene graphs for video document predictions and scene graphs for image descriptions predicted by large vision-language models Shared objects, properties, and relationships, where Used as human annotation, To measure the highest confidence.
8. The method for generating supervised data based on video semantic structured analysis according to claim 1, characterized in that: The S7 specifically includes the following steps: S71: Calculation Sample Overall score on , The calculation formula is as follows: ; in 、 、 is a hyperparameter; S72: In the active learning algorithm Stage, based on overall score Will The unlabeled videos are ranked, with the lower Indicates a greater amount of information, based on the budget selected at this stage choose The lowest batch of unlabeled samples ; S73: Feed samples to humans for annotation ; S74: Use Update labeled and unlabeled sets; S75: Use the newer Retrain the video transcript generation model .
9. The method for generating supervised data based on video semantic structured analysis according to claim 1, characterized in that: SPICE is a graph-based semantic representation method used to evaluate the quality of machine-generated image descriptions in image annotation tasks.
Citation Information
Patent Citations
Optical character recognition method in patent text scene
CN110674777A
Value chain network control tower and enterprise management platform
CN115699050A