De-duplication method for video material

By calculating the multi-level semantic tag matching degree and historical usage data of video clips, a weighted algorithm is used to select target clips, solving the problem of duplicate content in short video generation, realizing intelligent deduplication, and improving production efficiency and content uniqueness.

CN121722941AActive Publication Date: 2026-03-24LAINENG (HANGZHOU) E-COMMERCE CO LTD
View PDF 8 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-02-26
Publication Date
2026-03-24

AI Technical Summary

Technical Problem

Existing technologies rely on manual review after the video is generated to remove duplicate content, which is inefficient and produces inconsistent results, failing to meet the needs of large-scale production.

Method used

By acquiring multi-level semantic tag matching degree and historical usage data of script fragments and candidate video material fragments, tag matching degree, diversity score and cold start score are calculated, and a weighted algorithm is used to select target video material fragments to achieve intelligent deduplication.

Benefits of technology

By automatically avoiding duplicate content during the material matching stage, the automation level and content uniqueness of video production are improved, replacing inefficient manual review.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121722941A_ABST
    Figure CN121722941A_ABST
Patent Text Reader

Abstract

The invention relates to a video material deduplication method, which establishes an accurate range for subsequent processing by obtaining a script fragment and a semantic matched candidate video material fragment. The label matching degree is calculated based on the matching condition of the candidate video material segments and the scripts in the multi-level semantic label system, and it is ensured that the materials meet the script requirements in the subject and content categories. De-duplication related scores are calculated by analyzing the use history of candidate video material segments, the probability of repeatedly using the same or similar materials is effectively reduced through diversity scores, new materials are actively encouraged to be used through cold start scores, and the repetition risk and novelty of the materials are quantitatively evaluated from different dimensions. Finally, the target material is selected according to the tag matching degree and the weighting result of the deduplication correlation score, the content matching requirement and the deduplication target are organically combined, the repeated content is automatically avoided from the source in the material matching stage, a post-event manual auditing mode which is low in efficiency and high in subjectivity is replaced, and the intelligent level of video production is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of video content processing technology, and in particular to a method for deduplicating video footage. Background Technology

[0002] With the widespread adoption of short video applications, efficiently and automatically generating massive amounts of unique short videos has become a key challenge in content production. This is especially true in e-commerce and online education, where it's often necessary to generate short videos with diverse themes from the same batch of long video footage using different scripts. Automated video generation technologies, such as semantic tag matching of scripts and video clips, have significantly improved production efficiency. However, a new problem arises: within a limited video footage library, highly automated matching algorithms may repeatedly select the same or a few similar "high-quality" video clips, resulting in a large amount of visually repetitive content in the short videos generated for different scripts.

[0003] Currently, the main approach to addressing this content duplication issue is to rely on manual review and screening after the short videos are generated. Specifically, reviewers need to watch each generated short video, using visual inspection and subjective judgment to identify and replace or delete videos with excessively high levels of content duplication. While this post-processing method can alleviate the problem to some extent, its inherent flaws are quite obvious. First, it is extremely inefficient, heavily reliant on manpower, and cannot meet the needs of large-scale, batch production, significantly diminishing the efficiency advantages of automated generation. Second, the deduplication effect is difficult to guarantee with high stability and objectivity, relying entirely on the experience and focus of the reviewers. Different personnel, or even the same person at different times, may have different judgment standards, leading to inconsistent deduplication criteria and ultimately affecting the overall quality of the video output.

[0004] Therefore, there is an urgent need in the existing technology to provide a method that can be integrated into the automatic video generation process, which can proactively and intelligently avoid duplicate content in the material matching and selection stage, reduce the duplication of generated videos from the source, and overcome the inefficiency and unstable results caused by relying on post-event manual review. Summary of the Invention

[0005] Based on this, this application addresses the aforementioned shortcomings of traditional deduplication methods by providing a method for deduplicating video footage.

[0006] This application provides a method for deduplicating video footage, including: Retrieve script fragments and multiple candidate video clips that match their semantics; Based on the matching of the candidate video clips and script clips in the preset multi-level semantic tagging system, the tag matching degree is calculated; Based on the usage history of the candidate video clips, a deduplication-related score is calculated; the deduplication-related score includes at least a diversity score for reducing the repeated use of the same video clip or the use of highly similar video clips, and a cold start score for encouraging the use of new video clips. Based on the weighted result of the tag matching degree and deduplication correlation score, one or more target video clips are selected from the plurality of candidate video clips.

[0007] This application relates to a method for deduplicating video footage. The first step, acquiring script segments and multiple candidate video footage segments that semantically match them, defines a clear input range and processing object for subsequent intelligent deduplication. Then, the tag matching degree is calculated based on the matching between candidate video footage segments and script segments within a pre-defined multi-level semantic tagging system. This step ensures that the selected footage meets the basic requirements of the script in terms of theme and content, laying the foundation for quality. Next, a deduplication-related score is calculated based on the usage history of the candidate video footage segments. This score specifically includes a diversity score to reduce the repeated use of the same video footage segment or the use of highly similar video footage segments, and a cold start score to encourage the use of new video footage segments. These two scores quantify the risk of content reuse and novelty from different dimensions, directly intervening in the problem of content duplication. Finally, target video clips are selected from the candidate pool based on the weighted results of tag matching and deduplication-related scores. This step effectively integrates and balances content matching requirements with deduplication goals, enabling automatic and intelligent avoidance of duplicate content from the source during the material matching and selection stage. This fundamentally replaces the inefficient and subjective post-event manual review method, improving the automation level and content uniqueness of video production. Attached Figure Description

[0008] Figure 1 This is a flowchart illustrating a method for deduplicating video footage according to an embodiment of this application. Detailed Implementation

[0009] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0010] This application provides a method for deduplicating video footage.

[0011] Furthermore, the video material deduplication method provided in this application does not limit the executing entity. Optionally, the executing entity of the video material deduplication method provided in this application can be a video material deduplication terminal. Specifically, the executing entity of the video material deduplication method provided in this application can be one or more processors in the video material deduplication terminal.

[0012] like Figure 1 As shown, in one embodiment of this application, the method for deduplicating video materials includes: S100: Obtain the script fragment and multiple candidate video clips that match its semantics.

[0013] S200, calculate the tag matching degree based on the matching of the candidate video material fragments and script fragments in the preset multi-level semantic tag system.

[0014] S300, based on the usage history of the candidate video clips, calculate a deduplication relevance score. The deduplication relevance score includes at least a diversity score and a cold start score. The diversity score reduces the likelihood of repeatedly using the same video clip or using highly similar video clips. The cold start score encourages the use of new video clips.

[0015] S400, based on the weighted result of the tag matching degree and deduplication correlation score, select one or more target video clips from the plurality of candidate video clips.

[0016] Specifically, short videos need to be generated based on multiple script segments. Therefore, it is necessary to find one or more target video clips that match each script segment, and then aggregate these one or more target video clips to generate a complete short video. S100 to S400 in this embodiment describe the method for obtaining one or more target video clips corresponding to each script segment. It can be understood that S100 to S400 in this embodiment need to be executed for each script segment to obtain one or more target video clips corresponding to each script segment.

[0017] In S100, based on a multi-level semantic tag matching mechanism, multiple candidate video clips semantically related to the script segment are obtained. For example, for the script segment "This lipstick has a moisturizing texture and doesn't dry out the skin," the system will retrieve all candidate video clips with relevant multi-level tags such as "makeup-lipstick-texture display" from the video clip library, forming a candidate set C={v1, v2, ..., v...} m}, where m is the total number of candidate video clips that are semantically related to the current script clip.

[0018] The following describes the multi-level semantic tag matching mechanism in S100. Optionally, a semantic tag set is generated for each video clip in the video material library. For example, for the "arm swatch demonstration" clip, its semantic tag set may contain the following semantic tags: L1: Makeup, L2: Product Showcase, L3: Swatch. These semantic tags describe the same video clip from different dimensions such as product type, content function, and specific actions, and form a hierarchical relationship, such as Makeup > Product Showcase > Swatch, which is refined and specified layer by layer.

[0019] Scripts are generally complex in structure and contain a lot of content. For the sake of brevity, a simple example is used here. For example, the script is: "This lipstick has a silky smooth texture and doesn't dry out the skin. Our classic true red is very flattering." The script file format has been pre-labeled with multi-level semantic tags, meaning that the script file format has been pre-attached with a set of semantic tags consisting of multiple semantic tags. The script can be split into multiple script fragments, such as "This lipstick has a silky smooth texture and doesn't dry out the skin. Our classic true red is very flattering." The script can be split into two script fragments: "This lipstick has a silky smooth texture and doesn't dry out the skin" and "Our classic true red is very flattering." The semantic tag set of the script fragments has a isomorphic relationship with the semantic tag set of the video material fragments, that is, the generation principle is the same and the logic is the same. Otherwise, they would not be able to understand each other's semantic tags in the S100 semantic matching step. The process of finding multiple candidate video material fragments that semantically match the script fragments in S100 can be implemented in various ways. Specifically, it can be implemented using an LLM model (Large Language Model). The implementation method is not restricted and is not the focus of this application.

[0020] Next, the tag matching score (TagScore) is calculated using Formula 1.

[0021] TagScore=Tag1-Score×k1+secondaryTag×Tag2-Score×k2+...+Tagn-Score×kn Formula 1.

[0022] Here, TagScore represents the tag matching degree. Tag1-Score is the first-level tag matching score; if the first-level tags of the candidate video clip and the script clip match perfectly, the Tag1-Score is 1, otherwise it is 0. Tag2-Score is the second-level tag matching score; if the second-level tags of the candidate video clip and the script clip match perfectly, the Tag2-Score is 1, otherwise it is 0. Tagn-Score is the n-level tag matching score; if the n-level tags of the candidate video clip and the script clip match perfectly, the Tagn-Score is 1, otherwise it is 0. Since the candidate video clip is the main object being compared, n here represents the total number of tag levels within the candidate video clip. k1 is the weight of the first-level tag, k2 is the weight of the second-level tag, kn is the weight of the n-level tag, and k1+k2+...kn=1.

[0023] Optionally, a tree-like expansion can be set from the first-level tag to the n-level tag. That is, the first-level tag has the broadest semantic content, and the semantic content scope narrows progressively downwards. For example, the first-level tag might be the product category (beauty), the second-level tag might be the product name (lipstick), and the third-level tag might be the product texture. Since the first-level tag covers the theme of the material, it can be given the highest weight, for example, k1=0.5, to ensure that the subsequently selected candidate video clips (which will later become the target video clips) meet the script requirements at the theme level.

[0024] The S300 uses historical data to calculate a deduplication-related score. This score includes a diversity score (Novelty) and a cold start score (ColdStart).

[0025] S400 may include: A weighted algorithm is used to calculate the overall score. The formula for the weighted algorithm is Formula 2-1.

[0026] FinalScore=w1×TagScore+w2×Novelty+w3×ColdStart Formula 2-1.

[0027] Where FinalScore is the overall score. TagScore is the tag matching score. Novelty is the diversity score. ColdStart is the cold start score. w1 is the weight of tag matching score. w2 is the weight of diversity score. w3 is the weight of cold start score. w1+w2+w3=1.

[0028] In this embodiment, the step of acquiring script fragments and multiple candidate video clips that semantically match them defines a clear input range and processing object for subsequent intelligent deduplication processing. Then, the tag matching degree is calculated based on the matching between candidate video clips and script fragments in a preset multi-level semantic tagging system. This step ensures that the selected materials meet the basic requirements of the script in terms of theme and content scope, laying the foundation for quality. Next, a deduplication-related score is calculated based on the usage history of the candidate video clips. This score specifically includes a diversity score to reduce the repeated use of the same video clip or the use of highly similar video clips, and a cold start score to encourage the use of new video clips. These two scores quantify the risk of material reuse and novelty from different dimensions, directly intervening in the content duplication problem. Finally, target video clips are selected from the candidate pool based on the weighted results of tag matching and deduplication-related scores. This step effectively integrates and balances content matching requirements with deduplication goals, enabling automatic and intelligent avoidance of duplicate content from the source during the material matching and selection stage. This fundamentally replaces the inefficient and subjective post-event manual review method, improving the automation level and content uniqueness of video production.

[0029] In one embodiment of this application, S300 includes calculating a deduplication-related score based on the usage history of the candidate video material segments, including: S311: Calculate the historical usage frequency of candidate video clips within a preset time period. The higher the historical usage frequency, the lower the diversity score.

[0030] S312, calculate the similarity between the candidate video clip and the selected set of clips.

[0031] S313, Calculate a diversity score based on the historical usage frequency and the similarity. The higher the historical usage frequency, the lower the diversity score. The higher the similarity, the lower the diversity score.

[0032] Specifically, this step defines in detail how the diversity score is calculated.

[0033] Optionally, the diversity score can be calculated using Formula 3.

[0034] Novelty=(1-use_count / max_use)×(1-similarity_penalty) Formula 3.

[0035] Here, Novelty is the diversity score. use_count is the number of times the candidate video clip has been used within the past preset number of days; use_count is obtained from the system logs. max_use is the maximum allowed number of uses, used for normalization. similarity_penalty is the similarity penalty factor. / is a division sign.

[0036] The similarity penalty factor can be set as follows: when the maximum semantic similarity between the candidate video clip and the selected material set S is greater than the threshold of 0.9, similarity_penalty = 0.5; otherwise, it is 0.

[0037] For example, if a candidate video clip featuring a lipstick swatch has been used 8 times in the past 7 days (with a maximum allowed usage of 10 times) and has a maximum similarity of 0.92 with a selected clip, its diversity score would be: (1-8 / 10)×(1-0.5)=0.1. This mechanism ensures that the system considers the repetition with selected clips when selecting the current clip, avoiding the appearance of multiple similar clips in the same video.

[0038] In a feasible embodiment, the definitions of the selected material set S and the maximum semantic similarity are further described.

[0039] The selected material set S is the set of target video clips that have been selected and temporarily stored by other script segments in the current processing session. For example, when matching materials for a complete script containing 5 script segments, while the system is selecting target video clips for the 3rd script segment, the selected material set S contains all the target video clips that have been selected for the 1st and 2nd script segments.

[0040] The similarity penalty factor is calculated as follows: The semantic similarity between each candidate video clip and all target video clips in the selected clip set S is calculated, and the maximum value is taken as the maximum semantic similarity between the candidate video clip and the selected clip set S. When this maximum semantic similarity is greater than 0.9, similarity_penalty = 0.5; otherwise, it is 0. Semantic similarity is obtained by calculating the cosine similarity of the feature vectors; the specific algorithm will be mentioned later.

[0041] For example, when selecting footage for the third segment of a beauty script, "This lipstick is long-lasting and doesn't fade," the selected footage set S includes footage already selected for the first two segments: an "arm swatch" segment and a "lip close-up" segment. The maximum semantic similarity between the current candidate video footage segment and these two selected segments is 0.92, exceeding the threshold of 0.9, therefore similarity_penalty = 0.5.

[0042] The core objective of deduplication is to avoid any highly repetitive segments. Maximizing semantic similarity allows for the detection of situations where a candidate video clip is highly similar to even just one target video clip in set S, thus incurring a penalty. If averaged, a candidate video clip extremely similar to a segment in S (similarity 0.95) might escape penalty because other segments in S are dissimilar, resulting in a lower average value (e.g., 0.4), which defeats the purpose of deduplication. Therefore, the maximum value strategy ensures that the semantic similarity between all selected target video clips corresponding to the script segment is below a threshold, thus more effectively guaranteeing the diversity of the generated video content.

[0043] For example, suppose we are selecting materials for the third script segment, and the selected set S contains the materials selected for the first two script segments: S = {segment A (arm swatch), segment B (ingredient explanation)}.

[0044] There is now a candidate video clip C (lip swatch).

[0045] Calculate SemanticSim=cosine_similarity(C,A)=0.92 (because these are all color tests).

[0046] Calculate SemanticSim=cosine_similarity(C,B)=0.15 (the content difference is large).

[0047] Set the maximum value to max_SemanticSim = 0.92.

[0048] Since 0.92 > 0.9, therefore similarity_penalty = 0.5.

[0049] Where SemanticSim represents semantic similarity, cosine_similarity is the expression for calculating semantic similarity, max_SemanticSimy is the maximum semantic similarity, and similarity_penalty is the similarity penalty factor.

[0050] In this embodiment, a diversity score is calculated by combining historical usage frequency and similarity with the selected set. This imposes a double penalty on high-frequency and high-similarity materials, effectively dispersing the distribution of material usage and avoiding excessive concentration of local hot materials, thereby improving the diversity of material usage at the micro level.

[0051] In one embodiment of this application, S300 includes calculating a deduplication-related score based on the usage history of the candidate video material segments, including: S320 calculates the cold start score based on the historical usage frequency of candidate video clips. The lower the historical usage frequency, the higher the cold start score. Candidate video clips that have never been used are assigned the highest possible cold start score.

[0052] Alternatively, an exponential decay function can be used to calculate the cold start score, as shown in Formula 4.

[0053] ColdStart=e^(-λ·use_count) Formula 4.

[0054] Where ColdStart is the cold start score. use_count is the number of times the candidate video clip has been used in the past preset number of days. λ is the decay factor, which can be set to 0.5. When use_count=0, ColdStart=1, and the score decreases exponentially as the number of uses increases. e^(-λ·use_count) is e raised to the power of (-λ·use_count).

[0055] For example, a new product demonstration video that has never been used before has a cold start score of 1; a video that has been used 3 times has a cold start score of e^(-1.5) ≈ 0.22 when λ=0.5.

[0056] In practical applications, the cold start weight in Formula 2-1 can be adjusted according to the system's operating mode. During the cold start exploration period, the weight w3 can be set to 0.2 to significantly increase the probability of selecting new materials; in the normal mode, the weight w3 is set to 0.05 to moderately influence the selection results while ensuring content quality.

[0057] In this embodiment, by calculating a cold start score based on historical usage frequency and giving the highest reward to new materials, the system is incentivized to prioritize the discovery and use of underutilized or entirely new video clips in the material library, thereby activating the utilization rate of the stock materials and preventing duplication from the perspective of resource allocation.

[0058] In one embodiment of this application, the weighted result also incorporates semantic similarity. S400 includes selecting one or more target video clips from the plurality of candidate video clips based on the weighted result of the tag matching degree and deduplication relevance score, including: S411 converts the text of the script fragment and the text of the candidate video material fragment into feature vectors respectively.

[0059] S412, calculate the similarity between two feature vectors, and use the similarity between the two feature vectors as the semantic similarity score.

[0060] Specifically, this embodiment adds a semantic similarity dimension to the weighted scoring system. First, the script fragment text and candidate video material fragment text are converted into feature vectors using a pre-trained language model, such as using the SBERT model to generate a 300-dimensional vector.

[0061] The formula for calculating semantic similarity is: SemanticSim=cosine_similarity(Vs,Vp)=(Vs·Vp) / (||Vs||·||Vp||) Formula 5.

[0062] Where SemanticSim represents semantic similarity, Vs is the feature vector transformed from the text content of the script segment, and Vp is the feature vector transformed from the text content of the candidate video clip. ||Vs|| is the magnitude of Vs. ||Vp|| is the magnitude of Vp.

[0063] For example, the semantic similarity between the script fragment "This foundation has strong coverage" and the candidate video description "Model demonstrates the foundation's coverage effect" may reach 0.85, while the similarity with "Explanation of foundation ingredients" may only be 0.3.

[0064] S400 may include: A weighted algorithm is used to calculate the overall score. The formula for the weighted algorithm is Formula 2-2.

[0065] FinalScore=w1×TagScore+w2×Novelty+w3×ColdStart+w4×SemanticSim Formula 2-2.

[0066] Where FinalScore is the overall score. TagScore is the tag matching score. Novelty is the diversity score. ColdStart is the cold start score. SemanticSim is the semantic similarity score. w1 is the weight of tag matching score. w2 is the weight of diversity score. w3 is the weight of cold start score. w4 is the weight of semantic similarity score. w1+w2+w3+w4=1.

[0067] The semantic similarity weight w4 is set to 0.25 in normal mode, and can be increased to 0.35 when there are insufficient exact matching segments, to ensure semantic accuracy while avoiding repetition. What exact matching segments are will be explained later.

[0068] In this embodiment, by introducing semantic similarity scoring based on text feature vectors, a deeper layer of semantic understanding is added beyond tag matching. This ensures that while avoiding duplication, the selected materials and script content still maintain a high degree of semantic fit, thus guaranteeing the accuracy of the generated video content.

[0069] In one embodiment of this application, the weighted result also incorporates a coherence score. S400 includes selecting one or more target video clips from the plurality of candidate video clips based on the weighted result of the tag matching degree and deduplication correlation score, including: S421, calculate the similarity of visual features between the candidate video clip and the adjacent selected clips.

[0070] S422, calculate the similarity of the candidate video clip with the adjacent selected clip in terms of semantic features.

[0071] S423, a coherence score is obtained based on the similarity in visual features and the similarity in semantic features.

[0072] Specifically, this embodiment adds a coherence scoring dimension to the scoring system. Similarity in visual features is semantic similarity, and similarity in semantic features is semantic similarity.

[0073] The formula for calculating the consistency score is shown in Formula 6.

[0074] Continuity=A×visual_sim+(1-A)×semantic_sim Formula 6.

[0075] Here, Continuity represents the coherence score. visual_sim represents visual similarity, calculated by transcribing the Bach distance between the keyframe color histograms of the candidate video clip and its adjacent selected clips, and then converting this distance into a similarity score. semantic_sim represents semantic similarity, the calculation method of which has been mentioned above. A is the weighting coefficient, which can be set to 0.6.

[0076] The adjacent selected segment is the last target video clip corresponding to the script segment that is immediately adjacent to the current script segment and has already been selected. "Already selected" means that the adjacent selected segment is already a finalized target video clip, which is the output result of the previous S100-S400 loop, and a candidate video clip. "Immediately adjacent" refers to the immediately preceding position in the script segment order.

[0077] The reason it's the previous one and not the next is that while processing the current script segment, the next script segment hasn't started processing yet, and there's no selected target video clip segment to refer to. Therefore, the coherence evaluation is unidirectional and forward-looking; that is, it evaluates whether the current candidate video clip segment, once it becomes a target video clip segment, is coherent with the previously determined script segment.

[0078] The final key point is that, because a previous script segment may have multiple corresponding target video clips, it is essential to select the last corresponding target video clip as the adjacent selected segment for effective continuity scoring. At this point, the definition of adjacent selected segments has been clearly explained.

[0079] For example: A beauty script contains three script segments in sequence: Script snippet A: "Welcome to my beauty channel" (It selects one target video clip: an opening shot featuring the host).

[0080] Script Segment B: "Today we're reviewing this new lipstick." (It selects two target video clips, which are in the following semantic order: Clip A is a close-up of a lipstick product, and Clip B is a close-up of applying lipstick to the back of a hand.)

[0081] Script segment C: "Its texture is very silky smooth" (The script segment being processed currently has multiple candidate target video clips waiting to be selected, such as: segment C is a demonstration of applying it to the back of the hand, and segment D is a display of the ingredient list).

[0082] When the system calculates the coherence of segment A (hand back smear demonstration): the adjacent selected segment is segment B.

[0083] The system calculates the visual_sim of fragment C and fragment B in terms of color and composition, as well as the semantic_sim of the text describing both, and then obtains the coherence score through Formula 6.

[0084] The formula for calculating visual feature similarity is Equation 7.

[0085] visual_sim=1-Bhattacharyya_distance(hist1, hist2) Formula 7.

[0086] Where visual_sim represents the similarity in visual features, hist1 is the color histogram of keyframes in the candidate video clip, hist2 is the color histogram of keyframes in adjacent selected clips, and Bhattacharyya_distance is the expression for calculating Bhattacharyya distance.

[0087] For example, continuing from the above example, if segment C and segment B are consistent in color tone and composition, their visual similarity can reach 0.8, and their semantic similarity is 0.7. Then the coherence score is 0.6×0.8+0.4×0.7=0.76.

[0088] In this embodiment, the similarity in semantic features is called semantic similarity, and the calculation method is shown in Formula 5.

[0089] S400 may include: A weighted algorithm is used to calculate the overall score. The formula for the weighted algorithm is Formula 2-3. FinalScore=w1×TagScore+w2×Novelty+w3×ColdStart+w5×Continuity Formula 2-3.

[0090] Wherein, FinalScore is the overall score. TagScore is the tag matching score. Novelty is the diversity score. ColdStart is the cold start score. Continuity is the consistency score. w1 is the weight of tag matching score. w2 is the weight of diversity score. w3 is the weight of cold start score. w5 is the weight of consistency score. w1+w2+w3+w5=1.

[0091] In this embodiment, by calculating the visual and semantic coherence scores between candidate video clips and adjacent selected clips, the selection of materials not only considers the suitability of individual clips, but also takes into account the smooth transition between clips, thereby improving the visual fluency and viewing continuity of the final generated video.

[0092] In one embodiment of this application, the weighted result further incorporates a precise matching bonus. S400 includes selecting one or more target video clips from the plurality of candidate video clips based on the weighted result of the tag matching degree and deduplication correlation score, including: S431: When a candidate video clip and a script clip are perfectly matched in the preset multi-level semantic tagging system, the candidate video clip is given a precise match bonus.

[0093] Specifically, this embodiment adds a precise matching bonus mechanism to the weighted scoring system.

[0094] Optionally, the formula for calculating the bonus points for exact matching is shown in Formula 8.

[0095] PrecisionBoost=1, when the candidate video clip and the script clip are completely matched in the preset multi-level semantic tag system; For PrecisionBoost=0, use formula 8 for all other cases.

[0096] PrecisionBoost provides bonus points for accurate matching.

[0097] Exact match requires that the candidate video clip and the script clip be identical across all specified levels of tags. For example, if the script clip has the tags "makeup-lipstick-swatches" and a total of three levels of tags, the candidate video clip must also have the exact same three levels of tags to receive an exact match bonus.

[0098] S400 may include: A weighted algorithm is used to calculate the overall score. The formula for the weighted algorithm is Equation 2-4. FinalScore = w1 × TagScore + w2 × Novelty + w3 × ColdStart + w6 × PrecisionBoost (Equation 2-4).

[0099] Wherein, FinalScore is the overall score. TagScore is the tag matching score. Novelty is the diversity score. ColdStart is the cold start score. PrecisionBoost is the exact match bonus. w1 is the weight of tag matching score. w2 is the weight of diversity score. w3 is the weight of cold start score. w6 is the weight of exact match bonus. w1+w2+w3+w6=1.

[0100] The weight w6 for exact match bonus can be adjusted according to the mode: it is 0.1 in normal mode, and can be increased to 0.15 when there are insufficient exact match clips. Through the exact match mechanism mentioned later, it is ensured that the exact match clips in the final selected target video clips reach a certain preset proportion, such as the number of exact match clips in the total clips ≥ max(3, total number of clips × 0.4).

[0101] In this embodiment, by assigning a precision matching bonus to materials that fully match on multi-level semantic tags, a quality assurance mechanism is established during the deduplication optimization process. This ensures that the accuracy of the core content expression is not excessively sacrificed in pursuit of diversity, thus maintaining the basic quality minimum of the generated video.

[0102] In one embodiment of this application, S400 may include: A weighted algorithm is used to calculate the comprehensive score, and the formula of the weighted algorithm is Formula 2-5. FinalScore = w1×TagScore + w2×Novelty + w3×ColdStart + w4×SemanticSim + w5×Continuity + w6×PrecisionBoost Formula 2-5.

[0103] Among them, FinalScore is the comprehensive score. TagScore is the tag matching degree. Novelty is the diversity score. ColdStart is the cold start score. semantic_sim is the semantic similarity. Continuity is the coherence score. PrecisionBoost is the precision matching bonus. w1 is the weight of the tag matching degree. w2 is the weight of the diversity score. w3 is the weight of the cold start score. w4 is the weight of the semantic similarity. w5 is the weight of the coherence score. w6 is the weight of the precision matching bonus. w1 + w2 + w3 + w4 + w5 + w6 = 1.

[0104] Specifically, in this embodiment, all six dimensions are incorporated into the weighted calculation, so that the deduplication effect of the final weighted structure is optimized.

[0105] In an embodiment of the present application, S400 further includes selecting one or more target video material segments from the multiple candidate video material segments according to the weighted result of the tag matching degree and the deduplication-related score, and further includes: S441, set a scoring threshold to filter out candidate video material segments with a weighted result lower than the scoring threshold.

[0106] S442, select one or more with the highest weighted result from the remaining candidate video material segments as the target video material segments.

[0107] Specifically, this embodiment defines a screening mechanism based on the scoring threshold. Let the scoring threshold be T_score, and all candidate video material segments with FinalScore < T_score can be excluded. The scoring threshold T_score can be dynamically adjusted according to the application scenario. For example, it is set to 0.7 when the content of the video material library is rich, and set to 0.5 when the materials in the video material library are relatively scarce.

[0108] The screening process is divided into two stages: first, exclude candidate video material segments with a weighted result lower than T_score, and then select the top U with the highest score from the remaining candidate video material segments as the target video material segments. U is a positive integer, where U is determined by the script segment requirements.

[0109] For example, if a script segment needs to match three target video clips, the system selects the three candidate video clips with the highest overall score (FinalScore) from a candidate pool filtered by a scoring threshold. These three candidate video clips become the three target video clips. Assuming T_score = 0.6, and eight candidate video clips pass the screening in the candidate pool, the three with the highest overall score (FinalScore) are selected as the final target video clips.

[0110] In this embodiment, a step-by-step screening mechanism is constructed by setting a scoring threshold to filter low-scoring candidates and selecting from high-scoring candidates. This first eliminates obviously unsuitable materials and then selects the best ones, thereby improving the overall quality and efficiency of target material selection.

[0111] In one embodiment of this application, the deduplication method for the video material further includes: S500: Traverse all target video clips and determine whether the ratio of the number of target video clips that completely match the script clips in the multi-level semantic tag system to the total number of target video clips is greater than or equal to a preset ratio.

[0112] S610, if the ratio of the number of target video clips that completely match the script clips in the multi-level semantic tag system to the total number of target video clips is less than a preset ratio, then select candidate video clips that completely match the script clips in the multi-level semantic tag system from the candidate video clips whose weighted results are lower than the scoring threshold, and replace the non-completely matching clips in all target video clips, until the ratio is greater than or equal to the preset ratio.

[0113] Specifically, this step defines the precise matching ratio guarantee mechanism. A precisely matched segment is a target video clip that completely matches the script segment in the multi-level semantic tagging system. Precise matching means a complete match with the script segment in the multi-level semantic tagging system. The concept of precise matching has been mentioned repeatedly above and will not be repeated hereafter.

[0114] Set a preset ratio P_min, which usually requires that the proportion of precisely matched segments be no less than 40% or the absolute number be no less than 3.

[0115] This embodiment can use a ratio check formula, such as Formula 9, to determine whether the ratio of the number of target video clips that completely match the script clips in the multi-level semantic tag system to the total number of target video clips is greater than or equal to a preset ratio.

[0116] P = N_precise / N_total ≥ P_min (Formula 9)

[0117] Where P is the ratio of the number of target video material segments that exactly match the script segment in the multi-level semantic tag system to the total number of target video material segments. N_precise is the number of target video material segments that exactly match the script segment in the multi-level semantic tag system, and N_total is the total number of target video material segments. P_min is a preset ratio. / is the division sign.

[0118] When P < P_min, the system searches for those material segments with lower FinalScore but exactly matching the script segment from the candidate video material segments filtered by the threshold, and replaces the non-exactly matching segments in the selected targets. The replacement strategy is based on the score, and the non-exactly matching segment with the lowest score is replaced first.

[0119] For example, after S400 is executed, 5 target video material segments are initially selected. Among the 5 target video material segments, 2 are precisely matching segments (the ratio is 40%), which does not meet the preset ratio requirement of 50%. The system finds 2 precisely matching segments from the candidate video material segments with weighted results lower than the scoring threshold, and replaces the 2 non-precisely matching segments with the lowest scores in the selected targets, so that the precise matching ratio reaches 80%.

[0120] In this embodiment, by checking and forcibly ensuring the ratio of precisely matching segments, and supplementing precisely matching materials from the filtered candidate pool for replacement when the ratio is insufficient, the selection result is dynamically adjusted, and the consistency between the generated video and the script intention in the core content is reliably maintained.

[0121] In an embodiment of the present application, the deduplication method of the video material further includes: S710, after each script segment obtains one or more corresponding target video material segments, retain multiple target video material segments with the top weighted results for each script segment as alternative video material segments.

[0122] S720, based on the alternative video material segments of each script segment, generate multiple video material sequence combinations that conform to the script semantic order.

[0123] S730, calculate the overall duplication index of each video material sequence combination.

[0124] S740, according to the overall duplication index, determine the final video material sequence as the output from multiple video sequence combinations.

[0125] Specifically, this step defines a permutation and combination optimization mechanism. Given a script segment sequence [s_1, s_2, ..., s_i], where each script segment s_i has n_j candidate segments, generate all possible combinations of sequences and include them in the set {P_1, P_2, ..., P_t}.

[0126] For each video material sequence combination P_u, the overall repetition index R_u is calculated according to Formula 1: R_u=(1 / (k-1))·Σ_{j=1}^{k-1}(1-cosine_similarity(V_{u, j}, V_{u, j+1})) Formula 10.

[0127] Where k is the total number of script segments. P_u is the u-th video sequence combination. V_{u,j} is the text feature vector of the j-th video segment in the video segment sequence P_u. R_u is the average semantic difference between adjacent segments in the video segment combination P_u; a larger R_u value indicates a lower degree of repetition in P_u.

[0128] `cosine_similarity(V_{u, j}, V_{u, j+1})` calculates the semantic cosine similarity between the j-th and (j+1)-th segments in sequence P_u, with a value range of [0, 1]. `(1-cosine_similarity(...))` converts the similarity to a dissimilarity value. The more similar the two segments, the closer this value is to 0; the less similar they are, the closer this value is to 1. The average dissimilarity of all adjacent segments in the sequence is then calculated to obtain the overall repetition index of the sequence. A higher R_u value indicates better overall content diversity and lower repetition in the sequence.

[0129] For example, a lipstick review script is divided into 3 segments (k=3): s_1: "Introducing the product" (two optional clips: the host holding a lipstick, and a close-up of the lipstick packaging); s_2: "Show texture" (two optional clips: application on the back of the hand, glossy details); s_3: "Makeup effect" (two optional clips: close-up of lips, overall makeup); The system will then generate 2×2×2=8 possible video sequence combinations.

[0130] For one of the combinations P = [host holding lipstick (s_1), gloss details (s_2), overall makeup (s_3)], the system will: Calculate the difference between the text vectors of the segments "host holding lipstick" and "gloss details". Calculate the difference between the text vectors of the segments "gloss details" and "overall makeup". Average these two difference values ​​to obtain R. Finally, the system selects the combination with the largest R value from the eight combinations as the final output, thus ensuring that the generated video is the most content-rich and diverse.

[0131] In this embodiment, by retaining multiple alternative materials for each script segment and optimizing the repetition of global sequence combinations, under the constraint of a fixed script order, the optimal combination with the lowest overall repetition is selected from all possible feasible solutions, thereby maximizing the deduplication effect at the macroscopic video sequence level.

[0132] In one embodiment of this application, the deduplication method for the video material further includes: S810: After generating a short video based on one or more target video material segments corresponding to each script segment, calculate the similarity between the generated short video and the historically generated short videos.

[0133] S820, when the similarity between the generated short video and the historically generated short video is greater than or equal to the preset similarity threshold, adjust the weight allocation of each scoring dimension and return the obtained script fragment and multiple candidate video material fragments that match its semantics.

[0134] Specifically, this step defines a feedback adjustment mechanism. Video similarity is calculated using a perceptual hashing algorithm: keyframes of the generated short video are extracted, their pHash values ​​are calculated, and the Hamming distance is compared with the pHash values ​​of historical short videos.

[0135] To determine whether a newly generated short video is a duplicate of a historical video, an efficient and reliable video similarity calculation method is needed. As a preferred implementation, this application employs a comparison method based on perceptual hash (pHash).

[0136] The S810 includes: S811, Extract Keyframes. Extract one or more keyframes that represent the video content from both the newly generated short video and the historically generated short videos.

[0137] S812, Calculate the pHash value. For each keyframe, calculate its perceptual hash value (pHash). The pHash algorithm extracts the frequency features of the image through Discrete Cosine Transform (DCT) and generates a fixed-length (e.g., 64-bit) hash fingerprint. This fingerprint is insensitive to non-content-related changes such as image scaling and brightness adjustments, but is sensitive to changes in the content itself.

[0138] S813, Calculate the Hamming Distance. Compare the pHash value of the new video keyframe with the pHash values ​​of historical video keyframes, bit by bit. The Hamming distance is the number of different characters at the same position in two strings of equal length. In this scenario, it calculates the number of different bits between the two hash values ​​(pHash values). The smaller the Hamming distance, the more similar the two video frames are; a distance of 0 indicates they are completely identical.

[0139] S814 determines the overall similarity. If multiple keyframes are compared, the average Hamming distance is used for comparison. When the (average) Hamming distance is less than a preset threshold T_h (for example, for a 64-bit hash, T_h=10 corresponds to a similarity greater than approximately 85%), the generated short video is determined to have a similarity to historically generated short videos that exceeds a preset threshold S_max.

[0140] S820 means that when the similarity is greater than or equal to S_max, the weight allocation is adjusted. For example, when the generated video has a high degree of repetition, the weight of diversity score and cold start score is increased; when the content accuracy is insufficient, the weight of tag matching and semantic similarity score is increased.

[0141] See Formula 11 for the weight adjustment formula.

[0142] w_i'=w_i+Δ_i Formula 11.

[0143] Where w_i' is the adjusted weight, and w_i is the original weight. w_i can be replaced with any one of w1, w2, w3, w4, w5, and w6.

[0144] Δ_i is the weight adjustment amount, set according to the actual effect. After adjustment, the material selection process is re-executed to generate a new video version.

[0145] In this embodiment, a closed-loop optimization system is formed by calculating the similarity between the generated video and the historical video and adjusting the weights based on this feedback. This allows the deduplication strategy to be adaptively adjusted according to the actual generation effect, improving the robustness and long-term effectiveness of the method in different scenarios.

[0146] The technical features of the above embodiments can be combined arbitrarily, and the execution order of the method steps is not restricted. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0147] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.

Claims

1. A method for deduplicating video footage, characterized in that, include: Retrieve script fragments and multiple candidate video clips that match their semantics; Based on the matching of the candidate video clips and script clips in the preset multi-level semantic tagging system, the tag matching degree is calculated; Based on the usage history of the candidate video clips, a deduplication-related score is calculated; the deduplication-related score includes at least a diversity score for reducing the repeated use of the same video clip or the use of highly similar video clips, and a cold start score for encouraging the use of new video clips. Based on the weighted result of the tag matching degree and deduplication correlation score, one or more target video clips are selected from the plurality of candidate video clips.

2. The method for deduplicating video materials according to claim 1, characterized in that, The calculation of the deduplication-related score based on the usage history of the candidate video clips includes: The historical usage frequency of candidate video clips within a preset time period is counted; the higher the historical usage frequency, the lower the diversity score. Calculate the similarity between candidate video clips and the selected set of clips; A diversity score is calculated based on the historical usage frequency and the similarity; the higher the historical usage frequency, the lower the diversity score; the higher the similarity, the lower the diversity score.

3. The method for deduplicating video materials according to claim 1, characterized in that, The calculation of the deduplication-related score based on the usage history of the candidate video clips includes: The cold start score is calculated based on the historical usage frequency of candidate video clips; the lower the historical usage frequency, the higher the cold start score; and the candidate video clips that have never been used are assigned the highest cold start score.

4. The method for deduplicating video materials according to claim 1, characterized in that, The weighted result also incorporates semantic similarity. The selection of one or more target video clips from the plurality of candidate video clips based on the weighted result of the tag matching degree and deduplication relevance score includes: Convert the text of the script fragment and the text of the candidate video clip into feature vectors respectively; Calculate the similarity between two feature vectors and use the similarity between the two feature vectors as the semantic similarity score.

5. The method for deduplicating video materials according to claim 1, characterized in that, The weighted result also incorporates a coherence score. The selection of one or more target video clips from the plurality of candidate video clips based on the weighted result of the tag matching degree and deduplication correlation score includes: Calculate the similarity of visual features between candidate video clips and adjacent selected clips; Calculate the similarity of candidate video clips with adjacent selected clips in terms of semantic features; A coherence score is obtained based on the similarity in visual features and the similarity in semantic features.

6. The method for deduplicating video materials according to claim 1, characterized in that, The weighted result also incorporates a precise matching bonus. The selection of one or more target video clips from the plurality of candidate video clips, based on the weighted result of the tag matching degree and deduplication correlation score, includes: When a candidate video clip and a script clip perfectly match in a pre-defined multi-level semantic tagging system, the candidate video clip is given a bonus for accurate matching.

7. The method for deduplicating video materials according to claim 1, characterized in that, The step of selecting one or more target video clips from the plurality of candidate video clips based on the weighted result of the tag matching degree and deduplication correlation score further includes: Set a scoring threshold to filter out candidate video clips whose weighted results are below the scoring threshold; Select one or more of the remaining candidate video clips with the highest weighted results as the target video clips.

8. The method for deduplicating video materials according to claim 7, characterized in that, Also includes: Iterate through all target video clips and determine whether the ratio of the number of target video clips that completely match the script clips in the multi-level semantic tag system to the total number of target video clips is greater than or equal to a preset ratio. If the ratio of the number of target video clips that completely match the script clips in the multi-level semantic tag system to the total number of target video clips is less than a preset ratio, then candidate video clips that completely match the script clips in the multi-level semantic tag system are selected from candidate video clips whose weighted results are lower than the scoring threshold, to replace all non-completely matching clips in all target video clips, until the ratio is greater than or equal to the preset ratio.

9. The method for deduplicating video materials according to claim 8, characterized in that, Also includes: After obtaining one or more target video clips corresponding to each script segment, multiple target video clips with the highest weighted results are retained as candidate video clips for each script segment. Based on the candidate video clips for each script segment, generate multiple video clip sequences that conform to the semantic order of the script; Calculate the overall repetition index for each combination of video clip sequences; Based on the overall repetition index, the final video material sequence is determined as the output from multiple video sequence combinations.

10. The method for deduplicating video materials according to claim 1, characterized in that, Also includes: After generating a short video based on one or more target video clips corresponding to each script segment, the similarity between the generated short video and the historically generated short videos is calculated. When the similarity between the generated short video and the historically generated short video is greater than or equal to the preset similarity threshold, the weight allocation of each scoring dimension is adjusted and the obtained script fragment and multiple candidate video material fragments that match its semantics are returned.

Citation Information

Patent Citations

  • Lable feature near-duplicated video detection method based on convolutional neural network semantic classification

    CN111723692A

  • Video material screening method and device, computer equipment and storage medium

    CN114417058A

  • Video generation method and device, storage medium and electronic equipment

    CN116095422A

  • Material analysis method and system

    CN116258522A

  • Video generation method and device, equipment and storage medium

    CN117979088A